From research question to study design

From Data to Bedside · the full write-up

When this applies

Use this write-up when the study you’re planning hasn’t been formally designed yet. The work covered here is the front-end of any research project: from “I have a clinical hypothesis” to “I have a written protocol that an IRB or a journal reviewer can evaluate.”

A careful pass through this write-up should leave you with:

  • A research-question statement precise enough to be statistically actionable
  • A PICO or PECO table specifying population, intervention or exposure, comparator, and outcome
  • For observational work, a target-trial emulation framing
  • A primary endpoint definition with explicit alternatives and the rationale for choosing one
  • A pre-registration draft (where applicable) covering primary analysis, secondary analyses, and pre-specified subgroups
  • A first-draft statistical analysis plan (SAP) outline

If any of these is missing from your current planning, this write-up applies. If they’re all already in place, you can probably skip to populations and sample size (population definition and sample size).

The decision framework

Moving from a clinical hypothesis to a defensible study design is a sequence of seven decisions. Each closes off a degree of freedom that you don’t want to be choosing post-hoc.

Step 1. State the research question precisely

The ladder makes the single-sentence PICO/PECO test the gate for a framed study. The design work begins where that test bites: a compound question that passes as one sentence but is really two. “Does treatment X reduce mortality and improve quality of life in elderly patients with condition Y?” reads as one question and is two. The tell is that the halves need different sample sizes, often different primary endpoints, and sometimes different study populations, so no single design serves both well. Pick one as primary and demote the other to a pre-specified secondary aim; carrying both as co-primary splits the alpha and tends to leave the study underpowered for each.

Step 2. Specify the comparator explicitly

The C in PICO carries methodological weight that’s easy to skip. Treatment X versus what? Versus placebo, versus standard of care, versus an active comparator, versus no treatment, versus a different timing or dose of the same treatment. Each comparator answers a different real-world question, and the answers can diverge.

In observational designs especially, the comparator is where a causal effect quietly becomes the wrong one, and it is the single place the ladder’s target-trial emulation pays off most. The operational rule for this step: fix the comparator by the target trial’s eligibility and treatment-assignment rules (Hernán and Robins 2016) rather than letting the data define “untreated” for you. Get that one choice wrong and no downstream method recovers the contrast you meant.

Step 3. Choose primary and secondary endpoints

A primary endpoint is the one your sample-size calculation is built on and your headline statistical claim is reported against. Secondary endpoints are pre-specified, reported in the methods, but typically not the basis of the headline conclusion. Three discipline points:

  1. Designate exactly one primary endpoint, and fix the secondary hierarchy. The ladder’s rule (choose before you see the data) is the floor; the design decision is which one is primary, because the primary is what the sample size is powered for and what the headline alpha is spent on. Co-primary endpoints split that alpha and force a multiplicity correction before the study has run a single subject.
  2. For composite endpoints, document the components and the counting rule. Does a patient count as having the event on the first component event, on any event, on a hierarchical composite (Finkelstein–Schoenfeld, win ratio)? Each rule answers a slightly different clinical question.
  3. For continuous outcomes, decide upfront how change will be analyzed. Change-from-baseline, post-treatment value adjusted for baseline (ANCOVA), or a transformation of either, each has different statistical power and different interpretation, and the right choice depends on the within-subject correlation structure.

Step 4. Pick the study design that matches the question

The match between question and design is the highest-stakes decision here: a well-executed analysis of a mis-matched design produces precise estimates of the wrong thing. Match the design to what the question needs and to what the data and ethics allow.

Design Reach for it when Why it fits
Parallel-group RCT randomization is feasible and ethical and the question is efficacy or effectiveness randomization handles confounding by design; highest internal validity
Cluster RCT the intervention acts on a group (clinic, ward, community), not an individual prevents contamination between arms, at the cost of a larger N (the design effect)
Crossover trial the condition is chronic and stable and the effect washes out between periods each patient is their own control, so far fewer patients are needed
Stepped-wedge the intervention will be rolled out to everyone in phases regardless ethical and logistical fit for implementation; later-treated units act as controls
Factorial trial two or more interventions are tested at once, most efficient when they don’t interact, though it can also be powered to test the interaction itself gets multiple trials’ worth of answers from one sample
Adaptive / platform trial treatments compete or recruitment is long, and the adaptation rules can be pre-specified drop, re-allocate, or add arms mid-trial for efficiency, but the rules must be registered up front
N-of-1 trial the question is what works for this patient and the condition is chronic and stable repeated within-patient crossovers give individualized causal evidence
Prospective cohort you can define exposed and unexposed and follow them forward, and the outcome isn’t too rare measures incidence and supports multiple outcomes from one exposure
Retrospective cohort the exposure and outcome already sit in records (claims, EHR, registry) same logic as prospective, fast and cheap, but limited to what was recorded
Case-control the outcome is rare or slow to develop samples on the outcome and looks back at exposure, far more efficient for rare events
Nested case-control / case-cohort you have a cohort but the key measurement is expensive buys case-control efficiency inside a defined, well-characterized cohort
Cross-sectional you want prevalence or an association at a single time point a snapshot; cannot on its own establish that exposure preceded outcome
Case report / series, ecological you’re describing something new or generating a hypothesis, not testing one cheap and fast, but descriptive only; ecological data also risk the ecological fallacy (a group-level association need not hold at the individual level)
Quasi-experimental (DiD, RDD, IV, synthetic control) randomization is impossible but a policy, threshold, or instrument supplies exogenous variation approximates the experiment you couldn’t run; the causal-inference toolkit covers the methods and their assumptions
Systematic review / meta-analysis the primary evidence already exists across studies synthesizes it; the unit of analysis is studies, not patients

Two cross-cutting rules survive whichever row you land on: any observational design still needs a target-trial framing to name its estimand rather than back into one, and a design is only as good as the eligibility, comparator, and endpoint decisions from Steps 1–3 that feed it.

Step 5. Specify the analysis plan before seeing the data

A pre-specified statistical analysis plan covers, at minimum: the primary outcome and its statistical test; the handling of missing data; the sensitivity analyses; the subgroup analyses; and the stopping rules where applicable. Documenting the SAP before unblinding the data is what distinguishes a confirmatory analysis from an exploratory one, and it’s the single discipline that most reduces reviewer methodology objections downstream.

Registration makes the SAP public, which the ladder treats as the commitment that keeps a confirmatory analysis confirmatory. The operational detail is which registry and when: for US clinical trials, ClinicalTrials.gov is generally required and tied to FDAAA results-reporting deadlines; for observational work, the Open Science Framework or AsPredicted carry the same evidentiary weight without the regulatory hooks. The time-stamp only protects you if it predates unblinding, so register before the data are in hand, not after the first look.

Step 6. Define the population precisely enough to compute a sample size

Sample size calculation requires four inputs: the expected effect size, the variability of the outcome in the target population, the statistical power you want, and the significance threshold (alpha). Without a precisely-defined study population, the effect-size and variability estimates are guesses. Populations and sample size covers the calculation in detail; Step 6 of this write-up is the precondition for it.

Step 7. Plan the sensitivity analyses before you need them

Sensitivity and robustness covers sensitivity-analysis design in detail. The Step-7 discipline here is to list, while designing the primary analysis, the methodological choices that are most likely to be challenged at review. Missing-data handling, alternative endpoint definitions, alternative population specifications, alternative comparator definitions, and alternative model specifications are the usual suspects. Each is a candidate for a pre-specified sensitivity analysis, written into the SAP now so it reads as designed-in rather than reactive when the paper reaches review.

Worked example

The Part D insulin DiD case study walks through how this seven-step decision framework was applied to a policy-evaluation question: did the Inflation Reduction Act’s $35 insulin cap actually move utilization, or just shift cost between payers? The case study shows the target-trial-emulation reasoning, the primary endpoint choice (log 30-day fills as a clean utilization measure that survives the IRA’s contemporaneous cost-sharing redesigns), the pre-specified sensitivity analyses, and the placebo design that earned its place in the protocol.

The Medicaid outliers case study is a different application of the same framework: research question (which providers bill in patterns that warrant a second look), comparator (every other provider billing the same code, with the explicit caveat that the peer group is too broad), primary outcome (paid-per-beneficiary on a robust z-score scale), and the explicit pre-specification of robustness checks (isolation-forest triangulation, BH-FDR multiplicity, peer-group sensitivity).

Further reading

  • Hernán MA, Robins JM. Using big data to emulate a target trial when a randomized trial is not available. American Journal of Epidemiology 183(8): 758–764. 2016.
  • Schulz KF, Altman DG, Moher D, for the CONSORT Group. CONSORT 2010 statement: Updated guidelines for reporting parallel group randomised trials. BMJ 340: c332. 2010.
  • von Elm E, Altman DG, Egger M, Pocock SJ, Gøtzsche PC, Vandenbroucke JP, STROBE Initiative. The Strengthening the Reporting of Observational Studies in Epidemiology (STROBE) statement: Guidelines for reporting observational studies. Lancet 370(9596): 1453–1457. 2007.
  • Pocock SJ, Stone GW. The primary outcome is positive — is that good enough? New England Journal of Medicine 375(10): 971–979. 2016.
  • Ioannidis JPA. Why most published research findings are false. PLoS Medicine 2(8): e124. 2005.

← Back to the pathway · the full write-up behind the Framing the study rung.

Learn the methods. to follow new write-ups and traces as they go up, alongside the full From Data to Bedside pathway.