The Pathway
From Data to Bedside. One pathway from the datapoint to the bedside, worked two ways: building each rung soundly, and tracing whether a recommendation holds.
Every clinical recommendation is the end of a pathway. It begins with a raw datapoint as it was first defined and measured, runs up through the model that turned it into an estimate, the effect size and its uncertainty, the synthesis that treated it as evidence, and the decision rule, and arrives at the sentence a clinician acts on. A clinician works at the bedside end of that pathway. A statistician usually works near the data end. Whether the recommendation can bear the weight a clinician puts on it depends on links most readers of either kind never see together.
This page lays that pathway out as a set of rungs, numbered stages 00 to 06 that run from the raw datapoint (rung 00) up to the recommendation a clinician acts on (rung 06), with two cross-cutting rungs for defending and conducting a study. Each rung holds the methods that stage rests on. Open a rung to see them, and open any item inside to read it.
There are two ways to walk the pathway. To design a study, start at rung 00 and build upward, making each stage sound before the next. To appraise a recommendation that already exists, which is what a trace does, start at the top with the guideline sentence and work back down to the measurement it rests on, asking at each rung whether it really holds. It is the same pathway either way, and walking it in one direction sharpens the other.
Each rung opens with a short plain-language intro before its methods. The blue links cross-reference related concepts across rungs, each item says when to reach for a method and how to choose among the alternatives, and the formulas sit at the end of each item for anyone who wants them, and every term is collected and defined in the glossary. No account is needed to read the pathway.
The pathway
The rungs below hold the methods, stripped of any one dataset.
Create a free account → to follow new traces and methods as they go up.
Framing pins down exactly what the study is asking and what would count as an answer: the population, the exposure or intervention, the comparator, and the outcome, written as one PICO or PECO question before any data is touched. Every later rung inherits these choices, so a vague or doubled-up question is the first place a study quietly goes wrong.
The question then chooses the design, the most consequential call on this page, since the wrong kind of evidence cannot be rescued at a later rung. The Choosing a study design node just below routes each kind of question, whether to synthesize existing evidence, run a randomized trial, reach for an observational design, describe a distribution, or evaluate a test or a cost, to the design and the rung that handle it.
Framing also fixes what would count as a cause before any method is chosen. Two models sit underneath every later rung: the counterfactual model, which defines an effect as a contrast of potential outcomes, and its complement the sufficient-component (causal-pies) model, which explains how causes combine and why an effect size can shift between populations with no change in mechanism. When the task is to appraise a recommendation rather than build one, the Bradford Hill viewpoints are the lens for weighing whether an observed association should be read as causal at all.
Research question (PICO / PECO)
A study is only as clear as the sentence it answers, so a sharp research question is the first deliverable, written at the very start before any design choice commits the study to one answerable sentence. The framework you reach for depends on the kind of question, and time is the element that moves around, written as its own T in some hands and folded into the population in others, much as descriptive epidemiology bundles person, place, and time.
- Intervention question for a trial uses PICO (population, intervention, comparator, outcome), which forces the sentence to specify who, given what, compared with what, and measured how.
- Add an explicit time horizon uses PICOT, which appends a timeframe and is the convention in clinical-question teaching.
- Add study-design eligibility, as in a systematic reviewDohoo et al. 2012, uses PICOS, which appends the study-design element.
- Exposure question for observational work uses PECO, the observational cousin that swaps the intervention for an exposure.
A question that cannot be written this way is not yet ready to design around, and bundling two questions into one is the most common way the specification quietly fails. Once sharp, it becomes the first line of the protocol every other document hangs from.
Descriptive epidemiology: person, place, time
Before you can explain a health event, you have to describe it, and that is the work of descriptive epidemiology. Reach for it whenever you have a health event to characterize before asking why, or when you need to fix which frequency measure to report. The discipline organizes a health event along three axes: who is affected (age, sex, race or ethnicity, comorbidity), where it occurs (geography, clinical setting, urban or rural), and when it occurs (a secular trend, seasonality, a birth-cohort effect, the shape of an epidemic curveDohoo et al. 2012). This person, place, and time triad does double duty: it pins down the population in the research question above, the who, where, and when of its P, and it is hypothesis-generating, since a pattern in one of the three axes is often what first suggests or sharpens an analytic question. The same description also fixes the frequency measure you report, whether a prevalence, an incidence, or a rate, each quantified with the measures of disease frequency in the Estimate rung. Read with statistical rather than epidemiologic eyes, this is the same exercise as characterizing the distribution of a single variable. The pitfall to keep in view is that a pattern in person, place, or time only generates hypotheses; on its own it can never establish a cause, so the description is a starting point for the analytic question, not a substitute for one.
Choosing a study design
The design is the most consequential call on this page, since the wrong kind of evidence cannot be rescued at a later rung. Route the question to its design before anything else.
- The evidence already exists across studies → do not run a new one; synthesize it at the Synthesis rung (systematic reviewDohoo et al. 2012, meta-analysisDohoo et al. 2012).
- You can assign the intervention ethically and feasibly → a randomized trial, where randomization handles confoundingDohoo et al. 2012 by design (parallel, crossover, factorial, cluster, or adaptive, following the question).
- Assignment is not yours to make → an observational design, chosen by what is known and how rare the outcome is (cohort, case-control, cross-sectional), ideally specified first as a target trial.
- The question is descriptive, how common or how distributed → a primary survey on a probability sample, or standardized rates from existing aggregate data.
- The question is about a test or an intervention's worth → a diagnostic-accuracy study or an economic evaluation at the Decision rung.
Each branch below expands into the designs it names. Appraising a study that already exists runs this in reverse: identify the design, then hold it to the standard its branch sets.
Target-trial emulation
When a causal question arises and randomizing is impossible, the cleanest discipline for reining in an observational design is target-trial emulationHernán & Robins 2016. You imagine the randomized trial you would have run, write down its protocol (eligibility, treatment assignment, start of follow-up, outcome), and then build the observational analysis to match it element by element. When you are orienting to the overall idea, the target trial is the protocol you specify in full before touching the data, because making every implicit choice explicit is what surfaces the biases an informal observational design tends to hide. The element that most often goes wrong is the alignment of follow-up: when a misaligned start of follow-up creates bias, the culprit is usually immortal time, a stretch of follow-up during which the outcome could not yet have occurred, alongside prevalent-user selection. Skipping the explicit protocol is precisely what lets immortal-time and prevalent-user biasDohoo et al. 2012 slip in unnoticed, so the protocol is not paperwork but the safeguard itself. Once the protocol is specified, the next choice is which observational study design best emulates it.
Observational study designs
An observational study design is chosen to fit the question and to neutralize the bias that would otherwise dominate it. Reach for this step once the question is set, ideally framed as a target trial, then pick the design by what is known and how rare the outcome is. Each core design is worked in full below, on the same skeleton: study group, exposure, comparability, outcome, follow-up, analysis.
- Exposure is known and follow-up is possible → the cohort studyDohoo et al. 2012, which follows a group forward to the outcome.
- The outcome is rare or slow to develop → the case-control studyDohoo et al. 2012, which starts from the outcome and looks back at exposure.
- You need a prevalence snapshot → the cross-sectional studyDohoo et al. 2012, measuring exposure and outcome at one time.
- You can sample a cohort more cleverly, or compare a person with themselves → the hybrid and self-controlled designs (nested case-control, case-cohort, case-crossover, SCCS, case-time-control).
Two variants cut across these: the active-comparator new-user design restricts to initiators of one treatment versus an active alternative, curbing confoundingDohoo et al. 2012 by indication, and a disease registry supplies a standing source populationDohoo et al. 2012 many of these designs can be drawn from. Report any observational design to STROBEDohoo et al. 2012 (Strengthening the Reporting of Observational Studies in Epidemiology).
Cohort study
The cohort studyDohoo et al. 2012 takes a group defined by exposure and follows it forward to the outcome, prospectively when assembled before the outcomes occur or retrospectively when reconstructed from records. Reach for it when exposure is known and follow-up is possible, and when you want to estimate incidence directly. Its steps are the template every design here follows.
- Study group: assemble subjects free of the outcome at baseline and classify them by exposure; an inception cohort that enrolls everyone before the outcome can occur is what keeps survivorship out.
- Exposure: fix it at or before baseline, and prefer new (incident) users to prevalent users, since prevalent users are survivors of early events, the prevalent-user selection that target-trial emulationHernán & Robins 2016 is built to prevent.
- Comparability: exposed and unexposed differ for reasons other than exposure, which is confoundingDohoo et al. 2012; handle it by design (restriction, matching) or, more often, by the confounding tools on the Model rung.
- Follow-up: accrue person-timeDohoo et al. 2012 over a defined period; differential loss to follow-up is attrition bias, and a misaligned start of follow-up is immortal time.
- Outcome: ascertain it identically in both groups, since unequal ascertainment is detection bias.
- Analysis: a risk ratio or risk differenceDohoo et al. 2012 in a closed cohort, a rate ratio on person-time, or Cox or Poisson regression for time-to-event, always adjusted for the measured confounders.
The failures to watch are immortal time, loss to follow-up, and confounding by indication. Build the analytic cohort at the Measurement rung and report to STROBEDohoo et al. 2012.
Case-control study
The case-control studyDohoo et al. 2012 starts from the outcome and looks back at exposure, which makes it the efficient choice when the outcome is rare or slow to develop. Its validity rests almost entirely on one idea: cases and controls must come from the same source populationDohoo et al. 2012.
- The study baseDohoo et al. 2012 is the population and time that gives rise to the cases; naming it is the first move, because everything else is judged against it.
- Nested in a cohort, when that base is fully enumerable, the known sampling fractionsDohoo et al. 2012 let the design even recover disease frequency by exposure, which an ordinary case-control study cannot.
- The case seriesDohoo et al. 2012 is the incident cases arising from that base under a clear, consistently applied case definitionDohoo et al. 2012.
- Control selectionDohoo et al. 2012 samples from the same base independently of exposure, the principle that governs the whole design. How you sample controls decides what the odds ratio actually estimates:
- Sample controls through follow-up (density or risk-set sampling) and the odds ratio estimates a rate ratio, with no rare-disease assumptionDohoo et al. 2012 needed.
- Sample controls from those still disease-free at the start and it estimates a risk ratio.
- Sample survivors at the end and it estimates a risk ratio only when the disease is rare.
- Comparability: matching on strong confounders is common, but it forces a matched (conditional) analysis and can backfire as overmatchingDohoo et al. 2012 when the matching factor sits on the causal path.
- Exposure is measured retrospectively, so recall bias and reliance on memory or records are the standing threats.
- Analysis: an odds ratio from logistic regression, or conditional logistic regressionDohoo et al. 2012 when matched.
The characteristic failures are selection biasDohoo et al. 2012 (controls not drawn from the base), recall bias, and Berkson's bias when controls are hospital patients. Report to STROBEDohoo et al. 2012.
Cross-sectional study
The cross-sectional studyDohoo et al. 2012 takes a snapshot of a population at a single time, measuring exposure and outcome together. Reach for it to estimate prevalence or to describe how something is distributed, typically on a probability survey sample.
- Study group: a probability sample of the target population at one point in time, analyzed with its survey weights and design.
- Measurement: exposure and outcome are recorded simultaneously, which is what makes it fast and cheap.
- Analysis: a prevalence, a prevalence ratioDohoo et al. 2012 or prevalence odds ratio, and, under assumptions, an estimate of incidence.
The defining limitation is temporal ambiguity: because exposure and outcome are measured at once, you usually cannot tell which came first, so causal reading is weak and reverse causationDohoo et al. 2012 is live. Prevalent cases also over-represent long-lasting disease, the prevalence-incidence (Neyman) bias. Report to STROBEDohoo et al. 2012.
Hybrid and self-controlled designs
Between the cohort and the case-control sit hybrid and self-controlled designs, which either sample a cohort more cleverly or drop the between-person comparison altogether. Choose by what threatens your question.
- The nested case-control design samples controls from within a cohort at each case's event time (risk-set sampling), keeping the cohort's validity at a fraction of the measurement cost.
- The case-cohort design uses one random subcohort as the comparison for several outcomes at once.
- The case-crossover designDohoo et al. 2012 compares a person's exposure just before an acute event with their own earlier reference windows, so every stable characteristic is controlled by design; it suits transient triggers of abrupt events.
- The self-controlled case series (SCCS)Farrington 1995 compares event rates across exposed and unexposed person-timeDohoo et al. 2012 within each case, needing only the cases.
- The case-time-control design adds a control group to the case-crossover design to remove exposure time trends.
- The case-case design replaces controls with a second disease (a related pathogen, or the drug-susceptible strain), for surveillance data where ordinary controls are hard to define, so its odds ratio contrasts subtypes rather than estimating risk; the case-case-control design instead keeps a shared control series and adds a second case series, comparing each against the common controls to separate subtype-specific (say, resistance- vs susceptibility-specific) risk factors.
- The case-only design uses cases alone to estimate an exposure-by-covariate interaction, not the main effect, assuming the two are independent in the source populationDohoo et al. 2012.
- Two-stage samplingDohoo et al. 2012 subsamples for an expensive covariate and reweights.
The within-person designs assume the event does not itself change future exposure (and, for the SCCS, is not fatal) and that exposure carries no strong time trend; when those hold they are powerful, since they erase all fixed confoundingDohoo et al. 2012: each person is compared only with themselves at other times, so anything that stays constant within a person (genetics, stable habits, sex) cancels out and cannot confound. They connect to the g-methods on the Model rung.
Randomized-trial designs
A randomized trial assigns the intervention by chance, which is what lets it claim causation directly, and the variant follows the question. This overview picks the shape; the mechanics that follow, allocation, monitoring, and analysis, are their own nodes.
- The parallel-group trial compares arms in separate groups, the default for most comparisons.
- The crossover trial gives each subject both treatments in sequence with a washout, so each serves as their own control; it suits stable, chronic conditions where the effect reverses.
- The factorial trial tests two or more interventions at once in the same subjects, efficient when they do not interact.
- The cluster-randomized trial randomizes groups (clinics, villages) rather than individuals, needed when the intervention is delivered at the group level, and it pays for it with a design effectDohoo et al. 2012.
- Adaptive and platform trials add or drop arms under pre-specified rules, spending alpha carefully as they go.
Whichever the shape, the trial still needs eligible participants defined by clear criteria, a fully specified intervention and comparator, and a pre-specified primary endpoint. Assignment is protected by randomization and blinding, the size is fixed at sample size, early stopping runs through interim analyses, and the result is read on the intention-to-treat population. Report to CONSORT.
Adherence and follow-up
A trial's answer is only as clean as its follow-up, and adherence and follow-up is where a randomized comparison quietly degrades. Patients stop the drug, cross over, or drop out, and how you handle that is a design decision, not an afterthought.
- A run-in period before randomization can screen out non-adherers, which raises adherence but narrows generalizability.
- Adherence measurement (pill counts, pharmacy refills, drug levels) captures the gap between treatment assigned and treatment actually taken.
- Loss to follow-up is the main threat: differential dropout breaks the balance randomization bought, so minimizing and tracking it matters more than any analytic fix.
- Missing outcomes are handled under a stated assumption, with the mixed model for repeated measures the usual choice for a longitudinal endpoint under missing-at-random.
These choices are exactly what force the intention-to-treat versus per-protocol distinction: intention-to-treat keeps everyone as randomized and preserves the comparisonICH E9 1999, while per-protocol answers a different, non-randomized question. Report both when they might diverge.
Randomization and blinding
Randomization and blinding are what let a trial claim causation, and they matter most when assignment and assessment must be shielded from bias. How randomization is run depends on the balance you need.
- Hide the upcoming assignment so it cannot be foreseen and gamed through allocation concealment, the safeguard against foreknowledgeSchulz & Grimes 2002.
- Balance arms in small chunks as enrollment proceeds through block randomization, also called permuted-block randomization.
- Balance within prognostic strata through stratified randomization when a few strong prognostic factors must be balanced.
- Dynamically balance many factors through minimization, which adaptively assigns each patient to keep the arms balanced across all factors at oncePocock & Simon 1975.
- Mask treatment after allocation through blinding of patients, clinicians, and outcome assessors against the conscious and unconscious bias that knowing the arm introduces, with a placebo as its vehicle.
The whole point is fragile in a specific way: a trial that randomizes but then fails allocation concealment or blinding lets back in the very bias randomization was meant to remove, so the two safeguards are not optional trimmings but the conditions under which randomization actually buys causation.
Report this trial to CONSORT (Consolidated Standards of Reporting Trials); its protocol counterpart is SPIRITSchulz et al. 2010.
Interim analyses and group-sequential design
A long trial can answer early, for good or ill, and the machinery of interim analyses and group-sequential design is what makes that possible when a trial may resolve early for efficacy, futility, or harm. The governing concern is that unplanned peeking at accumulating data inflates the type-I errorDohoo et al. 2012, the false-positive risk you have to budget and spend deliberately rather than leak; the alpha-spending here is simply multiplicity control across time.
- A group-sequential design delivers planned looks with stopping rules, pre-specifying the interim analyses and spending the alpha across them with a boundary.
- The O'Brien-Fleming boundary is a stringent early-look boundary, demanding early and near-nominal at the end.
- The Pocock boundary holds a constant nominal level at every look.
- An independent data safety monitoring board, not the sponsor, reviews the accruing data and can halt the trial for overwhelming efficacy, for futility once significance is unreachable, or for harm.
The tradeoff worth stating plainly is that stopping early for benefit tends to overestimate the effect, so an early stop buys an answer at the cost of a likely inflated one.
Non-inferiority and equivalence
Not every trial aims to show that a new treatment is better; the family of non-inferiority and equivalence trials is what you reach for when the goal is to show a new option is not meaningfully worse while being cheaper, safer, or easier.
- The non-inferiority trial tests that the new option is not meaningfully worse, against a shifted null: instead of testing against "no difference," it tests against a margin set a fixed distance away, so the trial passes if the new treatment is no worse than the standard by more than that pre-specified amount, judged by whether the confidence interval for the difference stays on the acceptable side.
- The equivalence trial tests that new is neither worse nor better, bracketing the difference symmetrically on both sides.
- The non-inferiority margin is how much worse is tolerablePiaggio et al. 2012; it carries the whole argument and is set before the trial from what is clinically tolerable and from the active control's own established advantage over placebo.
- Assay sensitivity is the load-bearing assumption that the trial could have detected a real difference had one existed. The reason it matters is causal: anything that blurs the two arms toward looking alike (poor adherence, mismeasured outcomes, an underdosed comparator) drives the estimate toward the margin and so toward passing, so sloppiness is rewarded rather than punished.
That is the pitfall the margin guards against: it must be justified in advance, since without assay sensitivity a poorly run trial passes by default rather than on merit.
Dose-finding and early-phase designs
Before a confirmatory trial, dose-finding and early-phase designs find the dose and the go/no-go efficacy signal, with phase I asking what dose is tolerable and phase II screening whether the efficacy signal justifies a phase III.
- The 3+3 design is the classic dose-escalation rule, stepping up in small fixed cohortsDohoo et al. 2012 until toxicity appears.
- The continual reassessment methodO'Quigley et al. 1990 is model-based dose escalation, which estimates the dose more efficiently and with fewer patients overdosed.
- The maximum tolerated dose is the highest acceptably safe dose those escalations climb toward.
- Simon's two-stage designSimon 1989 is a common phase II efficacy screen, a small single-arm scheme that stops early when the first-stage responses are too few to be worth continuing, the same early-stopping logic the confirmatory interim analyses formalize later.
The tradeoff to keep in view across all of these is deliberate: they exchange the rigor of a large randomized comparison for speed and a close safety watch at the stage where the question is still go or no-go.
Endpoint logic and pre-registration
The work of endpoint logic and pre-registration belongs at the moment you define what the study will be judged on, and it applies to any design: a trial, a cohort, or a survey all need a pre-specified primary outcome and a plan committed before the outcome data are seen. Only the venue and the vocabulary shift with the design.
- The primary endpoint is the main pre-specified outcome, the one the sample size is built on and the headline claim is read against; everything else is secondary and labelled as such, and a composite endpointDohoo et al. 2012 has to be read component by component, since a significant composite can be driven entirely by its most frequent and least important part.
- A surrogate endpoint is a stand-in for the outcome, a lab marker or scan standing in for a clinical outcome, which buys speed but earns trust only once it is validated to capture the treatment's effect on what patients feel.
- Prentice's criteria are the formal test of whether a surrogate qualifies, which many surrogates fail.
- Pre-registration (ClinicalTrials.gov for trials, PROSPERO for systematic reviewsDohoo et al. 2012, the Open Science Framework for observational and other work) locks the analysis plan in advance, the public commitment that keeps a confirmatory analysis confirmatory, with the statistical analysis plan spelling it out before the outcome is seen, and for a regulated trial the same registration also satisfies the requirement to post results.
The pitfall the whole exercise guards against is the classic one: choosing or switching the endpoint after seeing the data is the surest route to a result that will not replicate.
Sample size and power
Before enrolling, fix the sample size, and the calculation forks on the goal: estimating a quantity to a target precision, or detecting a difference with adequate statistical powerDohoo et al. 2012.
One orientation before the algebra: the levers never change. N grows with the outcome's variability and with any demand for a smaller significance level or more power, and shrinks as the target effect grows.
To estimate to a target precision (a confidence interval of set width), match the calculation to what you are estimating:
- A single mean: the size is the squared z-value times the variance over the squared half-width.
- A single proportion: the same logic holds with the proportion's variance, and with no prior estimate use 0.5, which maximizes the variance and gives the most conservative N.
- A difference between two groups: size the interval around that difference by summing the two groups' variances.
- \(n\) is the required sample size and \(z\) (written \(z_{1-\alpha/2}\) for a two-sided interval) is the standard-normal value set by the confidence level, namely its \((1-\alpha/2)\) percentile, so \(z = 1.96\) at 95% confidence
- \(\sigma^2\) is the a priori variance and \(p\) the a priori proportion (whose variance is \(p(1-p)\), with \(q = 1-p\)); \(d\) is the half-width of the interval, the margin of error (also called the allowable error), so the interval has full width \(2d\). Each input is a guess made before the study, the mild paradox of sizing: you need a rough value of \(p\) or \(\sigma\) to plan for measuring it
To detect a difference with power, name the effect you want to detect. The two-group comparison of means is the template the others follow, at \(n\) per group:
- \(z_{1-\alpha/2}\) is the \((1-\alpha/2)\) percentile of the standard normal fixed by the significance level and \(z_{1-\beta}\) the \((1-\beta)\) percentile fixed by the power (1.96 and 0.84 for a two-sided \(\alpha = 0.05\) and 80% power)
- \(\sigma\) is the outcome standard deviation and \(\Delta\) the difference worth detecting; a one-sample test against a known value drops the factor of 2
Every other design is the same levers rearranged: a proportion's \(p(1-p)\) in place of \(\sigma^2\), the number of events rather than the head count for a time-to-event outcome, an events-per-variable rule (roughly 10 to 20) for a regression coefficient, and design-effect inflation \(1+(m-1)\rho\) for cluster randomization, with a simulated mock of the planned analysis (Monte Carlo) when no closed form fits. Read backwards, the same formulas size the effect a fixed sample can catch: when the data already exist (a registry extract, a public-use file, a natural experiment), solve instead for the power or the minimum detectable effect the sample supports and write that into the methods, rather than choosing N.
Either direction rests on two inputs, both guesses made before the study. The effect comes, in falling order of credibility, from prior literature (meta-analyses and close comparators), from pilot data (which tends to overstate it through regression to the mean), or from the smallest clinically meaningful difference when the literature is thin; document the source beside the number, a d of 0.3 from a named study rather than a bare 0.3. The variability (a standard deviation, a control-group event rate, a baseline hazard) comes from the same sources and is usually the more uncertain of the two, which is why underestimating it is the standard route to an underpowered study. Because N moves so sharply with both, report it under an optimistic, a best-estimate, and a conservative scenario, and put the conservative one in the protocol. For the arithmetic, pwr covers the textbook designs, WebPower and powerSurvEpi a wider catalog including survival, and simr or clusterPower the simulated cases.
A result is only as good as the act of measuring it. This rung is about where the raw data came from, how each variable was defined, and how much to trust it, before any model is run.
Data sources and their tradeoffs
Where the data came from bounds every question it can answer, so the first move with any dataset is identifying which of the data sources and their tradeoffs you are holding, since that fixes what it can support; each carries a characteristic strength and a characteristic bias.
- Survey data such as the National Health and Nutrition Examination Survey (NHANES) and the Behavioral Risk Factor Surveillance System (BRFSS) are sampled-population questionnaires, a probability sample built for population estimates, generalizing well once its weights and design are respected, though it is cross-sectional and self-reported in places.
- Electronic health record data is clinical detail from care records, rich with labs, vitals, and notes, but recorded for care rather than research, so it is messy, confined to one health system, missing in informative ways, and bound by privacy duties.
- Claims data are billing records across encounters, covering prescriptions and encounters broadly across a payer's population, but a code is a bill rather than a diagnosis and clinical detail is thin.
- Publicly available aggregate data such as the American Community Survey, CDC WONDER, and vital statistics supply population denominators and area-level rates, giving context and standardization but supporting only ecological analysis, which invites the ecological fallacy when a group-level association is read as an individual oneMorgenstern 1982.
- Registries are enrolled cohorts for a condition, sitting in between: purpose-built, deep but narrow.
The pitfall to hold onto is that ecological fallacy, since each source's bias is baked in.
Data feasibility, enrollment, and linkage
Choosing a data source is only the start; data feasibility, enrollment, and linkage is the step that comes after, confirming the database can actually answer your question, that patients are observable long enough to see both exposure and outcome, and that any joined datasets link without exposing identities. This is where the strengths and gaps from data sources and their tradeoffs become concrete.
- You need observable follow-up, so continuous enrollment requirements ensure a patient's claims are genuinely captured during baseline and follow-up, so that absence of a code means absence of care rather than a coverage gap.
- You must size the population, so database feasibility work counts how many patients survive each eligibility criterion, and that attrition funnel feeds directly into assembling the analytic cohort.
- You join multiple datasets, so privacy-preserving record linkage uses tokenized identifiers to match records across sourcesFellegi & Sunter 1969 while honoring data privacy and security obligations.
The reason none of this can be skipped is the pitfall it prevents: without feasibility checks an analysis can look complete yet rest on patients you could never have observed, because a coverage gap quietly masquerades as absence of care.
Data standards and provenance
Before any analysis, a datapoint arrives pre-shaped by the system that recorded it, and that system used a standard vocabulary whose reach you have to know; this is the heart of data standards and provenance. The overall idea is to ask which layer governs the data in hand: trial data follows the Clinical Data Interchange Standards Consortium (CDISC) regulatory models, while claims and records use clinical coding ontologies.
- CDISC SDTM, the Study Data Tabulation Model, governs collected trial data following the regulatory model, mirroring what was collected.
- ADaM is the analysis-ready dataset standard derived from it.
- SNOMED supplies clinical coding terminology for findings, alongside ontologies like ICD, HCPCS, and RxNorm.
Knowing what a code does and does not capture is half of real-world data (RWD) competence, and the pitfall to keep in front of you is that a billing code is not a diagnosis. A rule-out code makes the point: a claim carrying an ICD code for "chest pain, rule out myocardial infarction" records why a test was ordered, not that the patient had an infarction, yet a naive query counts it as a case. Provenance stays incomplete until you can say, for each field, which vocabulary it speaks and what that vocabulary was built to record rather than what was clinically true.
Claims and coding standards
Whenever you analyze claims fields, each is a coded value drawn from a specific vocabulary, and a field cannot be interpreted without knowing which vocabulary it speaks and what that vocabulary was built to capture; this is the concrete substrate of claims and coding standards beneath data standards and provenance. Resolve, field by field, which vocabulary encodes it.
- ICD-10-CM diagnosis codes capture diagnoses, the clinical modification used for morbidity coding.
- ICD-10-PCS procedure codes capture inpatient procedures.
- CPT/HCPCS codes capture professional services.
- The NDC (National Drug Code) identifies dispensed drugs, encoding manufacturer, product, and package rather than ingredient, so ingredient-level analysis requires mapping.
- LOINC codes capture labs and observations.
- The NPI identifies the provider.
- OMOP standardized vocabularies (OHDSI) enable cross-database mappingHripcsak et al. 2015, putting heterogeneous source codes onto standard concepts so a study runs unchanged across databases.
- Code crosswalks and mappings translate between systems more generally.
- The ATC and defined daily dose (DDD) classification measures drug utilization, grouping agents by anatomical and therapeutic class and expressing consumption in a comparable unit across products and countries.
The pitfall is that these vocabularies were built for billing and order entry, not research, and every crosswalk is lossy, so a code reflects what was payable or documented rather than what was necessarily clinically true; knowing which mapping you relied on, and what it dropped, is part of the provenance you owe the reader.
Survey research (primary data collection)
Survey research collects fresh individual-level data from a probability sample rather than reusing an existing dataset. Reach for it when the question is descriptive, how common something is, how it is distributed, what people report, and when you need a sample you can generalize from. The steps below build the survey in order; this node is how they fit together.
- Sampling: draw a probability sample so every unit has a known, non-zero chance of selection, choosing the scheme by cost and precision (sampling design).
- Instrument: write the questionnaire that measures exposure and outcome, since in a survey the instrument is the measurement (questionnaire design).
- Quality: pre-test the instrument, validate it, and push the response rateDohoo et al. 2012, because nonresponse is the survey's version of selection biasDohoo et al. 2012 (pre-testing and validationDohoo et al. 2012).
- Analysis: carry the sampling weightsDohoo et al. 2012, strata, and clusters into every estimate, or the standard errors are wrong (complex-sample analysis).
The design usually reads as a cross-sectional studyDohoo et al. 2012 when exposure and outcome are captured at once; repeated over time it becomes a panel or repeated cross-section. Treat a low response rate as the first threat to generalizability.
Survey sampling design
A survey's credibility comes from how the sample was drawn, so survey sampling designDohoo et al. 2012 is the step where you choose the scheme that lets a probability sample generalize to the population. Each choice leaves a trace the analysis must carry, so designing the sample and analyzing it are the same problem seen from two ends.
- A probability sample gives every unit a known nonzero chance of selection, the basis for generalization.
- Simple random samplingDohoo et al. 2012 is an equal-chance draw from one frame, every unit with equal probability.
- Stratified samplingDohoo et al. 2012 splits the frame and samples within each stratum, letting you oversample a small subgroup for precise estimation at the cost of unequal selection probabilities.
- Cluster samplingDohoo et al. 2012 draws whole groups, schools or blocks or clinics, when no list of individuals exists and field cost matters.
- Multistage samplingDohoo et al. 2012 samples in successive nested stages, drawing primary sampling units and then units within them, often with probability proportional to sizeDohoo et al. 2012.
- The design effectDohoo et al. 2012 is the variance inflation from clustering, by which the target sample size is scaled up to hold the effective sample sizeKish 1965 \(n_{\text{eff}}\) on target; unequal selection becomes the survey weight, and the strata and clusters become the design's strata and primary sampling units.
Questionnaire and instrument design
What a survey can measure is fixed before fieldwork by the instrument, so questionnaire and instrument designDohoo et al. 2012 is the craft you apply before a single response is collected, since item wording, response format, mode of administration, and skip logic all set what can later be analyzed. Each item's wording carries assumptions, and the flaws to watch for have names.
- A double-barreled question asks two things at once, so it cannot be answered cleanly.
- A leading questionDohoo et al. 2012 is wording that steers the answer, while an unbalanced response scale skews the result.
The response format decides what analysis is even possible later: categorical options must be mutually exclusive and exhaustiveDohoo et al. 2012, while a Likert scaleDohoo et al. 2012 is ordinal, so treating its scores as interval data (means and SDs) is only defensible with at least five points and roughly equal spacing, and otherwise ordinal or nonparametric methods fit better; individual items are often added into a summated scaleDohoo et al. 2012 that is then read as interval. The mode of administration, in person or phone or web or self-report, shifts both who responds and how candidly, since a face-to-face interviewer invites more social desirability shading on sensitive items than an anonymous form. Branching is designed here too: a gate question with explicit skip logic spares respondents irrelevant items, which later surfaces in the data as skip patterns, so a by-design blank is an instrument choice and not a data accident, a distinction that must be honored when the data are read. A partial questionnaire designDohoo et al. 2012 goes further, giving disjoint subsets of the secondary questions to random subgroups so the gaps are missing completely at random by design. Confirming the instrument measures what it claims is the reliability and validityDohoo et al. 2012 work that belongs to this same pre-fieldwork stage.
Pre-testing, validation, and response rate
A questionnaire is a measurement instrument, and like any instrument it has to be checked before and after fielding. Pre-testingDohoo et al. 2012, validation, and response rateDohoo et al. 2012 are the three quality steps that decide whether the answers mean anything.
- Pre-testing runs the draft on a small sample, often with cognitive interviewing, to catch items respondents read differently than intended before they contaminate the real data.
- Validation asks whether the instrument measures the construct it claims: content and construct validityDohoo et al. 2012 for meaning, criterion validityDohoo et al. 2012 against a reference, and reliability for consistency on repeat administration.
- The response rate is the share of the sampled who answer, and it is the survey's version of selection biasDohoo et al. 2012: a low rate threatens generalizability whenever nonresponders differ from responders, and the first lever on it is low response burdenDohoo et al. 2012, since a shorter questionnaire is completed more often.
- Nonresponse handling uses weighting adjustments or follow-up of a nonresponder subsample to gauge and correct the bias, not merely to lift the raw rate.
The pitfall is treating a high response rate as sufficient: a representative 50% can beat a skewed 80%, so what matters is whether nonresponse is related to the answers. These checks feed the weighted analysis that follows.
Complex-sample design and survey weighting
Surveys like the National Health and Nutrition Examination Survey (NHANES) are not simple random samples (SRS): they oversample some groups and cluster others by design, which is exactly when complex-sample design and survey weightingDohoo et al. 2012 applies. Design-aware analysis, using survey weights, strata, and primary sampling units, is what lets such a sample speak for the population it was drawn to represent.
- Scale respondents to the population with the survey weight, the device that undoes the unequal selection.
- Precision lost to the design is tracked by the effective sample sizeKish 1965, equal to \(n / \text{DEFF}\) where the design effectDohoo et al. 2012 \(\text{DEFF}\) is the penalty for clustering, so a design effect of 2 leaves the precision of only half the respondents.
In SAS the design is declared with the STRATA, CLUSTER, and WEIGHT statements of PROC SURVEYMEANS and its siblings, and a subpopulation is analyzed through a DOMAIN statement rather than by deleting rows, which would bias the variance; the same survey instruments also carry skip patterns. The pitfall the whole approach guards against is direct: ignoring the weights and design structure yields biased estimates, a form of selection biasDohoo et al. 2012, and confidence intervals that are too narrow.
To apply. Weight each observation by the inverse of its selection probability, then use design-based standard errors (Taylor linearization or replicate weights) that account for stratification and clustering; the design effect tells you how much the clustering inflates the variance.
Survey instruments: skip patterns and branching
When you operationalize a variable from a branching questionnaire, where a gate question routes respondents past items that do not apply to them, you have to decide how the resulting blanks are read, and survey skip patterns are the structure that answers it. The branching is fixed in the instrument design itself, so a skipped item is blank by design rather than missing. To reason about the overall idea, work with survey skip patterns as a single concept: a gate question routes a respondent past a set of items, and the analyst's job is to encode each route explicitly. When the element in play is the item that routes later questions, treat it as the gate question, because it is what determines which downstream items a person was never offered and therefore which blanks the branch implies. The discipline is to code a by-design blank as the value the branch implies rather than as missing, and to keep that skipped-by-design category strictly separate from the truly missing one of refused, did not know, or not reached. Reading a by-design blank as truly missing is the pitfall worth guarding against, since it discards information the instrument already gave you and can distort a denominator, inflating apparent nonresponse where there was none.
Operationalizing the variable
Operationalizing a variable means turning a clinical idea into data by writing a definition precise enough that two analysts produce the same cases, which is exactly when you reach for it: at the point where a concept must become codes, thresholds, and enrollment windows. "Adults with diabetes" is not a definition; an age range, a diagnosis code or lab threshold drawn from the coding standards, and an enrollment window are. The reason to be exacting is the pitfall loose definitions create: they invite reviewer questions, feed avoidable misclassification, and quietly change who the study is even about. When the variable being operationalized is a disease outcome in claims or records, the rules become a phenotype that must be validated, which is the subject of outcome phenotyping and algorithm validation. Once a single variable is defined this way, baseline health itself becomes a variable worth summarizing into a comorbidity or frailty score, so operationalization is the hinge between a clinical idea and every quantity the analysis will eventually estimate.
Comorbidity and frailty adjustment
Patients who receive a treatment often differ in how sick they already were, and that underlying illness drives outcomes on its own, so comorbidity and frailty adjustment is what you reach for to compress a messy claims history into a validated summary score and adjust for confoundingDohoo et al. 2012 by baseline health rather than guessing diagnosis by diagnosis. Each of these is an instance of operationalizing a variable from raw codes, so the same lookback and code-list choices apply.
- You need a mortality-weighted score, so the Charlson comorbidity index weights a short list of serious conditions toward mortality risk.
- You want broad comorbidity coverage, so the Elixhauser comorbidity measuresElixhauser et al. 1998 span a wider set and are often kept as separate indicators rather than collapsed into one number.
- Patients are older or frail with no direct functional assessment, so a claims-based frailty index approximates frailty from diagnosis and service codes.
The caution that travels with all of them is that adjusting on these scores blunts but does not erase confounding by indication, and an incomplete or miscoded diagnosis history feeds directly into the study bias you were trying to avoid.
To apply. Compute a weighted comorbidity score such as the Charlson index from diagnosis codes in the lookback window and enter it as a covariate, rather than adjusting for dozens of individual conditions.
Measurement error and misclassification
No instrument is exact, so measurement error and misclassificationDohoo et al. 2012 is worth considering whenever you judge how well the underlying quantity was actually measured, since the direction and structure of the error shape the bias in any estimate. The kind of error decides which way the bias runs.
- Non-differential misclassification is error unrelated to other variables, which for a binary exposure biases a single effect toward the null on average, that is, it shrinks the estimate toward "no effect" (a risk ratio toward 1, a difference toward 0), so a real effect looks weaker than it is. (Caveat: with three or more exposure categories even non-differential error can, less intuitively, bias away from the null.)
- Differential misclassification is error that differs by group, which can push the estimate either way, so its direction cannot be assumed.
- The reliability ratio \(\lambda\) is true variance over observed variance, the signal's share of total variance for a classically mismeasured continuous predictor; non-differential error multiplies the true slope by \(\lambda\), so an exposure with half its measured variance from noise has \(\lambda \approx 0.5\) and an estimated slope about half the truth.
The first question to ask of any estimate is therefore how well its quantity was measured, and reliability and validityDohoo et al. 2012 assessment is what quantifies and limits this before the analysis ever begins.
To prevent it. Reduce error at the source with validated instruments, standardized protocols, and blinded, identical measurement of exposure and outcome across groups, so any residual error is at least non-differential rather than differential. Run a validation substudy to estimate the sensitivity and specificity of the measurement, and use quantitative bias analysis to bound how much error remains and which way it pushes the estimate. Because non-differential error usually attenuates toward the null, a clear effect measured with noise is more likely an underestimate than an artifact.
Measurement-method effects
Measurement method effects arise when two devices or protocols measuring the same quantity disagree systematically, so consider them whenever you compare numbers produced by different methods. The canonical case is blood pressure: automated, rested, averaged office readings run lower than a single manual cuffBo et al. 2021, which means a value carries the fingerprint of how it was obtained, not just the underlying physiology. The practical consequence, and the pitfall, is that a threshold validated under one method does not transfer cleanly to another. When a guideline number and a clinic number were produced differently, they are not on the same scale, and treating them as interchangeable imports a differential measurement that is built into the method itself rather than into the patient. Reconciling the two means asking how each was measured before comparing them, so that a difference in protocol is not mistaken for a difference in the quantity.
Reliability and validity
A measurement has to be both reliable and valid, and because the two fail independently, reliability and validityDohoo et al. 2012 assessment is what you apply to confirm a measurement is reproducible and on-target before trusting the instrument. Reliability is reproducibility, getting the same answer when you measure the same quantity again, while validity covers content, construct, and criterion validityDohoo et al. 2012, whether the instrument measures the intended construct. The reliability statistic depends on the data.
- Reliability is consistency of measurement, the reproducibility that gives the same answer on remeasurement.
- Validity is measuring the intended construct, covering content, construct, and criterion validity.
- Cronbach's alphaCronbach 1951 gauges the internal consistency of items across a multi-item scale.
- The intraclass correlationShrout & Fleiss 1979 gauges agreement on continuous measures.
- Cohen's kappaCohen 1960 gauges categorical agreement between two raters, correcting agreement for what chance alone would produce; raters who agree 90% of the time but 80% by chance score \((0.90 - 0.80)/(1 - 0.80) = 0.5\).
- Weighted kappa gauges ordered-category agreement between two raters, crediting near-misses.
- Fleiss' kappa gauges categorical agreement among many raters, extending kappa past two raters scoring the same categories.
- The Bland-Altman plotBland & Altman 1986 charts method agreement and bias, plotting differences against means, which reveals a systematic offset that two methods can hide behind a near-perfect correlation.
The trap to keep in mind is that a reliable instrument can be precisely and repeatably wrong, a systematic error reliability alone never catches, and that agreement is not correlation.
To apply. Quantify reliability with the intraclass correlation (continuous measures) or a kappa (categorical), and internal consistency with Cronbach’s alpha; a reliable instrument can still be invalid, so check it against a reference standard.
Assembling the analytic cohort
A research database is not an analysis dataset, and turning one into a single analysis-ready table is the work of assembling the analytic cohort, fixing a single index date at which eligibility, exposure assignment, and follow-up start all align. Approach it as a sequence of construction steps.
- Pulling and reshaping raw source data uses extract-transform-load: draw from the source tables, derive the study variables from their operational definitions, and assemble one table at the right grain, one row per patient for a time-fixed cohort or one row per patient-interval when exposure or covariates change over time.
- Setting the time-zero anchor fixes the index date for each patient, exactly as target-trial emulationHernán & Robins 2016 prescribes, so that eligibility is assessed, exposure is assigned, and follow-up begins all at the same instant.
- Defining pre-index covariate history uses a lookback window, measuring confounders only before the index date so that adjustment captures baseline causes rather than post-exposure mediators or collidersDohoo et al. 2012, variables on or after the causal path that are introduced with the causal diagramsDohoo et al. 2012 in the Model rung.
- Restricting to treatment initiators applies a new-user design with a washout window so prevalent users do not contaminate the comparison.
The pitfall that voids the whole effort is getting the time alignment wrong: measuring confounders after T0 or letting prevalent users in manufactures immortal time before a single model runs, so each covariate should be stamped with the window in which it was measured and with how often it is missing.
Reading raw EHR and claims fields
EHR and claims data are recorded for care and billing, not analysis, so each column means what the source system meant by it, and the discipline of reading raw EHR and claims fields is checking how a field was populated before it becomes a variable. Three field-level traps recur.
- Which date is time-zero has to be pinned down first, because a single event carries several dates in separate fields: when the patient arrived or registered, when an order was placed, when a service actually started, and when the record was entered or the claim was filed. They routinely disagree by hours or days, and since the choice fixes the start of follow-up, the wrong one shifts the index date and can manufacture or erase immortal time. Reconcile the arrival, service, and recorded dates against the question rather than trusting whichever field loads first.
- Sequence and line numbers order the repeated records such as diagnoses, procedures, and claim lines, with the principal item usually at position 1 and secondaries after it. A sequence of 0 commonly marks a header, summary, or administrative placeholder rather than a real clinical record, so it is usually filtered out before analysis; counting it inflates totals and can double-count an encounter. Confirm what position 0 means in the source data dictionaryDohoo et al. 2012 instead of assuming.
- Ordered, dispensed, or administered are three very different things a medication field can mean: a prescription written, which is only intent; a pharmacy dispensing or fill, meaning the patient collected it; or an actual dose logged on an inpatient medication administration record. A written order is the weakest evidence of exposure, since the patient may never fill it (primary non-adherence) or never take it; a fill shows dispensing but not ingestion; the administration record is closest to a dose truly given. Settle which one you hold before fixing the exposure definition in RWD, since the medication possession ratio and proportion of days covered both assume dispensing rather than orders.
The common thread is that a field name rarely tells you how it was filled, so treating an order as a dose, a header row as an event, or a billing date as a clinical one are the quiet errors that survive every later model. Read the data dictionary before the data.
Defining exposure in real-world data
A prescription record or pharmacy claim is not exposure; it is a timestamped event, and exposure definition in RWD is the work of converting a string of such events into a variable with a start, an exposed window, and an end. The choices ripple straight into the analytic cohort and into a study's vulnerability to immortal time bias.
- Building one course uses exposure episode construction, stitching consecutive fills together.
- Tolerating gaps applies a grace period and permissible gap, extending coverage past the last day of supply and allowing a permissible gap between fills so a few late refills do not split one course into several.
- Shifting the clock when biology demands it uses induction, latency, and lag windowsDohoo et al. 2012, delaying the point at which exposure can plausibly cause the outcome.
- Summarizing an adherence metric uses the proportion of days covered (PDC) or the medication possession ratio (MPR), which both estimate the fraction of a period a patient had drug on hand and differ mainly in how they treat overlapping supplies and the denominator.
- Capturing how long a patient was treated uses persistence (time to discontinuation), treating the first gap beyond the permissible threshold as the endpoint.
- A standardized span reproducible across databases uses the drug era (OMOP), a derived span built from raw drug exposure records using an explicit persistence gap.
The pitfall is that mishandling permissible gaps, grace periods, or latency windows is a common route into immortal time, since classifying someone exposed only after surviving long enough to refill grants guaranteed survival time.
Outcome phenotyping and algorithm validation
An outcome in claims or electronic health record data is an algorithm, not a recorded fact: no one entered that a patient had a myocardial infarction, so outcome phenotyping and validation is the work of writing a rule that maps codes to a presumed event and then checking it. Because the rule is a classifier, treat it as such: the rule itself is a claims/EHR phenotype algorithm, governed by the logic of measurement error and the operating characteristics any comparison against truth produces.
- The 1-inpatient / 2-outpatient rule is a common coding rule, counting an event when there is one qualifying inpatient diagnosis or at least two outpatient diagnoses on separate dates, the second condition guarding against rule-out codes that appear once and never recur.
- Algorithm validation (the PPV and sensitivity tradeoff) is how you tune the rule: a specific rule yields high positive predictive value but misses true cases, whereas a broad rule captures more at the cost of false positives.
- Endpoint adjudication and chart review by clinicians blinded to exposure establish a reference standard, then report positive predictive value and, ideally, sensitivity.
- Composite endpointDohoo et al. 2012 construction builds a bundled outcome, stacking several phenotypes into one variable, where the weakest component algorithm tends to dominate the measurement error of the whole.
The pitfall is leaning on positive predictive value alone, because it depends on prevalence and ignores missed cases; and because phenotype misclassification is rarely symmetric across exposure groups, it can bias an estimate in either direction rather than simply attenuating it.
Immortal time bias
Immortal time bias arises when a span of follow-up during which the outcome could not have occurred is mistakenly assigned to the treated group, manufacturing a survival advantage out of bookkeeping. Recognize it in observational drug studies whenever follow-up that was guaranteed to be event-free risks being credited to the treated arm, which is why the start of follow-up has to be defined as carefully as the exposure itself. The mechanism is arithmetic: a rate is \(\text{events} / \text{person-timeDohoo et al. 2012}\), so assigning event-free immortal time to the treated arm enlarges its denominator and lowers its event rate by construction, regardless of any real effect. A concrete case: if patients count as treated only once they fill a prescription, they had to survive to reach the pharmacy, so that waiting stretch is event-free by definition. Crediting it to the treated group makes the treatment look protective when nothing about the drug caused it. The pitfall, then, is precisely that misallocation, and the remedy is built upstream: aligning the index date when assembling the cohort, so that eligibility, exposure assignment, and the start of follow-up coincide, is exactly what prevents the bias before any model runs.
To prevent it. Define time zero so that eligibility, exposure assignment, and the start of follow-up coincide, then handle the exposure with one of the standard fixes: model it as a time-varying covariate so the stretch before treatment counts as unexposed; use a landmark analysis, which fixes exposure status at a chosen landmark time and follows only those who survive to it; or, for a "treat within a grace period" question, use the clone-censor-weight approach. Target-trial emulationHernán & Robins 2016 builds all of this in by pinning time zero up front.
Missing data: MCAR, MAR, MNAR
When you need to decide whether listwise handling, imputation, or a sensitivity analysis is appropriate, the governing question is why values are missing, because the reason dictates what you are allowed to do about it. Treat missing data as the overarching concern, and then classify the mechanism.
- MCAR (missing completely at random) is missingness unrelated to anything, which is benign and leaves even simple handling unbiased.
- MAR (missing at random) is missingness explained by the observed data, which can be handled by multiple imputationRubin 1976 conditional on what you did observe.
- MNAR (missing not at random) is missingness that depends on the unseen values themselves, which cannot be fixed by imputation alone and instead needs a pattern-mixture or tipping-point sensitivity approach that shows how far from missing-at-random the data would have to be to overturn the result.
- Multiple imputation fills the gaps and pools estimates under MAR; in practice, multiple imputation by chained equations (R's mice, SAS PROC MI and MIANALYZE) is the workhorse, where you place the outcome and every analysis variable in the imputation model, generate twenty or more imputations, and pool with Rubin's rules.
Under a complex survey design the imputation has to be made design-consistent: put the survey weight, the strata, and the primary sampling units into the imputation model alongside the analysis variables, impute, then analyze each completed dataset with design-based standard errors and pool, so the variance carries both the design-based component within each imputation and the between-imputation component from Rubin’s rules (SAS PROC SURVEYIMPUTE, or PROC MI then PROC SURVEYMEANS by imputation then MIANALYZE). Two neighbors stay distinct from this: a value blanked by a questionnaire’s skip pattern is recoded rather than imputed, and unit nonresponse is absorbed by the survey weights rather than filled in.
The consequential shortcut to avoid is treating all missingness as ignorable, because that quietly assumes away the MNAR case that imputation cannot rescue.
Assembling a clinical trial dataset
A clinical trial dataset is assembled the opposite way from a real-world cohort, and assembling a clinical trial dataset is mostly standardization and traceability rather than bias-prevention, because the protocol already fixes eligibility, randomized assignment, and the start of follow-up by design. Raw case-report-form data is first mapped to CDISC SDTM, one domain per kind of data (DM for demographics, AE for adverse events, CM for concomitant medications, LB for labs, VS for vitals, EX for exposure), a tabulation that mirrors what was collected, and from the Study Data Tabulation Model analysis datasets are derived in the Analysis Data Model (ADaM).
- One row per subject means ADSL, the subject-level dataset, alongside basic-data-structure datasets carrying the derived endpoints the statistical analysis plan calls for.
- A medication-level dataset has one row per reported drug, coded to a dictionary such as WHO Drug and its Anatomical Therapeutic Chemical (ATC) classes so concomitant medications can be summarized by class, since not every analysis dataset is subject-level.
- Linking a result back to its source rests on traceability: every analysis value has to trace back through ADaM to its SDTM source and the original case reportDohoo et al. 2012 form (CRF).
The governing rule, and the pitfall if neglected, is traceability: without it a reviewer cannot follow any number in the results to the record it came from.
Characterizing the distribution
Before any model, look at what you measured; characterizing the distribution is the first exploratory step that keeps later choices honest, read from histograms, the interquartile range, and a LOESS-smoothed scatter. Work through the shape features it surfaces.
- Skewness summarizes the asymmetry of the distribution, telling you whether a mean is even the right summary.
- Kurtosis summarizes the heaviness of the tails, warning of the outliers that distort the mean and standard deviation, so spread is better read with the median and interquartile range (IQR) that survive them.
- A LOESS smoother fits a smooth curve over a scatter to reveal a nonlinear trend before you assume a relationship is linear.
Exploratory analysis is where you catch the heavy tail, the floor-or-ceiling effect, the multimodality that usually signals two subpopulations mixed together, and the nonlinearity, all of which decide which model and which summary are honest. The pitfall of skipping it is that these features slip past unnoticed, so a chosen mean or model may not honestly summarize the data, and you meet the problem later in a reviewer's question rather than now; what this step surfaces is precisely what the robust statistics and regression families downstream are there to handle.
To apply. Summarize center (mean, median), spread (standard deviation, interquartile range), and shape (skewness, kurtosis), and plot the data before modeling; a long tail or heavy kurtosis warns that mean-based methods may mislead.
What to compute, by data type. For one variable: a continuous measure calls for a histogram or density and a QQ plot to read the shape, with the median and interquartile range alongside the mean and SD (the former resist outliers), watching for heavy tails, floor or ceiling effects, and multimodality; a binary or categorical variable calls for a frequency table and bar chart, where for a binary outcome the event rate is the summary and doubles as a prevalence; a count or rate is usually right-skewed, so compare the mean and variance (a variance well above the mean is the overdispersionDohoo et al. 2012 that later sends Poisson regression to negative binomial regression) and check for excess zeros; a time-to-event variable calls for a Kaplan-Meier curve and the median survival with the censoringRothman et al. 2008 pattern described, since a mean is undefined under incomplete follow-up. For a relationship between two variables: two continuous variables call for a scatter with a LOESS smoother before assuming linearity, plus a Pearson correlation if linear or Spearman if monotone but curved; a continuous variable by group calls for boxplots or violins with group medians; and category by category calls for a cross-tabulation with row or column proportions.
Probability distributions and the CLT
Behind the empirical shape of the data sit the theoretical probability distributions that model it and drive inference, and the practical use is to pick a model family by matching the distribution to the outcome, which is the same decision as choosing the regression family.
- The normal distribution, also called Gaussian, models a continuous bell-shaped variable, whose standardized form is the z; the same logic extends to a gamma or lognormal for skewed positive data such as costs and lengths of stay, a beta for proportions bounded between zero and one, and the exponential or Weibull for survival times.
- The binomial distributionDohoo et al. 2012 models fixed-trial success counts.
- The Poisson distributionDohoo et al. 2012 models counts of rare events.
- The negative binomial distribution models overdispersed count data.
- The central limit theorem (CLT) explains why sample means turn normal: the mean of a large enough sample is approximately normal whatever the underlying shape, which is what lets z- and t-based inference work on data that is not itself normal, with the t, chi-square, and F serving as the reference distributions that give a p-value its meaning.
- The standard error is the spread of a sample estimate: averaging \(n\) observations shrinks the spread of the mean in proportion to \(1/\sqrt{n}\), so quadrupling \(n\) halves the standard error.
The pitfall is conflating the two roles: the CLT makes sample means approximately normal, but assuming an individual outcome is itself normal when it is skewed or a bounded proportion misspecifies the model.
To apply. Match the distribution to the outcome (binomial for yes/no counts, Poisson for rare event counts, normal for continuous), then lean on the central limit theorem: the sample mean is approximately normal for large samples even when the data are not.
A model turns raw data into a comparison: the effect of a treatment, an exposure, a risk factor. The catch is that the groups being compared usually differ for reasons other than the one you care about. Confounding is the name for those other reasons, a shared cause that makes an effect look larger or smaller than it really is, and much of this rung is machinery for handling it.
A model is fit for one of three reasons, and the reason sets the standard it is judged by: to explain an association (you read the coefficients, so keep it simple and check the assumptions), to estimate a cause (you want one unbiased effect, so the work goes into confounding rather than fit), or to predict an outcome (you want accuracy on unseen cases, so a black box is fine if it calibrates). The groups below follow that split, with clustered-data and Bayesian methods cutting across all three.
How a causal effect is estimated
These nodes are the steps of one workflow, not a glossary of terms. Estimating a causal effect from data runs in a fixed order, and the pieces below slot into it.
- Name the estimand. In potential-outcomes terms, state the exact causal quantity and in whom before choosing a method: the ATE, ATT, or LATE, and in a trial the intercurrent-event strategy and analysis population.
- Draw the structure. A causal diagram (DAG) sorts each variable into confounder, mediator, or colliderDohoo et al. 2012 and reads off the adjustment set: what to control, and what to leave alone.
- Get identification. Randomization identifies the effect by design; without it, choose the quasi-experimental design that neutralizes your dominant threat (difference-in-differencesDimick & Ryan 2014, instrumental variablesAngrist et al. 1996, regression discontinuityThistlethwaite & Campbell 1960, synthetic controlAbadie et al. 2010, interrupted time series) and name the one identifying assumption it rests on, a claim the data cannot test and you must argue.
- Estimate. With the adjustment set fixed, use stratification for a few confounders, or the causal estimators (propensity scoresRosenbaum & Rubin 1983 and g-methods, with their time-varying and claims-scale extensions) for the rest; confirm covariate balance and positivityHernán & Robins 2020 before trusting the number.
- Decompose, if the question is mechanism. When you need to know how an effect works, mediation analysis splits it into direct and indirect parts.
- Defend it. The estimate is only as sound as that untestable assumption, so it goes to the Defend rung for sensitivity analysis and bias quantification.
The spine is estimand → structure → identification → estimator → defense; every node below is one of those steps.
Potential outcomes and identifiability
Causal inference starts before any design, with a way to even define an effect, and potential outcomesHernán & Robins 2020 and identifiability supply it. The potential-outcomes framework imagines, for each unit, the outcome under treatment, \(Y(1)\), and under no treatment, \(Y(0)\); the causal effect is their contrast, and the fundamental problem of causal inference is that only one of the two is ever observed. Closing that gap from data takes four identifiability conditions, and you reach for each depending on which is at stake.
- Whether treated and untreated are comparable once confounders are controlled needs exchangeabilityDohoo et al. 2012 (no unmeasured confoundingDohoo et al. 2012).
- Whether every covariate stratum could have received either treatment needs positivityHernán & Robins 2020 (overlap).
- Whether the observed outcome equals the counterfactual under the treatment actually received needs consistency, so that the treatment is a well-defined intervention and \(Y(1)\) means something specific.
- Whether one unit's treatment affects another's outcome needs SUTVA, the stable unit treatment value assumption of no interference and a single version of treatment.
The design-specific identifying assumptions are how a particular study argues for exchangeability when it cannot simply be assumed. The pitfall to keep in view is that only one potential outcome is ever observed, and even when exchangeability holds in the sample, carrying the effect to another population is the separate problem of transportability.
A second model of causation. The counterfactual model says what an effect is; the sufficient-component cause modelDohoo et al. 2012 (Rothman's causal pies) says how causes combine to produce it. A sufficient cause is a complete set of conditions that inevitably produces the outcome, and most diseases have several; a component cause is one factor within such a set, and a factor can sit in several pies; a necessary cause is a component present in every sufficient cause, without which the outcome cannot occur. The co-factors that must join an exposure to complete a pie are its causal complement. Two consequences matter for appraisal. First, because a factor only acts when its causal complement is present, a measured effect size depends on how common those co-factors are, not on biology alone, so the same exposure can show a risk ratio of 4.8 in one population and 2.9 in another with no change in mechanism; this is why an effect is not a biological constant and why consistency across populations carries weight. Second, because component causes are shared across pies, the population attributable fractions of the separate factors can sum past 100%, each double-counted. This looks paradoxical but is not: a single component cause can belong to several sufficient causes at once, and removing it would prevent every case that needed it, so each factor is credited with the full case rather than a fraction, and the credits overlap. Biological synergism, two components in the same pie, shows up statistically as interaction, but the distribution of unmeasured co-factors can hide or exaggerate it, so absence of interaction in the data is not absence of synergism.
Choosing the estimand
Name the target before the method: choosing the estimand is stating the exact quantity to be estimated, which effect and in whom, before any technique is picked. Which one you mean changes both the magnitude and the policy meaning of the answer, and settling it first is what keeps the later method honest.
The causal estimand, which contrast and in whom.
- The ATE (average treatment effect) is the effect across the whole population.
- The ATT is the effect among the treated, \(E[Y(1) - Y(0) \mid \text{treated}]\).
- The CATE is the effect within a covariate stratum, \(\tau(x) = E[Y(1) - Y(0) \mid X = x]\), the basis of heterogeneous-effect and subgroup analysisDohoo et al. 2012.
- The LATE is the effect among compliers only.
These coincide only when the effect is constant across groups, so reporting a number without saying which one invites a policy meaning it does not carry.
In a trial, name the intercurrent-event strategy too. The International Council for Harmonisation (ICH) E9(R1) framework forces the estimand to be stated before post-randomization events, a patient stopping the drug, switching, taking rescue medication, or dying before the endpoint, can muddy it. Each is handled by a named strategy that fixes a different estimand: the treatment-policy strategy ignores the event and counts the outcome regardless (the intention-to-treat spirit); the hypothetical strategy estimates the outcome had it not occurred; the composite strategy folds the event into the endpoint; and the principal-stratum strategy restricts to a subpopulation such as those who would never have the event, with a while-on-treatment option as well.
Which subjects you analyze is the same choice. The analysis population is where the estimand becomes concrete. Intention-to-treat keeps every patient in the arm assigned, preserving randomization and answering the pragmatic question of offering a treatment; per-protocol restricts to those who followed protocol, answering the biological question of taking it as directed but breaking randomization and risking confoundingDohoo et al. 2012; as-treated groups by treatment actually received. The trap: intention-to-treat is conservative (makes it harder to declare a difference, so any error favors the null) for a superiority trial but anti-conservative (makes an error more likely, here a false claim of similarity) for a non-inferiority one, since dropout blurs the arms toward no difference. In a superiority trial that blurring works against you, so a win is credible; in a non-inferiority trial the same blurring flatters the drug by manufacturing the very sameness you are trying to prove, so both intention-to-treat and per-protocol are reported in the non-inferiority setting.
The pitfall across all three is failing to name the estimand and its strategy up front, which lets two analysts quietly answer different questions while believing they agree.
Causal diagrams (DAGs) and conceptual frameworks
A conceptual framework is a picture of what causes what, and causal diagramsDohoo et al. 2012 are its formal version, used to sort each covariate by its structural role and read off an adjustment set. The older informal version is the web of causationKrieger 1994, which pictures disease as chains of direct (proximal) and indirect (distal) causes; its enduring lesson is that you can often prevent disease by acting on a manipulable indirect cause without knowing the proximal mechanism, the way removing a water-pump handle stopped cholera before the organism was known. The graph notation itself is the DAG (directed acyclic graph): variables as nodes, assumed causal effects as arrows, no cycles allowed. Drawing it turns the vague worry of what might confound this into a decision you can defend, because the graph classifies each covariate.
- A confounder is a common cause of exposure and outcome, which you adjust for.
- A mediator is a variable on the causal path, which you leave alone when the total effect is the target.
- A colliderDohoo et al. 2012 is a common effect of two others, where conditioning actively opens bias rather than removing it.
- The back-door criterion tells you which adjustment set is sufficient, reading it straight off the graph.
The pitfall is that a DAG cannot prove itself: it encodes assumptions, and the arrow you left out, the unmeasured common cause, is exactly the identifying assumption the design then has to defend.
What to do with each variable, by its role. Read each candidate covariate off the diagram. A confounder, a common cause of exposure and outcome, must be adjusted for, since it is the back-door path you have to block. A mediator, on the causal path, should be left alone when the total effect is the target and conditioned only when you are explicitly decomposing direct and indirect effects. A collider, a common effect of two variables, must not be adjusted, because conditioning on it opens a spurious path and induces bias where none existed. A pure outcome predictor, related to the outcome but not the exposure, is optional and usually tightens precision. An instrument, which affects the exposure only, should be kept out of the outcome model, since adjusting for it can amplify residual confoundingDohoo et al. 2012. A post-treatment variable, measured after the exposure, is generally left out, as a descendant of the exposure is a mediator or collider in disguise, which is why covariates are defined in a pre-exposure window. When the goal is prediction rather than a causal effect, none of this applies: select variables for out-of-sample performance with regularization such as lassoTibshirani 1996 or ridge regression, not by causal role.
Causal designs without randomization
When you seek a causal effect without randomization, causal designs without randomization each neutralize a specific dominant threat, and the craft is matching the design to the threat that actually endangers your question rather than reaching for the most familiar tool.
- A before-after comparison across an exposed and a control group uses difference-in-differencesDimick & Ryan 2014, which removes fixed differences between groups and common time trends.
- A single group observed repeatedly before and after an intervention at a known time uses interrupted time series (segmented regression), fitting a level change and a slope change at the interruption when there is no control group to difference against.
- A haphazard nudge to exposure that is otherwise unrelated to the outcome uses instrumental variablesAngrist et al. 1996, which recovers an effect through that as-good-as-random variation.
- Assignment determined by a cutoff threshold on some running variable uses regression discontinuityThistlethwaite & Campbell 1960, comparing units just above and just below the cut.
- Weighted donors that can build a counterfactual for a single treated unit use synthetic controlAbadie et al. 2010, alongside propensity-score methods for balancing measured covariates.
The pitfall is reaching for the most familiar tool instead of the one matching the actual threat, which can leave you with a merely associational study you mistake for causal; if no design fits, recognizing that you have an associational study is itself worth knowing before you claim otherwise.
To defend it. Each design rests on an identifying assumption you must argue, not merely invoke. Difference-in-differences assumes parallel trends, that the groups would have moved together absent treatment: support it with pre-period trends, event-study leads, and placebo outcomes. Interrupted time series assumes the pre-intervention trend would have continued unchanged absent the intervention, so the extrapolated baseline is the counterfactual, and that no other change coincided with the interruption; it must model autocorrelation (Newey-West or an ARIMA error structure) or the standard errors run too small, and adding a comparison series (a controlled interrupted time series, the single-group cousin of difference-in-differences) sharpens it. Instrumental variables assume a strong first stage (relevance, testable by the first-stage F statistic) and the exclusion restriction, that the instrument affects the outcome only through the exposure, which is not testable and must be defended on substance. Regression discontinuity assumes potential outcomesHernán & Robins 2020 are continuous at the cutoff and that units cannot precisely manipulate the running variable, checked with a density test at the threshold; its estimate is local to that cutoff. Synthetic control assumes a close pre-treatment fit and a donor pool unaffected by the treatment, probed with in-space and in-time placebo tests.
Causal estimators (propensity scores, g-methods)
Once the design and adjustment set are fixed, a causal estimator turns them into a single number, the average contrast of potential outcomes. The choice follows from how you are willing to model the data and from whether treatment happens once or unfolds over time.
Point treatment, confounders measured. The workhorses when treatment is a single decision and the confounders are in hand.
- The propensity score models the probability of treatment, giving a score to match treated to untreated on.Rosenbaum & Rubin 1983
- IPTW reweights by the inverse treatment probability, building a pseudo-population where treatment is independent of the measured confounders; stabilized weights and truncation rein in the extreme weights a near-zero score produces.Robins et al. 2000
- The g-formula (g-computation) models the outcome under each treatment and averages over the covariates.Hernán & Robins 2020
- Doubly-robust estimators combine a treatment and an outcome model, staying consistent if either is right; TMLE brings machine learning to the nuisance models,van der Laan & Rubin 2006 and cross-fitting (fit the nuisance models on one fold, evaluate on another) lets such flexible models be used without overfitting bias.Chernozhukov et al. 2018
Trust nothing before the diagnostics agree: confirm covariate balance (standardized mean differences below 0.1), inspect propensity-score overlap for positivity, and truncate extreme weights so a few patients do not dominate. R: MatchIt, WeightIt, cobalt, tmle.
Treatment repeated over time. When a confounder both responds to the last dose and guides the next (treatment-confounder feedback, as a lab value does in a chronic-disease cohort), ordinary adjustment breaks: controlling the time-varying confounder blocks the path you want and opens a collider path you do not. Concretely, the lab value confounds the next dose, so you must adjust for it, but it also sits on the causal chain from the earlier dose to the outcome, so adjusting for it discards part of the effect you are after and conditions on a shared effect (a collider). Ordinary regression cannot do the first without also doing the second. The g-methods are the fix, effectively the only one, run on the person-time table (one row per patient-interval): a marginal structural model fitted by IPTW, g-estimation of a structural nested model, or the g-formula iterated over time.
Claims scale, evolving treatment. When there are thousands of candidate covariates and treatment, censoring, and confounding all shift over follow-up, reach for the real-world extensions: the high-dimensional propensity score (hdPS) screens thousands of claims codes as confounder proxies; a disease risk score summarizes outcome risk in place of treatment probability; clone-censor-weight estimation emulates a per-protocol target trial, cloning each patient into every strategy and censoring on deviation, to avoid immortal time bias, with inverse-probability-of-censoring weighting (IPCW) correcting the informative censoring it creates; and landmark analysis classifies exposure only as of a fixed later time.
To apply. For a point treatment, estimate the propensity score then match, weight (IPTW), or standardize (g-formula), preferring doubly-robust when either model might be wrong; when treatment is time-varying, switch to g-methods on the person-time table; check balance and positivity before trusting any of it.
Stratified analysis (Mantel-Haenszel)
Before regression made adjustment automatic, confoundingDohoo et al. 2012 was controlled by stratified analysis: you split the data by the confounder, estimate the association within each stratum where the confounder no longer varies, and pool the stratum-specific estimates into one summary. This is the tool to reach for when you want transparent, assumption-light confounding control across a few categories.
- Pooling across strata uses the Mantel-Haenszel estimatorDohoo et al. 2012, which combines stratum-specific odds ratios, risk ratios, or rate ratios into a single adjusted figure.
- Combining stratum log effects uses Woolf's methodDohoo et al. 2012, which does the same on the log scale.
- Testing for effect modification uses a homogeneity check, the step that earns its keep and comes first, before any pooling.
That homogeneity check is the discipline of the whole approach, because if the stratum-specific estimates differ by more than noise, that difference is effect modification, and collapsing them into a single number hides the very thing worth reporting. Stratification stays transparent and makes few modeling assumptions, but it runs out of room past a few confounders, which is where regression and propensity methods take over.
To apply. Compute the effect within each stratum, test homogeneity, and if the stratum effects agree pool them with the Mantel-Haenszel estimator; a pooled estimate that differs from the crude one is the signature of confounding by the stratifying variable.
Mediation analysis
Sometimes the question is not whether a treatment works but how, and mediation analysis is what you reach for to answer it, splitting a total effect into a direct effect and an indirect effect that runs through a mediator. This is exactly why a mediator must not be adjusted away when the total effect is what you want: doing so removes part of the very effect you are trying to measure. When the decomposition itself is the goal, the modern counterfactual framing decomposes the effect through the mediator into natural direct and indirect effectsRobins & Greenland 1992, and that is the concept in play whenever you ask which part of an effect travels through a given pathway. On a difference scale the split is additive, so a path whose \(a\) or \(b\) coefficient is near zero carries almost none of the total.
The decomposition carries its own identification price, and this is the pitfall to watch. On top of the usual conditions, it requires no unmeasured confoundingDohoo et al. 2012 of the mediator-outcome relationship, an assumption that randomizing the treatment cannot buy, because randomization assigns the treatment but not the mediator. Reported carefully, mediation analysis explains a mechanism; reported loosely, it smuggles in the very confounding it was meant to expose.
Bivariate tests (t-test, chi-square)
Before fitting a full model, the simplest question is whether two variables are associated, and a family of classical bivariate tests answers it unadjusted, with the test matched to the variable types.
- Mean across two groups calls for a t-test.
- Means across three or more groups calls for ANOVA, which extends it.
- Ranks across two groups calls for the Mann-Whitney test, the rank-based counterpart that carries the question without assuming normality.
- Ranks across three or more groups calls for the Kruskal-Wallis test, its many-group rank equivalent.
- Two categorical variables, large counts call for a chi-square testDohoo et al. 2012 of independence.
- Two categorical variables, small counts call for Fisher's exact testDohoo et al. 2012, which takes its place when cell counts are small.
- Linear correlation of two continuous variables uses Pearson correlation.
- Monotonic correlation, ranked uses Spearman correlation, the monotone version on ranks.
- Concordance-based rank correlation uses Kendall's tau, the concordance between two ordinal rankings.
- Effect size for two means uses Cohen's d.
- Effect size for ANOVA uses eta-squared, the share of variance the groups explain.
- Effect size for categorical association uses Cramer's V.
Each test also carries an effect-size measure the p-value hides, and reporting it is part of the work. The fact worth holding onto is that each of these is a special case of a regression model; a two-sample t-test is a linear regression on a single binary indicator, one-way ANOVA a regression on a categorical factor, and a test of independence a log-linear or logistic model. The bivariate test is the unadjusted answer, and the regression that generalizes it is the same comparison with covariates added.
To apply. Match the test to the data: a t-test (or the nonparametric Mann-Whitney) for two means, ANOVA (or Kruskal-Wallis) for several, a chi-square (Fisher exact when cells are small) for proportions, and a correlation coefficient for two continuous variables.
Regression families
The outcome dictates the model, and most of these regression families are one family underneath: a GLM is a choice of outcome distribution paired with a link function, which is what makes matching the model to the outcome a principle rather than a lookup table.
Choose by outcome type. The outcome fixes the model, and the model fixes the effect measure you report:
- Continuous → linear regression (mean difference).
- Binary → logistic regression (odds ratio).
- Ordered categories → ordinal, proportional-odds regression (odds ratio).
- Unordered categories → multinomial (polytomous) logistic regression (an odds ratio per non-reference category).
- Count or rate → Poisson regression (rate ratio), negative binomial when overdispersed, and a zero-inflated or hurdle model when zeros are in excess.
- Time-to-event → Cox or parametric survival (hazard ratio), or a discrete-time model with a complementary log-log link when events fall in intervals.
- Correlated or clustered → GEE for a population-average effect (the average change across the whole group, the quantity a policy question wants), or a mixed-effects model (the MMRM for a longitudinal trial endpoint) for a subject-specific one (the change within a given patient, the quantity a prognosis question wants).
Because each family emits its own kind of effect measure, the family you fit decides which estimate you are even reporting, and forcing the wrong one on an outcome produces estimates that are precise and wrong. The paragraphs below expand that table: how the families share one GLM backbone, and how to choose within a row.
To apply. Pick the link and distribution from the outcome (identity and Gaussian for continuous, logit and binomial for binary, log and Poisson or negative binomial for counts), fit by maximum likelihood, and read each coefficient on the model scale, exponentiating a log-odds or log-rate coefficient to get the ratio.
For clustered or repeated outcomes. GEE and mixed-effects models both relax the independence a plain GLM assumes, but on different terms. GEE posits a working correlation structure (exchangeable, autoregressive, or unstructured) and stays consistent even when that structure is misspecified, provided you use robust sandwich standard errors; it answers a population-average question. A mixed-effects model instead specifies random effects and a covariance structure, which is more efficient and gives subject-specific estimates but is biased if that structure is wrong. Pick GEE when you want a marginal effect robust to the correlation, a mixed model when you want subject-specific inference, and inspect the residual covariance either way.
Watch the load-bearing assumptions. The ordinal model rests on proportional odds; the multinomial model drops that ordering at the cost of a separate set of coefficients per category; and a time-to-event model rests on proportional hazards, checked by scaled Schoenfeld residualsDohoo et al. 2012 with restricted mean survival timeRoyston & Parmar 2013 as the fallback when it fails. And all of this is for explaining an interpretable effect: if the goal is instead predicting outcomes on new data, reach for flexible methods such as gradient boosting or regularized regression, judged on out-of-sample performance rather than coefficient plausibility. Either way, the fit is trustworthy only once its assumptions are checked.
Checking model assumptions
Distinct from the identifying assumption a causal design rests on, every regression carries statistical assumptions that are checkable with diagnostics, and checking model assumptions is what you do so the standard errors and p-values can be trusted; a model whose assumptions fail produces inference that cannot. A few checks run on the raw data before fitting, but most are post-fit because they read the residualsDohoo et al. 2012.
- Heteroscedasticity is non-constant residual variance, which shows up in a residual-versus-fitted plot, confirmed if needed with Breusch-Pagan or White.
- The variance inflation factorDohoo et al. 2012 (VIF), or the condition number, flags predictors that are too collinear, read on the raw data before fitting.
- Cook's distance catches single points driving the fit.
- The proportional hazards assumption that Cox models add holds the hazard ratio constant over time.
- Scaled Schoenfeld residuals test proportionality formally, alongside a log-log survival plot that should show parallel curves, or a Poisson model with a time-by-covariate interaction that should come out null.
- Heteroscedasticity-robust (sandwich) standard errors fix the variance without refitting, the modern default for non-constant variance.
One caution governs all of this: formal normality and heteroscedasticity tests are underpowered at small n and oversensitive at large n, so they make poor gates. The discipline is the remedy, not the test: a failed check sends you to a transformation, cluster-robust standard errors, or a different family, not to reporting the broken fit anyway.
Model modifications (splines, interactions)
A base regression is rarely the final model, and a few standard model modifications adapt its specification to the data and the question, each answering a specific signal rather than a reflex to lift the fit.
- Effect depends on another variable, so an interaction termDohoo et al. 2012 captures that effect modification, letting the estimate differ across subgroups instead of being averaged across them.
- Fixed exposure term in count models is carried by an offset, which holds the fixed exposure time or population at risk, turning a Poisson count into a rate.
- Smooth nonlinear flexible curves come from splines, restricted cubic or natural ones, which fit a smooth piecewise curve at a few knots, more stable than high-order polynomials.
- Additive smooth function components come from generalized additive models, which extend the same idea across several predictors.
When the outcome or a predictor is skewed or acts multiplicatively, a transformation such as a log or square root can restore the assumptions a linear model needs. Each modification answers a particular feature of the data or the design, so the discipline is to add one because the data or the question asked for it, not because it improved the fit statistic.
Linear combinations and contrasts
A regression returns coefficients one at a time, but the quantity worth reporting is often a linear combination of them, so reach for this whenever the answer is a contrast rather than a single coefficient. A contrast is a weighted sum of the coefficients: you put a weight on each one, usually +1, -1, or 0, and add them up, so the weights are simply a recipe for the comparison you want. Three recipes recur.
- A between-group contrast (the gap between two non-reference groups) is one group coefficient minus the other (weights +1 and -1), which cancels the shared reference category each was measured against.
- A predicted value at a chosen covariate setting adds the intercept to each coefficient times the value you fix for its predictor, returning a fitted mean rather than a difference.
- A subgroup treatment effect, when the model carries an interaction, is the main treatment term plus the interaction termDohoo et al. 2012 (a weight of +1 on each), because the subgroup effect is the baseline effect shifted by the interaction, not either term alone.
The pitfall is the standard error. Because a contrast adds coefficients that are correlated, its standard error is not the sum of their individual standard errors; it has to be computed from the model's full variance-covariance matrix, which carries those correlations. A contrast whose standard error ignores them is wrong no matter how clean the point estimate looks. The tools that return the contrast and its correct standard error together are lincom and contrast in Stata, or emmeans and glht in R.
Robust statistics for heavy tails
Means and standard deviations are fragile when the data have heavy tails or outliers, as cost and utilization data almost always do, and robust statisticsHuber 1964 are what you reach for so a few extreme points no longer set the scale. The reason ordinary summaries fail here is concrete: a handful of providers or patients can drive the totals, and a mean and standard deviation let those few values dominate every downstream cutoff.
- Robust spread of the data uses the median absolute deviation (MAD), which summarizes variability without being inflated by the tail.
- Robust outlier-resistant standardization uses the robust z-score, which recenters on the median and rescales by the MAD instead of the mean and standard deviation, so the usual cutoffs still apply while the extreme points no longer determine the scale.
Median-based summaries paired with MAD-scaled scores are the practical default whenever the distribution is heavy-tailed and a few observations would otherwise dictate the result.
Multiplicity control
Test enough hypotheses and some will look significant by chance, and multiplicity control is how you rein that in, with the right approach set by what a false positive costs.
- The family-wise error rate is the chance of any false positive at all; control it when even one false positive is costly.
- The false-discovery rateBenjamini & Hochberg 1995 is the expected share of false positives among the rejections; control it when screening or running a large flagging exercise.
- The Bonferroni correction is the simple conservative FWER divisor, dividing alpha by the number of tests.
- Holm's procedure is stepwise FWER control, achieving the same family-wise control step by step with more power.
- Benjamini-Hochberg and its step-up relatives give step-up FDR control.
- A gatekeeping procedure tests hypotheses in ordered families, spending alpha down the sequence and testing a secondary endpoint only if the primary already won, protecting the family-wise rate without dividing alpha equally.
The trade-offs cut both ways: family-wise procedures grow punishingly strict as the tests multiply, while the false-discovery procedures assume things about the null distribution that are not always true. And when some of the tested cells are nearly empty, controlling error also means handling sparse data with appropriate small-sample and resampling methods.
To apply. When testing many hypotheses, control the family-wise error rate (Bonferroni alpha over m, or Holm) when any single false positive is costly, or the false-discovery rate (Benjamini-Hochberg) when you can tolerate a known fraction of false leads.
Sparse data and resampling
With small cell counts or rare events, ordinary maximum likelihood estimation becomes biased or fails to converge, and the large-sample intervals from the regression families can mislead, which is when sparse data and resampling methods take over.
- Separation or small samples call for Firth penalized regressionFirth 1993, which adds a small bias-reducing penalty to the likelihood that keeps estimates finite even under complete or quasi-complete separation, the situation in which a covariate perfectly predicts the outcome and ordinary estimates run to infinity.
- Very sparse, exact inference calls for exact logistic regression, which conditions on sufficient statistics and enumerates the permutation distribution, giving valid inference without leaning on asymptotic approximations, at the cost of heavier computation as the data grow.
- No clean closed-form variance calls for bootstrap and resampling methodsEfron 1979, which repeatedly draw samples with replacement from the observed data and recompute the estimate, building an empirical sampling distribution that yields percentile or bias-corrected intervals.
Population-impact measures such as attributable risk and the population attributable fraction, defined among the effect measures in rung 03, are nonlinear functions of several estimates, so a bootstrap interval is often more trustworthy than a delta-method approximation for their confidence limits.
Prediction and machine learning
When the goal is to predict an outcome rather than to explain it or to estimate a cause, and the relationships are nonlinear or high-dimensional, prediction and machine learning earn their place: flexible learners such as gradient boosting and regularized regression, read where needed with interpretability tools like SHapley Additive exPlanations (SHAP). What sets this apart is the standard of judgment: a predictive model is graded only on how well it predicts cases it has never seen, not on whether its coefficients are interpretable or plausible.
That is a different goal from the other two reasons to fit a model, and the goal decides how the model is built and scored.
- To explain uses an inferential regression family, where the payoff is the coefficients themselves: their sign, size, and confidence interval, and whether the associations are plausible, so the model stays simple and its assumptions are checked.
- To estimate a cause uses the causal toolkit, where the payoff is one unbiased effect (a contrast of potential outcomesHernán & Robins 2020), so the effort goes into controlling confoundingDohoo et al. 2012 rather than into fit; a model that predicts beautifully can still return the wrong effect.
- To predict uses machine learning, where the payoff is accuracy on new cases, so a black-box learner with uninterpretable parameters is fine as long as it ranks and calibrates well. A model can predict well while explaining nothing, and a model whose coefficients explain a mechanism may predict poorly.
Within prediction, which learner fits depends first on whether the data carry a label, the supervised versus unsupervised split, and a classifier is then scored on its own performance metrics, not on its coefficients; SHAP attributes a single prediction across the input features when you need to read one. The failure that sinks most clinical prediction models is data leakage, so run every preprocessing and feature-selection step inside the cross-validation folds, validate on a temporally or geographically external sample before trusting the area under the ROC curve (AUC), and report to the TRIPOD standard. Skip those guards and the AUC looks strong on paper and collapses in deployment.
To apply. Hold out data or cross-validate, tune regularization (lassoTibshirani 1996 or ridge) to control overfitting, and judge a classifier by precision, recall, and the C-statistic on the held-out set rather than by in-sample fit.
Supervised and unsupervised learning
Machine learning splits by whether the data carry an outcome label, and supervised and unsupervised learningHastie et al. 2009 is that fork: decide the task by asking whether a label is available. When the data carry a known target, a diagnosis, a cost, a survival time, supervised learning learns to predict it, and it includes the regressions already covered plus the more flexible algorithms and ensembles. When there is no label, unsupervised learning instead finds structure: groups of similar patients through clustering, or a lower-dimensional summary of many variables through dimensionality reduction.
The line a biostatistician should hold across this fork is the explain-versus-predict one. Flexible machine learning earns its place when prediction is the goal and the relationships are nonlinear or high-dimensional, but a black-box predictor is not a causal model, and its coefficients, where it has any, are not effect estimates. Treating a predictive model as if its weights explained mechanism is the mistake to guard against.
Bias-variance and regularization
Every predictive model trades bias against variance, and bias-variance and regularizationHastie et al. 2009 is the framework for tuning that trade so out-of-sample error is what gets minimized. Expected prediction error decomposes into bias, variance, and irreducible noise, and the goal is the flexibility that minimizes their sum, not the training error, because minimizing training error chases noise and overfits.
- The bias-variance tradeoff is the underlying error tradeoff: a model too simple to capture the signal is biased and underfits, whereas a model flexible enough to chase noise has high variance.
- Overfitting is fitting noise at the cost of generalization, fitting the training data well but failing on new data.
- Regularization penalizes complexity broadly, buying the right flexibility.
- Cross-validationStone 1974 estimates out-of-sample error, reading error on held-out folds rather than trusting the in-sample fit, and so chooses the right amount of flexibility.
- Ridge regressionHoerl & Kennard 1970 shrinks coefficients while keeping all of them, applying an L2 penalty.
- LassoTibshirani 1996 shrinks and selects variables, applying an L1 penalty that drives some coefficients exactly to zero.
- Elastic netZou & Hastie 2005 blends selection and shrinkage, combining the two, while early stopping does the same job for iterative learners by halting before they overfit.
The right amount of flexibility is chosen by cross-validation, which estimates out-of-sample error on held-out folds rather than trusting the in-sample fit.
Learning algorithms and ensembles
Beyond regression lies the broader supervised toolkit, the learning algorithms and ensembles you reach for when prediction is the goal.
- A decision tree gives single interpretable splits, partitioning the predictors into regions, readable but unstable on its own.
- k-nearest neighbours classifies by closest neighbours, predicting from the majority or average of the k nearest, simple but sensitive to scaling and dimensionality.
- The support vector machineCortes & Vapnik 1995 finds the maximum-margin separating boundary, the widest margin between classes, and uses a kernel to bend it nonlinearly.
- BaggingBreiman 1996 averages parallel bootstrapped models, training many trees on bootstrap resamples and averaging them to lower variance.
- The random forestBreiman 2001 is many decorrelated bagged trees, the standard bagged ensemble.
- Boosting sequentially corrects prior errors, fitting trees in sequence to lower bias, with gradient boostingFriedman 2001 and XGBoost as the standard.
Ensembles routinely beat a single model, but at the cost of interpretability, which tools such as the SHapley Additive exPlanations (SHAP) values only partly restore.
Unsupervised learning
With no outcome to predict, the unsupervised learning half of the supervised/unsupervised split finds structure instead, and which method fits depends on whether you are grouping observations or compressing variables.
- Clustering groups similar observations, used for phenotyping, finding subtypes of a disease from a panel of measurements.
- Hierarchical clustering groups by linkage, building a nested tree of groupings without fixing k beforehand.
- K-meansLloyd 1982 partitions into k groups, minimizing within-cluster distance to the cluster mean, with k chosen in advance.
- Dimensionality reduction reduces the number of features, compressing many correlated variables into a few.
- Principal component analysisJolliffe 2002 finds orthogonal variance components, the orthogonal directions of greatest variance, while methods like UMAP or t-SNE do a nonlinear version for visualization.
The standing caution across all of these is that a discovered cluster is a pattern, not a diagnosis, so it needs external validation before it means anything clinical.
Classification performance metrics
A classifier's performance is read off the confusion matrix of predicted versus actual, and the classification metrics follow from it, with the metric you optimize set by the clinical cost of each kind of error.
- Precision is the share of predicted positives that are correct, the same quantity as positive predictive value.
- Recall is the share of true positives caught, the same as sensitivity.
- The F1 score balances precision and recall, their harmonic mean, high only when both are high.
- The ROC-AUCHanley & McNeil 1982 summarizes the tradeoff across all thresholds, but on a rare outcome it can look strong while precision is poor, so under class imbalance a precision-recall curve is the more honest summary of the same tradeoff.
Which metric to optimize is ultimately a clinical choice about the cost of a false positive versus a missed case, the same trade that a decision curve formalizes.
Clustered and longitudinal data
Many datasets are not flat tables of independent rows: patients sit within hospitals, measurements repeat within a patient, animals within a herd. Clustered and longitudinal data is the name for that structure, and it earns its own step because of one fact that runs under every model on this rung: observations inside a cluster are correlated, so they carry less information than their raw count suggests. Treat them as independent and the point estimate can be fine while the standard errors come out too small, the confidence intervals too narrow, and a null effect looks significant. Recognize the structure before choosing a model, whenever a sampling or measurement scheme groups observations.
- How much clustering there is is the intraclass correlationShrout & Fleiss 1979 (ICC), the share of total variance that sits between clusters rather than within them.
- What it costs you is the design effectDohoo et al. 2012 (roughly 1 + (m − 1)ρ, with m the cluster size and ρ the ICC), which inflates the sample size a clustered study needs and shrinks its effective n.
- The cheapest correction is cluster-robust (sandwich) standard errors, which widen the intervals to respect the clustering without touching the point estimate or the model.
- A population-average effect comes from generalized estimating equations (GEE), robust to a misspecified working correlation.
- A subject-specific effect, with the variance decomposed, comes from a mixed-effects (multilevel) model, the mixed model for repeated measures (MMRM) being its standard form for a longitudinal trial endpoint; its Bayesian twin is the hierarchical model, which borrows strength across clusters by partial pooling.
- Stable between-cluster confoundingDohoo et al. 2012 you want swept out calls for fixed effects, which absorb every time-invariant difference between clusters at the cost of any between-cluster effect.
The recurring pitfall is treating clustered data as if it were flat: the estimate may be right while its uncertainty is wrong, which is the failure cluster-randomized trials and repeated-measures designs are most often caught on. Match the response to the question, a marginal effect from GEE or a subject-specific one from a mixed model, but account for the clustering one way or another before trusting any interval.
Bayesian inference
Bayesian inferenceGelman et al. 2013 reverses the frequentist setup: where frequentist methods treat the parameter as fixed and the data as random, the Bayesian framework treats the parameter as a random quantity with a distribution that the data update. Reach for it when you want a full picture of belief about a parameter rather than a point estimate alone, and when you want an interval that says what people wrongly assume a confidence interval says.
- Bayes' theorem is the updating rule itself, under which the posterior is proportional to the likelihood times the prior, so what you believed before, weighted by what the data say, becomes what you believe after.
- The posterior distribution is your beliefs after seeing the data, summarized by its mean or median.
- The credible interval is the interval summary of the posterior; a 95 percent credible interval is a range the parameter lies in with 95 percent probability. That is the direct probability statement people mistakenly attach to a frequentist confidence interval: the credible interval genuinely licenses "95 percent probability the true value is in here," whereas the confidence interval only says intervals built this way capture the truth 95 percent of the time over repeated samples.
With abundant data the prior washes out and the two paradigms converge. The pitfall lives at the other extreme: with little data the prior does real work, for better or worse, so a poorly chosen prior can drive the conclusion rather than the evidence. The framework's deepest payoff is the partial pooling of a hierarchical model.
Choosing a prior
Choosing a priorGelman et al. 2013 is where a Bayesian analysis is won, lost, and most often attacked, so the question is which kind of prior the problem calls for and why.
- Strong external knowledge with sparse data calls for an informative prior, which earns its place by encoding that knowledge where the data alone would leave the estimate adrift.
- Light regularizing information calls for a weakly-informative priorGelman 2006, which supplies just enough information to stabilize the fit while letting the data dominate; a flat or non-informative prior tries to stay out of the way entirely, though truly uninformative priors are slippery.
- Computational convenience calls for a conjugate prior, chosen to match the likelihood so that the posterior has the same form as the prior and the update is closed-form; a beta prior with a binomial likelihood returns a beta posterior, a normal with a normal returns a normal, and a gamma with a Poisson returns a gamma.
Because the prior is the most attacked part of the analysis and can dominate sparse data, the honest discipline is a prior sensitivity analysis: refit under several defensible priors to show the conclusion does not hinge on one. That is what answers the subjectivity charge, since a Bayesian prior is an explicit, checkable assumption rather than a hidden one.
Bayesian computation (MCMC)
Most posteriors have no closed form, so Bayesian computation explores them by simulation, and which sampling or diagnostic tool you reach for depends on the stage you are at. This is the same machinery as Monte Carlo simulation pointed at the posterior, and it is what lets a Bayesian network meta-analysisDohoo et al. 2012 fit an entire evidence network at once.
- Sampling the posterior in general uses Markov chain Monte Carlo (MCMC)Dohoo et al. 2012, which draws a dependent sequence of samples whose long-run distribution is the posterior.
- A proposal-and-accept scheme uses Metropolis-HastingsHastings 1970, proposing a move and accepting or rejecting it.
- Sampling each parameter from its conditional uses Gibbs samplingGeman & Geman 1984, one parameter at a time.
- Efficient exploration of a high-dimensional posterior uses Hamiltonian Monte CarloHoffman & Gelman 2014, the gradient-guided engine of Stan, which mixes far more efficiently.
- Checking the chains have converged uses R-hat, which should sit near 1, meaning chains started far apart have mixed; the effective sample sizeKish 1965 should be large enough and trace plots should look like noise rather than trends.
- Checking the model reproduces the data uses a posterior predictive check, which asks whether data simulated from the fitted model resemble the real data.
Because the draws form a dependent chain rather than independent draws, convergence has to be checked before you trust the output.
Hierarchical (multilevel) Bayesian models
The deepest reason to go Bayesian is the hierarchical Bayesian modelGelman et al. 2013, also called the multilevel model, which you reach for when data come in groups: hospitals, studies, or patients with repeated measures. Rather than fitting one pooled estimate that ignores the groups or fully separate estimates that ignore each other, a hierarchical model estimates each group's parameter while letting the groups share a common prior whose variance is itself estimated.
The mechanism that does the work is partial pooling, also called shrinkage, and it is the idea to invoke when you want to borrow strength across groups. Each estimate is pulled toward the overall mean by an amount that depends on how noisy that group is, so partial pooling stabilizes small and sparse groups by borrowing strength from the rest, sitting between the extremes of one pooled estimate and fully separate estimates. The watch-point is exactly this pull: a noisy small group is shrunk substantially toward the overall mean rather than taken at face value, which is usually a feature but should be understood deliberately.
Structurally this is the same model as a random-effects meta-analysisDerSimonian & Laird 1986, with studies as the groups, and a mixed-effects model, with clusters as the groups, which is where the Bayesian and frequentist worlds meet.
Once a model runs it produces a number: how big the effect is. This rung is about reading that number honestly, in the right units, starting from the raw disease frequencies it is built on, and being equally honest about how uncertain it is.
Measures of disease frequency
Counting how often disease occurs comes before comparing groups, and the measures of disease frequency you reach for depend on the question and your follow-up. Underneath, every measure is a count, a proportion, an odds, or a rate, and the sharpest split is between a risk (a dimensionless proportion of people, valid only in a closed populationDohoo et al. 2012) and a rate (events per person-timeDohoo et al. 2012, unbounded); a risk is recovered from a rate by \(R = 1 - e^{-I\,\Delta t}\). The forms also connect by a rule of thumb: prevalence is roughly incidence times average duration, precisely \(P = ID/(1+ID)\), so a rise in prevalence can mean more new cases or simply longer survival. These are occurrence measures, distinct from the comparative effect measures a study later estimates.
- Prevalence is existing cases at a time point, the share of a population with the condition, at a point (point prevalenceDohoo et al. 2012) or over a window (period prevalenceDohoo et al. 2012), reflecting both how often the disease arises and how long it lasts.
- Incidence is new cases over follow-up, the rate at which new cases arise.
- Cumulative incidenceDohoo et al. 2012 is new cases as a proportion of those at risk, new cases over a fixed period divided by the population at risk.
- The incidence rateDohoo et al. 2012 is new cases per unit of follow-up time, dividing new cases by person-time at risk when follow-up varies.
- Person-time is the denominator of summed follow-up, each subject's observed time added across the cohort.
- A crude rateDohoo et al. 2012 is the unadjusted rate in a population.
- Age-standardizationDohoo et al. 2012 is what comparing rates across populations with different age structures needs, because the raw comparison is confounded, applied either directly (apply your group's age-specific rates to a standard population's age structure) or indirectly (apply a standard population's age-specific rates to your group's age structure) through the standardized mortality ratioDohoo et al. 2012.
- The standardized mortality ratio is observed over expected deaths, where "expected" is the death count you would see if your group experienced a reference population's age-specific mortality ratesDohoo et al. 2012, so a ratio above 1 means more deaths than that reference would predict.
- Risk versus rate turns on the population: a closed population followed for the full risk period yields a risk directly, whereas an open populationDohoo et al. 2012 that people enter and leave yields a rate, from which risk is then estimated.
- The case fatality rateDohoo et al. 2012 is deaths among cases, how deadly a disease is for those who get it, and despite its name is a risk, not a rate; contrast the mortality rate, deaths per person-time across the whole population.
- Attack ratesDohoo et al. 2012 carry this into outbreaks, cases over the number exposed, with a secondary attack rateDohoo et al. 2012 capturing spread to close contacts.
To apply. Use a proportion (cases over population) for prevalence and risk, a rate (events over person-time) when follow-up varies, and age-standardize before comparing populations with different age structures. Decide risk versus rate by whether the population is closed or open, and use an exact (binomial or Poisson) confidence interval when counts are small.
Effect measures
The same result reads differently depending on the scale, so the first question is which effect measure to report and on what scale it lives. Hazard ratios are relative on the rate scale.
- The risk ratio is the ratio of risks between groups, dividing the two risks, a relative measure.
- The incidence rate ratioDohoo et al. 2012 is the same comparison on the rate scale, dividing the two incidence ratesDohoo et al. 2012, which is what you reach for when follow-up varies.
- The odds ratio is the ratio of odds between groups, with the odds written p over one minus p, also relative.
- The risk differenceDohoo et al. 2012 is the absolute difference in risk, subtracting one risk from the other on the absolute scale.
- The number needed to treatLaupacis et al. 1988 is patients treated per outcome prevented, how many patients must be treated for one outcome to be prevented.
- Attributable risk is the excess risk among the exposed, the risk difference read as the portion of an exposed group's risk that the exposure accounts for.
- The population attributable fraction (PAF) is the share of population cases preventable, the proportion of cases across the whole population that would be removed if the exposure were eliminated.
The odds ratio and risk ratio diverge as the outcome risk climbs toward 1. A risk is the share who have the event (out of everyone), while an odds is the ratio of that share to its complement (event to no-event), so odds outrun risk as risk approaches 1. Concretely, a rise from 40 percent to 60 percent is a risk ratio of 1.5 but an odds ratio of \((60/40)/(40/60) = 2.25\). That divergence is the pitfall: odds ratios are routinely misread as risk ratios when the outcome is common, which overstates the effect. As a rule the three line up with the odds ratio furthest from 1, the rate ratio next, and the risk ratio closest.
To apply. Estimate each group's risk as events divided by total, then form the ratio (risk ratio, odds ratio) or the difference (risk difference); when the outcome is rare the odds ratio approximates the risk ratio, but as risk rises it moves farther from 1 and must not be reported as a risk ratio.
Uncertainty and inference
A point estimate without its uncertainty is half a result, so uncertainty and inference is what you attend to whenever you report an estimate. The everyday tool is the confidence interval, which shows the plausible range compatible with the data; reach for it on any point estimate. When observations cluster, such as patients within hospitals or repeated measures within a person, the standard errors have to account for that clustering or they will be too small, and a tight interval around a biased estimate is false comfort.
The companion lesson is that significance is not importance. A large enough sample makes a clinically trivial difference statistically significant, and a small one can miss an important effect, so what to report is the effect size and its interval, not the p-value alone. The confidence interval is also the frequentist quantity most often misread. The wrong reading: a 95 percent CI does not mean there is a 95 percent probability the true value lies in this particular interval. The right reading: if you repeated the study many times and built an interval each way, 95 percent of those intervals would capture the truth; any single interval either contains it or does notGreenland et al. 2016. The direct "95 percent probability the truth is in here" statement readers want is exactly what a Bayesian credible interval provides instead.
To apply. For ratio measures (risk ratio, odds ratio, hazard ratio) build the interval on the log scale and exponentiate the two ends; when observations cluster, use a robust or cluster-adjusted standard error, or the interval comes out falsely narrow.
Relative versus absolute
Given the effect measures, the relative versus absolute choice is which scale you lead with, and it is a communication decision with real stakes. A 50 percent relative reduction sounds dramatic and can still be a move from 2 percent to 1 percent. Patients and decisions live on the absolute scale, so when you want the benefit to be concrete you report the absolute risk reduction and the number needed to treatLaupacis et al. 1988, which make the size of the benefit and the baseline risk it depends on visible.
The number needed to treat is the reciprocal of the absolute risk reduction, so a 2-percent-to-1-percent move is an absolute risk reduction of 0.01 and a number needed to treat of 100, while the relative reduction reads 50 percent at any baseline.
The pitfall is one-sided reporting: leading with only the relative effect is the most common way a modest result is made to sound large, since a 50 percent reduction can be nothing more than a move from 2 percent to 1 percent.
To apply. A move from 2 percent to 1 percent gives \(\text{ARR}=0.01\), \(\text{RRR}=0.5\), and \(\text{NNT}=100\); report the absolute risk reduction and the number needed to treat alongside any relative figure so the baseline risk stays visible.
Hazard ratios and non-proportional hazards
A single hazard ratio assumes the treatment's effect on instantaneous risk is constant over time, and you reach for it to summarize a treatment effect when that proportional-hazards picture roughly holds. When a competing event blocks the outcome, or when the hazard ratio stops being constant, move on to competing risks and parametric survival models.
- Non-informative censoringRothman et al. 2008 is censoring unrelated to the outcome, the assumption underlying every survival estimate: those censored are representative of those who remain at risk. Informative dropout violates it, so when censoring is tied to prognosis the estimate is suspect.
- Restricted mean survival timeRoyston & Parmar 2013 is a summary for when hazards are non-proportional, a number that survives delayed effects, waning benefit, or crossing curves and that a patient can actually use.
The central pitfall is that when proportional hazards does not hold, the reported ratio becomes a weighted average that depends on the censoring pattern rather than the clinical story.
To apply. Estimate the hazard ratio from a Cox model, which assumes that ratio is constant; test that assumption (for example with Schoenfeld residualsDohoo et al. 2012), and when it fails report the restricted mean survival time, the area under the survival curve up to \(\tau\), as a difference in months.
Competing risks and parametric survival
A competing risk is an event that makes the event of interest impossible afterward; cardiovascular death, for instance, prevents a later cancer diagnosis. You move into competing risks and survival models whenever such an event removes patients along the way, and you pick cause-specific for mechanism and Fine-Gray for absolute risk, reporting both when the audience needs both.
- You want etiology calls for the cause-specific hazard, which describes the instantaneous rate among patients still at risk and speaks to whether the exposure changes the biological rate of the event.
- You want absolute risk calls for the cumulative incidenceDohoo et al. 2012 function (CIF), which gives the actual probability of experiencing the event by a given time, accounting for the competing events that remove patients.
- Modeling that absolute risk directly uses the Fine-Gray subdistribution hazardFine & Gray 1999, which models the cumulative incidence function, making it the tool for prognosis or resource planning. The two can disagree: a drug that raises the competing cause of death (say cardiovascular death) lowers the cumulative incidence of the event of interest (say a cancer recurrence) simply because fewer patients survive long enough to reach it, so the Fine-Gray absolute risk falls even while the cause-specific hazard, the biological rate of recurrence itself, is unchanged.
- Proportional hazards fails calls for accelerated failure time (AFT)Dohoo et al. 2012 models, which sidestep it by modeling the log of survival time, yielding a time ratio for how much exposure stretches or compresses survival.
- Some patients are cured calls for cure models, which split the population into a cured component and a susceptible component with its own survival distribution.
The pitfall to avoid is naive Kaplan-Meier, which treats those competing deaths as censoringRothman et al. 2008 and so assumes the patient could still have had the event later; that assumption is false here, and it overstates the cumulative risk of the event you care about.
To apply. For absolute risk when competing events are present, use the cumulative incidence function (Aalen-Johansen), not one minus Kaplan-Meier, which overstates risk; model the cause-specific hazard for etiology or the Fine-Gray subdistribution hazard for absolute risk.
Calibration versus discrimination
Judging a risk model that drives a bedside decision turns on two distinct aspects of performance, which is the calibration versus discrimination distinction.
- Discrimination is ranking cases above non-cases, whether the model ranks higher-risk patients above lower-risk ones.
- The area under the ROC curve (AUC) summarizes ranking across thresholds, equal to the probability that a randomly chosen case is given a higher predicted riskHanley & McNeil 1982 than a randomly chosen non-case, so 0.5 is chance and 1 is perfect ranking.
- Calibration is whether predicted risks match observed, read off a calibration plot against the 45-degree line and summarized by the calibration slope and calibration-in-the-largeSteyerberg et al. 2010.
- The Hosmer-Lemeshow statistic is the classic formal test of calibrationHosmer & Lemeshow 1980 and is still widely reported, though it is sensitive to how cases are grouped and to sample size, so it is best read alongside a calibration plot rather than relied on alone.
The pitfall is treating good ranking as enough: a model can discriminate well and still be badly miscalibrated, which is the failure that matters when a number drives a decision, since portability, whether the model still gives right numbers in a new population rather than only in the one it was built on, depends on calibration, and calibration is the more often neglected of the two. The continuous-outcome counterpart of this whole comparison is model fit and prediction error.
To apply. Read discrimination from the C-statistic (0.5 is chance, 1 is perfect ranking); check calibration separately by plotting observed event rates against predicted risk across bins, or the calibration slope (ideal is 1), since a model can discriminate well yet be poorly calibrated.
Model fit, comparison, and prediction error
Model fit, comparison, and prediction error are the continuous-outcome counterpart to calibration and discrimination, and they form the family most often misread, so the question is which fit or error measure the task needs. The F-testDohoo et al. 2012 answers a different question, whether the model as a whole beats an intercept-only null (a baseline model with no predictors that just guesses the overall mean for everyone), with the likelihood-ratio test and devianceDohoo et al. 2012 playing the same role for generalized linear models, where deviance measures how far the fitted model sits from a perfect fit.
- R-squared is the variance explained by a linear model, the share of outcome variance the model accounts for, but it climbs mechanically as predictors are added, so use adjusted or out-of-sample R-squared and never read a high value as proof the model is correct.
- McFadden's pseudo-R-squared is the pseudo-fit for a logistic model, the stand-in when the outcome is binary.
- The Akaike information criterion (AICAkaike 1974) compares models while penalizing parameters, trading goodness of fit against the number of parameters.
- The Bayesian information criterion (BIC) compares models while penalizing more heavily, charging each extra parameter more so it favors smaller models, with lower better for either.
- The mean absolute error is prediction error as average magnitude, the choice when a few big errors should not dominate.
- The root mean squared error (RMSE) is prediction error that penalizes large misses, sitting in the outcome's own units and punishing large misses hardest, alongside its squared form the mean squared error.
In-sample fit always flatters and AIC only approximates out-of-sample performance, so the figure that truly generalizes is the one computed on data the model never saw.
To apply. Compare candidate models by AIC or BIC (lower is better; BIC penalizes extra parameters more heavily), and judge predictive error on held-out data through cross-validation, because in-sample fit such as \(R^2\) is optimistic.
Safety and adverse-event analysis
Efficacy is only half a trial; safety analysis is the other half, and you analyze it differently whenever the task is to characterize trial harms. Adverse events are tabulated by type and graded by severity, and the counting is done on the safety population, meaning everyone who received any treatment rather than the randomized set, since harm follows exposure. Comparisons are made as risk differencesDohoo et al. 2012 or as exposure-adjusted incidence ratesDohoo et al. 2012 that account for differing time on drug.
The defining methodological choice is that safety is deliberately not corrected for multiplicity the way efficacy is, because missing a real harm signal is worse than a false alarm. So safety stays largely descriptive and hypothesis-generating, with rare serious events watched case by case rather than significance-tested. The asymmetry is the point and the pitfall: a trial is powered to detect benefit, not to rule out uncommon harm, which is why the first dose-tolerability read happens earlier, in the early-phase designs.
To apply. When follow-up differs between arms, use the exposure-adjusted incidence rate (for example events per 100 person-years) rather than a crude proportion, and compare arms with the incidence rate ratioDohoo et al. 2012.
One study is rarely the last word. Synthesis is how separate studies are combined into a single body of evidence, and how that combined evidence is graded for how far it can be trusted.
Conducting a systematic review
Before anything is pooled or appraised, conducting a systematic reviewDohoo et al. 2012 means a protocol-driven, pre-registered search, and the conduct is what separates it from a casual literature summary. The steps are a pre-registered question scoped in PICOS (population, intervention, comparator, outcome, study design), a reproducible search across several databases with the exact strings recorded, dual independent screening of titles and abstracts and then full texts against the eligibility criteria, structured data extraction, and a PRISMA flow diagram that accounts for every record from initial hits to included studies.
The decisive safeguard is registering the protocol on PROSPERO before screening begins; reach for it at the outset, because skipping it lets the review quietly become a search for the result you wanted. Only once that scaffold is in place do the pooled estimate, the risk-of-bias appraisal, and the certainty rating that follow it mean anything.
Meta-analysis and pooling
Meta-analysisDohoo et al. 2012 and pooling combine several studies estimating the same effect into a single weighted summary, but pooling can either sharpen an estimate or average away a real difference, depending on whether the studies are estimating the same thing. The base mechanism is inverse-variance pooling, which gives each study a weight and returns the weighted mean, so more precise studies receive more weight.
- A fixed-effect meta-analysis assumes one true effect, weighting by one over the variance alone.
- A random-effects meta-analysisDerSimonian & Laird 1986 allows effects to vary across studies, adding the between-study variance, tau-squared, to each weight (the DerSimonian-Laird estimator is the classic version), widening the interval and pulling the pooled estimate toward the simple average.
- Heterogeneity is how much effects vary, a measure of how much the studies' results actually disagree.
- Cochran's Q tests for heterogeneity.
- I-squared is the proportion of variance due to heterogeneity, the share of variation beyond chance.
- Tau-squared is the between-study variance estimate itself, on the same scale as the effect measure, so it says how much the true effect really varies from study to study rather than what fraction of the scatter is heterogeneity.
- The prediction interval is the range a new study's true effect might fall in, wider than the confidence interval.
- Meta-regression explains heterogeneity by covariates, attempting to account for it with study-level covariates.
- The funnel plot visualizes small-study effects, checking for publication bias.
- Egger's test tests funnel asymmetry.
The pitfall is misreading precision as agreement: a tight pooled estimate over heterogeneous studies, signaled by a high I-squared, can average away a real difference, so heterogeneity should warn you rather than reassure you. These pooled estimates are also the inputs a decision-analytic model runs on, so the care taken here propagates into any cost-effectiveness verdict downstream.
To apply. Weight each study by the inverse of its variance, pool to a summary effect, and quantify heterogeneity with Cochrans Q and I-squared; when heterogeneity is non-trivial use a random-effects model, which adds the between-study variance to the weights and widens the interval. Adding that shared variance to every study's weight makes the weights more equal, so the largest studies stop dominating the pooled estimate and it pulls toward the plain average of the studies rather than the precision-weighted one.
Report this review to PRISMA (Preferred Reporting Items for Systematic reviewsDohoo et al. 2012 and Meta-Analyses).
Network meta-analysis
Standard meta-analysisDohoo et al. 2012 pools head-to-head trials of two treatments. When a question involves several treatments and no trial has compared them all directly, network meta-analysis, also called mixed or indirect treatment comparison, combines the whole evidence network to estimate every pairwise contrast and rank the options; the arithmetic is subtraction through a shared comparator, and the network has to be connected.
- Transitivity is the key premise of comparability across the network, that the trials are similar enough in their populations and methods that an indirect comparison through a common comparator is valid.
- Node-splittingDias et al. 2010 checks direct-versus-indirect agreement, comparing the direct and indirect estimate for each contrast and flagging disagreement.
- The surface under the cumulative ranking curve (SUCRASalanti et al. 2011) ranks treatments overall, where 100 percent is certainly best and 0 percent certainly worst.
The pitfall sits on both the assumption and the ranking: the whole synthesis rests on transitivity, and a high SUCRA rank built on sparse evidence deserves caution. The estimation is often Bayesian, fitting the whole network by MCMC, and it feeds many economic models needing relative effects for treatments the trials never lined up against each other.
To apply. Fit the connected network with a frequentist model (for example the R netmeta package) or a Bayesian random-effects model, which pools direct and indirect evidence to estimate every pairwise contrast at once. Before trusting it, check transitivity by comparing effect modifiers across the comparisons, test consistency with node-splitting and a global inconsistency model, and rank options with SUCRA, treating ranks built on sparse or low-certainty evidence with caution.
Risk-of-bias appraisal
Not all evidence deserves equal weight, so risk-of-bias appraisal is what you reach for when you need to weigh a study's credibility domain by domain rather than on a gestalt impression. Which structured tool fits depends on the design.
- Randomized trials use RoB 2 (Risk of Bias 2)Sterne et al. 2019, which scores how the design and conduct threaten the result across its domains.
- Non-randomized intervention studies use ROBINS-I (Risk Of Bias In Non-randomized Studies of Interventions)Sterne et al. 2016, which plays the same role for observational evidence.
The value of either tool is that it scores design and conduct domain by domain rather than as an overall hunch, so skipping it lets an unreliable result carry the same weight as a sound one.
Certainty of evidence (GRADE)
Certainty of evidence rates how much confidence a body of evidence warrants, separately from the size of the effect, and you turn to it once individual studies have been appraised and pooled. The named framework is GRADE, the Grading of Recommendations Assessment, Development and Evaluation approachGuyatt et al. 2008: when you are rating confidence in a set of estimates, GRADE is the tool, downgrading for risk of bias, inconsistency, indirectness, imprecision, and publication bias.
The error this guards against is a common one in how findings get reported: a large effect from low-certainty evidence and a small effect from high-certainty evidence are different things, and conflating them misleads. Keeping certainty separate from effect size is the point, and this certainty rating is what the strength of a recommendation should track.
Certainty is only the first half of GRADE: the Recommendation rung takes it up through the evidence-to-decision framework and turns it into a strength of recommendation.
Generalizability and transportability
An effect estimated in one population does not automatically apply to another, so generalizability and transportabilityDohoo et al. 2012 is what you weigh whenever you ask whether a result carries to the population you actually care about. Generalizability asks whether the study sample represents the target, which is where selection biasDohoo et al. 2012 and an unrepresentative sampling design do their damage; transportability formalizes when and how an estimate can be carried to a different population.
Because that damage is easy to overlook, the honest default is the narrower claim, with extrapolation argued rather than assumed.
Evidence has to become an action: treat above this number, screen at this age, cover this drug. This rung turns an estimate into a rule someone can follow, and weighs what the rule costs against what it buys.
How a diagnostic or decision rule comes together
The nodes in this section are steps in building a rule you can act on, not a glossary. A test or score earns clinical use only when acting on it does more good than harm, and the pieces below assemble toward that verdict.
- Measure the test’s accuracy. Sensitivity and specificity are its intrinsic performance, estimated in a diagnostic-accuracy study against a reference standard while guarding against spectrum and verification biasDohoo et al. 2012; when no reference exists, latent-class estimation stands in.
- Read a result for the patient. Predictive valuesDohoo et al. 2012, what a positive or negative actually means, depend on prevalence, so the same test reads differently in a clinic than in a screening program.
- Turn the measurement into an action. A continuous test or risk score becomes a yes-or-no call at a threshold, which slides sensitivity against specificity and quietly encodes the cost of acting versus waiting.
- Combine predictors into a rule. A risk calculator merges several predictors into one probability, judged on calibration and discrimination and validated out-of-sample before use.
- Decide whether the rule is worth using. Decision-curve analysisVickers & Elkin 2006 weighs the true positives a rule catches against the false positives it triggers across the range of threshold probabilities a clinician might hold, showing whether it beats treating everyone and treating no one.
The through-line is net benefit: accuracy and calibration are necessary but not sufficient. A rule earns its place only if acting on it, at a plausible threshold, does more good than harm, which is exactly what the last step tests.
Operating characteristics
Turn to operating characteristics when you need to translate a test's sensitivity and specificity into what a positive or negative result actually means for the patient in front of you. The choice among the measures follows the question being asked.
- Sensitivity is true positives among the diseased, the fraction of truly diseased patients the test catches.
- Specificity is true negatives among the healthy, the fraction of the truly healthy it correctly clears.
- Predictive valuesDohoo et al. 2012 give the disease probability given a particular result, and shift with prevalence, since sensitivity and specificity describe a test only in the abstract.
That dependence on prevalence is just Bayes' rule, and it is the pitfall to watch: the same test that is reassuring in a high-prevalence clinic can generate mostly false positives in a low-prevalence screening setting, since at low prevalence the false-positive term, \((1 - \text{spec})(1 - \text{prev})\), comes to dominate. Sensitivity and specificity travel with the test; predictive values travel with the patient population, which is why a result must always be read against the prevalence in which it was obtained. Two further moves recur: combining tests, where series interpretation (positive only if every test is positive) raises specificity while parallel interpretation (any positive counts) raises sensitivity; and correcting an imperfect test's apparent prevalenceDohoo et al. 2012 \(AP = P\,\text{Se} + (1-P)(1-\text{Sp})\) back to the truth by the Rogan-Gladen formula \(P = (AP + \text{Sp} - 1)/(\text{Se} + \text{Sp} - 1)\).
Diagnostic-accuracy studies
Use a diagnostic-accuracy study when you are measuring how well an index test discriminates disease against a reference standard, comparing the two in a two-by-two table that yields sensitivity and specificity. The design has its own traps, and which concept you invoke depends on what you are after.
- How results shift disease odds calls for likelihood ratiosDohoo et al. 2012, which summarize how a result shifts disease odds independent of prevalence and update odds directly, \(\text{LR}^+ = \text{Se}/(1-\text{Sp})\), with post-test odds equal to the likelihood ratio times the pre-test odds.
- Index test informs the reference standard produces incorporation bias, where the index test is itself part of the reference standard.
- Only some get the reference standard produces verification, or work-up, bias, where only some patients, typically the test-positive ones, go on to the reference standard.
- Unrepresentative case mix produces spectrum bias, where floridly sick cases and plainly well controls inflate the estimates, so accuracy measured at a referral center overstates accuracy in primary care.
The recurring pitfall is that sensitivity and specificity are properties of the study sample, not of the test alone, so the reference standard must be applied to everyone regardless of the index result. An honest study enrolls a clinically realistic spectrum, verifies every patient, and reports to the Standards for Reporting Diagnostic accuracy studies (STARD).
To prevent it. Enroll a clinically realistic spectrum of patients (consecutive or representative sampling), not floridly sick cases against plainly well controls, so the estimate transfers to practice (against spectrum bias); apply the same reference standard to everyone regardless of the index result, or correct for partial verification, so accuracy is not read off a test-positive subset (against verification biasDohoo et al. 2012); and keep the index test out of the reference standard so it is not judged partly against itself (against incorporation bias). Report against STARD.
Evaluating a test without a gold standard
A diagnostic-accuracy study assumes the reference standard is itself perfect, but often the best available reference is imperfect, or there is no gold standardDohoo et al. 2012 at all, and evaluating a test without a gold standard is what you reach for then. Treating an imperfect reference as truth biases the new test's sensitivity and specificity, usually downward, because the test is penalized for the reference's own errors.
- An imperfect reference standard biases accuracy in a knowable direction: if you can estimate the reference's own sensitivity and specificity, you can adjust the two-by-two table for its misclassification.
- No reference at all calls for latent class analysisDohoo et al. 2012, which treats true disease status as an unobserved (latent) variable and estimates each test's sensitivity and specificity, along with prevalence, from the pattern of agreement across two or more tests applied to the same subjects.
- The load-bearing assumption is conditional independence: the tests must err independently given true status. When two tests share a failure mode, such as both being fooled by the same cross-reacting antibody, that assumption breaks and a conditional-dependence model is required.
The pitfall is that a latent-class estimate can look precise while resting entirely on the independence assumption and the number of tests: with only two tests the model is not identified without extra constraints, so more tests, or the same tests across populations with different prevalence, are what make the estimate trustworthy.
Thresholds and cut points
Turning a continuous risk or measurement into a yes-or-no action is convenient and lossy, which is the tension a threshold, or cut-point, has to resolve, and you confront it whenever a decision forces a continuous quantity into a binary call.
A cutoff treats a patient just below and just above as categorically different when they are nearly identical, and where you set it slides the test's sensitivity against its specificityYouden 1950. That placement encodes a value judgment about the costs of acting versus waiting, often a hidden one, which the absolute benefit at that risk level makes concrete.
Risk calculators and prediction tools
Reach for a risk calculator when you want a prediction model packaged for bedside use, turning a predicted risk into something a clinician can apply to the individual patient in front of them. The convenience comes with a hidden liability: a calculator carries its development population with it. Applied to patients who differ from the cohort it was built on, even a well-constructed tool can be systematically off, a failure of calibration rather than of the underlying logic. That is why external validation and recalibration matter moreCollins et al. 2015 than the elegance of the original model. A tool reproduced from a landmark cohort can look authoritative and still mislead in a population with a different baseline risk, so the question is never whether the model is clever but whether its predictions hold in the people you actually treat.
The discipline, then, is to validate before you act. Confirm that predicted and observed risks agree in your own setting, recalibrate where they diverge, and only then convert the output into a decision through a cut point. A calculator is a delivery mechanism for a model, and it inherits every limitation of the model and the data behind it.
Decision-curve analysis
Reach for decision-curve analysisVickers & Elkin 2006 when you want to judge whether acting on a test or model does more good than harm across the range of thresholds a clinician might reasonably hold, rather than relying on accuracy metrics that ignore the consequences of acting. When the question is the overall method, decision-curve analysis is the tool; when it narrows to the quantity that captures clinical utility weighted by the chosen threshold, you are working with net benefit. A test is only worth using if its benefits outweigh its harms, and accuracy alone cannot tell you that, because it counts true and false positives without weighing their very different consequences.
Net benefit closes that gap. At a threshold probability \(p_t\) it is \((\text{TP} - \text{FP} \times p_t/(1 - p_t)) / N\), where the odds \(p_t/(1 - p_t)\) set the exchange rate, weighting each false positive against a true positive by how reluctant a clinician at that threshold is to intervene. Plotting net benefit across the plausible threshold range shows whether the model beats the default strategies of treating everyone or no one, and over what range it does so. Adding cost to this same net-benefit logic is what cost-effectiveness does.
How a cost-effectiveness analysis fits together
The terms in this section are steps in one workflow, not a glossary. An economic evaluation asks whether an intervention is worth its costDrummond et al. 2015, and the pieces assemble into that answer in a fixed order.
- Frame the decision. Fix the comparators, the population, the time horizon, and the perspective whose costs and benefits count; the reference case is what makes analyses comparable.
- Build the model. Project lifetime costs and effects with the right decision-analytic model (a decision tree, a Markov model, or an estimate alongside a trial), discounted over a lifetime horizon.
- Populate it. Attach costs and QALYs to each state, read from the chosen perspective.
- Compute the ICER. The incremental cost-effectiveness ratio is the difference in cost over the difference in QALYs between an option and its comparator, judged against a willingness-to-pay threshold.
- Probe the uncertainty. Show how fragile the ratio is with a probabilistic sensitivity analysis and value-of-information analysis, not just a base-case number.
- Separate affordability from value. A budget-impact analysis and the HTA verdict decide what a payer can afford and will cover, which cost-effectiveness alone does not settle.
The through-line is the ICER: everything above either feeds it (the model, the costs, the QALYs), stress-tests it (uncertainty and value of information), or places it in context (budget impact, HTA). The nodes below unpack each step in that order.
Cost-effectiveness and the ICER
Turn to cost-effectiveness and the ICER when a decision must weigh resources against outcomes, putting cost and benefit on one page. Which framing fits depends on how the benefit is measured.
- The ICER (incremental cost-effectiveness ratio) is the headline number: the extra cost per extra unit of effect, incremental cost divided by incremental effect of one option over the next, judged against a willingness-to-pay threshold.
- Cost-benefit analysis values both costs and benefits in money.
- Cost-minimization analysis compares costs only, and applies when the effects are genuinely identical so only costs differ.
- Cost-utility analysis measures effects in quality-adjusted life years, which makes different conditions comparable.
- Net monetary benefit values each option at a willingness thresholdStinnett & Mullahy 1998; linear in its inputs, it handles dominance and extended dominance cleanly where ratios behave badly.
- The willingness-to-pay threshold is the maximum payable per unit of benefit, the bar the ICER must clear.
To apply. The ICER is always incremental, the cost difference over the effect difference between an option and the next-best alternative, never against a fixed baseline. With two options it is a single division. With several, the calculation is a ranking: order the options from least to most effective and drop any that are strongly dominated (costing more while delivering less than another); compute each survivor's ICER against the option just below it, then drop any that are extended-dominated (its own ICER is higher than that of a more effective option, so a blend of two other options, part-funding the cheaper and part the more effective, buys more health per dollar than this one does, even though nothing simply costs more while doing less); recompute the ratios along what remains, the efficiency frontier, and adopt the most effective option whose frontier ICER still falls below the willingness-to-pay threshold. Because the ICER is a ratio, a negative value is ambiguous (an option can be dominant or dominated), so for ranking and uncertainty the net monetary benefit (threshold times effect minus cost, highest value winning) is the cleaner equivalent; on the cost-effectiveness plane, the ICER is simply the slope of the line from one option to the next.
Costs and QALYs
To populate the model you attach a cost and an outcome weight to each health state. Both are sourced inputs whose provenance flows into every downstream conclusion, not constants to assume.
Costing the resources. The price weight can come from charges (the list prices a hospital bills, which overstate true cost and need a cost-to-charge ratio, a hospital-specific fraction that scales billed charges back down to estimated cost, to correct), from reimbursement, or from a true cost-accounting system; costs sort into direct medical, direct non-medical, and indirect productivity losses.
- Gross (top-down) costing values a whole episode with one aggregate weight such as a DRG payment (a diagnosis-related group is a fixed bundle price Medicare pays for an admission of a given type, regardless of the individual resources used); reach for it when itemized detail is unavailable or unnecessary.
- Micro-costing (bottom-up) counts each resource used times its unit price; reach for it when the intervention's cost difference lives in specific resources.
- Productivity loss uses the human-capital approach (all foregone lifetime earnings) or the friction-cost approach (only until a worker is replaced); the two can differ sharply, and a societal perspective also weighs caregiver time and the future medical costs of added life-years.
The outcome: QALYs. The quality-adjusted life year (QALY) is the common currency that lets a year with diabetes and a year after a stroke be compared: time in a state times a utility weight anchored at 0 (death) and 1 (full health), so a year at utility 0.7 is 0.7 QALYs. Utilities come from a preference-based instrument (the EQ-5D) or from time-trade-off or standard-gamble elicitation, and because the instrument and source populationDohoo et al. 2012 move the result, a utility is a named, sourced input. The QALY is the denominator the ICER is built on.
When costs come from real-world claims rather than a tidy model, they need the messier estimation the real-world cost methods supply.
Perspective and the reference case
Reach for perspective and the reference case when you must decide whose costs and benefits count, a choice that changes the answer.
- The societal perspective counts all costs to society, adding patient time, caregiving, and lost productivity to medical costs, and can flip the verdict for conditions whose burden falls outside the clinic, whereas a healthcare-sector perspective counts only medical costs to the payer.
- The reference case recommended by the Second Panel on Cost-Effectiveness in Health and MedicineSanders et al. 2016 sets standardized analysis conventions, a fixed set of methods reported alongside any analysis so results are comparable, together with an impact inventory listing every cost included or excluded.
- The Consolidated Health Economic Evaluation Reporting Standards (CHEERS) are the reporting checklist for economicsHusereau et al. 2022, the economic-evaluation member of the reporting-standards family.
- Opportunity cost is the value of foregone alternatives, capturing that every dollar spent is health some other patient could have had.
The pitfall is that the perspective choice can flip the verdict, so the reference case and impact inventory exist to keep results comparable and to avoid double-counting productivity. The trap: lost earnings from illness can be entered once on the cost side (a productivity cost) and again inside the quality-adjusted life year (QALY) weights if the people who rated the health state already priced in the effect of being unable to work, so counting both charges the same loss twice. This is one reason those QALY weights by rule reflect community rather than patient preferences. The Panel's harder calls live here too: whether to count the future medical costs of added life-years, which it now includes, and how widely to count whose health matters, out to caregiver QALYs and the distributional question a single incremental cost-effectiveness ratio (ICER) hides.
Building the model
Lifetime costs and QALYs are rarely observed directly, so a decision-analytic model projects them, matching the simplest structure that faithfully represents the decision.
Choose the structure by how the disease behaves.
- A decision tree for a one-off branching decision, clean for an acute choice but clumsy once events recur.
- A Markov (state-transition) model for recurring health states over cycles, the standard for chronic disease, advancing a cohort through a transition matrix each cycle.
- A partitioned-survival model when states are read off progression-free and overall survival curves, as in oncology.
- A dynamic transmission model for infection, where treating one person changes others' risk through herd immunityDohoo et al. 2012 that a fixed-risk cohort cannot capture.
Across all of these, future costs and effects are discounted below present ones over a horizon long enough to capture the consequences, usually a lifetime; a published rate must first be converted to a per-cycle probability (with background mortality from life tables) before it runs through the transition matrix.
Cycle length and the half-cycle correction. The cycle must be short enough that no important event is missed (often monthly, not annual). Because transitions happen throughout a cycle rather than at its boundary, tallying membership only at boundaries over- or under-counts costs and QALYs; the half-cycle correction (or Simpson's rule, or a life-table integration) fixes it, and the error grows with cycle length.
Validation and calibration. Verification confirms the implementation does the intended math; calibration tunes unobservable parameters until outputs match targets, carrying that uncertainty forward. Validation runs in layers, face, internal (reproduces its own inputs), external (matches independent data), and predictive (forecasts later events), with cross-validation against other published models; stopping at internal validityDohoo et al. 2012 is the trap.
Or estimate alongside a trial. When a randomized trial collects patient-level costs and outcomes, cost-effectiveness can be read directly from it by net-benefit regression (each patient's cost and effect become one net-benefit outcome at a chosen threshold, regressed on the arm, so covariate adjustment and subgroups come free). Watch the near-universal complications, right-skewed and censored costs, and correlated cost and effect, and the short horizon: a within-trial result is often only the seed for a lifetime model.
Handling uncertainty and value of information
A single ICER hides how fragile it is; handling uncertainty shows how robust the conclusion really is, and it starts by naming which kind of uncertainty you face, because each calls for a different tool.
- Parameter (second-order) uncertainty is doubt about an input's true value from finite data, propagated by probabilistic sensitivity analysis.
- Stochastic (first-order) uncertainty is random variation between identical individuals, averaged out by a microsimulation.
- Heterogeneity is variation explained by patient characteristics, handled by subgroups, not a distribution.
- Structural uncertainty is doubt about the model's own form, which states, which functional form, how to extrapolate, often larger than parameter uncertainty yet routinely ignored, so it needs scenario and structural sensitivity analysis.
Quantify it. One-way and tornado analyses vary each input across its range to find which move the conclusion. A probabilistic sensitivity analysis (PSA) instead draws every uncertain parameter from a distribution and runs the model thousands of times, producing a cloud of cost-effect pairs on the cost-effectiveness plane; the cost-effectiveness acceptability curve (CEAC)Briggs et al. 2006 summarizes that cloud as the probability each option is the best buy at each willingness-to-pay threshold. Because the ICER resists a tidy confidence interval (it is a ratio, and its denominator, the difference in effect, can be near zero or straddle zero across the uncertainty, which sends the ratio off to plus or minus infinity and makes an ordinary interval meaningless), results are restated as net monetary or net health benefit, which stay well-behaved because they are differences rather than ratios. When an interval on the ratio itself is still wanted, Fieller's theorem (an exact interval for a ratio) or a bootstrap (resampling to trace the ratio's spread) can supply one. Reporting the acceptability curve, not just a base-case ratio, is what a payer expects.
Then price the uncertainty: value of information. When decision uncertainty remains, the expected value of perfect information (EVPI) is the upper bound on what any further research could be worth, so an EVPI below the cost of a study means the evidence is already good enough to act on. The expected value of sample information (EVSI) values a specific study of a given design and size, and partial EVPI isolates which parameters drive the uncertainty and are worth measuring better, reframing the question from whether something is cost-effective to whether it is worth learning more before deciding.
Budget impact and health technology assessment
Cost-effectiveness rarely decides anything on its own; two further questions turn the analysis into a coverage verdict, and they answer to different decision-makers.
Budget impact: can it be afforded? Where cost-effectiveness asks about value per patient, a budget-impact analysis projects the total cost to a specific budget holderSullivan et al. 2014, a health plan or a national system, of adopting the intervention across the eligible population over a near-term horizon (usually one to five years) under realistic uptake. It is undiscounted and bounded to the adoption window, because the question is cash-flow burden, not lifetime value. The point it exists to expose: a therapy can be cost-effective per patient and still break a budget at population scale, since affordability turns on the size of the eligible population and the speed of uptake, not the cost per unit of benefit. Payers ask for both.
HTA and value frameworks: is it worth covering? A health technology assessment body weighs cost-effectiveness against clinical benefit, budget impact, and equity to reach a decision, and which framework applies depends on the system: the UK's NICE pairs the analysis with an explicit cost-per-QALY threshold, while the US Institute for Clinical and Economic Review publishes value assessments that anchor negotiations without a binding rule. The catch is that the willingness-to-pay threshold, which looked like a technical input, becomes a policy lever here, so different systems reach different verdicts from the same analysis. This same machinery now sits behind Medicare drug-price negotiation under the Inflation Reduction Act, the backdrop of the Part D trace, and it is where the pathway's analytic work finally meets a real choice about what to cover and at what price.
Real-world cost and HTA methods
Reach for real-world cost and HTA methods when real-world cost data break the assumptions of ordinary linear regression, with many patients incurring zero cost, the rest forming a long right tail driven by a few high utilizers, and variance growing with the mean.
- The two-part model handles zero-inflated costsMihaylova et al. 2011, separating whether any cost was incurred from how much conditional on some being incurred; a generalized linear model with a log link and a gamma family is another robust route.
- Winsorization and trimming of cost outliers cap or drop the tail when extreme but real outliers threaten to dominate the mean.
- Per-member-per-month costing (PMPM/PPPM) summarizes population spend, normalizing for enrollment time across people with different follow-up.
- Survival extrapolation projects beyond a trial that endsLatimer 2013 before the lifetime horizon, fitting parametric or flexible models to the observed Kaplan-Meier curve and extending it, with the chosen functional form itself a major source of uncertainty.
- Multi-criteria decision analysis (MCDA) handles value with multiple dimensions, weighting criteria such as equity and severity explicitly when a single ratio cannot capture them.
- Expected value of partial perfect information (EVPPI) prioritizes which uncertainty matters, pricing the resolution of specific parameters to tell you which uncertainty further research should target.
Health technology assessment then needs costs and outcomes projected over a lifetime, but trials end early, so the chosen extrapolation feeds directly into the decision-analytic models and the verdict that rests on them.
At the top of the pathway is the sentence a clinician actually reads in a guideline. This rung is about how that sentence is phrased, how firmly, and whether the confidence in the wording is matched by the evidence underneath it.
Evidence-to-decision
Moving from a body of evidence to a recommendation is never automatic, and an evidence-to-decision framework is what you apply whenever a body of evidenceAlonso-Coello et al. 2016 must become a recommendation. Its purpose is to make the step explicit, weighing benefits and harms alongside the values, feasibility, equity, and cost that the evidence itself cannot settle. These additional inputs are legitimate parts of the judgment, but they belong in the open rather than buried beneath a confident-sounding conclusion.
The framework's value is precisely that it forces each of those considerations to be named and reasoned through, so that a reader can see why a panel landed where it did. The point worth watching is that two panels reading identical evidence can legitimately land on different recommendations, because they weigh benefits, harms, values, feasibility, equity, and cost differently. That divergence is not a failure of rigor; it is the visible signature of value-laden choices that should be made transparent rather than concealed, so the reasoning can be inspected and, where one disagrees, contested.
This is the second half of GRADE (Grading of Recommendations Assessment, Development and Evaluation), the standard system for going from evidence to guidance. Its first half, done during synthesis, rates how certain the evidence is (high to very low) for each outcome; its second half, here, begins from that certainty rating and weighs it against the values, costs, and feasibility that the evidence cannot settle to arrive at a recommendation.
Strength of recommendation
Read the strength of a recommendation as the guideline body's signal of how firmly it is willing to speak. Different systems encode that signal differently: the ACC/AHA scheme pairs a class of recommendation with a level of evidence, while the Grading of Recommendations Assessment, Development and Evaluation (GRADE) system splits recommendations into strong and conditionalGuyatt et al. 2008. In each case the strength is a claim about confidence, and the question it resolves is how much weight a clinician should place on the recommendation given the evidence behind it.
The principle that ties these schemes together is that strength should track the certainty of the evidence: a strong or class I recommendation belongs on robust data, a conditional one on weaker data. The pitfall worth noticing is exactly the place where the two come apart. A confident recommendation resting on thin evidence is an evidence-recommendation gap, a signal to look past the label and ask what the underlying certainty actually supports before treating the recommendation as settled.
Consensus methods (Delphi, nominal group)
When the evidence underdetermines the answer but a panel must still converge on a recommendation or a core outcome set, you reach for formal consensus methods, the machinery that lets a group settle a question without the loudest voice deciding.
- The Delphi method reaches convergence through anonymous iterative roundsDalkey & Helmer 1963: a selected expert panel answers in repeated rounds, each member sees a statistical summary of the group between rounds and revises, and opinion converges through reasoned feedback rather than face-to-face pressure. Anonymity blunts dominance and bandwagon effects, consensus is defined in advance as a threshold such as 70% or 80% agreement, and the stopping rule is stability, the point at which further rounds stop shifting the distribution.
- The nominal group technique is structured in-person ranking, reaching the same convergence through silent ranking and then discussion, and the RAND/UCLA Appropriateness Method blends Delphi rating with a face-to-face round.
The same machinery defines a core outcome set and built the reporting checklists themselves. The catch to watch is that consensus is not evidence: a firmly worded recommendation resting on expert agreement rather than data is exactly the kind of evidence-recommendation gap an honest guideline should flag and label as consensus-based.
The evidence–recommendation gap
The most useful thing a statistically literate reader can do with any guideline is to ask which rung its confidence actually rests on, because this is where critical appraisal of a recommendation earns its keep. A recommendation can be worded firmly while leaning on an extrapolated threshold, a single trial, a measurement method the clinic does not reproduce, or nothing firmer than the expert consensus a panel reached. The evidence-recommendation gap is the distance between that phrasing and the support beneath it.
What to watch for is precisely that firm phrasing can outrun its foundation: confident language attaches just as easily to an extrapolated threshold, a lone study, an unreproducible measurement, or mere expert consensus as it does to a robust randomized result. Reading a guideline well therefore means looking past the strength of the wording to the strength of the evidence, and treating the two as separate questions. Where the recommendation is firm but the support is thin, the gap is where appraisal does its real work, and naming it honestly is what distinguishes a defensible guideline from one that merely sounds authoritative.
Reporting standards
You follow a reporting standard whenever a study's methods must be made auditable enough for a reader to judge them, and which checklist you reach for depends on what you are reporting. Each names the items that let a reader reconstruct and appraise the methods that produced the finding.
- A randomized controlled trial follows the Consolidated Standards of Reporting Trials (CONSORT).
- An observational study follows the Strengthening the Reporting of Observational Studies in Epidemiology (STROBEDohoo et al. 2012).
- A systematic reviewDohoo et al. 2012 follows PRISMA.
- A prediction model study follows TRIPOD.
These checklists are not bureaucracy; they are the difference between a result you can interrogate and one you have to take on faith, which is the pitfall to watch: a study reported without the relevant details cannot be judged on its merits, only trusted. The same family reaches back to the design stage, where the Standard Protocol Items: Recommendations for Interventional Trials (SPIRIT) govern the trial protocol before any result exists, so the discipline of auditable reporting begins before the first data are collected.
These moves do not belong to any single rung. They are the checks that ask, at every stage, whether a result would survive being pushed on: a hidden bias, a fragile assumption, a single influential data point.
Which move you reach for follows the threat most likely to be challenged, and each links straight to its node. When the question is whether an observed association should be read as causal at all, the Bradford Hill viewpoints are the lens a reviewer weighs it through. Before anything else, name the threats: the catalog of study biases, mapped to the rung each one enters, tells you what you are defending against. When the worry is unmeasured confounding, bias quantification puts a number on how strong a hidden confounder would have to be (E-values, Rosenbaum bounds), and negative controls and empirical calibration test for confounding you cannot see by checking associations that should be null. When a modeling or identification assumption is contestable, a pre-specified sensitivity analysis varies it on purpose and reports what moves. When a single observation might be driving the result, leave-one-out and specification curves show whether the finding survives dropping it or rerunning every defensible analytic choice. When the causal claim itself is in doubt, placebo and falsification tests look for an effect where none should exist. And when you need to know how an estimator or design behaves under a known truth, Monte Carlo simulation generates data you control and checks whether the method recovers it.
Bradford Hill viewpoints
When an association is observed and randomization was impossible, the question a reviewer presses is whether it should be read as causal, and the Bradford Hill viewpointsHill 1965 are the checklist that argument is usually run through. They are best understood as the appraisal layer that sits on top of this site's counterfactual and causal-diagram machinery, not a replacement for it: a set of considerations that make a causal reading more or less credible once confounding and bias have already been addressed. Reach for them whenever you are defending, or attacking, the causal interpretation of an observational finding.
- Temporality is the one genuinely necessary viewpoint: the cause must precede the effect, and a design that cannot fix that order cannot claim causation at all.
- Strength of association makes the case harder to dismiss, since a large effect is harder for an unmeasured confounder (a common cause of exposure and outcome that was never recorded, so it could not be adjusted for) to manufacture, though a small effect can still be causal.
- Biological gradient (dose-response) strengthens it further when risk rises with exposure and a monotone relationship is what the biology would predict.
- Consistency is the association reappearing across populations, designs, and investigators, which is exactly what a synthesis or a replication checks.
- Plausibility and coherence ask for a credible mechanism and a fit with what else is known, both bounded by the biology understood at the time.
- Experiment is the closest viewpoint to the counterfactual ideal (comparing what happened under the exposure with what would have happened to the same people without it): a change in exposure that changes the outcome.
- Specificity and analogy are the weakest two, one exposure mapping to one outcome and a parallel to an accepted causal relationship, and neither should carry much weight on its own.
The pitfall is treating these as a scoring sheet where enough ticked boxes prove causation. Hill offered them as viewpoints to weigh, not criteria to count, and only temporality is required; strength, specificity, and analogy in particular fail often enough that a mechanical tally invites both false positives and false negatives. Used well, they organize the argument that an explicit estimand, a DAG, and quantified bias then have to make rigorous.
Study biases, by rung
Bias is not a single defect but a family of distinct threatsDelgado-Rodríguez & Llorca 2004, each entering the research process at a particular rung, and the practical move is to map each named bias to where it arises so a design can anticipate the specific threat it faces. If you want the overall map, study biases organized by rung is that catalog; selection biases arise from who ends up in the analysis, and information biases arise from how variables are recorded.
- Sampling bias is distortion in who got sampled or enrolled.
- Selection bias is distorted entry into the study, the broader category.
- Berkson's bias is selection among hospitalized patients.
- Over-adjustment is conditioning on a common effect (a collider): adjusting for a variable that both the exposure and the outcome influence opens a spurious path between them, so controlling for it manufactures an association rather than removing one, adding bias while trying to remove it.
- The healthy-worker effect is employed groups appearing healthier than the general population.
- Survivorship bias is observing only the survivors.
- Attrition bias is differential loss to follow-up.
- Nonresponse bias arises when nonresponders differ systematically from responders.
- Confounding is a common cause of both exposure and outcome.
- Confounding by indication is its clinical archetype, treatment chosen by prognosis.
- Detection bias is unequal outcome ascertainment between groups.
- Recall bias is inaccurate recall of exposure.
- Interviewer bias is an interviewer shaping responses.
- Lead-time bias is earlier detection inflating apparent survival.
- Length-time bias is slow-progressing cases being preferentially detected.
- Overdiagnosis is detecting disease that would never have caused harm.
- Publication bias is selective reporting of positive results, the risk when single studies are pooled into a body of evidence.
Guarding against one member leaves the others, which is the pitfall to watch, including over-adjustment that adds bias while trying to remove it.
To prevent it. Match the fix to the family. Selection biases are designed out: draw controls from the same source population as the cases (against Berkson’s bias), assemble an inception cohort that includes units before they can drop out (against survivorship bias), minimize and model loss to follow-up with inverse-probability-of-censoring weighting (against attrition bias), and use a probability sample with follow-up of nonresponders (against sampling bias and nonresponse bias); the healthy-worker effect calls for an active-worker referent rather than the general population. Confounding by indication needs an active-comparator new-user design with propensity-score adjustment, or an instrumental variable when the indication is unmeasured. Information biases are controlled by blinding outcome assessors and applying identical, protocol-defined ascertainment (against detection bias and interviewer bias) and by using prospectively recorded or validated exposures rather than retrospective self-report (against recall bias). Screening biases need the right endpoint: compare mortality from a common zero rather than survival timed from diagnosis (against lead-time bias), and lean on randomized screening trials with mortality outcomes (against length-time bias).
Sensitivity analysis
A result is only as trustworthy as it is robust to the choices behind it, which is why you pre-specify a sensitivity analysisThabane et al. 2013 that deliberately varies the assumptions most likely to be challenged and reports what happens to the estimate. Designing these probes in advance is far more credible than running the same analyses only after a reviewer asks, and the contrast is the thing to watch: sensitivity analyses built into the plan are a strength, while ones bolted on afterward are a tell of weaker credibility, because a probe chosen after seeing the result is gameable: the analyst can keep the checks that happen to support the headline finding and quietly drop the ones that undercut it.
The choice of which assumption to stress depends on what is most contestable in the design. Where the vulnerable assumption is unmeasured confounding specifically, you move to bias quantification, which puts a number on how strong a hidden confounder would have to be to overturn the result rather than merely asserting that it is implausible. In that sense the general practice of pre-specified sensitivity analysis sets the frame, and the more specialized quantitative tools fill in the particular threat, so that a finding is shown to survive the challenges most likely to be raised against it instead of being defended only after the fact.
Bias quantification
Rather than asserting that the unmeasured confounding a causal estimate leaves behind is unlikely, you quantify it, treating bias quantificationVanderWeele & Ding 2017 as one pre-specified sensitivity analysis among others that puts a number on how strong hidden confounding would have to be to explain away the result.
- The E-value is the strength needed to explain away an estimate, asking how strong a hidden confounder would have to be to nullify the finding, and for a risk ratio (RR) it can be computed with a simple formula rather than by simulation. Read it concretely: an E-value of 3 means an unmeasured confounder would have to be associated with both the exposure and the outcome by a risk ratio of at least 3, above and beyond the covariates already adjusted for, to move the result to no effect. A large required strength means the finding is robust; a value near 1 means a weak confounder could overturn it.
- Rosenbaum bounds quantify hidden bias in a matched design, the analogous bound on how much hidden bias the matching could conceal.
The interpretive caution is that an E-value or Rosenbaum bound only describes the strength of confounding required, so a result that even a modest confounder could erase is correspondingly fragile, while one that survives a large E-value is sturdier. After putting numbers on plausible bias, you can go further and detect and correct residual confounding empirically with negative controls and empirical calibration.
Negative controls and empirical calibration
When you need to detect hidden residual confounding empirically, a negative controlLipsitch et al. 2010 is a variable whose true effect is known to be null, so any nonzero estimate recovered there reads off systematic error directly. Because the truth is null in these cases, whatever estimate appears is a direct readout of residual confounding, complementing the bounds set in bias quantification.
- Checking confounding on the exposure side uses a negative control exposure, which shares the confounding structure of the real exposure but has no plausible causal path to the outcome.
- Checking confounding on the outcome side uses a negative control outcome, which shares the confounding structure of the real outcome but cannot plausibly be caused by the exposure.
- Many negative controls available calls for empirical calibration. The intuition: analyses that should return exactly zero rarely do, and how far they scatter around zero reveals the study's systematic error. So it runs the same analysis across dozens of them, fits the spread of those should-be-null estimates, and uses that empirical null (the observed distribution of results when the truth is zero) as a yardstick to recalibrate the p-value and confidence interval, so the width reflects observed systematic error rather than sampling variance alone.
- Quantifying systematic error more broadly pairs this with quantitative bias analysis.
The caution to watch is that controls are only informative when they genuinely share the confounding pathways of the target question, making their choice a substantive, domain-driven exercise rather than a mechanical one.
Placebo and falsification tests
A good way to test a design is to look for an effect where none should exist, which is the logic of placebo and falsification testsPrasad & Jena 2013. You probe a design by examining a pre-treatment period, an outcome the intervention cannot plausibly affect, or an untreated group, places where the true effect must be zero. If the design nonetheless finds an effect there, that is the signal to watch: something is wrong with the design itself, because a result has appeared where there should be none.
The orientation that makes these tests useful is adversarial. Real falsification tests try to break the result rather than to confirm it, so they belong among the pre-specified sensitivity analyses run in good faith to stress a finding before others do. In that role they sit alongside leave-one-out checks, which probe whether a single unit is carrying the result, and bias quantification, which bounds the confounding needed to undo it. Together these defend a finding not by accumulating supportive evidence but by surviving deliberate attempts to expose it as artifact, and a design that passes a placebo or falsification test has earned more trust than one that was never challenged.
Leave-one-out and specification curves
If dropping a single site, year, or cohort overturns a finding, the result rests on that one unit rather than on the effect, and leave-one-out re-estimation is how you expose that fragility, re-fitting the analysis with each unit removed in turn. For the overall idea of probing specification robustness this leave-one-out and specification-curve approach is the tool. When the concern is instead results across many model choices, specification-curve analysisSimonsohn et al. 2020 does the analogous job across the many defensible modeling decisions, showing whether the conclusion holds broadly or only along one particular path.
You reach for these checks whenever you want to know whether a finding survives both dropping single units and varying the modeling choices, and they pair naturally with falsification tests as a form of robustness sensitivity analysis. The pitfall to watch is that running every specification revives the multiplicity problem: test hundreds of specifications and some will clear significance by chance alone, just as flipping enough coins eventually yields a run of heads, so the curve is something to read the conclusion against, gauging how consistently the effect appears across reasonable specifications, rather than treating it as a formal hypothesis test that the curve was never built to be.
Monte Carlo simulation
When closed-form theory does not fit the design, Monte Carlo simulationMorris et al. 2019 lets you generate data under a known process, run the planned analysis, and watch how it behaves over many replicates. You turn to it precisely when no analytic formula matches the estimator or the design, because simulation can answer empirically what theory cannot deliver in closed form: how much bias an estimator carries, whether its confidence intervals cover at the stated rate, and how much sample size a non-standard design actually needs.
The same machinery, run over input distributions rather than over observed data, becomes the probabilistic sensitivity analysis behind a cost-effectiveness result, propagating uncertainty in the inputs through to the conclusion. The caution to keep in view is that the results are only as trustworthy as the assumed data-generating process: the simulation tells you how the analysis behaves under the conditions you specified, so its conclusions hold under those simulated conditions rather than universally. Designing the generating process to reflect the real problem honestly, including the features most likely to trouble the estimator, is therefore as important as the replication itself.
A study is more than its analysis. Around the numbers sit the rules that make it legitimate: ethics review and consent, the written plans agreed in advance, and the privacy and record-keeping that let anyone trust what was done.
Research ethics and the IRB
No analysis is legitimate if the study should not have been run, which is the province of research ethics and the institutional review board (IRB)Emanuel et al. 2000, the body whose approval you secure before a study starts.
- The Belmont principles of respect for persons, beneficence, and justice are the foundational ethics of human-subjects research, operationalized in the US Common Rule (45 CFR 46); the Declaration of Helsinki is a separate, earlier international statement of research ethics.
- Clinical equipoise is the genuine uncertainty in the expert community about which arm is better, the specific license to randomize.
- Informed consent is a participant's voluntary agreement: a participant must understand the study, its risks, and their freedom to refuse or withdraw, with extra protection for vulnerable groups such as children, prisoners, and the cognitively impaired.
- The institutional review board reviews and approves studies, weighing risks against benefits and able to halt or modify a protocol.
The pitfall to watch is exactly here: assigning a patient to a treatment you believe is worse is unethical however clean the design, so the board's approval is only the first gate, with good clinical practice and a registered protocol carrying the same duties through the study.
Good clinical practice
Good Clinical PracticeICH E6 2016 (International Council for Harmonisation (ICH) E6) is the operational standard you follow in any trial whose data and participants must be trustworthy. It sets defined responsibilities for sponsor and investigator, requires a written protocol that is followed and whose deviations are documented, mandates monitoring and source-data verification so that a recorded value traces back to the medical record, and demands an audit trail for every change. Each of these elements exists so that the conduct of the study leaves a complete, inspectable record rather than an unverifiable summary.
The point is not bureaucracy but credibility: the standard exists so that a regulator or reviewer can reconstruct exactly what happened. That is also the pitfall to watch, because an undocumented change or a value that cannot be traced to its source undermines the credibility of the whole record, not just the single datum in question. This is the same provenance discipline that data standards and the CDISC pipeline enforce on the data themselves, applied to the conduct of the trial: a result is only as defensible as the documented chain that produced it, and Good Clinical Practice is what keeps that chain intact from protocol to final dataset.
Data privacy and security
Analyzing health data carries a duty to the people in it, the substance of data privacy and securityEl Emam et al. 2011, and which rule or method applies depends on the jurisdiction and the route to sharing.
- HIPAA is the US health privacy law, governing identifiable health information.
- The General Data Protection Regulation (GDPR) is the EU data protection law, imposing a stricter consent-and-purpose regime.
- The Safe Harbor method de-identifies by removing identifiers, stripping eighteen specified identifiers.
- Expert determination de-identifies by statistical opinion, having a statistician certify that the re-identification risk is very small; a limited dataset retains some dates and geography under a data use agreement.
- Synthetic data generates artificial substitute records rather than releasing real ones, fitting a generative model to the real records and drawing new ones that reproduce the joint distribution; a generative adversarial network (GAN), a diffusion model, or flow matching can do the fitting.
The cautions to watch are twofold: de-identification is never zero-risk because rich quasi-identifiers (fields that are not names but still narrow a person down, such as ZIP code plus birth date plus sex, which together single out most individuals) can be linked back to outside records, so the standard is reasonable risk rather than impossibility, and synthetic data inherits whatever bias and missingness the source carried and needs its own privacy and fidelity audits before it is trusted.
Regulatory pathways and registration
A study that informs a regulated decision lives inside a regulatory frame, and the practical task is to determine which regulatory pathway and registrationDe Angelis et al. 2004 applies before the first patient. Either way the trial must be registered on ClinicalTrials.gov before enrollment, and the FDAAA law requires summary results to be posted whether or not they are favorable, the public-record antidote to publication bias.
- An investigational drug usually needs an FDA investigational new drug (IND) application before the trial begins.
- An investigational device needs an investigational device exemption (IDE).
Registration does more than satisfy a rule: it timestamps the pre-specified endpoints, so a quietly swapped primary outcome becomes visible, and it forces results into the open regardless of how they came out. That is also the pitfall to watch, because skipping registration lets a swapped primary outcome or a buried null go unseen. Knowing which frame applies, securing the IND or IDE where needed, and registering before enrollment is therefore part of designing the study rather than an afterthought tacked on once the data are in.
The statistical analysis plan
The statistical analysis plan (SAP)Gamble et al. 2017 is the document you write before the data are unblinded to pre-commit exactly how the primary question will be answered. It fixes the estimand and its intercurrent-event strategy, the analysis population, the primary test and model, the handling of missing data, the multiplicity control across endpoints, the pre-specified subgroups, and the sensitivity analyses. You reach for it whenever a confirmatory analysis must be locked down in advance, because writing these choices down before seeing the data is precisely what makes a confirmatory analysis confirmatory.
The pitfall to watch follows directly: anything decided after seeing the data is exploratory and must be labelled so, since the credibility of a confirmatory result rests entirely on its having been specified before the data could influence the choice. A good SAP is detailed enough that two statisticians who run it independently land on the same result, which is the working test of whether it has truly removed analyst discretion. It also fixes the table, figure, and listing shells that the programming then fills, so the form of the reported output is settled before any number exists to shape it.
The study protocol (SPIRIT)
The study protocolChan et al. 2013 is the master plan every other document hangs from: the objectives, the eligibility criteria, the intervention and comparator, the outcomes, the sample-size justification, the analysis approach, the ethics and consent, and the dissemination plan. You draft it as that master plan whenever a study is being designed, so that every later document and decision can be checked against a single pre-committed source. The Standard Protocol Items: Recommendations for Interventional Trials (SPIRIT) is its reporting standard, the protocol counterpart to the CONSORT checklist for the finished trial, and when the situation calls for a reporting standard for protocols specifically, SPIRIT is the one you follow.
The discipline to watch is timing: writing and registering the protocol before enrollment is what lets a reviewer confirm that the published paper answers the question the study set out to answer, rather than the one the data happened to support. A protocol fixed only after the data are in cannot perform that function, because there is no longer a pre-data record to hold the final paper against. Registering it in advance therefore turns the protocol from a planning convenience into a verifiable commitment.
Data management and reproducibility
Between collection and analysis sits data managementPeng 2011, and its discipline decides whether a result can be trusted or even reproduced. For the overall practice, data management and reproducibility name this whole stage. On the data side it means a case report form (CRF), or its electronic version, designed to capture clean values, edit checks and queries that catch errors at entry, and a data dictionary. When the specific step is freezing the dataset before analysis, that is the database lock, a dated point after which no value changes silently, the clean source that cohort assembly then transforms into an analysis table.
On the analysis side the principle is that the whole analysis runs from code, under version control, with every dependency pinned so that the numbers regenerate from the raw data with one command (in practice: Git for version control, renv or conda to pin package versions, a fixed random seed for any stochastic step, and a literate workflow such as Quarto or targets that interleaves code with prose so the report and its computations stay in sync), and, increasingly, shared code and data so an outsider can rerun it. You build this discipline whenever a result must be defensible, because the pitfall to watch is that without it a finding may not be reproducible at all, which is what separates a genuine finding from an anecdote. In regulated trial analysis this same reproducibility is enforced by independent double-programming.
Statistical programming: TFLs and double-programming QC
The analysis is delivered as programmed outputs, and statistical programming and quality controlICH E6 2016 are what make those outputs defensible.
- The deliverable tables, figures, and listings (TFLs) are the reported output, whose shells are pre-specified in the statistical analysis plan so each layout is fixed before the data are seen; a table summarizes, a figure plots, and a listing prints records verbatim. All three are regenerated as the database moves toward lock, so you write the programs to rerun on a refreshed extract without editing.
- Double-programming is independent reproduction for quality control: a second programmer re-derives the dataset or output from the same source without seeing the first program, and the two are reconciled value by value, with a clean comparison serving as the sign-off.
This is the pitfall to watch, because an output produced by a single program without that adversarial re-derivation is not yet defensible for a regulated submission. Double-programming is, in effect, the reproducibility discipline made adversarial, meaning a second programmer works blind, re-deriving the result from the raw data without seeing the first program, so the two implementations catch each other's mistakes: the same number reached twice by independent paths rather than asserted once.
Learn the methods. Create a free account → to follow new methods and write-ups as they go up.