Sensitivity analysis and robustness as design defenses
From Data to Bedside · the full write-up
When this applies
Use this write-up when you have a causal study, or any observational study whose headline claim depends on assumptions that cannot be tested with the data alone. The work covered here defends the result against the most-likely methodological critiques by pre-specifying sensitivity analyses that try to break the design. The ladder’s designed-in-versus-bolted-on distinction is the why; this write-up is the how: which analyses, matched to which assumption, computed which way, and reported in what table.
The deliverables a careful pass produces:
- A map of the assumptions your design rests on
- An identification of which assumption is most vulnerable
- A sensitivity-analysis method matched to that assumption
- Pre-specified sensitivity analyses in the protocol
- Bias-adjusted estimates under plausible violation scenarios
- Leave-one-out and placebo diagnostics
- Transparent reporting of all sensitivity results, not just the favorable ones
The decision framework
Seven steps for designing sensitivity and robustness analyses.
Step 1. Map the assumptions
Every analysis rests on assumptions. For a causal study, the load-bearing one is the design’s identifying assumption, the claim the data cannot test. For non-causal studies, it is whichever of the measurement model, the missing-data mechanism, the model specification, or the population definition the result leans on hardest.
Write the assumptions down explicitly. A study where you cannot enumerate the assumptions does not have a well-defined methodology.
Step 2. Identify which assumption is most vulnerable
Not all assumptions are equal. Some are routinely satisfied; others are routinely violated. The discipline is to identify which assumption is most likely to be challenged at review.
For DiD, parallel trends is the most-common challenge. For RDD, manipulation around the cutoff. For IV, the exclusion restriction. For propensity-score methods, unmeasured confounding. For analyses with substantial missing data, the missing-at-random (MAR) assumption. The most-vulnerable assumption gets the most sensitivity-analysis attention.
Step 3. Choose a sensitivity-analysis method matched to the assumption
Match the method to the assumption most likely to be challenged:
| Assumption at risk | Sensitivity method | What it buys |
|---|---|---|
| Unmeasured confounding | E-values (VanderWeele–Ding 2017), Rosenbaum bounds (Rosenbaum 2002) | quantifies how strong a hidden confounder would have to be to overturn the result |
| Parallel trends (DiD) | event-study + F-test on pre-period leads, honest DiD (Rambachan–Roth 2023), placebo periods | tests the assumption the design rests on, and bounds the effect under mild violations |
| Missing data | the MCAR/MAR/MNAR menu (operationalized in Step 5) | an estimate under each plausible mechanism, not just the convenient MAR one |
| Outcome misclassification | probabilistic bias analysis (Lash–Fox–Fink 2009), bounds under non-differential error | corrects the estimate for assumed mismeasurement instead of ignoring it |
| Selection bias | inverse-probability-of-selection weighting, Heckman selection model | re-weights or models who was observed, so the estimate isn’t silently conditioned on it |
| Model misspecification | alternative model forms, leave-one-out, specification curves (Simonsohn et al. 2020) | shows whether the conclusion holds across defensible choices or hangs on one |
Step 4. Pre-specify the sensitivity analyses in the protocol
The single discipline that most increases sensitivity-analysis credibility is pre-specification. Document, before any data work, which sensitivity analyses you will run, under what assumed violation scenarios, and how you will report them. Reviewers and IRB members read pre-specified sensitivity-analysis plans as a credibility signal.
The protocol should specify:
- The sensitivity-analysis method(s) for each vulnerable assumption
- The violation scenarios (for example, “we will assess robustness to unmeasured confounding at E-value of 1.5, 2.0, and 3.0”)
- The reporting plan (a table of bias-adjusted estimates across scenarios)
Step 5. Compute bias-adjusted estimates under plausible violation scenarios
For unmeasured-confounding sensitivity, the E-value (the ladder’s headline bias-quantification tool) becomes decision-useful once you attach thresholds and a reference to it. For an observed risk ratio above 1, the E-value is
\[ \text{E-value} = \text{RR} + \sqrt{\text{RR}\,(\text{RR} - 1)} \]
where:
- \(\text{E-value}\) is the minimum strength of association, on the risk-ratio scale, that an unmeasured confounder would need with both treatment and outcome to fully explain away the observed effect
- \(\text{RR}\) is the observed risk ratio, with a risk ratio below 1 replaced by its inverse before the formula is applied
An E-value above roughly 2.0 means a confounder would have to be associated with both treatment and outcome more strongly than most measured covariates are, which usually reads as robust; below about 1.5 means a fairly ordinary unmeasured confounder could erase the effect, which reads as fragile. The honest move is to name a real candidate confounder and ask whether its plausible strength clears the E-value. Compute it twice, for the point estimate and for the confidence-interval limit nearest the null; the second is usually what settles the call, since it asks whether unmeasured confounding could pull the whole interval back across no effect. Rosenbaum bounds are the matched-design analog, reporting the same idea as a sensitivity parameter \(\Gamma\): the factor by which a hidden confounder would have to raise the odds of treatment to overturn the result, so \(\Gamma = 2\) takes a confounder that doubles treatment odds, and a finding that survives \(\Gamma \geq 3\) is hard to dismiss.
For missing-data sensitivity, the ladder’s MCAR/MAR/MNAR distinction sets the menu; the operational rule is to report the estimate under each plausible mechanism, not only the convenient MAR one, and to make the spread between MAR and a tipping-point MNAR the headline whenever it is wide enough to change the conclusion.
For measurement-error sensitivity, simex or probabilistic bias analysis produce bias-corrected estimates and intervals. Report both the corrected estimate and the magnitude of the correction relative to the headline.
Step 6. Run leave-one-out and placebo diagnostics
Two diagnostics that earn their place across most causal designs:
- Leave-one-out. The ladder’s fragility check, made concrete: re-estimate dropping one unit at a time (one drug, one state, one cohort year) and read off not just whether the result flips but how far the point estimate moves. A coefficient that swings 40% when one unit is dropped is fragile even if it never crosses zero, and that magnitude is the reportable number, not just the binary survives/fails.
- Placebo (or falsification) tests. Re-run the design where it should find nothing: a pre-treatment period, an untreated unit, an outcome the treatment cannot plausibly touch. The operational discipline the ladder doesn’t spell out is to pre-register the placebo and its pass/fail line, because a falsification test invented after the main result is too easy to design so it passes.
Both are pre-specifiable. Both are credibility-positive when reported.
Step 7. Report all sensitivity results, not just the favorable ones
The single most-common failure of sensitivity analysis in published work is selective reporting: the unfavorable sensitivity result is buried or omitted. Pre-specify the sensitivity table at design time, and report it in full at write-up time.
A transparent sensitivity table (with the headline estimate, the sensitivity-analysis-method-specific estimates, and the violation-scenario column) is more credible than a single headline number followed by a “robust to sensitivity analysis” assertion in the discussion.
Worked example
The NHANES cardiometabolic case study walks through case-definition sensitivity (NCEP ATP III, IDF 2005, JIS 2009) as a measurement-sensitivity exercise: how the headline prevalence shifts when the operationalization of the outcome shifts.
The Medicaid outliers case study walks through peer-group sensitivity (a related form of population-definition robustness) and methodological-procedure sensitivity (the opacity rule that handles “no information” answers in the BH-FDR setup, an explicit departure from textbook BH that the field note flags as v0.1’s biggest methodological stake).
The Part D insulin DiD case study is the most-complete sensitivity-analysis example in the portfolio: placebo cap-years, leave-one-out, control-group composition sensitivity (drop GLP-1 RAs to absorb the concurrent demand shock), and prescriber-FE robustness as a specification check.
Further reading
- VanderWeele TJ, Ding P. Sensitivity analysis in observational research: introducing the E-value. Annals of Internal Medicine 167(4): 268–274. 2017.
- Rosenbaum PR. Observational Studies. 2nd ed. Springer, 2002.
- Lash TL, Fox MP, Fink AK. Applying Quantitative Bias Analysis to Epidemiologic Data. Springer, 2009.
- Rambachan A, Roth J. A more credible approach to parallel trends. Review of Economic Studies 90(5): 2555–2591. 2023.
- van Buuren S. Flexible Imputation of Missing Data. 2nd ed. CRC Press, 2018.
- Simonsohn U, Simmons JP, Nelson LD. Specification curve analysis. Nature Human Behaviour 4(11): 1208–1214. 2020.
← Back to the pathway · the full write-up behind the Defend it rung.
Learn the methods. Create a free account → to follow new write-ups and traces as they go up, alongside the full From Data to Bedside pathway.