Protecting a Family of Hypotheses: FWER, FDR and Multiplicity Strategies
Testing many hypotheses increases false-positive risk. This guide compares family-wise error and false discovery rate methods with practical examples.
Why multiple testing matters
Every hypothesis test carries a chance of a false-positive result. If 20 independent true null hypotheses are tested at α=0.05, the probability of at least one false positive is approximately 1−0.9520=0.64. Dependence changes the exact calculation but not the central problem: searching across more tests makes selective “significance” increasingly likely.
FWER and FDR answer different questions
The family-wise error rate (FWER) is the probability of making at least one Type I error in a family of tests. It is often appropriate for confirmatory analyses. The false discovery rate (FDR) controls the expected proportion of false positives among findings declared significant and can be more suitable for high-dimensional exploratory research.
Bonferroni, Holm and Benjamini–Hochberg
Bonferroni tests each hypothesis at α/m or multiplies each raw p value by m. It is simple and valid under broad dependence structures but may be conservative. Holm's step-down procedure also controls FWER and is uniformly at least as powerful as simple Bonferroni. For many confirmatory settings, Holm is a strong default.
The Benjamini–Hochberg procedure compares ordered p values with i·q/m thresholds to control FDR. It is widely used in biomarker and omics screening. Its interpretation, dependence assumptions and the definition of the tested family should be prespecified.
Choosing a method
Few prespecified primary outcomes: limit the number of tests and consider Holm or a hierarchical strategy.
Many confirmatory secondary outcomes: retain FWER control.
Exploratory biomarker screening: FDR-adjusted results may be more useful.
Pairwise comparisons after ANOVA: use design-specific procedures such as Tukey, Dunnett or planned contrasts.
Common errors
Reporting only significant comparisons.
Assuming a significant omnibus test permits unlimited unadjusted pairwise tests.
Choosing the correction after viewing the results.
Confusing raw and adjusted p values.
Interpreting q=0.05 as proof that exactly 5% of reported findings are false.
Reporting example
“For eight prespecified secondary comparisons, FWER was controlled with Holm's procedure. Raw and adjusted p values are reported. For 120 exploratory biomarkers, the Benjamini–Hochberg procedure was applied at q=0.05.”
References
Holm S. A Simple Sequentially Rejective Multiple Test Procedure. 1979.
Benjamini Y, Hochberg Y. Controlling the False Discovery Rate. 1995.
Bender R, Lange S. Adjusting for multiple testing—when and how? 2001.
Accessed 20 June 2026.