Family-Wise Error
Info sheet · Statistics for Psychology & Neuroscience
Warning✎ Editing notes — to do / to check
Working notes for the author — not shown to students once collapsed; remove before publishing.
What you’ll get from this sheet
By the end you should be able to:
- Explain why running many tests on the same data inflates false positives.
- Count the pairwise comparisons among k groups and describe how that number grows.
- Compute the family-wise error rate — and play with it interactively below.
The more tests you run on the same data, the more likely at least one “significant” result is a false alarm. You have to budget for that risk before interpreting any single test.
The problem in one sentence
Roll a fair die once and your chance of a six is just \(1/6\). Roll it ten times and your chance of at least one six jumps to \(1-(5/6)^{10}\approx 0.84\) — about 84%. Nothing changed about any single roll; you just gave yourself more chances. Statistical tests behave the same way: every test carries a small risk of a false positive (a Type I error), and the more tests you run on the same data, the more likely at least one “significant” result is a fluke. The more you look, the more you find — meaningful or not. That accumulated risk is the family-wise error rate (FWE).
How many comparisons?
If you want to compare every pair of k groups, the number of comparisons is
\[ \binom{k}{2} = \frac{k(k-1)}{2}. \]
It’s worth being precise here: this grows quadratically with k (like \(k^2\)), not exponentially.
| Groups (k) | Pairwise comparisons |
|---|---|
| 2 | 1 |
| 3 | 3 |
| 4 | 6 |
| 5 | 10 |
| 6 | 15 |
The family-wise error rate
If you run m independent tests, each at significance level \(\alpha\), the chance of at least one false positive is
\[ \text{FWE} = 1 - (1-\alpha)^m. \]
This is the die-roll story from the top, made exact: with \(m=10\) chances and per-event probability \(\alpha=1/6\) you get the \(1-(5/6)^{10}\approx0.84\) we started with. For statistical tests at the usual \(\alpha=.05\), the family-wise error climbs just as relentlessly as \(m\) grows.
Try it yourself
Change alpha and m below and press Run. What value of m pushes the family-wise error past 50%?
alpha = 0.05; % per-test significance level
m = 10; % number of independent tests
fwe = 1 - (1 - alpha)^mReference code — identical in MATLAB and Octave. Browser cells can’t execute MATLAB/Octave, so this tab is static (copy it into MATLAB or free Octave to run it).
And here’s the whole curve — the reason “just run all the t-tests” is dangerous:
m = 1:20;
plot(m, 1 - (1 - 0.05).^m, 'o-'); hold on
plot([1 20], [0.05 0.05], '--') % the nominal per-test rate
xlabel('Number of comparisons')
ylabel('Family-wise error rate')
title('P(at least one false positive)')Static reference (MATLAB/Octave). The R and Python tabs run live in the page.
TipCheck your understanding
With 4 groups, how many pairwise comparisons are there, and what is the family-wise error rate at α = .05?
4 groups → \(4\times3/2 = \mathbf{6}\) comparisons. FWE \(= 1 - 0.95^{6} \approx \mathbf{0.26}\) — about a 26% chance of at least one false positive. (Confirm it by setting m <- 6 in the cell above.)
Controlling the family-wise error
The remedy is to spend your α budget carefully across the comparisons, so the family-wise error — not the per-test error — stays at 5%. A few standard tools, roughly from bluntest to smartest:
- Bonferroni. Divide α by the number of tests: judge each at \(\alpha/m\). Dead simple and guarantees FWE ≤ α, but conservative — with many tests the per-test threshold gets tiny and real effects can be missed.
- Holm–Bonferroni. A stepwise improvement: sort the p-values, compare the smallest to \(\alpha/m\), the next to \(\alpha/(m-1)\), and so on until one fails. It controls the family-wise error just as strictly as Bonferroni but is uniformly more powerful, so you lose fewer true effects — there’s rarely a reason to prefer plain Bonferroni over Holm.
- Tukey’s HSD. Built specifically for comparing all pairs of means after an ANOVA. Rather than a blanket α-split it uses the studentised range distribution, which accounts for the number of groups directly — giving tighter, better-calibrated comparisons than Bonferroni when you genuinely want every pair.
- Cluster-extent (for imaging). Exploits that true signals are spatially contiguous while noise is scattered — it keeps only blobs larger than a size threshold. That’s how fMRI tames tens of thousands of voxel tests.
- Or ask one question. Where a single omnibus test suffices, an ANOVA sidesteps the problem entirely — one test, one α.
(These get the full treatment in the ANOVA chapter; here they’re just enough to make sense of the demo below.)
See it: false positives across a “brain”
fMRI shows the problem at its most dramatic. A single scan carves the brain into tens of thousands of voxels, and the standard analysis runs a separate statistical test at every one. At an uncorrected α = .05, roughly 1 in 20 voxels crosses the threshold by chance alone — thousands of phantom “activations.” The famous demonstration is Bennett and colleagues’ dead Atlantic salmon, which appeared to show brain activity in response to emotional photographs when the authors deliberately skipped multiple‑comparisons correction.
The toy axial slice below hides one real activation (it shows up in green) in a field of otherwise pure noise (red = false positive). Your challenge: isolate only the real green activation — no red noise — by tuning the threshold and correction method. Watch what each approach keeps versus throws away:
With no correction, the single genuine cluster (green) is hard to pick out among the scattered red false positives. Bonferroni is safe but blunt — strict enough to erase a weak signal along with the noise. Cluster‑extent correction exploits the fact that true activations are spatially contiguous while noise is scattered: it keeps the real cluster and clears the rest — essentially how fMRI turns a field of noise into a trustworthy map.
Where this shows up next
The corrections above — Bonferroni, Holm, Tukey, and the ANOVA route — get their full treatment in Chapter 8 (Analysis of Variance); this sheet is the stand-alone primer.