Confidence Intervals

Info sheet · Statistics for Psychology & Neuroscience

Author

Andrew Bell

Published

September 22, 2026

Info sheet 0.2 (draft) · Prerequisites: the normal distribution; standard error · Give feedback ↗

Working notes for the author — not shown to students once collapsed; remove before publishing.

What you’ll get from this sheet

By the end you should be able to:

  1. Interpret a confidence interval correctly.
  2. Relate the 95% CI to the standard error and the 68–95–99.7 rule.

A 95% confidence interval is estimate ± ~2 standard errors. Correct reading: over many repeated samples, 95% of such intervals contain the true value — not ‘95% chance the truth is in this one’.

From a point estimate to an interval

A single number — “the mean improvement was 5 points” — hides how much that estimate would have wobbled with a different sample. A confidence interval puts a ruler around the estimate: a range of plausible values for the true quantity, given the noise in your data. It’s the honest way to report an estimate, because it shows both where you landed and how precise that landing was.

The standard error, and the ± rule

The width comes from the standard error (SE) — the standard deviation of the sampling distribution, i.e. how much the estimate itself would vary from sample to sample. For a mean, \(\text{SE}=\sigma/\sqrt{n}\). Because the estimate is (via the CLT) approximately normal, the 68–95–99.7 rule applies directly to it: about 95% of the time the estimate lands within ~2 SE of the truth. Turn that around and you get the 95% CI:

\[\text{95\% CI} = \text{estimate} \pm 1.96 \times \text{SE}.\]

Small SE (large n, low variability) → tight interval; large SE → wide one.

Interpreting it correctly — the dance of the CIs

Here’s the subtlety everyone trips over. “95% confidence” is a property of the procedure, not of one interval. It means: if you repeated the study many times, about 95% of the intervals you’d construct would contain the true value. Any single interval either contains the truth or it doesn’t — there’s no probability left once it’s computed. Watch 40 different samples each draw their own 95% interval around a known true mean:

Confidence level and width

Raising the confidence level doesn’t make you more right about the same interval — it makes the interval wider, because to be more certain of catching the truth you must cast a bigger net. Drag the level from 50% to 99.9% and watch the interval stretch (the 90 / 95 / 99% reference bars underneath show the standard choices):

So a 99% interval is always wider than a 95%, which is wider than a 90%, on the same data — there’s no free lunch, only a trade between confidence and precision.

CIs and significance

A confidence interval quietly contains a significance test. If a 95% CI for a difference excludes 0, then the difference is significant at α = .05 (you’d reject “no difference”); if it includes 0, it isn’t. So a CI does everything a p-value does and shows the effect size and its precision — which is why reporting guidelines increasingly prefer it.

Quick quiz: true or false?

For each statement, decide whether it’s a correct reading:

Answer each question — the feedback appears beneath it.

See it in code

rng(1);
x = 100 + 15*randn(30,1);
m = mean(x); se = std(x)/sqrt(numel(x));
[m - 1.96*se, m + 1.96*se]                 % by hand (~95%)
[~,~,ci95] = ttest(x);                     % exact 95% CI (t-distribution)
[~,~,ci99] = ttest(x, 0, 'Alpha', 0.01);   % 99% CI — wider
disp(ci95'); disp(ci99')

The R and Python tabs run live; MATLAB is a static reference.

A study reports a mean improvement of 5 points, 95% CI [1, 9]. Give one correct interpretation, and say whether the effect is significant at α = .05.

Correct interpretation: the procedure that produced this interval captures the true improvement 95% of the time (over many repeated studies) — not “there’s a 95% chance the truth is between 1 and 9.” Because the interval excludes 0, the improvement is significant at α = .05: “no improvement” is not among the plausible values. And usefully, the CI adds what a bare p-value wouldn’t — the effect is somewhere around 1 to 9 points, so you can judge whether it’s practically meaningful.

The seductive misreading is “95% probability the true value is in this interval” — wrong, because a computed interval no longer has a probability; the 95% belongs to the long-run procedure. Also, overlapping CIs don’t map neatly onto significance: two 95% intervals can overlap a little and the difference still be significant, so don’t eyeball two error bars and declare “no difference.” And a CI inherits every assumption of its SE — a biased sample gives a precise interval around the wrong value.

Where this shows up next

Confidence intervals are the estimation-flavoured twin of the p-value, and they lead straight into power and effect size (how precise an estimate your sample size can buy). See the inference chapter.