It All Comes Down to Noise

Info sheet · Statistics for Psychology & Neuroscience

Author

Andrew Bell

Published

September 22, 2026

Info sheet 0.2 (draft) · Prerequisites: the mean and variance · Give feedback ↗

Working notes for the author — not shown to students once collapsed; remove before publishing.

What you’ll get from this sheet

The single idea the rest of the book is built on. By the end you should be able to:

  1. Say why variance is the subject, not a nuisance.
  2. Explain why significance is always signal relative to noise.

Almost everything in statistics is an argument about noise. Every test — from a t‑test to a mixed model — compares the variation it can explain (signal) against the variation it can’t (noise). “Significant” means the signal is large relative to the noise.

Variance is the subject

If you take one idea from this book, let it be this: almost everything in statistics is an argument about noise. Measure the same thing twice — one person’s reaction times on two days, a neuron’s firing rate across identical trials, two classes taught the same way — and you never get the same numbers. The values wobble. Some of that wobble is interesting (a real effect of caffeine on reaction time); some is noise — the countless tiny, uncontrolled influences (attention, temperature, what they had for breakfast) that push a measurement around for reasons unrelated to your question.

That wobble has a name you know: variance. Variance isn’t a nuisance we tolerate on the way to the “real” statistics — variance is the subject. Every test in this book is a machine for one thing: comparing the variation we can explain against the variation we cannot. When people say a result is “significant,” they mean the signal is large relative to the noise. A big effect drowning in noise can be undetectable; a tiny effect measured with exquisite precision can be unmistakable. Signal alone tells you almost nothing — it’s always signal per unit of noise.

A concrete case: does the drug change heart rate?

Suppose we want to know whether a drug raises heart rate. The obvious move is to measure heart rate before and after and compare the two numbers — an A − B. With a single reading each, that subtraction looks clean and certain. But repeat the experiment on another day and the difference won’t be the same: heart rate is pushed around by temperature, time of day, caffeine, sleep, and mood — countless influences unrelated to the drug. That irreducible wobble is the noise, and the real question is never just “is A − B big?” but “is A − B bigger than the noise?”

This is exactly why one measurement can’t settle the question and many can. A single reading (n = 1) gives you a lone dot with no sense of how much it would have wobbled; a hundred readings (n = 100) trace out the whole distribution of that wobble, so you can finally judge whether a before/after gap really stands out from it. Drag the slider and watch a cloud of daily readings fill in a distribution:

Once you can see that spread of noise, the whole game becomes visual: is the gap between two conditions large compared with the width of these clouds?

Same effect, different noise

Two groups whose true means differ by a fixed amount. The signal (that gap) never changes — only the noise does. Slide it and watch whether the difference is visible or buried:

The gap between the two means is fixed at 3 the whole time — yet at low noise the groups are obviously different, and at high noise they melt into one blur. Nothing about the effect changed; only the noise did. That ratio, signal ÷ noise, is what every test is really measuring — here as Cohen’s d, elsewhere as an F‑ratio, a t‑statistic, or an R².

Keep the lens close

This is why, before running a single test, we spend so long thinking about where variance comes from and how to measure it. As each new method appears, ask: what variance is this technique trying to explain, and what variance is it treating as noise? The whole edifice of inferential statistics is built from that one question.

See it in code

rng(1); effect = 3;
for sd = [1 3 6]                           % same effect, more noise each time
    a = 10 + sd*randn(30,1); b = 10 + effect + sd*randn(30,1);
    d = (mean(b) - mean(a)) / sd;          % signal-to-noise (Cohen's d)
    [~, p] = ttest2(a, b);
    fprintf('noise SD %d  d = %.2f  p = %.4f\n', sd, d, p);
end

The R and Python tabs run live; MATLAB is a static reference.

Two studies find the exact same difference in means between conditions, but one is “significant” and the other isn’t. How is that possible?

Because significance depends on the difference relative to the noise, not the difference alone. The significant study measured with less noise — smaller within‑group variance and/or a larger sample (which shrinks the standard error). Same signal, less noise → bigger signal‑to‑noise ratio → detectable. It’s never the size of the effect on its own; it’s the effect per unit of noise.

Don’t read “not significant” as “no effect” — a real effect can be swamped by noise (low power), and “significant” as “big / important” — a trivial effect can clear the bar with a huge, precise sample. Report an effect size (like Cohen’s d or R²) alongside the p‑value so the signal and the noise are both visible.

Where this shows up next

This lens returns everywhere. Next, Signal and Error makes it an equation (Data = Signal + Error); the F‑ratio, R², and t‑statistic are all just versions of the same ratio.