Power & Effect Size

Info sheet · Statistics for Psychology & Neuroscience

Author

Andrew Bell

Published

September 22, 2026

Info sheet 0.2 (draft) · Prerequisites: what a p-value is; the normal distribution · Give feedback ↗

Working notes for the author — not shown to students once collapsed; remove before publishing.

What you’ll get from this sheet

By the end you should be able to:

  1. Define effect size and statistical power.
  2. Run a power analysis to plan a sample size.

Effect size measures how big an effect is, independent of n; power is the chance of detecting a real effect. Power rises with effect size, α, and — crucially — sample size. Underpowered studies miss real effects and produce flukier significant ones.

Effect size: how big, regardless of n

A p-value tells you whether an effect is detectable; an effect size tells you how big it is — on a scale that doesn’t depend on your sample size. The common measures:

  • Cohen’s d — a standardised difference between two means: \(d = \dfrac{\text{mean difference}}{\text{pooled SD}}\). Rough conventions: 0.2 small, 0.5 medium, 0.8 large.
  • Pearson’s r — the standardised strength of a linear association.
  • η² (eta-squared) — the proportion of variance explained, used with ANOVA.

Effect size is what makes results comparable across studies (the currency of meta-analysis) and what stops “significant” from being mistaken for “big.”

Power: the chance of catching a real effect

Statistical power is the probability that a study will detect an effect that truly exists — formally \(1-\beta\), where β is the Type II error rate (missing a real effect). The field’s conventional target is 80%. Power is tied to three other quantities in a fixed relationship — α, effect size, and sample size n — so that once you pin down any three, the fourth is determined. Two of them are the levers you actually control: bigger n and (if it’s genuinely there) a bigger effect both push power up. Watch the alternative distribution slide away from the null as you raise either:

A priori power analysis

The single most valuable habit here: run the power analysis before collecting data. Fix your α (usually .05), your target power (usually .80), and your best guess at the effect size, then solve for the sample size n you’ll need. It turns “how many participants?” from a guess into an answer — and often reveals that a study as planned has little chance of working, saving you from running it. (Where you can’t set n freely, run a sensitivity analysis instead: given the n you can get, what’s the smallest effect you’d have 80% power to detect?)

Why underpowered studies are doubly cursed

An underpowered study — too small for the effect it’s chasing — fails twice over. First, it misses real effects: low power means a genuine effect often doesn’t reach significance. Second, and less obviously, the significant results it does produce tend to overstate the effect — the only way a small, noisy study clears the bar is if the sampling noise happened to inflate the estimate. This “winner’s curse” is a major engine of non-replicable findings: the published effect is bigger than the truth, so the next study can’t reproduce it.

See it in code

% Statistics Toolbox: two-sample t-test, effect size d = 0.5 (sd = 1)
n  = sampsizepwr('t2', [0 1], 0.5, 0.80);     % n per group for 80% power
pw = sampsizepwr('t2', [0 1], 0.5, [], 30);   % power at n = 30/group
[n, pw]

The R and Python tabs run live; MATLAB is a static reference.

You expect a medium effect (Cohen’s d = 0.5) and want 80% power at α = .05 for a two-group comparison. Roughly how many participants per group do you need — and what happens if you run only 20 per group?

About 64 per group (≈128 total) — the standard result for d = 0.5, 80% power, two-sided α = .05. With only 20 per group you’d have roughly 35% power, so a real medium effect would fail to reach significance about two-thirds of the time — and any significant result you did get would likely overstate the effect (the winner’s curse). The fix is more participants (or a within-subjects design, which buys power by cutting noise), decided before you run.

Beware post-hoc (“observed”) power — power recomputed from the effect size your study happened to find. It’s circular (it’s just a re-expression of the p-value) and tells you nothing useful; plan power a priori instead. Don’t confuse a big effect size with a small p — a huge sample makes a trivial effect significant, and a tiny sample can miss a large one. And remember power analysis needs an assumed effect size: feed it an optimistic guess and it will happily under-power your study.

Where this shows up next

Power ties the whole inference chapter together — it’s the noise-vs-signal logic applied to planning, and it pairs naturally with confidence intervals (precision) and effect sizes (magnitude). It returns whenever we ask whether a study could actually answer its question.