Within- vs Between-Subjects Designs

Info sheet · Statistics for Psychology & Neuroscience

Author

Andrew Bell

Published

September 22, 2026

Info sheet 0.2 (draft) · Prerequisites: experiments vs observation · Give feedback ↗

Working notes for the author — not shown to students once collapsed; remove before publishing.

What you’ll get from this sheet

By the end you should be able to:

  1. Tell a within-subjects design from a between-subjects one.
  2. Say why within-subjects designs are more powerful but carry order effects.

Between-subjects: each participant sees one condition. Within-subjects: each sees all — more powerful (each person is their own baseline) but exposed to order/carryover effects, which is why we counterbalance.

Same people or different people?

When a study has more than one condition, you can compare either different people across conditions (a between-subjects design) or the same people across conditions (a within-subjects, or repeated-measures, design). It sounds like a logistical detail, but it changes the statistics profoundly — because it changes what counts as noise.

Why within-subjects is more powerful

People differ enormously from one another — in baseline reaction time, reading speed, brain volume, mood. In a between-subjects design all that person-to-person variability lands in the error term, and the effect you care about has to shout over it. In a within-subjects design each participant serves as their own control: because you compare each person to themselves, the huge differences between individuals cancel out, leaving a much cleaner view of the effect. Same effect, far less noise. Slide the between-person variability up and watch the between-subjects comparison drown while the within-subjects one holds steady:

Notice the effect is fixed at +3 the whole time. As between-person SD grows, the two point-clouds spread out vertically and their means get hard to tell apart — the between-subjects effect size shrinks — yet the connecting lines keep sloping up together, so the within-subjects effect barely budges. That’s the extra power of a repeated-measures design, drawn out.

The cost, and how we pay it

Within-subjects designs aren’t free. Because each participant meets every condition, they’re exposed to order and carryover effects: practice, fatigue, boredom, or sensitisation can make the second condition differ from the first for reasons that have nothing to do with your manipulation. The standard defence is counterbalancing — varying the order across participants (e.g. a Latin square) so that, averaged over the sample, order effects cancel rather than masquerade as the effect of interest.

See it in code

rng(1); n = 30; effect = 3;
person = 5*randn(n,1);                 % big between-person differences
A = 12 + person + 1.2*randn(n,1);
B = 12 + person + effect + 1.2*randn(n,1);
[~, p_ind]    = ttest2(A, B);          % between-subjects (independent) — weak
[~, p_paired] = ttest(A, B);           % within-subjects (paired) — strong
[p_ind, p_paired]

The R and Python tabs run live; MATLAB is a static reference.

You want to test whether a new font improves reading speed. You could (a) have each person read one passage in the old font and one in the new, or (b) split people into two groups, one font each. Which is more powerful, and what’s the main risk to manage?

Design (a) is within-subjects and more powerful: each person is their own control, so the large differences between individuals’ baseline reading speeds cancel out of the comparison, leaving a cleaner estimate of the font effect. The main risk is order/carryover effects — whichever passage or font comes second benefits from practice (or suffers fatigue). You manage it by counterbalancing: half the participants get the old font first, half the new.

Within-subjects power isn’t free, and counterbalancing has limits. It cancels symmetric order effects (practice, fatigue), but not irreversible ones: if a condition teaches a strategy, reveals the hypothesis, or changes the participant for good, they can’t return to baseline for the next condition. When exposure to one condition contaminates the other, a between-subjects design is the safer choice despite the power cost.

Where this shows up next

This choice dictates the analysis. A between-subjects comparison leads to an independent ANOVA; a within-subjects one leads to the repeated-measures ANOVA, which exists precisely to exploit the extra power a within-subjects design provides.