The Binomial Distribution

Info sheet · Statistics for Psychology & Neuroscience

Author

Andrew Bell

Published

September 22, 2026

Info sheet 0.2 (draft) · Prerequisites: probability basics; the idea of a distribution · Give feedback ↗

Working notes for the author — not shown to students once collapsed; remove before publishing.

What you’ll get from this sheet

The distribution for a fixed number of yes/no trials — coin flips, dice rolls, patients responding. By the end you should be able to:

  1. Recognise when the binomial applies, and give its mean and variance.
  2. Read a binomial coefficient \(\binom{n}{k}\) and say what it counts.

The binomial counts successes in a fixed number \(n\) of independent yes/no trials, each with success probability \(p\). Its PMF is \(P(X=k)=\binom{n}{k}p^k(1-p)^{n-k}\), with mean \(np\) and variance \(np(1-p)\). It’s a discrete distribution — \(k\) is a whole number from 0 to \(n\).

When the binomial applies

Reach for the binomial whenever you have a fixed number of trials, each an independent success/failure with the same probability \(p\), and you’re counting the successes:

  • how many of 10 coin flips come up heads (\(n=10\), \(p=0.5\));
  • how many sixes in 12 rolls of a die (\(n=12\), \(p=\tfrac{1}{6}\approx0.17\));
  • how many of 20 patients respond to a drug (\(n=20\), \(p=\) the response rate).

The three ingredients — a fixed \(n\), independent trials, and a constant \(p\) — are exactly the assumptions. Break any of them and the binomial no longer applies.

With replacement, or without?

The “constant \(p\), independent trials” assumption is really a statement about sampling with replacement. Flip a coin or roll a die and the odds reset every time — the coin has no memory. But draw without replacement from a small pool and \(p\) shifts as you go: deal cards from a 52-card deck and the chance of an ace changes with every card removed, because the trials aren’t independent. That situation is the hypergeometric distribution’s job, not the binomial’s. The saving grace: when the population is large relative to your sample, removing a few members barely moves \(p\), so the binomial is an excellent approximation even without replacement.

The binomial coefficient: \(\binom{n}{k}\)

The \(\binom{n}{k}\) (“n choose k”) out front counts the number of different ways \(k\) successes can be arranged among \(n\) trials:

\[\binom{n}{k}=\frac{n!}{k!\,(n-k)!}\]

For example, in 3 coin flips there are \(\binom{3}{2}=\frac{3!}{2!\,1!}=3\) ways to get exactly 2 heads — HHT, HTH, THH — so the probability of 2 heads is \(3\times0.5^2\times0.5^1=0.375\). In 5 flips there are \(\binom{5}{2}=10\) ways to get 2 heads. The coefficient is why the middle outcomes are more likely than the extremes: there are many arrangements that give a middling count, but only one that gives all successes (\(\binom{n}{n}=1\)) or none (\(\binom{n}{0}=1\)).

Explore the shape

Move \(n\) and \(p\) and watch the bars, the mean, and the variance respond (set \(p=0.5\) for a coin, \(p\approx0.17\) for rolling a six):

At \(p=0.5\) the bars are symmetric; push \(p\) toward 0 or 1 and they skew, and the variance \(np(1-p)\) is always less than the mean \(np\). Crank \(n\) up and (for middling \(p\)) the discrete bars trace out a bell shape — that’s the normal approximation to the binomial, valid when \(np\) and \(n(1-p)\) are both large.

Discrete, not continuous

Keep one thing straight: the binomial is discrete. Its outcomes are whole numbers of successes — you can get 3 heads or 4 heads, never 3.5 — so it’s drawn as separate bars, and each bar carries a genuine lump of probability (unlike the continuous normal, where probability is an area under a smooth curve and any exact value has probability zero). The normal can approximate a large-\(n\) binomial, but underneath it’s still counting whole events.

See it in code

binopdf(8, 20, 0.3)     % P(X = 8)
nchoosek(20, 8)         % C(20, 8)
[20*0.3, 20*0.3*0.7]    % mean, variance

The R and Python tabs run live; MATLAB is a static reference.

Quick quiz: which distribution?

Pick a scenario and choose the distribution you’d use:

Answer each question — the feedback appears beneath it.

You flip a fair coin 6 times. How many ways are there to get exactly 4 heads, and what’s the probability?

The number of arrangements is \(\binom{6}{4}=\frac{6!}{4!\,2!}=15\). Each specific sequence of 4 heads and 2 tails has probability \(0.5^6=\tfrac{1}{64}\), so \(P(\text{4 heads})=15\times\tfrac{1}{64}\approx0.234\). (Check the mean: \(np = 6\times0.5 = 3\), so 4 heads is just above average — plausible, hence the fairly high probability.)

The binomial’s assumptions are easy to break. It needs a fixed \(n\), independent trials, and a constant \(p\) — so sampling without replacement from a small pool (where \(p\) drifts) calls for the hypergeometric instead, and trials that influence each other (streaks, learning) violate independence. And don’t forget it’s discrete: a normal curve can approximate a large-\(n\) binomial, but it will happily assign probability to impossible fractional or negative counts.

Where this shows up next

The binomial is the outcome distribution behind logistic regression — the “binomial family” of the generalised linear model. Its rare-event cousin, the Poisson, is the very next sheet.