The Poisson Distribution

Info sheet · Statistics for Psychology & Neuroscience

Author

Andrew Bell

Published

September 22, 2026

Info sheet 0.2 (draft) · Prerequisites: probability basics; the binomial distribution · Give feedback ↗

Working notes for the author — not shown to students once collapsed; remove before publishing.

What you’ll get from this sheet

The distribution for counting events at a rate — spikes per second, emails per hour, rare incidents per year. By the end you should be able to:

  1. Recognise when the Poisson applies, and why its mean and variance are both \(\lambda\).
  2. Say why a normal curve is the wrong tool for rare-event counts.

The Poisson counts events in a fixed window (of time or space) occurring at an average rate \(\lambda\), with no fixed upper limit. Its PMF is \(P(X=k)=\dfrac{e^{-\lambda}\lambda^{k}}{k!}\), and its signature is that its mean and variance are both \(\lambda\). It’s discrete, and it’s the limit of the binomial for rare events (\(\lambda = np\) with large \(n\), small \(p\)).

Counting events at a rate

Where the binomial needs a fixed number of trials, the Poisson needs only a rate. You count how many events land in a window — a second, an hour, a page, a year — when there’s no natural upper limit and events occur independently at a constant average rate \(\lambda\):

  • spikes from a neuron per second;
  • emails arriving per hour;
  • typos per page, goals per match, earthquakes per decade.

Its defining quirk falls straight out of the maths: the mean equals the variance, both \(\lambda\). So if someone tells you a count process has mean 4 and variance 4, that’s a Poisson fingerprint.

The rare-event limit of the binomial

The Poisson is what a binomial becomes when trials are many and each success is unlikely. Imagine chopping an hour into thousands of tiny slices, each with a tiny chance of an event: a binomial with huge \(n\) and minuscule \(p\). As \(n\to\infty\) and \(p\to0\) with \(np=\lambda\) held fixed, the binomial converges to a Poisson(\(\lambda\)). That’s why the Poisson is the natural model for rare events counted over a continuous window.

Explore the shape

Move the rate \(\lambda\) and watch the bars, the mean, and the variance (which stay equal) respond:

At small \(\lambda\) the distribution is a sharply right-skewed pile near zero — most windows have few or no events, a few have several. As \(\lambda\) grows it becomes more symmetric and, for large \(\lambda\), approaches a normal shape. But the mean and variance stay locked together the whole way.

Why should I care?

Because a huge amount of real data is counts of things that happen at a rate, and modelling them as a bell curve gives wrong answers. Rare-event counting shows up in neuroscience (spike counts), epidemiology (cases per week), safety and public policy (accidents, incidents), quality control (defects per batch), and queueing (arrivals per minute). In every case the Poisson respects two facts a normal ignores: counts can’t be negative, and they come in whole numbers.

NoteIn the wild: counting rare events — Australia’s gun laws

Consider mass shootings. They are, mercifully, rare — discrete events happening at some low average rate over years, with no fixed number of “trials.” That’s exactly the setting the Poisson distribution was built for, and exactly where a normal distribution misleads: a bell curve would assign probability to negative and fractional shootings, and would badly misjudge the chance of a long run of zero-event years.

This is the reasoning behind Simon Chapman and colleagues’ influential study of Australia’s 1996 National Firearms Agreement — the sweeping reforms enacted after the Port Arthur massacre, in which 35 people were killed (Chapman, Alpers & Jones, 2016, JAMA 316(3):291–299). In the 18 years before the reforms Australia recorded 13 fatal mass shootings; in the 20 years after, none. To ask whether that drop is more than chance you don’t compare means with a t-test — you model the counts as a Poisson process (events in non-overlapping intervals treated as independent Poisson/negative-binomial variables) and ask how surprising a run of zeros would be under the pre-reform rate. Match the distribution to the phenomenon — rare-event counts are Poisson, not normal — and the statistics finally fit the question.

Poisson vs Gaussian: watch the bell curve fail

Same hypothetical count data — an average of λ rare events per period — modelled two ways: the Poisson (correct for counts, green bars) and a normal curve \(N(\lambda,\lambda)\) laid over it (the tempting but wrong “bell curve”, blue line). Drag λ down toward the rare-event range and watch the Gaussian break.

At a genuinely rare rate — under one event per period, roughly Australia’s pre-reform mass-shooting rate of about 0.7 per year — the two models disagree sharply. The Poisson says a zero-event period is common (\(e^{-\lambda}\)); the normal, forced to be symmetric around λ, understates that zero probability and hangs a slab of probability out over negative counts, which can’t occur. Only when λ is large do the bars and the bell line up. That mismatch is the whole case for choosing a distribution that matches the data-generating process.

Discrete, not continuous

Like the binomial, the Poisson is discrete — its outcomes are whole counts (0, 1, 2, …), each carrying a lump of probability, and none of them negative. The normal is continuous and unbounded, so at low rates it does two impossible things at once: it spreads probability over fractional counts and over negative ones. That’s why the demo above shows the bell curve leaking mass to the left of zero. Use the Poisson while events are rare; only once λ is large is the normal approximation safe.

See it in code

poisspdf(3, 4)          % P(X = 3)
[mean(poissrnd(4,1e5,1)), var(poissrnd(4,1e5,1))]  % ~ equal (= lambda)
poisspdf(0, 0.7)        % P(zero events)

The R and Python tabs run live; MATLAB is a static reference.

A neuron fires on average 4 times per second. Which distribution models the spike count in one second, and what are its mean and variance?

The Poisson, with \(\lambda = 4\) — it counts events in a fixed window with no upper limit and a constant average rate. Its mean and variance are both 4, so the SD is \(\sqrt{4}=2\). (If instead you had a fixed number of trials each with a success probability — say 20 stimuli and the chance of a response — you’d use the binomial.) A useful diagnostic: if your count data have variance much bigger than their mean, a plain Poisson is too simple.

Match the distribution to the process: the Poisson is for open-ended counts at a rate, not for a fixed number of trials (that’s the binomial). Don’t default to the normal for counts — it allows negative and fractional values and fits badly for small means (though for large λ it’s a fine approximation). And real count data are often overdispersed — variance noticeably bigger than the mean — which breaks the Poisson’s mean = variance assumption; reach for a negative-binomial or quasi-Poisson model then.

Where this shows up next

The Poisson is the outcome distribution behind Poisson regression — the count-data member of the generalised linear model family, alongside logistic regression for binomial outcomes.