Mathematical Notation & Key Formulas

Info sheet · Statistics for Psychology & Neuroscience

Author

Andrew Bell

Published

September 22, 2026

Info sheet 0.2 (draft) · Prerequisites: none · Give feedback ↗

Working notes for the author — not shown to students once collapsed; remove before publishing.

What you’ll get from this sheet

A decoder ring for the symbols and formulas used across the book — a recap of the conventions from RM1 and RM2. By the end you should be able to:

  1. Interpret standard mathematical notation, and tell population from sample statistics.
  2. Read the notation for a normal random variable and a general linear model.

Two habits unlock most of it: population quantities get Greek letters (\(\mu\), \(\sigma\)), their sample estimates get Latin letters (\(\bar{x}\), \(s\)); and the letters carry conventions — \(\alpha\) is the significance level, \(\beta\) a coefficient (or power), \(\varepsilon\) the error in a model. The big \(\Sigma\) just says add these up.

Letter conventions

A few habits make formulas readable:

  • Coordinates: \(X\) (horizontal), \(Y\) (vertical), \(Z\) (depth). Because of this, avoid using \(X, Y, Z\) as variable names unless you really mean locations.
  • Counting indices: \(i, j, k, n\) are the counters in formulas — \(x_i\) is “the \(i\)-th value”, \(n\) is how many there are.
  • \(\alpha\) (alpha): the significance level, typically 0.05 or 0.01.
  • \(\beta\) (beta): a regression coefficient — or the power of a test, depending on context.
  • \(\varepsilon\) (epsilon): the error term in a linear model.

Population vs sample — the distinction to learn first

Almost every symbol comes in a pair: the true value in the population (which we never see) and our estimate of it from a sample. Descriptive statistics summarise data (mean, median, mode, SD); inferential statistics let us generalise (t-values, p-values). The parameters and their estimates:

Quantity Population (parameter) Sample (estimate)
Mean \(\mu\) (mu) \(\bar{x}\) (x-bar)
Standard deviation \(\sigma\) (sigma) \(s\)
Variance \(\sigma^2\) \(s^2\)
Size \(N\) \(n\)

The pattern: Greek = the truth we’re after; Latin (or a hat) = our best guess from data. A statistic like \(\bar{x}\) estimates the parameter \(\mu\); a hat marks an estimate or prediction (\(\hat{y}\), \(\hat{\beta}\)).

Hypotheses and the decision rule

Inference is framed as a contest between two hypotheses:

  • Null hypothesis (\(H_0\)): no effect or difference.
  • Alternative hypothesis (\(H_1\)): an effect or difference exists.

The rule: reject \(H_0\) if the p-value \(< \alpha\); otherwise fail to reject \(H_0\) (we never “accept” it). A dependent variable (DV) is the outcome whose value depends on one or more independent variables (IVs).

Summation: the \(\Sigma\)

The most common symbol. \(\displaystyle\sum_{i=1}^{n} x_i\) means “add up the \(x\) values, from the first (\(i=1\)) to the \(n\)-th.” So the mean is a summation in disguise:

\[\bar{x} = \frac{1}{n}\sum_{i=1}^{n} x_i.\]

Drag the slider to watch the sum (and the mean) build term by term:

Once you see that \(\Sigma\) is “just a loop that adds,” the intimidating formulas (variance, the regression slope, the F-ratio) become readable: they’re sums of things, divided by other sums of things.

Normal random variables

We write “\(X\) is normally distributed with mean \(\mu\) and variance \(\sigma^2\)” as:

\[Y \sim N(\mu, \sigma^2)\]

reading piece by piece: \(Y\) the variable, \(\sim\) “is distributed as”, \(N\) normal, \(\mu\) the mean, \(\sigma^2\) the variance. (The binomial and Poisson use the same shape — \(X \sim \text{Binomial}(n,p)\), \(X \sim \text{Poisson}(\lambda)\).)

The general linear model in symbols

The skeleton behind most of the book:

\[Y = \beta_0 + \beta_1 X + \varepsilon\]

with \(Y\) the outcome (DV), \(\beta_0\) the intercept, \(\beta_1 X\) a coefficient times an independent variable, and \(\varepsilon\) the error. It’s the school-maths line \(y = b + mx\) wearing statistical clothes — \(\beta_0\) is the intercept \(b\), \(\beta_1\) is the slope \(m\). A larger \(|\beta_1|\) means that variable contributes more to \(Y\).

NoteA t-test is a linear model

An independent-samples t-test is this same equation with the group coded 0/1. Comparing apples and watermelons:

  • Apples (group 0): \(\;Y = \beta_0 + \beta_1 \cdot 0 + \varepsilon\;\) → the mean of the apples is \(\beta_0\).
  • Watermelons (group 1): \(\;Y = \beta_0 + \beta_1 \cdot 1 + \varepsilon\;\) → the mean of the watermelons is \(\beta_0 + \beta_1\).

So \(\beta_0\) is one group’s mean and \(\beta_1\) is the difference between the groups — testing “\(\beta_1 = 0\)” is exactly the t-test. This is why the whole book keeps returning to the general linear model: the familiar parametric tests are special cases of it.

A note on names: a General Linear Model (GLM) assumes the variables are normally distributed; a Generalized Linear Model relaxes that, allowing other outcome distributions (e.g. binomial for logistic regression).

See it in code

The notation maps directly onto one-word functions:

x = [4 7 3 8 5];
sum(x)      % ∑ xᵢ
mean(x)     % x̄
std(x)      % s   (divides by n − 1 by default)
var(x)      % s²
numel(x)    % n

The R and Python tabs run live; MATLAB is a static reference.

Quick quiz: population or sample?

Is each symbol a true population parameter, or a sample statistic we compute from data?

Answer each question — the feedback appears beneath it.

Watch the context-dependent letters: \(\alpha\) is the significance level here but a reliability coefficient elsewhere; \(\beta\) is a regression coefficient and the Type II error rate (power context); \(\varepsilon\) is the model error but the sphericity index in repeated-measures ANOVA. Don’t mix the population (Greek) and sample (Latin) symbols — writing \(\sigma\) when you mean the sample \(s\) quietly claims you know the truth. And don’t press \(X, Y, Z\) into service as arbitrary variable names — they’re reserved for coordinates.

Where this shows up next

This notation runs through every sheet. The general-linear-model equation is the backbone of the regression and ANOVA chapters; the population/sample split underlies Describing Data; and the bold, transpose, and \(A^{-1}\) conventions of linear algebra live in the Matrix Algebra appendix.