Skip to content
VibeFormer
56 min

Probability and Statistics, from the Ground Up

Written for someone relearning. Every term defined before use, the distribution-by-data-type table with plain-language meanings, the three tests your own paper runs with the arithmetic shown, and calibration as the bridge to a legal standard of proof.

Listen

0. Why this chapter first

Three reasons it comes before the machine learning chapters.

  1. Your own paper runs on it. Wilson intervals, an exact McNemar test, Fisher's exact test, and judge agreement rates. If anyone asks why Wilson rather than the ordinary formula, or what McNemar is doing, the answer has to come without hesitation — it is in your published work.
  2. Calibration is your best original bridge to law. A screen outputs a probability; the law demands sufficiently precise and consistent evidence. Connecting the two is a probability question, and it is the project nobody has done.
  3. Everything downstream is probability in different clothes. A loss function is a probability statement. Regularisation is a prior belief. A language model is a conditional distribution over words. Knowing that makes the later chapters shorter.

1. The foundations

TermDefinitionExample
OutcomeOne specific thing that could happenThis tender was rigged
Sample spaceThe set of everything that could happen{rigged, clean}
EventAny collection of outcomes you want to talk aboutThe tender was rigged *or* had fewer than three bidders
ProbabilityA number from 0 to 1 attached to an event. 0 means impossible, 1 means certainThe probability a randomly chosen tender is rigged is 0.005

Conditional probability is the probability of one thing *given that* you already know another. Written as the probability of A given B. It is the probability of both happening, divided by the probability of B.

P(A∣B)=P(A∩B)P(B)P(A \mid B) = \frac{P(A \cap B)}{P(B)}
The probability of A given B equals the probability of A and B both happening, divided by the probability of B.The vertical bar means *given*. The upside-down U means *and*. Dividing by the probability of B is the whole trick: you have thrown away every case where B did not happen, so you rescale what is left so it still totals one.

Independence means knowing one thing tells you nothing about the other — the conditional probability equals the unconditional one. Conditional independence means that holds once you account for some third thing, and it is the idea that makes causal graphs work.

RuleWhat it saysWhere it appears
Chain ruleThe probability of a sequence is the product of each step's probability given everything before itThis is literally how a language model assigns a probability to a sentence
Law of total probabilityTo get the overall probability of something, work it out separately in each situation and average, weighting by how likely each situation isOverall flag rate equals the flag rate in construction times the share of construction, plus the same for every other sector
Bayes' theoremFlips a conditional round. Given the chance of seeing this evidence if the firm is guilty, plus how common guilt is, you get the chance of guilt given the evidenceThe single most important idea in this chapter for your purposes
P(H∣E)=P(E∣H)⏞likelihood⋅P(H)⏞priorP(E)⏟how common the evidence isP(H \mid E) = \frac{\overbrace{P(E \mid H)}^{\text{likelihood}} \cdot \overbrace{P(H)}^{\text{prior}}}{\underbrace{P(E)}_{\text{how common the evidence is}}}
The probability of the hypothesis given the evidence equals the probability of the evidence given the hypothesis, times the prior probability of the hypothesis, divided by the overall probability of the evidence.This is the most important formula on the site for your purposes. H is the hypothesis — the firm colluded. E is the evidence — the screen fired. A screen gives you the top-left term, the probability of this evidence under honest competition. A court needs the left-hand side. The bridge between them is the prior, and that is the number no agency states.
P(x1,x2,…,xn)=∏t=1nP(xt∣x1,…,xt−1)P(x_1, x_2, \ldots, x_n) = \prod_{t=1}^{n} P(x_t \mid x_1, \ldots, x_{t-1})
The probability of a whole sequence equals the product of the probability of each item given everything that came before it.The chain rule. The large pi means *multiply all of these together*. Worth memorising, because this exact expression reappears in the language-model chapter as the thing a large language model computes — it is not an analogy, it is the same equation.

2. Random variables, and how to describe them

A random variable is just a number whose value depends on chance. Discrete means it takes separate values you could list — 0, 1, 2 bidders. Continuous means it can take any value in a range — a bid of 1,960,432.17 euro.

ObjectDiscrete versionContinuous version
How likely each value isProbability mass function. The probability that the variable equals exactly this value. Probability of exactly 3 bidders is 0.21Probability density function. Not a probability — a density. The probability of a bid being exactly 1,960,432.17 is zero. You ask about ranges instead: between 1.9 and 2.0 million
Probability of being at most xCumulative distribution function. A staircase that steps up at each possible valueCumulative distribution function. A smooth rising curve. Its slope is the density
Reverse lookupThe smallest value whose cumulative probability reaches pThe quantile. The 95th percentile bid is the value 95 percent of bids fall below

Expectation, or the mean, is the average you would get over endless repetitions — each value weighted by how likely it is. Variance is the average squared distance from that mean, so it measures spread. Standard deviation is its square root, which puts it back in the original units.

E[X]=∑ixi P(X=xi)and for a samplexˉ=1n∑i=1nxi\mathbb{E}[X] = \sum_{i} x_i \, P(X = x_i) \qquad\text{and for a sample}\qquad \bar{x} = \frac{1}{n}\sum_{i=1}^{n} x_i
The expectation of X equals the sum, over every possible value, of that value times its probability. For a sample, the mean is simply the total divided by the number of observations.The two are the same operation. A plain average is the weighted version where every observation happens to carry weight one over n. The E with a double stroke is standard notation for *expected value*.
Var(X)=E[(X−E[X])2]σ=Var(X)\mathrm{Var}(X) = \mathbb{E}\big[(X - \mathbb{E}[X])^2\big] \qquad\qquad \sigma = \sqrt{\mathrm{Var}(X)}
The variance of X is the expected squared distance from its own mean. The standard deviation, sigma, is the square root of the variance.Squaring does two things: it stops positive and negative deviations cancelling to zero, and it makes large deviations count disproportionately. That second effect is why one motorway contract can dominate the variance of 10,000 tenders.

Two further shape descriptions: skewness measures lopsidedness, and kurtosis measures how heavy the tails are — how often you get extreme values.

E[aX+bY]=a E[X]+b E[Y]alwaysE[g(X)]≠g(E[X])in general\mathbb{E}[aX + bY] = a\,\mathbb{E}[X] + b\,\mathbb{E}[Y] \quad\text{always} \qquad\qquad \mathbb{E}[g(X)] \neq g(\mathbb{E}[X]) \quad\text{in general}
Expectation always passes through sums and constant multiples. But the expectation of a function of X is not, in general, that function of the expectation of X.The left identity needs no assumptions at all — not even independence. The right-hand warning is Jensen's inequality, and the averaging-ratios example below is exactly it: the average of the ratios is not the ratio of the averages.

3. Which distribution for which kind of data

A distribution is a named mathematical shape describing how likely each value is. Choosing one is a claim about the mechanism that produced your data, not a convenience. The question to ask is always: what am I measuring, and what generated it?

Your data looks likeDistributionWhat it means, and what to watch
A single yes-or-no outcomeBernoulliOne parameter: the chance of yes. Variance is largest at a half, meaning a coin flip is the most uncertain case. One tender, flagged or not
How many yeses out of a fixed number of triesBinomialA sum of Bernoullis. Flagged or not across 15 runs. Mean is tries times chance
How many tries until the first yesGeometricMemoryless: having waited ten tries tells you nothing about the next one
Counts of events in a fixed window, arriving independently at a steady ratePoissonOne parameter, the rate. Mean equals variance — which is the assumption to test, not assume. Bidders per tender, complaints per month
Counts where the spread is bigger than the meanNegative binomialPoisson with a varying rate. Use this whenever count data is more spread out than Poisson allows, which real count data usually is
Time until something happens, at constant riskExponentialMemoryless, the continuous twin of geometric
Time until something happens, where risk rises or falls with ageWeibullA shape parameter says whether failure gets more or less likely over time. Standard for time-to-detection
A total of several waiting timesGammaGeneralises exponential
A quantity built by adding many small independent effectsNormal, the bell curveSymmetric, light-tailed, fully described by mean and variance. Justified by the central limit theorem, not by habit
A quantity built by multiplying many effects, always positive, right-skewedLog-normalIts logarithm is normal. Contract values, firm sizes, incomes. Model the log
A proportion or a probability, between 0 and 1BetaThe natural distribution for a flag rate. Two shape parameters give it very flexible form
One of several unordered categoriesCategorical, and Multinomial for countsGeneralises Bernoulli and binomial to more than two outcomes
A whole set of probabilities that must total 1DirichletA distribution over distributions
A quantity where the top few dwarf everything elsePareto, or Zipf for ranksA power law. Tails so heavy the mean may not exist at all. Market shares, platform traffic, city sizes
Equally likely anywhere in a rangeUniformThe honest choice when you know only the range
An average when you have few observations and do not know the true spreadStudent's tHeavier tails than normal, becoming normal as the sample grows
A sum of squared normals, or a varianceChi-squaredUnderlies variance tests and goodness-of-fit tests
A ratio of two variancesFUnderlies analysis of variance and model comparison
The leading digits of naturally occurring numbersBenford1 leads about 30 percent of the time, 9 about 4.6 percent. Deviation suggests fabricated figures, which is why it appears in bid-rigging screens

4. Two variables at once

TermDefinition
Joint distributionHow likely each *combination* of two or more values is
MarginalisingAdding up across one variable to get rid of it, leaving the distribution of the other
ConditioningFixing one variable at a known value and rescaling the rest so they still total 1
CovarianceWhether two variables tend to move together. Positive means both rise together. Measured in the product of their units, so the number itself is hard to read
CorrelationCovariance rescaled to sit between minus 1 and plus 1, so it is unitless and comparable

Pearson correlation measures straight-line association only. Spearman correlates the ranks instead, so it catches any consistently-increasing relationship even a curved one.

5. The two limit theorems

The law of large numbers says that as your sample grows, the sample average closes in on the true average. It tells you estimation eventually works.

The central limit theorem says something stronger and stranger: the distribution of a sample *average* approaches a bell curve, whatever shape the original data had, provided the variance is finite. It tells you what your uncertainty looks like, which is what lets you build an interval around an estimate.

xˉ  ≈  N ⁣(μ,  σ2n)soSE=σn\bar{x} \;\approx\; \mathcal{N}\!\left(\mu, \; \frac{\sigma^2}{n}\right) \qquad\text{so}\qquad \text{SE} = \frac{\sigma}{\sqrt{n}}
For a large enough sample, the sample mean behaves like a normal distribution centred on the true mean, with variance equal to the population variance divided by n. So the standard error is the standard deviation divided by the square root of n.The script N means *normal distribution*, written as centre and spread. The practical content is the square root: to halve your uncertainty you need four times the data. That is also why 15 runs is a hard ceiling — the square root of 15 is under 4.

6. Estimation

An estimator is a recipe for guessing an unknown quantity from data — for instance, use the sample average to estimate the true average. Estimators are judged on two things.

PropertyDefinitionPlain version
BiasThe difference between the estimator's average value and the truthDoes it systematically aim off-target?
VarianceHow much the estimate jumps around between samplesIs it reliable, or does it swing wildly?

Total error decomposes exactly into squared bias plus variance — which is the same decomposition that reappears as the bias–variance tradeoff in machine learning. They are the same idea.

MethodWhat it does
Maximum likelihoodPick the parameter value that makes the data you actually saw as probable as possible. Almost every loss function in machine learning is this in disguise
Maximum a posterioriThe same, but also weighted by what you believed beforehand. Ridge regression is this with a bell-curve prior; lasso is this with a different prior. Regularisation is a prior belief

7. Confidence intervals, and why yours are Wilson

A 95 percent confidence interval is a range built by a procedure that, across repeated samples, contains the true value 95 percent of the time.

MethodHow it worksHow it behaves
Wald, the normal approximationEstimate plus or minus roughly two standard errors. The formula everyone is taughtSimple, and badly behaved at small samples or near 0 and 1. It can produce bounds below 0 or above 1, and at an observed 0 of 5 or 5 of 5 it collapses to zero width — claiming perfect certainty from five observations
Wilson score intervalWorks backwards from a test rather than plugging in an estimated standard errorStays inside 0 to 1, keeps sensible width at the extremes, and has far better coverage at small samples. The standard recommendation over Wald
Wald:p^  ±  zp^(1−p^)n\text{Wald:}\quad \hat{p} \;\pm\; z\sqrt{\frac{\hat{p}(1 - \hat{p})}{n}}
The Wald interval is the observed proportion, plus or minus z times the square root of the observed proportion times one minus it, all divided by n.Here z is 1.96 for 95 percent confidence. Now put in an observed 5 out of 5: p-hat is 1, so one minus p-hat is 0, the whole square root is zero, and the interval is the single point 1.0. Five observations, and the formula claims certainty. That is the failure, visible in the algebra.
Wilson:p^+z22n  ±  zp^(1−p^)n+z24n21+z2n\text{Wilson:}\quad \frac{\hat{p} + \dfrac{z^2}{2n} \;\pm\; z\sqrt{\dfrac{\hat{p}(1-\hat{p})}{n} + \dfrac{z^2}{4n^2}}}{1 + \dfrac{z^2}{n}}
The Wilson interval adds z squared over two n to the observed proportion, adds a further z squared over four n squared inside the square root, and divides the whole thing by one plus z squared over n.You will not compute this by hand and nobody expects you to. What matters is why the extra terms exist: they pull the centre away from the boundary and keep the width non-zero there, so 5 of 5 returns roughly 57 to 100 percent instead of a false point estimate. Those added terms behave like a few imaginary extra observations.

8. Hypothesis testing, and the three tests in your paper

TermDefinition
Null hypothesisThe boring explanation you are trying to rule out — usually that nothing is going on
Test statisticA number computed from the data that would be large if something *is* going on
p-valueThe probability of seeing a result at least this extreme if the null hypothesis were true
Type I errorRejecting a true null. Crying wolf. Controlled by the significance level you set
Type II errorFailing to reject a false null. Missing a real effect
PowerOne minus the Type II error rate. The chance of detecting a real effect if there is one
TestWhen to use itWhy it was right for your design
McNemarPaired yes/no outcomes — the same unit measured under two conditions. It looks only at the cases that changedEvery scenario ran under all three conditions, so each scenario is its own control. That pairing is what gives you any power at all at 15 runs
Fisher's exactA small two-by-two table where the usual approximation failsComparing one model against another on five runs each
Chi-squaredLarger tables, goodness of fitNot usable here — the expected counts per cell are far too small

The paired table McNemar actually looks at

McNemar throws away the diagonal. Cells a and d are the runs that behaved the same under both conditions, and they say nothing about whether the condition mattered. Only the off-diagonal disagreements, b and c, enter the test. This is why a paired design extracts a real result from 15 runs: the test is about 9 discordant pairs, not 15 independent observations.
χMcNemar2=(b−c)2b+cor exactlyp=2∑k=0min⁡(b,c)(b+ck)(12)b+c\chi^2_{\text{McNemar}} = \frac{(b - c)^2}{b + c} \qquad\text{or exactly}\qquad p = 2\sum_{k=0}^{\min(b,c)} \binom{b+c}{k} \left(\tfrac{1}{2}\right)^{b+c}
The McNemar statistic is the squared difference between the two discordant counts, divided by their total. The exact version instead sums binomial probabilities at one half.Use the exact form at your sample sizes; the chi-squared approximation needs b plus c of at least about 25 and you have 9. The exact form is just a two-sided binomial test asking whether 9 changes in the same direction could be coin flips.

Multiple comparisons: testing many hypotheses inflates the chance of a false alarm somewhere. Bonferroni divides your threshold by the number of tests, which is safe but very conservative. Benjamini–Hochberg instead controls the expected *proportion* of your discoveries that are false, which is usually the more sensible target.

9. Bayesian and frequentist

FrequentistBayesian
A parameter isA fixed unknown numberItself uncertain, with a distribution over possible values
Probability meansLong-run frequency over repetitionsDegree of belief given what you know
You start withNothing but the dataA prior — your belief before seeing the data
You end withAn estimate, a confidence interval, a p-valueA posterior — your updated belief, and a credible interval
StrengthNo prior to defend, well-developed error controlNatural handling of uncertainty, works at small samples, composes cleanly

10. Calibration — your bridge to the standard of proof

A model is calibrated if, among all the cases it assigns 70 percent to, about 70 percent really turn out positive. Calibration asks whether the numbers mean what they say.

Discrimination is a separate property: whether the model puts positives above negatives in rank order. The two are independent. A model can rank perfectly while reporting nonsense probabilities, or report perfectly honest probabilities while ranking no better than chance.

MeasureWhat it does
Reliability diagramGroup predictions into bins, then plot what actually happened against what was predicted. Perfect calibration traces the diagonal. This is the picture to show a lawyer
Brier scoreAverage squared error of the probabilities. Splits into how calibrated you were and how much you discriminated
Expected calibration errorThe average gap between confidence and reality across bins. One number, but it hides where the problem sits
Log lossPunishes confident mistakes very heavily
Temperature scaling, Platt scaling, isotonic regressionFixes applied afterwards. Fit a correction on held-out data. Temperature scaling is the standard remedy for overconfident neural networks
Brier=1n∑i=1n(p^i−yi)2\text{Brier} = \frac{1}{n}\sum_{i=1}^{n}\left(\hat{p}_i - y_i\right)^2
The Brier score is the average squared gap between the probability you announced and what actually happened, where what happened is coded as one or zero.Lower is better; zero is perfect. Say you announce 0.9 on a tender that was rigged: the contribution is 0.01. Announce 0.9 on a clean one and it is 0.81 — eighty-one times worse. It is mean squared error applied to probabilities.
ECE=∑m=1M∣Bm∣n∣acc(Bm)⏟what happened−conf(Bm)⏟what you claimed∣\text{ECE} = \sum_{m=1}^{M} \frac{|B_m|}{n} \Big| \underbrace{\text{acc}(B_m)}_{\text{what happened}} - \underbrace{\text{conf}(B_m)}_{\text{what you claimed}} \Big|
Expected calibration error splits predictions into bins, takes the absolute gap between actual outcome rate and average claimed confidence in each bin, and averages those gaps weighted by how many predictions fall in each bin.B-m is the m-th bin, and the bars around it mean *how many predictions are in it*. In the overconfident example below, the 0.70-to-0.80 bin claimed 0.75 and delivered 0.20, so it contributes a gap of 0.55 weighted by 200 over n. **The weakness: a single number cannot tell you *where* the model is miscalibrated**, and for a legal argument the location is the point — overconfidence at the high end is what triggers a raid.

11. Causal inference, which is your differentiator

Correlation describes data you observed. Causation is a claim about what would happen if you intervened, which is a different and stronger thing. A directed acyclic graph draws your causal assumptions as arrows between variables — acyclic meaning no loops.

StructureShapeWhat it doesWhat to do
ConfounderOne cause pointing into both your variablesCreates a fake association between themAdjust for it. Otherwise you report something that is not there
MediatorSits on the path from cause to effectCarries part of the real effectDo not adjust for it, if you want the total effect. Adjusting removes part of the answer
ColliderBoth your variables point into itConditioning on it invents an association that does not existDo not adjust for it. Adjusting here creates bias rather than removing it

The backdoor criterion is the rule for which variables to adjust for: block every path that runs backwards into your cause, and never include anything downstream of it. Do-calculus is the formal system for deciding whether a causal question can be answered from observational data at all — and sometimes the answer is simply no.

DesignIdeaWhen you can use it
Randomised experimentAssign treatment by coin flip, so nothing can confound itThe benchmark. Rarely available in enforcement
Difference-in-differencesCompare the change over time in a treated group against the change in an untreated oneWhen something affected some units and not others at a known moment
Regression discontinuityCompare units just above a cutoff with units just below, since they are otherwise alikeThe right design for a flag threshold. Firms scoring 0.699 and 0.701 are near-identical, but only one got investigated
Instrumental variablesFind something that shifts treatment but could not affect the outcome any other wayWhen confounding is unmeasured but a valid instrument exists
Bayesian structural time seriesBuild a synthetic counterfactual from trend, seasonality and related series, then read off the gap after an interventionWhat you used on 22 years of crime and census data

12. How this connects to ATLANTIS

Both legal strands reduce to statistical questions once you push on them, and neither legal-track PhD will be equipped to answer them.

StrandThe statistical question underneath it
Accuracy, the data problemDoes the agency's data support the inference drawn from it? That is selection, measurement validity, and whether the sample says anything about the population. The dark figure problem is collider bias
FairnessDoes a flag mean what it claims? That is calibration. Is the selection reproducible or an artefact of the sample? That is variance. Do error rates differ across groups? That is the impossibility result
Institutional arrangementsWhat can an agency conclude from a pilot of thirty cases? That is inference at small samples, which is exactly the regime your own paper operates in

13. What the panel brings to this chapter

Panel memberWhere statistics meets their work
Thibault SchrepelComplexity science: non-linear dynamics, emergence, power laws rather than bell curves, systems that evolve rather than settle. He operationalised it by scoring eight Digital Acts against fourteen adaptivity criteria — a coding study with an explicit instrument, which is empirical social science rather than doctrine
Catalina GoantaLarge-scale empirical measurement. A multi-country longitudinal study across hundreds of creators and around a million posts raises every question in this chapter: sampling frame, measurement error, cross-country comparability, inter-coder reliability
Georgiana MirzaDigital ecosystems and market structure, where concentration measures and heavy-tailed share distributions are the basic descriptive tools
Tijmen WismanProportionality under Article 8 ECHR, which is at bottom a judgment about error rates and who bears them — how many innocent people a detection system may burden, and how badly

14. What is unexplored, and five projects

15. Your CV, mapped onto this chapter

What you haveWhere it landsWhy it fits
Statistical inference, probability and linear algebra at MSc levelAll five projectsThis is formal training the two legal-track PhDs will not have, and the vacancy says their empirical skills will be taught. Yours do not need to be
Wilson intervals, exact McNemar, Fisher's exact, and judge-agreement reporting in a published paperProject 3You have defended small-sample inference in print, including declining the comparisons the design could not support. The restraint is the credible part
Causal inference — DAGs, belief networks, do-calculus, structural time seriesProjects 1 and 5Collider bias is the dark-figure problem, and you have already separated a true signal from a collection artefact across 22 years of administrative data
Bayesian methods, in the crimes-against-women paperProjects 2 and 5A standard of proof attaches to a posterior, not a likelihood, so the legal question is natively Bayesian
Time series — ARIMA, LSTM, structural time seriesProjects 1 and 5Detection rates and market data are both temporal, and the interesting questions concern change around interventions
Model evaluation — calibration, cross-validation, hyperparameter searchProject 2Calibration is a measurement problem with mature tooling and no legal literature at all

16. Where this actually shows up in the interview

If they askReach for
Your sample is 15 runs. What can you honestly conclude?The paired design. Each scenario is its own control, McNemar uses only the cases that changed, pooled claims hold at p of 0.004 and 0.00006, and model-versus-model claims are not supported at all — Fisher gives 0.206
Why Wilson intervals?Boundary behaviour. The ordinary formula collapses to zero width at 0 of 5 and 5 of 5, asserting a certainty the data cannot support
How would you audit a screening tool?Calibration first, then flag-rate disparity across firm size and sector and country, then sensitivity to design choices, then the multiple-comparison problem at ten thousand tenders
Why should a lawyer care about calibration?Because a standard of proof is a statement about sufficient confidence, and an uncalibrated probability does not denote confidence. The number cannot be mapped onto the standard
What is wrong with training on past enforcement decisions?Detection is a collider. You are conditioning on where the agency looked, so you learn enforcement history rather than collusion
Explain something technical to a non-technical panelThe Bayes point with the numbers. A 90-percent-accurate flag at a one-in-200 base rate still leaves a 92 percent chance the firm did nothing wrong

17. If you remember twelve things

  1. Bayes flips a conditional, and it needs a base rate. Without one, a flag overstates guilt.
  2. Expectation is always linear. Variance adds only when uncorrelated. Nothing else is linear — Jensen.
  3. Choosing a distribution is a claim about mechanism: adding gives normal, multiplying gives log-normal, counting rare events gives Poisson, reinforcement gives a power law.
  4. Overdispersed counts need negative binomial, not Poisson. Check whether variance equals the mean.
  5. Beta is the natural distribution for a rate.
  6. Zero correlation does not imply independence.
  7. The central limit theorem needs finite variance and large samples. Neither holds at 15 runs or on power-law data.
  8. Total error equals squared bias plus variance, which is the same decomposition as the bias–variance tradeoff.
  9. Most loss functions are likelihoods; most regularisers are priors.
  10. Wilson over Wald at small samples and near the boundaries.
  11. Calibration and discrimination are independent, and calibration is the one a standard of proof needs.
  12. Detection is a collider. Conditioning on it is the dark-figure problem.