Probability and Statistics, from the Ground Up
Written for someone relearning. Every term defined before use, the distribution-by-data-type table with plain-language meanings, the three tests your own paper runs with the arithmetic shown, and calibration as the bridge to a legal standard of proof.
0. Why this chapter first
Three reasons it comes before the machine learning chapters.
- Your own paper runs on it. Wilson intervals, an exact McNemar test, Fisher's exact test, and judge agreement rates. If anyone asks why Wilson rather than the ordinary formula, or what McNemar is doing, the answer has to come without hesitation — it is in your published work.
- Calibration is your best original bridge to law. A screen outputs a probability; the law demands sufficiently precise and consistent evidence. Connecting the two is a probability question, and it is the project nobody has done.
- Everything downstream is probability in different clothes. A loss function is a probability statement. Regularisation is a prior belief. A language model is a conditional distribution over words. Knowing that makes the later chapters shorter.
1. The foundations
| Term | Definition | Example |
|---|---|---|
| Outcome | One specific thing that could happen | This tender was rigged |
| Sample space | The set of everything that could happen | {rigged, clean} |
| Event | Any collection of outcomes you want to talk about | The tender was rigged *or* had fewer than three bidders |
| Probability | A number from 0 to 1 attached to an event. 0 means impossible, 1 means certain | The probability a randomly chosen tender is rigged is 0.005 |
Conditional probability is the probability of one thing *given that* you already know another. Written as the probability of A given B. It is the probability of both happening, divided by the probability of B.
Independence means knowing one thing tells you nothing about the other — the conditional probability equals the unconditional one. Conditional independence means that holds once you account for some third thing, and it is the idea that makes causal graphs work.
| Rule | What it says | Where it appears |
|---|---|---|
| Chain rule | The probability of a sequence is the product of each step's probability given everything before it | This is literally how a language model assigns a probability to a sentence |
| Law of total probability | To get the overall probability of something, work it out separately in each situation and average, weighting by how likely each situation is | Overall flag rate equals the flag rate in construction times the share of construction, plus the same for every other sector |
| Bayes' theorem | Flips a conditional round. Given the chance of seeing this evidence if the firm is guilty, plus how common guilt is, you get the chance of guilt given the evidence | The single most important idea in this chapter for your purposes |
2. Random variables, and how to describe them
A random variable is just a number whose value depends on chance. Discrete means it takes separate values you could list — 0, 1, 2 bidders. Continuous means it can take any value in a range — a bid of 1,960,432.17 euro.
| Object | Discrete version | Continuous version |
|---|---|---|
| How likely each value is | Probability mass function. The probability that the variable equals exactly this value. Probability of exactly 3 bidders is 0.21 | Probability density function. Not a probability — a density. The probability of a bid being exactly 1,960,432.17 is zero. You ask about ranges instead: between 1.9 and 2.0 million |
| Probability of being at most x | Cumulative distribution function. A staircase that steps up at each possible value | Cumulative distribution function. A smooth rising curve. Its slope is the density |
| Reverse lookup | The smallest value whose cumulative probability reaches p | The quantile. The 95th percentile bid is the value 95 percent of bids fall below |
Expectation, or the mean, is the average you would get over endless repetitions — each value weighted by how likely it is. Variance is the average squared distance from that mean, so it measures spread. Standard deviation is its square root, which puts it back in the original units.
Two further shape descriptions: skewness measures lopsidedness, and kurtosis measures how heavy the tails are — how often you get extreme values.
3. Which distribution for which kind of data
A distribution is a named mathematical shape describing how likely each value is. Choosing one is a claim about the mechanism that produced your data, not a convenience. The question to ask is always: what am I measuring, and what generated it?
| Your data looks like | Distribution | What it means, and what to watch |
|---|---|---|
| A single yes-or-no outcome | Bernoulli | One parameter: the chance of yes. Variance is largest at a half, meaning a coin flip is the most uncertain case. One tender, flagged or not |
| How many yeses out of a fixed number of tries | Binomial | A sum of Bernoullis. Flagged or not across 15 runs. Mean is tries times chance |
| How many tries until the first yes | Geometric | Memoryless: having waited ten tries tells you nothing about the next one |
| Counts of events in a fixed window, arriving independently at a steady rate | Poisson | One parameter, the rate. Mean equals variance — which is the assumption to test, not assume. Bidders per tender, complaints per month |
| Counts where the spread is bigger than the mean | Negative binomial | Poisson with a varying rate. Use this whenever count data is more spread out than Poisson allows, which real count data usually is |
| Time until something happens, at constant risk | Exponential | Memoryless, the continuous twin of geometric |
| Time until something happens, where risk rises or falls with age | Weibull | A shape parameter says whether failure gets more or less likely over time. Standard for time-to-detection |
| A total of several waiting times | Gamma | Generalises exponential |
| A quantity built by adding many small independent effects | Normal, the bell curve | Symmetric, light-tailed, fully described by mean and variance. Justified by the central limit theorem, not by habit |
| A quantity built by multiplying many effects, always positive, right-skewed | Log-normal | Its logarithm is normal. Contract values, firm sizes, incomes. Model the log |
| A proportion or a probability, between 0 and 1 | Beta | The natural distribution for a flag rate. Two shape parameters give it very flexible form |
| One of several unordered categories | Categorical, and Multinomial for counts | Generalises Bernoulli and binomial to more than two outcomes |
| A whole set of probabilities that must total 1 | Dirichlet | A distribution over distributions |
| A quantity where the top few dwarf everything else | Pareto, or Zipf for ranks | A power law. Tails so heavy the mean may not exist at all. Market shares, platform traffic, city sizes |
| Equally likely anywhere in a range | Uniform | The honest choice when you know only the range |
| An average when you have few observations and do not know the true spread | Student's t | Heavier tails than normal, becoming normal as the sample grows |
| A sum of squared normals, or a variance | Chi-squared | Underlies variance tests and goodness-of-fit tests |
| A ratio of two variances | F | Underlies analysis of variance and model comparison |
| The leading digits of naturally occurring numbers | Benford | 1 leads about 30 percent of the time, 9 about 4.6 percent. Deviation suggests fabricated figures, which is why it appears in bid-rigging screens |
4. Two variables at once
| Term | Definition |
|---|---|
| Joint distribution | How likely each *combination* of two or more values is |
| Marginalising | Adding up across one variable to get rid of it, leaving the distribution of the other |
| Conditioning | Fixing one variable at a known value and rescaling the rest so they still total 1 |
| Covariance | Whether two variables tend to move together. Positive means both rise together. Measured in the product of their units, so the number itself is hard to read |
| Correlation | Covariance rescaled to sit between minus 1 and plus 1, so it is unitless and comparable |
Pearson correlation measures straight-line association only. Spearman correlates the ranks instead, so it catches any consistently-increasing relationship even a curved one.
5. The two limit theorems
The law of large numbers says that as your sample grows, the sample average closes in on the true average. It tells you estimation eventually works.
The central limit theorem says something stronger and stranger: the distribution of a sample *average* approaches a bell curve, whatever shape the original data had, provided the variance is finite. It tells you what your uncertainty looks like, which is what lets you build an interval around an estimate.
6. Estimation
An estimator is a recipe for guessing an unknown quantity from data — for instance, use the sample average to estimate the true average. Estimators are judged on two things.
| Property | Definition | Plain version |
|---|---|---|
| Bias | The difference between the estimator's average value and the truth | Does it systematically aim off-target? |
| Variance | How much the estimate jumps around between samples | Is it reliable, or does it swing wildly? |
Total error decomposes exactly into squared bias plus variance — which is the same decomposition that reappears as the bias–variance tradeoff in machine learning. They are the same idea.
| Method | What it does |
|---|---|
| Maximum likelihood | Pick the parameter value that makes the data you actually saw as probable as possible. Almost every loss function in machine learning is this in disguise |
| Maximum a posteriori | The same, but also weighted by what you believed beforehand. Ridge regression is this with a bell-curve prior; lasso is this with a different prior. Regularisation is a prior belief |
7. Confidence intervals, and why yours are Wilson
A 95 percent confidence interval is a range built by a procedure that, across repeated samples, contains the true value 95 percent of the time.
| Method | How it works | How it behaves |
|---|---|---|
| Wald, the normal approximation | Estimate plus or minus roughly two standard errors. The formula everyone is taught | Simple, and badly behaved at small samples or near 0 and 1. It can produce bounds below 0 or above 1, and at an observed 0 of 5 or 5 of 5 it collapses to zero width — claiming perfect certainty from five observations |
| Wilson score interval | Works backwards from a test rather than plugging in an estimated standard error | Stays inside 0 to 1, keeps sensible width at the extremes, and has far better coverage at small samples. The standard recommendation over Wald |
8. Hypothesis testing, and the three tests in your paper
| Term | Definition |
|---|---|
| Null hypothesis | The boring explanation you are trying to rule out — usually that nothing is going on |
| Test statistic | A number computed from the data that would be large if something *is* going on |
| p-value | The probability of seeing a result at least this extreme if the null hypothesis were true |
| Type I error | Rejecting a true null. Crying wolf. Controlled by the significance level you set |
| Type II error | Failing to reject a false null. Missing a real effect |
| Power | One minus the Type II error rate. The chance of detecting a real effect if there is one |
| Test | When to use it | Why it was right for your design |
|---|---|---|
| McNemar | Paired yes/no outcomes — the same unit measured under two conditions. It looks only at the cases that changed | Every scenario ran under all three conditions, so each scenario is its own control. That pairing is what gives you any power at all at 15 runs |
| Fisher's exact | A small two-by-two table where the usual approximation fails | Comparing one model against another on five runs each |
| Chi-squared | Larger tables, goodness of fit | Not usable here — the expected counts per cell are far too small |
The paired table McNemar actually looks at
Multiple comparisons: testing many hypotheses inflates the chance of a false alarm somewhere. Bonferroni divides your threshold by the number of tests, which is safe but very conservative. Benjamini–Hochberg instead controls the expected *proportion* of your discoveries that are false, which is usually the more sensible target.
9. Bayesian and frequentist
| Frequentist | Bayesian | |
|---|---|---|
| A parameter is | A fixed unknown number | Itself uncertain, with a distribution over possible values |
| Probability means | Long-run frequency over repetitions | Degree of belief given what you know |
| You start with | Nothing but the data | A prior — your belief before seeing the data |
| You end with | An estimate, a confidence interval, a p-value | A posterior — your updated belief, and a credible interval |
| Strength | No prior to defend, well-developed error control | Natural handling of uncertainty, works at small samples, composes cleanly |
10. Calibration — your bridge to the standard of proof
A model is calibrated if, among all the cases it assigns 70 percent to, about 70 percent really turn out positive. Calibration asks whether the numbers mean what they say.
Discrimination is a separate property: whether the model puts positives above negatives in rank order. The two are independent. A model can rank perfectly while reporting nonsense probabilities, or report perfectly honest probabilities while ranking no better than chance.
| Measure | What it does |
|---|---|
| Reliability diagram | Group predictions into bins, then plot what actually happened against what was predicted. Perfect calibration traces the diagonal. This is the picture to show a lawyer |
| Brier score | Average squared error of the probabilities. Splits into how calibrated you were and how much you discriminated |
| Expected calibration error | The average gap between confidence and reality across bins. One number, but it hides where the problem sits |
| Log loss | Punishes confident mistakes very heavily |
| Temperature scaling, Platt scaling, isotonic regression | Fixes applied afterwards. Fit a correction on held-out data. Temperature scaling is the standard remedy for overconfident neural networks |
11. Causal inference, which is your differentiator
Correlation describes data you observed. Causation is a claim about what would happen if you intervened, which is a different and stronger thing. A directed acyclic graph draws your causal assumptions as arrows between variables — acyclic meaning no loops.
| Structure | Shape | What it does | What to do |
|---|---|---|---|
| Confounder | One cause pointing into both your variables | Creates a fake association between them | Adjust for it. Otherwise you report something that is not there |
| Mediator | Sits on the path from cause to effect | Carries part of the real effect | Do not adjust for it, if you want the total effect. Adjusting removes part of the answer |
| Collider | Both your variables point into it | Conditioning on it invents an association that does not exist | Do not adjust for it. Adjusting here creates bias rather than removing it |
The backdoor criterion is the rule for which variables to adjust for: block every path that runs backwards into your cause, and never include anything downstream of it. Do-calculus is the formal system for deciding whether a causal question can be answered from observational data at all — and sometimes the answer is simply no.
| Design | Idea | When you can use it |
|---|---|---|
| Randomised experiment | Assign treatment by coin flip, so nothing can confound it | The benchmark. Rarely available in enforcement |
| Difference-in-differences | Compare the change over time in a treated group against the change in an untreated one | When something affected some units and not others at a known moment |
| Regression discontinuity | Compare units just above a cutoff with units just below, since they are otherwise alike | The right design for a flag threshold. Firms scoring 0.699 and 0.701 are near-identical, but only one got investigated |
| Instrumental variables | Find something that shifts treatment but could not affect the outcome any other way | When confounding is unmeasured but a valid instrument exists |
| Bayesian structural time series | Build a synthetic counterfactual from trend, seasonality and related series, then read off the gap after an intervention | What you used on 22 years of crime and census data |
12. How this connects to ATLANTIS
Both legal strands reduce to statistical questions once you push on them, and neither legal-track PhD will be equipped to answer them.
| Strand | The statistical question underneath it |
|---|---|
| Accuracy, the data problem | Does the agency's data support the inference drawn from it? That is selection, measurement validity, and whether the sample says anything about the population. The dark figure problem is collider bias |
| Fairness | Does a flag mean what it claims? That is calibration. Is the selection reproducible or an artefact of the sample? That is variance. Do error rates differ across groups? That is the impossibility result |
| Institutional arrangements | What can an agency conclude from a pilot of thirty cases? That is inference at small samples, which is exactly the regime your own paper operates in |
13. What the panel brings to this chapter
| Panel member | Where statistics meets their work |
|---|---|
| Thibault Schrepel | Complexity science: non-linear dynamics, emergence, power laws rather than bell curves, systems that evolve rather than settle. He operationalised it by scoring eight Digital Acts against fourteen adaptivity criteria — a coding study with an explicit instrument, which is empirical social science rather than doctrine |
| Catalina Goanta | Large-scale empirical measurement. A multi-country longitudinal study across hundreds of creators and around a million posts raises every question in this chapter: sampling frame, measurement error, cross-country comparability, inter-coder reliability |
| Georgiana Mirza | Digital ecosystems and market structure, where concentration measures and heavy-tailed share distributions are the basic descriptive tools |
| Tijmen Wisman | Proportionality under Article 8 ECHR, which is at bottom a judgment about error rates and who bears them — how many innocent people a detection system may burden, and how badly |
14. What is unexplored, and five projects
15. Your CV, mapped onto this chapter
| What you have | Where it lands | Why it fits |
|---|---|---|
| Statistical inference, probability and linear algebra at MSc level | All five projects | This is formal training the two legal-track PhDs will not have, and the vacancy says their empirical skills will be taught. Yours do not need to be |
| Wilson intervals, exact McNemar, Fisher's exact, and judge-agreement reporting in a published paper | Project 3 | You have defended small-sample inference in print, including declining the comparisons the design could not support. The restraint is the credible part |
| Causal inference — DAGs, belief networks, do-calculus, structural time series | Projects 1 and 5 | Collider bias is the dark-figure problem, and you have already separated a true signal from a collection artefact across 22 years of administrative data |
| Bayesian methods, in the crimes-against-women paper | Projects 2 and 5 | A standard of proof attaches to a posterior, not a likelihood, so the legal question is natively Bayesian |
| Time series — ARIMA, LSTM, structural time series | Projects 1 and 5 | Detection rates and market data are both temporal, and the interesting questions concern change around interventions |
| Model evaluation — calibration, cross-validation, hyperparameter search | Project 2 | Calibration is a measurement problem with mature tooling and no legal literature at all |
16. Where this actually shows up in the interview
| If they ask | Reach for |
|---|---|
| Your sample is 15 runs. What can you honestly conclude? | The paired design. Each scenario is its own control, McNemar uses only the cases that changed, pooled claims hold at p of 0.004 and 0.00006, and model-versus-model claims are not supported at all — Fisher gives 0.206 |
| Why Wilson intervals? | Boundary behaviour. The ordinary formula collapses to zero width at 0 of 5 and 5 of 5, asserting a certainty the data cannot support |
| How would you audit a screening tool? | Calibration first, then flag-rate disparity across firm size and sector and country, then sensitivity to design choices, then the multiple-comparison problem at ten thousand tenders |
| Why should a lawyer care about calibration? | Because a standard of proof is a statement about sufficient confidence, and an uncalibrated probability does not denote confidence. The number cannot be mapped onto the standard |
| What is wrong with training on past enforcement decisions? | Detection is a collider. You are conditioning on where the agency looked, so you learn enforcement history rather than collusion |
| Explain something technical to a non-technical panel | The Bayes point with the numbers. A 90-percent-accurate flag at a one-in-200 base rate still leaves a 92 percent chance the firm did nothing wrong |
17. If you remember twelve things
- Bayes flips a conditional, and it needs a base rate. Without one, a flag overstates guilt.
- Expectation is always linear. Variance adds only when uncorrelated. Nothing else is linear — Jensen.
- Choosing a distribution is a claim about mechanism: adding gives normal, multiplying gives log-normal, counting rare events gives Poisson, reinforcement gives a power law.
- Overdispersed counts need negative binomial, not Poisson. Check whether variance equals the mean.
- Beta is the natural distribution for a rate.
- Zero correlation does not imply independence.
- The central limit theorem needs finite variance and large samples. Neither holds at 15 runs or on power-law data.
- Total error equals squared bias plus variance, which is the same decomposition as the bias–variance tradeoff.
- Most loss functions are likelihoods; most regularisers are priors.
- Wilson over Wald at small samples and near the boundaries.
- Calibration and discrimination are independent, and calibration is the one a standard of proof needs.
- Detection is a collider. Conditioning on it is the dark-figure problem.