Machine Learning, from the Ground Up
Written for someone relearning rather than revising. Every term defined in plain words before it is used, with a worked example in procurement terms, and the audit question attached to each idea.
0. How to read this chapter
You are not being hired to train models. You are being hired to decide whether somebody else's model is sound enough to fine a firm ten percent of its worldwide turnover. So every idea below comes with two things: what it is, and what goes wrong with it in a way a lawyer needs to know about.
One running example throughout: 40,000 past public tenders, each with a handful of recorded facts, and the question of which ones involved collusion between bidders.
1. The objects, then the paradigms
Everything later is built from eight things. They are worth naming precisely, because most confusion about machine learning is confusion about which of these somebody means.
| Term | Definition | In the tender dataset |
|---|---|---|
| Dataset | A collection of examples you learn from | 40,000 past tenders |
| Instance, example, row | One single thing you want to say something about | One tender: the Valencia road contract of March 2019 |
| Feature, predictor, variable | One measured property of an instance. The input | Number of bidders — 3. Winning margin below estimate — 2 percent. Sector — construction. Country — Spain |
| Label, target, outcome | The thing you want to predict. The output. Only present in supervised learning | Was this tender rigged — yes or no |
| Model | A function that takes features and returns a prediction | A rule that takes those four features and returns a risk score between 0 and 1 |
| Parameters, weights | The numbers inside the model that get adjusted during learning. You do not choose these; the algorithm finds them | The weight on number of bidders comes out at minus 0.4, meaning more bidders lowers the score |
| Hyperparameters | Settings you choose before learning, which control how learning happens | How deep a decision tree may grow. How much regularisation to apply. Where the decision threshold sits |
| Training and inference | Training is finding the parameters from data. Inference, or prediction, is applying the finished model to something new | Train on 2010 to 2020 tenders; infer on a tender published this morning |
Two distinctions you will need constantly:
| Pair | Difference | Example |
|---|---|---|
| Classification versus regression | Both are supervised. Classification predicts which category something falls into; regression predicts a quantity | Classification: is this tender rigged, yes or no. Regression: what will the winning bid be, in euro |
| Binary, multiclass, multilabel | Two categories; more than two but only one can apply; or several that can apply at once | Binary: rigged or not. Multiclass: which of cover bidding, rotation or suppression. Multilabel: a tender showing rotation and suppression together |
Now the five paradigms. What separates them is entirely what you are given to start with.
And this is not one agency experimenting. On Schrepel's own account, all 27 EU national competition agencies, DG Competition, and most agencies worldwide now rely on computational tools. The field moved from experiment to routine in under a decade. Six deployments are publicly described, and they are doing four quite different jobs.
| Agency and tool | What it does | Which technique from this chapter |
|---|---|---|
| Spain — BRAVA | Screens public procurement for bid rigging, mapping relationships between firms, bids and individuals | Supervised classification, with network features built from the relationship graph. Labels come from past enforcement |
| Brazil — Cerebro | Runs procurement documents through analysis to surface signs of collusion | Document processing and classification — closer to the NLP chapter than to tabular modelling |
| France | Queries its own case database in natural language, through a retrieval-augmented generation system built on large language models | Retrieval plus generation. Not a screen at all — it finds nothing new, it helps case handlers find what the agency already decided |
| Greece | Analyses email metadata seized in dawn raids to establish which companies communicated with each other | Network analysis over communications metadata. Who talked to whom, how often, and when |
| Chile | Monitors more than 80,000 products for price anomalies | Unsupervised anomaly detection over time series. No labels, so it reports *unusual* |
| Poland | Investigates dark patterns using eye tracking and other neuromarketing methods | Human-subject experimental research. Barely machine learning at all, and a different legal animal entirely |
2. What learning actually is
Stripped of vocabulary, supervised learning is four decisions.
| Decision | Plain meaning | In the example |
|---|---|---|
| Hypothesis class | Which *shapes* of rule you are willing to consider. You are choosing a family of candidate models before seeing which one wins | Only straight-line relationships. Or: any decision tree up to depth six. Or: any combination of 500 trees |
| Loss function | A number saying how bad a single prediction was. Smaller is better | The model said 0.9 and the tender was clean. That costs a lot. It said 0.52 and the tender was rigged. That costs a little |
| Empirical risk minimisation | Search the hypothesis class for the model with the lowest average loss on the data you have. This is what training *is* | Try parameter settings until the average cost across 40,000 tenders stops falling |
| Generalisation | Whether that model still works on tenders it has never seen. The only thing that actually matters | Trained on 2010 to 2020, and the question is whether it works in 2026 |
A loss function is just a formula for the cost of being wrong. The common ones:
| Loss | What it penalises | Used for |
|---|---|---|
| Squared error | The squared distance between prediction and truth. Being wrong by 10 costs a hundred times being wrong by 1 | Regression. Badly affected by outliers, which heavy-tailed contract values are full of |
| Absolute error | The plain distance. Being wrong by 10 costs ten times being wrong by 1 | Regression where outliers should not dominate. Targets the median rather than the mean |
| Cross-entropy, log loss | How surprised the model should be by the truth it was given. Confident and wrong is punished very heavily | Classification. This is the standard one |
| Hinge | Only whether the prediction was on the right side of the boundary, by a margin | Support vector machines |
| Pair | Definition | Example |
|---|---|---|
| Parametric | The model has a fixed number of parameters whatever the data size. Learning means setting those numbers | Logistic regression with four features has about five numbers, whether you train on 400 tenders or 4 million |
| Non-parametric | The model's size grows with the data. It can keep adding structure | A decision tree adds branches as needed. Nearest-neighbour methods effectively store the whole dataset |
3. Overfitting, underfitting, and the tradeoff
Two different ways a model can be wrong, and they need opposite fixes — which is why diagnosing which one you have is the first thing you do.
| Term | Plain meaning | Cause |
|---|---|---|
| Bias | The model is too simple to capture the real pattern, so it is wrong in the same direction for everyone. Called underfitting | Hypothesis class too restrictive, too few features, too much regularisation |
| Variance | The model is so flexible it has fitted the accidents of this particular sample. Called overfitting | Model too flexible for the amount of data, too many features, too little regularisation |
| Irreducible noise | Genuine randomness no model can remove. Two identical tenders can have different outcomes | The world |
Reducing one term often raises the other, which is why it is a tradeoff rather than a problem to be solved.
| What you see | What it means | What to do |
|---|---|---|
| Both numbers poor and close together | Underfitting | More capacity, better features, less regularisation |
| Training number excellent, test number poor | Overfitting | More data, more regularisation, fewer features, a simpler model |
| Both reasonable and close | A defensible fit | Now check calibration and performance per subgroup before believing it |
4. Regularisation, or deliberately handicapping the model
Regularisation means adding a penalty for complexity, so that training has to trade off fitting the data against staying simple. You are accepting slightly worse performance on the data you have, in exchange for better performance on data you do not.
To follow this you need one idea: a weight is the number a model multiplies a feature by. A large weight means that feature strongly moves the output; a weight of zero means the feature is ignored entirely.
| Method | What it does | Effect |
|---|---|---|
| L2, ridge | Adds a penalty equal to the sum of the squared weights | Shrinks every weight smoothly toward zero without reaching it. All features stay in, none dominates |
| L1, lasso | Adds a penalty equal to the sum of the absolute weights | Pushes some weights to exactly zero, which removes those features entirely. It performs feature selection for you |
| Elastic net | Both penalties together | Sparsity, with correlated features handled more gracefully |
| Early stopping | Stop training when performance on held-out data starts getting worse | Caps complexity without changing the model |
| Dropout | Randomly switch off part of a neural network during each training step | Stops the model relying on any single path |
| Data augmentation | Train on altered copies of the input that should not change the answer | Teaches an invariance you believe in |
5. The algorithms, explained rather than listed
6. Unsupervised methods
No labels, so the goal is structure. Two families: grouping similar things together, and compressing many features into few.
| Method | What it does | Watch out for |
|---|---|---|
| k-means | Pick k centre points, assign every instance to the nearest, move each centre to the average of what it caught, repeat until nothing moves | You must choose k yourself. Assumes roughly round, similar-sized groups. Sensitive to feature scaling and to where the centres start |
| Hierarchical clustering | Repeatedly merge the two closest groups, producing a tree of nested groupings | No k needed in advance, and the tree is readable. Expensive on large datasets |
| DBSCAN | Groups points that sit in dense regions, and labels points in sparse regions as noise rather than forcing them into a group | Finds irregular shapes and identifies outliers directly. Struggles when some genuine groups are much sparser than others |
| Gaussian mixture model | Assumes the data came from several overlapping bell curves and works out which, giving each instance a probability of belonging to each group | Soft membership rather than hard assignment, which is often more honest |
| Principal component analysis | Finds the directions along which the data varies most and describes each instance by its position along those directions | Reduces 200 features to 10 while keeping most of the variation. The 10 are mixtures of the originals, so interpretability drops sharply |
| t-SNE and UMAP | Squash high-dimensional data to two dimensions so you can look at it | For looking only. Distances between clusters in the picture are not meaningful, and people over-read these plots constantly |
| Isolation forest, one-class methods, autoencoders | Anomaly detection: learn what normal looks like and score how far each instance sits from it | This is what an unlabelled screen actually is. It outputs *unusual*, which is not *unlawful* |
7. Preparing the data
Most of the real work, and most of the real errors.
| Step | What it means | Example |
|---|---|---|
| Scaling | Putting features on comparable numeric ranges. Standardising subtracts the average and divides by the spread | Contract value runs to millions, bidder count to single digits. Without scaling, distance-based methods see only the value. Trees do not care |
| Categorical encoding | Turning categories into numbers. One-hot creates one yes/no column per category | Country becomes 27 columns. For something with thousands of categories, such as buyer identity, one-hot is impractical and you encode by frequency or by average outcome instead |
| Missing data | Deciding what to do about blanks: drop the row, fill with the average, predict the value, or add a column recording that it was missing | In enforcement data, missingness is almost never random — the fact that a field is blank is itself informative, so an explicit missing flag usually beats filling it in |
| Transforming skew | Taking logarithms of quantities with a long right tail, so the few huge values stop dominating | Contract values. Model the logarithm, not the euro |
| Feature construction | Building the informative features by hand from the raw ones | Margin as a fraction of estimate. Bidder count relative to the sector norm. Gap between lowest and second-lowest bid. These are where the signal lives, and none of them is in the raw data |
| Curse of dimensionality | With many features, all points become roughly equally far apart and distance stops being informative | Why nearest-neighbour methods degrade as features are added |
8. Evaluation, which matters most to you
Everything starts from a table of four counts. These four words appear constantly, so they are worth fixing properly.
| Term | Definition | In enforcement terms |
|---|---|---|
| True positive | Flagged, and genuinely rigged | A cartel caught |
| False positive | Flagged, but actually clean | An innocent firm raided |
| True negative | Not flagged, and clean | Correctly left alone |
| False negative | Not flagged, but rigged | A cartel missed |
Those four counts arranged in a two-by-two grid are called the confusion matrix. Every metric below is a ratio of them.
The confusion matrix
| Metric | Definition in words | When it misleads |
|---|---|---|
| Accuracy | Of everything, the fraction got right | Useless when one class is rare. If 1 tender in 1,000 is rigged, predicting *never rigged* scores 99.9 percent |
| Precision | Of those you flagged, the fraction that really were rigged | This is the number firms care about. Low precision means innocent firms raided |
| Recall, sensitivity | Of those really rigged, the fraction you flagged | The number the agency cares about. Trades directly against precision |
| Specificity | Of those really clean, the fraction correctly cleared | Looks reassuringly high under imbalance even when precision is dreadful |
| F1 | A single score blending precision and recall | Hides which of the two is bad, and weights them equally when their real costs are nothing like equal |
| ROC AUC | The chance that a randomly chosen rigged tender scores above a randomly chosen clean one | Over-optimistic when positives are rare, because the false positive rate has an enormous denominator |
| Precision–recall AUC | The same idea, but focused on the flagged set | The right summary for rare positives. Use this for screening, not ROC |
| Calibration | Whether a stated 0.7 really corresponds to 70 percent | Independent of everything above. A model can top every leaderboard and still be badly calibrated |
9. Splitting the data honestly
You cannot judge a model on data it learned from, any more than you can mark your own exam. So you hold data back. How you hold it back is where the errors are.
| Method | What it is | Why |
|---|---|---|
| Train, validation, test | Three portions. Fit on the first, tune choices on the second, and touch the third exactly once at the very end | If you tune against the test set you have learned from it, and the reported number stops being an estimate of anything |
| k-fold cross-validation | Split into k parts. Train on k minus one, measure on the one left out, rotate so each part is left out once, average the results | Uses all the data for both purposes and gives a spread rather than a single number. Usually k of 5 or 10 |
| Stratified k-fold | The same, but each fold keeps the same proportion of rigged tenders | Essential with rare positives. Otherwise a fold may contain none at all |
| Grouped split | Keep all rows belonging to the same entity on the same side of the split | If a firm appears in 400 tenders, splitting by tender puts the same firm in both train and test and inflates the result |
| Time-based split | Train on earlier periods, test on later ones | The real task is predicting the future. A random split lets the model see tomorrow, which is leakage |
| Nested cross-validation | An inner loop for tuning, an outer loop for measuring | The correct procedure when you must both tune and report from the same data |
10. When the interesting class is rare
Class imbalance means one outcome is far more common than the other. Cartels are rare, so a model can score brilliantly by never flagging anything. Five standard responses:
| Approach | What it does | Cost |
|---|---|---|
| Class weights | Tell the loss function that getting a rigged tender wrong costs more than getting a clean one wrong | Simplest, and usually first choice. It distorts the output probabilities, so recalibrate afterwards |
| Undersampling | Throw away most of the clean tenders so the classes balance | Discards real information |
| Oversampling | Duplicate the rigged tenders until the classes balance | Encourages memorising the few positives you have |
| SMOTE | Invent new synthetic rigged tenders by interpolating between real ones | Can produce implausible cases. Averaging two real cartels does not give you a third |
| Threshold moving | Leave the model alone and choose explicitly how high the score must be before you act | Often the cleanest, because it makes the tradeoff a visible decision rather than a hidden default |
11. Explaining a model
Two different ideas, routinely confused. Interpretability is a property of the model: you can read it directly. Explanation is something generated afterwards about a model you cannot read.
| Method | What it gives you | Limits |
|---|---|---|
| Interpretable models | Read the thing itself: a linear equation, a printable tree, a rule list | May cost accuracy — though on tabular data far less often than people assume |
| Permutation importance | Shuffle one feature and see how much performance drops. Big drop means it mattered | Misleading with correlated features, which share the credit arbitrarily |
| Partial dependence | Vary one feature across its range, holding the rest, and plot the average effect | Assumes you can vary one feature independently of the others, which is often false |
| LIME | Fit a simple readable model in the small neighbourhood of one prediction | Unstable. Re-running with different random samples can change the explanation |
| SHAP | Attribute the prediction across features, with a principled grounding in cooperative game theory | Expensive to compute exactly. And faithful to the model, not the world — it says what the model used, not what caused anything |
| Counterfactual explanation | The smallest change that would have flipped the decision | Often the most legally useful form, because it tells the firm what would have had to be different |
12. Bias and fairness, properly
*Bias* means two completely different things in this field, and conflating them wastes most discussions. There is statistical bias from section 3, meaning a model too simple to fit the pattern. And there is social bias, meaning systematically worse treatment of some group. This section is the second. It enters at five distinct points.
| Source | Definition | In a screen |
|---|---|---|
| Historical | The world the data records was already unequal | Past enforcement concentrated on construction, because that is where the agency built expertise |
| Representation | Some groups are under-sampled | Small firms rarely bid for above-threshold contracts, so they are thin in the data |
| Measurement | The thing you recorded differs from the thing you care about | The label is *was prosecuted*, not *did collude* |
| Aggregation | One model for groups that behave differently | Construction and IT services have unlike normal bid spreads, so IT gets scored against construction's baseline |
| Deployment | Used differently from how it was validated | Validated on national tenders, run on municipal ones |
| Fairness criterion | What it requires |
|---|---|
| Demographic parity | Equal flag rates across groups |
| Equal opportunity | Equal recall across groups — you catch the same proportion of real cartels in each |
| Equalised odds | Equal recall and equal false positive rates across groups |
| Calibration within groups | A score of 0.7 means 70 percent in every group |
| Individual fairness | Similar cases treated similarly |
13. Four questions you asked
14. How this connects to ATLANTIS
This chapter is the technical substrate of the fairness strand, and a good part of the accuracy strand too.
| Schrepel's framing of the fairness problem | The machine learning content of it |
|---|---|
| AI systems inherit the biases of their training data | The five sources in section 12, and measurement bias above all: the label is *was prosecuted*, not *did collude* |
| When a model flags one market rather than another, the selection reflects choices about data and design that nobody outside the agency can examine | Feature construction, class weighting, the decision threshold, and variance under resampling. Every one is a design choice with a distributional consequence, and none appears in the decision |
| The duty to give reasons meets an explainability problem the case law never anticipated | Section 11: interpretable models against post-hoc explanation, and the fact that faithful, stable and causal are three different properties |
| The AI Act adds obligations but its application raises questions the text does not settle | Article 10 says examine for bias without saying how, and the impossibility result means no choice of metric discharges it neutrally |
15. What the panel brings to this chapter
| Panel member | Where machine learning meets their work |
|---|---|
| Thibault Schrepel | Built a knowledge graph of every Commission competition decision from 1977 to 2025. Constructed an openness metric with Pott and applied it across GPT-4, Llama 3, Gemini, Mistral and MidJourney, then argued policy from the scores. Co-authored with Alex Pentland at MIT on foundation model competition. The pattern throughout: build an instrument, measure something with it, then argue law from the measurement. That is the house style and it is what your position is for |
| Catalina Goanta | Founded the Maastricht Law and Tech Lab with computer scientists in residence at a law school, and runs classification and measurement studies at scale. She has lived the practical problem your appointment creates — how a lawyer and a modeller divide work without one becoming the other's service desk. Expect a question about it, and have a view |
| Georgiana Mirza | Digital ecosystems and market regulation, where the measurement questions are about structure and concentration |
| Tijmen Wisman | SyRI, where a state detection model was struck down partly because its risk indicators and its operation were not sufficiently knowable from outside. That is an interpretability finding handed down by a court |
16. What is unexplored, and five projects
17. Your CV, mapped onto this chapter
| What you have | Where it lands | Why it fits |
|---|---|---|
| Supervised learning across the standard algorithms — logistic regression, trees, forests, SVM, KNN, boosting | Projects 1 and 3 | Auditing a screen means being able to rebuild it, and tabular procurement data is where these methods live |
| Unsupervised fraud detection at Infosys — insurance fraud, identity theft, phantom billing | Projects 1 and 3 | Rare positives, no reliable labels, and a flag meaning anomalous rather than unlawful. You have met the failure mode already |
| Model evaluation discipline — cross-validation, grid search, calibration, interval estimation at small n, agreement reporting | Project 1, which is essentially this skill written down as a protocol | A protocol is only credible from someone who has been burned by a bad evaluation. You have, and you published the post-mortem |
| The v1 defect analysis — a grader fed reasoning traces instead of output, empty responses scored as results, and an omitted control arm that invalidated the headline | Projects 1 and 3 | A leakage and measurement-validity case study with your name on it, and the most persuasive evidence you own that you audit your own work |
| Feature engineering and preprocessing | Projects 1 and 3 | The informative features in bid screens are constructed, not given |
| Production deployment at 28,000 monthly users with guardrails and edge cases found in live logs | Project 5 | Documentation from someone who has operated a system differs from documentation by someone who has only specified one |
| Legal training in evidence and procedure | Projects 2 and 4 | Both are normative arguments needing the doctrinal half done properly rather than gestured at |
18. Where this shows up in the interview
| If they ask | Reach for |
|---|---|
| How would you audit a screening tool? | Leakage first, then the split design — grouped and time-based — then precision–recall rather than ROC, then the base-rate calculation, then calibration, then subgroup flag rates, then stability under resampling |
| What model would you use? | Logistic regression as the baseline you can explain, boosting as the comparison, and the gap between them is the price of interpretability stated as a number |
| Why is this harder than ordinary classification? | Rare positives, no clean labels, labels that come from past enforcement, and a false positive that is a dawn raid on an innocent firm |
| Can you make a model fair? | Name the five sources, then the impossibility result. Fairness is not one property, and choosing between criteria is normative rather than technical |
| Explain overfitting to a lawyer | A high-variance model trained on slightly different history would have flagged different firms, so the choice of who gets investigated is partly arbitrary — and arbitrariness is a legal defect |
| Explain any of this to a non-technical panel | Use the threshold. One number decides how many innocent firms are disturbed per cartel caught, and nobody chose it deliberately |
19. If you remember twelve things
- A feature is an input, a label is the answer, parameters are found by training, hyperparameters are chosen by you.
- Supervised has labels; unsupervised does not; self-supervised invents them; reinforcement has only a reward.
- Error equals bias squared plus variance plus noise. Variance is the one with legal consequences.
- The gap between training and test performance is the diagnosis of over- or underfitting, every time.
- L1 produces sparsity, and sparsity is what makes a model explainable in a decision.
- Logistic regression is the baseline worth defending in a regulated setting; boosting usually wins on tabular data.
- Anomaly detection flags *unusual*, not *unlawful*.
- Leakage is the first thing to check in an audit, and the classic case is a feature that exists only because an investigation happened.
- Accuracy is useless under imbalance. Use precision–recall, not ROC, for rare positives.
- The base rate destroys precision: 95 percent recall and 95 percent specificity at one in a thousand gives under two percent precision.
- Split by group and by time, never randomly, on enforcement data.
- Calibration within groups and equalised odds are provably incompatible when base rates differ.