Skip to content
VibeFormer
58 min

Machine Learning, from the Ground Up

Written for someone relearning rather than revising. Every term defined in plain words before it is used, with a worked example in procurement terms, and the audit question attached to each idea.

Listen

0. How to read this chapter

You are not being hired to train models. You are being hired to decide whether somebody else's model is sound enough to fine a firm ten percent of its worldwide turnover. So every idea below comes with two things: what it is, and what goes wrong with it in a way a lawyer needs to know about.

One running example throughout: 40,000 past public tenders, each with a handful of recorded facts, and the question of which ones involved collusion between bidders.

1. The objects, then the paradigms

Everything later is built from eight things. They are worth naming precisely, because most confusion about machine learning is confusion about which of these somebody means.

TermDefinitionIn the tender dataset
DatasetA collection of examples you learn from40,000 past tenders
Instance, example, rowOne single thing you want to say something aboutOne tender: the Valencia road contract of March 2019
Feature, predictor, variableOne measured property of an instance. The inputNumber of bidders — 3. Winning margin below estimate — 2 percent. Sector — construction. Country — Spain
Label, target, outcomeThe thing you want to predict. The output. Only present in supervised learningWas this tender rigged — yes or no
ModelA function that takes features and returns a predictionA rule that takes those four features and returns a risk score between 0 and 1
Parameters, weightsThe numbers inside the model that get adjusted during learning. You do not choose these; the algorithm finds themThe weight on number of bidders comes out at minus 0.4, meaning more bidders lowers the score
HyperparametersSettings you choose before learning, which control how learning happensHow deep a decision tree may grow. How much regularisation to apply. Where the decision threshold sits
Training and inferenceTraining is finding the parameters from data. Inference, or prediction, is applying the finished model to something newTrain on 2010 to 2020 tenders; infer on a tender published this morning

Two distinctions you will need constantly:

PairDifferenceExample
Classification versus regressionBoth are supervised. Classification predicts which category something falls into; regression predicts a quantityClassification: is this tender rigged, yes or no. Regression: what will the winning bid be, in euro
Binary, multiclass, multilabelTwo categories; more than two but only one can apply; or several that can apply at onceBinary: rigged or not. Multiclass: which of cover bidding, rotation or suppression. Multilabel: a tender showing rotation and suppression together

Now the five paradigms. What separates them is entirely what you are given to start with.

And this is not one agency experimenting. On Schrepel's own account, all 27 EU national competition agencies, DG Competition, and most agencies worldwide now rely on computational tools. The field moved from experiment to routine in under a decade. Six deployments are publicly described, and they are doing four quite different jobs.

Agency and toolWhat it doesWhich technique from this chapter
Spain — BRAVAScreens public procurement for bid rigging, mapping relationships between firms, bids and individualsSupervised classification, with network features built from the relationship graph. Labels come from past enforcement
Brazil — CerebroRuns procurement documents through analysis to surface signs of collusionDocument processing and classification — closer to the NLP chapter than to tabular modelling
FranceQueries its own case database in natural language, through a retrieval-augmented generation system built on large language modelsRetrieval plus generation. Not a screen at all — it finds nothing new, it helps case handlers find what the agency already decided
GreeceAnalyses email metadata seized in dawn raids to establish which companies communicated with each otherNetwork analysis over communications metadata. Who talked to whom, how often, and when
ChileMonitors more than 80,000 products for price anomaliesUnsupervised anomaly detection over time series. No labels, so it reports *unusual*
PolandInvestigates dark patterns using eye tracking and other neuromarketing methodsHuman-subject experimental research. Barely machine learning at all, and a different legal animal entirely

2. What learning actually is

Stripped of vocabulary, supervised learning is four decisions.

DecisionPlain meaningIn the example
Hypothesis classWhich *shapes* of rule you are willing to consider. You are choosing a family of candidate models before seeing which one winsOnly straight-line relationships. Or: any decision tree up to depth six. Or: any combination of 500 trees
Loss functionA number saying how bad a single prediction was. Smaller is betterThe model said 0.9 and the tender was clean. That costs a lot. It said 0.52 and the tender was rigged. That costs a little
Empirical risk minimisationSearch the hypothesis class for the model with the lowest average loss on the data you have. This is what training *is*Try parameter settings until the average cost across 40,000 tenders stops falling
GeneralisationWhether that model still works on tenders it has never seen. The only thing that actually mattersTrained on 2010 to 2020, and the question is whether it works in 2026

A loss function is just a formula for the cost of being wrong. The common ones:

LossWhat it penalisesUsed for
Squared errorThe squared distance between prediction and truth. Being wrong by 10 costs a hundred times being wrong by 1Regression. Badly affected by outliers, which heavy-tailed contract values are full of
Absolute errorThe plain distance. Being wrong by 10 costs ten times being wrong by 1Regression where outliers should not dominate. Targets the median rather than the mean
Cross-entropy, log lossHow surprised the model should be by the truth it was given. Confident and wrong is punished very heavilyClassification. This is the standard one
HingeOnly whether the prediction was on the right side of the boundary, by a marginSupport vector machines
MSE=1n∑i=1n(yi−y^i)2\text{MSE} = \frac{1}{n}\sum_{i=1}^{n}\left(y_i - \hat{y}_i\right)^2
Mean squared error equals the average, over all cases, of the squared difference between the true value and the prediction.The hat on y means *predicted*. Squaring means an error of 10 costs a hundred times an error of 1, which is why outliers dominate it.
Log loss=−1n∑i=1n[yilog⁡p^i+(1−yi)log⁡(1−p^i)]\text{Log loss} = -\frac{1}{n}\sum_{i=1}^{n}\Big[ y_i \log \hat{p}_i + (1 - y_i)\log(1 - \hat{p}_i) \Big]
Log loss equals minus the average of: the true label times the log of the predicted probability, plus one minus the label times the log of one minus the predicted probability.Also called binary cross-entropy. Only one of the two terms survives per case, because the label is either 1 or 0 — so it reduces to minus the log of the probability you assigned to the right answer.
PairDefinitionExample
ParametricThe model has a fixed number of parameters whatever the data size. Learning means setting those numbersLogistic regression with four features has about five numbers, whether you train on 400 tenders or 4 million
Non-parametricThe model's size grows with the data. It can keep adding structureA decision tree adds branches as needed. Nearest-neighbour methods effectively store the whole dataset

3. Overfitting, underfitting, and the tradeoff

Two different ways a model can be wrong, and they need opposite fixes — which is why diagnosing which one you have is the first thing you do.

TermPlain meaningCause
BiasThe model is too simple to capture the real pattern, so it is wrong in the same direction for everyone. Called underfittingHypothesis class too restrictive, too few features, too much regularisation
VarianceThe model is so flexible it has fitted the accidents of this particular sample. Called overfittingModel too flexible for the amount of data, too many features, too little regularisation
Irreducible noiseGenuine randomness no model can remove. Two identical tenders can have different outcomesThe world
E[(y−f^(x))2]=(Bias[f^(x)])2⏟too simple+Var[f^(x)]⏟too sample-specific+σ2⏟irreducible\mathbb{E}\big[(y - \hat{f}(x))^2\big] = \underbrace{\big(\text{Bias}[\hat{f}(x)]\big)^2}_{\text{too simple}} + \underbrace{\text{Var}[\hat{f}(x)]}_{\text{too sample-specific}} + \underbrace{\sigma^2}_{\text{irreducible}}
Expected squared error decomposes exactly into bias squared, plus variance, plus irreducible noise.An identity, not an approximation. It is why you cannot reduce both terms at once by tuning alone — and why the only way to improve both is more data or better features.

Reducing one term often raises the other, which is why it is a tradeoff rather than a problem to be solved.

What you seeWhat it meansWhat to do
Both numbers poor and close togetherUnderfittingMore capacity, better features, less regularisation
Training number excellent, test number poorOverfittingMore data, more regularisation, fewer features, a simpler model
Both reasonable and closeA defensible fitNow check calibration and performance per subgroup before believing it

4. Regularisation, or deliberately handicapping the model

Regularisation means adding a penalty for complexity, so that training has to trade off fitting the data against staying simple. You are accepting slightly worse performance on the data you have, in exchange for better performance on data you do not.

To follow this you need one idea: a weight is the number a model multiplies a feature by. A large weight means that feature strongly moves the output; a weight of zero means the feature is ignored entirely.

MethodWhat it doesEffect
L2, ridgeAdds a penalty equal to the sum of the squared weightsShrinks every weight smoothly toward zero without reaching it. All features stay in, none dominates
L1, lassoAdds a penalty equal to the sum of the absolute weightsPushes some weights to exactly zero, which removes those features entirely. It performs feature selection for you
Elastic netBoth penalties togetherSparsity, with correlated features handled more gracefully
Early stoppingStop training when performance on held-out data starts getting worseCaps complexity without changing the model
DropoutRandomly switch off part of a neural network during each training stepStops the model relying on any single path
Data augmentationTrain on altered copies of the input that should not change the answerTeaches an invariance you believe in
w^=arg⁡min⁡w  L(w)⏟fit the data  +  λ R(w)⏟stay simple\hat{w} = \arg\min_{w} \; \underbrace{L(w)}_{\text{fit the data}} \; + \; \underbrace{\lambda \, R(w)}_{\text{stay simple}}
The chosen weights are those minimising the loss plus lambda times a complexity penalty.Lambda is the dial. At zero there is no regularisation; large lambda forces every weight toward zero and the model underfits. Lambda is a hyperparameter, so a human chooses it.
RL2(w)=∑j=1pwj2RL1(w)=∑j=1p∣wj∣R_{L2}(w) = \sum_{j=1}^{p} w_j^2 \qquad\qquad R_{L1}(w) = \sum_{j=1}^{p} \lvert w_j \rvert
The L2 penalty is the sum of squared weights. The L1 penalty is the sum of absolute weights.The whole difference is squared against absolute — and that is why L1 drives weights to exactly zero while L2 only shrinks them. Squaring makes the penalty on a tiny weight negligible, so L2 never bothers finishing the job.

5. The algorithms, explained rather than listed

6. Unsupervised methods

No labels, so the goal is structure. Two families: grouping similar things together, and compressing many features into few.

MethodWhat it doesWatch out for
k-meansPick k centre points, assign every instance to the nearest, move each centre to the average of what it caught, repeat until nothing movesYou must choose k yourself. Assumes roughly round, similar-sized groups. Sensitive to feature scaling and to where the centres start
Hierarchical clusteringRepeatedly merge the two closest groups, producing a tree of nested groupingsNo k needed in advance, and the tree is readable. Expensive on large datasets
DBSCANGroups points that sit in dense regions, and labels points in sparse regions as noise rather than forcing them into a groupFinds irregular shapes and identifies outliers directly. Struggles when some genuine groups are much sparser than others
Gaussian mixture modelAssumes the data came from several overlapping bell curves and works out which, giving each instance a probability of belonging to each groupSoft membership rather than hard assignment, which is often more honest
Principal component analysisFinds the directions along which the data varies most and describes each instance by its position along those directionsReduces 200 features to 10 while keeping most of the variation. The 10 are mixtures of the originals, so interpretability drops sharply
t-SNE and UMAPSquash high-dimensional data to two dimensions so you can look at itFor looking only. Distances between clusters in the picture are not meaningful, and people over-read these plots constantly
Isolation forest, one-class methods, autoencodersAnomaly detection: learn what normal looks like and score how far each instance sits from itThis is what an unlabelled screen actually is. It outputs *unusual*, which is not *unlawful*

7. Preparing the data

Most of the real work, and most of the real errors.

StepWhat it meansExample
ScalingPutting features on comparable numeric ranges. Standardising subtracts the average and divides by the spreadContract value runs to millions, bidder count to single digits. Without scaling, distance-based methods see only the value. Trees do not care
Categorical encodingTurning categories into numbers. One-hot creates one yes/no column per categoryCountry becomes 27 columns. For something with thousands of categories, such as buyer identity, one-hot is impractical and you encode by frequency or by average outcome instead
Missing dataDeciding what to do about blanks: drop the row, fill with the average, predict the value, or add a column recording that it was missingIn enforcement data, missingness is almost never random — the fact that a field is blank is itself informative, so an explicit missing flag usually beats filling it in
Transforming skewTaking logarithms of quantities with a long right tail, so the few huge values stop dominatingContract values. Model the logarithm, not the euro
Feature constructionBuilding the informative features by hand from the raw onesMargin as a fraction of estimate. Bidder count relative to the sector norm. Gap between lowest and second-lowest bid. These are where the signal lives, and none of them is in the raw data
Curse of dimensionalityWith many features, all points become roughly equally far apart and distance stops being informativeWhy nearest-neighbour methods degrade as features are added

8. Evaluation, which matters most to you

Everything starts from a table of four counts. These four words appear constantly, so they are worth fixing properly.

TermDefinitionIn enforcement terms
True positiveFlagged, and genuinely riggedA cartel caught
False positiveFlagged, but actually cleanAn innocent firm raided
True negativeNot flagged, and cleanCorrectly left alone
False negativeNot flagged, but riggedA cartel missed

Those four counts arranged in a two-by-two grid are called the confusion matrix. Every metric below is a ratio of them.

The confusion matrix

Precision reads across the flagged row; recall reads down the actually-rigged column. That is the whole distinction, and it is the one people reverse under pressure. Precision asks: of those I accused, how many were guilty. Recall asks: of those who were guilty, how many did I catch.
Accuracy=TP+TNTP+TN+FP+FN\text{Accuracy} = \frac{TP + TN}{TP + TN + FP + FN}
Accuracy equals true positives plus true negatives, divided by the total number of cases.Everything you got right, over everything. Useless under imbalance — see the base-rate box below.
Precision=TPTP+FP\text{Precision} = \frac{TP}{TP + FP}
Precision equals true positives divided by true positives plus false positives.Of everything you flagged, the share that was genuinely rigged. This is the number the firm cares about.
Recall=TPTP+FN\text{Recall} = \frac{TP}{TP + FN}
Recall equals true positives divided by true positives plus false negatives.Of everything genuinely rigged, the share you caught. This is the number the agency cares about. Also called sensitivity, or the true positive rate.
Specificity=TNTN+FP\text{Specificity} = \frac{TN}{TN + FP}
Specificity equals true negatives divided by true negatives plus false positives.Of everything genuinely clean, the share you correctly cleared. Its complement, FP over FP plus TN, is the false positive rate.
F1=2⋅Precision⋅RecallPrecision+RecallF_1 = 2 \cdot \frac{\text{Precision} \cdot \text{Recall}}{\text{Precision} + \text{Recall}}
F one equals two times precision times recall, divided by precision plus recall.The harmonic mean, which punishes imbalance between the two — 0.9 and 0.1 gives 0.18, not 0.5. It also weights them equally, which is wrong whenever their real costs differ.
Precision=Recall⋅πRecall⋅π+(1−Specificity)(1−π)\text{Precision} = \frac{\text{Recall} \cdot \pi}{\text{Recall} \cdot \pi + (1 - \text{Specificity})(1 - \pi)}
Precision equals recall times the base rate, divided by recall times the base rate plus the false positive rate times one minus the base rate.Bayes' theorem written in confusion-matrix terms, where pi is the base rate. This is the formula behind the base-rate argument: hold recall and specificity fixed, vary pi, and precision moves anyway.
MetricDefinition in wordsWhen it misleads
AccuracyOf everything, the fraction got rightUseless when one class is rare. If 1 tender in 1,000 is rigged, predicting *never rigged* scores 99.9 percent
PrecisionOf those you flagged, the fraction that really were riggedThis is the number firms care about. Low precision means innocent firms raided
Recall, sensitivityOf those really rigged, the fraction you flaggedThe number the agency cares about. Trades directly against precision
SpecificityOf those really clean, the fraction correctly clearedLooks reassuringly high under imbalance even when precision is dreadful
F1A single score blending precision and recallHides which of the two is bad, and weights them equally when their real costs are nothing like equal
ROC AUCThe chance that a randomly chosen rigged tender scores above a randomly chosen clean oneOver-optimistic when positives are rare, because the false positive rate has an enormous denominator
Precision–recall AUCThe same idea, but focused on the flagged setThe right summary for rare positives. Use this for screening, not ROC
CalibrationWhether a stated 0.7 really corresponds to 70 percentIndependent of everything above. A model can top every leaderboard and still be badly calibrated

9. Splitting the data honestly

You cannot judge a model on data it learned from, any more than you can mark your own exam. So you hold data back. How you hold it back is where the errors are.

MethodWhat it isWhy
Train, validation, testThree portions. Fit on the first, tune choices on the second, and touch the third exactly once at the very endIf you tune against the test set you have learned from it, and the reported number stops being an estimate of anything
k-fold cross-validationSplit into k parts. Train on k minus one, measure on the one left out, rotate so each part is left out once, average the resultsUses all the data for both purposes and gives a spread rather than a single number. Usually k of 5 or 10
Stratified k-foldThe same, but each fold keeps the same proportion of rigged tendersEssential with rare positives. Otherwise a fold may contain none at all
Grouped splitKeep all rows belonging to the same entity on the same side of the splitIf a firm appears in 400 tenders, splitting by tender puts the same firm in both train and test and inflates the result
Time-based splitTrain on earlier periods, test on later onesThe real task is predicting the future. A random split lets the model see tomorrow, which is leakage
Nested cross-validationAn inner loop for tuning, an outer loop for measuringThe correct procedure when you must both tune and report from the same data

10. When the interesting class is rare

Class imbalance means one outcome is far more common than the other. Cartels are rare, so a model can score brilliantly by never flagging anything. Five standard responses:

ApproachWhat it doesCost
Class weightsTell the loss function that getting a rigged tender wrong costs more than getting a clean one wrongSimplest, and usually first choice. It distorts the output probabilities, so recalibrate afterwards
UndersamplingThrow away most of the clean tenders so the classes balanceDiscards real information
OversamplingDuplicate the rigged tenders until the classes balanceEncourages memorising the few positives you have
SMOTEInvent new synthetic rigged tenders by interpolating between real onesCan produce implausible cases. Averaging two real cartels does not give you a third
Threshold movingLeave the model alone and choose explicitly how high the score must be before you actOften the cleanest, because it makes the tradeoff a visible decision rather than a hidden default

11. Explaining a model

Two different ideas, routinely confused. Interpretability is a property of the model: you can read it directly. Explanation is something generated afterwards about a model you cannot read.

MethodWhat it gives youLimits
Interpretable modelsRead the thing itself: a linear equation, a printable tree, a rule listMay cost accuracy — though on tabular data far less often than people assume
Permutation importanceShuffle one feature and see how much performance drops. Big drop means it matteredMisleading with correlated features, which share the credit arbitrarily
Partial dependenceVary one feature across its range, holding the rest, and plot the average effectAssumes you can vary one feature independently of the others, which is often false
LIMEFit a simple readable model in the small neighbourhood of one predictionUnstable. Re-running with different random samples can change the explanation
SHAPAttribute the prediction across features, with a principled grounding in cooperative game theoryExpensive to compute exactly. And faithful to the model, not the world — it says what the model used, not what caused anything
Counterfactual explanationThe smallest change that would have flipped the decisionOften the most legally useful form, because it tells the firm what would have had to be different

12. Bias and fairness, properly

*Bias* means two completely different things in this field, and conflating them wastes most discussions. There is statistical bias from section 3, meaning a model too simple to fit the pattern. And there is social bias, meaning systematically worse treatment of some group. This section is the second. It enters at five distinct points.

SourceDefinitionIn a screen
HistoricalThe world the data records was already unequalPast enforcement concentrated on construction, because that is where the agency built expertise
RepresentationSome groups are under-sampledSmall firms rarely bid for above-threshold contracts, so they are thin in the data
MeasurementThe thing you recorded differs from the thing you care aboutThe label is *was prosecuted*, not *did collude*
AggregationOne model for groups that behave differentlyConstruction and IT services have unlike normal bid spreads, so IT gets scored against construction's baseline
DeploymentUsed differently from how it was validatedValidated on national tenders, run on municipal ones
Fairness criterionWhat it requires
Demographic parityEqual flag rates across groups
Equal opportunityEqual recall across groups — you catch the same proportion of real cartels in each
Equalised oddsEqual recall and equal false positive rates across groups
Calibration within groupsA score of 0.7 means 70 percent in every group
Individual fairnessSimilar cases treated similarly

13. Four questions you asked

14. How this connects to ATLANTIS

This chapter is the technical substrate of the fairness strand, and a good part of the accuracy strand too.

Schrepel's framing of the fairness problemThe machine learning content of it
AI systems inherit the biases of their training dataThe five sources in section 12, and measurement bias above all: the label is *was prosecuted*, not *did collude*
When a model flags one market rather than another, the selection reflects choices about data and design that nobody outside the agency can examineFeature construction, class weighting, the decision threshold, and variance under resampling. Every one is a design choice with a distributional consequence, and none appears in the decision
The duty to give reasons meets an explainability problem the case law never anticipatedSection 11: interpretable models against post-hoc explanation, and the fact that faithful, stable and causal are three different properties
The AI Act adds obligations but its application raises questions the text does not settleArticle 10 says examine for bias without saying how, and the impossibility result means no choice of metric discharges it neutrally

15. What the panel brings to this chapter

Panel memberWhere machine learning meets their work
Thibault SchrepelBuilt a knowledge graph of every Commission competition decision from 1977 to 2025. Constructed an openness metric with Pott and applied it across GPT-4, Llama 3, Gemini, Mistral and MidJourney, then argued policy from the scores. Co-authored with Alex Pentland at MIT on foundation model competition. The pattern throughout: build an instrument, measure something with it, then argue law from the measurement. That is the house style and it is what your position is for
Catalina GoantaFounded the Maastricht Law and Tech Lab with computer scientists in residence at a law school, and runs classification and measurement studies at scale. She has lived the practical problem your appointment creates — how a lawyer and a modeller divide work without one becoming the other's service desk. Expect a question about it, and have a view
Georgiana MirzaDigital ecosystems and market regulation, where the measurement questions are about structure and concentration
Tijmen WismanSyRI, where a state detection model was struck down partly because its risk indicators and its operation were not sufficiently knowable from outside. That is an interpretability finding handed down by a court

16. What is unexplored, and five projects

17. Your CV, mapped onto this chapter

What you haveWhere it landsWhy it fits
Supervised learning across the standard algorithms — logistic regression, trees, forests, SVM, KNN, boostingProjects 1 and 3Auditing a screen means being able to rebuild it, and tabular procurement data is where these methods live
Unsupervised fraud detection at Infosys — insurance fraud, identity theft, phantom billingProjects 1 and 3Rare positives, no reliable labels, and a flag meaning anomalous rather than unlawful. You have met the failure mode already
Model evaluation discipline — cross-validation, grid search, calibration, interval estimation at small n, agreement reportingProject 1, which is essentially this skill written down as a protocolA protocol is only credible from someone who has been burned by a bad evaluation. You have, and you published the post-mortem
The v1 defect analysis — a grader fed reasoning traces instead of output, empty responses scored as results, and an omitted control arm that invalidated the headlineProjects 1 and 3A leakage and measurement-validity case study with your name on it, and the most persuasive evidence you own that you audit your own work
Feature engineering and preprocessingProjects 1 and 3The informative features in bid screens are constructed, not given
Production deployment at 28,000 monthly users with guardrails and edge cases found in live logsProject 5Documentation from someone who has operated a system differs from documentation by someone who has only specified one
Legal training in evidence and procedureProjects 2 and 4Both are normative arguments needing the doctrinal half done properly rather than gestured at

18. Where this shows up in the interview

If they askReach for
How would you audit a screening tool?Leakage first, then the split design — grouped and time-based — then precision–recall rather than ROC, then the base-rate calculation, then calibration, then subgroup flag rates, then stability under resampling
What model would you use?Logistic regression as the baseline you can explain, boosting as the comparison, and the gap between them is the price of interpretability stated as a number
Why is this harder than ordinary classification?Rare positives, no clean labels, labels that come from past enforcement, and a false positive that is a dawn raid on an innocent firm
Can you make a model fair?Name the five sources, then the impossibility result. Fairness is not one property, and choosing between criteria is normative rather than technical
Explain overfitting to a lawyerA high-variance model trained on slightly different history would have flagged different firms, so the choice of who gets investigated is partly arbitrary — and arbitrariness is a legal defect
Explain any of this to a non-technical panelUse the threshold. One number decides how many innocent firms are disturbed per cartel caught, and nobody chose it deliberately

19. If you remember twelve things

  1. A feature is an input, a label is the answer, parameters are found by training, hyperparameters are chosen by you.
  2. Supervised has labels; unsupervised does not; self-supervised invents them; reinforcement has only a reward.
  3. Error equals bias squared plus variance plus noise. Variance is the one with legal consequences.
  4. The gap between training and test performance is the diagnosis of over- or underfitting, every time.
  5. L1 produces sparsity, and sparsity is what makes a model explainable in a decision.
  6. Logistic regression is the baseline worth defending in a regulated setting; boosting usually wins on tabular data.
  7. Anomaly detection flags *unusual*, not *unlawful*.
  8. Leakage is the first thing to check in an audit, and the classic case is a feature that exists only because an investigation happened.
  9. Accuracy is useless under imbalance. Use precision–recall, not ROC, for rare positives.
  10. The base rate destroys precision: 95 percent recall and 95 percent specificity at one in a thousand gives under two percent precision.
  11. Split by group and by time, never randomly, on enforcement data.
  12. Calibration within groups and equalised odds are provably incompatible when base rates differ.