Skip to content
VibeFormer
54 min

REVISION 5 — Everything Technical

AI, machine learning, deep learning, neural networks, RNNs, LSTMs, transformers, attention, encoders and decoders, LLMs, RAG, evaluation and data-science craft. Every term defined in plain words, with a diagram and an example. Written for someone who has not read any of this in two years.

Listen

0. How to read this, and the honest framing

1. The map — how the fields nest

Start here. Almost everyone muddles these four words, and getting the nesting right in your first sentence buys you real credibility. They are not competing things — they are boxes inside boxes, plus one that sits sideways.

The nesting

Every large language model is deep learning, every deep-learning system is machine learning, and all of it is artificial intelligence — but never the reverse. Rule-based systems are AI and are not machine learning, which matters enormously, because most tools competition agencies actually deploy are rule-based.

2. The vocabulary spine — every term, one line each

TermPlain definitionConcrete example
ModelA function with adjustable numbers in it, which you tune until its outputs match known answers*Given a tender's features, output the probability it was rigged*
Parameter / weightOne of those adjustable numbers. Learned, not chosen by youA 1.1-billion-parameter model has 1.1 billion of them
HyperparameterA setting you choose *before* training, which is not learnedHow fast to learn; how many layers; the `r = 16` in your own fine-tune
FeatureOne input the model is allowed to look atFirm size, sector, number of bidders, winning margin
Label / targetThe known answer you are training it to reproduce*This tender was found to be rigged* — and note that is a record of enforcement, not of conduct
TrainingAdjusting the parameters until outputs fit labelsWhat ran on the rented GPU in your thesis
InferenceUsing the finished model on new data. No learning happensEvery request your Astroformer chat makes
Training / test setData it learns from / data held back to check itSkip the split and you cannot tell learning from memorising
OverfittingIt memorised the training data and fails on anything newPerfect on 200 past cases, useless on the 201st
UnderfittingIt is too simple to capture the pattern at allFitting a straight line to something curved
GeneralisationPerforming well on data never seen in training. The entire goal—
Neural network / ANNA model made of layers of simple units passing numbers forward. *Artificial neural network*Section 4
Deep**Just means *many layers*.** There is no deeper meaningTwo layers is shallow; fifty is deep
RNN*Recurrent neural network* — reads a sequence one step at a time, carrying a memorySection 6
LSTM*Long short-term memory* — an RNN with gates that let it hold information longerSection 6
TransformerThe architecture behind every modern language model. Reads a whole sequence at onceSection 7
AttentionThe mechanism letting a model decide which other words matter for the word it is onSection 7
Encoder / decoderAn encoder reads; a decoder writes. Models use one or bothSection 8
TokenThe chunk of text a model actually processes — a word or part of one*competition* → compet + ition
EmbeddingA list of numbers representing meaning, positioned so similar meanings sit close togetherSection 11
Context windowThe maximum number of tokens it can consider at onceYours was 512, which is small
PretrainingThe first, enormous, general training run. Costs millions; you will never do itDone by others for TinyLlama
Fine-tuningTaking a trained model and adjusting it on your own smaller, specific dataYour thesis
RAG*Retrieval-augmented generation* — look it up first, then answer using what you foundSection 9
HallucinationFluent, confident falsehoodInventing a case citation
TemperatureA randomness dial. Low is repetitive and safe; high is creative and unreliableFor legal work, low — and say why: reproducibility
Loss functionA number measuring how wrong the model currently is. Training minimises itSection 5
Gradient descentThe method: nudge every parameter in whichever direction lowers the lossSection 5
BackpropagationThe bookkeeping that works out which direction each parameter should be nudgedSection 5
EpochOne complete pass through all the training dataYours ran two
BatchHow many examples it sees before updating the parametersYours: 4, accumulated 8 times, so effectively 32
QuantisationStoring parameters at lower numeric precision to save memoryThe *Q* in QLoRA — 4-bit instead of 16-bit
Data leakageInformation about the answer sneaking into the inputs, inflating performanceSection 10 — the most important failure on this page
DriftThe world changes, so a model trained on the past gets worse over timeA screen calibrated before a market restructured

3. The four ways a machine can learn

ParadigmWhat it needsWhat it givesThe catch
SupervisedLabelled examples — inputs *and* known answersPrediction and classification. The workhorseLabels are expensive, and in law they are usually records of enforcement rather than of conduct
UnsupervisedJust the data, no answersClusters, groups, anomalies, structure**No ground truth, so you cannot call it *right* — only coherent.** Interpretation is on you
Self-supervisedRaw data only — it invents its own task, like predicting the next wordHow every LLM is built. Turns the whole internet into training data with nobody labelling anythingYou get general capability, not the specific thing you wanted
ReinforcementAn environment and a reward signalBehaviour learned by trial and errorNeeds enormous numbers of trials. Rare in legal settings

4. The neuron and the neural network

y=f ⁣(∑i=1nwixi+b)y = f\!\left(\sum_{i=1}^{n} w_i x_i + b\right)
The output of a neuron is a function applied to the sum of each input multiplied by its own weight, plus a bias term.Term by term. The x values are the inputs. The w values are the weights — how much each input counts, and these are what training learns. b is the bias, a standing nudge up or down. **f is the *activation function*, and it is the only reason any of this works: it is a simple non-linear step, and without it a stack of layers would collapse mathematically into one straight line however many you added. So the activation function is what lets depth buy you anything at all.**

One neuron, then a network

A neural network is a stack of layers of these units where every connection is one learned number. Hidden is unfortunate jargon for the middle layers — it means structurally internal, not unknowable. Depth is simply how many middle layers there are.

5. How it actually learns

w  ←  w  −  η ∂L∂ww \;\leftarrow\; w \;-\; \eta \, \frac{\partial L}{\partial w}
Each weight is replaced by itself, minus the learning rate times the slope of the loss with respect to that weight.In words: find out whether nudging this weight up would make things worse, and if so nudge it down instead. L is the loss. The fraction is the slope — how much the loss changes if this one weight changes. **η (*eta*) is the learning rate, the step size, and you choose it. The minus sign is the entire algorithm: move against the slope.** Repeat a few million times.

The training loop, and the only graph that matters

The loop is mechanical; the judgement is in the second graph. Training loss falling while held-out loss rises is the signature of overfitting, and the moment they diverge is when to stop. This is why a held-out set is not optional — without one you cannot see this picture at all.

6. Sequences — RNNs, LSTMs, and why they lost

The network in section 4 takes a fixed set of inputs. Language is not fixed — it is a sequence of varying length where order changes meaning. *The Commission fined the firm* and *The firm fined the Commission* contain identical words. So sequence models were the next step, and the story of how they failed is why transformers exist.

An RNN reads one word at a time, carrying a memory

The recurrent design reuses one box at every step, carrying a fixed-size memory forward. Both failures follow from that: the memory has a fixed budget, and the training signal decays as it travels back. The second has a name worth knowing — the vanishing gradient problem.

7. Transformers and attention

Sequential reading against attention

Attention replaces a walk along the sentence with a direct weighted lookup from every position to every other. Multi-head attention runs several such lookups in parallel so different relationship types are captured at once. This is the single architectural idea behind every modern language model.
Attention(Q,K,V)=softmax ⁣(QK⊤dk)V\mathrm{Attention}(Q,K,V) = \mathrm{softmax}\!\left(\frac{QK^{\top}}{\sqrt{d_k}}\right) V
Attention compares each query against every key to get scores, scales them down, turns them into weights that sum to one, and uses those weights to average the values.The library analogy makes this readable. The query is what the current word is looking for. The keys are the labels on every other word — what each offers. The values are the actual content. So: match your question against every label, see which fit best, then take a blend of the contents weighted by fit. softmax just turns raw scores into weights adding to 1. Dividing by the square root of the dimension is housekeeping — without it the scores grow as the model widens and the weights collapse onto one word. You do not need this formula in the interview. You need the library analogy.

8. Encoder, decoder, and the three model families

The three families

Encoder-only models read and classify. Decoder-only models write. Encoder-decoder models transform one text into another. You have used all three: BERT-family thinking in your NLP work, TinyLlama as a decoder in your thesis, and BART as an encoder-decoder in both your thesis and your summarisation paper.

9. LLMs and RAG

TermMeaningYour own instance
PretrainingThe vast general run using next-word prediction. Costs millionsDone by others
Fine-tuningAdjusting a trained model on your own smaller data346,920 rows on TinyLlama
LoRA / parameter-efficient tuningFreeze the original, train a small set of extra parameters alongside. Far cheaper, and you can keep several task add-ons for one base model`r = 16`, `alpha = 32`
QuantisationLower numeric precision to fit in less memoryThe *Q* in QLoRA — and it needs CUDA, which is why it could not run on your MacBook
Context windowHow many tokens it can consider at once512 — which is why stuffing in three documents plus three web snippets was too much
TemperatureRandomness dialSay *low, for anything legal*, and say why: reproducibility
System promptStanding instructions the user does not seeYour Astroformer chat: *use only the chart data above*
Instruction tuning / RLHFExtra training that makes a model follow instructions. *Reinforcement learning from human feedback*Why a chat model answers rather than continuing your sentence
SycophancyThe tendency to agree with the user rather than hold a position — a known side effect of training on human approvalDirectly relevant to your Mens Rea result
EmergenceTwo different meanings. Be careful. In ML: abilities appearing suddenly with scale. In complexity science: system behaviour not predictable from the partsSchrepel means the second. The first is contested — one well-known paper argues the apparent jumps are largely an artefact of all-or-nothing metrics

RAG — the pipeline

Retrieval-augmented generation separates the factual layer from the writing layer. The facts come from documents you control and can show; the model only phrases them. Chunking is the step that decides retrieval quality — and splitting a holding away from its reasoning is the classic way to ruin a legal corpus.

10. Evaluation — and why accuracy lies

This is the most important section on the page for your interview, because the project's framing is about accuracy and safeguards — and the sharpest thing you can demonstrate is knowing when a number is meaningless.

The confusion matrix — everything comes from four boxes

Every evaluation metric is built from these four counts. The asymmetry between the two errors is where law enters, because the trade-off between investigating the innocent and missing the guilty is a normative choice that a default parameter should not be making.
Precision=TPTP+FPRecall=TPTP+FNF1=2⋅P⋅RP+R\text{Precision} = \frac{TP}{TP + FP} \qquad \text{Recall} = \frac{TP}{TP + FN} \qquad F_1 = 2\cdot\frac{P \cdot R}{P + R}
Precision is the share of things you flagged that really were positive. Recall is the share of actual positives you managed to flag. F1 is the harmonic mean of the two.In plain words. Precision: when it raises the alarm, how often is it right? — the question an investigated firm cares about. Recall: of all the real cartels, how many did it find? — the question the agency cares about. They trade off: flag everything and recall is perfect while precision collapses. F1 combines them, and the *harmonic* mean is used because it punishes imbalance — you cannot score well by being excellent at one and terrible at the other.

11. NLP and embeddings

TermDefinitionWorked instance
CorpusThe body of text you are working withAll Commission decisions 1977–2025
DocumentOne item in the corpusOne decision
TokenOne processed unit of text*self-preferencing* → self, -, prefer, encing
TokenisationSplitting text into tokensWhere legal text breaks: *Art 102 TFEU* can shatter into five meaningless pieces
Stop wordsVery common words often removed — *the*, *of*, *and*Careful: *not* is a stop word in some lists, and removing it inverts legal meaning
LemmatisationReducing words to a base form*concealed*, *concealing*, *conceals* → *conceal*
Named entity recognitionFinding and typing names in text*Google Ireland Ltd* → organisation
TF-IDFScores a word highly if frequent here and rare elsewhere, so it finds distinctive words. Order is discardedFast, interpretable, and a genuinely hard baseline to beat on legal text
Static embeddingOne vector per word, always the sameGloVe — so *consideration* is identical in contract doctrine and ordinary speech. The limitation in your summarisation paper
Contextual embeddingThe vector changes with the surrounding sentenceBERT-family. Solves exactly that
Cosine similarityThe angle between two vectors, ignoring length1 = same direction, 0 = unrelated
cos⁡(θ)=a⋅b∥a∥ ∥b∥\cos(\theta) = \frac{\mathbf{a} \cdot \mathbf{b}}{\|\mathbf{a}\|\,\|\mathbf{b}\|}
Cosine similarity is the dot product of two vectors divided by the product of their lengths.Why divide by the lengths: it removes document size from the comparison. A long decision and a short one on the same subject should count as similar; without this division the long one looks different merely for being long. And a precise point you can make: if you use squared Euclidean distance on vectors that have not been normalised, you are not computing cosine similarity, and the two rank results differently. That exact mismatch is in your own thesis — the text says cosine, the code builds a distance index on unnormalised vectors.

What an embedding space actually is

An embedding places text in a space where proximity reflects contextual similarity. That is what makes semantic search work without shared keywords, and it is also the limitation: distributional similarity is not legal equivalence, and two concepts can sit adjacent while requiring entirely different things to be proved.

12. Data-science craft — the ten things that break projects

  1. About sixty percent of the work happens before any model, and cleaning is the largest block. Saying this is a credibility signal — it says you have run a project rather than read about one.
  2. The unit of analysis decides what your result can mean. One row = one firm, one case, or one tender are three different studies. **And *case* measures enforcement, not conduct.**
  3. The ecological fallacy: a correlation at district or country level licenses no claim about individuals. **Say *at district level* inside the sentence and the fallacy cannot sneak in.**
  4. Plot rows per time period first, before anything else. Real-world processes change gradually; data-collection processes change abruptly. A step in the series is a change in administration, not in the world.
  5. Three kinds of missing data. *Completely at random* — safe. *At random* — means random given what you already observe, so fixable by imputation or weighting. *Not at random* — depends on the missing value itself, and no fix exists from the data alone. **Your own crime data is the third kind: the dark figure *is* the selection.**
  6. Rates or counts? Counts scale with population, so the biggest unit always looks worst. Per-capita rates are unstable with small denominators. And changing the denominator can reverse the finding with the numerator untouched — so declare it.
  7. Simpson's paradox: a pattern can hold in every subgroup and reverse when pooled, with every number correct. Check the obvious subgroups before reporting any pooled comparison.
  8. Entity resolution is a legal question, not a cleaning step — competition law acts on the *undertaking*, a single economic unit, not the legal person. So your name-matching threshold affects attribution of liability and the lawful size of a fine.
  9. Measurement validity: name the construct, name the proxy, name the direction of the error. *I am measuring detected cartels, which is a sample selected on concealment, so my prevalence estimate is a lower bound.* Three clauses, and it closes off most attacks.
  10. The design is the argument, not the regression. Difference-in-differences rests on parallel trends; regression discontinuity on nothing else jumping at the threshold; synthetic control on no spillover. **And all give *average* effects, which cannot carry individual attribution — that gap is the real frontier.**

13. The twelve sentences that cover almost anything

  1. *Every LLM is deep learning, every deep-learning system is machine learning, all of it is AI — but never the reverse. Rule-based systems are AI and are not machine learning, and most tools agencies actually deploy are rule-based.*
  2. *Statistics asks whether a pattern is real. Machine learning asks whether I can predict the next one. Data science is the craft around both — and in practice the hard part is usually that the same firm appears under five different names.*
  3. *Nobody labelled the training data for a language model. It was shown text with the next word hidden and asked to predict it, billions of times — which is why the amount of data stopped being the constraint.*
  4. *RNNs and LSTMs read a sentence strictly left to right, carrying what they can remember. Transformers look at the whole page at once. And the reason that mattered was not mainly accuracy but parallelism, which is what made internet-scale training possible.*
  5. *Attention means that for every word, the model computes how much every other word matters to it and rebuilds that word's meaning as a weighted blend. But attention weights are not an explanation — they can be changed substantially while the output stays the same.*
  6. *Encoder-only models read and classify; decoder-only models write; encoder-decoder models transform one text into another. Reaching for a generative model when you need a classifier is the most common and most expensive mistake in applied legal NLP.*
  7. *Ground and generate: compute or retrieve the facts in a way you can show and reproduce, and let the model only do the prose. Never let it invent the input. I have built that twice.*
  8. *Accuracy is the wrong metric for anything rare. If one tender in a thousand is rigged, a model that flags nothing is 99.9 percent accurate and worthless. Report precision and recall separately.*
  9. *The two errors are not symmetric — investigating an innocent undertaking and missing a live cartel are different kinds of harm — so the threshold between them is a legal decision currently made by typing a number. Making that choice visible is most of the safeguard question.*
  10. *The clearest documented leakage failure is in legal prediction: about 79 percent accuracy from judgment text written by the court after it decided, falling to roughly 58 to 68 percent when redone as genuine forecasting — and high accuracy from the judges' names alone. Competition law is about to make the same mistake.*
  11. *Police records measure reporting, not offending. A cartel screen measures detection, not collusion. Same defect, and I have already built the toolkit for separating the two.*
  12. *These methods can tell a regulator where to look and whether a rule worked. They cannot carry the burden of proof against an individual undertaking — and being clear about which of those two jobs a tool is doing is most of the safeguard question.*