REVISION 5 — Everything Technical
AI, machine learning, deep learning, neural networks, RNNs, LSTMs, transformers, attention, encoders and decoders, LLMs, RAG, evaluation and data-science craft. Every term defined in plain words, with a diagram and an example. Written for someone who has not read any of this in two years.
0. How to read this, and the honest framing
1. The map — how the fields nest
Start here. Almost everyone muddles these four words, and getting the nesting right in your first sentence buys you real credibility. They are not competing things — they are boxes inside boxes, plus one that sits sideways.
The nesting
2. The vocabulary spine — every term, one line each
| Term | Plain definition | Concrete example |
|---|---|---|
| Model | A function with adjustable numbers in it, which you tune until its outputs match known answers | *Given a tender's features, output the probability it was rigged* |
| Parameter / weight | One of those adjustable numbers. Learned, not chosen by you | A 1.1-billion-parameter model has 1.1 billion of them |
| Hyperparameter | A setting you choose *before* training, which is not learned | How fast to learn; how many layers; the `r = 16` in your own fine-tune |
| Feature | One input the model is allowed to look at | Firm size, sector, number of bidders, winning margin |
| Label / target | The known answer you are training it to reproduce | *This tender was found to be rigged* — and note that is a record of enforcement, not of conduct |
| Training | Adjusting the parameters until outputs fit labels | What ran on the rented GPU in your thesis |
| Inference | Using the finished model on new data. No learning happens | Every request your Astroformer chat makes |
| Training / test set | Data it learns from / data held back to check it | Skip the split and you cannot tell learning from memorising |
| Overfitting | It memorised the training data and fails on anything new | Perfect on 200 past cases, useless on the 201st |
| Underfitting | It is too simple to capture the pattern at all | Fitting a straight line to something curved |
| Generalisation | Performing well on data never seen in training. The entire goal | — |
| Neural network / ANN | A model made of layers of simple units passing numbers forward. *Artificial neural network* | Section 4 |
| Deep | **Just means *many layers*.** There is no deeper meaning | Two layers is shallow; fifty is deep |
| RNN | *Recurrent neural network* — reads a sequence one step at a time, carrying a memory | Section 6 |
| LSTM | *Long short-term memory* — an RNN with gates that let it hold information longer | Section 6 |
| Transformer | The architecture behind every modern language model. Reads a whole sequence at once | Section 7 |
| Attention | The mechanism letting a model decide which other words matter for the word it is on | Section 7 |
| Encoder / decoder | An encoder reads; a decoder writes. Models use one or both | Section 8 |
| Token | The chunk of text a model actually processes — a word or part of one | *competition* → compet + ition |
| Embedding | A list of numbers representing meaning, positioned so similar meanings sit close together | Section 11 |
| Context window | The maximum number of tokens it can consider at once | Yours was 512, which is small |
| Pretraining | The first, enormous, general training run. Costs millions; you will never do it | Done by others for TinyLlama |
| Fine-tuning | Taking a trained model and adjusting it on your own smaller, specific data | Your thesis |
| RAG | *Retrieval-augmented generation* — look it up first, then answer using what you found | Section 9 |
| Hallucination | Fluent, confident falsehood | Inventing a case citation |
| Temperature | A randomness dial. Low is repetitive and safe; high is creative and unreliable | For legal work, low — and say why: reproducibility |
| Loss function | A number measuring how wrong the model currently is. Training minimises it | Section 5 |
| Gradient descent | The method: nudge every parameter in whichever direction lowers the loss | Section 5 |
| Backpropagation | The bookkeeping that works out which direction each parameter should be nudged | Section 5 |
| Epoch | One complete pass through all the training data | Yours ran two |
| Batch | How many examples it sees before updating the parameters | Yours: 4, accumulated 8 times, so effectively 32 |
| Quantisation | Storing parameters at lower numeric precision to save memory | The *Q* in QLoRA — 4-bit instead of 16-bit |
| Data leakage | Information about the answer sneaking into the inputs, inflating performance | Section 10 — the most important failure on this page |
| Drift | The world changes, so a model trained on the past gets worse over time | A screen calibrated before a market restructured |
3. The four ways a machine can learn
| Paradigm | What it needs | What it gives | The catch |
|---|---|---|---|
| Supervised | Labelled examples — inputs *and* known answers | Prediction and classification. The workhorse | Labels are expensive, and in law they are usually records of enforcement rather than of conduct |
| Unsupervised | Just the data, no answers | Clusters, groups, anomalies, structure | **No ground truth, so you cannot call it *right* — only coherent.** Interpretation is on you |
| Self-supervised | Raw data only — it invents its own task, like predicting the next word | How every LLM is built. Turns the whole internet into training data with nobody labelling anything | You get general capability, not the specific thing you wanted |
| Reinforcement | An environment and a reward signal | Behaviour learned by trial and error | Needs enormous numbers of trials. Rare in legal settings |
4. The neuron and the neural network
x values are the inputs. The w values are the weights — how much each input counts, and these are what training learns. b is the bias, a standing nudge up or down. **f is the *activation function*, and it is the only reason any of this works: it is a simple non-linear step, and without it a stack of layers would collapse mathematically into one straight line however many you added. So the activation function is what lets depth buy you anything at all.**One neuron, then a network
5. How it actually learns
L is the loss. The fraction is the slope — how much the loss changes if this one weight changes. **η (*eta*) is the learning rate, the step size, and you choose it. The minus sign is the entire algorithm: move against the slope.** Repeat a few million times.The training loop, and the only graph that matters
6. Sequences — RNNs, LSTMs, and why they lost
The network in section 4 takes a fixed set of inputs. Language is not fixed — it is a sequence of varying length where order changes meaning. *The Commission fined the firm* and *The firm fined the Commission* contain identical words. So sequence models were the next step, and the story of how they failed is why transformers exist.
An RNN reads one word at a time, carrying a memory
7. Transformers and attention
Sequential reading against attention
softmax just turns raw scores into weights adding to 1. Dividing by the square root of the dimension is housekeeping — without it the scores grow as the model widens and the weights collapse onto one word. You do not need this formula in the interview. You need the library analogy.8. Encoder, decoder, and the three model families
The three families
9. LLMs and RAG
| Term | Meaning | Your own instance |
|---|---|---|
| Pretraining | The vast general run using next-word prediction. Costs millions | Done by others |
| Fine-tuning | Adjusting a trained model on your own smaller data | 346,920 rows on TinyLlama |
| LoRA / parameter-efficient tuning | Freeze the original, train a small set of extra parameters alongside. Far cheaper, and you can keep several task add-ons for one base model | `r = 16`, `alpha = 32` |
| Quantisation | Lower numeric precision to fit in less memory | The *Q* in QLoRA — and it needs CUDA, which is why it could not run on your MacBook |
| Context window | How many tokens it can consider at once | 512 — which is why stuffing in three documents plus three web snippets was too much |
| Temperature | Randomness dial | Say *low, for anything legal*, and say why: reproducibility |
| System prompt | Standing instructions the user does not see | Your Astroformer chat: *use only the chart data above* |
| Instruction tuning / RLHF | Extra training that makes a model follow instructions. *Reinforcement learning from human feedback* | Why a chat model answers rather than continuing your sentence |
| Sycophancy | The tendency to agree with the user rather than hold a position — a known side effect of training on human approval | Directly relevant to your Mens Rea result |
| Emergence | Two different meanings. Be careful. In ML: abilities appearing suddenly with scale. In complexity science: system behaviour not predictable from the parts | Schrepel means the second. The first is contested — one well-known paper argues the apparent jumps are largely an artefact of all-or-nothing metrics |
RAG — the pipeline
10. Evaluation — and why accuracy lies
This is the most important section on the page for your interview, because the project's framing is about accuracy and safeguards — and the sharpest thing you can demonstrate is knowing when a number is meaningless.
The confusion matrix — everything comes from four boxes
11. NLP and embeddings
| Term | Definition | Worked instance |
|---|---|---|
| Corpus | The body of text you are working with | All Commission decisions 1977–2025 |
| Document | One item in the corpus | One decision |
| Token | One processed unit of text | *self-preferencing* → self, -, prefer, encing |
| Tokenisation | Splitting text into tokens | Where legal text breaks: *Art 102 TFEU* can shatter into five meaningless pieces |
| Stop words | Very common words often removed — *the*, *of*, *and* | Careful: *not* is a stop word in some lists, and removing it inverts legal meaning |
| Lemmatisation | Reducing words to a base form | *concealed*, *concealing*, *conceals* → *conceal* |
| Named entity recognition | Finding and typing names in text | *Google Ireland Ltd* → organisation |
| TF-IDF | Scores a word highly if frequent here and rare elsewhere, so it finds distinctive words. Order is discarded | Fast, interpretable, and a genuinely hard baseline to beat on legal text |
| Static embedding | One vector per word, always the same | GloVe — so *consideration* is identical in contract doctrine and ordinary speech. The limitation in your summarisation paper |
| Contextual embedding | The vector changes with the surrounding sentence | BERT-family. Solves exactly that |
| Cosine similarity | The angle between two vectors, ignoring length | 1 = same direction, 0 = unrelated |
What an embedding space actually is
12. Data-science craft — the ten things that break projects
- About sixty percent of the work happens before any model, and cleaning is the largest block. Saying this is a credibility signal — it says you have run a project rather than read about one.
- The unit of analysis decides what your result can mean. One row = one firm, one case, or one tender are three different studies. **And *case* measures enforcement, not conduct.**
- The ecological fallacy: a correlation at district or country level licenses no claim about individuals. **Say *at district level* inside the sentence and the fallacy cannot sneak in.**
- Plot rows per time period first, before anything else. Real-world processes change gradually; data-collection processes change abruptly. A step in the series is a change in administration, not in the world.
- Three kinds of missing data. *Completely at random* — safe. *At random* — means random given what you already observe, so fixable by imputation or weighting. *Not at random* — depends on the missing value itself, and no fix exists from the data alone. **Your own crime data is the third kind: the dark figure *is* the selection.**
- Rates or counts? Counts scale with population, so the biggest unit always looks worst. Per-capita rates are unstable with small denominators. And changing the denominator can reverse the finding with the numerator untouched — so declare it.
- Simpson's paradox: a pattern can hold in every subgroup and reverse when pooled, with every number correct. Check the obvious subgroups before reporting any pooled comparison.
- Entity resolution is a legal question, not a cleaning step — competition law acts on the *undertaking*, a single economic unit, not the legal person. So your name-matching threshold affects attribution of liability and the lawful size of a fine.
- Measurement validity: name the construct, name the proxy, name the direction of the error. *I am measuring detected cartels, which is a sample selected on concealment, so my prevalence estimate is a lower bound.* Three clauses, and it closes off most attacks.
- The design is the argument, not the regression. Difference-in-differences rests on parallel trends; regression discontinuity on nothing else jumping at the threshold; synthetic control on no spillover. **And all give *average* effects, which cannot carry individual attribution — that gap is the real frontier.**
13. The twelve sentences that cover almost anything
- *Every LLM is deep learning, every deep-learning system is machine learning, all of it is AI — but never the reverse. Rule-based systems are AI and are not machine learning, and most tools agencies actually deploy are rule-based.*
- *Statistics asks whether a pattern is real. Machine learning asks whether I can predict the next one. Data science is the craft around both — and in practice the hard part is usually that the same firm appears under five different names.*
- *Nobody labelled the training data for a language model. It was shown text with the next word hidden and asked to predict it, billions of times — which is why the amount of data stopped being the constraint.*
- *RNNs and LSTMs read a sentence strictly left to right, carrying what they can remember. Transformers look at the whole page at once. And the reason that mattered was not mainly accuracy but parallelism, which is what made internet-scale training possible.*
- *Attention means that for every word, the model computes how much every other word matters to it and rebuilds that word's meaning as a weighted blend. But attention weights are not an explanation — they can be changed substantially while the output stays the same.*
- *Encoder-only models read and classify; decoder-only models write; encoder-decoder models transform one text into another. Reaching for a generative model when you need a classifier is the most common and most expensive mistake in applied legal NLP.*
- *Ground and generate: compute or retrieve the facts in a way you can show and reproduce, and let the model only do the prose. Never let it invent the input. I have built that twice.*
- *Accuracy is the wrong metric for anything rare. If one tender in a thousand is rigged, a model that flags nothing is 99.9 percent accurate and worthless. Report precision and recall separately.*
- *The two errors are not symmetric — investigating an innocent undertaking and missing a live cartel are different kinds of harm — so the threshold between them is a legal decision currently made by typing a number. Making that choice visible is most of the safeguard question.*
- *The clearest documented leakage failure is in legal prediction: about 79 percent accuracy from judgment text written by the court after it decided, falling to roughly 58 to 68 percent when redone as genuine forecasting — and high accuracy from the judges' names alone. Competition law is about to make the same mistake.*
- *Police records measure reporting, not offending. A cartel screen measures detection, not collusion. Same defect, and I have already built the toolkit for separating the two.*
- *These methods can tell a regulator where to look and whether a rule worked. They cannot carry the burden of proof against an individual undertaking — and being clear about which of those two jobs a tool is doing is most of the safeguard question.*