Skip to content
VibeFormer
62 min

LLMs, RAG, RLHF and Reasoning

How a language model actually works, from tokens to decoding, then fine-tuning and RLHF, then retrieval-augmented generation and its failure modes. The chapter where your own paper is the content rather than a footnote.

Listen

0. Why this is the chapter that matters most

Three reasons. It is your actual domain. It is what the panel is likeliest to probe, because it is where your published work sits. And France's competition authority already runs retrieval-augmented generation over its own case database, so this is not speculative — it is a deployed enforcement tool nobody has audited.

1. Tokens, the unit everything is built from

A model does not see words. It sees tokens — pieces of text from a fixed vocabulary, usually somewhere between a character and a word. Tokenisation is the process of cutting text into them.

TermDefinition
VocabularyThe fixed set of tokens the model knows, typically 30,000 to 150,000 of them
Byte-pair encodingThe common method. Start from characters and repeatedly merge the most frequent adjacent pair until you have your vocabulary. Frequent words become one token; rare ones get split
Token IDEach token is really just an integer index into the vocabulary
Rule of thumbIn English, roughly 4 characters or 0.75 words per token

2. What the model does with those tokens

You do not need the mathematics. You need the four ideas, in order.

StepWhat happensPlain version
EmbeddingEach token ID is replaced by a list of numbers — a vector — learned during trainingEvery token becomes a point in a space where nearby points mean similar things
Positional encodingInformation about *where* in the sequence each token sits is addedOtherwise the model sees a bag of tokens and *A sued B* looks identical to *B sued A*
AttentionFor each token, the model computes how much to weight every other token, then builds a new representation as a weighted blend of themEach word gets to look at every other word and decide which ones matter to its own meaning
Feed-forward layers, stackedThe blended representations pass through ordinary layers, and the whole attention-plus-layers block repeats dozens of timesEach repetition builds a slightly more abstract reading of the text
Attention(Q,K,V)=softmax ⁣(QK⊤dk)V\text{Attention}(Q, K, V) = \mathrm{softmax}\!\left(\frac{QK^{\top}}{\sqrt{d_k}}\right) V
Attention equals softmax of Q times K transposed, divided by the square root of d-k, all multiplied by V.Three roles per token: the query is what this token is looking for, the key is what each token offers, the value is what it contributes if chosen. Q times K transposed scores every token against every other. Softmax turns those scores into weights that total one. Multiplying by V produces the weighted blend. Dividing by the square root of d-k keeps the scores from growing with vector size, which would otherwise push softmax into saturation where almost all the weight lands on one token and gradients die.

One transformer block

Attention mixes information across positions; the feed-forward layer processes each position on its own. Those are the only two operations, and a large model is this block repeated. The two plus-signs are *residual connections* — the input is added back to the output, so each block learns a correction rather than a replacement, which is what makes a hundred-layer stack trainable at all.
ArchitectureReadsUsed for
Encoder-only, BERT-styleThe whole text at once, in both directionsClassification, extraction, labelling a corpus. Still the right and far cheaper tool for coding 30 agency reports
Decoder-only, GPT-styleLeft to right, each token seeing only what came beforeGeneration. Every model you think of as a chatbot
Encoder-decoderEncodes an input, then generates an outputTranslation, summarisation

3. Pretraining, and why it works at all

Pretraining is the self-supervised step from the machine learning chapter, at enormous scale: take text, hide the next token, and train the model to predict it. The answer was free because you deleted it yourself. Repeat across trillions of tokens.

TermDefinition
ParametersThe adjustable numbers inside the model. A 7-billion-parameter model has 7 billion of them
Compute, FLOPsTotal arithmetic used in training. The AI Act's systemic-risk threshold is set at 10 to the 25th FLOPs
Scaling lawsEmpirical finding that performance improves predictably as parameters, data and compute grow together. Chinchilla showed most early models were badly undertrained — they needed far more data for their size
EmergenceCapabilities appearing abruptly beyond a certain scale rather than improving smoothly. Contested in detail, and central to Schrepel's argument that emergence defeats prediction
Mixture of expertsOnly a fraction of the network activates per token, so a model can be large in parameters but cheap per token. One of your three subject models, gpt-oss-20b, is this design

4. Decoding — turning probabilities into text

The model gives a probability for every possible next token. Decoding is how you choose one. This is a setting, not a property of the model, and it changes the output substantially.

P(x1,…,xn)=∏t=1nP(xt∣x<t)whereP(xt∣x<t)=softmax(zt)P(x_1, \ldots, x_n) = \prod_{t=1}^{n} P(x_t \mid x_{<t}) \qquad\text{where}\qquad P(x_t \mid x_{<t}) = \mathrm{softmax}(z_t)
The probability a model assigns to a whole passage is the product of the probability of each token given all preceding tokens, and each of those conditional probabilities comes from a softmax over the model's raw scores.This is the chain rule from the probability chapter, unchanged. That is the entire mathematical content of *a language model*: it is a device for evaluating one factor of this product at a time. Note what the formula does not contain — no term for truth, for intent, or for whether the passage describes the world correctly.
P(xt=i)=exp⁡(zi/T)∑jexp⁡(zj/T)P(x_t = i) = \frac{\exp(z_i / T)}{\sum_{j} \exp(z_j / T)}
The probability of token i equals e to the power of its score divided by temperature, divided by the sum of that same quantity over every token in the vocabulary.The z values are logits — raw unbounded scores. Dividing by T before exponentiating is the whole temperature mechanism. T below 1 magnifies the gaps between scores, so the leader dominates. T above 1 shrinks them toward uniform. As T approaches 0 the largest logit takes essentially all the probability, which is why temperature 0 is described as greedy. Exponentiating guarantees positives; dividing by the total guarantees they sum to one.
MethodWhat it doesEffect
GreedyAlways take the single most probable tokenDeterministic in principle, and often repetitive and flat
Beam searchKeep several candidate continuations alive and pick the best overall sequenceBetter for translation. Tends to produce bland text for open generation
TemperatureFlattens or sharpens the distribution before sampling. Below 1 sharpens toward the likeliest; above 1 flattens toward variety; 0 is effectively greedyThe single most important knob. Your paper used 0.0 throughout to minimise variation
Top-kSample only from the k most probable tokensCuts off the long tail of nonsense
Top-p, nucleusSample from the smallest set of tokens whose probabilities sum to pAdapts how many options to consider based on how confident the model is
Repetition penaltyDown-weight tokens already usedStops loops
TermDefinition
Context windowThe maximum number of tokens the model can attend to at once — prompt plus generated output. Anything beyond it is invisible
KV cacheStored intermediate results so generating token 500 does not recompute tokens 1 to 499. Why the first token is slow and the rest are fast
Lost in the middleMeasured tendency to use information at the start and end of a long context better than material buried in the middle. Directly relevant to stuffing 20 retrieved documents into a prompt

5. Adapting a model: fine-tuning, SFT, RLHF and DPO

A pretrained model predicts plausible continuations. It does not follow instructions, refuse harmful requests, or adopt a tone. Those are added afterwards, in stages.

StageWhat it isData needed
PretrainingNext-token prediction over a vast corpusTrillions of tokens, no labels
Supervised fine-tuning (SFT)Continue training on example prompt-and-ideal-response pairs, so the model learns the shape of being helpfulThousands to hundreds of thousands of written demonstrations
Reward modellingHumans rank several candidate responses. A separate model is trained to predict those rankings — it learns to score responses as a human wouldHuman preference comparisons
RLHFReinforcement learning from human feedback. The language model is optimised to score highly on the reward model, using reinforcement learning — usually PPOThe reward model, plus compute
DPODirect preference optimisation. Skips the separate reward model and the reinforcement learning loop, optimising directly on preference pairs. Simpler, more stable, now very widely usedThe same preference pairs
Constitutional AI / RLAIFUses a written set of principles and model-generated critiques in place of much of the human labellingPrinciples, plus a capable model
Efficient methodWhat it does
Full fine-tuningUpdate every parameter. Expensive, and needs the whole model in memory
LoRAFreeze the original model and train small added matrices alongside it. A tiny fraction of the parameters, and the result is a small file you can swap in
QLoRALoRA on a quantised model, so a large model fits on a single consumer GPU
QuantisationStore parameters at lower numerical precision — 8-bit or 4-bit instead of 16. Smaller and faster, with some quality loss
DistillationTrain a small model to imitate a large one's outputs

6. Prompting

TechniqueDefinition
System promptInstructions placed before the conversation, setting role and constraints. Usually hidden from the user
Zero-shotJust ask
Few-shotInclude several worked examples in the prompt, so the model infers the pattern
Chain of thoughtAsk the model to reason step by step before answering. Often improves accuracy on multi-step problems
Self-consistencySample several reasoning paths and take the majority answer
Structured outputConstrain the response to a schema, so it can be parsed rather than interpreted

7. RAG — retrieval-augmented generation

RAG means fetching relevant documents and placing them in the prompt, so the model answers from them rather than from memory. Two phases.

The RAG pipeline, both phases

Everything upstream of the model determines the answer, and only the last box is the part people discuss. Note where the failure points sit: three of the four are retrieval and formatting problems, not generation problems. Note also that the final verification step is a parser rather than a model, which is the same instinct as the forced-choice probe in your own harness.
Index time, done onceWhat happens
LoadCollect the documents — every decision from 1977 onward
ChunkSplit each into passages small enough to retrieve usefully
EmbedConvert each chunk into a vector capturing its meaning
StorePut the vectors in a database built for similarity search
Query time, every questionWhat happens
Embed the queryTurn the question into a vector in the same space
RetrieveFind the k nearest chunks by vector similarity
Rerank, optionalRe-score the candidates with a slower, more accurate model that reads query and chunk together
AugmentPaste the surviving chunks into the prompt
GenerateThe model answers using what it was given
TermDefinition
EmbeddingA vector representing meaning, where similar texts sit close together
Cosine similarityThe standard closeness measure between two vectors
Vector databaseStorage optimised for nearest-neighbour search
ANN searchApproximate nearest neighbour. Trades exactness for speed, which means retrieval itself is approximate
BM25The classical keyword ranking function. Still strong, and the baseline any semantic method should be compared against
Hybrid searchCombining keyword and semantic retrieval, which usually beats either alone
Cross-encoder rerankerReads query and candidate together rather than separately. Much more accurate, far slower, so used on a shortlist
Chunking strategiesFixed-size; recursive, respecting paragraph boundaries; semantic, splitting on topic shifts; or parent-document, retrieving a small chunk but supplying its larger parent

8. Reasoning in RAG

Basic RAG retrieves once and answers. That fails whenever the question needs several steps, or when the right search terms are not in the question.

TechniqueWhat it does
Query rewritingRephrase the user's question into better search terms before retrieving
Multi-queryGenerate several reformulations, retrieve for each, and pool the results
HyDEHave the model write a hypothetical ideal answer, then retrieve documents similar to *that* rather than to the question. Answers resemble documents more than questions do
Step-back promptingAsk a more general question first to establish the governing principle, then the specific one
Multi-hopRetrieve, read, work out what is still missing, retrieve again. Necessary for chained questions
Self-RAGThe model decides for itself whether to retrieve, and critiques whether what came back actually supports its answer
Corrective RAGAssess retrieval quality and fall back — to a different query or to the open web — when it is poor
GraphRAGRetrieve over a knowledge graph of entities and relations rather than over loose passages
Agentic RAGThe model plans a sequence of searches and tool calls rather than following a fixed pipeline

9. The failure modes — and this is your section

FailureDefinition
HallucinationFluent, confident output that is false. Follows directly from the training objective: plausibility was rewarded, truth was not
Prompt injection, directA user instructs the model to ignore its instructions
Prompt injection, indirectThe malicious instruction arrives inside retrieved content, so the attacker never touches the prompt
RAG poisoning, corpus poisoningPlanting material in the retrieval corpus to steer outputs
SycophancyShifting the stated answer toward what the user seems to believe
Unfaithful chain of thoughtStated reasoning that does not describe the actual computation
Lost in the middleMaterial buried in a long context being effectively ignored
Context conflictRetrieved documents contradicting each other, or contradicting the model's training
Groundedness failureAn answer not supported by the documents it was given, while appearing to be

10. Evaluating these systems

MeasureWhat it asks
Faithfulness, groundednessIs every claim in the answer supported by the retrieved documents?
Answer relevanceDoes the answer address the question asked?
Context precisionOf what was retrieved, how much was actually relevant?
Context recallOf what was needed, how much was retrieved?
Citation accuracyDo the cited sources exist, and do they say what is claimed?
Refusal appropriatenessDoes it decline when it should, and only then?
Perplexity=exp⁡ ⁣(−1n∑t=1nlog⁡P(xt∣x<t))\text{Perplexity} = \exp\!\left(-\frac{1}{n}\sum_{t=1}^{n} \log P(x_t \mid x_{<t})\right)
Perplexity is e raised to minus the average log probability the model assigned to each actual next token.Read it as the effective number of options the model was choosing between. Perplexity 1 means it knew the next token exactly; perplexity 50 means it was as uncertain as someone picking from 50 equally likely words. It is the exponential of average log loss, so it is a *fluency* measure. A model can have excellent perplexity on legal text and still invent a case citation, because fluent and true are different properties — which is why nothing in the table above is a perplexity.

LLM-as-a-judge means using a model to grade outputs. It beats surface-overlap metrics on semantic tasks, and it carries known biases.

Judge biasWhat it is
Position biasPreferring whichever candidate appears first
Verbosity biasPreferring longer answers regardless of quality
Self-preferencePreferring text generated by itself or its own family
Grading instabilityThe same input scored differently across calls, even at temperature zero

11. Vocabulary, including the ones used without definition

TermMeaning
Foundation model, frontier modelA large general-purpose model trained at scale, adapted to many tasks downstream
Open weights versus open sourceOpen weights means the parameters are downloadable. Open source properly also implies the training data and code, which almost no model offers. Schrepel and Pott built a metric precisely because the distinction is routinely blurred
InferenceRunning a trained model. Distinguished from training
GuardrailsChecks around a model constraining what it accepts or emits. On your CV from Astroformer
Red teamingDeliberately attacking your own system to find failures. Now a legal obligation under AI Act Article 55 for systemic-risk models
AlignmentMaking a model behave as intended. Alignment tax is capability lost in the process
JailbreakA prompt defeating a model's safety training
Agent, agenticA model that plans, calls tools, and acts over several steps rather than answering once
Tool use, function callingLetting a model invoke external functions — a search, a calculator, a database query
Regulatory sandboxA legal term, not a technical one. A supervised environment where a firm may test an innovation under regulator oversight with some requirements relaxed. AI Act Articles 57 to 59 require Member States to establish them; the Digital Omnibus pushed that deadline to December 2027
Model card, datasheetStandardised documentation of a model's intended use, performance and limitations. Voluntary practice, and the ancestor of AI Act Annex IV technical documentation
BenchmarkA fixed dataset and metric for comparing systems. Benchmark contamination is when test data leaked into training
AblationRemoving one component to measure its contribution. Your middle condition is an ablation: the fake document with the instruction removed
DriftPerformance decaying as the world moves away from the training data. Legal corpora drift constantly, because law changes

12. How this connects to ATLANTIS

StrandThe content of it here
Accuracy, the data problemFrance's retrieval system is a deployed enforcement tool whose answers depend entirely on what its corpus contains and how it was chunked. Corpus completeness and chunk integrity are accuracy questions with no current standard
FairnessArticle 296 requires reasons and Article 14 requires meaningful oversight. Your finding is that a system's self-account can be true, checkable and incomplete — which attacks both at once, empirically
Institutional arrangementsWho audits an agency's retrieval corpus? Nobody, currently. There is no obligation, no methodology, and after the Annex III gap possibly no legal hook either

13. What the panel brings to this chapter

Panel memberWhy this chapter is their territory
Thibault SchrepelTeaches The Law of Artificial Intelligence. Wrote Decoding the AI Act. Built the EC decisions knowledge graph and offers to help agencies build queryable systems from their own records — which is this pipeline. Co-authored on foundation model competition with Pentland at MIT and on measuring model openness with Pott. His podcast has an episode on two language models learning to fix prices in a repeated duopoly, where some agents learn to cover their tracks — which is your research domain arriving from the substantive side
Catalina GoantaMeasurement over unstructured text at scale, and lawful research access to data held by someone unwilling to share it. Both are your problems
Georgiana MirzaDigital ecosystems and data spaces, where model access and data access shape market structure
Tijmen WismanA system whose reasoning cannot be inspected from outside is the SyRI problem. He has seen a court strike one down for exactly that

14. What is unexplored, and six projects

15. Your CV, mapped onto this chapter

What you haveWhere it landsWhy it fits
The Mens Rea Evaluator — controlled ground truth, three conditions, clean control, six probes, enforced measurement invariantsProjects 1, 2 and 6. The single most transferable asset you ownYou have built and debugged the instrument, and the failure mode you found is the one Article 14 assumes away
RAG pipelines in production at 28,000 monthly users, with guardrails and edge cases found in live logsProjects 1, 3 and 4Chunking and retrieval failures are things you have debugged rather than theorised
Adversarial prompting and red teaming, professionally at Outlier and in a published harnessProjects 1 and 2Article 55 makes this a statutory obligation. No translation needed
Fine-tuning on open-weight models — TinyLlama, Hugging Face, Gemma, Llama 3, NVIDIA NIMProjects 1 and 6Credibility on the GPAI chapter and on the openness debate Schrepel writes about
Evaluation methodology — majority-voted judges with agreement recorded, deterministic parsing where no model is needed, invariants enforced in codeProjects 2 and 5This is the discipline the whole field is missing, and you have it in print
The v1 post-mortem — traces fed to a grader, empty outputs scored, an omitted control arm that invalidated the headlineEvery project hereIt is your best evidence that you audit your own work, and the best answer to a time something broke
Multilingual RAG across six Indian languagesProjects 1 and 3A pan-EU corpus is 24 equally authentic languages, and cross-language divergence is where it breaks first
Legal training in evidence and procedureProjects 1, 2 and 5Why you can see that Article 86 and Article 27 of Regulation 1/2003 ask different questions about the same file

16. If you remember twelve things

  1. A language model is a conditional distribution over the next token. Plausibility was rewarded; truth was not. Every failure follows from that.
  2. Tokens are not words, and legal citations fragment into meaningless debris.
  3. Attention lets every token weigh every other, which is what represents a negated agreement rather than an agreement.
  4. Temperature is the main decoding knob, and temperature zero is not determinism through an API.
  5. Pretraining, then supervised fine-tuning, then preference optimisation by RLHF or DPO.
  6. RLHF optimises what rankers preferred, which is where sycophancy comes from — and why persona-shift probes are necessary.
  7. Prompt, fine-tune, or retrieve. Only retrieval can produce a citation, which is why a body under a duty to give reasons needs it.
  8. RAG is chunk, embed, store, then retrieve, rerank, generate. Retrieval is approximate at every stage.
  9. Chain of thought can be unfaithful — it is not an explanation, and feeding it to a grader is what invalidated your v1.
  10. Indirect injection and corpus poisoning reach the model through retrieved content, never touching the prompt.
  11. Your finding: the document channel alone sufficed in 9 of 15, the clean control at zero is what licensed the claim, and self-report failed at every level.
  12. The real result is partial attribution — a true, checkable, incomplete reason that survives testing and points away from the cause.