Skip to content
VibeFormer
42 min

Frontier Topics & Jargon

The words that get used in these conversations without ever being defined — agents, digital twins, mixture of experts, quantisation, sandboxes, model cards, drift. Each defined plainly, with what it means for enforcement and where the hype outruns the evidence.

Listen

0. Why this chapter exists

Everything so far has been a subject with a textbook. This chapter is different: it is the vocabulary that circulates in exactly the kind of conversation you are walking into, usually undefined, often by people who are themselves unsure. You will hear several of these words on 7 October.

1. Agents and agentic systems

An agent, in current usage, is a language model placed in a loop where it can call tools, observe the results, and decide what to do next — rather than producing a single response and stopping.

The agent loop

The defining feature is that the model chooses its own next action. Every step compounds: an error at step two is the input to step three, and nothing in the architecture notices. This is why agent reliability falls off sharply as the number of steps grows, and why the honest version of this technology is still mostly demonstrations.
TermWhat it means
Tool use, function callingThe model emits a structured request — a search query, a database call — which the surrounding program executes and feeds back. The model does not *do* anything itself; it asks
ReActReason, then act, then observe, then repeat. The pattern most agent frameworks implement
Multi-agent systemSeveral models with different assigned roles passing messages. Sometimes genuinely useful, often a way of hiding that no single step is reliable
Planner-executorOne model breaks a goal into steps, another carries them out
Human in the loopA person approves actions before they take effect. The design choice with actual legal significance
GuardrailsRules outside the model constraining what it may do. A parser, a permission list, a cost limit. Not a prompt

2. Digital twins, and why the term travels badly

A digital twin is a computational model of a specific real system, kept synchronised with it by a live data feed, used to simulate what would happen under conditions you have not tried. The term comes from engineering: a twin of a particular jet engine, fed by that engine's sensors.

RequirementWhy it mattersDoes a market have it?
A specific real referentA twin of a thing, not a class of thingsA market is a contested legal construct, not an object. Market definition is itself litigated
A live data feedWhat makes it a twin rather than a modelAgencies get data in response to requests, with delays measured in months
A validated mechanismYou must know the system's physics to simulate itFirm behaviour has no equivalent of physics. Competing economic models give opposite predictions
Counterfactual fidelityThe simulation must be right about conditions never observedThis is the whole point and the whole problem. Nobody can validate a counterfactual that did not occur

3. How large models are actually built and shipped

These five terms come up constantly and are rarely explained. Each is an engineering choice with a consequence for whether a system can be audited.

TermDefinitionWhy a regulator should care
Mixture of expertsThe network contains many parallel sub-networks, and a routing layer sends each token to only a few of them. A model with 400 billion parameters might use 30 billion per tokenThe model's behaviour depends on routing, so two near-identical inputs can be processed by different sub-networks. Reproducibility and explanation both get harder, and parameter count stops being a meaningful measure of the compute actually used — which matters because the AI Act's systemic-risk threshold is stated in training compute
DistillationTrain a small model to imitate a large one's outputs. The small model inherits much of the behaviour at a fraction of the costObligations attach to the model that was trained, not to its imitation. A distilled copy of a regulated model may fall under thresholds while reproducing the capability. Also the main route by which a provider's safety work is or is not inherited
QuantisationStore the weights at lower numerical precision — 8-bit or 4-bit instead of 16-bit — making the model smaller and fasterIt changes the outputs. A quantised model is not the model that was evaluated, so conformity assessment performed on the full-precision version does not straightforwardly cover what was actually deployed. Nobody has litigated this
Fine-tuning a released modelTaking an open-weights model and training it further on your own dataShifts responsibility. Under the AI Act a party who substantially modifies a system can become its provider, and fine-tuning is the ordinary case of that
Retrieval instead of trainingRather than fine-tune, give the model the documents at query timeThe reason RAG dominates in legal applications: the knowledge is inspectable and replaceable and you can point at the source. It is an auditability decision disguised as an engineering one
TermDefinition
Open weightsThe trained parameters are downloadable. Not the same as open source — the training data and code usually are not released, so you can run and modify the model but not reproduce it
Context lengthHow many tokens fit in one prompt. Long-context models reduce but do not remove the need for retrieval, because attention cost grows with the square of length
MultimodalHandles more than text — images, audio, video. Relevant here for interface screenshots in dark-pattern work and for scanned documents
InferenceRunning a trained model. Distinguished from training, which produced it. Most cost and most regulatory surface is inference, because that is where it touches people
Scaling lawsThe empirical finding that loss falls predictably with model size, data and compute. The basis of the whole industry's capital allocation, and an extrapolation rather than a theorem

4. Regulatory sandboxes, and the dates that changed

A regulatory sandbox is a supervised environment where a provider may develop and test a system under a regulator's oversight, with some flexibility on how rules apply during the test. The AI Act requires them, which is unusual — most sandboxes are optional national initiatives.

ProvisionWhat it does
Art 57Each Member State must establish at least one AI regulatory sandbox at national level. A duty, not an option — and it is one of the few places the Act obliges states to build capacity rather than obliging firms to comply
Art 58Detailed arrangements, set out in Commission implementing acts
Art 59Processing of personal data lawfully collected for other purposes, in the sandbox, for developing certain public-interest AI systems — subject to conditions
Art 60Testing of high-risk systems in real-world conditions outside a sandbox, with safeguards and informed consent
Art 62Measures for SMEs and start-ups, including priority sandbox access

5. Documentation artefacts, and what each is worth

A cluster of proposals exists for documenting models and datasets. All are voluntary in origin; several are now partially mandatory through the AI Act. They matter to you because documentation is where the technical and legal tracks meet, and because your own methodological instincts already point this way.

ArtefactWhat it recordsStatus
Model cardIntended use, out-of-scope use, training data summary, evaluation results broken down by subgroup, known limitationsVoluntary, industry-standard, and the ancestor of the AI Act's technical documentation requirement
Datasheet for datasetsHow the data was collected, by whom, with what consent, what it excludes, what it should not be used forVoluntary. The most useful of the lot for legal purposes, because most failures trace to the data rather than the model
System cardDocuments the deployed system rather than the model — including the surrounding guardrails and human reviewVoluntary, and closer to what Art 14 oversight actually needs
Annex IV technical documentationThe AI Act's mandatory file for high-risk systems: design, architecture, data, metrics, oversight measuresMandatory, from the dates above
Art 11 and 12Technical documentation and automatic logging of events over the system's lifetimeMandatory. Logging is the provision with real forensic potential, and the least discussed
Art 13Instructions for use, sufficient for the deployer to interpret outputMandatory, and the hinge between provider and deployer responsibility

6. Drift, contamination and the ways evaluations lie

Four failure modes that sit between a reported benchmark number and what a system does in use. Each has a precise name, and each is a reason not to believe a performance claim.

FailureWhat happensConsequence
Data driftThe input distribution changes after deployment — new procurement rules, a new tendering platform, inflation moving contract valuesThe model is unchanged and its accuracy falls anyway. The model did not break; the world moved
Concept driftThe relationship between inputs and the right answer changes. Cartels adapt to the screen that was catching themWorse than data drift, because the historical labels are now actively misleading. And in enforcement it is adversarial: deploying the screen causes the drift
Benchmark contaminationThe evaluation set was in the training data, so the reported score measures memorisationA model can score highly on a legal benchmark and fail on a case decided after its training cut-off. This is why a published leaderboard number tells an agency almost nothing
Goodhart's lawOnce a metric becomes the target it stops measuring what it measured. Optimise for flag precision and the screen learns to flag only obvious casesThe metric improves and the enforcement outcome worsens. The bid-rigging version: a screen tuned to look precise stops finding novel schemes

7. Privacy and verification techniques you should be able to name

These appear whenever the answer to a governance problem is *technical*. Each solves something real, and each has a limitation worth knowing, because the limitation is usually where the legal question lives.

TechniqueWhat it doesThe catch
Differential privacyAdds calibrated noise so that any one individual's presence in the data cannot be detected from the output. Comes with a tunable parameter, epsilon, quantifying the leakageNoise costs accuracy, and it costs it most on rare cases. Cartels are rare cases. The privacy-utility tradeoff is not a detail here, it is the finding
Federated learningTrain across several parties' data without moving it; only model updates are sharedUpdates leak information, so it needs differential privacy on top. And it does not solve the legal problem of who may process what — it changes where the processing happens, not its lawfulness
Synthetic dataGenerate artificial records with the real data's statistical structure. The GAN route from the deep learning chapterFaithful enough to be useful may be faithful enough to re-identify. Whether synthetic data is personal data under the GDPR is genuinely unsettled, and that is a live question rather than a settled one
Secure multi-party computation, homomorphic encryptionCompute over data nobody can read, including the party doing the computingExpensive, and limited in the operations supported. Real, but rarely the practical answer today
Zero-knowledge proofsProve a statement is true without revealing why. In principle: prove a screen was run correctly without disclosing the screenEarly for this use, and the proof covers the computation, not whether the computation was the right one to run
Audit by behaviourQuery the system and study the outputs, without access to weights or data. What your own harness doesCannot distinguish mechanisms that produce the same behaviour — which is exactly the limitation you already state in your own paper

8. Fast glossary — the rest of it

Terms you may hear once, needing only a sentence each. Grouped so you can find them.

TermOne sentence
Foundation modelA large model trained broadly and then adapted to many tasks. The AI Act's near-equivalent is general-purpose AI model, which is the term to use in a legal sentence
Systemic riskAn AI Act category for general-purpose models above a compute threshold, carrying extra obligations. The threshold is stated in training FLOPs, which is why mixture-of-experts routing complicates it
Emergent capabilityA capability that appears abruptly with scale rather than improving gradually. Contested — some apparent emergence is an artefact of the metric used
AlignmentMaking a model behave as intended. RLHF and DPO are alignment techniques; the word also names the broader research field
Red teamingDeliberately attacking a system to find failures before deployment. Required in substance for systemic-risk models
JailbreakA prompt that defeats a model's restrictions. Relevant to you because the poisoned condition in your own study is structurally a jailbreak with a legal payload
Prompt injectionInstructions hidden in retrieved content that the model follows as if they came from the user. The security problem RAG creates, and it is unsolved
Hallucination, confabulationFluent output that is false. The second word is better because it does not imply perception
SycophancyAgreeing with the user rather than being correct. Measured, and a direct product of training on human preference
Chain of thoughtIntermediate reasoning text produced before an answer. Not a reliable account of how the answer was produced
Reasoning modelA model trained to spend more computation at inference on intermediate steps. Better at verifiable tasks; the faithfulness problem is unchanged
Interpretability, explainabilityUnderstanding how a model computes versus producing an account a human can use. These are different, and the law almost always means the second
Mechanistic interpretabilityReverse-engineering the actual circuits inside a network. Real progress, nowhere near courtroom-ready
Saliency, SHAP, LIMEPost-hoc methods attributing an output to input features. Widely used, and known to be unstable — different methods disagree on the same prediction, which is itself the legal problem
Counterfactual explanation*Had your turnover been below X, you would not have been flagged.* The most legally useful explanation format, and closest to what a court wants
Vector databaseStorage built for similarity search over embeddings. The retrieval half of RAG
EmbeddingA list of numbers representing meaning, where nearby lists mean similar things
TokenThe unit a language model reads and writes — roughly a word piece, not a word
Edge, on-deviceRunning a model locally rather than on a server. Relevant because it keeps data in the agency's hands
MLOpsThe engineering practice of deploying and monitoring models. Where drift detection actually lives, and usually absent in public-sector deployments
ReproducibilitySame code, same data, same result. Harder than it sounds with these systems, and the reason your harness pins configuration and logs every call

9. What is unexplored here, and four projects

Each of these comes out of this chapter rather than from a literature review, each is doable in a PhD, and each has a legal half and a technical half.

ProjectThe questionHow you would actually do it
Where is the decision in an agent loop?If an agency's pipeline has twenty automated steps and one human approval, has a human meaningfully overseen anything? Art 22 GDPR and Art 14 AI Act both assume a single identifiable decision pointBuild a realistic multi-step enforcement pipeline. Instrument every step. Measure how often a reviewer approving only the final output would have caught an error injected at step three. An empirical test of a legal standard, and the natural next experiment after your own paper
Does conformity assessment survive quantisation?A system is evaluated at full precision and deployed quantised. Does Annex IV documentation cover what was deployed?Take an open-weights model, evaluate a legal classification task at 16-bit, 8-bit and 4-bit, and report where the outputs diverge. Then read the conformity provisions against the measurements. Cheap to run, and nobody has run it
Who sandboxes the regulator?Arts 57 to 60 assume a firm testing a product. Enforcement screening is the state deploying a high-risk system against undertakings, with no sandbox and no supervisorDoctrinal, plus interviews. Map the sandbox provisions onto the agency-as-deployer case, compare how the six existing agency tools were validated before use, and propose an institutional form. Sits directly on the institutional-arrangements strand
The privacy-utility tradeoff as an enforcement policy choiceDifferential privacy costs the most accuracy on rare cases, and cartels are rare cases. What epsilon may an enforcement authority lawfully choose?Measurable: synthetic procurement data at several privacy settings, measure screen performance and re-identification risk at each, then ask what the law has to say about selecting a point on that curve. Joins the GDPR chapter's synthetic-data project to this one

10. Your CV, mapped onto this chapter

From this chapterWhat you can genuinely claim
Agent reliability and step compoundingYour harness is a multi-step pipeline — generate, rule, probe, judge, vote — and you found a step-level error in v1 that corrupted the result. You have debugged exactly the failure mode this section describes, which is a stronger claim than having read about it
Decision-level provenanceYour JSONL log records prompt, configuration, output and three judge votes per run. Built out of necessity, and it is the artefact the documentation literature lacks
Prompt injection and jailbreaksYour poisoned condition is a prompt injection carrying a legal payload, and 15 of 15 runs were susceptible. Direct experimental experience of the mechanism
Chain-of-thought faithfulnessYour v1 error was caused by trusting reasoning traces as evidence of reasoning. You diagnosed it, re-ran, and published the corrected design. That is the literature's central caution, learned the hard way
Benchmark contamination and evaluation designYou tried keyword scoring, measured that it failed, and discarded it for a judged label with majority voting. A documented negative result, which is rarer and more credible than a clean positive one
Synthetic data and GANsYou have built diffusion and generative models on the technical side. Do not claim a privacy-preserving synthetic-data pipeline — you have not built one. Claim the generative-model experience and the project idea separately
Drift as an adversarial processYour three-condition design exists to separate outcomes that look identical from outside. That is the same methodological instinct the drift problem needs, and you can say so in one sentence

11. If you remember ten things

  1. An agent is a model in a loop choosing its own next action. At 95 percent per-step reliability, 20 steps is 36 percent — compounding, not capability, is the binding constraint.
  2. A digital twin needs a specific referent, a live feed and a validated mechanism. A market has none of the three. Agent-based modelling is the honest name for what people mean, and it supports possibility claims, not predictions.
  3. Mixture of experts means parameter count overstates compute per token, which complicates a systemic-risk threshold stated in FLOPs. Distillation means capability can be copied out from under the obligations attached to its source.
  4. Quantisation changes the outputs, so the deployed model is not the evaluated model, and conformity assessment has not caught up.
  5. Art 57 obliges Member States to run sandboxes. The Digital Omnibus, in force 27 July 2026, moved Annex III high-risk to 2 December 2027 and Annex I to 2 August 2028.
  6. Sandboxes assume a firm testing a product. Nobody sandboxes the agency's own screen, which is an institutional gap sitting on the ATLANTIS third strand.
  7. Model cards document systems; evidence law needs per-decision provenance. Art 12 logging is the closest EU hook. Your own JSONL log is already the right shape.
  8. Concept drift in enforcement is adversarial — deploying the screen causes the drift. Falling detections are equally consistent with deterrence and with evasion, and no amount of model monitoring tells them apart.
  9. Every privacy technique trades accuracy for protection, and differential privacy costs most on rare cases, which is what cartels are. Choosing epsilon is an enforcement policy decision made by default parameter.
  10. Interpretability is not explainability. The law wants an account a human can act on, and the counterfactual form — *had your turnover been lower you would not have been flagged* — is the one closest to what a court needs.