Frontier Topics & Jargon
The words that get used in these conversations without ever being defined — agents, digital twins, mixture of experts, quantisation, sandboxes, model cards, drift. Each defined plainly, with what it means for enforcement and where the hype outruns the evidence.
0. Why this chapter exists
Everything so far has been a subject with a textbook. This chapter is different: it is the vocabulary that circulates in exactly the kind of conversation you are walking into, usually undefined, often by people who are themselves unsure. You will hear several of these words on 7 October.
1. Agents and agentic systems
An agent, in current usage, is a language model placed in a loop where it can call tools, observe the results, and decide what to do next — rather than producing a single response and stopping.
The agent loop
| Term | What it means |
|---|---|
| Tool use, function calling | The model emits a structured request — a search query, a database call — which the surrounding program executes and feeds back. The model does not *do* anything itself; it asks |
| ReAct | Reason, then act, then observe, then repeat. The pattern most agent frameworks implement |
| Multi-agent system | Several models with different assigned roles passing messages. Sometimes genuinely useful, often a way of hiding that no single step is reliable |
| Planner-executor | One model breaks a goal into steps, another carries them out |
| Human in the loop | A person approves actions before they take effect. The design choice with actual legal significance |
| Guardrails | Rules outside the model constraining what it may do. A parser, a permission list, a cost limit. Not a prompt |
2. Digital twins, and why the term travels badly
A digital twin is a computational model of a specific real system, kept synchronised with it by a live data feed, used to simulate what would happen under conditions you have not tried. The term comes from engineering: a twin of a particular jet engine, fed by that engine's sensors.
| Requirement | Why it matters | Does a market have it? |
|---|---|---|
| A specific real referent | A twin of a thing, not a class of things | A market is a contested legal construct, not an object. Market definition is itself litigated |
| A live data feed | What makes it a twin rather than a model | Agencies get data in response to requests, with delays measured in months |
| A validated mechanism | You must know the system's physics to simulate it | Firm behaviour has no equivalent of physics. Competing economic models give opposite predictions |
| Counterfactual fidelity | The simulation must be right about conditions never observed | This is the whole point and the whole problem. Nobody can validate a counterfactual that did not occur |
3. How large models are actually built and shipped
These five terms come up constantly and are rarely explained. Each is an engineering choice with a consequence for whether a system can be audited.
| Term | Definition | Why a regulator should care |
|---|---|---|
| Mixture of experts | The network contains many parallel sub-networks, and a routing layer sends each token to only a few of them. A model with 400 billion parameters might use 30 billion per token | The model's behaviour depends on routing, so two near-identical inputs can be processed by different sub-networks. Reproducibility and explanation both get harder, and parameter count stops being a meaningful measure of the compute actually used — which matters because the AI Act's systemic-risk threshold is stated in training compute |
| Distillation | Train a small model to imitate a large one's outputs. The small model inherits much of the behaviour at a fraction of the cost | Obligations attach to the model that was trained, not to its imitation. A distilled copy of a regulated model may fall under thresholds while reproducing the capability. Also the main route by which a provider's safety work is or is not inherited |
| Quantisation | Store the weights at lower numerical precision — 8-bit or 4-bit instead of 16-bit — making the model smaller and faster | It changes the outputs. A quantised model is not the model that was evaluated, so conformity assessment performed on the full-precision version does not straightforwardly cover what was actually deployed. Nobody has litigated this |
| Fine-tuning a released model | Taking an open-weights model and training it further on your own data | Shifts responsibility. Under the AI Act a party who substantially modifies a system can become its provider, and fine-tuning is the ordinary case of that |
| Retrieval instead of training | Rather than fine-tune, give the model the documents at query time | The reason RAG dominates in legal applications: the knowledge is inspectable and replaceable and you can point at the source. It is an auditability decision disguised as an engineering one |
| Term | Definition |
|---|---|
| Open weights | The trained parameters are downloadable. Not the same as open source — the training data and code usually are not released, so you can run and modify the model but not reproduce it |
| Context length | How many tokens fit in one prompt. Long-context models reduce but do not remove the need for retrieval, because attention cost grows with the square of length |
| Multimodal | Handles more than text — images, audio, video. Relevant here for interface screenshots in dark-pattern work and for scanned documents |
| Inference | Running a trained model. Distinguished from training, which produced it. Most cost and most regulatory surface is inference, because that is where it touches people |
| Scaling laws | The empirical finding that loss falls predictably with model size, data and compute. The basis of the whole industry's capital allocation, and an extrapolation rather than a theorem |
4. Regulatory sandboxes, and the dates that changed
A regulatory sandbox is a supervised environment where a provider may develop and test a system under a regulator's oversight, with some flexibility on how rules apply during the test. The AI Act requires them, which is unusual — most sandboxes are optional national initiatives.
| Provision | What it does |
|---|---|
| Art 57 | Each Member State must establish at least one AI regulatory sandbox at national level. A duty, not an option — and it is one of the few places the Act obliges states to build capacity rather than obliging firms to comply |
| Art 58 | Detailed arrangements, set out in Commission implementing acts |
| Art 59 | Processing of personal data lawfully collected for other purposes, in the sandbox, for developing certain public-interest AI systems — subject to conditions |
| Art 60 | Testing of high-risk systems in real-world conditions outside a sandbox, with safeguards and informed consent |
| Art 62 | Measures for SMEs and start-ups, including priority sandbox access |
5. Documentation artefacts, and what each is worth
A cluster of proposals exists for documenting models and datasets. All are voluntary in origin; several are now partially mandatory through the AI Act. They matter to you because documentation is where the technical and legal tracks meet, and because your own methodological instincts already point this way.
| Artefact | What it records | Status |
|---|---|---|
| Model card | Intended use, out-of-scope use, training data summary, evaluation results broken down by subgroup, known limitations | Voluntary, industry-standard, and the ancestor of the AI Act's technical documentation requirement |
| Datasheet for datasets | How the data was collected, by whom, with what consent, what it excludes, what it should not be used for | Voluntary. The most useful of the lot for legal purposes, because most failures trace to the data rather than the model |
| System card | Documents the deployed system rather than the model — including the surrounding guardrails and human review | Voluntary, and closer to what Art 14 oversight actually needs |
| Annex IV technical documentation | The AI Act's mandatory file for high-risk systems: design, architecture, data, metrics, oversight measures | Mandatory, from the dates above |
| Art 11 and 12 | Technical documentation and automatic logging of events over the system's lifetime | Mandatory. Logging is the provision with real forensic potential, and the least discussed |
| Art 13 | Instructions for use, sufficient for the deployer to interpret output | Mandatory, and the hinge between provider and deployer responsibility |
6. Drift, contamination and the ways evaluations lie
Four failure modes that sit between a reported benchmark number and what a system does in use. Each has a precise name, and each is a reason not to believe a performance claim.
| Failure | What happens | Consequence |
|---|---|---|
| Data drift | The input distribution changes after deployment — new procurement rules, a new tendering platform, inflation moving contract values | The model is unchanged and its accuracy falls anyway. The model did not break; the world moved |
| Concept drift | The relationship between inputs and the right answer changes. Cartels adapt to the screen that was catching them | Worse than data drift, because the historical labels are now actively misleading. And in enforcement it is adversarial: deploying the screen causes the drift |
| Benchmark contamination | The evaluation set was in the training data, so the reported score measures memorisation | A model can score highly on a legal benchmark and fail on a case decided after its training cut-off. This is why a published leaderboard number tells an agency almost nothing |
| Goodhart's law | Once a metric becomes the target it stops measuring what it measured. Optimise for flag precision and the screen learns to flag only obvious cases | The metric improves and the enforcement outcome worsens. The bid-rigging version: a screen tuned to look precise stops finding novel schemes |
7. Privacy and verification techniques you should be able to name
These appear whenever the answer to a governance problem is *technical*. Each solves something real, and each has a limitation worth knowing, because the limitation is usually where the legal question lives.
| Technique | What it does | The catch |
|---|---|---|
| Differential privacy | Adds calibrated noise so that any one individual's presence in the data cannot be detected from the output. Comes with a tunable parameter, epsilon, quantifying the leakage | Noise costs accuracy, and it costs it most on rare cases. Cartels are rare cases. The privacy-utility tradeoff is not a detail here, it is the finding |
| Federated learning | Train across several parties' data without moving it; only model updates are shared | Updates leak information, so it needs differential privacy on top. And it does not solve the legal problem of who may process what — it changes where the processing happens, not its lawfulness |
| Synthetic data | Generate artificial records with the real data's statistical structure. The GAN route from the deep learning chapter | Faithful enough to be useful may be faithful enough to re-identify. Whether synthetic data is personal data under the GDPR is genuinely unsettled, and that is a live question rather than a settled one |
| Secure multi-party computation, homomorphic encryption | Compute over data nobody can read, including the party doing the computing | Expensive, and limited in the operations supported. Real, but rarely the practical answer today |
| Zero-knowledge proofs | Prove a statement is true without revealing why. In principle: prove a screen was run correctly without disclosing the screen | Early for this use, and the proof covers the computation, not whether the computation was the right one to run |
| Audit by behaviour | Query the system and study the outputs, without access to weights or data. What your own harness does | Cannot distinguish mechanisms that produce the same behaviour — which is exactly the limitation you already state in your own paper |
8. Fast glossary — the rest of it
Terms you may hear once, needing only a sentence each. Grouped so you can find them.
| Term | One sentence |
|---|---|
| Foundation model | A large model trained broadly and then adapted to many tasks. The AI Act's near-equivalent is general-purpose AI model, which is the term to use in a legal sentence |
| Systemic risk | An AI Act category for general-purpose models above a compute threshold, carrying extra obligations. The threshold is stated in training FLOPs, which is why mixture-of-experts routing complicates it |
| Emergent capability | A capability that appears abruptly with scale rather than improving gradually. Contested — some apparent emergence is an artefact of the metric used |
| Alignment | Making a model behave as intended. RLHF and DPO are alignment techniques; the word also names the broader research field |
| Red teaming | Deliberately attacking a system to find failures before deployment. Required in substance for systemic-risk models |
| Jailbreak | A prompt that defeats a model's restrictions. Relevant to you because the poisoned condition in your own study is structurally a jailbreak with a legal payload |
| Prompt injection | Instructions hidden in retrieved content that the model follows as if they came from the user. The security problem RAG creates, and it is unsolved |
| Hallucination, confabulation | Fluent output that is false. The second word is better because it does not imply perception |
| Sycophancy | Agreeing with the user rather than being correct. Measured, and a direct product of training on human preference |
| Chain of thought | Intermediate reasoning text produced before an answer. Not a reliable account of how the answer was produced |
| Reasoning model | A model trained to spend more computation at inference on intermediate steps. Better at verifiable tasks; the faithfulness problem is unchanged |
| Interpretability, explainability | Understanding how a model computes versus producing an account a human can use. These are different, and the law almost always means the second |
| Mechanistic interpretability | Reverse-engineering the actual circuits inside a network. Real progress, nowhere near courtroom-ready |
| Saliency, SHAP, LIME | Post-hoc methods attributing an output to input features. Widely used, and known to be unstable — different methods disagree on the same prediction, which is itself the legal problem |
| Counterfactual explanation | *Had your turnover been below X, you would not have been flagged.* The most legally useful explanation format, and closest to what a court wants |
| Vector database | Storage built for similarity search over embeddings. The retrieval half of RAG |
| Embedding | A list of numbers representing meaning, where nearby lists mean similar things |
| Token | The unit a language model reads and writes — roughly a word piece, not a word |
| Edge, on-device | Running a model locally rather than on a server. Relevant because it keeps data in the agency's hands |
| MLOps | The engineering practice of deploying and monitoring models. Where drift detection actually lives, and usually absent in public-sector deployments |
| Reproducibility | Same code, same data, same result. Harder than it sounds with these systems, and the reason your harness pins configuration and logs every call |
9. What is unexplored here, and four projects
Each of these comes out of this chapter rather than from a literature review, each is doable in a PhD, and each has a legal half and a technical half.
| Project | The question | How you would actually do it |
|---|---|---|
| Where is the decision in an agent loop? | If an agency's pipeline has twenty automated steps and one human approval, has a human meaningfully overseen anything? Art 22 GDPR and Art 14 AI Act both assume a single identifiable decision point | Build a realistic multi-step enforcement pipeline. Instrument every step. Measure how often a reviewer approving only the final output would have caught an error injected at step three. An empirical test of a legal standard, and the natural next experiment after your own paper |
| Does conformity assessment survive quantisation? | A system is evaluated at full precision and deployed quantised. Does Annex IV documentation cover what was deployed? | Take an open-weights model, evaluate a legal classification task at 16-bit, 8-bit and 4-bit, and report where the outputs diverge. Then read the conformity provisions against the measurements. Cheap to run, and nobody has run it |
| Who sandboxes the regulator? | Arts 57 to 60 assume a firm testing a product. Enforcement screening is the state deploying a high-risk system against undertakings, with no sandbox and no supervisor | Doctrinal, plus interviews. Map the sandbox provisions onto the agency-as-deployer case, compare how the six existing agency tools were validated before use, and propose an institutional form. Sits directly on the institutional-arrangements strand |
| The privacy-utility tradeoff as an enforcement policy choice | Differential privacy costs the most accuracy on rare cases, and cartels are rare cases. What epsilon may an enforcement authority lawfully choose? | Measurable: synthetic procurement data at several privacy settings, measure screen performance and re-identification risk at each, then ask what the law has to say about selecting a point on that curve. Joins the GDPR chapter's synthetic-data project to this one |
10. Your CV, mapped onto this chapter
| From this chapter | What you can genuinely claim |
|---|---|
| Agent reliability and step compounding | Your harness is a multi-step pipeline — generate, rule, probe, judge, vote — and you found a step-level error in v1 that corrupted the result. You have debugged exactly the failure mode this section describes, which is a stronger claim than having read about it |
| Decision-level provenance | Your JSONL log records prompt, configuration, output and three judge votes per run. Built out of necessity, and it is the artefact the documentation literature lacks |
| Prompt injection and jailbreaks | Your poisoned condition is a prompt injection carrying a legal payload, and 15 of 15 runs were susceptible. Direct experimental experience of the mechanism |
| Chain-of-thought faithfulness | Your v1 error was caused by trusting reasoning traces as evidence of reasoning. You diagnosed it, re-ran, and published the corrected design. That is the literature's central caution, learned the hard way |
| Benchmark contamination and evaluation design | You tried keyword scoring, measured that it failed, and discarded it for a judged label with majority voting. A documented negative result, which is rarer and more credible than a clean positive one |
| Synthetic data and GANs | You have built diffusion and generative models on the technical side. Do not claim a privacy-preserving synthetic-data pipeline — you have not built one. Claim the generative-model experience and the project idea separately |
| Drift as an adversarial process | Your three-condition design exists to separate outcomes that look identical from outside. That is the same methodological instinct the drift problem needs, and you can say so in one sentence |
11. If you remember ten things
- An agent is a model in a loop choosing its own next action. At 95 percent per-step reliability, 20 steps is 36 percent — compounding, not capability, is the binding constraint.
- A digital twin needs a specific referent, a live feed and a validated mechanism. A market has none of the three. Agent-based modelling is the honest name for what people mean, and it supports possibility claims, not predictions.
- Mixture of experts means parameter count overstates compute per token, which complicates a systemic-risk threshold stated in FLOPs. Distillation means capability can be copied out from under the obligations attached to its source.
- Quantisation changes the outputs, so the deployed model is not the evaluated model, and conformity assessment has not caught up.
- Art 57 obliges Member States to run sandboxes. The Digital Omnibus, in force 27 July 2026, moved Annex III high-risk to 2 December 2027 and Annex I to 2 August 2028.
- Sandboxes assume a firm testing a product. Nobody sandboxes the agency's own screen, which is an institutional gap sitting on the ATLANTIS third strand.
- Model cards document systems; evidence law needs per-decision provenance. Art 12 logging is the closest EU hook. Your own JSONL log is already the right shape.
- Concept drift in enforcement is adversarial — deploying the screen causes the drift. Falling detections are equally consistent with deterrence and with evasion, and no amount of model monitoring tells them apart.
- Every privacy technique trades accuracy for protection, and differential privacy costs most on rare cases, which is what cartels are. Choosing epsilon is an enforcement policy decision made by default parameter.
- Interpretability is not explainability. The law wants an account a human can act on, and the counterfactual form — *had your turnover been lower you would not have been flagged* — is the one closest to what a court needs.