The Panel, Topic by Topic
All four of them, on every topic. Who each one is, what they actually probe, and the exact answer to speak — fourteen topics times four panellists, so no question arrives from a direction you have not already heard.
Listen
0. How to use this chapter
Earlier chapters give one probe per topic. This gives four, because four people are in the room and each of them hears the same answer differently. Every table below has one row per panellist: their likely question, and the words to say back.
1. Who the four of them actually are
| Panellist | What they work on | What they are testing | Their tell |
|---|---|---|---|
| Thibault Schrepel — chair, VU Amsterdam | Competition law and digital markets through complexity science: markets as systems that evolve rather than settle. Founded Stanford Computational Antitrust (80+ agencies). Built a knowledge graph of EC decisions 1977–2025. *Complexity-Minded Antitrust*; *The Future-Proof Fantasy of AI Regulation*; *Competition between AI foundation models* with Pentland; openness measurement with Pott; Dynamic Competition Initiative with Nicolas Petit. Hosts *Scaling Theory*. Currently working with agentic workflows, MCP and graphs | Whether you can carry an idea rather than repeat one. He is critical of *both* neoclassical models and the neo-Brandeisian critique, so he is not looking for a camp — he is looking for someone who reasons about dynamics | Will take something you said and push it one step further to see whether you can follow. Follow. Do not retreat to safety |
| Georgiana Mirza — VU Amsterdam | Digital ecosystems, competition law and market regulation. Has worked on EU data spaces at the intersection of fundamental rights, competition and innovation in digital ecosystems | Project fit and regulatory design. Whether your work serves the legal strands, and whether you understand the instrument as a designed system rather than a set of rules | Questions about how things would actually operate — who does what, under which instrument, with what data |
| Tijmen Wisman — VU Amsterdam | Privacy and surveillance law. Teaches legality and proportionality — *necessary in a democratic society*. Long record on IoT, RFID and smart-meter privacy. Connected to the civil-rights platform that brought the SyRI case, where the District Court of The Hague struck down a Dutch state fraud risk-scoring system in February 2020 for breaching Article 8 ECHR | Whether you think about the person on the receiving end. He will not be impressed by detection rates. He will want to know what happens to the firm or the individual who was wrongly flagged | Shifts the frame from *does it work* to *what does it do to people, and is that proportionate* |
| Catalina Goanta — Utrecht University | Associate Professor of Private Law and Technology. PI of the ERC Starting Grant HUMANads on content monetisation and platform-governance fairness. Previously ran the Maastricht Law and Tech Lab, with computer scientists resident in a law school. Proposed a legal compliance API for enforcing the DSA. Runs large multi-country empirical studies — on the order of a million social-media posts across four countries | Method. She does this work herself, at scale. She will ask how you collected it, what your ground truth is, and whether it reproduces | Interrupts the narrative to ask a concrete methodological question. This is the panellist most likely to catch an overclaim |
Where each of them sits relative to your work
2. Competition law
| Who | Their probe | Say this |
|---|---|---|
| Schrepel | *Is the object–effect distinction still doing useful work in digital markets?* | *For cartels, yes — it is what makes enforcement tractable, because by-object conduct skips the economic-proof stage entirely. For the dynamic conduct you write about, I think it strains. Object categories were drawn from conduct types that were stable enough to list, and algorithmic coordination does not obviously belong to any of them — it produces a supra-competitive outcome without an agreement to characterise. So the distinction is not wrong, it is under-inclusive at the frontier, and the question is whether you extend the object list, which is slow and legislative, or accept that the hard cases fall to effect and are therefore much more expensive to prove.* |
| Mirza | *Where does computational work actually fit in the enforcement workflow?* | *Three places, and only one is glamorous. Case selection — screening to decide where to look, which is where the screens sit. Investigation support — retrieval over a decisional corpus, as the French authority built. And ex post evaluation — did the remedy work, which is the stage almost nobody resources. I would argue the third is where computation is most underused and least legally risky, because it touches no individual firm's rights. Most of the attention goes to the first, which is the one with the gravest due-process exposure.* |
| Wisman | *A screen flags a firm. What has happened to that firm?* | *Something substantial, before any finding. A dawn raid under Article 20 of Regulation 1/2003, which is unannounced and covers premises, books and records. Legal costs. Internal disruption. Potentially a reputational hit if it becomes public. And the firm cannot see what triggered it — so it cannot contest the trigger, only the conclusion. That is the part that troubles me: the rights of defence attach to the proceeding, but the flag precedes the proceeding, and nothing in the instrument requires the authority to disclose what the model relied on. Two firms can receive identical flags with very different probabilities of guilt, and the output does not reveal which.* |
| Goanta | *Where would you get the data for any of this?* | *The decisions are public, so a corpus of Commission decisions is straightforward — and I would log per-document provenance rather than just downloading, because the finding is only auditable if the collection is. The procurement data that would actually make a screen testable is not public: it is confidential business information and often personal data. So the honest answer is that the interesting questions are access-constrained, not method-constrained, and that is why synthetic data becomes a workstream rather than a side project. I would want to know early what access the project has, because it determines what is answerable.* |
3. The AI Act
| Who | Their probe | Say this |
|---|---|---|
| Schrepel | *What do you make of the Act?* | *I would take your own framing, if I may — the future-proofing claim does not survive the text. It is like an immune system that detects pathogens but cannot produce antibodies: mechanisms for noticing change, very few for responding. Article 7 lets the Commission add to Annex III but not remove. Article 112(10) says may, not shall. And machine-readable reporting is essentially absent, so the Act cannot be monitored by the methods it regulates. The Digital Omnibus then moved the high-risk dates — Annex III to December 2027, Annex I to August 2028 — which is itself evidence that adaptation is happening through amendment rather than through the Act's own machinery.* |
| Mirza | *Where does an enforcement screen sit in the risk tiers?* | *Walk the pyramid and it falls out the bottom. Not prohibited — Article 5(1)(d) covers predicting criminal offences by natural persons from traits, and a screen assesses undertakings from conduct. Probably not high risk — Annex III point 6 covers law enforcement, but EU competition enforcement is administrative rather than criminal and an undertaking is not a natural person. Article 50 transparency is about telling a human they are talking to an AI, which is not the issue. So it plausibly lands in minimal risk, unregulated — the tier built for spam filters, holding a tool that triggers dawn raids and nine-figure fines. I think that is the most important gap in the instrument for this project.* |
| Wisman | *Does Article 14 human oversight actually protect anyone?* | *Only if the reviewer can do what it assumes. Article 14 requires that the person can correctly interpret the output and resist automation bias. My own work suggests the first condition can fail even with a completely willing reviewer. In fifteen of fifteen runs I steered a model's legal conclusion, and when I asked it why it ruled as it did, not one gave a cognizable account — the explanations were consistent, verifiable and incomplete in the same way. A reviewer handed a true but incomplete account has nothing to resist with. So Article 14 is drafted as a design requirement and it is really an empirical claim, and nobody is testing it.* |
| Goanta | *How would you test an Article 14 claim empirically?* | *Build a realistic multi-step pipeline, instrument every step, inject a known error at an intermediate step, and measure how often a reviewer approving only the final output catches it. The dependent variable is detection rate; the manipulation is error position and salience. It is a within-subjects design so each reviewer is their own control, which is what makes it viable at a sample size a PhD can actually recruit. The limitation I would state up front is that recruited reviewers are not case officers with career stakes, so the external validity is the weak point, not the design.* |
4. The DMA and the DSA
| Who | Their probe | Say this |
|---|---|---|
| Schrepel | *Does the DMA's ex ante design solve the adaptivity problem or restate it?* | *It relocates it. Under 101 and 102 you prove market definition and dominance case by case, which is slow; the DMA replaces that with counting — turnover, 45 million end users, 10,000 business users, sustained over three years. So the thresholds are the theory of harm, pre-committed in the text. That is what makes it tractable and computable. But a number fixed in legislation while the market moves is exactly the adaptivity problem in a new place, and Article 19 review requires an amendment. The layer-interaction point applies: if a training-efficiency advance changes what it costs to be a credible entrant, the user thresholds are measuring something different from what they were set to measure.* |
| Mirza | *Which of the two instruments is the better model for regulating agency tools?* | *The DSA, and specifically Article 40. It is the only place in EU law that compels a private actor to open its data to vetted outside researchers, with a procedure for who qualifies and what they may see. That is the template the agency-transparency argument needs, because the objection to auditing a screen is always confidentiality, and Article 40 is a worked example of a regime that grants access without publishing. The DMA is the wrong model here because it is addressed to firms by designation, and an authority is not designated.* |
| Wisman | *The DSA's systemic risk provisions — real or performative?* | *Partly performative as drafted, for a structural reason: the platform assesses its own systemic risk under Article 34, mitigates it under 35, and the audit under 37 is procured by the audited party. The independence is contractual, not institutional. What makes it more than performative is Article 40, because outside researchers are not paid by the subject. So my reading is that the accountability in the DSA comes almost entirely from the data-access provision rather than from the risk-assessment machinery — and if that is right, the same lesson applies to any oversight regime built for agency tools.* |
| Goanta | *You mention Article 40 — have you looked at what researchers actually get?* | *Not in detail, and I would not pretend otherwise — I know the provision and the delegated-act structure, not the practical experience of applicants. I would want to talk to people who have used it, because the gap between a data-access right on paper and a usable dataset is where this kind of provision succeeds or fails, and that gap is itself a research finding. If the project has contacts there I would treat that as a priority rather than reading more about it.* |
5. The GDPR — and this is Wisman's territory
| Who | Their probe | Say this |
|---|---|---|
| Wisman | *Is a cartel screen different from SyRI?* | *Structurally it is very close, and I think that is the most useful comparison available. In both cases the state builds a risk model, applies it to subjects who did not consent and cannot see it, and uses the output to decide whom to investigate. SyRI failed on transparency and proportionality under Article 8(2) rather than on accuracy — the court did not need to find the model wrong. Three differences matter. The subjects are undertakings rather than natural persons, so Article 8 and the GDPR engage differently and sometimes not at all. The consequence is a dawn raid rather than a benefits investigation, which is serious but differently so. And competition authorities have an express legal basis for investigation that the SyRI architecture lacked. But the core defect SyRI was struck down for — an opaque risk model the subject cannot interrogate — is unaddressed in the competition context, and nobody has brought that challenge. I think that is a paper.* |
| Wisman, follow-up | *So would a screen survive an Article 8 challenge?* | *I genuinely do not know, and I would rather say so than guess at a conclusion. What I can say is which way the pressure runs. The fact that the subject is a legal person weakens the Article 8 claim considerably, and business premises get less protection than a home — though not none, on the Strasbourg case law. The stronger route is probably not Article 8 at all but the rights of defence and equality of arms: the undertaking cannot contest a trigger it cannot see. That is a procedural-fairness argument rather than a privacy one, and it does not depend on the subject being human. I would want a legal-track colleague to pressure-test that, because it is at the edge of what I can responsibly claim.* |
| Schrepel | *Is the GDPR an obstacle to computational antitrust?* | *It is a constraint with an exception that is rarely used well. Article 6(1)(e) and Article 6(3) give a public authority a route, but the task must have a basis in law, and the purpose-limitation and data-minimisation principles bite against exactly the kind of broad exploratory analysis screening involves. The more interesting constraint is Article 5's accuracy principle, which is usually read as being about correct records — but a risk score is also personal data where individuals are identifiable, and an uncalibrated score is arguably inaccurate personal data. I have not seen that argument made, and it would be a GDPR route to the calibration point.* |
| Goanta | *Is synthetic data a real answer to the access problem?* | *A partial one, and the honest framing is that it trades a known problem for an unsettled one. It is genuinely useful for auditing — you can let outsiders test a screen's behaviour without exposing a single real firm. But whether synthetic data is personal data under the GDPR is unresolved, because faithful enough to be useful may be faithful enough to permit re-identification, and the answer probably depends on the generator and the privacy budget rather than on the category. And the specific failure mode matters here: if the generator has quietly learned only the common cases — mode collapse — then an audit conducted on synthetic data systematically misses the rare ones, which in this field are the cartels.* |
6. Probability and statistics
| Who | Their probe | Say this |
|---|---|---|
| Goanta | *Fifteen runs. What can that actually support?* | *Pooled claims, and nothing between models. It works at all because the design is paired — every scenario appears under control, ablation and poisoned conditions, so the scenario is its own control and the comparison happens within a unit rather than between groups. That is what McNemar exploits: nine runs changed from unbiased to biased and none went the other way, which gives p of about 0.004, and fifteen against zero gives about 0.00006. Between models it is unpaired and five runs each, and Fisher's exact gives 0.206, which licenses nothing. So I report pooled susceptibility and I do not rank the models.* |
| Goanta, follow-up | *Why Wilson intervals?* | *Because several proportions sat on the boundary — zero of five in control, five of five in poisoned. The ordinary normal approximation has poor coverage there, and at an observed five of five it collapses to zero width, which would assert certainty from five observations. Wilson keeps the interval inside zero to one and keeps it honest at the edges: four of five gives roughly 38 to 96 percent, one of five roughly 4 to 62 percent. Those overlap heavily, which is the formal reason I say the models are not separable rather than just declining to rank them.* |
| Schrepel | *Why should an economist care about calibration?* | *Because calibration is what makes a probability a usable input to a decision rule, and enforcement is a decision rule. An uncalibrated 0.9 and a calibrated 0.9 are different objects — only the second lets you compute the expected cost of acting. And it bears on your dynamics point: a screen calibrated on last year's market is calibrated to a data-generating process that has since changed, so calibration is not a property you establish once. It has a shelf life, and nobody states one.* |
| Wisman | *Explain to me why a ninety-percent-accurate screen is a problem.* | *With numbers, because it is counterintuitive. Suppose one tender in two hundred is rigged, the screen catches ninety percent of rigged tenders and wrongly flags five percent of clean ones. Take ten thousand tenders: fifty are rigged and it catches forty-five; of the nine thousand nine hundred and fifty clean ones it flags about four hundred and ninety-eight. So five hundred and forty-three firms are flagged and forty-five did something — about eight percent. A flag from a screen described as ninety percent accurate means a ninety-two percent chance the firm did nothing wrong. That is the prosecutor's fallacy with a procurement dataset attached, and the number that fixes it is the base rate, which the agency must supply and almost never states.* |
| Mirza | *Does that mean screens should not be used?* | *No — it means the threshold is a policy decision that is currently being made by a default parameter. The statistics are well understood: testing ten thousand tenders at a five percent level gives about five hundred false flags before a single cartel exists, and Benjamini–Hochberg or Bonferroni will control that. The unanswered question is legal, not statistical: what false discovery rate may an enforcement authority lawfully operate at? That is a question for an institution to answer and no instrument currently asks it, which is why I think it sits on the third strand rather than the first.* |
7. Machine learning
| Who | Their probe | Say this |
|---|---|---|
| Goanta | *What model would you actually use for a bid screen?* | *Gradient-boosted trees as the baseline, because on rows-and-columns data they generally beat neural networks and are cheaper and more reviewable. Logistic regression alongside it, not as a competitor but as the model you can explain — and I would report the gap between them as a number, because that gap is the price of interpretability stated honestly rather than asserted. But I would resist treating it as pure classification, because the labels are convictions, and convictions are detections rather than occurrences. A supervised model trained on them learns what was previously caught.* |
| Goanta, follow-up | *How would you validate it?* | *Leakage first, and leakage in this setting is mostly temporal and structural — so grouped splits by firm or sector so the same undertaking cannot be in train and test, and time-based splits so you are predicting forward rather than interpolating. Precision–recall rather than ROC, because ROC flatters you under heavy class imbalance. Then the base-rate calculation, calibration, flag rates by subgroup, and stability under resampling. If the set of flagged firms changes substantially when I refit on a slightly different sample, that is not a tuning issue — it means who gets investigated is partly arbitrary, and arbitrariness is a legal defect rather than a statistical one.* |
| Wisman | *Can you make such a model fair?* | *Not in the sense the question implies, and I think that is the important answer. There are at least five separate sources of unfairness — the historical data, the labels, the features, the objective, and the deployment — and then there is an impossibility result: demographic parity, equalised odds and calibration cannot generally all hold at once unless base rates are equal across groups, which they are not. So fairness is not one property you achieve; it is a set of mutually incompatible criteria and choosing between them is a normative decision. An engineer who picks one has made a political choice while appearing to make a technical one. The useful thing I can contribute is making that choice visible rather than pretending to dissolve it.* |
| Schrepel | *Explain overfitting to a lawyer.* | *A high-variance model trained on slightly different history would have flagged a different set of firms. So the choice of who gets investigated depends partly on an accident of which sample the model happened to see. That is not a performance problem, it is an arbitrariness problem, and arbitrariness in the exercise of public power is a legal defect independent of accuracy. I find that framing lands with lawyers immediately, because they already have a concept for decisions that vary without reason.* |
8. NLP and unstructured data
| Who | Their probe | Say this |
|---|---|---|
| Goanta | *How would you code a corpus of decisions?* | *Not with a generative model, which I think is the answer that signals judgment. A fine-tuned encoder — BERT-family — or TF-IDF with a linear classifier, on a hand-labelled sample, with a written codebook and a reported inter-annotator agreement statistic. If two humans cannot agree on the label, no model score computed against it means anything, so the agreement figure comes before the accuracy figure. Cheaper, reproducible, inspectable, and defensible in a methods section. Reaching for an LLM where a classifier suffices is the most common failure of judgment in applied legal NLP right now.* |
| Goanta, follow-up | *Your offence classifier reports 84 percent. On what?* | *336 of 400 benchmark sentences, and the caveat comes first: I wrote those sentences, so the test set and the author share a vocabulary and 84 percent is optimistic. A fair evaluation would use descriptions written by people who had not seen the provisions — ideally real complaint text. I would also report macro rather than micro F1, because across 358 sections micro is dominated by the frequent offences and can look strong while the model is blind to whole categories.* |
| Schrepel | *What breaks when you apply NLP to legal text?* | *Negation, party roles and outcome — which are the three things a legal summary exists to convey, and all three are invisible to overlap metrics. *The appeal was dismissed* and *the appeal was allowed* share almost every n-gram, so a summary that reverses the holding scores well on ROUGE. I tried keyword scoring in my own harness, measured that it failed because a model will name a party while ruling against them, and discarded it. That discarded approach is more informative than the one I kept, because it shows the metric and the legal meaning are orthogonal.* |
| Mirza | *Where is NLP genuinely ready for regulatory use?* | *Extraction with a fixed surface form, and retrieval with human verification. Dates, article numbers, case citations, fine amounts — those are regex and exact-match territory and the answer is either right or visibly wrong. Retrieval that returns a span the reviewer can see in the original document is also ready, because the system is pointing rather than asserting. What is not ready is anything where the output is a characterisation — *this conduct amounts to an abuse* — because there the fluency of the output is uncorrelated with its correctness and the reviewer has nothing to check it against.* |
9. LLMs, RAG and reasoning
| Who | Their probe | Say this |
|---|---|---|
| Schrepel | *Do these models reason?* | *Partly a definitional argument, so I would rather make the operational claim. The sceptical case is well put in Graepel's recent piece: AlphaGo had two mechanisms, a network supplying intuition and an explicit search over a game tree supplying deliberation — and move 37 was chosen by the search against the network's own instinct, which rated it about one in ten thousand. A language model is the first mechanism only, and chain-of-thought is not a second mechanism, it is the same next-token prediction run longer. His three objections are no persistent inspectable epistemic state, no separation between knowledge and its manipulation, and chains of thought composed after the fact. The third one is what I measured. So my position is: whether it reasons is contestable, but the claim that the visible reasoning is not reliably the causal reasoning is testable, and it failed in fifteen of fifteen runs.* |
| Schrepel, follow-up | *So what follows for regulation?* | *That the duty to give reasons and the machine-learning critique are the same objection. Graepel wants an auditable chain of evidence and inference rather than a persuasive account assembled afterwards, because in a high-stakes setting you must be able to locate what went wrong. Article 296 TFEU wants that, for the same reason, and Article 14 of the AI Act requires a reviewer who can correctly interpret the output. All three converge on one property: the account of a decision must be its actual provenance, not a plausible reconstruction. **The engineering literature calls it unfaithful chain of thought; administrative law calls it a defective statement of reasons. I have not seen anyone say they are the same defect.* |
| Goanta | *Describe your RAG pipeline and where it fails.* | *Load, chunk, embed, index, then at query time embed, retrieve, rerank, insert, generate, verify. Chunking decides most of the quality — for legal text I would chunk on structure, by paragraph and recital, because a fixed token window splits a holding from its reasoning. Failures in the order they actually bite: chunking severed the relevant passage; retrieval returned plausible but wrong material; the model answered from memory and ignored the context; and the citation points at a real document that does not say what is claimed. Three of those four are retrieval and formatting problems, not generation problems, which is the opposite of where attention goes.* |
| Wisman | *What is the danger of an agency using one of these?* | *That it produces an account which is true, verifiable and incomplete — and that is worse than an unstable one. An unstable system contradicts itself and a contradiction is somewhere to start looking. What I got instead was consistency: every answer agreeing with every other, all of them omitting the thing that actually drove the outcome. Article 296 may be satisfied on the face of a reason like that, and an Article 14 reviewer cannot resist what they cannot detect. For someone on the receiving end, a system that is coherently incomplete leaves them nothing to appeal against.* |
10. Logic and the symbolic layer
| Who | Their probe | Say this |
|---|---|---|
| Schrepel | *Why a symbolic layer rather than a bigger model?* | *Because the property I want is not accuracy, it is a derivation. A classifier outputs 0.84 and cannot tell you why that rather than 0.79. A derivation either goes through or it does not, and when it fails it names the element that was unestablished. In my own work the engine could say that `affectsTrade` was the missing conjunct — that is an instruction to a case team rather than a score. And it maps onto the structure of law directly, because a statutory test is a conjunction of elements each of which somebody must prove. Graepel arrives at the same place from the machine-learning side: his second objection is that knowledge and its manipulation are entangled in the weights, and a symbolic layer is precisely the separation he is asking for.* |
| Schrepel, follow-up | *And the limitation?* | *Somebody has to write the rules, which needs legal expertise per provision, and it degrades on genuinely ambiguous language — my own paper documents both. Which is why the interesting architecture is neurosymbolic rather than symbolic: a language model to extract the elements from messy text, a logic layer to check whether they are satisfied. The model handles language, the logic handles law. I have built both halves separately and found the failure mode of each, which is a reasonable basis for proposing the combination, though I have not built the combination.* |
| Mirza | *Is formalising provisions realistic at scale?* | *Not across a whole corpus, and I would not propose it. It is realistic for narrow, high-volume, element-structured tests — the ones where the same conjunction gets applied repeatedly and the cost of formalising is amortised. Jurisdictional thresholds, limitation periods, notification triggers. What formalisation reliably produces even at small scale is a consistency check: encode the findings in a decision and a solver will tell you if they contradict, which either passes or fails rather than returning a confidence. That is a modest claim and it is deliverable.* |
| Wisman | *Does formalisation remove discretion that ought to be human?* | *It can, and that is the real risk rather than inaccuracy. Deontic logic makes the distinction sharp: an obligation, a prohibition, and the state where both an act and its omission are permitted — which is a discretion. Prose drafting constantly leaves that third state unclear, and a formalisation has to pick one. If it resolves a genuine discretion into a rule, it has made a legislative choice disguised as an implementation detail. So I would treat formalisation as a tool for surfacing where the discretion sits, not for closing it — and the fact that Article 112(10) of the AI Act says *may* rather than *shall* is exactly that third state, which is why it matters so much.* |
11. Deep learning and graphs
| Who | Their probe | Say this |
|---|---|---|
| Schrepel | *You mention graphs. Which sense?* | *Both, and they are different tools. A knowledge graph is a queryable store of entities and typed relations — no training, exact answers. Your Commission-decisions graph is that, and the questions it answers are ones retrieval handles badly: *which undertakings were fined more than once in the same sector* is a traversal, and a RAG system will miscount it confidently because counting is not what next-token prediction does. A graph neural network is a learned model passing messages over a graph. The rule I would state: a graph for questions about relations across documents, retrieval for questions about what a document says. I have built retrieval, not graphs, and I think the distinction is where a lot of current disappointment comes from.* |
| Mirza | *Why would a graph model fit cartel detection?* | *Because the signal is relational and therefore invisible to a tabular model by construction. Nodes are firms, edges are co-bidding weighted by frequency. After one round of message passing a firm knows who it competes with; after two it knows about its competitors' competitors — and that is where a closed bidding ring shows up, as a group that bids against each other constantly and almost never against anyone outside. That property is not an attribute of any single firm or any single tender, so no row of data contains it. The constraint is real though: two or three rounds is the limit, because beyond that over-smoothing makes every node's representation converge and destroys the structure you were looking for.* |
| Goanta | *Have you built one?* | *No. I have not built a graph neural network and I would not want to imply otherwise. What I can defend is why the architecture fits the problem and what the depth constraint is. I have built retrieval pipelines in production and an experiment that breaks one, so the retrieval side I can take several questions deep; graphs I understand structurally and have not implemented.* |
| Wisman | *A relational model implicates people who were not investigated.* | *Yes — and I think that is the sharpest objection to the architecture, sharper than accuracy. A graph model scores a firm partly on the basis of who it is connected to, so a firm with no suspicious conduct of its own can be flagged because of its neighbourhood. That is guilt by association, formalised, and it is close to what the UN Special Rapporteur objected to about SyRI — targeting by neighbourhood without individual suspicion. The counterfactual explanation gets much harder too: *had your turnover been lower you would not have been flagged* is answerable; *had your competitors been differently connected to each other* is not something the firm could have acted on. I would want that in the paper, not as a caveat at the end.* |
12. Agents, MCP and the GenAI stack
| Who | Their probe | Say this |
|---|---|---|
| Schrepel | *What do you make of MCP?* | *That it is a privately supplied interoperability standard doing voluntarily what the DMA tries to mandate. Articles 6(4) and 6(7) compel interoperability by law, after litigation; MCP achieved cross-vendor tool interoperability in months by convention, because every vendor benefited from the same ecosystem of servers. That is a live natural experiment in standards against regulation, and it cuts both ways — voluntary is faster and unenforceable, a mandate is slower and binding. The question I would want to ask rather than answer is who governs the protocol: it is open but not neutral, and whoever shapes the specification shapes what models can easily reach. That is standard-setting power without the standard-setting institutions.* |
| Schrepel, follow-up | *And the layer-interaction argument — do you buy it?* | *Yes, and I would extend it rather than repeat it. Your Nvidia illustration is that a share figure at the chip layer measures something the model layer is busy redefining — DeepSeek reportedly trained on around two thousand chips for about 5.6 million dollars against tens of thousands and hundreds of millions for comparable models. The extension I think is mine: the same argument applies to the agency's own instruments. If a screen is calibrated on a market whose technical substrate is shifting underneath it, the screen's error rate is a quantity with a shelf life — and no agency states one. Your point about market analysis becomes a point about enforcement tools, which is the cross-project position's job.* |
| Mirza | *Are agentic systems usable by a regulator today?* | *Not for multi-step work, and the reason is arithmetic rather than capability. At ninety-five percent reliability per step — which is generous — five steps gives about seventy-seven percent, ten gives sixty, twenty gives thirty-six. Compounding eats the result before capability does, and because the output is fluent either way the failed runs look exactly like the successful ones. The sharper regulatory question is where the decision is: if a pipeline has twenty automated steps and one human approval, Article 22 of the GDPR and Article 14 of the AI Act both assume a single identifiable decision point, and an agent loop does not have one. That is a formally compliant system in which no human reviewed anything that mattered.* |
| Wisman | *Who is accountable when an agent acts?* | *Unresolved, and I think the protocol boundary is where the answer has to live. MCP separates resources, which are read, from tools, which act — and that distinction is the one Article 14 needs and does not make. Reading a file is not issuing a request for information. An oversight regime that treats them alike is either too strict to be workable or too loose to protect anyone. Practically, the client can allow reads and require human approval for actions, which is the technical form of meaningful oversight; and because every call crosses the protocol as structured data, that boundary is also the natural place to log. So the accountability question has an architectural answer available, and nobody has drafted it.* |
13. Your own work, and the questions about you
| Who | Their probe | Say this |
|---|---|---|
| Goanta | *Walk me through your harness.* | *Five scenarios, three conditions, fifteen runs. Control is a clean context. Ablation is partial manipulation. Poisoned plants fabricated authority in the retrieved context. Each run produces a ruling, then a forced-choice probe with four options of which exactly one is an admission — scored by reading which letter came back, with no model involved. Bias is graded by a judge model on the final answer only, never the reasoning trace, with a majority vote over three calls and the agreement rate recorded. Temperature fixed at zero, and I still vote, because API serving is not fully deterministic even at zero. **Every run writes a JSONL row with the prompt, configuration, raw output and all three votes.* |
| Goanta, follow-up | *What went wrong in version one?* | *The graders saw the subject's full reasoning trace rather than its final ruling, and the trace restated the planted instruction almost verbatim. So any model that produced visible reasoning was nearly guaranteed a biased grade — I was measuring verbosity, not steering. I found it because the numbers were suspiciously clean and tracked which models were verbose rather than which were manipulated, and that correlation did not belong there. Version two shows graders final answers only. The result changed, and the corrected numbers are the ones I quote. The discarded version is the more informative half, because it is a clean instance of the chain-of-thought faithfulness problem hitting my own instrumentation.* |
| Schrepel | *Does mens rea transplant to AI systems?* | *No, and I want to be clear my paper is not arguing that it does. EU competition liability is largely objective — the undertaking is liable for conduct, intention bears on fine-setting rather than on establishing the infringement, and the AI Act places duties on providers and deployers precisely because it does not treat a model as a subject. What my result actually bears on is the duty to give reasons, not culpability. The finding is that the system's account of its own decision is consistently true and incomplete, which is a problem for Article 296 and for Article 14 oversight. The mens rea framing was my route into the question; it is not the conclusion.* |
| Wisman | *Why does your second finding matter more than the first?* | *Because susceptibility is a known problem with a known direction of travel, and the second finding is a problem with no obvious fix. Zero of fifteen runs gave a cognizable account of why they ruled as they did — and critically, the two that did admit anything were both in the five rows where the judge was also the subject; in the ten rows graded by a different model there were none. An unstable confession would at least leave a contradiction to investigate. Consistent partial disclosure leaves a reviewer with nothing to pull on. **For the person or firm affected, that is the difference between a decision you can challenge and one you cannot.* |
| Mirza | *Why the computational position rather than a legal one?* | *Because the gap I keep hitting is not doctrinal. The legal questions in this project are well posed. What is missing is someone who can say how much a technical result actually supports — which claims survive contact with the method and which do not. I am more useful producing the measurements the legal strands need, and being the person who says when a measurement will not bear the argument. That is also why I think the three strands connect more tightly than they look: accuracy is not a property of a model, it is a property of a model inside a monitoring regime, which makes it partly institutional.* |
14. Fit, commitment and the awkward ones
| Who | Their probe | Say this |
|---|---|---|
| Any of them | *You have no law degree — wait, you do. Explain the three degrees.* | *Computer applications, then law, then data science. It looks like switching and it was not — I kept circling the same question from different sides. The law degree is a full LL.B., so I read instruments properly; what I do not have is the training to read a case the way a litigator does, and I would expect to learn that from the legal tracks. What I bring is the ability to design the measurement and then say honestly what it supports.* |
| Schrepel | *Why are you committed for four years?* | *Three reasons and the third is the real one. It is the only thing that has held my interest across three degrees. I want to leave something behind rather than ship more tools. And I think legal AI is genuinely underexplored, and I think I know why — it needs real knowledge of both sides, which is a barrier most people do not cross, because the computational people cannot read a statute and the lawyers cannot read a model. Legal data is also not like healthcare data: no structured ground truth, no labels, no equivalent of a diagnosis. And the task is reasoning rather than prediction, which is why I do not think scaling a language model gets there and why I think the answer involves putting a symbolic layer back in. That is a four-year question. It is also why I came to VU for the logic course in 2024, before any of this existed.* |
| Goanta | *What would you need to learn?* | *Doctrine, with no pretence otherwise — I read instruments well and I have not been trained to read a case like a lawyer. And scale: my own work has been small and careful, fifteen runs with exact tests, and this project needs corpus-level work where the statistical problems are different and multiple comparisons become existential rather than a footnote. I would rather say that now than discover it in year two. What I would not need to learn is how to be sceptical about my own measurements.* |
| Wisman | *You built a tool that predicts judge behaviour.* | *I did, for India, and I am glad you raised it because it is the most interesting thing on my CV from a comparative angle. France criminalised exactly that — Article 33 of the 2019 justice reform law prohibits reusing judges' identity data to evaluate, analyse, compare or predict their professional practices, carrying up to five years' imprisonment, and it is reported to be the first ban of its kind. So the same technique is a commercial product in the United States, a criminal offence in France, and unregulated in India. Having built one, my view is that the French position is more defensible than it first looks, because judicial analytics does not just describe a judge, it creates an incentive to forum-shop and eventually to perform consistency — the measurement changes the thing measured. That is a reason to regulate the use rather than the technique, and it is a tension I would want to work on rather than defend.* |
| Any of them | *Do you have experience with [something you have not built]?* | *No, followed by something useful.* *No, I have not built one. My understanding is that it does X, and for this problem I think it would or would not fit because Y.* Never guess. Not a digital twin, not an agent-based model, not a graph neural network, not blockchain, not federated learning, not differential privacy, not zero-knowledge anything, not a knowledge graph, not an MCP server. **This panel could take any of those apart in two questions.* |
15. If you remember ten things
- Four people, four frames. Schrepel: dynamics. Mirza: regulatory design. Wisman: the person affected. Goanta: the method. Same fact, different framing.
- Calibration is the one answer responsive to all four. It is a dynamics problem, a regulatory gap, a proportionality problem and a measurement problem simultaneously.
- Know SyRI. The Hague District Court, February 2020, Article 8 ECHR, struck down for opacity about the risk indicators and the model, not for inaccuracy. Raise it yourself with Wisman — it is your thesis in his jurisdiction.
- Goanta will test method, not knowledge. Lead every answer to her with the limitation: you wrote the benchmark sentences; 84 percent is optimistic; macro not micro.
- Schrepel will push one step further. Follow him. The extension that is yours: layer interaction applies to the agency's own instruments, so a screen's error rate has a shelf life nobody states.
- **Wisman reframes from *does it work* to *what does it do to someone*.** For the graph question, the honest answer is that relational scoring is guilt by association formalised.
- Mirza asks how it would operate. Name the instrument, the actor and the data. Article 40 DSA is your template for agency transparency.
- Mens rea does not transplant. Liability is largely objective. Your result bears on Article 296 and Article 14, not on culpability. Say so before anyone pushes.
- On reasoning: make the operational claim, not the definitional one. Not *LLMs cannot reason* but *the visible reasoning is not reliably the causal reasoning, and I measured that*.
- Raise the France judge-analytics point yourself. US product, French crime, Indian gap — and you built one, so you have a view. It converts your biggest liability into your most interesting answer.