The Interview Brief
Twenty-five minutes, four people, one opener. The question bank with answers, the story about the time something broke, the three questions to ask them, and the six traps to avoid.
0. The facts of the day
| When | Wednesday 7 October 2026, 09:00 Amsterdam time. That is 12:30 IST — check it again the evening before, because European clocks change in late October and your margin for error here is zero |
| How long | 25 minutes. Short. This is a fit-and-judgment conversation, not an examination |
| Who | Thibault Schrepel chairing, with Georgiana Mirza, Tijmen Wisman, and Catalina Goanta from Utrecht |
| What for | One of three fully funded PhD positions on ATLANTIS — ERC Consolidator Grant, roughly €2 million, running 2026 to 2031, at the VU Amsterdam Faculty of Law. Vacancy 5706, grant 101228709 |
| Which position | The cross-project computational one. The other two are legal-track: one on the data problem, one on the fairness problem. Yours is the position that serves all three strands |
| Terms | Start date 1 December 2026. €3,059 to €3,881 per month |
1. The ninety-second opener
They will open with some version of *tell us about yourself* or *why this project*. This is the only answer you can fully script, so script it. Four beats, ninety seconds, ending on a question-shaped hook rather than a full stop.
2. The five numbers, and the five limitations
If they engage with your paper at all, they will engage with the numbers. Know these cold, and know the limitations better than the results — a candidate who volunteers the weaknesses of their own study is read as a researcher rather than an applicant.
| Result | The number | Say it like this |
|---|---|---|
| Control condition | 0 of 15 biased | Clean context, no bias. The design has a working baseline, which is what makes the rest interpretable |
| Ablation condition | 9 of 15 pooled — 4, 4 and 1 of 5 across the three models | Partial manipulation produced partial effect. The spread across models is wide and I cannot separate them at this sample size |
| Poisoned condition | 15 of 15 biased | Every run. Fabricated authority in the context produced a biased ruling every single time |
| Admissions | 2 of 15, and both self-judged | Two admissions, both in the five rows where the judge model was also the subject. In the ten rows graded by a different model, none |
| Cognizable intent | 0 of 15 | Not one run gave an account that identified what actually drove the ruling. This is the finding, not the susceptibility rate |
| Statistic | Value | What it licenses |
|---|---|---|
| McNemar, control to ablation | p ≈ 0.004 | 9 changes in one direction, 0 back. The paired design is what makes this possible at 15 runs |
| McNemar, control to poisoned | p ≈ 0.00006 | 15 changes, 0 reversals |
| Fisher's exact, model versus model | p = 0.206 | Nothing. This is why I make pooled claims and no model rankings |
| Wilson, 4 of 5 | ≈ 38 to 96 percent | Why I used Wilson rather than the normal approximation — several proportions sit on the boundary |
| Wilson, 1 of 5 | ≈ 4 to 62 percent | These two intervals overlap heavily, which is the formal reason the models are not separable |
3. The time something broke — and this is your best story
Almost every research interview asks some version of *tell us about a time something went wrong*. You have an unusually good answer, because the failure was subtle, you found it yourself, and the fix changed a published result.
4. The question bank
Grouped by who is most likely to ask. Each answer is a skeleton to deliver in your own words, not a script to recite — and each is deliberately short.
5. The three questions to ask them
They will leave two or three minutes for this and it is not a formality — the questions you choose tell them what you think the project is. Ask two. Keep the third in reserve in case one gets answered earlier.
6. The six traps
| Trap | Why it is a trap | What to do instead |
|---|---|---|
| Arguing mens rea transplants | It does not. EU competition liability is largely objective, intention bears on fines rather than on the infringement, and the AI Act deliberately regulates providers and deployers rather than models. Pushing it makes you look like you misread the field | Say the framing was your route in, and that the result bears on the duty to give reasons — Art 296, Art 14 — not on culpability |
| Claiming a method you have not used | One follow-up question exposes it, and the panel can ask that question | *No, I have not built one. Here is what I understand it to be and whether I think it fits* |
| Saying *scrape* | It signals carelessness about terms of use and lawful basis to a panel that includes people who study exactly that | *I assemble a corpus from published sources, log how, and keep per-document provenance* |
| Overclaiming your own results | 15 runs, 5 scenarios. Any model-versus-model claim is unsupported — Fisher gives 0.206 | Pooled claims only, limitations volunteered first |
| Monologuing | 25 minutes, four panellists. A four-minute answer costs an entire exchange and reads as poor judgment | Ninety seconds, then stop. Let them ask for more |
| Agreeing with everything | A research panel is partly testing whether you will push back. Pure agreement reads as either thin or deferential | Disagree once, carefully, on something you actually know — the sandbox gap, or system-level documentation versus decision-level provenance |
7. The three projects to have ready
Do not arrive with a thesis plan — it signals you have decided without them. Arrive with three, offer one when asked, and ask whether it fits.
| Project | The pitch, in three sentences |
|---|---|
| Calibration and the standard of proof | *Legal standards of proof are statements about sufficient confidence, and modern classifiers are systematically overconfident. So I would measure the calibration of enforcement-style classifiers and ask whether an uncalibrated probability can be mapped onto *sufficiently precise and consistent evidence* at all. It is measurable, it is a joint law-and-statistics paper, and I think the answer is no in principle rather than in practice.* |
| Who sandboxes the regulator? | *Articles 57 to 60 of the AI Act assume a firm bringing a product to a regulator. But enforcement screening is the state deploying a high-risk system against undertakings, and there is no sandbox for that and no obvious supervisor. I would map the provisions onto the agency-as-deployer case, compare how the existing agency tools were actually validated before use, and propose an institutional form.* |
| Decision-level provenance | *Model cards, datasheets and Annex IV all document the system. None of them records what a specific output was based on — which documents were retrieved, what the configuration was, what the confidence was. Evidence law needs a per-decision record and the entire documentation literature is organised around per-system records, and I have already built the artefact that fills the gap because my harness needed it.* |
8. If you remember ten things
- Ninety seconds, four beats, end on a question. Who you are, what you built, why it is a legal result, why ATLANTIS — then hand it back.
- 0 of 15, 9 of 15, 15 of 15, 2 admissions both self-judged, 0 cognizable intent. McNemar 0.004 and 0.00006. Fisher 0.206, which licenses nothing between models.
- Volunteer the five limitations before anyone asks. It converts every weakness into evidence of judgment and it is the specific competence the position is for.
- Mens rea does not transplant — liability is largely objective. Your result bears on the duty to give reasons, Art 296 and Art 14, not on culpability. Say so first.
- The v1 failure is your best story. Graders saw reasoning traces, traces restated the planted instruction, so verbosity was being measured rather than bias. You found it, fixed it, and the result changed.
- Never claim a method you have not built. *No, but here is what I understand it to be and whether it fits* is a strong answer. No digital twin, no agent-based model, no GNN, no blockchain.
- Never say scrape. You assemble a corpus from published sources, log how, and keep per-document provenance.
- Lead with calibration if they ask what you would work on. Keep the sandbox gap and decision-level provenance in reserve.
- Ask how the computational role sits beside the two legal tracks, and ask what data access the project has. Both are real questions and both signal you are thinking about the work rather than about yourself.
- Disagree once, carefully. A research panel is partly testing whether you can. And then stop talking — in 25 minutes across four people, the silence you leave is what lets them hire you.