Skip to content
VibeFormer
50 min

REVISION 6 — Every Question

Every question they could realistically ask, with the answer, grouped by type and ranked by likelihood. Plus the opener, the questions to ask them, the six traps, and what to do when you do not know.

Listen

0. The shape of the 25 minutes

The likely shape

The motivation block is the most likely and the most heavily weighted, and it is also the one you can rehearse to completion. The technical and legal probes are possible rather than probable, which is the opposite of how most candidates allocate their preparation.

1. The opener — the ninety seconds you fully control

2. Background and motivation — the most likely block

QuestionThe answerOdds
*Tell us about yourself.*Section 1. Four beats, ninety seconds, then stop.Near certain
*Why law and data science?***Five beats, and lead with *it was not a plan*.** *I came to a question rather than a field. My father has been in a fraud matter for roughly twenty years — not a complicated case, twenty years. And what stayed with me was not the unfairness but that the delay was nobody's decision. No one chose it; it is a property of how the system processes volume. A property of a system is measurable rather than a grievance — which is where this stopped being personal. I did the law degree expecting to practise and concluded I would be one more person inside a process whose problem is structural. I would rather measure the structure than add to the queue*Very likely
*Why antitrust specifically?*Be honest about the route — the dishonest version will not survive. *I came to a question, not to antitrust. What I care about is whether the measurements a legal system runs on are sound. Competition enforcement turns out to be the place where that question is sharpest right now, because agencies have started computing at scale while the law still assumes they read — and because the gap between a pattern and a legal finding is unusually explicit there. Wood Pulp says parallel conduct proves concertation only if nothing else explains it. That is a rule about statistical evidence, written by a court, and I find that genuinely interesting*Likely
*Why a PhD, and why four years?*Your answer is documented, which is unusual. *I came to VU in 2024 for the summer school on logic as a tool for modelling, on a scholarship — two years before this vacancy existed. So the interest is not a reaction to a job posting. And what I took from it is the thing I keep returning to: formalising a provision forces you to choose a reading. *Every firm agreed with some firm* versus *there is one firm every firm agreed with* is the difference between scattered bilateral deals and a hub-and-spoke cartel — same words, reordered. English leaves it ambiguous; the formalisation cannot*Likely
*Why the computational position rather than a legal one?**Because I would be a mediocre third legal researcher and I think I would be a useful first computational one. And the honest reason is that what I am good at is being suspicious of measurements, which is only useful if someone is producing measurements*Likely
*What do you know about our project?*Use their own words. *The central question as you put it is not whether these tools will be used, but how — and that safeguards are owed to companies as well as individuals. It combines legal analysis, computational science and institutional economics, and the announcement notes computational researchers in the law faculty would be a first for the faculty. More than 75 agencies in the network. I read that as the gap being institutional rather than technical*Likely
*You said you are not familiar with EU competition law. Why should we take you?*Answer directly; do not soften. *Because the combination is rarer than it looks. A competition lawyer who codes usually codes at the level of a notebook, and this position needs someone who has run a live system, designed an evaluation and found their own error in it. And doctrine is learnable in a way that the instinct to distrust your own measurement is not. I have spent the weeks since applying on the instruments; I have spent three years learning to doubt my own numbers*Medium-high
*What is your biggest weakness?*One answer, true, and it doubles as a reason you want the post. *I have worked alone. The consequence is that nobody has ever told me a measurement would not bear the argument I was building on it, so I have had to be my own referee — and that is a worse system than having one*Medium-high
*What would you need to learn?**Three things. EU competition law doctrine properly, which I said in my letter and have started. Institutional economics, which is your third named pillar and which I do not have — applied econometrics is not the same thing. And formal annotation methodology — agreement statistics, codebook design — which is standard in empirical legal studies, is not on my CV, and is a few days of reading*Medium
*Where do you see yourself after the PhD?*Do not over-promise academia. *Honestly, either a research role or an agency. What I would not want is to go back to building products. The thing I want to be doing is deciding whether a measurement is good enough to act on*Medium
*Your work is all Indian law. Does any of it transfer?**The doctrine does not and I would not pretend otherwise. The structure does. My crime paper is a measurement problem wearing the clothes of a behavioural finding, which is the same shape as a screen measuring detection rather than collusion*Medium
*How do you work with non-technical colleagues?*Answer with evidence, not enthusiasm. *I have not done a single-discipline project since my undergraduate degree — every paper I have written joins law to computation. And the direction people forget is the second one: the useful thing is not explaining a model to a lawyer, it is telling an engineer that the threshold they just typed is a legal decision*Medium
*You applied one day before the deadline.**I found the post late and would rather have had longer, but I did not want to miss it.* Then stop. Nobody cares unless you make it a topicLow-medium

3. Your own work — the second most likely block

QuestionThe answerOdds
*Tell us about your most recent paper.*The thirty-second version. *A twenty-year district panel from national crime records plus interpolated census data. Three findings: no spatial clustering, Moran's I 0.095 at p 0.151. State policy explains under one percent of local variance — an ICC of 0.6 percent. And the 2013 amendment produced about 88,879 additional recorded cases a year against the counterfactual, which is an administrative reporting shock rather than a behavioural epidemic. The paper's real subject is that police records measure reporting, not offending*High
*What are the weaknesses of that paper?*Volunteer the real one. *The literacy gap is not statistically significant — p of 0.399 and 0.156 — and my abstract uses the word *proving* about its causal effect. Identification and precision are separate questions: closing a backdoor path tells me which quantity I am estimating, not how precisely. So the correct statement is a point estimate of 0.1487 that cannot exclude zero, and urbanisation at p 0.004 is what I would actually defend. I would write that abstract more carefully now*High if the paper comes up
*Tell us about the Mens Rea paper.**A paired three-condition design testing whether fabricated legal authority in a model's context produces a biased ruling. Zero of fifteen in the clean condition, nine of fifteen with partial manipulation, fifteen of fifteen when poisoned. And the finding is not the susceptibility rate — it is zero of fifteen showing any cognizable intent, while the model's account of its own decision stayed fluent and plausible.* Then the correction: *mens rea does not transplant, and the paper is not arguing that it does. EU competition liability is largely objective. What the result bears on is the duty to give reasons*Medium-high
*Is your sample not very small?*Concede instantly and completely. *Yes. Fifteen runs over five scenarios, produced in a hackathon sprint. And in five of those rows the model judging the output was also a subject — both of my two admissions fall in exactly those rows. The Fisher test at p 0.206 licenses no model ranking and I would not claim one. What fifteen runs can support is that the effect exists and is large in the poisoned condition; what they cannot support is any comparison between models*Medium-high
*Does mens rea transplant to AI systems?*Be careful — the title invites this and the answer is no. **See the row above. Say *no* first, then explain what the result is actually about**Medium
*What is the through-line across your work?*The best answer you have. *One question I keep returning to without planning it: what is the measurement actually measuring? Crime records measure reporting, not offending. A model's stated reasons are not its actual reasons. Overlap metrics score a summary that dropped the operative facts. An accuracy figure over two hundred classes says nothing about the rare ones. Five pieces of work, one question — and it is the question your accuracy strand is asking*Medium
*Tell us about your master's thesis.**An end-to-end system: a school topic goes in, a narrated video explaining it at a chosen grade level comes out, in English or six Indian languages. Ten stages, built solo in four months. The design constraint was cost — use the expensive model once offline to manufacture a dataset, then distil it into a model small enough that a school could run it. The build works; the evaluation does not, and I know why: the training answers and the reference answers both came from the same teacher model, so my metrics measured imitation rather than accuracy — and the thing the system actually promises, grade-appropriateness, is the one thing I never measured*Medium — they have the document
*So is any of your work reliable?*The trap after you volunteer weaknesses. Do not retreat. *Yes — differently in each case, and being able to say which is the point. The ICC I would defend as it stands, because a variance partition is a direct computation. The reporting shock I would defend with the caveat that a counterfactual time series assumes nothing else changed at that moment. The literacy-gap causal estimate I would present as a point estimate that cannot exclude zero. Being able to say which claims survive at which strength is the skill, and it is the one this position is for*Medium
*You are second author on some of these. What was yours?*Be precise and do not inflate. *State what was mine and what was not. If I cannot remember the division clearly I would rather say so than construct one.* An honest gap in recollection is harmless; a claimed contribution the first author would describe differently is notLow-medium

4. Technical probes — possible, not probable

QuestionThe answer
*What would you build first?*Deliberately unglamorous, and that is the point. *A corpus and an entity layer. Before any of this is measurable you need the decisions in one place with provenance per document, and firms resolved to undertakings with a reported matching threshold. That is most of the first year and everything else depends on it*
*How would you detect a cartel?*Three layers, and keep them separate. *Exploratory analysis asks whether a pattern is unusual; graph analysis asks who is unusual together; and the law asks whether concertation is the only plausible explanation. The strongest single screen is not the price level but the price variance — collusive prices are steadier, not just higher. Graphs then add closed cliques, a coordinating node, and best of all a structure that forms and dissolves around a known date. But Wood Pulp means a screen generates hypotheses and can never be the conclusion*
*What is RAG and why does it matter here?**Retrieval-augmented generation: look the facts up first, then answer using what you found. It matters legally because the answer can be traced back to a document. I have built it twice — once over an encyclopaedia corpus, once where the retrieval step was a deterministic calculation instead of a lookup. The generalisation is ground-and-generate: compute or retrieve the facts in a way you can show, and let the model only do the prose*
*How would you audit an enforcement tool?*And here the standard advice is wrong. Do not offer SHAP or LIME — explanation methods routinely disagree and are unstable under resampling, with explanations differing across repeated runs with the model and input fixed. *What I would do instead: design-level interpretability where the task allows it, per-decision input logging, counterfactual explanations because that form is testable by a court, and a stability test — refit on a slightly different sample and check whether the same firms are flagged. If the flagged set moves, who gets investigated is partly arbitrary, and that is a legal defect rather than a tuning issue*
*What is the difference between a neural network and a graph?**Both use the word *node*, which is the confusion. In a neural network the nodes are arithmetic; in a graph the nodes are things in the world and the structure is what you are studying. And agencies mostly use plain graph analysis, because it needs no training data and the output can be explained to a lawyer in one sentence*
*Explain transformers.**Earlier models read a sentence strictly left to right, carrying what they could remember. Transformers look at the whole page at once — for every word, the model computes how much every other word matters to it. And the reason that mattered was not mainly accuracy but parallelism, which made internet-scale training possible*
*What is QLoRA?*The encyclopaedia analogy. *Instead of reprinting the whole encyclopaedia to add a subject, leave it untouched and write a thin booklet of additions read alongside it — and first reprint it in smaller type so it fits on the shelf. The small type is quantisation, the booklet is the low-rank adapter. About one percent of the parameters train*
*How is Astroformer deployed?**Vercel for the frontend, Render for the Python API as a Docker container* — Docker rather than a buildpack because the astrology engine compiles a native C extension. *No GPUs; every model call goes to a hosted provider, chosen by environment variable.* Then volunteer the free plan's cold starts, the permissive CORS setting and the default secret in the config
*Which metric would you report?**Precision and recall separately, never accuracy. If one tender in a thousand is rigged, a model that flags nothing is 99.9 percent accurate and worthless. And I would state the baseline, because 84 percent means nothing until you know that guessing the commonest class gives 80*
*What is data leakage?**Information about the answer sneaking into the inputs. The clearest documented case is in legal prediction: about 79 percent accuracy from judgment text written by the court after it decided, falling to roughly 58 to 68 percent when redone as genuine forecasting — and high accuracy from the judges' names alone. Competition law is about to make the same mistake*
*Do LLMs reason?*Do not take either confident position. *It is contested. They produce outputs consistent with reasoning on a wide range of tasks and also fail in ways that suggest pattern-matching — the same problem rephrased can flip the answer. For a legal process what matters is not which it is but whether the output can be checked against a source, which is an argument for building verifiability rather than for settling the question*
*Your CV says TensorFlow and PyTorch — which do you use?**PyTorch in practice, through Hugging Face rather than writing training loops. The TensorFlow work was coursework. I would not claim framework-level depth in either; what I have is the applied layer above them*

6. Who is likely to ask what

PanellistTheir territoryTheir likely questionWhere your answer is
Thibault Schrepel (chair)Computational antitrust, complexity economics, the knowledge graph of Commission decisions*What would you actually build first?* and *how do you see the strands fitting?*Section 4, row one. The corpus-and-entity-layer answer. Plus the identification-versus-forecasting distinction
Catalina GoantaLarge-scale empirical study of online content; multilingual measurement; a legal compliance API*How would you handle the data side?* — method, not doctrineValidation: a written codebook, two coders, an agreement statistic fixed in advance. And volunteer that you wrote your own benchmark sentences
Tijmen WismanPrivacy and surveillance; legality and proportionality; IoT and connected devices; connected to the platform behind SyRI*What about the people subject to these systems?*Section 5, the SyRI row. And the IoT paper is the closest thing you have written to his field
Georgiana MirzaDigital ecosystems, competition law, market regulationRegulatory design — thresholds, designation, market definitionThe DMA-threshold-as-research-design point, including the bunching problem. And that *market* is a legal conclusion smuggled in as a data field

7. The questions to ask them

Ask thisWhy it worksWhat you will probably hear
*You describe this as the cross-project computational position. I am trying to understand how you picture it day to day — closer to a service role producing the measurements the legal strands specify, or a parallel line of inquiry that sets some of its own questions? Both are workable and they would shape my first year quite differently.*The best one. It is a real question, the answer genuinely changes your plans, and it signals you are thinking about collaboration rather than your own thesis**Most likely *both*, or *it will evolve* — an honest answer, not a dodge. The structural fact underneath: an ERC-funded PhD must produce a dissertation, so pure service is not possible. Do not look disappointed by *we will work it out*; at a first interview that is the normal answer**
*The project description mentions that the agency networks give access to practices and challenges not visible in public documents. I am curious what form that takes in practice — closer to institutional knowledge and interviews, or does it extend to the underlying data? I ask because it changes what is answerable, and if it is mostly the former then synthetic data becomes a workstream rather than a side project.*Framed so it does not reveal you missed the announcement — which already says the networks give that access, so the open version would show you had not read itMost likely mostly the former. Real agency data is confidential business information and rarely shared even with academics. If it is largely qualitative, say on the spot that your synthetic-data ideas move to the centre — it shows you can replan rather than deflate
*If I joined, what would a good first six months look like? I would rather spend them building something the legal tracks actually need than something I find interesting.*Practical, modest, service-oriented without being servile, and the answer is immediately usefulReading into EU competition law, learning the legal tracks' questions, some scoping work. If he names a specific first task, write it down — that is the brief
*What would you want me to learn in the first year that I do not have now?*Impossible to answer insincerely, and it lets you say you have already startedAlmost certainly EU competition law doctrine, possibly institutional economics. This is him naming the gap he is actually worried about — the most valuable thing you can extract
*Who would supervise the computational side day to day? I am conscious a computational researcher in a law faculty may not have an obvious technical supervisor, and I would want to know where methodological pushback comes from.*Real and under-asked, and it connects to your honest weaknessHe may name a co-supervisor or say it is being arranged. Raise it as a question, never as a concern about the post

8. The six traps

TrapHow it appearsThe escape
Overclaiming a tier*So you are comfortable with deep learning frameworks?*Name the tier immediately. *Applied layer, yes. Framework internals, no*
Being drawn into criticising the project's premise*Do you think agencies should be using these tools at all?*The premise is that they will. Frame every criticism as a condition of proper use, never as an objection
Offering SHAP and LIME as the audit answer*How would you make a model explainable?*Say why they are the problem, then offer counterfactual explanations and the stability test
Performing humility, then collapsing*So is any of your work reliable?*Section 3, last row. Say which claims survive at which strength. Do not retreat
Letting *mens rea* stand as a claim*Do you think models can have intent?*Say no first. *The framing was the route in, not the conclusion. The result is about the duty to give reasons*
Filling silenceA pause after your answerLet it sit. A panel pausing is often writing, not waiting. Adding a third example after a good answer is how strong answers get diluted

9. The ten things to carry into the room

  1. It is a fit interview from nearly 400 applications. Six to eight questions. The syllabus is your CV and your letter, not EU law.
  2. The opener: four beats, ninety seconds, one number — 88,879 — then stop.
  3. *Police records measure reporting, not offending. A cartel screen measures detection, not collusion.* Your single strongest sentence. Use it early.
  4. You disclosed the competition-law gap in writing and were interviewed anyway. *I said I would need to learn it, and I have spent the time since doing exactly that.*
  5. Volunteer the literacy-gap problem before anyone asks: not significant, abstract says *proving*, identification and precision are separate, urbanisation is what you defend.
  6. *Mens rea does not transplant.* Say it first. The result is about the duty to give reasons.
  7. Accuracy lies on rare events. Precision and recall separately, and always state the baseline.
  8. ***Wood Pulp*: parallel conduct proves concertation only where nothing else explains it. So a screen generates hypotheses, never conclusions.**
  9. Your weakness is working alone, and it doubles as why you want the post.
  10. *I do not know that well enough to give you a real answer* is a winning move, not a loss. Use it the moment you need it.