Skip to content
VibeFormer
58 min

Data Science: The Actual Craft

The part nobody writes papers about and every project dies on. The lifecycle, what exploratory analysis actually means, the three kinds of missing data, rates against counts, Simpson's paradox worked numerically, entity resolution as a legal question, the four quasi-experimental designs, whether EDA and graph theory can actually find a cartel, and reproducibility as evidence.

Listen

0. Why this chapter, and the distinction to open with

Every other technical chapter here is about a method — a model, a measure, a formalism. This one is about the craft around the method, and it is the chapter that matters most for a project like ATLANTIS. The reason is blunt: competition agencies do not fail at modelling, they fail at data. Firm names do not match, the denominator is wrong, the missing rows are missing for a reason, and nobody can rerun last year's number.

1. The lifecycle, with honest proportions

Textbooks draw this as a tidy loop with equal boxes. It is not equal. The proportions below are the ones practitioners actually report, and saying them out loud is a credibility signal — it tells a panel you have run a project rather than read about one.

Where the time actually goes

Roughly sixty percent of the work happens before any model is fitted, and the modelling step — the only one that appears in a paper — is about a tenth of it. The loop at the bottom is not failure, it is the normal shape of empirical work: the data tells you the question was imprecise.

2. The unit of analysis — the first decision, and the one most often wrong

The unit of analysis is the thing that each row of your table represents. It sounds trivial. It is the decision that silently determines what your result can mean, and getting it wrong invalidates everything downstream no matter how good the model is.

Candidate unitOne row is…So the result is a claim about…
FirmOne companyFirms. *Larger firms are more likely to be investigated*
Firm-yearOne company in one yearFirm behaviour over time. Lets you ask whether a rule change altered conduct
CaseOne investigationEnforcement, not conduct. This is the switch that catches people out
Decision paragraphOne passage of reasoningReasoning style — what agencies cite, how they argue
TenderOne procurement competitionBidding patterns. The natural unit for cartel screens
Firm-pairTwo companies considered togetherRelationships — co-bidding, interlocking directorates, common ownership
MarketOne product-and-geography combinationMarket structure. But *market* is a legal conclusion, not a data field

3. Exploratory data analysis — what you actually look at

Exploratory data analysis, or EDA, means looking at the data before modelling it — not to test anything, but to find out what you are holding. It has a reputation as the soft part of the job. It is in fact where almost every serious error is caught, and it is mostly a checklist rather than an art.

CheckWhat you are looking forThe failure it catches
Row and column countDoes it match what the source claimed?A truncated download. Depressingly common
One row, read in fullDoes a single record make sense as a sentence?Columns shifted by one; a header row read as data
Type of every columnIs the number column actually numeric?Numbers stored as text, so sorting puts 10 before 9
Min and max of every numberImpossible valuesA negative turnover; a fine of 0; a date in 1900 or 2099
Count of unique valuesToo few or too manyA *country* column with 340 values — spelling variants
Missing count per columnHow much is absent, and whereA column that is 90 percent empty and about to be used
Exact duplicate rowsIdentical recordsA double-loaded file inflating every count
Distribution of each numberShape: one hump, two humps, a long tailA bimodal column usually means two populations mixed together
Counts over timeRows per month or yearA reporting change. A step in the series is almost never real
Cross-tabs of key pairsCombinations that should not existDecisions dated before the investigation opened

4. Missing data — the three kinds, and why deleting rows is a choice

Almost every real dataset has holes. The critical insight is that the holes have a cause, and the cause determines what you are allowed to do about them. There is a standard three-way classification, due to the statistician Donald Rubin, and the names are unfortunately confusing — so here they are with the confusion removed.

The three kinds of missingness, as a decision

The names are about what the missingness depends on, not about the values. Missing At Random does not mean random — it means random once you condition on what you have already observed. That single clarification is worth more than the rest of the taxonomy.
KindA legal-data exampleWhat you may do
MCARA digitisation batch was skipped, and which batch was arbitraryDrop the rows, or impute. Estimates stay unbiased; standard errors widen
MARSmaller firms report turnover less completely — and you observe firm sizeMultiple imputation, or inverse-probability weighting, *conditioning on size*. Complete-case analysis is biased here
MNARUndetected cartels are absent precisely because they were good at concealment. Unreported crime is unreported because of the nature of the offenceNothing, from the data alone. Report a bound, model the selection explicitly, or find an external source. Imputation here manufactures a result

On imputation. *Imputation* means filling a hole with an estimate. The naive version — substitute the column mean — is almost always wrong, because it pretends you know the value exactly and so shrinks your uncertainty artificially. Multiple imputation is the defensible version: fill the holes several times with different plausible draws, run the analysis on each completed dataset, then combine. The combination rule is what makes it honest.

θ^=1m∑i=1mθ^iT=Uˉ⏟within+(1+1m)B⏟between\hat{\theta} = \frac{1}{m}\sum_{i=1}^{m} \hat{\theta}_i \qquad\qquad T = \underbrace{\bar{U}}_{\text{within}} + \left(1 + \tfrac{1}{m}\right)\underbrace{B}_{\text{between}}
The combined estimate is the average of the estimates from each imputed dataset. The combined variance is the average variance within datasets, plus a term for how much the estimates disagreed across datasets.The second term is the entire point. It adds the disagreement between imputations back into your uncertainty, so a quantity you had to guess at ends up with wider error bars than one you observed. Mean-substitution omits that term, which is why it reports false confidence. Five to ten imputations is the usual range.

5. Rates against counts — denominators, and the trap in both directions

A count is how many. A rate is how many per something. Choosing between them is not presentation, it changes the finding — and both choices have a failure mode, which is why this is a section rather than a footnote.

r=countpopulation at risk×100,000r = \frac{\text{count}}{\text{population at risk}} \times 100{,}000
A rate equals the count divided by the population at risk, multiplied by one hundred thousand to give a readable number.**The phrase doing the work is *population at risk* — the set of units that could have produced an event. For a crime rate, residents. For a cartel-detection rate, not** the population: the number of firms, or of tenders, or of markets. Dividing enforcement counts by population is a common and meaningless move.
UseProblemExample of the problem biting
Raw countsScale with the size of the unit, so the biggest unit always looks worstUttar Pradesh records more offences than Goa because it has fifty times the people. Germany has more cartel cases than Malta. Neither is a finding
Per-capita ratesSmall denominators make unstable rates. With a tiny population, one extra event swings the rate wildlyA district of 5,000 with 2 events is at 40 per 100,000; a third event takes it to 60. A fifty percent jump from one incident — and such districts will fill both tails of your ranking
Either, across a reporting changeThe denominator or the definition shifted mid-seriesA boundary redrawn, a category merged, a threshold changed. The series breaks and the break looks like an effect

6. Simpson's paradox — worked numerically, because it has to be

Simpson's paradox is when a pattern holds in every subgroup and reverses when the subgroups are pooled. It sounds like a curiosity. It is the single most dangerous arithmetic fact in applied work, because both the pooled number and the subgroup numbers are correct — so no amount of checking the computation will reveal it. The example below is in Catalina Goanta's research area deliberately, since she is on your panel.

The question: does labelling a post as an advertisement reduce engagement? Take 2,000 posts, split by influencer size, and count how many achieved high engagement.

GroupDisclosed as an adNot disclosedWhich wins?
Small influencers90 of 100 = 90%720 of 900 = 80%Disclosed, by 10 points
Large influencers540 of 900 = 60%45 of 100 = 45%Disclosed, by 15 points
Pooled630 of 1,000 = 63%765 of 1,000 = 76.5%Not disclosed, by 13.5 points

Why the reversal happens, structurally

The arrow from size to disclosure and the arrow from size to engagement together form what is called a backdoor path. Conditioning on size — analysing the groups separately, or including size in a regression — blocks it. This is the same logic as controlling for a variable, drawn rather than asserted.

7. Entity resolution — and why it is a legal question, not a cleaning step

Entity resolution is deciding which records refer to the same real-world thing. It is the least glamorous task in this chapter and the one most likely to decide whether a competition-law dataset is usable at all, because the records are firm names and firm names are chaos.

The pipeline

Step 2 is what makes the task tractable and step 6 is what makes it defensible. Most published pipelines describe steps 1 and 3 and quietly omit the threshold in step 4, which is the only number a reader would need to assess the result.
n(n−1)2son=50,000  ⇒  1,249,975,000 pairs\frac{n(n-1)}{2} \qquad\text{so}\qquad n = 50{,}000 \;\Rightarrow\; 1{,}249{,}975{,}000 \text{ pairs}
The number of pairs to compare is n times n minus one, over two. For fifty thousand firms that is just under one and a quarter billion comparisons.This is why blocking exists. Comparing every record to every other grows with the square of the dataset, so doubling the firms quadruples the work. Blocking on, say, the first four characters of the normalised name cuts it to millions. **The cost is that a genuine match whose names differ at the start — *Google* against *Alphabet* — will never be compared at all.** Blocking trades recall for feasibility, and the trade should be stated.
ToolWhat it measuresWhere it fails
Levenshtein distanceHow many single-character edits turn one string into the otherGoogel→Google is 1 edit, good. But two unrelated short names can also be 1 edit apart
Jaro-WinklerCharacter overlap, weighted towards agreement at the start of the stringThe usual default for names, because company names rarely differ at the beginning — which is also exactly why it misses Alphabet/Google
Token-set comparisonTreats the name as a bag of words, so word order stops matteringBank of Ireland and Ireland Bank match — sometimes right, sometimes two different banks
Identifier joinMatch on a registration number, VAT number or LEI instead of a nameAlways prefer this when it exists. It fails only by being absent, which in public legal corpora it usually is

8. Measurement validity — the gap between what you want and what you have

The construct is what you want to measure. The operationalisation is what you actually measured. Validity is the size of the gap. Almost every criticism worth making of an empirical legal paper is a validity criticism, and almost none of them can be fixed by a better model.

What you want (construct)What is in the dataThe gap, stated plainly
CollusionCartel decisionsYou are measuring enforcement. Detection depends on leniency programmes, agency resources and concealment skill. A fall in cases may mean less collusion or a worse-funded agency
Harm to consumersFines imposedFines track turnover and procedure, not harm. A large fine on a small harm is routine
Market powerMarket shareShare is a proxy, and the legal threshold is dominance — which depends on entry barriers and countervailing buyer power that no share captures
Non-compliance with disclosure rulesPosts without a visible ad labelMisses paid posts you cannot identify as paid. The denominator is unknown, which is the hardest version of this problem
CrimeRecorded offencesRecording depends on reporting, and reporting depends on the offence. The dark figure, and it is MNAR
Algorithmic discriminationOutcome differences across groupsDifferences may reflect the decision rule, the data, or genuine differences in the underlying population — and the three need different remedies

9. The four quasi-experimental designs — how to claim a cause without an experiment

An experiment randomises who gets the treatment, which is why it supports causal claims. You cannot randomise which firms are investigated or which member states get a regulation. A *quasi-experiment* exploits a situation where something close to random assignment happened anyway. These four designs are the workhorses of empirical legal research, and knowing what each one assumes is more useful than knowing how to fit it — because the assumption is where the argument lives.

DesignThe counterfactual it buildsThe assumption it dies on
Difference-in-differencesWhat the treated group's trend would have been, taken from an untreated group's trendParallel trends — absent treatment, both groups would have moved together. *Not* that they were at the same level
Regression discontinuityUnits just below a threshold, compared with units just aboveNothing else jumps at the cutoff, and units cannot precisely place themselves on either side
Synthetic controlA weighted blend of untreated units, chosen to track the treated unit's pre-treatment pathNo spillover onto the donor units, and a long enough pre-period to fit the blend honestly
Event studyThe same unit, just before the eventNothing else happened in the window, and the date is correct and not anticipated
τ^DiD=(Yˉafter treated−Yˉbefore treated)−(Yˉafter control−Yˉbefore control)\hat{\tau}_{\text{DiD}} = \big(\bar{Y}^{\,\text{treated}}_{\text{after}} - \bar{Y}^{\,\text{treated}}_{\text{before}}\big) - \big(\bar{Y}^{\,\text{control}}_{\text{after}} - \bar{Y}^{\,\text{control}}_{\text{before}}\big)
The difference-in-differences estimate is the change in the treated group, minus the change in the control group over the same period.Two subtractions, and each removes one threat. The first removes everything about the treated group that was constant over time. The second removes everything that happened to both groups — a recession, a price shock. What survives is the part that happened to the treated group and not the other one.

Difference-in-differences, and what parallel trends means

Difference-in-differences does not require the two groups to be alike in level, only alike in trend. Testing the pre-treatment period for parallel movement is the standard supporting evidence, and a visible divergence before treatment is the standard reason a difference-in-differences paper is rejected.
τ^RD=lim⁡x→c+E[Y∣X=x]  −  lim⁡x→c−E[Y∣X=x]\hat{\tau}_{\text{RD}} = \lim_{x \to c^{+}} \mathbb{E}[Y \mid X = x] \;-\; \lim_{x \to c^{-}} \mathbb{E}[Y \mid X = x]
The regression-discontinuity estimate is the expected outcome just above the cutoff, minus the expected outcome just below it.The logic is that a firm at 7.4 billion and a firm at 7.6 billion are nearly identical in everything except which side of the rule they fall on — so the jump in outcomes at the line is attributable to the rule. The cost is that the estimate is local: it is the effect *at the threshold*, and says nothing about a firm at 30 billion.
Y^t 0=∑j=2J+1wj Yjtwithwj≥0,∑jwj=1\hat{Y}^{\,0}_{t} = \sum_{j=2}^{J+1} w_j\, Y_{jt} \qquad\text{with}\qquad w_j \geq 0, \quad \sum_{j} w_j = 1
The synthetic counterfactual at each time is a weighted sum of the untreated units' outcomes, where the weights are non-negative and sum to one.The weights are chosen so the blend tracks the treated unit's path before treatment. Forcing them to be non-negative and to total one is what keeps the comparison interpretable — you are building a portfolio of real units, not extrapolating. Use this when exactly one unit was treated, which is the common case for a single member state or a single market.

10. Putting it together — can EDA and graph theory find a cartel?

This section is the applied synthesis of the chapter, and it is the most likely thing you would actually be asked to design. Short answer to the question in the heading: **EDA and graph theory can tell an agency *where to look*. Neither can establish that a cartel exists.** Holding that line is what makes the rest credible, so it is stated at the start and again at the end.

What a cartel leaves behind. Collusion is an agreement, which is not in the data. What *is* in the data are its consequences:

  • Prices that behave differently from competitive prices — higher, and crucially *steadier*, which is the basis of the main screen below.
  • Bids that do not look like independent attempts to win — because under an allocation agreement, most bidders are trying to lose.
  • Market shares that are too stable — allocation agreements have to be monitored, and monitoring shows up as shares that barely move.
  • Costs that stop passing through to prices — a competitive firm must follow its input costs; a cartel can absorb a shock and hold the price.
  • Relationships that are too consistent — the same firms appearing together, in the same roles, across many tenders. This is the part only a graph can see.
  • A structural break at a known date — a dawn raid, a leniency application, a cartel collapse. The most legally useful pattern in the list, because it gives you a before and an after, which is the difference-in-differences setup from section 9.
CV=σxˉand the screen looks forCVsuspect  ≪  CVbenchmark\mathrm{CV} = \frac{\sigma}{\bar{x}} \qquad\text{and the screen looks for}\qquad \mathrm{CV}_{\text{suspect}} \;\ll\; \mathrm{CV}_{\text{benchmark}}
The coefficient of variation is the standard deviation divided by the mean. The variance screen looks for a period, market or group of firms whose coefficient of variation is much lower than a comparable benchmark.Dividing by the mean is what makes it comparable across markets — a raw standard deviation is bigger wherever prices are bigger. **The direction is the counter-intuitive part and it is the thing to remember: collusive prices are *less* variable, not more.** A cartel sets a price and defends it; competing firms chase each other's costs and promotions, and that chasing is visible as variance.

The EDA screens, as a checklist. Each is a plot or a statistic, and each has a legitimate alternative explanation — which is the point of the limits section below.

  • Coefficient of variation of price, by market and by period. Look for low-variance pockets. This is the primary screen.
  • Price level against a benchmark — a comparable region, a comparable product, or the same market before and after. Never a bare level; always a comparison.
  • Cost pass-through. Regress price on an input-cost index. A pass-through coefficient near zero where it should be substantial is a strong signal, because holding price against a cost shock requires coordination or market power.
  • Price series plotted around known dates — raids, leniency applications, arrests. A drop in mean and a jump in variance immediately after is the signature of a cartel ending.
  • Market shares over time. Plot each firm's share as a line. Lines that are suspiciously flat, or that swap places in a regular pattern, suggest allocation.
  • The distribution of losing bids. In genuine competition, bids cluster near cost with a right tail. Under cover bidding, losers sit implausibly far above the winner — and the gap between first and second is informative in a way the winning bid alone is not.
  • The winner sequence. Test whether the identity of the winner across successive tenders is more regular than chance. Rotation is a pattern randomness does not produce.
  • Bid-to-cost or bid-to-estimate ratios. Screens out the effect of tenders simply being different sizes.
  • Round numbers and digit patterns. Genuine cost-based bids rarely land on round figures as often as coordinated ones do.
  • Entry and exit counts. A market with no new entrants over a long period is both a precondition for a stable cartel and a consequence of one.
  • Participation patterns. Which firms bid for which tenders, and — more telling — which firms consistently decline to bid where they plainly could have. Bid suppression is an absence, and absences are invisible unless you plot them deliberately.

What the bid distribution looks like, competitive against coordinated

Competitive bidding produces a spread because firms have different costs and different appetites for the work. Cover bidding produces a gap followed by a cluster, because the losing bids are not attempts to win but a service performed for the winner. Plotting the distribution rather than the mean is what makes this visible at all.

Now the graphs. EDA treats each firm or tender as an independent row. Graph theory is what you reach for once you suspect the rows are related — and cartels are, by definition, a relationship. Five graphs are worth building from procurement data:

  • Firm–tender bipartite graph. Two kinds of node, edges only between kinds. This is what procurement data literally is, so build this first and keep it — it is the ground truth the others are derived from.
  • Firm–firm co-bidding graph. An edge when two firms bid for the same tender, weighted by how often. The workhorse. Derived from the bipartite graph by projection — see the warning below.
  • Winner–loser directed graph. An edge from loser to winner. Rotation shows up as cycles, and a hub-and-spoke arrangement shows up as a node with unusual in-degree.
  • Subcontracting graph. An edge from winner to the firms it subcontracts to. This is the strong one, because compensating the losers is a cartel's hardest practical problem — and a firm that reliably loses to a given winner and is then reliably subcontracted by that winner is a pattern with few innocent readings.
  • Ownership and interlocking-directorate graph. Firms linked by common shareholders, or by sharing a director. Nodes here include people, and the same entity-resolution problem from section 7 applies with the extra difficulty that people change names.

What graph theory then tells you, measure by measure, with the legal reading beside it:

  • Connected components — a group reachable only within itself. Reads as market allocation or a genuinely separate regional market. Separate components in a market you believed was national is a finding worth checking.
  • Cliques — every member connected to every other. A closed bidding ring makes exactly this shape, so enumerating large cliques is a direct search for candidate rings.
  • k-cores — the subgraph where every node has at least k connections within it. More robust than cliques, because a single missing edge destroys a clique but not a core.
  • Community detection and modularity — groups denser inside than between. Candidate cartel membership, discovered rather than assumed. The honest framing is that modularity gives you a partition for *every* graph, so a high score needs comparing against a randomised benchmark.
  • Local clustering with low overall density — a tight group inside an otherwise sparse market. The signature of a closed group operating within a larger market.
  • Betweenness centrality — who sits on the paths between others. The candidate coordinator, trade-association hub, or spoke-connecting party in a hub-and-spoke arrangement.
  • Weighted-edge outliers — two firms that co-bid far more often than their sizes and sectors would predict. The single cheapest thing to compute and often the most informative.
  • Temporal graphs — rebuild the graph per year and watch edges appear and vanish. A structure that forms at one date and dissolves at another, especially around a raid, is the most persuasive graph evidence available, because it adds a before and an after.
  • Bipartite co-occurrence — firms that appear in tender sets together far more often than chance. Computed on the bipartite graph directly, this avoids the projection problem below.
  • The bipartite projection trap. Projecting a firm–tender graph into a firm–firm graph mechanically creates a clique for every tender — ten firms bidding for one contract become ten mutually connected nodes. So clustering coefficients computed on a projected graph are inflated by construction and mean nothing without a null model. This is the single most common technical error in this area, and it is covered in the network analysis chapter.
  • Everything has a lawful explanation. Specialisation, capacity limits, geography, prequalification rules, consortium bidding and joint ventures all produce co-bidding structure legitimately. Firms that only bid for what they can actually build will look clustered.
  • The base-rate problem. Cartels are rare. A screen with excellent accuracy still produces mostly false positives when the underlying rate is low — and converting *this pattern would be unusual if they were competing* into *they are probably colluding* is the prosecutor's fallacy. It needs a base rate nobody has.
  • Entity resolution decides the graph. Get the corporate group wrong and you have invented or destroyed edges. Section 7's point lands here with full force: the matching threshold is upstream of every graph measure.
  • Multiple testing. Screen 500 markets at a five percent threshold and about 25 will flag by chance. Any screening programme needs a correction or a stated expected false-positive count, in advance.
  • Screens are gameable once known. A published screen is a specification for evading it — add noise to prices and the variance screen goes quiet. So screens detect unsophisticated collusion, which makes the detected set a biased sample of the real set. This is section 4's MNAR problem arriving in operational form.
  • Data access is the binding constraint, not method. Bid-level procurement data with firm identifiers is often unavailable to researchers. Say this plainly — it is why synthetic data and public tender portals matter.

The pipeline, in order, if you were asked to design one. This is the answer to *what would you build first* and it is deliberately unglamorous:

  1. Get the bid-level data with firm identifiers, and record provenance per tender. Without identifiers there is no graph.
  2. Resolve entities to undertakings, report the threshold, hand-check a sample and print the largest cluster size. Section 7 — and this step is where the project is won or lost.
  3. Do the EDA before anything clever. Rows per period first, then the price and bid distributions, then the unique-value checks.
  4. Build the bipartite graph and keep it as the base object. Derive everything else from it rather than from a projection.
  5. Run the variance and pass-through screens to flag candidate markets and periods.
  6. Build the graph only for the flagged markets, and compare every structural measure against a degree-preserving randomised null model. Without the null model the numbers are decoration.
  7. Rebuild per year and look for structures that form and dissolve, especially around known enforcement dates.
  8. Enumerate the innocent explanations for each flag — capacity, geography, specialisation, lawful consortia — and try to eliminate them. This is the Wood Pulp step and it is the one that decides whether any of it is usable.
  9. Hand the survivors to a case team as hypotheses with stated error rates, never as conclusions. And log what the pipeline saw on the day it flagged them, so the selection can be reconstructed and contested.

11. Reproducibility — and why it is a legal requirement here, not hygiene

Reproducibility means someone else, with your data and your code, gets your number. In most fields this is professional good manners. In this project it is closer to an evidentiary standard, because a measurement that feeds an enforcement decision may have to be explained to a court.

PracticeWhat it preventsThe minimum version
Raw data is never editedThe silent, unrecoverable changeKeep raw/ read-only. Every transformation writes to interim/ or processed/
Every step is a script*I fixed that one in the spreadsheet and cannot remember how*If a step was done by hand, it is not reproducible. No exceptions
Random seeds fixedA different answer on rerunSet the seed once at the top. And report that results are seed-dependent if they are
Environment pinnedThe result changing because a library updatedA requirements.txt with exact versions, not ranges
Data versioned or hashedNot knowing which vintage produced the numberRecord the download date and a checksum of the file
One command runs everythingA pipeline only its author can operateA single script that goes from raw/ to every figure in the paper
Decisions logged with reasonsUnexplainable choices at review timeThe matching threshold, the exclusion rules, the imputation model — a short written record of each

12. The seven ways an empirical legal result dies

A compression of the whole chapter. If a result of yours is going to be attacked, it will almost certainly be on one of these seven, and each has a pre-emptive sentence.

Cause of deathThe attackYour pre-emptive sentence
Wrong unit*Your district-level finding does not support a claim about people*State the level in the result itself. *At district level…*
Selected sample*You measured detection, not conduct*Name it as MNAR and give the direction of the bias
Bad denominator*Per capita is the wrong base for enforcement counts*Report the result under two or three denominators
Hidden confounder*This reverses if you split by firm size*Check the obvious subgroups before publishing; present the design, not just the regression
Entity errors*Your market shares are wrong because the group was mis-assembled*Report the threshold, the largest cluster size, and a hand-checked sample
Overclaimed significance*Your p-value is 0.156 and your abstract says* provingFix the abstract. This one is live in your own paper and you should raise it first
Not reproducible*Nobody can rerun this, including you, in six months*One command from raw data to every figure

13. How this maps onto ATLANTIS, and the probes per panellist

The project combines legal analysis, computational science and institutional economics, and the announced central question is not whether these tools will be used, but how. That phrasing matters for this chapter: **the *how* is overwhelmingly a data-craft question rather than a modelling one.** The safeguards owed to companies and individuals are, operationally, most of the checklist above.

PanellistLikely probe from this chapterWhere your answer comes from
Thibault Schrepel*What would you actually have to build before any of this is measurable?* — a scoping question, not a technical oneSection 1's proportions and section 7. The honest answer is that the first year is a corpus and an entity layer, and that this is the unglamorous precondition for everything else
Catalina Goanta*How would you handle the data side?* — her own work is measurement over messy multilingual text at scaleSections 3, 7 and 8. The disclosure example in section 6 is in her area — use it as an illustration of why you check subgroups, not as a comment on her findings
Tijmen Wisman*What about the people on the other side of these systems?*Section 10 and SyRI: reproducibility as a condition of a reviewable proportionality assessment. And section 9's closing limit — average effects cannot carry individual attribution
Georgiana Mirza*Regulatory design* — thresholds, designation, market definitionSection 9's discontinuity material and section 2's point about market definition. Threshold rules are both a research design and an avoidance-behaviour problem

14. If you remember eight things

  1. Sixty percent of the work happens before the model, and cleaning is the single largest block. Saying so is a credibility signal.
  2. The unit of analysis decides what your result can mean. *Case* measures enforcement; *firm* measures conduct. They are not interchangeable, and *market* is a legal conclusion smuggled in as a data field.
  3. Missing At Random does not mean random — it means random given what you already observe. And your own crime data is Not At Random, because the dark figure is the selection. That sentence is the best methodological point you have.
  4. Plot rows per time period first. A step in the series is a change in administration, never a change in the world.
  5. Simpson's paradox: a pattern can hold in every subgroup and reverse when pooled, with every number correct. Check subgroups before reporting a pooled comparison.
  6. Entity resolution is a legal question here, because EU competition law acts on the undertaking rather than the legal person — so the matching threshold affects attribution of liability and the lawful size of a fine.
  7. The design is the argument. Difference-in-differences lives on parallel trends, discontinuity on nothing else jumping at the cutoff, synthetic control on no spillover. And all of them give average effects, which cannot carry individual attribution — which is the real frontier, and the best sentence in this chapter.
  8. On cartel screens: variance, not level. When the frozen-perch conspiracy ended the mean price fell about 16 percent while the standard deviation rose over 200 percent — collusive prices are steadier, not just higher. Graphs add who is anomalous *together*, and a structure that forms and dissolves around a raid is the most persuasive version. **But *Wood Pulp* (C-89/85, 31 March 1993) holds that parallel conduct proves concertation only where no other explanation is plausible — so a screen generates hypotheses and can never be the conclusion.**