Data Science: The Actual Craft
The part nobody writes papers about and every project dies on. The lifecycle, what exploratory analysis actually means, the three kinds of missing data, rates against counts, Simpson's paradox worked numerically, entity resolution as a legal question, the four quasi-experimental designs, whether EDA and graph theory can actually find a cartel, and reproducibility as evidence.
0. Why this chapter, and the distinction to open with
Every other technical chapter here is about a method — a model, a measure, a formalism. This one is about the craft around the method, and it is the chapter that matters most for a project like ATLANTIS. The reason is blunt: competition agencies do not fail at modelling, they fail at data. Firm names do not match, the denominator is wrong, the missing rows are missing for a reason, and nobody can rerun last year's number.
1. The lifecycle, with honest proportions
Textbooks draw this as a tidy loop with equal boxes. It is not equal. The proportions below are the ones practitioners actually report, and saying them out loud is a credibility signal — it tells a panel you have run a project rather than read about one.
Where the time actually goes
2. The unit of analysis — the first decision, and the one most often wrong
The unit of analysis is the thing that each row of your table represents. It sounds trivial. It is the decision that silently determines what your result can mean, and getting it wrong invalidates everything downstream no matter how good the model is.
| Candidate unit | One row is… | So the result is a claim about… |
|---|---|---|
| Firm | One company | Firms. *Larger firms are more likely to be investigated* |
| Firm-year | One company in one year | Firm behaviour over time. Lets you ask whether a rule change altered conduct |
| Case | One investigation | Enforcement, not conduct. This is the switch that catches people out |
| Decision paragraph | One passage of reasoning | Reasoning style — what agencies cite, how they argue |
| Tender | One procurement competition | Bidding patterns. The natural unit for cartel screens |
| Firm-pair | Two companies considered together | Relationships — co-bidding, interlocking directorates, common ownership |
| Market | One product-and-geography combination | Market structure. But *market* is a legal conclusion, not a data field |
3. Exploratory data analysis — what you actually look at
Exploratory data analysis, or EDA, means looking at the data before modelling it — not to test anything, but to find out what you are holding. It has a reputation as the soft part of the job. It is in fact where almost every serious error is caught, and it is mostly a checklist rather than an art.
| Check | What you are looking for | The failure it catches |
|---|---|---|
| Row and column count | Does it match what the source claimed? | A truncated download. Depressingly common |
| One row, read in full | Does a single record make sense as a sentence? | Columns shifted by one; a header row read as data |
| Type of every column | Is the number column actually numeric? | Numbers stored as text, so sorting puts 10 before 9 |
| Min and max of every number | Impossible values | A negative turnover; a fine of 0; a date in 1900 or 2099 |
| Count of unique values | Too few or too many | A *country* column with 340 values — spelling variants |
| Missing count per column | How much is absent, and where | A column that is 90 percent empty and about to be used |
| Exact duplicate rows | Identical records | A double-loaded file inflating every count |
| Distribution of each number | Shape: one hump, two humps, a long tail | A bimodal column usually means two populations mixed together |
| Counts over time | Rows per month or year | A reporting change. A step in the series is almost never real |
| Cross-tabs of key pairs | Combinations that should not exist | Decisions dated before the investigation opened |
4. Missing data — the three kinds, and why deleting rows is a choice
Almost every real dataset has holes. The critical insight is that the holes have a cause, and the cause determines what you are allowed to do about them. There is a standard three-way classification, due to the statistician Donald Rubin, and the names are unfortunately confusing — so here they are with the confusion removed.
The three kinds of missingness, as a decision
| Kind | A legal-data example | What you may do |
|---|---|---|
| MCAR | A digitisation batch was skipped, and which batch was arbitrary | Drop the rows, or impute. Estimates stay unbiased; standard errors widen |
| MAR | Smaller firms report turnover less completely — and you observe firm size | Multiple imputation, or inverse-probability weighting, *conditioning on size*. Complete-case analysis is biased here |
| MNAR | Undetected cartels are absent precisely because they were good at concealment. Unreported crime is unreported because of the nature of the offence | Nothing, from the data alone. Report a bound, model the selection explicitly, or find an external source. Imputation here manufactures a result |
On imputation. *Imputation* means filling a hole with an estimate. The naive version — substitute the column mean — is almost always wrong, because it pretends you know the value exactly and so shrinks your uncertainty artificially. Multiple imputation is the defensible version: fill the holes several times with different plausible draws, run the analysis on each completed dataset, then combine. The combination rule is what makes it honest.
5. Rates against counts — denominators, and the trap in both directions
A count is how many. A rate is how many per something. Choosing between them is not presentation, it changes the finding — and both choices have a failure mode, which is why this is a section rather than a footnote.
| Use | Problem | Example of the problem biting |
|---|---|---|
| Raw counts | Scale with the size of the unit, so the biggest unit always looks worst | Uttar Pradesh records more offences than Goa because it has fifty times the people. Germany has more cartel cases than Malta. Neither is a finding |
| Per-capita rates | Small denominators make unstable rates. With a tiny population, one extra event swings the rate wildly | A district of 5,000 with 2 events is at 40 per 100,000; a third event takes it to 60. A fifty percent jump from one incident — and such districts will fill both tails of your ranking |
| Either, across a reporting change | The denominator or the definition shifted mid-series | A boundary redrawn, a category merged, a threshold changed. The series breaks and the break looks like an effect |
6. Simpson's paradox — worked numerically, because it has to be
Simpson's paradox is when a pattern holds in every subgroup and reverses when the subgroups are pooled. It sounds like a curiosity. It is the single most dangerous arithmetic fact in applied work, because both the pooled number and the subgroup numbers are correct — so no amount of checking the computation will reveal it. The example below is in Catalina Goanta's research area deliberately, since she is on your panel.
The question: does labelling a post as an advertisement reduce engagement? Take 2,000 posts, split by influencer size, and count how many achieved high engagement.
| Group | Disclosed as an ad | Not disclosed | Which wins? |
|---|---|---|---|
| Small influencers | 90 of 100 = 90% | 720 of 900 = 80% | Disclosed, by 10 points |
| Large influencers | 540 of 900 = 60% | 45 of 100 = 45% | Disclosed, by 15 points |
| Pooled | 630 of 1,000 = 63% | 765 of 1,000 = 76.5% | Not disclosed, by 13.5 points |
Why the reversal happens, structurally
7. Entity resolution — and why it is a legal question, not a cleaning step
Entity resolution is deciding which records refer to the same real-world thing. It is the least glamorous task in this chapter and the one most likely to decide whether a competition-law dataset is usable at all, because the records are firm names and firm names are chaos.
The pipeline
| Tool | What it measures | Where it fails |
|---|---|---|
| Levenshtein distance | How many single-character edits turn one string into the other | Googel→Google is 1 edit, good. But two unrelated short names can also be 1 edit apart |
| Jaro-Winkler | Character overlap, weighted towards agreement at the start of the string | The usual default for names, because company names rarely differ at the beginning — which is also exactly why it misses Alphabet/Google |
| Token-set comparison | Treats the name as a bag of words, so word order stops mattering | Bank of Ireland and Ireland Bank match — sometimes right, sometimes two different banks |
| Identifier join | Match on a registration number, VAT number or LEI instead of a name | Always prefer this when it exists. It fails only by being absent, which in public legal corpora it usually is |
8. Measurement validity — the gap between what you want and what you have
The construct is what you want to measure. The operationalisation is what you actually measured. Validity is the size of the gap. Almost every criticism worth making of an empirical legal paper is a validity criticism, and almost none of them can be fixed by a better model.
| What you want (construct) | What is in the data | The gap, stated plainly |
|---|---|---|
| Collusion | Cartel decisions | You are measuring enforcement. Detection depends on leniency programmes, agency resources and concealment skill. A fall in cases may mean less collusion or a worse-funded agency |
| Harm to consumers | Fines imposed | Fines track turnover and procedure, not harm. A large fine on a small harm is routine |
| Market power | Market share | Share is a proxy, and the legal threshold is dominance — which depends on entry barriers and countervailing buyer power that no share captures |
| Non-compliance with disclosure rules | Posts without a visible ad label | Misses paid posts you cannot identify as paid. The denominator is unknown, which is the hardest version of this problem |
| Crime | Recorded offences | Recording depends on reporting, and reporting depends on the offence. The dark figure, and it is MNAR |
| Algorithmic discrimination | Outcome differences across groups | Differences may reflect the decision rule, the data, or genuine differences in the underlying population — and the three need different remedies |
9. The four quasi-experimental designs — how to claim a cause without an experiment
An experiment randomises who gets the treatment, which is why it supports causal claims. You cannot randomise which firms are investigated or which member states get a regulation. A *quasi-experiment* exploits a situation where something close to random assignment happened anyway. These four designs are the workhorses of empirical legal research, and knowing what each one assumes is more useful than knowing how to fit it — because the assumption is where the argument lives.
| Design | The counterfactual it builds | The assumption it dies on |
|---|---|---|
| Difference-in-differences | What the treated group's trend would have been, taken from an untreated group's trend | Parallel trends — absent treatment, both groups would have moved together. *Not* that they were at the same level |
| Regression discontinuity | Units just below a threshold, compared with units just above | Nothing else jumps at the cutoff, and units cannot precisely place themselves on either side |
| Synthetic control | A weighted blend of untreated units, chosen to track the treated unit's pre-treatment path | No spillover onto the donor units, and a long enough pre-period to fit the blend honestly |
| Event study | The same unit, just before the event | Nothing else happened in the window, and the date is correct and not anticipated |
Difference-in-differences, and what parallel trends means
10. Putting it together — can EDA and graph theory find a cartel?
This section is the applied synthesis of the chapter, and it is the most likely thing you would actually be asked to design. Short answer to the question in the heading: **EDA and graph theory can tell an agency *where to look*. Neither can establish that a cartel exists.** Holding that line is what makes the rest credible, so it is stated at the start and again at the end.
What a cartel leaves behind. Collusion is an agreement, which is not in the data. What *is* in the data are its consequences:
- Prices that behave differently from competitive prices — higher, and crucially *steadier*, which is the basis of the main screen below.
- Bids that do not look like independent attempts to win — because under an allocation agreement, most bidders are trying to lose.
- Market shares that are too stable — allocation agreements have to be monitored, and monitoring shows up as shares that barely move.
- Costs that stop passing through to prices — a competitive firm must follow its input costs; a cartel can absorb a shock and hold the price.
- Relationships that are too consistent — the same firms appearing together, in the same roles, across many tenders. This is the part only a graph can see.
- A structural break at a known date — a dawn raid, a leniency application, a cartel collapse. The most legally useful pattern in the list, because it gives you a before and an after, which is the difference-in-differences setup from section 9.
The EDA screens, as a checklist. Each is a plot or a statistic, and each has a legitimate alternative explanation — which is the point of the limits section below.
- Coefficient of variation of price, by market and by period. Look for low-variance pockets. This is the primary screen.
- Price level against a benchmark — a comparable region, a comparable product, or the same market before and after. Never a bare level; always a comparison.
- Cost pass-through. Regress price on an input-cost index. A pass-through coefficient near zero where it should be substantial is a strong signal, because holding price against a cost shock requires coordination or market power.
- Price series plotted around known dates — raids, leniency applications, arrests. A drop in mean and a jump in variance immediately after is the signature of a cartel ending.
- Market shares over time. Plot each firm's share as a line. Lines that are suspiciously flat, or that swap places in a regular pattern, suggest allocation.
- The distribution of losing bids. In genuine competition, bids cluster near cost with a right tail. Under cover bidding, losers sit implausibly far above the winner — and the gap between first and second is informative in a way the winning bid alone is not.
- The winner sequence. Test whether the identity of the winner across successive tenders is more regular than chance. Rotation is a pattern randomness does not produce.
- Bid-to-cost or bid-to-estimate ratios. Screens out the effect of tenders simply being different sizes.
- Round numbers and digit patterns. Genuine cost-based bids rarely land on round figures as often as coordinated ones do.
- Entry and exit counts. A market with no new entrants over a long period is both a precondition for a stable cartel and a consequence of one.
- Participation patterns. Which firms bid for which tenders, and — more telling — which firms consistently decline to bid where they plainly could have. Bid suppression is an absence, and absences are invisible unless you plot them deliberately.
What the bid distribution looks like, competitive against coordinated
Now the graphs. EDA treats each firm or tender as an independent row. Graph theory is what you reach for once you suspect the rows are related — and cartels are, by definition, a relationship. Five graphs are worth building from procurement data:
- Firm–tender bipartite graph. Two kinds of node, edges only between kinds. This is what procurement data literally is, so build this first and keep it — it is the ground truth the others are derived from.
- Firm–firm co-bidding graph. An edge when two firms bid for the same tender, weighted by how often. The workhorse. Derived from the bipartite graph by projection — see the warning below.
- Winner–loser directed graph. An edge from loser to winner. Rotation shows up as cycles, and a hub-and-spoke arrangement shows up as a node with unusual in-degree.
- Subcontracting graph. An edge from winner to the firms it subcontracts to. This is the strong one, because compensating the losers is a cartel's hardest practical problem — and a firm that reliably loses to a given winner and is then reliably subcontracted by that winner is a pattern with few innocent readings.
- Ownership and interlocking-directorate graph. Firms linked by common shareholders, or by sharing a director. Nodes here include people, and the same entity-resolution problem from section 7 applies with the extra difficulty that people change names.
What graph theory then tells you, measure by measure, with the legal reading beside it:
- Connected components — a group reachable only within itself. Reads as market allocation or a genuinely separate regional market. Separate components in a market you believed was national is a finding worth checking.
- Cliques — every member connected to every other. A closed bidding ring makes exactly this shape, so enumerating large cliques is a direct search for candidate rings.
- k-cores — the subgraph where every node has at least k connections within it. More robust than cliques, because a single missing edge destroys a clique but not a core.
- Community detection and modularity — groups denser inside than between. Candidate cartel membership, discovered rather than assumed. The honest framing is that modularity gives you a partition for *every* graph, so a high score needs comparing against a randomised benchmark.
- Local clustering with low overall density — a tight group inside an otherwise sparse market. The signature of a closed group operating within a larger market.
- Betweenness centrality — who sits on the paths between others. The candidate coordinator, trade-association hub, or spoke-connecting party in a hub-and-spoke arrangement.
- Weighted-edge outliers — two firms that co-bid far more often than their sizes and sectors would predict. The single cheapest thing to compute and often the most informative.
- Temporal graphs — rebuild the graph per year and watch edges appear and vanish. A structure that forms at one date and dissolves at another, especially around a raid, is the most persuasive graph evidence available, because it adds a before and an after.
- Bipartite co-occurrence — firms that appear in tender sets together far more often than chance. Computed on the bipartite graph directly, this avoids the projection problem below.
- The bipartite projection trap. Projecting a firm–tender graph into a firm–firm graph mechanically creates a clique for every tender — ten firms bidding for one contract become ten mutually connected nodes. So clustering coefficients computed on a projected graph are inflated by construction and mean nothing without a null model. This is the single most common technical error in this area, and it is covered in the network analysis chapter.
- Everything has a lawful explanation. Specialisation, capacity limits, geography, prequalification rules, consortium bidding and joint ventures all produce co-bidding structure legitimately. Firms that only bid for what they can actually build will look clustered.
- The base-rate problem. Cartels are rare. A screen with excellent accuracy still produces mostly false positives when the underlying rate is low — and converting *this pattern would be unusual if they were competing* into *they are probably colluding* is the prosecutor's fallacy. It needs a base rate nobody has.
- Entity resolution decides the graph. Get the corporate group wrong and you have invented or destroyed edges. Section 7's point lands here with full force: the matching threshold is upstream of every graph measure.
- Multiple testing. Screen 500 markets at a five percent threshold and about 25 will flag by chance. Any screening programme needs a correction or a stated expected false-positive count, in advance.
- Screens are gameable once known. A published screen is a specification for evading it — add noise to prices and the variance screen goes quiet. So screens detect unsophisticated collusion, which makes the detected set a biased sample of the real set. This is section 4's MNAR problem arriving in operational form.
- Data access is the binding constraint, not method. Bid-level procurement data with firm identifiers is often unavailable to researchers. Say this plainly — it is why synthetic data and public tender portals matter.
The pipeline, in order, if you were asked to design one. This is the answer to *what would you build first* and it is deliberately unglamorous:
- Get the bid-level data with firm identifiers, and record provenance per tender. Without identifiers there is no graph.
- Resolve entities to undertakings, report the threshold, hand-check a sample and print the largest cluster size. Section 7 — and this step is where the project is won or lost.
- Do the EDA before anything clever. Rows per period first, then the price and bid distributions, then the unique-value checks.
- Build the bipartite graph and keep it as the base object. Derive everything else from it rather than from a projection.
- Run the variance and pass-through screens to flag candidate markets and periods.
- Build the graph only for the flagged markets, and compare every structural measure against a degree-preserving randomised null model. Without the null model the numbers are decoration.
- Rebuild per year and look for structures that form and dissolve, especially around known enforcement dates.
- Enumerate the innocent explanations for each flag — capacity, geography, specialisation, lawful consortia — and try to eliminate them. This is the Wood Pulp step and it is the one that decides whether any of it is usable.
- Hand the survivors to a case team as hypotheses with stated error rates, never as conclusions. And log what the pipeline saw on the day it flagged them, so the selection can be reconstructed and contested.
11. Reproducibility — and why it is a legal requirement here, not hygiene
Reproducibility means someone else, with your data and your code, gets your number. In most fields this is professional good manners. In this project it is closer to an evidentiary standard, because a measurement that feeds an enforcement decision may have to be explained to a court.
| Practice | What it prevents | The minimum version |
|---|---|---|
| Raw data is never edited | The silent, unrecoverable change | Keep raw/ read-only. Every transformation writes to interim/ or processed/ |
| Every step is a script | *I fixed that one in the spreadsheet and cannot remember how* | If a step was done by hand, it is not reproducible. No exceptions |
| Random seeds fixed | A different answer on rerun | Set the seed once at the top. And report that results are seed-dependent if they are |
| Environment pinned | The result changing because a library updated | A requirements.txt with exact versions, not ranges |
| Data versioned or hashed | Not knowing which vintage produced the number | Record the download date and a checksum of the file |
| One command runs everything | A pipeline only its author can operate | A single script that goes from raw/ to every figure in the paper |
| Decisions logged with reasons | Unexplainable choices at review time | The matching threshold, the exclusion rules, the imputation model — a short written record of each |
12. The seven ways an empirical legal result dies
A compression of the whole chapter. If a result of yours is going to be attacked, it will almost certainly be on one of these seven, and each has a pre-emptive sentence.
| Cause of death | The attack | Your pre-emptive sentence |
|---|---|---|
| Wrong unit | *Your district-level finding does not support a claim about people* | State the level in the result itself. *At district level…* |
| Selected sample | *You measured detection, not conduct* | Name it as MNAR and give the direction of the bias |
| Bad denominator | *Per capita is the wrong base for enforcement counts* | Report the result under two or three denominators |
| Hidden confounder | *This reverses if you split by firm size* | Check the obvious subgroups before publishing; present the design, not just the regression |
| Entity errors | *Your market shares are wrong because the group was mis-assembled* | Report the threshold, the largest cluster size, and a hand-checked sample |
| Overclaimed significance | *Your p-value is 0.156 and your abstract says* proving | Fix the abstract. This one is live in your own paper and you should raise it first |
| Not reproducible | *Nobody can rerun this, including you, in six months* | One command from raw data to every figure |
13. How this maps onto ATLANTIS, and the probes per panellist
The project combines legal analysis, computational science and institutional economics, and the announced central question is not whether these tools will be used, but how. That phrasing matters for this chapter: **the *how* is overwhelmingly a data-craft question rather than a modelling one.** The safeguards owed to companies and individuals are, operationally, most of the checklist above.
| Panellist | Likely probe from this chapter | Where your answer comes from |
|---|---|---|
| Thibault Schrepel | *What would you actually have to build before any of this is measurable?* — a scoping question, not a technical one | Section 1's proportions and section 7. The honest answer is that the first year is a corpus and an entity layer, and that this is the unglamorous precondition for everything else |
| Catalina Goanta | *How would you handle the data side?* — her own work is measurement over messy multilingual text at scale | Sections 3, 7 and 8. The disclosure example in section 6 is in her area — use it as an illustration of why you check subgroups, not as a comment on her findings |
| Tijmen Wisman | *What about the people on the other side of these systems?* | Section 10 and SyRI: reproducibility as a condition of a reviewable proportionality assessment. And section 9's closing limit — average effects cannot carry individual attribution |
| Georgiana Mirza | *Regulatory design* — thresholds, designation, market definition | Section 9's discontinuity material and section 2's point about market definition. Threshold rules are both a research design and an avoidance-behaviour problem |
14. If you remember eight things
- Sixty percent of the work happens before the model, and cleaning is the single largest block. Saying so is a credibility signal.
- The unit of analysis decides what your result can mean. *Case* measures enforcement; *firm* measures conduct. They are not interchangeable, and *market* is a legal conclusion smuggled in as a data field.
- Missing At Random does not mean random — it means random given what you already observe. And your own crime data is Not At Random, because the dark figure is the selection. That sentence is the best methodological point you have.
- Plot rows per time period first. A step in the series is a change in administration, never a change in the world.
- Simpson's paradox: a pattern can hold in every subgroup and reverse when pooled, with every number correct. Check subgroups before reporting a pooled comparison.
- Entity resolution is a legal question here, because EU competition law acts on the undertaking rather than the legal person — so the matching threshold affects attribution of liability and the lawful size of a fine.
- The design is the argument. Difference-in-differences lives on parallel trends, discontinuity on nothing else jumping at the cutoff, synthetic control on no spillover. And all of them give average effects, which cannot carry individual attribution — which is the real frontier, and the best sentence in this chapter.
- On cartel screens: variance, not level. When the frozen-perch conspiracy ended the mean price fell about 16 percent while the standard deviation rose over 200 percent — collusive prices are steadier, not just higher. Graphs add who is anomalous *together*, and a structure that forms and dissolves around a raid is the most persuasive version. **But *Wood Pulp* (C-89/85, 31 March 1993) holds that parallel conduct proves concertation only where no other explanation is plausible — so a screen generates hypotheses and can never be the conclusion.**