Skip to content
VibeFormer
38 min

Legal Analytics

Quantitative methods applied to legal materials. The three things people mean by prediction and why only one is forecasting, the leakage result that reframed the field, precedent as a citation network, and the selection effect that makes litigated cases the worst possible sample.

Listen

0. What it is, and three things it is not

Legal analytics is the quantitative study of legal materials and legal institutions — judgments, dockets, citations, filings, outcomes — treated as data. It is the closest named field to what ATLANTIS's computational strand would be doing, and it has a twenty-year methodological literature that has already found most of the traps. Knowing that literature is cheap credibility, because it means you arrive with the field's known failure modes rather than rediscovering them.

Not thisWhy the distinction matters
Not legal techContract-review products and e-discovery platforms are tools sold to firms. Legal analytics is a research method. Overlapping techniques, different purpose, and conflating them makes you sound like a vendor
Not *AI judging*Nobody serious proposes replacing judges, and the press coverage of this field has been consistently worse than the field. If you bring this up, bring it up to dismiss it
Not natural-language processing with a legal datasetThe method is NLP, graph analysis and statistics. What makes it legal analytics is taking the legal process seriously as a data-generating process — which is the whole content of sections 2 and 4

1. The four data objects

Legal materials are not one kind of data, and each kind supports different questions. Being able to say which object you would need for a given question is a strong, concrete answer.

What a single case actually offers as data

The split that matters most is between the docket and the text. The docket exists before the decision; the text is written afterwards by the institution that decided. That single fact is what section 3 is about, and it is the most consequential methodological point in the field.

2. The three things people call *prediction* — and only one is forecasting

This is the spine of the chapter. Medvedeva, Wieling and Vols argued, in a paper on rethinking automatic prediction of court decisions, that the field had been lumping together three genuinely different tasks under one word — and that most published *prediction* results are not prediction in the sense a reader assumes. The distinction is simple, it is devastating when applied, and it is the single most useful thing in this chapter.

TaskWhat information the model getsWhat it actually demonstrates
Outcome identificationThe published judgment, including the court's own account of the factsThat the outcome is recoverable from a document written after, and by, the deciding body. Useful for automatic annotation of large archives. It is not prediction of anything
Outcome-based categorisationSome part of the judgment, with the explicit result removedThat textual features correlate with outcome. Interesting as a study of judicial reasoning and framing. Still not forecasting, because the document postdates the decision
Outcome forecastingOnly information that existed before the decision — the application, the docket, the parties, the dateGenuine prediction. This is the only one that would be usable prospectively, and it is much harder and much less accurate

3. The ECtHR story, worked through — the field's cautionary tale

This sequence is the best-documented example of the whole problem and it is worth knowing in order, because it runs from a headline result to a careful rebuttal and ends somewhere more honest.

StepThe finding
Aletras, Tsarapatsanis, Preoţiuc-Pietro and Lampos (2016), *PeerJ Computer Science*Textual features from roughly 600 European Court of Human Rights cases predicted whether the Court found a violation, at about 79 percent average accuracy. The paper reports that the formal facts section was the most predictive part, and reads that as consistent with legal realism — that judicial decisions are substantially shaped by the facts
The pressReported as an AI predicting human-rights trials. The paper's own authors pushed back on the AI-judge framing in interviews at the time
A reviewer on the published peer-review recordNoted that the acknowledged limitations leave the paper reliant on crude proxies, and that its arguments would struggle in a legal setting. The caveat was visible from the start
Medvedeva, Vols and Wieling (2020), *Artificial Intelligence and Law*Reported about 75 percent average accuracy across nine Convention articles — and then did the test that mattered: training on past cases and predicting future ones dropped performance to roughly 58 to 68 percent. They also found they could classify outcomes at relatively high accuracy using only the surnames of the judges hearing the case
Medvedeva and colleagues (2021)Moved to forecasting for pending applications, using only information available before the judgment — the methodologically correct framing, and much harder

Why the 79 percent is not forecasting

This is information leakage, and it is the most common fatal flaw in applied machine learning generally. The fix is a strict temporal split and a feature set restricted to pre-decision information, which is exactly what the later work did — and performance fell by ten to twenty points, which is the honest number.

4. Precedent as a network

Citation analysis is the most mature and least controversial part of legal analytics, because it does not predict anything — it measures a structure that genuinely exists. A judgment citing another judgment is a directed edge, and a body of case law is therefore a graph.

A(c)=1−dN+d∑c′→cA(c′)L(c′)A(c) = \frac{1-d}{N} + d \sum_{c' \to c} \frac{A(c')}{L(c')}
The authority of a case equals a small baseline, plus a share of the authority of every case that cites it, divided by how many cases each of those cites in turn.This is the PageRank idea applied to precedent, and the recursion is the point: being cited by an authoritative case counts for more than being cited by an obscure one. Dividing by the number of outgoing citations stops a case that cites two hundred others from passing on full weight to each. Plain-language version: a precedent is authoritative if important cases rely on it, and importance is defined the same way all the way down.
MeasureLegal readingHonest caveat
In-degree — how many cases cite itCrude popularityScales with age. A 1950 case has had seventy years to accumulate citations
Recursive authority (above)Cited *by cases that matter*Better, and still age-dependent
BetweennessA case that bridges two doctrinal areasOften the genuinely interesting one — bridges are where doctrine transfers between fields
Community detectionClusters of mutually citing cases = doctrinal areas, discovered rather than assumedThe appealing use: watching a doctrinal area split or merge over time
Citation decayA case stops being citedCould be overruled, superseded, or simply the question stopped arising. The graph cannot tell you which

5. The selection effect — and why litigated cases are the worst available sample

This is the catchment problem, and it has a classic formal statement. Priest and Klein, in *The Selection of Disputes for Litigation* (1984), argued that the disputes which reach adjudication are neither a random nor a representative sample of the disputes that exist — because parties settle the cases whose outcomes they can both predict.

(Pp−Pd) J  >  Cp+Cd⟺Pp−Pd  >  Cp+CdJ(P_{p} - P_{d})\, J \;>\; C_{p} + C_{d} \qquad\Longleftrightarrow\qquad P_{p} - P_{d} \;>\; \frac{C_{p} + C_{d}}{J}
A case goes to trial rather than settling when the gap between the plaintiff's and the defendant's estimates of the plaintiff's chance of winning, multiplied by the amount at stake, exceeds the combined cost of litigating.**Read the right-hand form: trial requires the two sides to *disagree* by more than the cost ratio. If both sides agree the plaintiff will probably win, there is a settlement both prefer. So litigation selects for cases where the parties disagree — which means cases that are genuinely close. Priest and Klein's famous corollary is a tendency toward 50 percent plaintiff victories among litigated cases, with the party having more at stake winning more often. Note the status of that result: later formal work (Lee and Klerman, 2016) shows the exact 50 percent figure is sensitive to the bargaining and distributional assumptions. So treat it as a strong tendency with known conditions, not a law — and saying that is the correct level of confidence.**

The funnel — and your dataset is the last layer

Every layer of this funnel is a selection step, and the settlement layer selects on predictability itself. This is why a plaintiff win rate computed from litigated cases cannot be read as the strength of plaintiffs' claims in general, and it is the single most important thing to understand before building any dataset of decided cases.

6. Judicial behaviour analytics — where the law has already drawn a line

Analysing how individual judges decide is the most legally fraught application in the field, and it is the one place where a jurisdiction has banned the research outright.

7. Legal analytics in competition law specifically

This is the part that connects the chapter to the actual job, and the chair of your panel is one of the people building it.

What existsWhat it is
Stanford Computational AntitrustFounded by Thibault Schrepel. The project that named the field, and it runs an annual cross-agency report — the fourth, with Teodora Groza, drew on 25 agencies
A knowledge graph of European Commission decisions, 1977 to 2025His own project, and it is legal analytics in the precise sense of this chapter — the decisional record as a structured, queryable graph rather than a pile of PDFs
The *digital brain* proposal, April 2026Offered to agencies: consistency-check a draft decision against the full decisional record. This is the single most interesting artefact to have an opinion about, because it is citation-and-consistency analysis applied to an agency's own output rather than prediction of a third party
His Kluwer post, July 2026Titled *Competition Agencies Compute. The Law Still Assumes They Read.* States that all 27 EU national competition authorities and DG Competition rely on computational tools. It is a title, so you can attribute it precisely

8. The probes, and what to say

PanellistLikely probeYour answer
Thibault Schrepel*What would you do with the decisions corpus?*Section 7's consistency-check answer: retrieval recall, not prediction accuracy. Plus section 2's identification-versus-forecasting distinction, which is the framing contribution
Catalina Goanta*How would you validate a coding scheme over legal text?*The NLP chapter's validation plan: a defined codebook, two coders, an agreement statistic fixed in advance, and a reported threshold. **Not *I would fine-tune a model***
Tijmen Wisman*What about the people being analysed?*Section 6. The French prohibition, the institution-versus-individual line, and the fact that your own site sits on the wrong side of it — volunteered
Georgiana Mirza*Doctrinal change and regulatory design*Section 4's community detection and citation decay as a way of watching a doctrinal area split — and the caveat that decay cannot distinguish overruled from irrelevant

9. If you remember seven things

  1. Three tasks, one word. Outcome *identification* reads a decided judgment; outcome-based *categorisation* removes the explicit result; only outcome *forecasting* uses pre-decision information. Most headline accuracy figures are the first.
  2. Aletras and colleagues, 2016, PeerJ Computer Science: about 79 percent on roughly 600 ECtHR cases, with the facts section most predictive — but the facts section is written after and by the deciding court.
  3. Medvedeva, Vols and Wieling, 2020: about 75 percent, falling to roughly 58 to 68 percent when forecasting future cases from past ones — and high accuracy from judges' surnames alone. That pair of numbers is the most useful thing you can quote in this area.
  4. **Fowler and Jeon, 2008, *Social Networks*: 30,288 US Supreme Court majority opinions, 1754 to 2002, with centrality used to quantify precedent authority. Caveat: citation is partly rich-get-richer, so centrality measures salience rather than quality. And a citation can be a rejection**.
  5. Priest and Klein, 1984: litigated cases are neither random nor representative, because parties settle what they can both predict — so litigation selects for close cases. The tendency toward 50 percent is real but assumption-sensitive.
  6. The enforcement funnel has the same shape as the dark figure in your own paper. Detected cartels are selected on concealment competence. Say this — it is the bridge between your work and theirs.
  7. France, Article 33 of Law 2019-222 of 23 March 2019: analysing or predicting individual judges' practices is prohibited, with penalties up to five years. Study the institution, not the individual — and your own LawReformer judge-behaviour feature is on the wrong side of that line.