NLP and Unstructured Data
Preprocessing through embeddings, classification and sequence models, the evaluation metrics, and the part that matters most for this project: how you validate unstructured data, and why legal text is the hardest case there is.
0. Why this chapter is load-bearing
Almost everything a competition agency holds is unstructured. Decisions, Statements of Objections, seized emails, meeting notes, internal memoranda, leniency applications, tender documents. The structured parts — bid values, dates, firm identifiers — are the minority, and they are the part everyone already knows how to handle.
1. Structured, semi-structured, unstructured
| Kind | What it is | Agency example |
|---|---|---|
| Structured | Fixed schema, typed fields, one meaning per column | A tender award table: contract reference, buyer, winner, value, date, CPV code |
| Semi-structured | Self-describing but irregular. Tags or keys exist, presence and nesting vary | An Akoma Ntoso or XML judgment with marked-up articles and recitals; a JSON export of case metadata |
| Unstructured | No schema. Meaning is carried entirely by the content | The body of a decision, a seized email thread, a handwritten meeting note, a scanned PDF annex |
2. Preprocessing, and where it silently destroys meaning
- Sentence and word segmentation. Splitting text into sentences and tokens. Harder than it looks: abbreviations, citations and decimals all contain full stops.
- Case folding. Lowercasing everything. Cheap, and it conflates distinctions that matter.
- Stop word removal. Dropping high-frequency words like *the*, *of*, *not*.
- Stemming. Chopping suffixes by rule. Fast, crude, produces non-words — *organisation* becomes *organis*.
- Lemmatisation. Mapping to the dictionary form using morphology and part of speech. Slower, correct — *was* becomes *be*.
- Normalisation. Unifying quotes, dashes, whitespace, Unicode forms, and digits.
- OCR. Turning scanned pages into text, with an error rate that is never zero.
3. Representing text as numbers
| Representation | How it works | Strength and limit |
|---|---|---|
| Bag of words | Count each vocabulary term, discard order | Trivially interpretable. Loses word order entirely, so *A sued B* and *B sued A* are identical |
| TF-IDF | Term frequency times inverse document frequency, so a word scores highly when it is frequent here and rare across the corpus | Still the right default for document retrieval and a strong classification baseline. Sparse, fast, and inspectable — you can print the top-weighted terms and read them |
| n-grams | Count short sequences rather than single words | Recovers some local order. Vocabulary explodes |
| Word2Vec and GloVe | Dense vectors learned from co-occurrence, so similar words sit near each other | Captures semantic similarity. One vector per word regardless of context, so the bank of a river and a commercial bank share a vector |
| Contextual embeddings | A transformer produces a different vector for a word in each context | Resolves polysemy. Opaque, and expensive relative to TF-IDF |
| Sentence and document embeddings | A single vector for a whole passage | The basis of semantic search and of retrieval in RAG |
Once documents are vectors, *similar* needs a definition. The standard one is cosine similarity: the angle between two vectors, ignoring their lengths.
BM25 is the retrieval function that actually wins in practice, and it is TF-IDF with two corrections: term frequency saturates, and document length is normalised explicitly.
4. The core tasks
| Task | What it produces |
|---|---|
| Part-of-speech tagging | A grammatical label per token: noun, verb, modal, determiner |
| Named entity recognition | Spans labelled as person, organisation, location, date, money, and in legal work: court, statute, case citation, party role |
| Dependency parsing | The grammatical relations in a sentence — who did what to whom |
| Coreference resolution | Linking *it*, *the undertaking*, *the applicant* back to the entity meant |
| Relation extraction | Typed links between entities: Acme subcontracted-to Beta |
| Text classification | A label for a document or passage |
| Sequence labelling | A label per token, which is how NER is actually implemented |
| Summarisation | Extractive, selecting original sentences, or abstractive, generating new text |
| Question answering | Extractive from a passage, or generative |
5. Models, from classifiers to sequence to sequence
- Linear classifiers over TF-IDF. Logistic regression or a linear SVM. Still the baseline to beat, interpretable, and often within a few points of anything fancier on well-defined document classification.
- Naive Bayes. Fast, works oddly well on text, badly calibrated because the independence assumption is false.
- Recurrent models. An LSTM or GRU reads a sequence left to right, carrying a hidden state. Handles order; struggles with long dependencies and cannot be parallelised across the sequence.
- Convolutional text models. Filters over word windows, good at local pattern detection such as key phrases.
- Sequence to sequence. An encoder compresses the input, a decoder generates the output. Translation, summarisation, and question answering all fit this frame.
- Attention, then transformers. The decoder looks back at all encoder positions rather than relying on one compressed state. Removing recurrence entirely gives the transformer — which is chapter 9.
- Encoder-only transformers. BERT-style, bidirectional, fine-tuned for classification and extraction. Still the right tool for labelling a corpus, and far cheaper than a generative model.
6. Evaluation metrics, and what each hides
| Metric | Measures | Limitation |
|---|---|---|
| Accuracy | Fraction of correct labels | Useless under class imbalance, which legal corpora always have |
| Precision, recall, F1 | Per class, then averaged | Macro average treats every class equally; micro is dominated by frequent classes. Which you report changes the story, so state it |
| Exact match | Whole output identical to reference | Brutal and appropriate for extracting a citation or a date |
| Token-level F1 | Overlap between predicted and reference spans | Standard for extractive question answering and NER |
| BLEU | n-gram overlap with a reference, precision-oriented | Built for translation. Rewards surface similarity and punishes correct paraphrase |
| ROUGE | n-gram and longest-common-subsequence overlap, recall-oriented | Standard for summarisation, and only weakly related to whether the summary is true |
| BERTScore | Similarity in embedding space rather than surface form | Credits paraphrase. Opaque, and still says nothing about factual correctness |
| Perplexity | How surprised a language model is by held-out text | A fluency measure. Not a correctness measure |
| Inter-annotator agreement | Cohen's or Fleiss' kappa, Krippendorff's alpha | Measures whether the task is well defined. If humans cannot agree, no model score means anything |
7. Validating unstructured data
This is the section you asked for, and it is the one with the least written about it. Structured validation is a solved engineering problem. Unstructured validation is mostly unsolved, and in law it is barely posed.
| Check | Structured data | Unstructured data |
|---|---|---|
| Schema conformance | Types, ranges, required fields, foreign keys. Deterministic pass or fail | No schema exists. The nearest equivalent is whether an extraction pipeline produced the expected fields, which validates the pipeline rather than the data |
| Completeness | Count nulls per column | Unknowable. You cannot count the clauses a document failed to contain, or the emails never produced |
| Consistency | Cross-field rules — award date after publication date | Requires extracting the facts first, so every inconsistency check inherits the extraction error rate |
| Uniqueness | Primary keys and deduplication | Near-duplicates everywhere: drafts, translations, annexes, forwarded email chains. Exact matching misses almost all of it |
| Accuracy | Compare against a register or a source system | Compare against what? Usually against human annotation, which is itself fallible and which is where inter-annotator agreement becomes the real measurement |
| Provenance | Row-level lineage | Document-level at best. Which paragraph a claim came from is itself an extraction task |
8. Why legal text is the hardest unstructured data there is
Not rhetoric. Nine specific properties, each of which breaks a standard assumption.
- Open texture. Legal concepts have deliberately indeterminate boundaries. *Reasonable*, *necessary*, *appreciable*, *dominant* are not vague by accident; their content is settled by adjudication over time. So the ground truth is itself contested, and there may be no label to learn.
- Negation and scope carry the meaning. *Shall not*, *unless*, *save where*, *without prejudice to*. Models handle negation poorly and these constructions nest several deep.
- Sentences are extremely long and deeply subordinated. A single provision can run three hundred words with six embedded conditions. Parsers and chunkers both degrade.
- Cross-reference is pervasive. *Within the meaning of Article 4(7)* means the sentence cannot be understood in isolation, so chunking a statute destroys it.
- Defined terms override ordinary meaning. *Undertaking* means one thing in competition law and another in ordinary English, and a contract may define *Agreement* to mean something idiosyncratic. Context windows and embeddings both assume ordinary usage.
- Temporal versioning. Provisions are amended. The text that applied in 2019 is not the text in force today, and a model trained on current consolidated law will misread a 2019 decision.
- Multilingualism with legal equivalence. Twenty-four EU languages, all equally authentic. Translations are not paraphrases; divergence between language versions is itself a ground of interpretation.
- Severe class imbalance and tiny positive sets. The interesting category — the rigged tender, the abusive clause — is rare, and sometimes you have a dozen labelled examples in total.
- The cost of error is asymmetric and legally cognisable. A false positive is a dawn raid on an innocent firm. A mislabelled film review is nothing. Standard benchmarks assume symmetric costs.
9. Vocabulary you will hear
| Term | What it means |
|---|---|
| Corpus | The body of text you are working on. Plural corpora |
| Token | The unit a model actually processes. Roughly a word-piece, not a word |
| Vocabulary | The set of tokens a model knows |
| Out of vocabulary | A term the model has no token for, historically a major failure source |
| Gold standard, ground truth | Human-annotated labels treated as correct for evaluation |
| Annotation, labelling | The human process of producing those labels |
| Inter-annotator agreement | How often independent humans assign the same label. Kappa or alpha |
| Codebook | The written rules annotators follow. If it is not written down, the labels are not reproducible |
| Adjudication | Resolving disagreements between annotators, and recording how |
| Distant supervision | Generating noisy labels automatically from an existing resource rather than by hand |
| Zero-shot and few-shot | Performing a task with no, or very few, labelled examples |
| Transfer learning | Taking a model trained on one task and adapting it to another |
| Domain shift | The deployment data differs from the training data. Legal corpora shift constantly as law changes |
| Embedding space | The geometry in which semantic similarity becomes distance |
| Cosine similarity | The standard distance measure between embeddings |
| Sparse and dense retrieval | Term-matching such as BM25 against embedding-based semantic search |
| BM25 | The strong classical ranking function. Still competitive, and the baseline any dense method should be compared against |
| Hybrid search | Combining sparse and dense retrieval, which usually beats either alone |
10. How this connects to ATLANTIS
| Strand | The NLP content of it |
|---|---|
| Accuracy, the data problem | Most of what an agency collects is text. Whether its inferences are sound depends on extraction accuracy, chunking integrity, and whether the corpus is complete — and completeness is the one property unstructured data cannot report about itself |
| Fairness | Article 296 asks whether a decision explains itself. That is a property of text, and it is measurable: specificity, template reuse, whether the stated ground accounts for the outcome. The DSA statement-of-reasons corpus is the instrument |
| Institutional arrangements | Twenty-four equally authentic languages and 27 national authorities. Cross-language consistency is not a nicety; it determines whether a tool deployed in one Member State means the same thing in another |
11. What the panel brings to this chapter
| Panel member | Where NLP meets their work |
|---|---|
| Thibault Schrepel | Built a knowledge graph of every Commission competition decision from 1977 to 2025, with decisions as nodes and doctrinal links as edges. Extracting those links from decision text is relation extraction over legal prose — this chapter is how that artefact was made. He also offers to help agencies build queryable systems from their own decisional records, which is the same pipeline |
| Catalina Goanta | The strongest fit on the panel. Her multi-country longitudinal study of influencer disclosure across around a million posts is classification and measurement over unstructured multilingual text at scale, with every validation problem in section 7. Her legal compliance API proposal is about making compliance machine-checkable, which is the same instinct |
| Georgiana Mirza | Digital ecosystems and data spaces, where the governing question is what makes data usable and by whom |
| Tijmen Wisman | Seized communications and metadata are unstructured personal data, and proportionality depends on how much can be inferred from them |
12. What is unexplored, and five projects
13. Your CV, mapped onto this chapter
| What you have | Where it lands | Why it fits |
|---|---|---|
| The offence classifier — TF-IDF, GloVe and WordNet over 358 sections of the Bhartiya Nyaya Sanhita, mapping plain language to statutory offences and returning punishment, cognizability, bailability and jurisdiction | Projects 1 and 5 | A document-classification and retrieval system over statute, built and shipped. It is the closest thing on any candidate's CV to what France's retrieval system does |
| The hybrid summarisation and entity extraction paper — regex, POS tagging, word embeddings | Projects 3 and 5 | The hybrid instinct is the correct one: rules where surface form is reliable, learning where meaning is open. That judgment is what the work needs |
| Multilingual production systems across six Indian languages | Project 4 | Cross-language consistency is where pan-EU tooling breaks first, and you have operated it rather than read about it |
| Rubric-based evaluation at Outlier across law, programming and linguistics | Projects 1 and 3 | Professional practice at applying a consistent codebook and recording where documents fail. That is literally the validation standard |
| The Mens Rea harness — judged labels, majority voting over three calls, agreement recorded, and a discarded keyword-counting attempt | Projects 1, 2 and 3 | You reached the right evaluation design empirically and documented the approach you abandoned. That is the credible part |
| RAG pipelines in production, including guardrails and edge cases found in live logs | Project 2 | Chunking failures are something you have debugged, not theorised |
| Legal training in statutory interpretation and evidence | Projects 1 and 4 | Open texture, defined terms and cross-reference are legal problems before they are NLP problems, and most NLP people cannot see them |
14. If you remember ten things
- Most of what an agency holds is unstructured, so the data problem is largely a text problem.
- Structured data validates against a schema; unstructured validation becomes interpretation, and interpretation is contestable.
- Stop word removal deletes not, case folding merges defined terms with ordinary ones, and stemming collapses distinctions. All three are defaults and all three can invert a legal sentence.
- TF-IDF discovers uninformative words from the corpus, which is why it needs no stop list. It remains the baseline to beat.
- Static embeddings give one vector per word; contextual embeddings resolve polysemy at the cost of transparency.
- ROUGE and BLEU are blind to negation and party roles, so a summary can score well while reversing the holding.
- Inter-annotator agreement measures whether the task is well defined. If humans cannot agree, no model score is interpretable.
- Validating a text pipeline means a codebook, an independent gold set, per-class metrics with the averaging stated, stability under reruns, and a hand audit of errors.
- Legal text breaks standard assumptions nine distinct ways, of which open texture, nested negation and cross-reference are the worst.
- Chunking can sever a did not from its clause and silently invert a holding, with nothing in the output disclosing it — the same failure as poisoning, with no adversary required.