Skip to content
VibeFormer
44 min

NLP and Unstructured Data

Preprocessing through embeddings, classification and sequence models, the evaluation metrics, and the part that matters most for this project: how you validate unstructured data, and why legal text is the hardest case there is.

Listen

0. Why this chapter is load-bearing

Almost everything a competition agency holds is unstructured. Decisions, Statements of Objections, seized emails, meeting notes, internal memoranda, leniency applications, tender documents. The structured parts — bid values, dates, firm identifiers — are the minority, and they are the part everyone already knows how to handle.

1. Structured, semi-structured, unstructured

KindWhat it isAgency example
StructuredFixed schema, typed fields, one meaning per columnA tender award table: contract reference, buyer, winner, value, date, CPV code
Semi-structuredSelf-describing but irregular. Tags or keys exist, presence and nesting varyAn Akoma Ntoso or XML judgment with marked-up articles and recitals; a JSON export of case metadata
UnstructuredNo schema. Meaning is carried entirely by the contentThe body of a decision, a seized email thread, a handwritten meeting note, a scanned PDF annex

2. Preprocessing, and where it silently destroys meaning

  • Sentence and word segmentation. Splitting text into sentences and tokens. Harder than it looks: abbreviations, citations and decimals all contain full stops.
  • Case folding. Lowercasing everything. Cheap, and it conflates distinctions that matter.
  • Stop word removal. Dropping high-frequency words like *the*, *of*, *not*.
  • Stemming. Chopping suffixes by rule. Fast, crude, produces non-words — *organisation* becomes *organis*.
  • Lemmatisation. Mapping to the dictionary form using morphology and part of speech. Slower, correct — *was* becomes *be*.
  • Normalisation. Unifying quotes, dashes, whitespace, Unicode forms, and digits.
  • OCR. Turning scanned pages into text, with an error rate that is never zero.

3. Representing text as numbers

RepresentationHow it worksStrength and limit
Bag of wordsCount each vocabulary term, discard orderTrivially interpretable. Loses word order entirely, so *A sued B* and *B sued A* are identical
TF-IDFTerm frequency times inverse document frequency, so a word scores highly when it is frequent here and rare across the corpusStill the right default for document retrieval and a strong classification baseline. Sparse, fast, and inspectable — you can print the top-weighted terms and read them
n-gramsCount short sequences rather than single wordsRecovers some local order. Vocabulary explodes
Word2Vec and GloVeDense vectors learned from co-occurrence, so similar words sit near each otherCaptures semantic similarity. One vector per word regardless of context, so the bank of a river and a commercial bank share a vector
Contextual embeddingsA transformer produces a different vector for a word in each contextResolves polysemy. Opaque, and expensive relative to TF-IDF
Sentence and document embeddingsA single vector for a whole passageThe basis of semantic search and of retrieval in RAG
tf-idf(t,d)=ft,d∑t′ft′,d⏟term frequency  ×  log⁡Nnt⏟inverse document frequency\text{tf-idf}(t, d) = \underbrace{\frac{f_{t,d}}{\textstyle\sum_{t'} f_{t',d}}}_{\text{term frequency}} \;\times\; \underbrace{\log\frac{N}{n_t}}_{\text{inverse document frequency}}
The TF-IDF weight of a term in a document is how often that term appears in the document, divided by the document's total length, multiplied by the log of the total number of documents divided by the number of documents containing that term.N is the corpus size and little-n-t is how many documents contain the term. Note what happens when a term appears in every document: N over n-t is 1, the log of 1 is 0, and the weight vanishes. That is the mechanism that silently deletes *the* without anyone writing a stop-word list.

Once documents are vectors, *similar* needs a definition. The standard one is cosine similarity: the angle between two vectors, ignoring their lengths.

cos⁡(θ)=a⋅b∥a∥ ∥b∥=∑iaibi∑iai2  ∑ibi2\cos(\theta) = \frac{\mathbf{a} \cdot \mathbf{b}}{\lVert \mathbf{a} \rVert \, \lVert \mathbf{b} \rVert} = \frac{\sum_{i} a_i b_i}{\sqrt{\sum_i a_i^2}\;\sqrt{\sum_i b_i^2}}
Cosine similarity is the dot product of two vectors divided by the product of their lengths. Equivalently: multiply the vectors element by element and add up, then divide by the square root of each vector's sum of squares.Runs from 1, meaning identical direction, through 0, meaning unrelated, to minus 1, meaning opposite. Dividing by the lengths is the important part: it makes a two-page decision and a two-hundred-page one comparable, because only the mix of words matters, not how many there are. This is the similarity function underneath almost every vector database.

BM25 is the retrieval function that actually wins in practice, and it is TF-IDF with two corrections: term frequency saturates, and document length is normalised explicitly.

BM25(q,d)=∑t∈qidf(t)⋅ft,d⋅(k1+1)ft,d+k1(1−b+b∣d∣avgdl)\text{BM25}(q, d) = \sum_{t \in q} \text{idf}(t) \cdot \frac{f_{t,d} \cdot (k_1 + 1)}{f_{t,d} + k_1\left(1 - b + b\dfrac{|d|}{\text{avgdl}}\right)}
BM25 sums, over each query term, the inverse document frequency times a saturating function of how often the term occurs in the document, adjusted by the document's length relative to the average length.k-one and b are tuning constants, usually about 1.2 and 0.75. The fraction is the whole idea: as term frequency grows, the expression flattens towards k-one plus 1 rather than growing without limit. **The twentieth mention of *cartel* adds almost nothing over the tenth** — which is right, and which raw TF-IDF gets wrong. This is the baseline a dense retriever has to beat, and frequently does not.

4. The core tasks

TaskWhat it produces
Part-of-speech taggingA grammatical label per token: noun, verb, modal, determiner
Named entity recognitionSpans labelled as person, organisation, location, date, money, and in legal work: court, statute, case citation, party role
Dependency parsingThe grammatical relations in a sentence — who did what to whom
Coreference resolutionLinking *it*, *the undertaking*, *the applicant* back to the entity meant
Relation extractionTyped links between entities: Acme subcontracted-to Beta
Text classificationA label for a document or passage
Sequence labellingA label per token, which is how NER is actually implemented
SummarisationExtractive, selecting original sentences, or abstractive, generating new text
Question answeringExtractive from a passage, or generative

5. Models, from classifiers to sequence to sequence

  • Linear classifiers over TF-IDF. Logistic regression or a linear SVM. Still the baseline to beat, interpretable, and often within a few points of anything fancier on well-defined document classification.
  • Naive Bayes. Fast, works oddly well on text, badly calibrated because the independence assumption is false.
  • Recurrent models. An LSTM or GRU reads a sequence left to right, carrying a hidden state. Handles order; struggles with long dependencies and cannot be parallelised across the sequence.
  • Convolutional text models. Filters over word windows, good at local pattern detection such as key phrases.
  • Sequence to sequence. An encoder compresses the input, a decoder generates the output. Translation, summarisation, and question answering all fit this frame.
  • Attention, then transformers. The decoder looks back at all encoder positions rather than relying on one compressed state. Removing recurrence entirely gives the transformer — which is chapter 9.
  • Encoder-only transformers. BERT-style, bidirectional, fine-tuned for classification and extraction. Still the right tool for labelling a corpus, and far cheaper than a generative model.

6. Evaluation metrics, and what each hides

MetricMeasuresLimitation
AccuracyFraction of correct labelsUseless under class imbalance, which legal corpora always have
Precision, recall, F1Per class, then averagedMacro average treats every class equally; micro is dominated by frequent classes. Which you report changes the story, so state it
Exact matchWhole output identical to referenceBrutal and appropriate for extracting a citation or a date
Token-level F1Overlap between predicted and reference spansStandard for extractive question answering and NER
BLEUn-gram overlap with a reference, precision-orientedBuilt for translation. Rewards surface similarity and punishes correct paraphrase
ROUGEn-gram and longest-common-subsequence overlap, recall-orientedStandard for summarisation, and only weakly related to whether the summary is true
BERTScoreSimilarity in embedding space rather than surface formCredits paraphrase. Opaque, and still says nothing about factual correctness
PerplexityHow surprised a language model is by held-out textA fluency measure. Not a correctness measure
Inter-annotator agreementCohen's or Fleiss' kappa, Krippendorff's alphaMeasures whether the task is well defined. If humans cannot agree, no model score means anything
F1macro=1K∑k=1KF1(k)F1micro=∑kTPk∑kTPk+12∑k(FPk+FNk)F_1^{\text{macro}} = \frac{1}{K}\sum_{k=1}^{K} F_1^{(k)} \qquad\qquad F_1^{\text{micro}} = \frac{\sum_k TP_k}{\sum_k TP_k + \tfrac{1}{2}\sum_k (FP_k + FN_k)}
Macro F1 computes F1 separately for each class and then takes a plain average. Micro F1 instead pools all the true positives, false positives and false negatives across classes before computing a single F1.This single choice can move a reported score by thirty points. Macro gives a class with 4 examples the same weight as a class with 4,000, so rare-offence performance is visible. Micro is dominated by the frequent classes, so a model that ignores every rare class entirely can still look strong. For a 358-section statute, macro is the honest number and micro is the flattering one — state which you used.
κ=po−pe1−pe\kappa = \frac{p_o - p_e}{1 - p_e}
Cohen's kappa equals the observed agreement minus the agreement expected by chance, divided by one minus the chance agreement.Zero means your annotators did no better than guessing; one means perfect agreement. Why subtract chance at all: if 95 percent of tenders are clean, two annotators who both label everything clean agree 95 percent of the time while contributing no information. Kappa correctly scores that at about zero. Your three-call majority vote with recorded agreement is this idea applied to model judges.

7. Validating unstructured data

This is the section you asked for, and it is the one with the least written about it. Structured validation is a solved engineering problem. Unstructured validation is mostly unsolved, and in law it is barely posed.

CheckStructured dataUnstructured data
Schema conformanceTypes, ranges, required fields, foreign keys. Deterministic pass or failNo schema exists. The nearest equivalent is whether an extraction pipeline produced the expected fields, which validates the pipeline rather than the data
CompletenessCount nulls per columnUnknowable. You cannot count the clauses a document failed to contain, or the emails never produced
ConsistencyCross-field rules — award date after publication dateRequires extracting the facts first, so every inconsistency check inherits the extraction error rate
UniquenessPrimary keys and deduplicationNear-duplicates everywhere: drafts, translations, annexes, forwarded email chains. Exact matching misses almost all of it
AccuracyCompare against a register or a source systemCompare against what? Usually against human annotation, which is itself fallible and which is where inter-annotator agreement becomes the real measurement
ProvenanceRow-level lineageDocument-level at best. Which paragraph a claim came from is itself an extraction task

9. Vocabulary you will hear

TermWhat it means
CorpusThe body of text you are working on. Plural corpora
TokenThe unit a model actually processes. Roughly a word-piece, not a word
VocabularyThe set of tokens a model knows
Out of vocabularyA term the model has no token for, historically a major failure source
Gold standard, ground truthHuman-annotated labels treated as correct for evaluation
Annotation, labellingThe human process of producing those labels
Inter-annotator agreementHow often independent humans assign the same label. Kappa or alpha
CodebookThe written rules annotators follow. If it is not written down, the labels are not reproducible
AdjudicationResolving disagreements between annotators, and recording how
Distant supervisionGenerating noisy labels automatically from an existing resource rather than by hand
Zero-shot and few-shotPerforming a task with no, or very few, labelled examples
Transfer learningTaking a model trained on one task and adapting it to another
Domain shiftThe deployment data differs from the training data. Legal corpora shift constantly as law changes
Embedding spaceThe geometry in which semantic similarity becomes distance
Cosine similarityThe standard distance measure between embeddings
Sparse and dense retrievalTerm-matching such as BM25 against embedding-based semantic search
BM25The strong classical ranking function. Still competitive, and the baseline any dense method should be compared against
Hybrid searchCombining sparse and dense retrieval, which usually beats either alone

10. How this connects to ATLANTIS

StrandThe NLP content of it
Accuracy, the data problemMost of what an agency collects is text. Whether its inferences are sound depends on extraction accuracy, chunking integrity, and whether the corpus is complete — and completeness is the one property unstructured data cannot report about itself
FairnessArticle 296 asks whether a decision explains itself. That is a property of text, and it is measurable: specificity, template reuse, whether the stated ground accounts for the outcome. The DSA statement-of-reasons corpus is the instrument
Institutional arrangementsTwenty-four equally authentic languages and 27 national authorities. Cross-language consistency is not a nicety; it determines whether a tool deployed in one Member State means the same thing in another

11. What the panel brings to this chapter

Panel memberWhere NLP meets their work
Thibault SchrepelBuilt a knowledge graph of every Commission competition decision from 1977 to 2025, with decisions as nodes and doctrinal links as edges. Extracting those links from decision text is relation extraction over legal prose — this chapter is how that artefact was made. He also offers to help agencies build queryable systems from their own decisional records, which is the same pipeline
Catalina GoantaThe strongest fit on the panel. Her multi-country longitudinal study of influencer disclosure across around a million posts is classification and measurement over unstructured multilingual text at scale, with every validation problem in section 7. Her legal compliance API proposal is about making compliance machine-checkable, which is the same instinct
Georgiana MirzaDigital ecosystems and data spaces, where the governing question is what makes data usable and by whom
Tijmen WismanSeized communications and metadata are unstructured personal data, and proportionality depends on how much can be inferred from them

12. What is unexplored, and five projects

13. Your CV, mapped onto this chapter

What you haveWhere it landsWhy it fits
The offence classifier — TF-IDF, GloVe and WordNet over 358 sections of the Bhartiya Nyaya Sanhita, mapping plain language to statutory offences and returning punishment, cognizability, bailability and jurisdictionProjects 1 and 5A document-classification and retrieval system over statute, built and shipped. It is the closest thing on any candidate's CV to what France's retrieval system does
The hybrid summarisation and entity extraction paper — regex, POS tagging, word embeddingsProjects 3 and 5The hybrid instinct is the correct one: rules where surface form is reliable, learning where meaning is open. That judgment is what the work needs
Multilingual production systems across six Indian languagesProject 4Cross-language consistency is where pan-EU tooling breaks first, and you have operated it rather than read about it
Rubric-based evaluation at Outlier across law, programming and linguisticsProjects 1 and 3Professional practice at applying a consistent codebook and recording where documents fail. That is literally the validation standard
The Mens Rea harness — judged labels, majority voting over three calls, agreement recorded, and a discarded keyword-counting attemptProjects 1, 2 and 3You reached the right evaluation design empirically and documented the approach you abandoned. That is the credible part
RAG pipelines in production, including guardrails and edge cases found in live logsProject 2Chunking failures are something you have debugged, not theorised
Legal training in statutory interpretation and evidenceProjects 1 and 4Open texture, defined terms and cross-reference are legal problems before they are NLP problems, and most NLP people cannot see them

14. If you remember ten things

  1. Most of what an agency holds is unstructured, so the data problem is largely a text problem.
  2. Structured data validates against a schema; unstructured validation becomes interpretation, and interpretation is contestable.
  3. Stop word removal deletes not, case folding merges defined terms with ordinary ones, and stemming collapses distinctions. All three are defaults and all three can invert a legal sentence.
  4. TF-IDF discovers uninformative words from the corpus, which is why it needs no stop list. It remains the baseline to beat.
  5. Static embeddings give one vector per word; contextual embeddings resolve polysemy at the cost of transparency.
  6. ROUGE and BLEU are blind to negation and party roles, so a summary can score well while reversing the holding.
  7. Inter-annotator agreement measures whether the task is well defined. If humans cannot agree, no model score is interpretable.
  8. Validating a text pipeline means a codebook, an independent gold set, per-class metrics with the averaging stated, stability under reruns, and a hand audit of errors.
  9. Legal text breaks standard assumptions nine distinct ways, of which open texture, nested negation and cross-reference are the worst.
  10. Chunking can sever a did not from its clause and silently invert a holding, with nothing in the output disclosing it — the same failure as poisoning, with no adversary required.