Deep Learning
From a single neuron to transformers, with worked arithmetic. Includes GANs and graph neural networks, and why graphs are the architecture that actually matters for competition work.
0. What you need and what you do not
Nobody will ask you to derive backpropagation. What they might ask is whether a deep model is appropriate for a given enforcement task, and what its use costs in reviewability. So this chapter covers the mechanics well enough that you can reason about them, and spends more time on the two architectures that genuinely matter here: transformers, and graphs.
1. One neuron, with arithmetic
A neuron does three things: multiply each input by a weight, add them up with a bias, then pass the total through a non-linear function.
One neuron, matching the worked example below
| Activation | What it does | Used where |
|---|---|---|
| Sigmoid | Squashes to 0–1 | Final layer for binary classification. Saturates at the extremes, which kills gradients |
| Tanh | Squashes to −1 to 1 | Older recurrent networks. Same saturation problem |
| ReLU | Returns the input if positive, otherwise zero | The default hidden-layer choice. Cheap, and does not saturate on the positive side |
| GELU | A smoother ReLU | Standard inside transformers |
| Softmax | Turns a vector of numbers into probabilities summing to 1 | Final layer for multiclass, and inside attention |
2. From one neuron to a network
| Term | Definition |
|---|---|
| Layer | A group of neurons all reading the same inputs |
| Multilayer perceptron | Fully connected layers stacked. Every neuron reads every output of the layer below |
| Hidden layer | Any layer that is not the input or output |
| Depth and width | Number of layers, and neurons per layer |
| Forward pass | Pushing input through to a prediction |
| Backward pass | Computing how each weight should change |
| Epoch | One complete sweep through the training data |
| Batch | The handful of examples processed before each weight update |
3. How training works
Four steps, repeated millions of times: predict, measure the error, work out which direction each weight should move, move it a little.
| Concept | Plain meaning |
|---|---|
| Gradient | For each weight, which direction and how steeply the loss would change if you nudged it |
| Backpropagation | An efficient way to compute every gradient in one backward sweep, by applying the chain rule from the output back toward the input. It is not a learning method — it only computes gradients |
| Gradient descent | The actual learning: move each weight a small step against its gradient |
| Learning rate | How big that step is. The most important hyperparameter in deep learning |
| Stochastic gradient descent | Using a small batch rather than the whole dataset per step. Noisier, far faster, and the noise itself helps escape poor solutions |
| Momentum | Keep some of the previous step's direction, so the path smooths out |
| Adam | Adapts the step size per weight using running averages. The default optimiser |
| Learning rate schedule | Lowering the rate over training, so you take big steps early and settle precisely later |
| Problem | What happens | Standard fix |
|---|---|---|
| Vanishing gradients | In a deep stack, gradients shrink toward zero on the way back, so early layers barely learn | ReLU instead of sigmoid, residual connections, careful initialisation |
| Exploding gradients | Gradients grow instead, and weights diverge | Gradient clipping, lower learning rate |
| Bad initialisation | Start all weights at zero and every neuron in a layer computes the same thing forever | Random initialisation scaled to layer size — Xavier or He |
| Internal covariate shift | Each layer's input distribution drifts as the layer below learns | Batch normalisation, or layer normalisation in transformers |
| Dead ReLUs | A neuron stuck outputting zero for every input learns nothing again | Leaky ReLU, lower learning rate |
4. Regularisation in deep networks
| Method | What it does |
|---|---|
| Dropout | Randomly switch off a fraction of neurons on each training step, so no single pathway can be relied on. Switched off at prediction time |
| Weight decay | The L2 penalty from the machine learning chapter, applied to network weights |
| Early stopping | Halt when held-out performance stops improving |
| Data augmentation | Train on label-preserving variations of the input |
| Label smoothing | Train toward 0.9 rather than 1.0, which reduces overconfidence |
5. Convolutional networks
Built for data with local spatial structure — images above all, where nearby pixels are related and the same pattern can appear anywhere.
| Term | Definition |
|---|---|
| Convolution | Slide a small window of weights, a filter, across the input, computing a weighted sum at each position |
| Filter, kernel | The small weight patch. One filter detects one kind of local pattern |
| Feature map | The output of applying one filter everywhere |
| Stride and padding | How far the window moves each step, and whether the edges are extended |
| Pooling | Shrink a feature map by taking the maximum or average of each small region |
| Receptive field | How much of the original input one deep unit ultimately sees |
| Parameter sharing | The same filter weights are reused at every position — far fewer parameters than a fully connected layer |
6. Recurrent networks, and why they were replaced
Built for sequences. A recurrent network reads one element at a time, carrying a hidden state that summarises everything so far.
| Term | Definition |
|---|---|
| Hidden state | A fixed-size vector carrying the memory of everything read so far |
| Backpropagation through time | Unrolling the sequence and backpropagating through every step |
| LSTM | Long short-term memory. Adds gates — learned switches deciding what to forget, what to store, and what to output — so information can survive many steps |
| GRU | A simpler gated design with fewer parameters and similar performance |
| Bidirectional | Two passes, forward and backward, concatenated. Only possible when you have the whole sequence up front |
7. Attention and transformers
Chapter 9 covers this from the language side. Here is the mechanism itself, because it is the architecture behind almost everything current.
| Term | Plain meaning |
|---|---|
| Query, key, value | Each token produces three vectors. The query is what this token is looking for; the key is what each token offers; the value is what it contributes if selected |
| Scaled dot-product attention | Match every query against every key to get relevance scores, divide by a scaling factor to keep them stable, softmax them into weights, then take the weighted sum of the values |
| Self-attention | Queries, keys and values all come from the same sequence |
| Multi-head | Several attention operations in parallel, each free to track a different kind of relationship |
| Causal mask | Prevents a token attending to anything after it. What makes generation possible |
| Positional encoding | Injects position, because attention alone is order-blind |
8. Generative models, including GANs
Models that produce new data resembling their training distribution, rather than labelling existing data.
| Model | How it works |
|---|---|
| Autoencoder | An encoder compresses input to a small code, a decoder rebuilds it. Trained to reproduce its own input, so the code must capture what matters |
| Variational autoencoder | The code is a distribution rather than a point, so you can sample it to generate new examples |
| GAN | Two networks in competition. See below |
| Diffusion model | Learn to reverse a gradual noising process: start from noise and denoise step by step into a sample. Behind current image generation, and the design of one of your own subject models |
| Normalising flows | Invertible transformations from a simple distribution to a complex one, with exact likelihoods |
The GAN training loop
| GAN failure | What happens |
|---|---|
| Mode collapse | The generator finds a few outputs that reliably fool the discriminator and produces only those. Technically successful, practically useless — it has learned a corner of the distribution, not the distribution |
| Training instability | The two networks oscillate instead of converging, because each is chasing a moving target |
| Vanishing discriminator gradient | If the discriminator becomes too good too fast, the generator receives no useful signal. Wasserstein GANs address this with a different objective |
| No likelihood | A GAN cannot tell you how probable a given sample is, unlike a VAE or a flow. It can only generate |
9. Graph neural networks — the architecture that actually fits
Most data in competition work is relational. Who owns whom, who bid against whom, who sat on whose board, which decision cites which. Those are graphs, and a graph is the one data shape where the structure carries the signal.
| Term | Definition |
|---|---|
| Node, vertex | An entity. A firm, a person, a decision, a tender |
| Edge | A relationship. *owns*, *bid against*, *cites*, *shares a director with* |
| Node features | Attributes of each node — turnover, sector, country |
| Degree | How many edges a node has |
| Neighbourhood | The nodes one step away |
| Directed, weighted | Edges may have direction (*A owns B*) and strength (*47 shared tenders*) |
| Homophily | Whether connected nodes tend to be similar. The assumption most graph networks rely on |
Two rounds of message passing on a co-bidding graph
| Variant | How it aggregates |
|---|---|
| GCN | Averages neighbour representations, normalised by degree. The simplest version |
| GraphSAGE | Samples a fixed number of neighbours rather than using all, so it scales to large graphs |
| GAT | Uses attention to weight neighbours, so the model learns which connections matter rather than treating them equally |
| Message passing networks | The general framework all of the above are instances of |
| Problem | What it is |
|---|---|
| Over-smoothing | After too many rounds every node's representation converges toward the same thing, because everyone has aggregated everyone. Usually two or three rounds is the limit |
| Scalability | Real graphs are huge and neighbourhoods explode. Hence sampling approaches |
| Homophily assumption | These methods assume neighbours are similar. Sometimes the signal is the opposite — a firm connected to very *unlike* firms may be the interesting one |
| Explainability | Harder than for tabular models. The answer depends on a subgraph, so the explanation has to be a subgraph |
10. The limits worth stating out loud
- Data hunger. Deep models need far more examples than tree ensembles. Confirmed cartels number in the dozens, not the millions.
- Opacity. Millions of parameters with no individually meaningful interpretation. The interpretability methods from the machine learning chapter are approximations, and unstable ones.
- Adversarial examples. Tiny, deliberately chosen input changes can flip a confident prediction. A firm that understood the screen could in principle structure a bid to avoid it.
- Distribution shift. Performance decays as the world moves away from the training data, and legal and market conditions move constantly.
- Overconfidence. Reliably miscalibrated without correction, which is the precondition problem from section 4.
- Cost and environmental footprint. Training and serving are expensive, which matters for an agency procurement decision and is a fair question to be asked about.
11. How this connects to ATLANTIS
| Strand | What deep learning contributes, and costs |
|---|---|
| Accuracy, the data problem | Transformers are what make unstructured evidence usable at scale — seized documents, decisions, correspondence. That is the capability driving the data appetite the framework was not built for |
| Fairness | The opacity is the fairness problem's technical core. Article 296 wants retraceable reasons; this is the architecture least able to supply them, and the one most likely to be deployed |
| Institutional arrangements | Agencies themselves name explainability of deployed tools and human–machine interaction among their three biggest difficulties. Both are consequences of this chapter |
12. What the panel brings to this chapter
| Panel member | Connection |
|---|---|
| Thibault Schrepel | Complexity science is about non-linear dynamics, emergence and feedback — and a deep network is a non-linear system whose capabilities appear discontinuously with scale. His *emergence defeats prediction* argument is partly a claim about this architecture. He also co-authored with Pentland at MIT on foundation models, and his knowledge graph is a graph learning problem waiting to happen |
| Catalina Goanta | Measurement over unstructured content at scale, which is what these architectures are for |
| Georgiana Mirza | Digital ecosystems — and an ecosystem is a graph, so section 9 is the formal treatment of her object of study |
| Tijmen Wisman | SyRI failed on transparency and verifiability. An opaque architecture is the hard case for that holding, and the question of whether a deep model can ever satisfy it is unresolved |
13. What is unexplored, and four projects
14. Your CV, mapped onto this chapter
| What you have | Where it lands | Why it fits |
|---|---|---|
| Deep learning foundations — ANNs, CNNs, RNNs, TensorFlow and Keras | All four projects | Enough to build, and more importantly enough to judge somebody else's architecture choice |
| LSTM and time-series work — ARIMA, LSTM | Project 1 | Honest framing: the previous generation for text, still reasonable for short numeric series. Do not oversell it for language |
| Computer vision — OpenCV, image classification, Pillow | Dark patterns, which is Poland's tool | Interface classification is a vision task with a legal standard attached |
| Fine-tuning open-weight transformers — TinyLlama, Hugging Face, Gemma, Llama 3 | Projects 3 and 4 | You have worked inside these models rather than only calling them, which is the difference between discussing and auditing |
| Model evaluation discipline and calibration awareness | Projects 2 and 4 | Overconfidence is the property that breaks the legal argument, and you already report agreement and intervals |
| Production serving at 28,000 monthly users, including inference latency | Any adoption question | Cost and latency are real procurement constraints for an agency, and you have met them |
15. If you remember ten things
- A neuron is a weighted sum plus a bias through a non-linearity, and the non-linearity is what makes depth mean anything.
- Backpropagation computes gradients; gradient descent does the learning. They are different things.
- The learning rate is the hyperparameter that most often decides whether training works at all.
- Residual connections are why networks could go deep, by giving gradients a short path back.
- CNNs win through parameter sharing — 320 weights where a dense layer needed 40 million — because the inductive bias is built in.
- Recurrent networks squeeze history through one fixed state; attention looks directly at any position, which is why transformers replaced them for text.
- Attention cost grows with the square of sequence length, which is why context is limited and why RAG exists.
- A GAN is a forger and an inspector improving together, and mode collapse is its characteristic failure — which matters because it would make synthetic audit data miss the rare cases.
- A graph network passes messages between neighbours, and two rounds reveal cluster structure invisible in any single row. Three of the field's real tools are graph problems.
- On tabular data boosted trees usually win. Deep learning earns its place on text, images and graphs — and knowing where not to use it is the stronger signal.