Skip to content
VibeFormer
54 min

Deep Learning

From a single neuron to transformers, with worked arithmetic. Includes GANs and graph neural networks, and why graphs are the architecture that actually matters for competition work.

Listen

0. What you need and what you do not

Nobody will ask you to derive backpropagation. What they might ask is whether a deep model is appropriate for a given enforcement task, and what its use costs in reviewability. So this chapter covers the mechanics well enough that you can reason about them, and spends more time on the two architectures that genuinely matter here: transformers, and graphs.

1. One neuron, with arithmetic

A neuron does three things: multiply each input by a weight, add them up with a bias, then pass the total through a non-linear function.

z=∑j=1pwjxj+b=w⋅x+ba=σ(z)z = \sum_{j=1}^{p} w_j x_j + b = \mathbf{w} \cdot \mathbf{x} + b \qquad\qquad a = \sigma(z)
The pre-activation z equals the sum of each input times its weight, plus a bias. Equivalently, the dot product of the weight vector and the input vector, plus b. The output a is an activation function applied to z.The bold letters mean *the whole list at once*, which is only notational shorthand. This is the entire neuron. The weights and the bias are learned; the activation function is chosen by a human. A network is this expression nested inside itself a few hundred thousand times.

One neuron, matching the worked example below

Three inputs, three weights, one bias, one squashing function, one number out. The negative weight on bidder count is the interesting part — the network learned that more bidders lowers the risk score, which is the direction a competition economist would predict.
ActivationWhat it doesUsed where
SigmoidSquashes to 0–1Final layer for binary classification. Saturates at the extremes, which kills gradients
TanhSquashes to −1 to 1Older recurrent networks. Same saturation problem
ReLUReturns the input if positive, otherwise zeroThe default hidden-layer choice. Cheap, and does not saturate on the positive side
GELUA smoother ReLUStandard inside transformers
SoftmaxTurns a vector of numbers into probabilities summing to 1Final layer for multiclass, and inside attention

2. From one neuron to a network

TermDefinition
LayerA group of neurons all reading the same inputs
Multilayer perceptronFully connected layers stacked. Every neuron reads every output of the layer below
Hidden layerAny layer that is not the input or output
Depth and widthNumber of layers, and neurons per layer
Forward passPushing input through to a prediction
Backward passComputing how each weight should change
EpochOne complete sweep through the training data
BatchThe handful of examples processed before each weight update

3. How training works

Four steps, repeated millions of times: predict, measure the error, work out which direction each weight should move, move it a little.

ConceptPlain meaning
GradientFor each weight, which direction and how steeply the loss would change if you nudged it
BackpropagationAn efficient way to compute every gradient in one backward sweep, by applying the chain rule from the output back toward the input. It is not a learning method — it only computes gradients
Gradient descentThe actual learning: move each weight a small step against its gradient
Learning rateHow big that step is. The most important hyperparameter in deep learning
Stochastic gradient descentUsing a small batch rather than the whole dataset per step. Noisier, far faster, and the noise itself helps escape poor solutions
MomentumKeep some of the previous step's direction, so the path smooths out
AdamAdapts the step size per weight using running averages. The default optimiser
Learning rate scheduleLowering the rate over training, so you take big steps early and settle precisely later
wj  ←  wj−η ∂L∂wjw_j \;\leftarrow\; w_j - \eta \, \frac{\partial L}{\partial w_j}
Each weight is replaced by itself, minus the learning rate eta times the partial derivative of the loss with respect to that weight.This one line is all of learning. The arrow means *becomes*. The curly-d term is the gradient: how much the loss would rise if this weight rose. Subtracting it moves the weight the other way, downhill. Eta is the learning rate. Everything else — Adam, momentum, schedules — is a refinement of how far and in what direction to step.
∂L∂w(1)=∂L∂a(3)⋅∂a(3)∂a(2)⋅∂a(2)∂w(1)\frac{\partial L}{\partial w^{(1)}} = \frac{\partial L}{\partial a^{(3)}} \cdot \frac{\partial a^{(3)}}{\partial a^{(2)}} \cdot \frac{\partial a^{(2)}}{\partial w^{(1)}}
The gradient of the loss with respect to a first-layer weight is the product of the gradients at each layer between them, multiplied together.This is backpropagation, and it is just the chain rule from calculus. The multiplication explains both failure modes in the table below: if each factor is a bit under 1, a long product shrinks toward zero — vanishing gradients. If each is a bit over 1, it explodes. Residual connections work by giving this product a path where one of the factors is exactly 1.
ProblemWhat happensStandard fix
Vanishing gradientsIn a deep stack, gradients shrink toward zero on the way back, so early layers barely learnReLU instead of sigmoid, residual connections, careful initialisation
Exploding gradientsGradients grow instead, and weights divergeGradient clipping, lower learning rate
Bad initialisationStart all weights at zero and every neuron in a layer computes the same thing foreverRandom initialisation scaled to layer size — Xavier or He
Internal covariate shiftEach layer's input distribution drifts as the layer below learnsBatch normalisation, or layer normalisation in transformers
Dead ReLUsA neuron stuck outputting zero for every input learns nothing againLeaky ReLU, lower learning rate

4. Regularisation in deep networks

MethodWhat it does
DropoutRandomly switch off a fraction of neurons on each training step, so no single pathway can be relied on. Switched off at prediction time
Weight decayThe L2 penalty from the machine learning chapter, applied to network weights
Early stoppingHalt when held-out performance stops improving
Data augmentationTrain on label-preserving variations of the input
Label smoothingTrain toward 0.9 rather than 1.0, which reduces overconfidence

5. Convolutional networks

Built for data with local spatial structure — images above all, where nearby pixels are related and the same pattern can appear anywhere.

TermDefinition
ConvolutionSlide a small window of weights, a filter, across the input, computing a weighted sum at each position
Filter, kernelThe small weight patch. One filter detects one kind of local pattern
Feature mapThe output of applying one filter everywhere
Stride and paddingHow far the window moves each step, and whether the edges are extended
PoolingShrink a feature map by taking the maximum or average of each small region
Receptive fieldHow much of the original input one deep unit ultimately sees
Parameter sharingThe same filter weights are reused at every position — far fewer parameters than a fully connected layer
Hout=⌊Hin+2p−ks⌋+1H_{\text{out}} = \left\lfloor \frac{H_{\text{in}} + 2p - k}{s} \right\rfloor + 1
The output size equals the input size, plus twice the padding, minus the kernel size, all divided by the stride, rounded down, plus one.The half-brackets mean *round down*. A 200-pixel input with a 3-by-3 kernel, no padding and stride 1 gives 198 — you lose one pixel at each edge because the window cannot hang off the end. Set padding to 1 and you get 200 back, which is why padding exists. Set stride to 2 and you get 99: stride is how a network downsamples. Worth being able to compute, because a mismatch here is the most common reason a vision model refuses to run.

6. Recurrent networks, and why they were replaced

Built for sequences. A recurrent network reads one element at a time, carrying a hidden state that summarises everything so far.

TermDefinition
Hidden stateA fixed-size vector carrying the memory of everything read so far
Backpropagation through timeUnrolling the sequence and backpropagating through every step
LSTMLong short-term memory. Adds gates — learned switches deciding what to forget, what to store, and what to output — so information can survive many steps
GRUA simpler gated design with fewer parameters and similar performance
BidirectionalTwo passes, forward and backward, concatenated. Only possible when you have the whole sequence up front

7. Attention and transformers

Chapter 9 covers this from the language side. Here is the mechanism itself, because it is the architecture behind almost everything current.

TermPlain meaning
Query, key, valueEach token produces three vectors. The query is what this token is looking for; the key is what each token offers; the value is what it contributes if selected
Scaled dot-product attentionMatch every query against every key to get relevance scores, divide by a scaling factor to keep them stable, softmax them into weights, then take the weighted sum of the values
Self-attentionQueries, keys and values all come from the same sequence
Multi-headSeveral attention operations in parallel, each free to track a different kind of relationship
Causal maskPrevents a token attending to anything after it. What makes generation possible
Positional encodingInjects position, because attention alone is order-blind

8. Generative models, including GANs

Models that produce new data resembling their training distribution, rather than labelling existing data.

ModelHow it works
AutoencoderAn encoder compresses input to a small code, a decoder rebuilds it. Trained to reproduce its own input, so the code must capture what matters
Variational autoencoderThe code is a distribution rather than a point, so you can sample it to generate new examples
GANTwo networks in competition. See below
Diffusion modelLearn to reverse a gradual noising process: start from noise and denoise step by step into a sample. Behind current image generation, and the design of one of your own subject models
Normalising flowsInvertible transformations from a simple distribution to a complex one, with exact likelihoods

The GAN training loop

The generator never sees the real data. Its only signal is the discriminator's reaction, passed back as a gradient. That is why the equilibrium is informative — the generator can only succeed by matching the real distribution well enough to be indistinguishable — and also why training is unstable, since both networks are chasing a target that moves whenever the other improves.
min⁡Gmax⁡D  Ex∼pdata[log⁡D(x)]  +  Ez∼pz[log⁡(1−D(G(z)))]\min_{G}\max_{D} \; \mathbb{E}_{x \sim p_{\text{data}}}\big[\log D(x)\big] \; + \; \mathbb{E}_{z \sim p_{z}}\big[\log\big(1 - D(G(z))\big)\big]
The discriminator maximises, and the generator minimises, the same quantity: the expected log probability the discriminator assigns to real data, plus the expected log probability it assigns to fakes being fake.A minimax objective — one expression, two players, opposite directions. The squiggle means *drawn from*. Only the second term involves G, and G wants D of G of z to be large, so it wants that term small. There is no single loss being minimised, which is exactly why a GAN has no clean convergence criterion and you cannot tell it is done by watching a loss curve fall.
GAN failureWhat happens
Mode collapseThe generator finds a few outputs that reliably fool the discriminator and produces only those. Technically successful, practically useless — it has learned a corner of the distribution, not the distribution
Training instabilityThe two networks oscillate instead of converging, because each is chasing a moving target
Vanishing discriminator gradientIf the discriminator becomes too good too fast, the generator receives no useful signal. Wasserstein GANs address this with a different objective
No likelihoodA GAN cannot tell you how probable a given sample is, unlike a VAE or a flow. It can only generate

9. Graph neural networks — the architecture that actually fits

Most data in competition work is relational. Who owns whom, who bid against whom, who sat on whose board, which decision cites which. Those are graphs, and a graph is the one data shape where the structure carries the signal.

TermDefinition
Node, vertexAn entity. A firm, a person, a decision, a tender
EdgeA relationship. *owns*, *bid against*, *cites*, *shares a director with*
Node featuresAttributes of each node — turnover, sector, country
DegreeHow many edges a node has
NeighbourhoodThe nodes one step away
Directed, weightedEdges may have direction (*A owns B*) and strength (*47 shared tenders*)
HomophilyWhether connected nodes tend to be similar. The assumption most graph networks rely on
hv(k+1)=ϕ(hv(k)⏟itself,    ⨁u∈N(v)ψ(hu(k))⏟its neighbours, aggregated)h_v^{(k+1)} = \phi\Big( \underbrace{h_v^{(k)}}_{\text{itself}}, \;\; \underbrace{\textstyle\bigoplus_{u \in \mathcal{N}(v)} \psi\big(h_u^{(k)}\big)}_{\text{its neighbours, aggregated}} \Big)
The new representation of node v is a function of two things: its own current representation, and an order-independent aggregation of a transformation of each of its neighbours' representations.h-v is node v's vector, and the superscript is the round number. The circled plus is any order-independent aggregation — sum, mean, or max — and that constraint is the whole point: a node has no first neighbour, so the operation must give the same answer regardless of order. Swap in mean and you have a GCN; swap in attention weights and you have a GAT. Every graph network is this one line with different choices for the two functions and the aggregator.

Two rounds of message passing on a co-bidding graph

Each round extends how far a node can see by one hop. The signal of a rotation scheme is not in any firm's attributes and not in any single tender — it is the closure of the group, which only appears at round two. The depth limit is a real constraint, not a tuning preference: over-smoothing means more rounds actively destroy the pattern.
VariantHow it aggregates
GCNAverages neighbour representations, normalised by degree. The simplest version
GraphSAGESamples a fixed number of neighbours rather than using all, so it scales to large graphs
GATUses attention to weight neighbours, so the model learns which connections matter rather than treating them equally
Message passing networksThe general framework all of the above are instances of
ProblemWhat it is
Over-smoothingAfter too many rounds every node's representation converges toward the same thing, because everyone has aggregated everyone. Usually two or three rounds is the limit
ScalabilityReal graphs are huge and neighbourhoods explode. Hence sampling approaches
Homophily assumptionThese methods assume neighbours are similar. Sometimes the signal is the opposite — a firm connected to very *unlike* firms may be the interesting one
ExplainabilityHarder than for tabular models. The answer depends on a subgraph, so the explanation has to be a subgraph

10. The limits worth stating out loud

  • Data hunger. Deep models need far more examples than tree ensembles. Confirmed cartels number in the dozens, not the millions.
  • Opacity. Millions of parameters with no individually meaningful interpretation. The interpretability methods from the machine learning chapter are approximations, and unstable ones.
  • Adversarial examples. Tiny, deliberately chosen input changes can flip a confident prediction. A firm that understood the screen could in principle structure a bid to avoid it.
  • Distribution shift. Performance decays as the world moves away from the training data, and legal and market conditions move constantly.
  • Overconfidence. Reliably miscalibrated without correction, which is the precondition problem from section 4.
  • Cost and environmental footprint. Training and serving are expensive, which matters for an agency procurement decision and is a fair question to be asked about.

11. How this connects to ATLANTIS

StrandWhat deep learning contributes, and costs
Accuracy, the data problemTransformers are what make unstructured evidence usable at scale — seized documents, decisions, correspondence. That is the capability driving the data appetite the framework was not built for
FairnessThe opacity is the fairness problem's technical core. Article 296 wants retraceable reasons; this is the architecture least able to supply them, and the one most likely to be deployed
Institutional arrangementsAgencies themselves name explainability of deployed tools and human–machine interaction among their three biggest difficulties. Both are consequences of this chapter

12. What the panel brings to this chapter

Panel memberConnection
Thibault SchrepelComplexity science is about non-linear dynamics, emergence and feedback — and a deep network is a non-linear system whose capabilities appear discontinuously with scale. His *emergence defeats prediction* argument is partly a claim about this architecture. He also co-authored with Pentland at MIT on foundation models, and his knowledge graph is a graph learning problem waiting to happen
Catalina GoantaMeasurement over unstructured content at scale, which is what these architectures are for
Georgiana MirzaDigital ecosystems — and an ecosystem is a graph, so section 9 is the formal treatment of her object of study
Tijmen WismanSyRI failed on transparency and verifiability. An opaque architecture is the hard case for that holding, and the question of whether a deep model can ever satisfy it is unresolved

13. What is unexplored, and four projects

14. Your CV, mapped onto this chapter

What you haveWhere it landsWhy it fits
Deep learning foundations — ANNs, CNNs, RNNs, TensorFlow and KerasAll four projectsEnough to build, and more importantly enough to judge somebody else's architecture choice
LSTM and time-series work — ARIMA, LSTMProject 1Honest framing: the previous generation for text, still reasonable for short numeric series. Do not oversell it for language
Computer vision — OpenCV, image classification, PillowDark patterns, which is Poland's toolInterface classification is a vision task with a legal standard attached
Fine-tuning open-weight transformers — TinyLlama, Hugging Face, Gemma, Llama 3Projects 3 and 4You have worked inside these models rather than only calling them, which is the difference between discussing and auditing
Model evaluation discipline and calibration awarenessProjects 2 and 4Overconfidence is the property that breaks the legal argument, and you already report agreement and intervals
Production serving at 28,000 monthly users, including inference latencyAny adoption questionCost and latency are real procurement constraints for an agency, and you have met them

15. If you remember ten things

  1. A neuron is a weighted sum plus a bias through a non-linearity, and the non-linearity is what makes depth mean anything.
  2. Backpropagation computes gradients; gradient descent does the learning. They are different things.
  3. The learning rate is the hyperparameter that most often decides whether training works at all.
  4. Residual connections are why networks could go deep, by giving gradients a short path back.
  5. CNNs win through parameter sharing — 320 weights where a dense layer needed 40 million — because the inductive bias is built in.
  6. Recurrent networks squeeze history through one fixed state; attention looks directly at any position, which is why transformers replaced them for text.
  7. Attention cost grows with the square of sequence length, which is why context is limited and why RAG exists.
  8. A GAN is a forger and an inspector improving together, and mode collapse is its characteristic failure — which matters because it would make synthetic audit data miss the rare cases.
  9. A graph network passes messages between neighbours, and two rounds reveal cluster structure invisible in any single row. Three of the field's real tools are graph problems.
  10. On tabular data boosted trees usually win. Deep learning earns its place on text, images and graphs — and knowing where not to use it is the stronger signal.