Skip to content
VibeFormer
46 min

Your Technical Stack

Every library, model and method named on your CV, defined in one line each, with what you actually used it for and the one follow-up question each invites. Built so that nothing on your own CV can freeze you.

Listen

0. How to use this chapter

Your CV names roughly a hundred technologies. A panel will not test them, but one of them may come up, and the danger is not ignorance — it is the blank pause. This chapter exists so that every item has a sentence attached.

1. The honest tiering of your own CV

Read this section first and be clear with yourself about which tier each claim sits in, because the panel's follow-up question will land somewhere and you need to know in advance which ones you can take two questions deep.

TierWhat it meansItems
Deep — can go three questions downYou built it, debugged it, and know why it failedRAG pipelines, LLM behavioural evaluation, adversarial prompting and RAG poisoning, prompt engineering, TF-IDF, cosine similarity, GloVe, WordNet, NLTK, First-Order Logic for legal elements, QLoRA fine-tuning of TinyLlama, Hugging Face Transformers, FAISS, SentenceTransformers, Next.js, FastAPI, Python generally
Real but one project deepYou used it once, successfully. Fine to claim, say *on one project*BART, GPT-4o mini as a data-generation tool, gTTS and MoviePy, DoWhy and do-calculus, Bayesian Belief Networks, Bayesian Structural Time Series, Moran's I, hierarchical mixed-effects models, scikit-learn classifiers, ARIMA, LSTM, OpenCV, Docker, SWI-Prolog, R
Listed but thin — do not volunteerCoursework or incidental. True, but will not survive probing, so do not lead with itTensorFlow and Keras as frameworks you have architected in, SQLAlchemy, Alembic, Boto3, Razorpay, Framer Motion, PySwisseph, Pillow, caret and tidyverse specifics, SVM and KNN internals

2. The numerical core — NumPy, Pandas, SciPy, statsmodels

LibraryWhat it isThe follow-up it invites
NumPyFast arrays and matrix maths. Everything else in Python's data stack is built on it. The key object is the ndarray, a typed n-dimensional block of memory*Why is it fast?* Because loops run in compiled C over contiguous memory rather than in Python — that is what vectorisation means. A NumPy operation on a million rows is one C call, not a million Python steps
PandasTables with labelled rows and columns — the DataFrame. Loading, joining, grouping, reshaping. Where most real analysis time goes*What goes wrong?* Chained indexing producing a copy rather than a view, silent dtype promotion to object, and groupby dropping NaN keys by default. Also memory: a DataFrame is roughly several times the size of the CSV
SciPyScientific computing on top of NumPy — optimisation, interpolation, sparse matrices, and scipy.stats for distributions and hypothesis tests*Which tests?* fisher_exact and the binomial test are the two in your own paper's territory. Sparse matrices matter too: a TF-IDF matrix is overwhelmingly zeros and must be stored sparse
statsmodelsClassical statistics with proper inference — regression with standard errors, p-values, confidence intervals, mixed-effects models, time series*Why not scikit-learn?* Because scikit-learn gives you predictions and statsmodels gives you inference. If you need a coefficient's standard error and a p-value, you need statsmodels. The crimes-against-women work is in this register

3. scikit-learn and the classical models on your CV

The library, then the six models you list. For each: the one-line mechanism and the one honest limitation.

ModelHow it works, in one lineLimitation to volunteer
Logistic regressionFits a weighted sum of features, squashed through a sigmoid into a probability between 0 and 1Linear in the features, so it cannot capture interactions unless you build them by hand. But it is interpretable and reasonably calibrated, which is why it is the right default for anything a lawyer must read
Decision treeRepeatedly splits the data on the single feature and threshold that best separates the classes, forming a flowchartVery high variance — a slightly different sample gives a visibly different tree. Interpretable only while shallow
Random forestHundreds of trees, each on a bootstrap sample of rows and a random subset of features, then voteStrong default on tabular data. The randomness is the point: decorrelating the trees is what makes averaging reduce variance. Not interpretable — feature importances are not explanations, and they are biased toward high-cardinality features
KNNNo training at all. To classify a point, find its k nearest neighbours in the training set and take the majority labelCollapses in high dimensions, because distance stops discriminating. Also stores the entire training set and is slow at prediction time
SVMFinds the boundary with the widest margin between classes; the kernel trick lets it do this in a higher-dimensional space without ever computing the coordinatesScales badly past tens of thousands of rows, needs features scaled, and produces no natural probability — the probabilities scikit-learn reports are a Platt-scaling fit bolted on afterwards
Gradient boosting — XGBoost, LightGBMTrees added one at a time, each fitted to the errors the previous ones left behind**Not on your CV, and it is the honest answer to *what would you actually use on tabular data*.** Say that: boosted trees usually beat both random forests and neural networks on rows-and-columns data
scikit-learn pieceWhat it does
`fit` / `predict` / `transform`The uniform API. Every estimator exposes the same three verbs, which is the library's actual contribution
`Pipeline`Chains preprocessing and model into one object. The correct defence against leakage — fit the scaler inside the pipeline so cross-validation refits it on each fold instead of leaking the full dataset's mean
`TfidfVectorizer`Text to sparse TF-IDF matrix. What the offence classifier ran on
`cosine_similarity`Pairwise similarity over rows. The retrieval step in that same classifier
`train_test_split`, `cross_val_score`, `StratifiedKFold`Splitting. Stratified matters under class imbalance, which legal data always has
`GridSearchCV`Exhaustive hyperparameter search with cross-validation. RandomizedSearchCV is usually the better choice once the grid is large
`classification_report`, `confusion_matrix`Per-class precision, recall and F1. The chapter-7 formulas, implemented

4. The deep learning stack — TensorFlow, Keras, PyTorch, Hugging Face

ToolWhat it isYour use of it
TensorFlowGoogle's deep learning framework. Builds a computation graph and runs automatic differentiation over itOn your CV; realistically coursework plus the Bi-LSTM offence classifier. Say coursework and the LSTM, not more
KerasThe high-level API that sits on TensorFlow. Sequential, model.add(...), compile, fit. Deliberately simpleThe Bi-LSTM was almost certainly this: an Embedding layer, a Bidirectional(LSTM(...)), dense layers, softmax output, trained 50 epochs to 86.89 percent validation accuracy
PyTorchMeta's framework and the research standard. Dynamic graphs, so the model is ordinary Python you can step through in a debuggerUnderneath the thesis work — Hugging Face runs on it. You used it through Transformers rather than writing training loops, and that is the honest framing
Hugging Face `transformers`Thousands of pretrained models behind one interface, plus tokenizers and a Trainer that handles the training loopCore to the thesis: loading TinyLlama and BART, tokenising, and the fine-tuning run
Hugging Face `datasets`Memory-mapped dataset handling with map for batched preprocessing. Handles data larger than RAMUsed to hold the 346,920-row training set and tokenise it in batches
`peft`Parameter-Efficient Fine-Tuning. Implements LoRA and friends — freeze the base model, train small adaptersLoraConfig with r of 16, alpha 32, dropout 0.05, targeting the q_proj and v_proj attention projections, wrapped with get_peft_model
`bitsandbytes`4-bit and 8-bit quantisation. What the Q in QLoRA refers toload_in_4bit=True on the TinyLlama load
`accelerate`Device placement and distributed training. What device_map="auto" is doing under the hoodInstalled and used implicitly rather than configured directly
W′=W0+ΔW≈W0+BA⏟trainedB∈Rd×r,  A∈Rr×d,  r≪dW' = W_0 + \Delta W \approx W_0 + \underbrace{BA}_{\text{trained}} \qquad B \in \mathbb{R}^{d \times r},\; A \in \mathbb{R}^{r \times d},\; r \ll d
The adapted weights equal the frozen original weights plus a low-rank update, which is the product of two thin matrices B and A whose inner dimension r is much smaller than the model dimension d.W-nought is frozen; only A and B receive gradients. With d of 4096 and r of 16, the two thin matrices together hold about 131,000 numbers against the full matrix's 16.8 million — roughly 0.8 percent. lora_alpha is a scaling factor applied to BA, and the ratio of alpha to r is what actually controls the update's strength, which is why 32 over 16 is a common choice.

5. The NLP and retrieval stack

ToolWhat it isThe follow-up it invites
NLTKThe teaching toolkit for NLP. Tokenisers, POS tagger, stemmers, lemmatiser, and the WordNet interface*Why NLTK over spaCy?* Honest answer: NLTK for WordNet access and for POS tagging in the FOL work; spaCy when speed and a production pipeline matter. NLTK is slower and more academic by design
spaCyProduction NLP. Fast tokenisation, POS, dependency parsing, named entity recognition, in one pipeline object*What is a dependency parse?* A tree of grammatical relations — which word is the subject of which verb. Directly relevant to your FOL work, because *A sued B* and *B sued A* differ only in the dependency structure
WordNetA hand-built lexical database: words grouped into synsets of synonyms, linked by hypernym and hyponym relations. Not a model — a human-curated graph*Limitation?* No legal sense inventory. *Consideration* in contract law is not any of WordNet's senses of the word, so synonym expansion over legal text introduces noise as well as recall
GloVePretrained static word vectors, learned from a global word-co-occurrence matrix. One fixed vector per word*Why not BERT?* Because GloVe is a lookup table — free, instant, no GPU. And its limit is the thing to say first: one vector per word regardless of context, so *bail* has a single vector across every sense
SentenceTransformersProduces one vector for a whole sentence or passage, tuned so that cosine similarity means semantic similarity. all-MiniLM-L6-v2 is the small fast default*Why not average GloVe vectors?* Averaging destroys word order and composition. These models are trained directly on the similarity objective, which is why they work for retrieval
FAISSFacebook's vector similarity search library. Indexes millions of embeddings for fast nearest-neighbour lookup*Which index?* Your thesis uses IndexFlatL2 — exact brute-force search with L2 distance. Fine at thesis scale. At corpus scale you move to an approximate index like IVF or HNSW, trading a little recall for large speed gains
BARTAn encoder-decoder transformer trained by corrupting text and reconstructing it. Strong at summarisation and rewriting*Why BART and not GPT?* Because the task was transformation with a clear input and output — take stiff Wikipedia prose, return fluent prose. That is what an encoder-decoder is for

6. The LLM and safety vocabulary on your CV

Term as you list itWhat you mean by itEvidence you can point to
Retrieval-Augmented GenerationFetch relevant documents, place them in the prompt, generate from them rather than from memoryTwo production systems plus the thesis. And the poisoning experiment, which is the rarer half
Prompt engineeringStructuring instructions, roles, output formats and examples to make behaviour reliable and parseableThe forced-choice probe with four options and exactly one admission, scored by reading the returned letter
Adversarial prompting and red teamingDeliberately attacking a system to surface failures before deploymentThe poisoned condition — fabricated authority planted in retrieved context, 15 of 15 runs biased
LLM behavioural evaluation (RAG poisoning)Measuring what a model does under controlled manipulation of its context, treating the model as a black boxThis is the phrase that describes your paper most precisely, and it is the strongest line on your CV
AI safety and alignmentMaking systems behave as intended. RLHF, DPO and guardrails are the techniques; the field is broaderOutput guardrails on Astroformer, preventing deterministic advice. Be modest here — this is applied engineering, not alignment research
Fine-tuning LLMsContinuing training on your own data to adapt behaviourTinyLlama with QLoRA on 346,920 examples
LLM API integrationCalling hosted models and handling retries, rate limits, cost and non-determinismGemini and NVIDIA NIM on Astroformer, OpenAI SDK in the thesis pipeline, batch processing with ThreadPoolExecutor and exponential backoff
Knowledge representation, logic programmingEncoding facts and rules so they can be reasoned over rather than pattern-matchedThe FOL offence work and SWI-Prolog, plus the VU logic course

7. The causal and Bayesian stack

This is the crimes-against-women paper's machinery, and it is the most distinctive technical cluster you have, because almost nobody in legal AI works in this register.

Method or toolWhat it doesWhy it was the right choice there
DAG — directed acyclic graphA diagram of assumed causal structure. Arrows mean *causes*; the absence of an arrow is also a claimThe assumptions become explicit and criticisable. That is the methodological point, and it is the one a lawyer understands immediately
Pearl's do-calculusRules for working out whether a causal effect is identifiable from observational data, and which variables to condition onLets you ask *what would happen if urbanisation changed* rather than *what correlates with reporting*
DoWhyMicrosoft's Python causal inference library. Four stages: model the DAG, identify the estimand, estimate it, then refute it with sensitivity checksThe refutation stage is the selling point — it is built around trying to break your own result
Hierarchical / mixed-effects modelsModels with both fixed effects shared across groups and random effects varying by group. Districts nested in statesYielded the striking finding: state-level variation explains under 1 percent of the variance, so the action is local, not in state policy
Moran's IA spatial autocorrelation statistic. Tests whether nearby areas have similar values, or whether the pattern is randomReturned no clustering — a negative result, and those are more credible than positive ones
Bayesian Structural Time SeriesDecomposes a series into trend, seasonality and regression components, then estimates a counterfactual *what if the intervention had not happened*Isolated the 2013 Criminal Law Amendment Act as a reporting shock of about 88,879 annual cases rather than a change in underlying behaviour
Bayesian Belief NetworkA DAG plus conditional probability tables. Supports inference in any direction, handles missing values nativelyModels nonlinear sociological interactions and gives a probability rather than a point prediction

8. R and Prolog — the two that are differentiators

ItemWhat it isWhat to say
RA language built for statistics rather than adapted to it. Regression, mixed models and survival analysis are first-class, and the inferential output is richer by default*I use Python by default and R where the statistical tooling is better.* Name ggplot2 for grammar-of-graphics plotting, dplyr and tidyverse for data manipulation, caret for model training workflows, randomForest for forests. Do not overclaim depth here
SWI-PrologA logic programming language. You state facts and rules; the engine searches for values that satisfy a query. Declarative — you describe *what*, not *how*This is a genuine differentiator. Almost no computational candidate has it, and it is the natural implementation language for the symbolic layer you want to argue for

9. Visualisation and infrastructure, briefly

Low interview risk. One line each is sufficient; nobody is hiring a PhD on Docker.

ToolOne line
MatplotlibThe foundational Python plotting library. Verbose, and everything else wraps it
SeabornStatistical plots over Matplotlib with far better defaults. Distributions, heatmaps, regression plots
PlotlyInteractive charts that render in a browser. Good for anything a reader should explore
StreamlitTurns a Python script into a web app with no front-end code. The thesis interface. Ideal for research demos, not for production
FastAPIModern Python web framework. Async, and generates OpenAPI docs from type hints automatically
PydanticRuntime data validation from type annotations. What makes FastAPI's type checking real rather than decorative
UvicornThe ASGI server that actually runs FastAPI
DockerPackages an application with its dependencies into a reproducible image. Worth one sentence on reproducibility if research infrastructure comes up
Next.js / ReactThe front-end framework behind both your platforms, and this site
GitVersion control. Relevant because provenance and reproducibility are themes you care about
SQLQuerying relational databases. JOIN, GROUP BY, window functions
Vercel, Render, FirebaseHosting and deployment platforms

10. The technical questions they would actually ask

11. If you remember ten things

  1. Three beats per tool: what it is, what you used it for, where it stops. The third beat is what makes it sound like experience.
  2. scikit-learn predicts; statsmodels explains. Enforcement questions are explanatory, so they need standard errors, not just accuracy.
  3. Boosted trees beat neural networks on tabular data. Say it unprompted — it signals judgment over enthusiasm.
  4. LoRA trains a low-rank update with the base model frozen; QLoRA loads that frozen base in 4-bit. r equals 16, alpha 32, about 0.8 percent of the parameters.
  5. GloVe gives one vector per word regardless of context; sentence-transformers embed whole passages against a similarity objective. That is why retrieval uses the second.
  6. FAISS `IndexFlatL2` on unnormalised vectors partly ranks by length. Normalise. This is your best *what would you fix* answer.
  7. Encoder reads both ways and cannot generate; decoder is masked and can. For coding a corpus, the encoder is the right and cheaper tool.
  8. Fine-tuning changes behaviour, retrieval changes knowledge — and retrieval wins in law because documents can be inspected and cited.
  9. The crimes-against-women paper is a measurement-versus-behaviour problem, and a cartel screen is the same problem. Detection is not occurrence. That is your strongest research transfer.
  10. Volunteer the two weaknesses: L2-versus-cosine, and a thesis evaluated against its own teacher model. And revoke the API key on page 37.