Your Technical Stack
Every library, model and method named on your CV, defined in one line each, with what you actually used it for and the one follow-up question each invites. Built so that nothing on your own CV can freeze you.
0. How to use this chapter
Your CV names roughly a hundred technologies. A panel will not test them, but one of them may come up, and the danger is not ignorance — it is the blank pause. This chapter exists so that every item has a sentence attached.
1. The honest tiering of your own CV
Read this section first and be clear with yourself about which tier each claim sits in, because the panel's follow-up question will land somewhere and you need to know in advance which ones you can take two questions deep.
| Tier | What it means | Items |
|---|---|---|
| Deep — can go three questions down | You built it, debugged it, and know why it failed | RAG pipelines, LLM behavioural evaluation, adversarial prompting and RAG poisoning, prompt engineering, TF-IDF, cosine similarity, GloVe, WordNet, NLTK, First-Order Logic for legal elements, QLoRA fine-tuning of TinyLlama, Hugging Face Transformers, FAISS, SentenceTransformers, Next.js, FastAPI, Python generally |
| Real but one project deep | You used it once, successfully. Fine to claim, say *on one project* | BART, GPT-4o mini as a data-generation tool, gTTS and MoviePy, DoWhy and do-calculus, Bayesian Belief Networks, Bayesian Structural Time Series, Moran's I, hierarchical mixed-effects models, scikit-learn classifiers, ARIMA, LSTM, OpenCV, Docker, SWI-Prolog, R |
| Listed but thin — do not volunteer | Coursework or incidental. True, but will not survive probing, so do not lead with it | TensorFlow and Keras as frameworks you have architected in, SQLAlchemy, Alembic, Boto3, Razorpay, Framer Motion, PySwisseph, Pillow, caret and tidyverse specifics, SVM and KNN internals |
2. The numerical core — NumPy, Pandas, SciPy, statsmodels
| Library | What it is | The follow-up it invites |
|---|---|---|
| NumPy | Fast arrays and matrix maths. Everything else in Python's data stack is built on it. The key object is the ndarray, a typed n-dimensional block of memory | *Why is it fast?* Because loops run in compiled C over contiguous memory rather than in Python — that is what vectorisation means. A NumPy operation on a million rows is one C call, not a million Python steps |
| Pandas | Tables with labelled rows and columns — the DataFrame. Loading, joining, grouping, reshaping. Where most real analysis time goes | *What goes wrong?* Chained indexing producing a copy rather than a view, silent dtype promotion to object, and groupby dropping NaN keys by default. Also memory: a DataFrame is roughly several times the size of the CSV |
| SciPy | Scientific computing on top of NumPy — optimisation, interpolation, sparse matrices, and scipy.stats for distributions and hypothesis tests | *Which tests?* fisher_exact and the binomial test are the two in your own paper's territory. Sparse matrices matter too: a TF-IDF matrix is overwhelmingly zeros and must be stored sparse |
| statsmodels | Classical statistics with proper inference — regression with standard errors, p-values, confidence intervals, mixed-effects models, time series | *Why not scikit-learn?* Because scikit-learn gives you predictions and statsmodels gives you inference. If you need a coefficient's standard error and a p-value, you need statsmodels. The crimes-against-women work is in this register |
3. scikit-learn and the classical models on your CV
The library, then the six models you list. For each: the one-line mechanism and the one honest limitation.
| Model | How it works, in one line | Limitation to volunteer |
|---|---|---|
| Logistic regression | Fits a weighted sum of features, squashed through a sigmoid into a probability between 0 and 1 | Linear in the features, so it cannot capture interactions unless you build them by hand. But it is interpretable and reasonably calibrated, which is why it is the right default for anything a lawyer must read |
| Decision tree | Repeatedly splits the data on the single feature and threshold that best separates the classes, forming a flowchart | Very high variance — a slightly different sample gives a visibly different tree. Interpretable only while shallow |
| Random forest | Hundreds of trees, each on a bootstrap sample of rows and a random subset of features, then vote | Strong default on tabular data. The randomness is the point: decorrelating the trees is what makes averaging reduce variance. Not interpretable — feature importances are not explanations, and they are biased toward high-cardinality features |
| KNN | No training at all. To classify a point, find its k nearest neighbours in the training set and take the majority label | Collapses in high dimensions, because distance stops discriminating. Also stores the entire training set and is slow at prediction time |
| SVM | Finds the boundary with the widest margin between classes; the kernel trick lets it do this in a higher-dimensional space without ever computing the coordinates | Scales badly past tens of thousands of rows, needs features scaled, and produces no natural probability — the probabilities scikit-learn reports are a Platt-scaling fit bolted on afterwards |
| Gradient boosting — XGBoost, LightGBM | Trees added one at a time, each fitted to the errors the previous ones left behind | **Not on your CV, and it is the honest answer to *what would you actually use on tabular data*.** Say that: boosted trees usually beat both random forests and neural networks on rows-and-columns data |
| scikit-learn piece | What it does |
|---|---|
| `fit` / `predict` / `transform` | The uniform API. Every estimator exposes the same three verbs, which is the library's actual contribution |
| `Pipeline` | Chains preprocessing and model into one object. The correct defence against leakage — fit the scaler inside the pipeline so cross-validation refits it on each fold instead of leaking the full dataset's mean |
| `TfidfVectorizer` | Text to sparse TF-IDF matrix. What the offence classifier ran on |
| `cosine_similarity` | Pairwise similarity over rows. The retrieval step in that same classifier |
| `train_test_split`, `cross_val_score`, `StratifiedKFold` | Splitting. Stratified matters under class imbalance, which legal data always has |
| `GridSearchCV` | Exhaustive hyperparameter search with cross-validation. RandomizedSearchCV is usually the better choice once the grid is large |
| `classification_report`, `confusion_matrix` | Per-class precision, recall and F1. The chapter-7 formulas, implemented |
4. The deep learning stack — TensorFlow, Keras, PyTorch, Hugging Face
| Tool | What it is | Your use of it |
|---|---|---|
| TensorFlow | Google's deep learning framework. Builds a computation graph and runs automatic differentiation over it | On your CV; realistically coursework plus the Bi-LSTM offence classifier. Say coursework and the LSTM, not more |
| Keras | The high-level API that sits on TensorFlow. Sequential, model.add(...), compile, fit. Deliberately simple | The Bi-LSTM was almost certainly this: an Embedding layer, a Bidirectional(LSTM(...)), dense layers, softmax output, trained 50 epochs to 86.89 percent validation accuracy |
| PyTorch | Meta's framework and the research standard. Dynamic graphs, so the model is ordinary Python you can step through in a debugger | Underneath the thesis work — Hugging Face runs on it. You used it through Transformers rather than writing training loops, and that is the honest framing |
| Hugging Face `transformers` | Thousands of pretrained models behind one interface, plus tokenizers and a Trainer that handles the training loop | Core to the thesis: loading TinyLlama and BART, tokenising, and the fine-tuning run |
| Hugging Face `datasets` | Memory-mapped dataset handling with map for batched preprocessing. Handles data larger than RAM | Used to hold the 346,920-row training set and tokenise it in batches |
| `peft` | Parameter-Efficient Fine-Tuning. Implements LoRA and friends — freeze the base model, train small adapters | LoraConfig with r of 16, alpha 32, dropout 0.05, targeting the q_proj and v_proj attention projections, wrapped with get_peft_model |
| `bitsandbytes` | 4-bit and 8-bit quantisation. What the Q in QLoRA refers to | load_in_4bit=True on the TinyLlama load |
| `accelerate` | Device placement and distributed training. What device_map="auto" is doing under the hood | Installed and used implicitly rather than configured directly |
lora_alpha is a scaling factor applied to BA, and the ratio of alpha to r is what actually controls the update's strength, which is why 32 over 16 is a common choice.5. The NLP and retrieval stack
| Tool | What it is | The follow-up it invites |
|---|---|---|
| NLTK | The teaching toolkit for NLP. Tokenisers, POS tagger, stemmers, lemmatiser, and the WordNet interface | *Why NLTK over spaCy?* Honest answer: NLTK for WordNet access and for POS tagging in the FOL work; spaCy when speed and a production pipeline matter. NLTK is slower and more academic by design |
| spaCy | Production NLP. Fast tokenisation, POS, dependency parsing, named entity recognition, in one pipeline object | *What is a dependency parse?* A tree of grammatical relations — which word is the subject of which verb. Directly relevant to your FOL work, because *A sued B* and *B sued A* differ only in the dependency structure |
| WordNet | A hand-built lexical database: words grouped into synsets of synonyms, linked by hypernym and hyponym relations. Not a model — a human-curated graph | *Limitation?* No legal sense inventory. *Consideration* in contract law is not any of WordNet's senses of the word, so synonym expansion over legal text introduces noise as well as recall |
| GloVe | Pretrained static word vectors, learned from a global word-co-occurrence matrix. One fixed vector per word | *Why not BERT?* Because GloVe is a lookup table — free, instant, no GPU. And its limit is the thing to say first: one vector per word regardless of context, so *bail* has a single vector across every sense |
| SentenceTransformers | Produces one vector for a whole sentence or passage, tuned so that cosine similarity means semantic similarity. all-MiniLM-L6-v2 is the small fast default | *Why not average GloVe vectors?* Averaging destroys word order and composition. These models are trained directly on the similarity objective, which is why they work for retrieval |
| FAISS | Facebook's vector similarity search library. Indexes millions of embeddings for fast nearest-neighbour lookup | *Which index?* Your thesis uses IndexFlatL2 — exact brute-force search with L2 distance. Fine at thesis scale. At corpus scale you move to an approximate index like IVF or HNSW, trading a little recall for large speed gains |
| BART | An encoder-decoder transformer trained by corrupting text and reconstructing it. Strong at summarisation and rewriting | *Why BART and not GPT?* Because the task was transformation with a clear input and output — take stiff Wikipedia prose, return fluent prose. That is what an encoder-decoder is for |
6. The LLM and safety vocabulary on your CV
| Term as you list it | What you mean by it | Evidence you can point to |
|---|---|---|
| Retrieval-Augmented Generation | Fetch relevant documents, place them in the prompt, generate from them rather than from memory | Two production systems plus the thesis. And the poisoning experiment, which is the rarer half |
| Prompt engineering | Structuring instructions, roles, output formats and examples to make behaviour reliable and parseable | The forced-choice probe with four options and exactly one admission, scored by reading the returned letter |
| Adversarial prompting and red teaming | Deliberately attacking a system to surface failures before deployment | The poisoned condition — fabricated authority planted in retrieved context, 15 of 15 runs biased |
| LLM behavioural evaluation (RAG poisoning) | Measuring what a model does under controlled manipulation of its context, treating the model as a black box | This is the phrase that describes your paper most precisely, and it is the strongest line on your CV |
| AI safety and alignment | Making systems behave as intended. RLHF, DPO and guardrails are the techniques; the field is broader | Output guardrails on Astroformer, preventing deterministic advice. Be modest here — this is applied engineering, not alignment research |
| Fine-tuning LLMs | Continuing training on your own data to adapt behaviour | TinyLlama with QLoRA on 346,920 examples |
| LLM API integration | Calling hosted models and handling retries, rate limits, cost and non-determinism | Gemini and NVIDIA NIM on Astroformer, OpenAI SDK in the thesis pipeline, batch processing with ThreadPoolExecutor and exponential backoff |
| Knowledge representation, logic programming | Encoding facts and rules so they can be reasoned over rather than pattern-matched | The FOL offence work and SWI-Prolog, plus the VU logic course |
7. The causal and Bayesian stack
This is the crimes-against-women paper's machinery, and it is the most distinctive technical cluster you have, because almost nobody in legal AI works in this register.
| Method or tool | What it does | Why it was the right choice there |
|---|---|---|
| DAG — directed acyclic graph | A diagram of assumed causal structure. Arrows mean *causes*; the absence of an arrow is also a claim | The assumptions become explicit and criticisable. That is the methodological point, and it is the one a lawyer understands immediately |
| Pearl's do-calculus | Rules for working out whether a causal effect is identifiable from observational data, and which variables to condition on | Lets you ask *what would happen if urbanisation changed* rather than *what correlates with reporting* |
| DoWhy | Microsoft's Python causal inference library. Four stages: model the DAG, identify the estimand, estimate it, then refute it with sensitivity checks | The refutation stage is the selling point — it is built around trying to break your own result |
| Hierarchical / mixed-effects models | Models with both fixed effects shared across groups and random effects varying by group. Districts nested in states | Yielded the striking finding: state-level variation explains under 1 percent of the variance, so the action is local, not in state policy |
| Moran's I | A spatial autocorrelation statistic. Tests whether nearby areas have similar values, or whether the pattern is random | Returned no clustering — a negative result, and those are more credible than positive ones |
| Bayesian Structural Time Series | Decomposes a series into trend, seasonality and regression components, then estimates a counterfactual *what if the intervention had not happened* | Isolated the 2013 Criminal Law Amendment Act as a reporting shock of about 88,879 annual cases rather than a change in underlying behaviour |
| Bayesian Belief Network | A DAG plus conditional probability tables. Supports inference in any direction, handles missing values natively | Models nonlinear sociological interactions and gives a probability rather than a point prediction |
8. R and Prolog — the two that are differentiators
| Item | What it is | What to say |
|---|---|---|
| R | A language built for statistics rather than adapted to it. Regression, mixed models and survival analysis are first-class, and the inferential output is richer by default | *I use Python by default and R where the statistical tooling is better.* Name ggplot2 for grammar-of-graphics plotting, dplyr and tidyverse for data manipulation, caret for model training workflows, randomForest for forests. Do not overclaim depth here |
| SWI-Prolog | A logic programming language. You state facts and rules; the engine searches for values that satisfy a query. Declarative — you describe *what*, not *how* | This is a genuine differentiator. Almost no computational candidate has it, and it is the natural implementation language for the symbolic layer you want to argue for |
9. Visualisation and infrastructure, briefly
Low interview risk. One line each is sufficient; nobody is hiring a PhD on Docker.
| Tool | One line |
|---|---|
| Matplotlib | The foundational Python plotting library. Verbose, and everything else wraps it |
| Seaborn | Statistical plots over Matplotlib with far better defaults. Distributions, heatmaps, regression plots |
| Plotly | Interactive charts that render in a browser. Good for anything a reader should explore |
| Streamlit | Turns a Python script into a web app with no front-end code. The thesis interface. Ideal for research demos, not for production |
| FastAPI | Modern Python web framework. Async, and generates OpenAPI docs from type hints automatically |
| Pydantic | Runtime data validation from type annotations. What makes FastAPI's type checking real rather than decorative |
| Uvicorn | The ASGI server that actually runs FastAPI |
| Docker | Packages an application with its dependencies into a reproducible image. Worth one sentence on reproducibility if research infrastructure comes up |
| Next.js / React | The front-end framework behind both your platforms, and this site |
| Git | Version control. Relevant because provenance and reproducibility are themes you care about |
| SQL | Querying relational databases. JOIN, GROUP BY, window functions |
| Vercel, Render, Firebase | Hosting and deployment platforms |
10. The technical questions they would actually ask
11. If you remember ten things
- Three beats per tool: what it is, what you used it for, where it stops. The third beat is what makes it sound like experience.
- scikit-learn predicts; statsmodels explains. Enforcement questions are explanatory, so they need standard errors, not just accuracy.
- Boosted trees beat neural networks on tabular data. Say it unprompted — it signals judgment over enthusiasm.
- LoRA trains a low-rank update with the base model frozen; QLoRA loads that frozen base in 4-bit. r equals 16, alpha 32, about 0.8 percent of the parameters.
- GloVe gives one vector per word regardless of context; sentence-transformers embed whole passages against a similarity objective. That is why retrieval uses the second.
- FAISS `IndexFlatL2` on unnormalised vectors partly ranks by length. Normalise. This is your best *what would you fix* answer.
- Encoder reads both ways and cannot generate; decoder is masked and can. For coding a corpus, the encoder is the right and cheaper tool.
- Fine-tuning changes behaviour, retrieval changes knowledge — and retrieval wins in law because documents can be inspected and cited.
- The crimes-against-women paper is a measurement-versus-behaviour problem, and a cartel screen is the same problem. Detection is not occurrence. That is your strongest research transfer.
- Volunteer the two weaknesses: L2-versus-cosine, and a thesis evaluated against its own teacher model. And revoke the API key on page 37.