Skip to content
VibeFormer
44 min

REVISION 1 — Your Technical Stack

Everything technical on your CV, plus the whole Astroformer system, in one place. What each thing is in plain words, why you chose it, what you would say if asked, and the honest line between what you have built and what you have only touched.

Listen

0. The one rule for this whole chapter

TierWhat it meansWhat is in it
Tier 1 — built it, broke it, fixed itYou can go three questions deep. Lead with theseRAG pipelines, LLM behavioural evaluation, adversarial prompting and RAG poisoning, prompt engineering, TF-IDF and cosine similarity, GloVe, WordNet, NLTK, first-order logic for legal elements, QLoRA fine-tuning of TinyLlama, Hugging Face Transformers, FAISS, SentenceTransformers, Next.js, FastAPI, Python generally
Tier 2 — working knowledgeYou have used it, can discuss it, would check the detailsPandas, NumPy, scikit-learn, matplotlib, statsmodels, DoWhy causal inference, BSTS counterfactual time series, mixed-effects models, Bayesian networks, Postgres, Docker, Streamlit
Tier 3 — listed, honestly shallowName the tier immediately if askedTensorFlow and Keras (the Bi-LSTM offence classifier was coursework), R, Prolog, deep framework internals

1. Your CV, two levels deep

TermKindSAY THISIF PUSHED
PythonLanguage*The language almost all data work is written in. I use it for everything.*Interpreted and dynamically typed. It is dominant not because of the language itself but because its numerical libraries are a thin wrapper over compiled C, so you get easy code at near-native speed
NumPyLibrary*It lets Python do arithmetic on huge lists of numbers quickly. Almost everything else I used sits on top of it.*Provides the ndarray — a fixed-type, contiguous-memory array with vectorised operations running in compiled code, plus broadcasting for operating across mismatched shapes
PandasLibrary*It holds data in a table and lets me filter, group and join it with code instead of a mouse. I used it at every stage of the thesis and both papers.*DataFrame and Series over NumPy arrays. The distinguishing features are automatic alignment on index labels, split-apply-combine via groupby, relational joins, and built-in missing-data and time-series handling
scikit-learnLibrary*A box of ready-made machine-learning models that all work the same way, so you can swap one for another. I used it for the eight models in DiagnoChat.*Uniform fit/predict/transform API, which is what makes pipelines and grid search composable. It deliberately gives you no p-values or standard errors — its object is prediction, not inference
statsmodelsLibrary*The one you use when you care about whether a result is statistically significant rather than just predictive. I used it for the regressions in my crime paper.*Regression, generalised linear and mixed-effects models with standard errors, confidence intervals, hypothesis tests and residual diagnostics. Use it when the coefficient is the result; use scikit-learn when the prediction is
TF-IDFMethod*A way of scoring words so that words which are common in this document but rare everywhere else score highest — so it finds the distinctive words rather than *the* and *of*.*Term frequency times inverse document frequency, idf typically log(N / document frequency). Produces sparse bag-of-words vectors, so word order is discarded and synonyms are unrelated
Cosine similarityMethod*A way of measuring how similar two pieces of text are, which ignores how long they are — so a short note and a long judgment about the same thing still count as similar.*The inner product divided by the product of the norms — the cosine of the angle. On length-normalised vectors, cosine ranking and Euclidean ranking are identical, which is exactly why unnormalised distance search is *not* cosine search
GloVeModel / resource*A ready-made set of number-lists, one per word, built so that words used in similar ways have similar numbers. Its weakness is that each word gets only one, so* consideration *is the same in contract law as in ordinary speech.*Learns vectors by factorising a weighted log co-occurrence matrix over the whole corpus, rather than from local prediction windows as word2vec does. Static: one vector per type, no contextual disambiguation, and no representation at all for unseen words
WordNetResource*A dictionary built by hand that records how words relate to each other — which ones mean the same thing, and which ones are kinds of other things. I used it so that someone writing* broke in *could still reach the right offence.***Its unit is the *synset* — a set of synonyms standing for one concept — not the word. Synsets are linked by explicit relations: hypernymy and hyponymy (is-a), meronymy (part-of), antonymy, verb entailment. That makes it a graph, which is what supports the standard similarity measures. Being hand-built, it has high precision and thin coverage of domain terms** — which is where statutory vocabulary lives
NLTKLibrary*A basic toolkit for text: splitting it into words and sentences, labelling which word is a noun or a verb, cutting words back to their root.*Tokenisation, POS tagging against the Penn Treebank tagset, Porter and Snowball stemmers, WordNet lemmatiser, chunking, bundled corpora. Built for teaching, so it exposes components separately and runs slowly
spaCyLibrary*The faster, more industrial version of the same kind of text toolkit.*A single-pass pipeline — tokeniser, tagger, dependency parser, entity recogniser — in Cython. The practical difference from NLTK is that spaCy gives you a parse tree
Hugging Face TransformersLibrary*The place you download pretrained language models from, and the code that runs them. It is plumbing, not a model — worth saying, because people confuse the two.*Resolves a checkpoint name to an architecture, weights and the matching tokeniser via AutoModel / AutoTokenizer, and supplies the training loop and generation utilities
SentenceTransformersLibrary*It turns a whole sentence into one list of numbers representing its meaning, so you can compare meanings numerically rather than comparing words.*Fine-tunes an encoder with siamese or triplet objectives plus pooling so cosine distances become meaningful. It exists because raw encoder outputs are poor at sentence similarity. Mine was all-MiniLM-L6-v2 — six layers, 384 dimensions, 256 word-piece input limit, which caused my index bug
FAISSLibrary*A filing cabinet for those number-lists that can find the closest ones very fast. A search index for meaning instead of keywords.*Exact indexes (IndexFlatL2, IndexFlatIP) and approximate ones (IVF, HNSW, product quantisation) trading recall for speed and memory. `Flat` means exhaustive — exact, and linear in corpus size
QLoRAMethod*A cheap way to retrain a big model: shrink it to use less memory, freeze it, and train a small add-on alongside it instead of changing the whole thing.*4-bit quantisation of the frozen base plus trainable low-rank adapters, with double quantisation and paged optimiser states. Trains roughly one percent of the parameters a full update would touch, which is what brings it inside one GPU
First-order logicFormalism*A way of writing statements so precisely that ambiguity is impossible — it has* for all *and* there exists*. Writing a rule out this way forces you to choose one reading of a sentence that English leaves vague.*Predicates over terms with quantification over individuals — stronger than propositional logic, which cannot express relations; weaker than higher-order, which quantifies over predicates. Validity is semi-decidable
DoWhyLibrary*You draw a diagram of what you think causes what, and it tells you which variables you need to control for — then lets you test whether the answer survives.*Four stages: model the assumptions as a graph, identify an estimand via do-calculus, estimate, then refute with placebo and sensitivity tests. Its real contribution is forcing identification and estimation apart
BSTSMethod*It predicts what a line on a graph would have done if an event had never happened, then measures the gap between that and what actually happened. That gap is how I got the 88,879 figure.*A state-space model with trend, seasonality and covariate components, fitted by MCMC. The identifying assumption is that the covariates were themselves unaffected by the intervention
Mixed-effects modelMethod*A regression that knows my districts sit inside states, so it does not treat 700 districts in one state as 700 unrelated observations.*Fixed effects estimated as population parameters, random effects as draws from a distribution, producing partial pooling — small groups shrink toward the global mean. The intra-class correlation is between-group variance over total variance, hence my 0.6 percent
Bayesian networkFormalism*A diagram of variables with probabilities on the arrows, which lets you ask: given this and this, what is the chance of that?*A directed acyclic graph with conditional probability tables factorising a joint distribution. Independencies are readable off the graph via d-separation; inference is exact or sampled
FastAPIFramework*The Python tool I used to build the back end — the part that receives a request from the website and answers it.*ASGI, with Pydantic deriving request validation, serialisation and an OpenAPI schema from type hints. Async, so the process is not blocked waiting on an upstream model API
React / Next.jsLibrary / Framework*The tools for building the part of a website people actually see and click.*React maintains a virtual representation of the interface and reconciles changes. Next.js adds routing, server-side rendering and static generation, and server-side API routes
DockerTool*It packs an application with everything it needs into one image so it runs identically anywhere. Reproducibility, but for deployment.*Containers share the host kernel rather than booting a guest OS, so they start instantly. Images build in layers from a declarative Dockerfile. I needed it because `pyswisseph` compiles a native C extension
StreamlitFramework*It turns a plain Python script into a simple web app with almost no extra work. I used it for the thesis demo.*Re-executes the entire script top to bottom on every interaction, with widgets returning values and decorators caching expensive work. That model is why it is trivial for prototypes and wrong for anything stateful
RLanguage*A language built for statistics from the start rather than adapted to it. Honestly tier three for me — but worth mentioning, because empirical legal scholars often use it.*Vectorised by default, model formulas as a language-level construct, and an ecosystem where diagnostics and visual grammar are first-class
PrologLanguage*A language where you write down facts and rules, and it works out the conclusions itself — you describe what is true rather than how to compute it.*Horn clauses resolved by unification and backtracking search. Two properties matter legally: the closed-world assumption, so anything unproven is treated as false, and negation-as-failure, which is not classical negation — both are the wrong defaults for a burden of proof, and saying so is the interesting observation
Also on your CVKindSAY THIS
PyTorchFramework*The deep-learning framework I actually use, and the one Hugging Face runs on.*
TensorFlow / KerasFramework*Coursework only — the Bi-LSTM offence classifier. I would not claim depth.*
LSTMArchitecture*An older kind of network for sequences, with gates that let it remember things for longer than the plain version could.*
ARIMAMethod*A classical forecasting method for a single series over time.*
Cross-validation / GridSearchCVMethod*Hold back part of the data to check the model on data it has not seen; grid search repeats that across many settings. And note DiagnoChat tuned only one of eight models, so the comparison was not fair.*
SQLAlchemy / AlembicLibrary*Lets me work with database rows as Python objects, and keeps a versioned history of schema changes.*
PydanticLibrary*Checks that incoming data is the right shape before my code touches it.*
boto3Library*The Amazon storage client. I used it against Cloudflare's storage, which speaks the same protocol but charges nothing to download.*
OpenCV / PillowLibrary*Image handling. I used the slimmed-down build to keep the container small.*
JWT / bcryptStandard / Algorithm*A signed token that keeps someone logged in without the server storing a session, and a deliberately slow way of hashing passwords. My one-time codes are hashed too, which most side projects skip.*
pyswissephLibrary*Computes real astronomical positions. It is the deterministic factual layer in Astroformer — the part the model is not allowed to invent.*
Matplotlib / Seaborn / PlotlyLibrary*Charts: the basic one, the statistical one, and the interactive one.*
Git / GitHubTool*Version control. Worth mentioning because a reproducibility claim means nothing without it.*
SQLLanguage*The language for querying relational databases.*
TypeScriptLanguage*JavaScript with type checking, so mistakes are caught before it runs.*

2. The one technical thing to be able to explain properly — QLoRA

If any single technical item gets probed, it will be this, because it is the most advanced thing you have personally done. So it gets its own section, built up from nothing.

The letterWhat it stands forWhat it does, plainly
QQuantisedStore the frozen original model at 4-bit precision instead of 16-bit. Roughly a quarter of the memory, with modest loss of fidelity. This is why it needs an NVIDIA GPU — the 4-bit arithmetic requires CUDA and does not run on Apple Silicon
LoLow-RankInstead of one huge grid of new numbers, use two thin grids whose product has the same shape. Far fewer numbers for nearly the same effect
RAdaptationThe original model is frozen. Only the thin grids learn. So you are adapting rather than retraining
Wnew=W0⏟frozen+BA⏟learnedwhere A∈Rr×d,  B∈Rd×r,  r≪dW_{\text{new}} = \underbrace{W_0}_{\text{frozen}} + \underbrace{BA}_{\text{learned}} \qquad\text{where } A \in \mathbb{R}^{r \times d},\; B \in \mathbb{R}^{d \times r},\; r \ll d
The effective weights are the original frozen weights plus the product of two thin learned matrices, where the shared inner dimension r is much smaller than the model's width.The arithmetic that makes the point. If the model's width d is 4096 and you set r to 16, the full grid would hold about 16.8 million numbers. The two thin grids together hold about 131,000 — roughly 0.8 percent. `lora_alpha` is a scaling factor applied to that product, and the ratio of alpha to r controls how strongly the adapter speaks, which is why 32 over 16 is a common choice. Your own settings were `r = 16`, `alpha = 32`, dropout 0.05, applied to the query and value projections.

3. The thesis system, as a stack

Ten stages, four external models, four APIs

Only one of the ten stages is model training. The other nine are data engineering and systems work, which is the honest description of what data science actually is — and the right thing to emphasise, because it is the skill this position needs.
ResultValueWhat it means
Perplexity9.72How surprised the model was, on average, at each word — roughly as uncertain as choosing between ten options. A respectable fluency figure for a small model. But it measures fluency, not correctness — say that when you quote it
BLEU0.0032 in the text, 0.032 in the tableWord-overlap with a reference. Very low — and there is a tenfold inconsistency between two places in your own document
ROUGE-1 / 2 / L F10.121 / 0.0227 / 0.1076Other overlap measures. All low
Dataset5,782 × 12 × 5 = 346,920The real contribution of the thesis

4. Astroformer — the system you actually shipped

This is the only production system you have built with payments, authentication, storage and real users. It is worth being precise about, because *I shipped a thing that takes money and does not fall over* is a different claim from *I ran a notebook*.

The whole system, by layer

Two hosts, no GPUs, and a deliberate split between a deterministic factual layer and a generative presentation layer. That split is the part worth talking about in a law interview; the astrology is not.
LayerWhat it isThe decision behind it
FrontendVercel, Next.js 16 / React 19 / Tailwind 4A CDN for static and cached content. Security headers set at the edge, so they apply before any application code runs and cannot be forgotten in a route
BackendRender, Docker container, FastAPI + UvicornDocker rather than a buildpack because `pyswisseph` compiles a native C extension. A buildpack is a guess; a Dockerfile is reproducible — the same argument as research reproducibility
GPUNoneSay this plainly; it is correct, not a gap. Serving a 70-billion-parameter model yourself means renting a GPU by the hour whether anyone visits or not; an API is priced per token so cost tracks usage. The trade you accepted: your data leaves your infrastructure — which is exactly why an agency with confidential business information would choose the opposite
Model providersGemini, NVIDIA's hosted endpoint, Groq. Which one is live is set by environment variable, not in codeProvider redundancy, and two NVIDIA keys for rotation. Free tiers have hard ceilings, so an application-level guard holds requests to 14 per minute
DatabasePostgres via psycopg3 in production, SQLite locally. SQLAlchemy 2.0, Alembic. Nine modelsThe detail that shows you debug rather than copy: Render hands out a postgres:// URL and SQLAlchemy 2 with psycopg3 requires postgresql+psycopg://, so the session module rewrites the scheme at startup. Chat sessions are persisted, not held in memory, so a conversation survives a container restart
StorageCloudflare R2 via boto3 with S3 signature v4, bucket astroformer-assets. PDFs server-side and client-sideR2 speaks the S3 API so the same client works, but charges no egress fees — and a purchased report gets downloaded repeatedly. A pure cost decision, priced between two options
AuthJWT, HS256, seven-day expiry. bcrypt for passwords. Hashed OTPsHashing the one-time codes is the detail worth mentioning — most hobby projects store them in plain text
ServicesRazorpay and PayPal, Notion API as the back office, SendGrid, Firebase token verification, OpenCV headlessNotion as a back office is the judgment call to lead with: a real admin dashboard is weeks of work and non-technical order review was needed immediately. You bought time with an off-the-shelf interface
Chat endpoints/astro-chat, /love-astro-chat, /tarot/premium-chatThree separate tools sharing one credit pool

5. LawReformer, and the one live risk on it

AspectWhat it is
What it doesSix tools covering India, the UK and the US. **Rule-based, and the site says so — *not a chatbot*.** Disclaimers present
The sister siteai.lawreformer.com, which points users to free legal aid through the District Legal Services Authority
Why you built itCharging aggrieved people for legal information felt wrong. That is the honest motive and it is a good one
Why it is rule-based, and this is the strong point*A generative model that invents a procedural deadline for someone with no lawyer does active harm. A rule-based tool can only output what I wrote, which is checkable and wrong in predictable ways rather than unpredictable ones.* That is a deliberate safety decision, not a technical limitation, and you should present it as one
What you learned from itVery few users — and that is the informative part. It is free, so cost was not the barrier. *People do not look for legal information until they are already in trouble, and by then they want someone to act rather than a tool to read. So access to justice is not primarily an information-supply problem, and I had assumed it was.* A finding about your own assumption

6. Outlier, and the two framings for it

7. The technical questions, with answers

QuestionThe answer
*How is Astroformer deployed? What is it hosted on?*Vercel for the Next.js frontend, Render for the Python API as a Docker container. Docker rather than a buildpack because the astrology engine compiles a native C extension. No GPUs — every model call goes to a hosted provider, chosen by environment variable. Postgres via psycopg3 in production, SQLite locally, Cloudflare R2 for files, JWT sessions. Then volunteer the free plan, the open CORS and the default secret
*What would you build first for an agency?*A corpus and an entity layer, and it is deliberately unglamorous. *Before any of this is measurable you need the decisions in one place with provenance per document, and firms resolved to undertakings with a reported matching threshold. That is most of the first year and everything else depends on it*
*Walk me through a retrieval system over Commission decisions.**Collect with provenance per document. Chunk on structure — article, paragraph, recital — not a fixed token count, because splitting a holding from its reasoning destroys the passage. Embed, index, normalise the vectors so cosine and distance agree, retrieve more than you need and rerank. Keep a keyword baseline, because on legal text with precise terminology it is hard to beat and sometimes wins. Verify citations with a parser rather than a model, and log per query what was retrieved — that last part matters most, because without it you cannot reconstruct why the system said what it said*
*You have no GPUs and no infrastructure. Is that a problem?*Answer it straight. *For this project I do not think so. The work is document-scale rather than training-scale — classification, retrieval, graph analysis over tens of thousands of documents, which runs on a laptop or a modest server. If we needed to train a model from scratch I would need compute I have never had. But most of what I would propose does not*
*What is the most technically difficult thing you have done?*The thesis fine-tune, honestly stated. *Getting a 4-bit QLoRA run working on a rented GPU against a 346,920-row dataset, and then diagnosing why the retrieval half underperformed — the index embedded whole articles against a 256-word-piece limit, so each article was represented by roughly its opening paragraph. I had reported the symptom in the thesis without knowing the cause*
*What would you need to learn?*Three things, and name them in this order. *EU competition law doctrine properly, which I said in my letter and have started on. Institutional economics, which is the project's third named pillar and which I do not have — applied econometrics is not the same thing. And formal annotation methodology — agreement statistics, codebook design, adjudication protocols — which is standard in empirical legal studies, is not on my CV, and is a few days of reading*
*Do you code every day?**Yes. Two live sites and two papers' worth of analysis in the past year.* Do not elaborate. This question is checking for a hobbyist, and the answer is a fact rather than an argument