REVISION 1 — Your Technical Stack
Everything technical on your CV, plus the whole Astroformer system, in one place. What each thing is in plain words, why you chose it, what you would say if asked, and the honest line between what you have built and what you have only touched.
Listen
0. The one rule for this whole chapter
| Tier | What it means | What is in it |
|---|---|---|
| Tier 1 — built it, broke it, fixed it | You can go three questions deep. Lead with these | RAG pipelines, LLM behavioural evaluation, adversarial prompting and RAG poisoning, prompt engineering, TF-IDF and cosine similarity, GloVe, WordNet, NLTK, first-order logic for legal elements, QLoRA fine-tuning of TinyLlama, Hugging Face Transformers, FAISS, SentenceTransformers, Next.js, FastAPI, Python generally |
| Tier 2 — working knowledge | You have used it, can discuss it, would check the details | Pandas, NumPy, scikit-learn, matplotlib, statsmodels, DoWhy causal inference, BSTS counterfactual time series, mixed-effects models, Bayesian networks, Postgres, Docker, Streamlit |
| Tier 3 — listed, honestly shallow | Name the tier immediately if asked | TensorFlow and Keras (the Bi-LSTM offence classifier was coursework), R, Prolog, deep framework internals |
1. Your CV, two levels deep
| Term | Kind | SAY THIS | IF PUSHED |
|---|---|---|---|
| Python | Language | *The language almost all data work is written in. I use it for everything.* | Interpreted and dynamically typed. It is dominant not because of the language itself but because its numerical libraries are a thin wrapper over compiled C, so you get easy code at near-native speed |
| NumPy | Library | *It lets Python do arithmetic on huge lists of numbers quickly. Almost everything else I used sits on top of it.* | Provides the ndarray — a fixed-type, contiguous-memory array with vectorised operations running in compiled code, plus broadcasting for operating across mismatched shapes |
| Pandas | Library | *It holds data in a table and lets me filter, group and join it with code instead of a mouse. I used it at every stage of the thesis and both papers.* | DataFrame and Series over NumPy arrays. The distinguishing features are automatic alignment on index labels, split-apply-combine via groupby, relational joins, and built-in missing-data and time-series handling |
| scikit-learn | Library | *A box of ready-made machine-learning models that all work the same way, so you can swap one for another. I used it for the eight models in DiagnoChat.* | Uniform fit/predict/transform API, which is what makes pipelines and grid search composable. It deliberately gives you no p-values or standard errors — its object is prediction, not inference |
| statsmodels | Library | *The one you use when you care about whether a result is statistically significant rather than just predictive. I used it for the regressions in my crime paper.* | Regression, generalised linear and mixed-effects models with standard errors, confidence intervals, hypothesis tests and residual diagnostics. Use it when the coefficient is the result; use scikit-learn when the prediction is |
| TF-IDF | Method | *A way of scoring words so that words which are common in this document but rare everywhere else score highest — so it finds the distinctive words rather than *the* and *of*.* | Term frequency times inverse document frequency, idf typically log(N / document frequency). Produces sparse bag-of-words vectors, so word order is discarded and synonyms are unrelated |
| Cosine similarity | Method | *A way of measuring how similar two pieces of text are, which ignores how long they are — so a short note and a long judgment about the same thing still count as similar.* | The inner product divided by the product of the norms — the cosine of the angle. On length-normalised vectors, cosine ranking and Euclidean ranking are identical, which is exactly why unnormalised distance search is *not* cosine search |
| GloVe | Model / resource | *A ready-made set of number-lists, one per word, built so that words used in similar ways have similar numbers. Its weakness is that each word gets only one, so* consideration *is the same in contract law as in ordinary speech.* | Learns vectors by factorising a weighted log co-occurrence matrix over the whole corpus, rather than from local prediction windows as word2vec does. Static: one vector per type, no contextual disambiguation, and no representation at all for unseen words |
| WordNet | Resource | *A dictionary built by hand that records how words relate to each other — which ones mean the same thing, and which ones are kinds of other things. I used it so that someone writing* broke in *could still reach the right offence.* | **Its unit is the *synset* — a set of synonyms standing for one concept — not the word. Synsets are linked by explicit relations: hypernymy and hyponymy (is-a), meronymy (part-of), antonymy, verb entailment. That makes it a graph, which is what supports the standard similarity measures. Being hand-built, it has high precision and thin coverage of domain terms** — which is where statutory vocabulary lives |
| NLTK | Library | *A basic toolkit for text: splitting it into words and sentences, labelling which word is a noun or a verb, cutting words back to their root.* | Tokenisation, POS tagging against the Penn Treebank tagset, Porter and Snowball stemmers, WordNet lemmatiser, chunking, bundled corpora. Built for teaching, so it exposes components separately and runs slowly |
| spaCy | Library | *The faster, more industrial version of the same kind of text toolkit.* | A single-pass pipeline — tokeniser, tagger, dependency parser, entity recogniser — in Cython. The practical difference from NLTK is that spaCy gives you a parse tree |
| Hugging Face Transformers | Library | *The place you download pretrained language models from, and the code that runs them. It is plumbing, not a model — worth saying, because people confuse the two.* | Resolves a checkpoint name to an architecture, weights and the matching tokeniser via AutoModel / AutoTokenizer, and supplies the training loop and generation utilities |
| SentenceTransformers | Library | *It turns a whole sentence into one list of numbers representing its meaning, so you can compare meanings numerically rather than comparing words.* | Fine-tunes an encoder with siamese or triplet objectives plus pooling so cosine distances become meaningful. It exists because raw encoder outputs are poor at sentence similarity. Mine was all-MiniLM-L6-v2 — six layers, 384 dimensions, 256 word-piece input limit, which caused my index bug |
| FAISS | Library | *A filing cabinet for those number-lists that can find the closest ones very fast. A search index for meaning instead of keywords.* | Exact indexes (IndexFlatL2, IndexFlatIP) and approximate ones (IVF, HNSW, product quantisation) trading recall for speed and memory. `Flat` means exhaustive — exact, and linear in corpus size |
| QLoRA | Method | *A cheap way to retrain a big model: shrink it to use less memory, freeze it, and train a small add-on alongside it instead of changing the whole thing.* | 4-bit quantisation of the frozen base plus trainable low-rank adapters, with double quantisation and paged optimiser states. Trains roughly one percent of the parameters a full update would touch, which is what brings it inside one GPU |
| First-order logic | Formalism | *A way of writing statements so precisely that ambiguity is impossible — it has* for all *and* there exists*. Writing a rule out this way forces you to choose one reading of a sentence that English leaves vague.* | Predicates over terms with quantification over individuals — stronger than propositional logic, which cannot express relations; weaker than higher-order, which quantifies over predicates. Validity is semi-decidable |
| DoWhy | Library | *You draw a diagram of what you think causes what, and it tells you which variables you need to control for — then lets you test whether the answer survives.* | Four stages: model the assumptions as a graph, identify an estimand via do-calculus, estimate, then refute with placebo and sensitivity tests. Its real contribution is forcing identification and estimation apart |
| BSTS | Method | *It predicts what a line on a graph would have done if an event had never happened, then measures the gap between that and what actually happened. That gap is how I got the 88,879 figure.* | A state-space model with trend, seasonality and covariate components, fitted by MCMC. The identifying assumption is that the covariates were themselves unaffected by the intervention |
| Mixed-effects model | Method | *A regression that knows my districts sit inside states, so it does not treat 700 districts in one state as 700 unrelated observations.* | Fixed effects estimated as population parameters, random effects as draws from a distribution, producing partial pooling — small groups shrink toward the global mean. The intra-class correlation is between-group variance over total variance, hence my 0.6 percent |
| Bayesian network | Formalism | *A diagram of variables with probabilities on the arrows, which lets you ask: given this and this, what is the chance of that?* | A directed acyclic graph with conditional probability tables factorising a joint distribution. Independencies are readable off the graph via d-separation; inference is exact or sampled |
| FastAPI | Framework | *The Python tool I used to build the back end — the part that receives a request from the website and answers it.* | ASGI, with Pydantic deriving request validation, serialisation and an OpenAPI schema from type hints. Async, so the process is not blocked waiting on an upstream model API |
| React / Next.js | Library / Framework | *The tools for building the part of a website people actually see and click.* | React maintains a virtual representation of the interface and reconciles changes. Next.js adds routing, server-side rendering and static generation, and server-side API routes |
| Docker | Tool | *It packs an application with everything it needs into one image so it runs identically anywhere. Reproducibility, but for deployment.* | Containers share the host kernel rather than booting a guest OS, so they start instantly. Images build in layers from a declarative Dockerfile. I needed it because `pyswisseph` compiles a native C extension |
| Streamlit | Framework | *It turns a plain Python script into a simple web app with almost no extra work. I used it for the thesis demo.* | Re-executes the entire script top to bottom on every interaction, with widgets returning values and decorators caching expensive work. That model is why it is trivial for prototypes and wrong for anything stateful |
| R | Language | *A language built for statistics from the start rather than adapted to it. Honestly tier three for me — but worth mentioning, because empirical legal scholars often use it.* | Vectorised by default, model formulas as a language-level construct, and an ecosystem where diagnostics and visual grammar are first-class |
| Prolog | Language | *A language where you write down facts and rules, and it works out the conclusions itself — you describe what is true rather than how to compute it.* | Horn clauses resolved by unification and backtracking search. Two properties matter legally: the closed-world assumption, so anything unproven is treated as false, and negation-as-failure, which is not classical negation — both are the wrong defaults for a burden of proof, and saying so is the interesting observation |
| Also on your CV | Kind | SAY THIS |
|---|---|---|
| PyTorch | Framework | *The deep-learning framework I actually use, and the one Hugging Face runs on.* |
| TensorFlow / Keras | Framework | *Coursework only — the Bi-LSTM offence classifier. I would not claim depth.* |
| LSTM | Architecture | *An older kind of network for sequences, with gates that let it remember things for longer than the plain version could.* |
| ARIMA | Method | *A classical forecasting method for a single series over time.* |
| Cross-validation / GridSearchCV | Method | *Hold back part of the data to check the model on data it has not seen; grid search repeats that across many settings. And note DiagnoChat tuned only one of eight models, so the comparison was not fair.* |
| SQLAlchemy / Alembic | Library | *Lets me work with database rows as Python objects, and keeps a versioned history of schema changes.* |
| Pydantic | Library | *Checks that incoming data is the right shape before my code touches it.* |
| boto3 | Library | *The Amazon storage client. I used it against Cloudflare's storage, which speaks the same protocol but charges nothing to download.* |
| OpenCV / Pillow | Library | *Image handling. I used the slimmed-down build to keep the container small.* |
| JWT / bcrypt | Standard / Algorithm | *A signed token that keeps someone logged in without the server storing a session, and a deliberately slow way of hashing passwords. My one-time codes are hashed too, which most side projects skip.* |
| pyswisseph | Library | *Computes real astronomical positions. It is the deterministic factual layer in Astroformer — the part the model is not allowed to invent.* |
| Matplotlib / Seaborn / Plotly | Library | *Charts: the basic one, the statistical one, and the interactive one.* |
| Git / GitHub | Tool | *Version control. Worth mentioning because a reproducibility claim means nothing without it.* |
| SQL | Language | *The language for querying relational databases.* |
| TypeScript | Language | *JavaScript with type checking, so mistakes are caught before it runs.* |
2. The one technical thing to be able to explain properly — QLoRA
If any single technical item gets probed, it will be this, because it is the most advanced thing you have personally done. So it gets its own section, built up from nothing.
| The letter | What it stands for | What it does, plainly |
|---|---|---|
| Q | Quantised | Store the frozen original model at 4-bit precision instead of 16-bit. Roughly a quarter of the memory, with modest loss of fidelity. This is why it needs an NVIDIA GPU — the 4-bit arithmetic requires CUDA and does not run on Apple Silicon |
| Lo | Low-Rank | Instead of one huge grid of new numbers, use two thin grids whose product has the same shape. Far fewer numbers for nearly the same effect |
| R | Adaptation | The original model is frozen. Only the thin grids learn. So you are adapting rather than retraining |
d is 4096 and you set r to 16, the full grid would hold about 16.8 million numbers. The two thin grids together hold about 131,000 — roughly 0.8 percent. `lora_alpha` is a scaling factor applied to that product, and the ratio of alpha to r controls how strongly the adapter speaks, which is why 32 over 16 is a common choice. Your own settings were `r = 16`, `alpha = 32`, dropout 0.05, applied to the query and value projections.3. The thesis system, as a stack
Ten stages, four external models, four APIs
| Result | Value | What it means |
|---|---|---|
| Perplexity | 9.72 | How surprised the model was, on average, at each word — roughly as uncertain as choosing between ten options. A respectable fluency figure for a small model. But it measures fluency, not correctness — say that when you quote it |
| BLEU | 0.0032 in the text, 0.032 in the table | Word-overlap with a reference. Very low — and there is a tenfold inconsistency between two places in your own document |
| ROUGE-1 / 2 / L F1 | 0.121 / 0.0227 / 0.1076 | Other overlap measures. All low |
| Dataset | 5,782 × 12 × 5 = 346,920 | The real contribution of the thesis |
4. Astroformer — the system you actually shipped
This is the only production system you have built with payments, authentication, storage and real users. It is worth being precise about, because *I shipped a thing that takes money and does not fall over* is a different claim from *I ran a notebook*.
The whole system, by layer
| Layer | What it is | The decision behind it |
|---|---|---|
| Frontend | Vercel, Next.js 16 / React 19 / Tailwind 4 | A CDN for static and cached content. Security headers set at the edge, so they apply before any application code runs and cannot be forgotten in a route |
| Backend | Render, Docker container, FastAPI + Uvicorn | Docker rather than a buildpack because `pyswisseph` compiles a native C extension. A buildpack is a guess; a Dockerfile is reproducible — the same argument as research reproducibility |
| GPU | None | Say this plainly; it is correct, not a gap. Serving a 70-billion-parameter model yourself means renting a GPU by the hour whether anyone visits or not; an API is priced per token so cost tracks usage. The trade you accepted: your data leaves your infrastructure — which is exactly why an agency with confidential business information would choose the opposite |
| Model providers | Gemini, NVIDIA's hosted endpoint, Groq. Which one is live is set by environment variable, not in code | Provider redundancy, and two NVIDIA keys for rotation. Free tiers have hard ceilings, so an application-level guard holds requests to 14 per minute |
| Database | Postgres via psycopg3 in production, SQLite locally. SQLAlchemy 2.0, Alembic. Nine models | The detail that shows you debug rather than copy: Render hands out a postgres:// URL and SQLAlchemy 2 with psycopg3 requires postgresql+psycopg://, so the session module rewrites the scheme at startup. Chat sessions are persisted, not held in memory, so a conversation survives a container restart |
| Storage | Cloudflare R2 via boto3 with S3 signature v4, bucket astroformer-assets. PDFs server-side and client-side | R2 speaks the S3 API so the same client works, but charges no egress fees — and a purchased report gets downloaded repeatedly. A pure cost decision, priced between two options |
| Auth | JWT, HS256, seven-day expiry. bcrypt for passwords. Hashed OTPs | Hashing the one-time codes is the detail worth mentioning — most hobby projects store them in plain text |
| Services | Razorpay and PayPal, Notion API as the back office, SendGrid, Firebase token verification, OpenCV headless | Notion as a back office is the judgment call to lead with: a real admin dashboard is weeks of work and non-technical order review was needed immediately. You bought time with an off-the-shelf interface |
| Chat endpoints | /astro-chat, /love-astro-chat, /tarot/premium-chat | Three separate tools sharing one credit pool |
5. LawReformer, and the one live risk on it
| Aspect | What it is |
|---|---|
| What it does | Six tools covering India, the UK and the US. **Rule-based, and the site says so — *not a chatbot*.** Disclaimers present |
| The sister site | ai.lawreformer.com, which points users to free legal aid through the District Legal Services Authority |
| Why you built it | Charging aggrieved people for legal information felt wrong. That is the honest motive and it is a good one |
| Why it is rule-based, and this is the strong point | *A generative model that invents a procedural deadline for someone with no lawyer does active harm. A rule-based tool can only output what I wrote, which is checkable and wrong in predictable ways rather than unpredictable ones.* That is a deliberate safety decision, not a technical limitation, and you should present it as one |
| What you learned from it | Very few users — and that is the informative part. It is free, so cost was not the barrier. *People do not look for legal information until they are already in trouble, and by then they want someone to act rather than a tool to read. So access to justice is not primarily an information-supply problem, and I had assumed it was.* A finding about your own assumption |
6. Outlier, and the two framings for it
7. The technical questions, with answers
| Question | The answer |
|---|---|
| *How is Astroformer deployed? What is it hosted on?* | Vercel for the Next.js frontend, Render for the Python API as a Docker container. Docker rather than a buildpack because the astrology engine compiles a native C extension. No GPUs — every model call goes to a hosted provider, chosen by environment variable. Postgres via psycopg3 in production, SQLite locally, Cloudflare R2 for files, JWT sessions. Then volunteer the free plan, the open CORS and the default secret |
| *What would you build first for an agency?* | A corpus and an entity layer, and it is deliberately unglamorous. *Before any of this is measurable you need the decisions in one place with provenance per document, and firms resolved to undertakings with a reported matching threshold. That is most of the first year and everything else depends on it* |
| *Walk me through a retrieval system over Commission decisions.* | *Collect with provenance per document. Chunk on structure — article, paragraph, recital — not a fixed token count, because splitting a holding from its reasoning destroys the passage. Embed, index, normalise the vectors so cosine and distance agree, retrieve more than you need and rerank. Keep a keyword baseline, because on legal text with precise terminology it is hard to beat and sometimes wins. Verify citations with a parser rather than a model, and log per query what was retrieved — that last part matters most, because without it you cannot reconstruct why the system said what it said* |
| *You have no GPUs and no infrastructure. Is that a problem?* | Answer it straight. *For this project I do not think so. The work is document-scale rather than training-scale — classification, retrieval, graph analysis over tens of thousands of documents, which runs on a laptop or a modest server. If we needed to train a model from scratch I would need compute I have never had. But most of what I would propose does not* |
| *What is the most technically difficult thing you have done?* | The thesis fine-tune, honestly stated. *Getting a 4-bit QLoRA run working on a rented GPU against a 346,920-row dataset, and then diagnosing why the retrieval half underperformed — the index embedded whole articles against a 256-word-piece limit, so each article was represented by roughly its opening paragraph. I had reported the symptom in the thesis without knowing the cause* |
| *What would you need to learn?* | Three things, and name them in this order. *EU competition law doctrine properly, which I said in my letter and have started on. Institutional economics, which is the project's third named pillar and which I do not have — applied econometrics is not the same thing. And formal annotation methodology — agreement statistics, codebook design, adjudication protocols — which is standard in empirical legal studies, is not on my CV, and is a few days of reading* |
| *Do you code every day?* | *Yes. Two live sites and two papers' worth of analysis in the past year.* Do not elaborate. This question is checking for a hobbyist, and the answer is a fact rather than an argument |