Skip to content
VibeFormer
44 min

Your Master's Thesis, Stage by Stage

All seventy-seven pages, in order: how you gathered the data, how you cleaned it, every technique you tried, exactly what you trained and where, what the figures show, what the numbers mean — and an honest account of what is good, what lacks, and why the absence of novelty is not a problem.

Listen

0. The honest frame — start here

1. The facts of the document

Title*Fine-Tuning LLM with RAG for Generating Educational Audio-Visual Content*
DegreeMSc Data Science, School of Advanced Sciences, Vellore Institute of Technology. Registration 23MDT0077
Internal guideDr Sathyanarayana Sharma K, Department of Mathematics — and note he also supervised DiagnoChat
External guideMr Kushal Rastogi
Period on the certificate1 January to 5 May 2025 — so four months, start to finish, alone. Hold that number; it is the right frame for everything below
Length77 pages, 14 figures, 7 tables, 14 references
What it doesA topic name goes in; a grade-level explanation comes out, in English or one of six Indian languages, as text and as a narrated video
Scope claimedCBSE and ICSE curricula, Classes 1 to 12

2. Every term in your own title, defined

**Your thesis is called *Fine-Tuning LLM with RAG for Generating Educational Audio-Visual Content*. Four of those words are technical, and if someone asks *what does fine-tuning actually mean?* you need a real answer rather than a pipeline description. Each definition below is three or four sentences, with an example, and you should be able to say any of them cold.**

Prompting, retrieval and fine-tuning — what each one actually changes

This is the single most useful diagram for defending the thesis, because the obvious question is why you did both. Fine-tuning taught the model the shape of the task; retrieval gave it facts it had never seen. They solve different problems and the distinction is the answer.

How searching by meaning works — the part that sounds like magic

An embedding converts text into coordinates so that similar meanings sit close together, which is what lets retrieval work without shared keywords. The limitation is the same fact seen from the other side: contextual similarity is not legal equivalence.
The other terms in your methodThree-sentence definition
TokenA token is the unit of text a model actually processes — usually a word, or a fragment of one. Models do not read letters or words; they read a fixed vocabulary of pieces, so *competition* might arrive as compet + ition. This matters practically because every limit in the system is counted in tokens: my sequences were capped at 512 and generation at 100
Context windowThe maximum number of tokens the model can hold in front of it at once — question, retrieved passages, instructions and answer, all together. Exceed it and the earliest material is simply dropped. Mine was 512 tokens, which is small, and it is why stuffing in three retrieved documents plus three web snippets was too much
QuantisationStoring the model's numbers at lower precision to make it fit in less memory — 4 bits per number instead of 16. The analogy is reprinting a book in smaller type on thinner paper: the same content, a quarter of the shelf space, slightly harder to read. **It is the *Q* in QLoRA, and it is the reason the fine-tune needed an NVIDIA GPU rather than my laptop**
Adapter / LoRAInstead of changing all 1.1 billion numbers, you freeze them and train a small set of extra numbers alongside — about one percent as many. The analogy is leaving the encyclopaedia untouched and writing a thin booklet of additions read alongside it. The practical benefit is that you can keep several booklets for several tasks against one shared base model
Distillation (what you actually did, though the thesis does not use the word)Using a large expensive model to generate training data, then training a small cheap model on it, so the small one inherits some of the large one's behaviour. The large model is the *teacher*, the small one the *student*. My pipeline is a distillation: GPT-4o Mini wrote the explanations and TinyLlama learned to imitate them — which is also exactly why my evaluation was circular, because the teacher's answers were also my references
PerplexityA measure of how surprised the model was, on average, by the text it was asked to predict. Lower is better, and a perplexity of about 10 means it was roughly as uncertain as choosing between ten equally likely options. Mine was 9.72 — and the thing to say alongside it is that perplexity measures fluency, not correctness
BLEU and ROUGEBoth measure how much word overlap there is between the model's output and a reference answer written by a human or another model. BLEU was built for translation and counts matching word sequences; ROUGE was built for summarisation and counts matching words and longest common sequences. The shared weakness is decisive for legal work: a summary that reverses the holding can still score well, because the words overlap

3. The whole pipeline, in one picture

Ten stages — memorise the shape, not the details

Only stage six is model training. The other nine are data engineering, integration and systems work — which is the honest description of what data science jobs actually are, and the right thing to emphasise because it is the skill this position needs.

4. How you gathered the data

StepWhat you actually didThe honest note
The topic list5,782 unique topics, compiled by manual extraction plus public datasets from academic databases and educational websites, spanning Mathematics, Science, Social Science, Language Arts and Environmental Studies across Classes 1 to 12The weakest link in the whole provenance chain. *Manual extraction plus public datasets* is not a reproducible sourcing method, and the thesis does not list which datasets. If asked how you would redo it: start from the published syllabus documents themselves, record the source per topic, and version the list
RetrievalWikipedia API, programmatically, per topic. Deliberately not just the lead paragraph — full article content, then filtered text blocks by keyword density and semantic similarity to the topic nameThis part is genuinely good and worth saying. Most student projects grab the summary field. You went for subsections because that is where the substance is
Quality filteringDropped disambiguation pages, stub articles, and articles flagged as lacking sourcesAlso good, and also worth volunteering — it is an explicit data-quality decision with a stated rationale
Redirect handlingThe code detects #REDIRECT in the returned content and recursively re-fetches the target titleA small, correct engineering detail
StorageRaw material stored as JSON, each topic mapped to its extract, then annotated for later stages. Thesis describes the raw corpus as *millions of words*No checksum, no retrieval date recorded per article. Wikipedia changes continuously, so the corpus is not reconstructible — name this if reproducibility comes up

5. How you cleaned it

This is the stage most people skip describing, and the one that is most worth describing, because it is where sixty percent of real data work lives. Yours was explicit and multi-step.

The cleaning pipeline, and what each step removes

Eight distinct operations, in order, each removing a different class of noise. The paragraph-assembly rule at the end is the only non-obvious one: lines are accumulated until there are at least four sentences, which produces chunks of roughly consistent length rather than splitting on arbitrary newlines.

6. The BART stage — what it was for, and the honest question about it

What you used`facebook/bart-large-cnn` through the Hugging Face summarization pipeline
WhyCleaned Wikipedia prose is still stiff and encyclopaedic. BART was used to rewrite each paragraph into something more readable while keeping the meaning
How, exactlyText chunked at 1024 tokens with 100-token overlap (because that is BART's input limit), then max_length=300, min_length=50, do_sample=False. Chunk outputs rejoined. Failures caught per chunk and skipped
What BART isAn encoder-decoder model — it reads the whole input, then writes an output. The -cnn suffix means this checkpoint was fine-tuned on news summarisation

7. Grade adaptation and augmentation — how 5,782 became 346,920

The multiplication, and where each factor comes from

The two figures in the thesis show this as exponential growth, but the honest reading is that the underlying information content never increased after stage two. Five paraphrases of one question against the same answer are five rows and one fact. This is the single most important thing to understand about your own dataset.
Grade adaptation model`gpt-4o-mini-2024-07-18`, via the OpenAI API
Settingstemperature=0.7, max_tokens=600, system prompt *you are an expert teacher adapting explanations for different grade levels*
EngineeringThreadPoolExecutor with 5 workers, batches of 50, 3 retries with exponential backoff, progress saved to CSV after each batch. This is competent production-style code and worth mentioning
The question sideTemplated variants — *Explain to a Grade 4 student about Photosynthesis*, *In Grade 4, Photosynthesis*, *A student in Class 4 might ask...*, *As per the Grade 4 syllabus...*, *A Class 4 student is asking about...*

8. The training — exactly what you did, and where it ran

SettingValueWhat it means in plain words
Base modelTinyLlama/TinyLlama-1.1B-Chat-v1.0A 1.1-billion-parameter open model on the LLaMA architecture. Chosen for footprint, not quality — the design constraint was that a school could run it
Quantisationload_in_4bit=True, torch_dtype=float16Store the frozen base at 4-bit instead of 16-bit — about a quarter of the memory. This is what needs CUDA
`r`16The rank — how thick the booklet is. With a model width of 4096, the full update grid would hold about 16.8 million numbers; two thin grids at r=16 hold about 131,000 — roughly 0.8 percent
`lora_alpha`32A scaling factor on the adapter's contribution. The ratio of alpha to r is what controls how loudly the adapter speaks, which is why 32 over 16 is a common pairing
`lora_dropout`0.05Randomly ignore 5 percent of adapter connections during training, to discourage memorising
`target_modules`["q_proj", "v_proj"]Only the query and value projections inside attention get adapters. The standard minimal choice — cheapest place to intervene
Epochs2Two complete passes over all 346,920 rows
Batch / accumulation4 and 8Effective batch of 32 — process 4 at a time but only update the weights every 8 batches, which simulates a larger batch on small memory
`max_length`512, padding="max_length"Every sequence padded to 512 tokens. Wasteful, and see the warning below
Precisionfp16=True16-bit arithmetic during training — faster, less memory
Output./tinyllama-qa, then saved to a Google Drive pathThe Drive path is the evidence of where it actually ran

9. The retrieval stage — and the bug you can now diagnose

Corpusload_dataset("wikipedia", "20220301.simple", split="train[:10%]") — **10 percent of *Simple English* Wikipedia. The thesis text says *English* Wikipedia, which is wrong and a much larger claim**
Embedding modelall-MiniLM-L6-v2 — a small, fast sentence-embedding model producing 384-number vectors
Indexfaiss.IndexFlatL2 — exact search by squared Euclidean distance. Flat means no approximation: it compares against everything
Retrievaltop k = 3 documents, concatenated into the prompt
Generation capmax_new_tokens=100

The bug — and it explains a result the thesis already reports

This is the strongest thing you can say about the thesis. You reported an anomaly honestly without knowing its cause, and you can now state the mechanism and the fix. That is what diagnosis looks like, and it is a different and better claim than having built something that worked.

10. Web search, translation, and the video — including three unsupported claims

StageWhat the code actually doesWhat the thesis claims
Web searchapi.duckduckgo.com/?q={query}&format=json, takes the first 3 RelatedTopics text fieldsClaims it is a conditional fallback for low-confidence answers. It is not — it runs every time
TranslationA free Google-Translate wrapper (reference 14 is the GitHub project)Three different mechanisms named in three places. §3.2 says the MyMemory API; §4.2.6 says M2M100 and IndicTrans; §4.3 says the free wrapper. **And §4.2.7 confusingly describes the MyMemory API as the *text-to-speech* engine, which it is not. This is the most findable inconsistency in the document**
Round-trip checkTranslates out, translates back, computes a BLEU score and displays it to the user**Claims the system *alerts for human adjustment or re-translation* when drift is large. There is no threshold and no alert** — a claimed feature that was not built
SpeechA single gTTS(text=..., lang=...) callClaims every sentence was scanned for pause, pitch and rhythm, with prosody *manually tweaked for every language*. gTTS exposes no prosody control at all
VideoMoviePy TextClip per character, concatenated, duration = audio length ÷ character count; rendered MP4 via FFmpeg, libx264 video, aac audio, 24 fpsMatches. This part is accurately described
InterfaceStreamlit: text input, language dropdown, answer, translated answer, back-translation BLEU, narrated videoMatches, and Figures 12–14 are screenshots of it running

11. What the figures actually show — redrawn so you remember them

Fourteen figures. Four of them are worth being able to describe, because they are the only exploratory analysis in the thesis — and if anyone asks *did you look at your data*, these are the answer.

Figures 6 and 7 — question and answer length distributions

The question distribution is narrow because a template produced it. The answer distribution is broad and long-tailed, which collides directly with the hundred-token generation cap — and that collision is a mechanical reason the overlap scores were low, independent of whether the content was any good.

Figures 10 and 11 — part-of-speech distributions

The two distributions tell opposite stories. The question side is dominated by proper nouns and numbers because it was template-generated. The answer side looks like normal English prose, which is the right shape. Being able to read your own figure as evidence about your own pipeline is the useful skill here.

12. The results, and what they actually mean

MetricValueWhat it means, plainly
Perplexity9.72How surprised the model was, on average, at each word — roughly as uncertain as picking between ten options. A respectable fluency figure for a 1.1-billion-parameter model. But it measures fluency, not correctness — say that when you quote it
BLEU0.0032 in the text, 0.032 in Table 7Word-overlap with a reference. Very low. And there is a tenfold inconsistency between two places in your own document
ROUGE-1 F10.1210Single-word overlap. The text also cites a ROUGE-1 recall of 0.529 — a different quantity, so not strictly contradictory, but precision is never reported
ROUGE-2 F10.0227Two-word-phrase overlap. Near zero
ROUGE-L F10.1076Longest common subsequence
Back-translation BLEUHindi 0.6168, Kannada 0.5363, Telugu 0.5325, Bengali 0.5324, Punjabi 0.5084, Tamil 0.5040From a single example sentence. See the warning below
PPL=exp⁡ ⁣(−1N∑t=1Nlog⁡P(xt∣x1,…,xt−1))\mathrm{PPL} = \exp\!\left(-\frac{1}{N}\sum_{t=1}^{N} \log P(x_t \mid x_1, \ldots, x_{t-1})\right)
Perplexity is e raised to the average negative log probability the model assigned to each token it actually saw.In plain terms: how surprised the model was, expressed as an effective number of equally likely choices. 9.72 means that at each word it was about as uncertain as choosing among ten options. The key point for the interview: a fluent model and a correct model are different things, and perplexity only sees the first.

Why the overlap scores cannot mean what they appear to mean

Any one of these would weaken the numbers. Together they mean the reported scores cannot support a claim about quality — which is the single most important honest thing to say about this thesis, and it is a conclusion you reached about your own work.

13. What is good, what lacks — the honest ledger

Genuinely goodWhy it counts
It is end-to-end and it worksMost student LLM projects stop at a notebook. This one ships an app that takes a question and returns a narrated video in a chosen language, and Figures 12 to 14 are screenshots of it running
The dataset construction346,920 rows through a five-stage pipeline with explicit arithmetic at every step. This is data engineering, and it is the skill the position actually needs
Explicit data-quality decisionsDropping disambiguation pages, stubs and unsourced articles; retrieving subsections rather than just lead paragraphs. Both are deliberate and both are defensible
Production-quality engineering in the API stageThread pool, batching, retries with exponential backoff, incremental saves. Not student code
The cost reasoningSmall open model, parameter-efficient tuning, free knowledge source, free search. **Every choice follows from *a school with no budget must be able to run this***
Retrieval added for the correct reasonThe thesis states it plainly: a fine-tuned model will hallucinate or go stale, so ground it in retrieved text. That is the right motivation, stated before it was fashionable
An honest limitations sectionEight limitations including hardware, dataset bias, unvalidated grade alignment, API dependency, and explicitly that evaluation with students and teachers could not be done. Point at it
A relevant literature reviewFourteen references and they are the right ones — LaMDA, LLaMA, REALM, the original RAG paper, the hallucination survey, long-tail knowledge, retrieval-quality evaluation. Not padding
What lacksThe honest statement
No methodological noveltyEvery component is published work. *The contribution is integration and the dataset, not a new method, and I would not claim otherwise*
No held-out split describedThe most serious methodological gap. *I cannot rule out that my metrics include training data*
Circular evaluation*Training answers and reference answers both came from GPT-4o Mini, so BLEU and ROUGE measured imitation of the teacher rather than correctness*
The central claim is never tested*The system promises grade-appropriate explanations, and grade-appropriateness is the one thing I never measured.* This is the best sentence available about this work
A claimed metric that is never reported§4.2.3 says periodic tests used *specially defined metrics for grade-level appropriateness using linguistic readability measures*. No such metric appears anywhere in the results
User testing claimed and deniedThe discussion claims it; Limitation 8 says it could not be done. Volunteer this one
Three features described but not builtConfidence-gated fallback, round-trip alerting, prosody tuning
Prompt tokens not masked in the lossThe model was partly trained to echo questions
Retrieval index truncatedWhole articles embedded against a 256-token limit
Not reproducibleHardcoded local paths, a live API key in the document, no corpus checksum or retrieval dates, no pinned environment

14. What to say, and the likely questions

QuestionThe answer
*What was novel about it?*Answer straight. *Nothing methodologically, and I would not pretend otherwise. What it demonstrates is integration under constraint — ten stages, four external models, four APIs, four months, alone. For a master's thesis I think that is the right thing to have demonstrated, and the thing I actually learned was how badly I had designed the evaluation*
*How did you get the data?**A list of 5,782 curriculum topics, then the Wikipedia API per topic — full articles rather than just lead paragraphs, filtered by keyword density and similarity to the topic name, with disambiguation pages, stubs and unsourced articles dropped. The weak link is the topic list itself: manual extraction plus public datasets, which is not a reproducible sourcing method, and I would start from the published syllabus documents if I did it again*
*How did you clean it?**Eight steps: HTML stripping, then regex for references, templates, citation markers and links, then whitespace and punctuation normalisation, then sentence segmentation, then paragraph assembly at four sentences. And you can see in my own output figure that nested templates survived the regex — the right tool is a wikitext parser, not regular expressions, and I never quantified what fraction of rows carried markup through*
*What did you train, and on what hardware?**A 4-bit QLoRA fine-tune of TinyLlama-1.1B — rank 16, alpha 32, adapters on the query and value projections, two epochs, effective batch of 32, 512-token sequences. Development and the whole pipeline ran on my MacBook; the fine-tune itself needed CUDA so it ran on rented GPUs, Colab and RunPod. The thesis says MacBook in several places and that is imprecise*
*Why such a small model?**Cost and deployability. GPT-4o Mini was used once, offline, to build a dataset; TinyLlama is what would actually serve students, and it is small enough to run without renting a GPU per request. The pattern is to use the expensive model to manufacture training data and distil it into a cheap one*
*Your scores are very low. What happened?*Three reasons, in order of honesty. *The metrics were the wrong family — overlap cannot see a factually wrong summary. Generation was capped at 100 new tokens while my answers averaged 500 to 600 characters, so outputs were truncated mechanically. And the references came from the same model that wrote the training data, so the metric measured imitation. What the task needed was a factual-accuracy rubric and a readability measure*
*Did you validate the grade levels?*The hardest question, and the answer is no. *No. My own limitations section says the grade adaptation may not align with curriculum standards, and section 4.2.3 claims readability metrics were used in periodic testing which are never reported. So the central promise of the system is the one claim with no evidence behind it, and I found that myself*

15. If you remember nine things

  1. Revoke the API key on page 37. Today, separately from the interview.
  2. Lead with *there is no novelty and I would not claim any* — then say what it does demonstrate: integration under constraint, ten stages, four months, alone.
  3. The numbers that are safe: 5,782 topics × 12 grades × 5 paraphrases = 346,920 rows; TinyLlama-1.1B with QLoRA at r = 16, alpha = 32; two epochs; perplexity 9.72.
  4. Perplexity measures fluency, not correctness. Say that whenever you quote 9.72.
  5. The dataset was multiplied, not grown — all 346,920 rows descend from 5,782 Wikipedia articles.
  6. Hardware: three environments. MacBook for development, pipeline, index, video and app; Colab and RunPod for the fine-tune, because 4-bit training needs CUDA.
  7. The retrieval bug is your best story: whole articles embedded against a 256-word-piece limit, so the index held only opening paragraphs — **and that is exactly the *tangent retrievals* symptom you had already reported honestly without knowing the cause.**
  8. The evaluation is circular and there is no described held-out split. Volunteer it, then name the replacement: held-out topics, independent references, a factual-accuracy rubric, a readability measure.
  9. The six language BLEU scores come from one sentence and measure a translation API rather than your model — and grade-appropriateness, the system's central promise, was never tested. That is the most honest and most impressive thing you can say.