Your Master's Thesis, Stage by Stage
All seventy-seven pages, in order: how you gathered the data, how you cleaned it, every technique you tried, exactly what you trained and where, what the figures show, what the numbers mean — and an honest account of what is good, what lacks, and why the absence of novelty is not a problem.
0. The honest frame — start here
1. The facts of the document
| Title | *Fine-Tuning LLM with RAG for Generating Educational Audio-Visual Content* |
|---|---|
| Degree | MSc Data Science, School of Advanced Sciences, Vellore Institute of Technology. Registration 23MDT0077 |
| Internal guide | Dr Sathyanarayana Sharma K, Department of Mathematics — and note he also supervised DiagnoChat |
| External guide | Mr Kushal Rastogi |
| Period on the certificate | 1 January to 5 May 2025 — so four months, start to finish, alone. Hold that number; it is the right frame for everything below |
| Length | 77 pages, 14 figures, 7 tables, 14 references |
| What it does | A topic name goes in; a grade-level explanation comes out, in English or one of six Indian languages, as text and as a narrated video |
| Scope claimed | CBSE and ICSE curricula, Classes 1 to 12 |
2. Every term in your own title, defined
**Your thesis is called *Fine-Tuning LLM with RAG for Generating Educational Audio-Visual Content*. Four of those words are technical, and if someone asks *what does fine-tuning actually mean?* you need a real answer rather than a pipeline description. Each definition below is three or four sentences, with an example, and you should be able to say any of them cold.**
Prompting, retrieval and fine-tuning — what each one actually changes
How searching by meaning works — the part that sounds like magic
| The other terms in your method | Three-sentence definition |
|---|---|
| Token | A token is the unit of text a model actually processes — usually a word, or a fragment of one. Models do not read letters or words; they read a fixed vocabulary of pieces, so *competition* might arrive as compet + ition. This matters practically because every limit in the system is counted in tokens: my sequences were capped at 512 and generation at 100 |
| Context window | The maximum number of tokens the model can hold in front of it at once — question, retrieved passages, instructions and answer, all together. Exceed it and the earliest material is simply dropped. Mine was 512 tokens, which is small, and it is why stuffing in three retrieved documents plus three web snippets was too much |
| Quantisation | Storing the model's numbers at lower precision to make it fit in less memory — 4 bits per number instead of 16. The analogy is reprinting a book in smaller type on thinner paper: the same content, a quarter of the shelf space, slightly harder to read. **It is the *Q* in QLoRA, and it is the reason the fine-tune needed an NVIDIA GPU rather than my laptop** |
| Adapter / LoRA | Instead of changing all 1.1 billion numbers, you freeze them and train a small set of extra numbers alongside — about one percent as many. The analogy is leaving the encyclopaedia untouched and writing a thin booklet of additions read alongside it. The practical benefit is that you can keep several booklets for several tasks against one shared base model |
| Distillation (what you actually did, though the thesis does not use the word) | Using a large expensive model to generate training data, then training a small cheap model on it, so the small one inherits some of the large one's behaviour. The large model is the *teacher*, the small one the *student*. My pipeline is a distillation: GPT-4o Mini wrote the explanations and TinyLlama learned to imitate them — which is also exactly why my evaluation was circular, because the teacher's answers were also my references |
| Perplexity | A measure of how surprised the model was, on average, by the text it was asked to predict. Lower is better, and a perplexity of about 10 means it was roughly as uncertain as choosing between ten equally likely options. Mine was 9.72 — and the thing to say alongside it is that perplexity measures fluency, not correctness |
| BLEU and ROUGE | Both measure how much word overlap there is between the model's output and a reference answer written by a human or another model. BLEU was built for translation and counts matching word sequences; ROUGE was built for summarisation and counts matching words and longest common sequences. The shared weakness is decisive for legal work: a summary that reverses the holding can still score well, because the words overlap |
3. The whole pipeline, in one picture
Ten stages — memorise the shape, not the details
4. How you gathered the data
| Step | What you actually did | The honest note |
|---|---|---|
| The topic list | 5,782 unique topics, compiled by manual extraction plus public datasets from academic databases and educational websites, spanning Mathematics, Science, Social Science, Language Arts and Environmental Studies across Classes 1 to 12 | The weakest link in the whole provenance chain. *Manual extraction plus public datasets* is not a reproducible sourcing method, and the thesis does not list which datasets. If asked how you would redo it: start from the published syllabus documents themselves, record the source per topic, and version the list |
| Retrieval | Wikipedia API, programmatically, per topic. Deliberately not just the lead paragraph — full article content, then filtered text blocks by keyword density and semantic similarity to the topic name | This part is genuinely good and worth saying. Most student projects grab the summary field. You went for subsections because that is where the substance is |
| Quality filtering | Dropped disambiguation pages, stub articles, and articles flagged as lacking sources | Also good, and also worth volunteering — it is an explicit data-quality decision with a stated rationale |
| Redirect handling | The code detects #REDIRECT in the returned content and recursively re-fetches the target title | A small, correct engineering detail |
| Storage | Raw material stored as JSON, each topic mapped to its extract, then annotated for later stages. Thesis describes the raw corpus as *millions of words* | No checksum, no retrieval date recorded per article. Wikipedia changes continuously, so the corpus is not reconstructible — name this if reproducibility comes up |
5. How you cleaned it
This is the stage most people skip describing, and the one that is most worth describing, because it is where sixty percent of real data work lives. Yours was explicit and multi-step.
The cleaning pipeline, and what each step removes
6. The BART stage — what it was for, and the honest question about it
| What you used | `facebook/bart-large-cnn` through the Hugging Face summarization pipeline |
|---|---|
| Why | Cleaned Wikipedia prose is still stiff and encyclopaedic. BART was used to rewrite each paragraph into something more readable while keeping the meaning |
| How, exactly | Text chunked at 1024 tokens with 100-token overlap (because that is BART's input limit), then max_length=300, min_length=50, do_sample=False. Chunk outputs rejoined. Failures caught per chunk and skipped |
| What BART is | An encoder-decoder model — it reads the whole input, then writes an output. The -cnn suffix means this checkpoint was fine-tuned on news summarisation |
7. Grade adaptation and augmentation — how 5,782 became 346,920
The multiplication, and where each factor comes from
| Grade adaptation model | `gpt-4o-mini-2024-07-18`, via the OpenAI API |
|---|---|
| Settings | temperature=0.7, max_tokens=600, system prompt *you are an expert teacher adapting explanations for different grade levels* |
| Engineering | ThreadPoolExecutor with 5 workers, batches of 50, 3 retries with exponential backoff, progress saved to CSV after each batch. This is competent production-style code and worth mentioning |
| The question side | Templated variants — *Explain to a Grade 4 student about Photosynthesis*, *In Grade 4, Photosynthesis*, *A student in Class 4 might ask...*, *As per the Grade 4 syllabus...*, *A Class 4 student is asking about...* |
8. The training — exactly what you did, and where it ran
| Setting | Value | What it means in plain words |
|---|---|---|
| Base model | TinyLlama/TinyLlama-1.1B-Chat-v1.0 | A 1.1-billion-parameter open model on the LLaMA architecture. Chosen for footprint, not quality — the design constraint was that a school could run it |
| Quantisation | load_in_4bit=True, torch_dtype=float16 | Store the frozen base at 4-bit instead of 16-bit — about a quarter of the memory. This is what needs CUDA |
| `r` | 16 | The rank — how thick the booklet is. With a model width of 4096, the full update grid would hold about 16.8 million numbers; two thin grids at r=16 hold about 131,000 — roughly 0.8 percent |
| `lora_alpha` | 32 | A scaling factor on the adapter's contribution. The ratio of alpha to r is what controls how loudly the adapter speaks, which is why 32 over 16 is a common pairing |
| `lora_dropout` | 0.05 | Randomly ignore 5 percent of adapter connections during training, to discourage memorising |
| `target_modules` | ["q_proj", "v_proj"] | Only the query and value projections inside attention get adapters. The standard minimal choice — cheapest place to intervene |
| Epochs | 2 | Two complete passes over all 346,920 rows |
| Batch / accumulation | 4 and 8 | Effective batch of 32 — process 4 at a time but only update the weights every 8 batches, which simulates a larger batch on small memory |
| `max_length` | 512, padding="max_length" | Every sequence padded to 512 tokens. Wasteful, and see the warning below |
| Precision | fp16=True | 16-bit arithmetic during training — faster, less memory |
| Output | ./tinyllama-qa, then saved to a Google Drive path | The Drive path is the evidence of where it actually ran |
9. The retrieval stage — and the bug you can now diagnose
| Corpus | load_dataset("wikipedia", "20220301.simple", split="train[:10%]") — **10 percent of *Simple English* Wikipedia. The thesis text says *English* Wikipedia, which is wrong and a much larger claim** |
|---|---|
| Embedding model | all-MiniLM-L6-v2 — a small, fast sentence-embedding model producing 384-number vectors |
| Index | faiss.IndexFlatL2 — exact search by squared Euclidean distance. Flat means no approximation: it compares against everything |
| Retrieval | top k = 3 documents, concatenated into the prompt |
| Generation cap | max_new_tokens=100 |
The bug — and it explains a result the thesis already reports
10. Web search, translation, and the video — including three unsupported claims
| Stage | What the code actually does | What the thesis claims |
|---|---|---|
| Web search | api.duckduckgo.com/?q={query}&format=json, takes the first 3 RelatedTopics text fields | Claims it is a conditional fallback for low-confidence answers. It is not — it runs every time |
| Translation | A free Google-Translate wrapper (reference 14 is the GitHub project) | Three different mechanisms named in three places. §3.2 says the MyMemory API; §4.2.6 says M2M100 and IndicTrans; §4.3 says the free wrapper. **And §4.2.7 confusingly describes the MyMemory API as the *text-to-speech* engine, which it is not. This is the most findable inconsistency in the document** |
| Round-trip check | Translates out, translates back, computes a BLEU score and displays it to the user | **Claims the system *alerts for human adjustment or re-translation* when drift is large. There is no threshold and no alert** — a claimed feature that was not built |
| Speech | A single gTTS(text=..., lang=...) call | Claims every sentence was scanned for pause, pitch and rhythm, with prosody *manually tweaked for every language*. gTTS exposes no prosody control at all |
| Video | MoviePy TextClip per character, concatenated, duration = audio length ÷ character count; rendered MP4 via FFmpeg, libx264 video, aac audio, 24 fps | Matches. This part is accurately described |
| Interface | Streamlit: text input, language dropdown, answer, translated answer, back-translation BLEU, narrated video | Matches, and Figures 12–14 are screenshots of it running |
11. What the figures actually show — redrawn so you remember them
Fourteen figures. Four of them are worth being able to describe, because they are the only exploratory analysis in the thesis — and if anyone asks *did you look at your data*, these are the answer.
Figures 6 and 7 — question and answer length distributions
Figures 10 and 11 — part-of-speech distributions
12. The results, and what they actually mean
| Metric | Value | What it means, plainly |
|---|---|---|
| Perplexity | 9.72 | How surprised the model was, on average, at each word — roughly as uncertain as picking between ten options. A respectable fluency figure for a 1.1-billion-parameter model. But it measures fluency, not correctness — say that when you quote it |
| BLEU | 0.0032 in the text, 0.032 in Table 7 | Word-overlap with a reference. Very low. And there is a tenfold inconsistency between two places in your own document |
| ROUGE-1 F1 | 0.1210 | Single-word overlap. The text also cites a ROUGE-1 recall of 0.529 — a different quantity, so not strictly contradictory, but precision is never reported |
| ROUGE-2 F1 | 0.0227 | Two-word-phrase overlap. Near zero |
| ROUGE-L F1 | 0.1076 | Longest common subsequence |
| Back-translation BLEU | Hindi 0.6168, Kannada 0.5363, Telugu 0.5325, Bengali 0.5324, Punjabi 0.5084, Tamil 0.5040 | From a single example sentence. See the warning below |
Why the overlap scores cannot mean what they appear to mean
13. What is good, what lacks — the honest ledger
| Genuinely good | Why it counts |
|---|---|
| It is end-to-end and it works | Most student LLM projects stop at a notebook. This one ships an app that takes a question and returns a narrated video in a chosen language, and Figures 12 to 14 are screenshots of it running |
| The dataset construction | 346,920 rows through a five-stage pipeline with explicit arithmetic at every step. This is data engineering, and it is the skill the position actually needs |
| Explicit data-quality decisions | Dropping disambiguation pages, stubs and unsourced articles; retrieving subsections rather than just lead paragraphs. Both are deliberate and both are defensible |
| Production-quality engineering in the API stage | Thread pool, batching, retries with exponential backoff, incremental saves. Not student code |
| The cost reasoning | Small open model, parameter-efficient tuning, free knowledge source, free search. **Every choice follows from *a school with no budget must be able to run this*** |
| Retrieval added for the correct reason | The thesis states it plainly: a fine-tuned model will hallucinate or go stale, so ground it in retrieved text. That is the right motivation, stated before it was fashionable |
| An honest limitations section | Eight limitations including hardware, dataset bias, unvalidated grade alignment, API dependency, and explicitly that evaluation with students and teachers could not be done. Point at it |
| A relevant literature review | Fourteen references and they are the right ones — LaMDA, LLaMA, REALM, the original RAG paper, the hallucination survey, long-tail knowledge, retrieval-quality evaluation. Not padding |
| What lacks | The honest statement |
|---|---|
| No methodological novelty | Every component is published work. *The contribution is integration and the dataset, not a new method, and I would not claim otherwise* |
| No held-out split described | The most serious methodological gap. *I cannot rule out that my metrics include training data* |
| Circular evaluation | *Training answers and reference answers both came from GPT-4o Mini, so BLEU and ROUGE measured imitation of the teacher rather than correctness* |
| The central claim is never tested | *The system promises grade-appropriate explanations, and grade-appropriateness is the one thing I never measured.* This is the best sentence available about this work |
| A claimed metric that is never reported | §4.2.3 says periodic tests used *specially defined metrics for grade-level appropriateness using linguistic readability measures*. No such metric appears anywhere in the results |
| User testing claimed and denied | The discussion claims it; Limitation 8 says it could not be done. Volunteer this one |
| Three features described but not built | Confidence-gated fallback, round-trip alerting, prosody tuning |
| Prompt tokens not masked in the loss | The model was partly trained to echo questions |
| Retrieval index truncated | Whole articles embedded against a 256-token limit |
| Not reproducible | Hardcoded local paths, a live API key in the document, no corpus checksum or retrieval dates, no pinned environment |
14. What to say, and the likely questions
| Question | The answer |
|---|---|
| *What was novel about it?* | Answer straight. *Nothing methodologically, and I would not pretend otherwise. What it demonstrates is integration under constraint — ten stages, four external models, four APIs, four months, alone. For a master's thesis I think that is the right thing to have demonstrated, and the thing I actually learned was how badly I had designed the evaluation* |
| *How did you get the data?* | *A list of 5,782 curriculum topics, then the Wikipedia API per topic — full articles rather than just lead paragraphs, filtered by keyword density and similarity to the topic name, with disambiguation pages, stubs and unsourced articles dropped. The weak link is the topic list itself: manual extraction plus public datasets, which is not a reproducible sourcing method, and I would start from the published syllabus documents if I did it again* |
| *How did you clean it?* | *Eight steps: HTML stripping, then regex for references, templates, citation markers and links, then whitespace and punctuation normalisation, then sentence segmentation, then paragraph assembly at four sentences. And you can see in my own output figure that nested templates survived the regex — the right tool is a wikitext parser, not regular expressions, and I never quantified what fraction of rows carried markup through* |
| *What did you train, and on what hardware?* | *A 4-bit QLoRA fine-tune of TinyLlama-1.1B — rank 16, alpha 32, adapters on the query and value projections, two epochs, effective batch of 32, 512-token sequences. Development and the whole pipeline ran on my MacBook; the fine-tune itself needed CUDA so it ran on rented GPUs, Colab and RunPod. The thesis says MacBook in several places and that is imprecise* |
| *Why such a small model?* | *Cost and deployability. GPT-4o Mini was used once, offline, to build a dataset; TinyLlama is what would actually serve students, and it is small enough to run without renting a GPU per request. The pattern is to use the expensive model to manufacture training data and distil it into a cheap one* |
| *Your scores are very low. What happened?* | Three reasons, in order of honesty. *The metrics were the wrong family — overlap cannot see a factually wrong summary. Generation was capped at 100 new tokens while my answers averaged 500 to 600 characters, so outputs were truncated mechanically. And the references came from the same model that wrote the training data, so the metric measured imitation. What the task needed was a factual-accuracy rubric and a readability measure* |
| *Did you validate the grade levels?* | The hardest question, and the answer is no. *No. My own limitations section says the grade adaptation may not align with curriculum standards, and section 4.2.3 claims readability metrics were used in periodic testing which are never reported. So the central promise of the system is the one claim with no evidence behind it, and I found that myself* |
15. If you remember nine things
- Revoke the API key on page 37. Today, separately from the interview.
- Lead with *there is no novelty and I would not claim any* — then say what it does demonstrate: integration under constraint, ten stages, four months, alone.
- The numbers that are safe: 5,782 topics × 12 grades × 5 paraphrases = 346,920 rows; TinyLlama-1.1B with QLoRA at r = 16, alpha = 32; two epochs; perplexity 9.72.
- Perplexity measures fluency, not correctness. Say that whenever you quote 9.72.
- The dataset was multiplied, not grown — all 346,920 rows descend from 5,782 Wikipedia articles.
- Hardware: three environments. MacBook for development, pipeline, index, video and app; Colab and RunPod for the fine-tune, because 4-bit training needs CUDA.
- The retrieval bug is your best story: whole articles embedded against a 256-word-piece limit, so the index held only opening paragraphs — **and that is exactly the *tangent retrievals* symptom you had already reported honestly without knowing the cause.**
- The evaluation is circular and there is no described held-out split. Volunteer it, then name the replacement: held-out topics, independent references, a factual-accuracy rubric, a readability measure.
- The six language BLEU scores come from one sentence and measure a translation API rather than your model — and grade-appropriateness, the system's central promise, was never tested. That is the most honest and most impressive thing you can say.