REVISION 3 — Your Papers
Every paper in one place: what it set out to do, what actually worked, what did not, the numbers to quote, and the one weakness to volunteer. Plus the single question all of your work turns out to be asking.
Listen
0. Start here — the one sentence that ties them together
The same defect, five times
1. Beyond the Dark Figure — your lead publication
| What it is | SocArXiv preprint, DOI 10.31235/osf.io/4fxsm_v2, posted 19 September 2026. Sole author. A twenty-year district-level panel study of recorded crime in India, built from National Crime Records Bureau data with interpolated census demographics |
|---|---|
| What it set out to do | Ask whether recorded crime is driven by demographic factors — urbanisation, a male-female literacy gap, gender ratio — and whether the pattern is geographic |
| What worked | Three findings, and all three are defensible. No spatial clustering. State-level factors explain almost nothing. And a legislative change produced a large, measurable jump in *recording* |
| What did not | The literacy gap, which was the variable the paper was most interested in, is not statistically significant — and the abstract claims more than the statistics support |
| Result | The number | What it actually supports |
|---|---|---|
| Moran's I (spatial clustering) | I = 0.095, p = 0.151 | No evidence of geographic clustering. High-crime districts are not systematically next to other high-crime districts |
| ICC from the two-level model | state variance 3.895, district residual 662.59 → ICC = 0.6% | State policy and governance explain under 1 percent of variation in local rates. The other 99.4 percent is local. This is the headline and it is robust — a variance partition is a direct computation, not an inference |
| Urbanisation | β = 0.371, p = 0.004 | The one clearly significant demographic driver. The result you would actually defend |
| Literacy gap (mixed model) | β = 0.148, p = 0.156 | Not significant |
| Literacy gap (OLS, clustered errors) | β = 0.2049, p = 0.399 | Also not significant |
| Causal effect of literacy gap | ATE = 0.1487 | A point estimate after closing the urbanisation backdoor path — and note it is essentially the same number as the non-significant coefficient |
| Robustness check | re-estimated at 0.1485, p = 0.96 | The estimate barely moved when a random confounder was injected |
| BSTS, 2013 Criminal Law (Amendment) Act | +88,879 cases per year against the counterfactual | A reporting shock, not a behavioural epidemic. Your single most quotable result |
| Kidnapping → rape | r = 0.43; causal β = 0.26 | Offence types are linked rather than independent — an escalation pathway |
| Risk matrix | low urbanisation + high gap → P(high crime) = 49.0%; high urbanisation + high gap → 5.9% | A strongly non-linear interaction — the literacy gap's effect reverses with urbanisation |
| VIF diagnostics | all below 5 | No multicollinearity problem. Worth mentioning unprompted; it shows you ran diagnostics |
2. The Mens Rea Evaluator — the most ATLANTIS-relevant thing you have
| What it is | SSRN, DOI 10.2139/ssrn.7446198, dated 16 August 2026. Written during an Apart Research hackathon sprint — which is your legitimate answer on sample size |
|---|---|
| What it set out to do | Test whether a language model, given fabricated legal authority in its context, would produce a biased legal ruling — and whether it would acknowledge having done so |
| The design, and this is the strong part | A paired three-condition experiment. A clean *control*; an *ablation* with partial manipulation; a *poisoned* condition with fabricated authority planted in the retrieved context. Plus a parser-scored forced-choice probe and a per-run provenance log |
| What worked | The design worked and the result is clean. A clean baseline makes everything else interpretable, and the effect was total in the poisoned condition |
| What did not | 15 runs over 5 scenarios. And in 5 of those rows the model judging the output was also a subject — both of the two admissions fall in exactly those rows |
| Result | Value | What to say |
|---|---|---|
| Control | 0 of 15 biased | A clean baseline — which is what makes the rest mean anything |
| Ablation | 9 of 15 (4, 4, 1 of 5) | Partial manipulation, partial effect. Wide spread across scenarios, so not separable |
| Poisoned | 15 of 15 biased | Fabricated authority in context produced a biased ruling on every single run |
| Admissions | 2 of 15, both self-judged | In the 10 rows graded by a different model, none. Volunteer this |
| Cognizable intent | 0 of 15 | This is the finding — not the susceptibility rate |
| McNemar, control → ablation | p ≈ 0.004 | 9 changes one way, 0 back |
| McNemar, control → poisoned | p ≈ 0.00006 | 15 changes, 0 reversals |
| Fisher, model against model | p = 0.206 | Licenses nothing. No model ranking — say so |
| Wilson intervals | ≈38–96% and ≈4–62% | They overlap heavily — the formal reason the models are not separable |
3. Hybrid Legal Text Summarization — the honest negative result
| What it is | SSRN, DOI 10.2139/ssrn.6669602. Ghosh and Vetriselvi. A three-stage design study on summarising Indian judgments |
|---|---|
| What it set out to do | Produce summaries of judgments that keep the legally operative facts |
| What worked | The diagnosis. You tried the obvious thing, watched it fail, tried the pure-logic thing, watched it fail differently, and built a hybrid. The negative result is the contribution and it is a real one |
| What did not | There is no quantitative evaluation of any kind, and the abstract claims experimental results |
Three approaches, and why each failed
4. DiagnoChat — and the move that turns it into an asset
| Model | Accuracy | Note |
|---|---|---|
| Naive Bayes + hyperparameter search | 90.11% | The paper's headline — and the only model that was tuned |
| KNN | 88.94% | Untuned, second best |
| Logistic regression | 88.23% | Untuned |
| Random forest | 87.95% | Untuned |
| SVM | 86.81% | Untuned |
| Decision tree | 81.78% | Untuned |
| Weighted Bernoulli NB | 70.91% | The paper's custom variant |
| Plain Bernoulli NB | 66.73% | Worst of the eight |
5. The IoT paper — your only journal article, and closest to Wisman
| What it does | Detail |
|---|---|
| The argument | India has no law specifically regulating connected devices. The IT Act 2000 does not define cybercrime at all, and its provisions are bailable unless read with the penal code — so the paper maps IT Act sections onto penal-code sections to show how prosecutions actually have to be constructed. Conclusion: India needs a dedicated framework |
| The comparative section — mention this part | EU: the GDPR for data generated by devices, the NIS Directive for infrastructure, the Cybersecurity Act 2019. US: the IoT Cybersecurity Improvement Act 2020 for federal procurement, and California's device-security law, the first in the US. This is comparative EU/US/India law, which is exactly this panel's register |
| The empirical part | A questionnaire to 50 respondents in Kolkata, distributed by messaging apps. 98 percent daily internet use, 30.6 percent reusing passwords, 69.4 percent claiming to understand connected devices, 100 percent believing cybercrime had increased, 100 percent supporting separate legislation |
| The case material | A 2018 baby-monitor intrusion; a regulator's recall of around 500,000 pacemakers over hackability; a casino database reached through an aquarium thermostat; the 2016 botnet that took down major services; heating cut to Finnish apartment blocks for a week; and an Indian case where a complaint filed in November 2021 remained untraced |
6. The trap, and the answer that beats every individual result
7. If you remember eight things
- The through-line: all of your work asks what the measurement is actually measuring. Say it. It turns a scattered CV into a research programme and it is true.
- **Strike the word *proves*** and say unprompted that you would rewrite those sentences.
- Dark Figure's one real hole: the literacy gap is not significant — p = 0.399 and p = 0.156 — yet the abstract claims its causal effect is proven. Identification and precision are separate; concede the wording and defend urbanisation at β = 0.371, p = 0.004.
- Defend these three numbers without hesitation: Moran's I 0.095 at p = 0.151, ICC 0.6 percent, and +88,879 recorded cases a year from the 2013 amendment.
- Mens Rea: 0 of 15 control, 9 of 15 ablation, 15 of 15 poisoned, and 0 cognizable intent — which is the finding. Volunteer that both admissions were self-judged and that Fisher at p = 0.206 licenses no model ranking. And correct the framing: mens rea does not transplant; the result is about the duty to give reasons.
- The summarisation paper has no experiment despite the abstract claiming one. Reframe it as a design study — and its negative result, that hand-built logic does not scale across offence types, is the actual contribution.
- DiagnoChat: only one model was tuned, so the superiority claim does not follow — and accuracy is the wrong metric over 200-plus classes. Concede it, then make the AI Act and Medical Device Regulation point, which converts it into an asset.
- The IoT paper mis-names the GDPR. It is the General Data Protection Regulation, 2016/679, applicable from 25 May 2018. Concede it instantly if raised, and never repeat the error.