Skip to content
VibeFormer
46 min

REVISION 3 — Your Papers

Every paper in one place: what it set out to do, what actually worked, what did not, the numbers to quote, and the one weakness to volunteer. Plus the single question all of your work turns out to be asking.

Listen

0. Start here — the one sentence that ties them together

The same defect, five times

Every item in the left column is the thing you cared about. Every item in the right column is the thing that was available. The gap between the two columns is called measurement validity, and the bottom row is the problem this project exists to solve.

1. Beyond the Dark Figure — your lead publication

What it isSocArXiv preprint, DOI 10.31235/osf.io/4fxsm_v2, posted 19 September 2026. Sole author. A twenty-year district-level panel study of recorded crime in India, built from National Crime Records Bureau data with interpolated census demographics
What it set out to doAsk whether recorded crime is driven by demographic factors — urbanisation, a male-female literacy gap, gender ratio — and whether the pattern is geographic
What workedThree findings, and all three are defensible. No spatial clustering. State-level factors explain almost nothing. And a legislative change produced a large, measurable jump in *recording*
What did notThe literacy gap, which was the variable the paper was most interested in, is not statistically significant — and the abstract claims more than the statistics support
ResultThe numberWhat it actually supports
Moran's I (spatial clustering)I = 0.095, p = 0.151No evidence of geographic clustering. High-crime districts are not systematically next to other high-crime districts
ICC from the two-level modelstate variance 3.895, district residual 662.59 → ICC = 0.6%State policy and governance explain under 1 percent of variation in local rates. The other 99.4 percent is local. This is the headline and it is robust — a variance partition is a direct computation, not an inference
Urbanisationβ = 0.371, p = 0.004The one clearly significant demographic driver. The result you would actually defend
Literacy gap (mixed model)β = 0.148, p = 0.156Not significant
Literacy gap (OLS, clustered errors)β = 0.2049, p = 0.399Also not significant
Causal effect of literacy gapATE = 0.1487A point estimate after closing the urbanisation backdoor path — and note it is essentially the same number as the non-significant coefficient
Robustness checkre-estimated at 0.1485, p = 0.96The estimate barely moved when a random confounder was injected
BSTS, 2013 Criminal Law (Amendment) Act+88,879 cases per year against the counterfactualA reporting shock, not a behavioural epidemic. Your single most quotable result
Kidnapping → raper = 0.43; causal β = 0.26Offence types are linked rather than independent — an escalation pathway
Risk matrixlow urbanisation + high gap → P(high crime) = 49.0%; high urbanisation + high gap → 5.9%A strongly non-linear interaction — the literacy gap's effect reverses with urbanisation
VIF diagnosticsall below 5No multicollinearity problem. Worth mentioning unprompted; it shows you ran diagnostics

2. The Mens Rea Evaluator — the most ATLANTIS-relevant thing you have

What it isSSRN, DOI 10.2139/ssrn.7446198, dated 16 August 2026. Written during an Apart Research hackathon sprint — which is your legitimate answer on sample size
What it set out to doTest whether a language model, given fabricated legal authority in its context, would produce a biased legal ruling — and whether it would acknowledge having done so
The design, and this is the strong partA paired three-condition experiment. A clean *control*; an *ablation* with partial manipulation; a *poisoned* condition with fabricated authority planted in the retrieved context. Plus a parser-scored forced-choice probe and a per-run provenance log
What workedThe design worked and the result is clean. A clean baseline makes everything else interpretable, and the effect was total in the poisoned condition
What did not15 runs over 5 scenarios. And in 5 of those rows the model judging the output was also a subject — both of the two admissions fall in exactly those rows
ResultValueWhat to say
Control0 of 15 biasedA clean baseline — which is what makes the rest mean anything
Ablation9 of 15 (4, 4, 1 of 5)Partial manipulation, partial effect. Wide spread across scenarios, so not separable
Poisoned15 of 15 biasedFabricated authority in context produced a biased ruling on every single run
Admissions2 of 15, both self-judgedIn the 10 rows graded by a different model, none. Volunteer this
Cognizable intent0 of 15This is the finding — not the susceptibility rate
McNemar, control → ablationp ≈ 0.0049 changes one way, 0 back
McNemar, control → poisonedp ≈ 0.0000615 changes, 0 reversals
Fisher, model against modelp = 0.206Licenses nothing. No model ranking — say so
Wilson intervals≈38–96% and ≈4–62%They overlap heavily — the formal reason the models are not separable

3. Hybrid Legal Text Summarization — the honest negative result

What it isSSRN, DOI 10.2139/ssrn.6669602. Ghosh and Vetriselvi. A three-stage design study on summarising Indian judgments
What it set out to doProduce summaries of judgments that keep the legally operative facts
What workedThe diagnosis. You tried the obvious thing, watched it fail, tried the pure-logic thing, watched it fail differently, and built a hybrid. The negative result is the contribution and it is a real one
What did notThere is no quantitative evaluation of any kind, and the abstract claims experimental results

Three approaches, and why each failed

Read Approach B's diagnosis as the finding. You did not fail to make logic work — you demonstrated why hand-built logic does not scale across offence types, which is precisely the argument for pairing a learned model with a logic layer rather than choosing one. That is the neurosymbolic case, reached empirically rather than asserted.

4. DiagnoChat — and the move that turns it into an asset

ModelAccuracyNote
Naive Bayes + hyperparameter search90.11%The paper's headline — and the only model that was tuned
KNN88.94%Untuned, second best
Logistic regression88.23%Untuned
Random forest87.95%Untuned
SVM86.81%Untuned
Decision tree81.78%Untuned
Weighted Bernoulli NB70.91%The paper's custom variant
Plain Bernoulli NB66.73%Worst of the eight

5. The IoT paper — your only journal article, and closest to Wisman

What it doesDetail
The argumentIndia has no law specifically regulating connected devices. The IT Act 2000 does not define cybercrime at all, and its provisions are bailable unless read with the penal code — so the paper maps IT Act sections onto penal-code sections to show how prosecutions actually have to be constructed. Conclusion: India needs a dedicated framework
The comparative section — mention this partEU: the GDPR for data generated by devices, the NIS Directive for infrastructure, the Cybersecurity Act 2019. US: the IoT Cybersecurity Improvement Act 2020 for federal procurement, and California's device-security law, the first in the US. This is comparative EU/US/India law, which is exactly this panel's register
The empirical partA questionnaire to 50 respondents in Kolkata, distributed by messaging apps. 98 percent daily internet use, 30.6 percent reusing passwords, 69.4 percent claiming to understand connected devices, 100 percent believing cybercrime had increased, 100 percent supporting separate legislation
The case materialA 2018 baby-monitor intrusion; a regulator's recall of around 500,000 pacemakers over hackability; a casino database reached through an aquarium thermostat; the 2016 botnet that took down major services; heating cut to Finnish apartment blocks for a week; and an Indian case where a complaint filed in November 2021 remained untraced

6. The trap, and the answer that beats every individual result

7. If you remember eight things

  1. The through-line: all of your work asks what the measurement is actually measuring. Say it. It turns a scattered CV into a research programme and it is true.
  2. **Strike the word *proves*** and say unprompted that you would rewrite those sentences.
  3. Dark Figure's one real hole: the literacy gap is not significant — p = 0.399 and p = 0.156 — yet the abstract claims its causal effect is proven. Identification and precision are separate; concede the wording and defend urbanisation at β = 0.371, p = 0.004.
  4. Defend these three numbers without hesitation: Moran's I 0.095 at p = 0.151, ICC 0.6 percent, and +88,879 recorded cases a year from the 2013 amendment.
  5. Mens Rea: 0 of 15 control, 9 of 15 ablation, 15 of 15 poisoned, and 0 cognizable intent — which is the finding. Volunteer that both admissions were self-judged and that Fisher at p = 0.206 licenses no model ranking. And correct the framing: mens rea does not transplant; the result is about the duty to give reasons.
  6. The summarisation paper has no experiment despite the abstract claiming one. Reframe it as a design study — and its negative result, that hand-built logic does not scale across offence types, is the actual contribution.
  7. DiagnoChat: only one model was tuned, so the superiority claim does not follow — and accuracy is the wrong metric over 200-plus classes. Concede it, then make the AI Act and Medical Device Regulation point, which converts it into an asset.
  8. The IoT paper mis-names the GDPR. It is the General Data Protection Regulation, 2016/679, applicable from 25 May 2018. Concede it instantly if raised, and never repeat the error.