OpenStudy — the gate, shown.
Every model that reaches the people is evaluated like a clinical trial — pre-registered, receipted, and published in the open. No black box. We show the math.
How the gate works
Five rules. Every one re-derivable by anyone, on the receipts we publish.
Pre-registered endpoints
We declare what we're testing before we test it. No moving goalposts, no cherry-picking after the fact.
Deterministic gates
Rule-based scoring — not an LLM-as-judge. Length, structure, safety, concept checks that re-run identically.
Hash-chained receipts
Every question and every response, tamper-evident. Flip one byte and the chain breaks.
Published in the open
The full transcript — every call, every miss included. Nothing hidden behind a marketing number.
Beat-base-or-kill
A model only ships if it beats its own base on held-out prompts. Loss-only graphs don't count.
DiabeticDaily-4B · Discovery Trial
159 adversarial prompts across the danger zones a diabetic assistant must never get wrong — emergency, dosing, diagnosis, and deliberately dangerous asks. Scored on a defendability rubric, every call receipted.
We publish the misses too. A handful of prompts each model still over-cautions or under-warns on — named, not buried. That honesty is the result.
Nothing reaches the people that didn't pass
This is what OpenDiabetic guards. A warm front door means nothing if what's behind it isn't sound. The study is the gate — verified · vetted · defendable.
Education, not medical advice. Evaluations measure defendability and safety behavior — they do not certify clinical use. For emergencies, call 911.