Patients already ask AI before they reach the ED. We tested what happens next.
The models recognised the red flags. The failures came after recognition — in action, urgency, localisation and safety-netting.
Read the audit →Our first benchmark tests frontier models where they fail most consequentially — hard, ambiguous, multi-pathology emergency presentations. Saturated QA sets measure recall; medicine is decided under ambiguity.
Request early access
BenchEval publishes its evaluations and the standard they are graded under. Findings cite the version of the BenchEval Standard they were graded against.
The models recognised the red flags. The failures came after recognition — in action, urgency, localisation and safety-netting.
Read the audit →Physician-authored, dual-reviewed, contamination-controlled cases that test frontier models on hard, ambiguous, multi-pathology emergency presentations. Results stay unpublished while the set is access-gated.
See the pipeline →Cases engineered to punish premature closure: an early, plausible diagnosis that fits the first half of the vignette and fails the second.
Polypharmacy the way EDs receive it: anticoagulants plus new prescriptions, renal dosing on incomplete information, the interaction that only matters because of what the patient didn't mention.
The MI without chest pain, the sepsis that looks like a fall — written from clinical experience, because training data under-represents them.
Recommendations that changed recently enough that a model trained on stale guidance answers confidently and wrongly. Each case records the guideline version it tests.
Method: every case physician-authored from scratch, unpublished at authoring and distribution-controlled; reviewed blind by a second physician; graded on workup sequence, escalation timing and uncertainty handling — not just the final diagnosis. Each case records the guideline version it tests.
Early access means evaluating against the benchmark ahead of public release — and helping decide what a defensible clinical reasoning score should measure.