Benchmark · In development
EM-ARC: adversarial reasoning cases from emergency medicine.
Our first benchmark — working title EM-ARC (Emergency Medicine: Adversarial Reasoning Cases) — tests frontier models where they fail most consequentially: hard, ambiguous, multi-pathology emergency presentations.
Thesis
Saturated QA sets measure recall. Medicine is decided under ambiguity.
Frontier models score impressively on public medical QA — and still fail on the presentations that fill an emergency department: ambiguous symptoms, more than one active pathology, incomplete histories, and decisions that can't wait for certainty. EM-ARC is built from exactly those cases: physician-authored adversarial presentations, each independently reviewed by a second clinician and graded against a rubric that scores the reasoning — workup, escalation, uncertainty handling — not just the final diagnosis.
What makes it hard
Four failure modes, deliberately provoked
Each case targets at least one of the ways confident models — and tired clinicians — get emergency medicine wrong.
Anchoring
Cases are engineered to punish premature closure: an early, plausible diagnosis that fits the first half of the vignette and fails the second. Built so that pattern-matching to the obvious answer fails the way anchoring fails clinicians.
Medication interactions
Polypharmacy the way emergency departments actually receive it: anticoagulants plus new prescriptions, renal dosing on incomplete information, the interaction that only matters because of what the patient didn't mention.
Atypical presentations
The MI without chest pain, the sepsis that looks like a fall, the paediatric presentation textbooks describe in adults. Written from clinical experience, precisely because training data under-represents them.
Guideline recency
Recommendations that changed recently enough that a model trained on stale guidance answers confidently and wrongly. Each case records which guideline version it tests, so scores stay interpretable as guidance moves.
Method
Authored, reviewed, rubric-graded
Physician-authored
Every case is written from scratch by a practising physician — original, unpublished, and absent from any training set.
Dual review
A second, independent physician reviews each case blind: clinical accuracy, a defensible single best pathway, and no unintended giveaways.
Rubric grading
Scoring goes beyond the final answer: workup sequence, escalation timing, uncertainty handling, and the safety of the reasoning that got there.
Reporting format · Illustrative
What results will look like
A layout preview of the eventual report card. No evaluation runs exist yet — every cell below is deliberately empty.
Illustrative layout — no results yet. Benchmark in development.
| Model | Diagnostic accuracy | Escalation safety | Uncertainty handling | Overall |
|---|---|---|---|---|
| Frontier model A | No data | No data | No data | No data |
| Frontier model B | No data | No data | No data | No data |
| Frontier model C | No data | No data | No data | No data |
Early access
Labs and health-AI teams will run EM-ARC before it's public.
Early access means evaluating against the benchmark as it hardens — and helping decide what a defensible clinical reasoning score should measure. Physicians who want to author cases can apply to the network.