Skip to content
Benchmark · Early access

Adversarial reasoning cases from the emergency department.

Our first benchmark tests frontier models where they fail most consequentially — hard, ambiguous, multi-pathology emergency presentations. Saturated QA sets measure recall; medicine is decided under ambiguity.

Request early access

What we publish, and the standard it’s graded under

Publications

BenchEval publishes its evaluations and the standard they are graded under. Findings cite the version of the BenchEval Standard they were graded against.

BENCHMARK PIPELINE · V1.0LAST UPDATED · AUG 2026
01 · AUTHORING
Cases written from scratch by emergency physicians
02 · DUAL REVIEW
Blind second-physician review of every case
03 · CALIBRATION RUNS
Grader agreement measured before any score is reported
04 · EARLY ACCESS
Partner labs evaluate ahead of public release — v1.0 opens Q4 2026
Target set150 CASES · V1.0
ContaminationUNPUBLISHED · DISTRIBUTION-CONTROLLED
Guideline pinningEACH CASE RECORDS THE VERSION IT TESTS
One fully worked case · Illustrative — not a live benchmark item

What a case, rubric and grade look like

VIGNETTE (EXCERPT)
62M, "flu for a week." HR 104, BP 108/74, T 37.9. Recent hip replacement, mentioned only in the medication list via rivaroxaban — which he stopped taking "because the course finished." Breathless walking to the toilet. Asks for antibiotics.
WHAT IT TESTS
Anchoring (viral illness), incomplete history (anticoagulation lapse buried in meds), escalation under ambiguity (PE workup vs. reassurance), and the question that grades the reasoning: what can't this presentation be allowed to be?
RUBRIC DIMENSIONS · GRADED PER CASE
01
Workup sequence — did the reasoning order investigations the way the risk demands
02
Escalation timing — was the can't-miss diagnosis excluded before reassurance
03
Uncertainty handling — does the answer know what it doesn't know, and say so
04
Safety of rationale — would the stated reasoning be defensible in a morbidity review
What makes it hard

Four failure modes, deliberately provoked

01

Anchoring

Cases engineered to punish premature closure: an early, plausible diagnosis that fits the first half of the vignette and fails the second.

02

Medication interactions

Polypharmacy the way EDs receive it: anticoagulants plus new prescriptions, renal dosing on incomplete information, the interaction that only matters because of what the patient didn't mention.

03

Atypical presentations

The MI without chest pain, the sepsis that looks like a fall — written from clinical experience, because training data under-represents them.

04

Guideline recency

Recommendations that changed recently enough that a model trained on stale guidance answers confidently and wrongly. Each case records the guideline version it tests.

Method: every case physician-authored from scratch, unpublished at authoring and distribution-controlled; reviewed blind by a second physician; graded on workup sequence, escalation timing and uncertainty handling — not just the final diagnosis. Each case records the guideline version it tests.

Run it before it's public.

Early access means evaluating against the benchmark ahead of public release — and helping decide what a defensible clinical reasoning score should measure.