Started in the emergency department. Aimed at the frontier.
BenchEval exists because the person evaluating a medical AI should be qualified to disagree with it.

Between them: benchmark tasks authored, models trained and clinical AI red-teamed for frontier labs. The gap was obvious from inside — labs and healthcare AI teams need expert evaluation at frontier quality, and physicians who can produce it have no serious front door into the work. BenchEval is that front door, built from both sides.
Two physicians. Both sides of the exam.
Founders
Dr Omer Atli
CO-FOUNDER · EMERGENCY MEDICINEPractising emergency physician; authors benchmark tasks for frontier AI labs — the cases that probe what a model genuinely understands, and where a confident answer quietly goes wrong.

Dr Semsettin Atli
CO-FOUNDER · NEUROLOGYNeurologist, 15+ years across the UK and Türkiye, most recently Senior Clinical Fellow at the Royal London Hospital. Clinically red-teams a frontier AI lab's medical model — signing off every judgement as clinician of record.
The thinking, on the record
In publicThe “shadow formulary”: the ungoverned general-purpose AI already in clinical use at 3am — and why the dangerous failure is fluency, not error. You can't ban it, so you have to govern it.
Listen to the episode →20 physician-authored emergency scenarios, run verbatim through ChatGPT, Claude and Gemini — 60 answers graded by the physicians who would be liable for acting on them. The failure that mattered: the confidently incomplete answer.
Read the full audit →Make expert judgment the standard AI is measured against
MissionAdversarial, not adversary
We try to break models because that's how they get safe. The point of finding a failure is that a patient never does.
No inflated numbers
No invented headcounts, logos or testimonials. When we cite a figure, it will be real, and we'll show where it came from.
Physicians are the product
Verification, specialty-matching and fair rates aren't overheads — they're the entire reason the work is worth buying.
Publish what survives scrutiny
We write everything as if a sceptical clinician and a sceptical researcher will both check it — because they will.
Founder-led, by design
Founder-led by design: a growing physician network, a first benchmark moving through early access, and scoped engagements where senior attention is a feature rather than a bottleneck. If you value depth over headcount, this is the right moment to talk.
BenchEval publishes its evaluation method openly — read the Standard.