Skip to content
For AI labs

Your model is good at medicine. Find out where it isn't.

Public medical benchmarks saturate, leak into training data, and lean on non-expert raters. BenchEval supplies what they can't: fresh adversarial tasks and blinded expert judgment from verified, specialty-matched physicians.

Scope an evaluation
What we deliver

Evaluation, benchmarks, red-teaming, data

01

Model Evaluation

Blinded, rubric-driven grading of model outputs by verified physicians, with written clinical rationale on every judgment. Designed for the questions that decide launches: did behaviour meet the rubric under tested conditions, is it improving, would a specialist sign their name under it.

02

Adversarial Benchmarks

Original task sets authored by practising clinicians — atypical presentations, incomplete information, time pressure, distractors chosen because they fool models. Unpublished at authoring, distribution-controlled, monitored for contamination risk.

03

Clinical Red-Teaming

Structured attack campaigns run by physicians who know where medicine goes wrong: dosing edge cases, escalation thresholds, unsafe reassurance, specialty handoffs. Every finding reproducible and ranked by clinical severity.

04

Post-Training Data

Expert demonstrations, preference comparisons and error corrections produced by specialty-matched physicians under calibrated guidelines — authored by people qualified to overrule the model.

Labs don’t buy findings, they buy method. Ours is published in both directions: a completed evaluation of three frontier models showing exactly what the method surfaces and how each finding is evidenced, and the BenchEval Standard that governs it — taxonomy, severity scale and grading rules locked before a single prompt runs, so a finding can’t be chosen after the fact. Evaluations are graded under a published, versioned protocol, so findings are auditable and citable.

Built by the kind of expert your evaluation needs

Authored at the frontier

Our founders — an emergency physician and a neurologist — have authored benchmark tasks and trained models for frontier AI labs. The work is designed by people who have written to that bar.

Clinicians, not crowds

Every evaluator is a verified physician, matched to the specialty the task demands. A triage question goes to an emergency physician, not to whoever is online.

Adversarial by default

Our job is to find the failure before your users do. Tasks are written to break models — and refreshed as capabilities move.

Honest about scale

Founder-led with a deliberately grown physician network. You get senior attention on a scoped engagement — and a plain answer about what we will and won't take on.

Bring us the capability claim you need to test. We'll design the evaluation that could falsify it.

First conversations are with a founder — a working clinician, not an account manager. Scoped pilots before anything bigger.

Talk to us