Your model is good at medicine. Find out where it isn't.
Public medical benchmarks saturate, leak into training data, and lean on non-expert raters. BenchEval supplies what they can't: fresh adversarial tasks and blinded expert judgment from verified, specialty-matched physicians.
Scope an evaluation




Evaluation, benchmarks, red-teaming, data
Model Evaluation
Blinded, rubric-driven grading of model outputs by verified physicians, with written clinical rationale on every judgment. Designed for the questions that decide launches: did behaviour meet the rubric under tested conditions, is it improving, would a specialist sign their name under it.
Adversarial Benchmarks
Original task sets authored by practising clinicians — atypical presentations, incomplete information, time pressure, distractors chosen because they fool models. Unpublished at authoring, distribution-controlled, monitored for contamination risk.
Clinical Red-Teaming
Structured attack campaigns run by physicians who know where medicine goes wrong: dosing edge cases, escalation thresholds, unsafe reassurance, specialty handoffs. Every finding reproducible and ranked by clinical severity.
Post-Training Data
Expert demonstrations, preference comparisons and error corrections produced by specialty-matched physicians under calibrated guidelines — authored by people qualified to overrule the model.
Labs don’t buy findings, they buy method. Ours is published in both directions: a completed evaluation of three frontier models showing exactly what the method surfaces and how each finding is evidenced, and the BenchEval Standard that governs it — taxonomy, severity scale and grading rules locked before a single prompt runs, so a finding can’t be chosen after the fact. Evaluations are graded under a published, versioned protocol, so findings are auditable and citable.
Built by the kind of expert your evaluation needs
Authored at the frontier
Our founders — an emergency physician and a neurologist — have authored benchmark tasks and trained models for frontier AI labs. The work is designed by people who have written to that bar.
Clinicians, not crowds
Every evaluator is a verified physician, matched to the specialty the task demands. A triage question goes to an emergency physician, not to whoever is online.
Adversarial by default
Our job is to find the failure before your users do. Tasks are written to break models — and refreshed as capabilities move.
Honest about scale
Founder-led with a deliberately grown physician network. You get senior attention on a scoped engagement — and a plain answer about what we will and won't take on.
Bring us the capability claim you need to test. We'll design the evaluation that could falsify it.
First conversations are with a founder — a working clinician, not an account manager. Scoped pilots before anything bigger.
Talk to us