Skip to content
BenchEval

For AI labs

Your model is good at medicine. Find out where it isn't.

Public medical benchmarks saturate, leak into training data, and lean on non-expert raters. BenchEval supplies what they can't: fresh adversarial tasks and blinded expert judgment from verified, specialty-matched physicians.

What we deliver

Evaluation, benchmarks, red-teaming, data

Four instruments, one discipline: physician-authored tasks, calibrated expert grading, and findings you can defend internally.

01

Model Evaluation

Blinded, rubric-driven grading of model outputs by verified physicians, with written clinical rationale on every judgment. Designed for the questions that decide launches: is this behaviour safe, is it improving, and would a specialist sign their name under it.

02

Adversarial Benchmarks

Original, contamination-resistant task sets authored by practising clinicians — atypical presentations, incomplete information, time pressure, and distractors chosen because they fool models, not because they fill a category. Built to stay hard after the leaderboard settles.

03

Clinical Red-Teaming

Structured attack campaigns run by physicians who know where medicine actually goes wrong: dosing edge cases, escalation thresholds, unsafe reassurance, specialty handoffs. Every finding is reproducible and ranked by clinical severity.

04

Post-Training Data

Expert demonstrations, preference comparisons and error corrections produced by specialty-matched physicians under calibrated guidelines — the data for teaching models clinical judgment, authored by people qualified to overrule them.

Why BenchEval

Built by the kind of expert your evaluation needs

Authored at the frontier

Our founder is a practising emergency physician who has authored benchmark tasks for frontier AI labs. The work is designed by someone who has written to that bar.

Clinicians, not crowds

Every evaluator is a verified physician, matched to the specialty the task demands. A triage question goes to an emergency physician, not to whoever is online.

Adversarial by default

Our job is to find the failure before your users do. Tasks are written to break models — saturation-resistant, contamination-aware, and refreshed as capabilities move.

Honest about scale

We are founder-led and building our physician network deliberately. You get senior attention on a scoped engagement — and we will tell you plainly what we can and cannot take on yet.

Start a conversation

Bring us the capability claim you need to test. We'll design the evaluation that could falsify it.

First conversations are with the founder — a working clinician, not an account manager. Scoped pilots before anything bigger.

Talk to us