Physician-authored · Adversarial · Specialty-matched
Expert evaluation for AI that can't afford to be wrong.
BenchEval builds physician-authored evaluations, adversarial benchmarks, red-teaming and post-training data for AI labs and healthcare-AI companies — starting where the stakes are highest: medicine.
What we do
Four ways to find out what your model actually knows
Every engagement is designed and reviewed by a practising physician. The work is adversarial by default — built to surface failure, not to certify comfort.
Model Evaluation
Physician-authored rubrics and blinded expert grading of model behaviour on real clinical tasks — not crowd-rated plausibility. You get defensible judgments about where a model is safe, and where it only sounds safe.
Adversarial Benchmarks
Benchmarks built the way hard cases actually present: atypical, time-pressured, incomplete. Written fresh by clinicians, so performance measures reasoning — not memorised training data.
Clinical Red-Teaming
Practising physicians probing for the failures that matter in medicine: confident wrong dosing, missed escalation, unsafe reassurance, brittle triage. Documented, reproducible, prioritised by clinical risk.
Post-Training Data
Specialty-matched physicians producing the data that moves frontier models: expert demonstrations, graded comparisons, and corrections written by people qualified to disagree with the model.
How it works
From your risk to reviewable evidence
Define the failure that matters
We start from your risk, not a generic test set — the claim you need to make, the deployment you need to defend, the failure mode that would end trust in your product.
Match the right clinicians
Verified, specialty-matched physicians are briefed and calibrated on your task. Emergency physicians for triage. Oncologists for oncology. No general-purpose raters on specialist questions.
Deliver evidence you can act on
Structured findings with expert rationale attached to every judgment — what failed, why it's dangerous, and what data would fix it. Results you can hand to your safety team, not just a score.
Coverage
Building a specialty-matched physician network
We recruit and verify practising clinicians across the specialties that high-stakes medical AI touches first.
- Emergency Medicine
- Internal Medicine
- General Practice
- Cardiology
- Radiology
- Paediatrics
- Psychiatry
- Oncology
- Anaesthetics
- Critical Care
- Obstetrics & Gynaecology
- Neurology
- Infectious Diseases
- Surgery
For AI labs & healthcare AI
Put your model in front of physicians whose job is to break it.
Tell us what failure would cost you, and we'll scope an evaluation against it. Founder-led, so you talk to the person who designs the work.
Talk to usFor physicians
Your clinical judgment, applied where it shapes frontier AI.
Paid, remote, specialty-matched work — writing cases, grading model outputs and red-teaming systems that will meet patients.
Apply as an Expert