Your product will meet patients. Evaluate it like that matters.
Scribes, triage tools, symptom checkers, decision support — deployed clinical AI fails in ways generic evals never surface. We test your product against the failures that would actually reach a patient, and give you reviewable evidence for your internal safety case.
Scope an evaluation







Our published audit of ChatGPT, Claude and Gemini on 20 emergency-care scenarios found the failure patterns benchmarks miss — read it.
The failures we're built to find
Confident wrong dosing
Renal adjustment missed, interactions unflagged, paediatric weight-band errors — delivered fluently enough that a busy clinician might not re-check.
Missed escalation
The presentation that needed a red flag and got reassurance. We probe escalation thresholds the way patients actually describe symptoms — vaguely, late, out of order.
Unsafe reassurance
Symptom checkers and scribes that normalise the abnormal: the 'likely benign' that anchors a workup away from the diagnosis that can't be missed.
Brittle triage
Rules that hold on textbook inputs and break on multimorbidity, polypharmacy, atypical presentations and incomplete histories.
Evidence you can put in front of a regulator's questions
Structured findings with clinical rationale, severity ranking and reproduction steps — written so they hold up when a sceptical clinician and a sceptical reviewer both read them.
Results are bounded evidence about observed behaviour under tested conditions — not certification, regulatory approval, or a guarantee of clinical safety.
Engagements — graded under BenchEval Standard v1.0
How to work with usHow it works: scenario design tuned to your product and population → locked grading protocol (pre-registered failure taxonomy F1–F8, severity scale S0–S3, three clinical dimensions) → runs with variance testing → severity-graded findings with verbatim receipts → a prioritised risk-reduction backlog your team can ship: escalation wording, mandatory next-step instructions, location handling, dispositional guardrails.
Pulse check
10 scenarios on your highest-risk flow · repeated-run variance on the worst finding · short written report + 45-min call · 1 week.
The easy yes: find out what a clinician sees in your product in seven days.
Focused red-team
30–50 scenarios tailored to your use case and populations (elderly, pregnancy, polypharmacy, mental health) · full taxonomy + severity grading · verbatim receipts · remediation backlog · readout call with your product/eng team · 2–3 weeks.
The standard engagement: a defensible answer to “how do you know it’s safe?”
Pre-enterprise / deep red-team
100+ scenarios · special populations and medication traps · adversarial phrasings · before/after re-test following your fixes · board/investor-ready report aligned to your risk register and regulatory pathway · scoped per product.
For teams heading into hospital pilots, enterprise procurement, or regulatory conversations.
Every engagement is graded under BenchEval Standard v1.0, and the mark and its evidence are yours to publish — see what the mark says.
Pricing is scoped per engagement, not listed.
Scope an evaluationNot a regulatory certification, not a clinical validation study, and not a substitute for your quality-management process — it’s the structured clinical risk-discovery layer that makes all three of those conversations easier. Device-classification and formal regulatory strategy: we’ll tell you when you need it and won’t pretend to be it.
Tell us what failure would cost you. We'll scope an evaluation against it.
Founder-led, so you talk to the people who design the work.
Talk to us