Skip to content
For healthcare AI teams

Your product will meet patients. Evaluate it like that matters.

Scribes, triage tools, symptom checkers, decision support — deployed clinical AI fails in ways generic evals never surface. We test your product against the failures that would actually reach a patient, and give you reviewable evidence for your internal safety case.

Scope an evaluation

Our published audit of ChatGPT, Claude and Gemini on 20 emergency-care scenarios found the failure patterns benchmarks miss — read it.

Where deployed clinical AI fails

The failures we're built to find

01

Confident wrong dosing

Renal adjustment missed, interactions unflagged, paediatric weight-band errors — delivered fluently enough that a busy clinician might not re-check.

02

Missed escalation

The presentation that needed a red flag and got reassurance. We probe escalation thresholds the way patients actually describe symptoms — vaguely, late, out of order.

03

Unsafe reassurance

Symptom checkers and scribes that normalise the abnormal: the 'likely benign' that anchors a workup away from the diagnosis that can't be missed.

04

Brittle triage

Rules that hold on textbook inputs and break on multimorbidity, polypharmacy, atypical presentations and incomplete histories.

Evidence you can put in front of a regulator's questions

Structured findings with clinical rationale, severity ranking and reproduction steps — written so they hold up when a sceptical clinician and a sceptical reviewer both read them.

Results are bounded evidence about observed behaviour under tested conditions — not certification, regulatory approval, or a guarantee of clinical safety.

SPECIALTY COVERAGE — RECRUITING & VERIFYING NOW
Emergency MedicineInternal MedicineGeneral PracticeCardiologyRadiologyPaediatricsPsychiatryOncologyAnaestheticsCritical CareObstetrics & GynaecologyNeurologyInfectious DiseasesSurgery

Engagements — graded under BenchEval Standard v1.0

How to work with us

How it works: scenario design tuned to your product and population → locked grading protocol (pre-registered failure taxonomy F1–F8, severity scale S0–S3, three clinical dimensions) → runs with variance testing → severity-graded findings with verbatim receipts → a prioritised risk-reduction backlog your team can ship: escalation wording, mandatory next-step instructions, location handling, dispositional guardrails.

01

Pulse check

10 scenarios on your highest-risk flow · repeated-run variance on the worst finding · short written report + 45-min call · 1 week.

The easy yes: find out what a clinician sees in your product in seven days.

02

Focused red-team

30–50 scenarios tailored to your use case and populations (elderly, pregnancy, polypharmacy, mental health) · full taxonomy + severity grading · verbatim receipts · remediation backlog · readout call with your product/eng team · 2–3 weeks.

The standard engagement: a defensible answer to “how do you know it’s safe?”

03

Pre-enterprise / deep red-team

100+ scenarios · special populations and medication traps · adversarial phrasings · before/after re-test following your fixes · board/investor-ready report aligned to your risk register and regulatory pathway · scoped per product.

For teams heading into hospital pilots, enterprise procurement, or regulatory conversations.

Every engagement is graded under BenchEval Standard v1.0, and the mark and its evidence are yours to publish — see what the mark says.

Pricing is scoped per engagement, not listed.

Scope an evaluation
What this is not

Not a regulatory certification, not a clinical validation study, and not a substitute for your quality-management process — it’s the structured clinical risk-discovery layer that makes all three of those conversations easier. Device-classification and formal regulatory strategy: we’ll tell you when you need it and won’t pretend to be it.

Tell us what failure would cost you. We'll scope an evaluation against it.

Founder-led, so you talk to the people who design the work.

Talk to us