Ask a capable generalist rater to compare two model responses to a medication question. Both answers are fluent, structured, and reassuring. One of them is wrong in a way that matters — an interaction missed, a contraindication glossed, a dosing edge case handled with confidence instead of caution. The rater cannot see it. Not because they are careless, but because seeing it is what a clinical education is.
This is the structural problem with medical AI evaluation, and it is why healthcare is the sharpest case for verified expertise: the failure modes that matter most are precisely the ones that are invisible to a non-clinician. Fluency and correctness diverge, and only one of them is on the surface.
Where generalist evaluation runs out
Generalist raters are genuinely good at a lot: instruction-following, tone, formatting, obvious refusals. Clinical work leaves that territory fast. Judging whether a triage recommendation is safe requires knowing what a dangerous presentation looks like. Judging a differential requires knowing what should be on it. Grading against current standards of care requires knowing what current is — guidelines move, and last decade's confident answer can be this year's error.
The deeper issue is asymmetry: a wrong label from a non-expert does not look wrong. It looks like a clean, agreeing label. Aggregating more non-expert opinions does not fix this — consensus among people who cannot see the error simply launders it. The model then learns, with high confidence, the exact shape of plausible-but-unsafe.
What the rubric work actually looks like
The productive form of clinical evaluation is not a thumbs-up from someone in scrubs. It is rubric-graded assessment built with clinicians: what a safe answer to this class of question must contain, what it must never contain, and when it must escalate to "see a doctor now." Clinicians are essential twice — once in writing those rubrics, where the edge cases live, and again in applying them, where judgment fills the space rubrics cannot enumerate.
- Rubric construction — clinicians define what safe, current, complete answers look like per specialty and per risk class.
- Graded evaluation — licensed professionals apply the rubrics, with inter-rater reliability tracked and disagreements adjudicated, not averaged away.
- Red-teaming — clinicians probe for the failures only they can construct: the plausible drug interaction, the reassuring answer to an emergency presentation.
- Escalation review — the highest-risk judgments route to the most senior verified specialists, because seniority is a verified property, not a screen name.
Why verification is the load-bearing wall
Everything above assumes one thing: the clinician is a clinician. That assumption is exactly what standard marketplace vetting cannot deliver. A profile that says "MD" is a text field. A white coat on a webcam is a costume. If the graders' expertise is the product, then the graders' identity and credentials are the supply chain — and an unverified supply chain fails silently, in the labels, where no one can see it.
This is why we verify clinicians the way we verify everyone, and why it matters most here: identity, employer, role, and seniority, at source, against systems of record — medical licensure being one of the few credentials with genuine public registries behind it. The verification attaches to the work as provenance, which means a health-system buyer or a safety team can take any graded judgment and answer the question their own auditors will ask: who said so, and were they qualified to?
In clinical evaluation, the grader's credential is not a nice-to-have. It is the measurement instrument.
Medical AI is moving from demos into care pathways, and the evaluation behind it is becoming the evidence base regulators, hospitals, and patients will implicitly rely on. Evidence needs provenance. That is the standard we run clinical engagements to — and the same logic extends to every domain where wrong answers cost the most, which is exactly where we have chosen to work.
Expert human data, verified at source.