Griffin Farrow, Lily Sijia Li, Jack Johnson, Tingyan Wang, Philip Torr, William Bolton, Fabio J. Fehr
Rigorous, timely critique of the field's dominant medical-LLM evaluation method, with reusable taxonomy/pipeline and a credible mechanistic account, bounded by injected-error framing and a preliminary remedy.
Hallucinations can undermine clinician trust in LLMs, making it important that evaluation methods capture clinically relevant errors. Rubric-based evaluation has become the leading approach for assessing LLMs in medicine, but it is unclear whether rubric scores reflect such errors. We first study this in a controlled setting using MedHallu, finding that more specific rubrics better distinguish correct from hallucinated responses. To test this systematically, we develop a taxonomy of medical hallucination types and a clinician-validated error-injection pipeline that creates matched correct and error-injected responses. Across HealthBench, HealthBench Professional, and LiveMedBench, our clinically relevant hallucinations are missed by rubrics, often leaving scores unchanged. We find that rubrics are most effective when explicitly checking facts, and are less effective for additional or unexpected errors they do not anticipate. A preliminary retrieval-based factuality check recovers some of the rubric-blind errors, suggesting a complementary approach. These findings reveal systematic blind spots in current medical evaluation of LLMs and suggest that rubric scores alone are insufficient to establish clinical reliability, potentially undermining clinician trust and confidence in clinical deployment.
Core Contribution. This paper interrogates a widely-adopted but under-scrutinized practice: rubric-based evaluation of medical LLMs (e.g., HealthBench). Its central claim is that rubric scores systematically fail to register clinically meaningful hallucinations—particularly errors that a rubric author could not have anticipated and encoded in advance (additive claims, fabricated citations, dosage/threshold errors). The authors solve the meta-evaluation problem cleverly: rather than trying to detect naturally-occurring hallucinations, they *inject* controlled, single, clinician-validated errors into otherwise-fixed correct responses, then measure whether the benchmark's own rubric+grader distinguishes the matched pair (paired AUROC). This isolates the effect of a known error while holding style/quality constant. Two durable artifacts result: a literature-grounded, clinician-validated 13-type taxonomy of medical hallucinations, and a quality-controlled adversarial injection pipeline.
Methodological Rigor. The design is unusually thorough for an evaluation-critique paper. The controlled MedHallu study establishes the mechanism (fact-specific criteria drive discrimination; generic criteria contribute little; RubricHub fact-check criteria are ~2× more discriminative), and this is corroborated across three frontier benchmarks with consistent error-type rankings (Kendall's τ_b ≈ 0.69 between HealthBench and LiveMedBench). Supporting analyses are strong: bootstrap CIs throughout, a grader-stability re-run (κ=0.765, no significant AUROC shifts), sensitivity across three injection models (highly consistent) and three answer models (revealing an interesting length/capability confound), programmatic + LLM + clinician triple validation of injections, and a 12-clinician review finding 80% of injected errors diagnosis/management-changing. The authors are candid about limitations: a single grader per main benchmark, a 500-question HealthBench subset, and—most importantly—reliance on *injected* rather than *naturally occurring* errors, which trades ecological validity for causal cleanliness. The MedHallu "context leakage" analysis (Appendix L), showing RubricHub barely beats an NLI baseline because knowledge fields leak the answer, is a sophisticated self-critique that strengthens credibility.
Potential Impact. Medical LLM evaluation is a rapidly growing area, and HealthBench-style rubric grading has become a de facto standard that increasingly feeds training signals (rubrics-as-rewards). Demonstrating that strong rubric scores provide only partial evidence of clinical reliability—and that HealthBench Professional performs near chance at distinguishing injected errors—is a consequential cautionary result that could reshape how benchmarks are built, interpreted, and used for safety claims. The taxonomy and injection pipeline are reusable building blocks for stress-testing any rubric-based benchmark. The retrieval-grounded (SAFE-based) complement points toward a concrete remedy, though it remains preliminary.
Timeliness & Relevance. Highly timely. Clinician trust, hallucinations, and the scaling of LLM-as-judge/rubric evaluation are all active bottlenecks. The paper directly addresses whether the field's leading evaluation methodology can support deployment claims—a question regulators, hospitals, and model developers are actively confronting.
Strengths & Limitations. Key strengths: causal, matched-pair design; extensive sensitivity/stability analyses; genuine clinician involvement; mechanistic explanation (rubrics catch only pre-specifiable, verifiable facts) rather than a bare empirical result; and intellectual honesty about confounds. Weaknesses: (1) injected errors may not mirror the distribution or subtlety of natural hallucinations; (2) the retrieval remedy has a 55% false-positive rate, limiting immediate utility; (3) single-grader main runs; (4) English-centric analysis for the clinical review; (5) inter-rater agreement is only moderate for subtler error types (evidence fabrication, overconfidence), which the authors handle transparently but which weakens claims for those categories. The core finding—rubrics miss errors they cannot anticipate—is also somewhat intuitive once stated, tempering surprise, though the *magnitude* (near-chance on Professional; fabricated citations never penalized despite explicit criteria) is striking.
Other observations. Reproducibility is aided by data/code links, exhaustive appendices (prompts, thresholds, model configs, compute costs of ~$760), and public benchmarks; however, reliance on many proprietary frontier models complicates exact replication. Resource intensity is modest (API-scale plus clinician volunteers). The work challenges an emerging load-bearing assumption—that rubric performance certifies clinical safety—giving it moderate refutation value within the medical-eval subfield.
Overall, this is a well-executed, timely critique with reusable artifacts and a credible mechanistic account, likely to be cited and built upon by the medical-AI-evaluation community, though its influence is bounded by the injected-error framing and the preliminary nature of the proposed fix.
Generated Sep 14, 2026
Rigorous, timely critique of the field's dominant medical-LLM evaluation method, with reusable taxonomy/pipeline and a credible mechanistic account, bounded by injected-error framing and a preliminary remedy.