about
IatroBench: Pre-Registered Evidence of Iatrogenic Harm from AI Safety Measures (arxiv.org)
2 points by NavinF 148 days ago | hide | past | pdf | 1 comment on HN

In plain words: Sixty clinical questions were each asked twice—once as a patient and once as a doctor—to check whether AI models hide information from patients. Models withheld more from patients, with an average gap of +0.22 on a 0–1 harm scale.

Abstract · IatroBench: A Pre-Registered Benchmark of Clinical Omission in Language Models

We introduce IatroBench, a benchmark with two axes of harm (commission and omission), comprising 60 pre-registered clinical scenarios, tested on 6 models. Matched scenarios are framed as a patient query and a doctor consultation, differing in register and request (with the implication of supervision by a treating physician in the latter). We analyse the responses of five different models and find that all share more information in the doctor framing than the patient framing (which we call "framing-contingent withholding"). For example, a model with strong safety training provides a benzodiazepine tapering schedule to a doctor, but does not provide this schedule to a patient who requests it. We use Claude Opus 4.6 for structured evaluation, and Gemini 3 Flash as our primary judge, to score model responses against a physician's rubrics. Our primary judge agrees with physicians' omission scores about as well as physicians agree with each other. We find a decoupling gap of +0.38 (p = 0.003) on average across models. With our primary judge (checked by physicians) the decoupling gap is +0.22 (95% CI 0.10-0.36, p = 0.0014). We find three distinct patterns underlying this gap, exemplified by each of the models below. In the doctor framing, Claude Opus demonstrates that it has the information, and withholds it in the patient framing. Llama 4 performs poorly in both framings, meaning the decoupling gap cannot distinguish between withholding and incompetence. Finally, GPT-5.2 (excluded from this analysis) failed to return text for 33.2% of doctor responses, compared to 0% of layperson responses. In 86.6% of cases that we score (through our structured evaluation) as having omission harms, our primary judge (Gemini 3 Flash) scores zero omission harm. Because our scenarios are designed to pit safety against helpfulness, these statistics hold only for this distribution.

David Gringras
arXiv:2604.07709 · cs.AI, cs.CL, cs.CY, cs.LG · submitted Apr 9, 2026 · updated Sep 30, 2026
abstract · pdf · html · 28 pages, 3 figures, 15 tables. Pre-registered on OSF (DOI: https://doi.org/10.17605/OSF.IO/G6VMZ). Code and derived results: https://github.com/davidgringras/iatrobench. v6 completes the revision begun in v5: physician validation reported against the primary judge; pair-by-model cluster tests added; examples, rubrics and reference excerpts moved to ancillary files; Figure 1 redrawn

add comment on HN
Also discussed: Jun 2026 (3 points, 0 comments)

Navin, don't leave a comment 'selling' the content. It's a good way to get people to assume it's a spam submission. Best to delete this and re-submit.