about
LLM Judges Verify Presence, Not Absence: Omission Blindness in AI Clinical Notes (arxiv.org)
20 points by sbulaev 31 days ago | hide | past | pdf | 14 comments on HN

In plain words: AI-written doctor notes often leave out facts, and a second AI comparing note to recording barely spots missing ones, only added or altered ones. Listing the facts the recording establishes, then checking each, catches 36.9% of omissions where usual checks catch almost none.

Abstract · LLM Judges Verify Presence, Not Absence: Omission Blindness in AI Clinical Notes and What Recovers It

Ambient AI scribes draft clinical notes, and published audits find their dominant error is omission: information the encounter established that the note fails to record. The standard check is an LLM judge: a second model reads the note against the transcript and flags problems. We ask whether judges detect omissions. Public corpora cannot supply the answer key: their clinician reference notes and transcripts are materially discrepant. Our benchmark has 500 single-error note pairs from audited fact sheets, 298 with a named fact certainly absent and 202 added-or-altered controls. Across eight judge designs, paired discrimination (the flawed note below its clean twin, 0.5 a coin flip) reads 0.79-0.94 on added or altered content and 0.50-0.63 on omissions. On single notes, no design flags omissions reliably more often than perfect notes. Wording changes, voting and GEPA prompt optimisation move the operating point without creating usable detection. Restructuring the task recovers it: list the facts the transcript establishes, then check the note for each. Two methods reach it independently and trade off: a per-fact pipeline, and a GEPA-evolved prompt doing the same in one call. The pipeline's flags name the missing fact and its severity at 2.7% false alarms. The single call detects more (36.9% against 24.6%, p=0.002) at 6.2% false alarms and a tenth of the cost per note. A physician author validated 70 items and, where the two routes disagree, sided with the pipeline on 10 of 10 (p=0.002). A second clinician, not an author, graded the severity rubric blind and agrees to within a grade. On real vendor notes from a companion census no benchmark threshold transfers, but the re-calibrated single call detects more than the best of the eight at half its false-alarm rate. Omissions whose fact is restated elsewhere defeat both routes. We release the benchmark, prompts and judgements.

Sebastian Fox, Luke Markham, Ryan Lail, Michael Karotsieris
arXiv:2608.31016 · cs.CL, cs.AI · submitted Aug 31, 2026
abstract · pdf · html · 97 pages, 8 figures. Dataset: https://huggingface.co/datasets/ComposoAI/OmissionBench . Code: https://github.com/composo-ai/omission-bench . Companion paper: "One note in three: a verified census of three deployed AI scribes, and the instrument that counted it"

add comment on HN

The fact that even the abstract is very clearly AI-generated does not fill me with confidence in the rigour of the research.
"Verify Presence, Not Absence" is very LLM-coded all on its own.
I understand the answer to the question I'm about to ask, but how does any human read that abstract and not think, "This is entirely too many colons."

EDIT: And Pangram agrees that the abstract is 100% AI generated.

While I agree I have to say that AI-detectors are complete bullshit
Max from Pangram has done a bunch of interviews recently. I'd give them a listen; he might change your mind.
Interesting. My initial thought was “this is so horribly written it must be human”
The entire paper is almost certainly Claude-generated, given the clear Claude-speak throughout. It seems like they didn’t even try to have Claude write a polished paper: it reads like Claude writing up a lengthy report, down to the needless sectioning with idiosyncratic title language. There appears to be no disclosure, and the author contribution section appears to falsely claim that a particular author wrote the text.

I know arXiv has taken some measures to combat spam like this, but it seems like they’ll need to do more. There’s just very little barrier now to creating giant slop papers like this and then dumping them anywhere that won’t reject them. It is an insult to everyone’s time, and I can’t imagine they expect people to actually read this. If the expectation is that everyone will use an LLM to interpret it, then maybe they should have at least had a few more rounds of tightening and polishing the paper, even via LLM, to save the redundant token use.

I had to use an AI to undertand the abstract made by AI…
Not sure about the paper but the results make sense, we see it in PDF extraction too. Fields that aren't in the document are being made up 11 to 40% of the time depending on the API.

The whole thing is nasty partly because it isn't just AI problem. When checking the human labels that we used in our evals 40 out of 142 answer keys claiming absence were wrong. Tricky one.

> Fields that aren't in the document are being made up 11 to 40% of the time depending on the API.

Do you have an example?

I'm in healthcare and ambient documentation is obviously a huge thing now but I don't have any experience with it. We have anywhere from 5-10 companies reach out a week trying to sell us on their product and the demos are mostly okay (though you can tell they're rely on the happy path through a lot of it), but we haven't actually pulled the trigger on anything. Thanks in advance.

Sure. Here's two good examples from the published outputs. Receipts where there's no subtotal line, only a total. One model returns a subtotal anyway, "63.000", "88.000", numbers it copied or computed from elsewhere in the document. Ours invented a 2% discount on a receipt that has no discount line.

TV ad contracts where the station's address isn't on the contract. Several models fill it with a real address that is on the page, just the agency's or the advertiser's. Given your industry this pattern should worry you most, since these aren't made up from nothing. Just the model tripping and taking it from the wrong place, which might look plausible on the surface.

Here's what I'd suggest for your vendor demos. Run them on a handful of your own documents where you know a field is genuinely absent, and count how many come back filled. Their own samples won't tell you that.

All the raw outputs are public, so you can check the examples above yourself. My reply to you got filtered so I'm omitting the link. If you want to read the writeup, search for "velrim" and "fabrication on absent fields" for the article plus the repo.

Sure. Here's two good examples from the published outputs.

Receipts where there's no subtotal line, only a total. One model returns a subtotal anyway, "63.000", "88.000", numbers it copied or computed from elsewhere in the document. Ours invented a 2% discount on a receipt that has no discount line.

TV ad contracts where the station's address isn't on the contract. Several models fill it with a real address that is on the page, just the agency's or the advertiser's. Given your industry this pattern should worry you most, since these aren't made up from nothing. Just the model tripping and taking it from the wrong place, which might look plausible on the surface.

Here's what I'd suggest for your vendor demos. Run them on a handful of your own documents where you know a field is genuinely absent, and count how many come back filled. Their own samples won't tell you that.

All the raw outputs are here if you want to look: https://velrim.com/research/fabrication-on-absent-fields

May be AI written. Used my AI to get the gist of it. One thing that I believe is LLMs need to be considered a pure play tool and can be guided to point out misses as well. The reason is simple. To find whats missing, one needs to ask the right questions on how to judge and when that is there.
"the floored note before its clean twin"

People using AI like this should be run out of their jobs.