about
KnowBench: Evaluating clinical AI with effort reduction (arxiv.org)
5 points by kangjl888 18 days ago | hide | past | pdf | 2 comments on HN

In plain words: A benchmark scores clinical AI by how much doctor work it removes: the share of generated notes, codes, and orders a clinician signs off unchanged. Unlike tests that compare text to a reference, it found 97.99% of work accepted in real clinics.

Abstract · KnowBench: Effort Reduction as a Unified, Deployment-Grounded Benchmark for Clinical AI

Clinical AI systems are evaluated with instruments built for research settings (reference-based similarity metrics and expert rubric panels) that measure resemblance to an artifact rather than reduction of a burden. We introduce KnowBench, pioneered by Knowtex, whose unifying metric is Effort Reduction (ER): the proportion of system-generated clinical work product accepted by the responsible clinician under expert and safety review. ER is defined once and instantiated per task across the administrative workload clinical AI automates: visit notes, diagnosis and billing codes, orders, EHR chart summarization, patient after-visit summaries, and clinical decision support. In every instantiation the construction is identical: the clinician's review-and-attestation event is the ground truth, every accepted unit is work the system completed, and every correction is residual effort returned to the clinician. The primary contribution of this paper is the benchmark itself: the metric, its degenerate cases, and a reporting protocol under which ER claims are auditable and cross-system comparable. Alongside it we report an initial headline measurement from the documentation instantiation: over one million signed encounters across a production window exceeding six months and thirteen medical specialties, Knowtex's proprietary fine-tuned clinical foundation models operating inside a closed feedback architecture achieve an aggregate ER of 97.99%, with per-specialty aggregates spanning 96.8-98.9%. This release reports the protocol's checklist partially, and states which companion statistics are withheld; the benchmark is offered so that this figure, and every figure reported after it, can be held to the same standard.

Jocelyn Kang, Caroline Zhang
arXiv:2609.15794 · cs.AI · submitted Sep 14, 2026
abstract · pdf · html · 7 pages, 1 table. Contact: [email protected]

add comment on HN

Knowtex (YC S22) is the first frontier AI Lab for healthcare. We build AI that automates clinical work and provides clinical intelligence to providers. Standard evals in this space use medical exams or use expert rubric panels. These do not measure whether a clinician's workload actually was decreased.

Our observation is that this measure already exists in the workflow. Before a note enters the medical record, a clinician reviews the draft and fixes whatever is wrong, then signs and owns it legally. The edits represents the measurement in that whatever survives review is effort Knowtex reduced, and whatever got edited by the clinician is effort returned.

We formalized this as our benchmark KnowBench measuring Effort Reduction: the fraction of generated content accepted without content edits. The same construction works for billing codes, orders, and chart summaries, not just notes. Our arXiv paper defines the metric, its failure modes (you can game it with short drafts, clinicians can rubber-stamp, etc.), and a six-item reporting checklist.

Our result showed 97.99% aggregate across 1M+ signed encounters and 13 specialties.

Closest prior art is HTER from machine translation and Copilot's acceptance rate, with one difference: our reviewer signs the output into a legal record.

Paper: https://arxiv.org/abs/2609.15794

Its more imp now than ever to have mechanisms to build trust in AI. Glad to see Knowtex is doing great work in this space!