about
CURE: A Dataset for Clinical Understanding and Retrieval Evaluation (arxiv.org)
1 point by fzliu on Jun 17, 2025 | hide | past | pdf | discuss on HN

In plain words: Built with doctors, it holds 2,000 search questions across 10 medical areas, in English and French/Spanish answered by English passages, to test how well search systems rank useful text for doctors at the bedside. Baseline runs show it can tell systems apart on these questions.

Abstract · CURE: A Dataset for Clinical Understanding & Retrieval Evaluation

Given the dominance of dense retrievers that do not generalize well beyond their training dataset distributions, domain-specific test sets are essential in evaluating retrieval. There are few test datasets for retrieval systems intended for use by healthcare providers in a point-of-care setting. To fill this gap we have collaborated with medical professionals to create CURE, an ad-hoc retrieval test dataset for passage ranking with 2000 queries spanning 10 medical domains with a monolingual (English) and two cross-lingual (French/Spanish -> English) conditions. In this paper, we describe how CURE was constructed and provide baseline results to showcase its effectiveness as an evaluation tool. CURE is published with a Creative Commons Attribution Non Commercial 4.0 license and can be accessed on Hugging Face and as a retrieval task on MTEB.

Nadia Athar Sheikh, Daniel Buades Marcos, Anne-Laure Jousse, Akintunde Oladipo, Olivier Rousseau, Jimmy Lin
arXiv:2412.06954 · cs.IR · submitted Dec 9, 2024 · updated Jun 27, 2025
abstract · pdf · html

add comment on HN