Why Retrieval-Based Medical Fact Checking Breaks on Open-Ended Answers
Introduction
Retrieval-augmented factuality checking has become an attractive way to evaluate medical answers. Instead of asking a language model to judge a response from memory, a system retrieves passages from medical literature or other authoritative sources and then asks a verifier to determine whether each claim is supported. The approach appears more transparent and easier to scale, but a new study shows that good benchmark scores can conceal serious weaknesses.
The central findings
The researchers focus on the open-ended MedExpert dataset and compare it with three closed-ended datasets. They test four retrieval methods, six verifier models, and three corpus settings. The results reveal a sharp gap between benchmark types. A top score on the closed biomedical benchmarks reaches 0.78 F1, while the strongest tested pipeline—Qwen3 as retriever and GPT-5.4 as verifier—achieves only 0.06 F1 on MedExpert, which contains clinician-annotated open-ended medical answers.
The first problem is that retrieval is not the same as evidence. Across the tested retrievers, only 18.5% to 42% of returned passages directly affirm or contradict the claim being checked. Many passages are merely related to the topic. Others concern a different patient population, clinical condition, or level of evidence. Passing such material to a verifier creates a difficult reasoning problem before verification even begins.
The study therefore separates retrieval quality from verifier reasoning. It identifies five dimensions of retrieval-stage quality and six consecutive steps in verification: selecting evidence, interpreting it, making the clinical inference, grounding the inference in the evidence, calibrating confidence, and keeping the final label consistent. A model may fail at any one of these stages. It may retrieve a relevant-looking passage but miss a population mismatch, or turn general medical knowledge into a claim that the cited source does not actually establish.
Why scaling does not solve the problem
An LLM-assisted pattern-induction pipeline was used to analyze 800 claim–evidence pairs and 576 reasoning traces, with human guidance and clinician adjudication. The experiments show that increasing reasoning effort consumes substantially more tokens without improving recall in the tested settings. Larger models, medical fine-tuning, and expanded authoritative web corpora also fail to eliminate the dominant error patterns.
This suggests that the limitation is not simply a shortage of model capacity. Open-ended medical answers contain many fine-grained claims, each with its own population, conditions, evidence strength, and clinical context. A retrieve-then-verify pipeline must align all of these elements, but an aggregate F1 score cannot reveal whether a failure came from missing evidence, misreading evidence, weak clinical inference, or overconfident labeling.
Implications
For teams building medical RAG systems or LLM-as-judge evaluators, the main lesson is methodological. Accuracy, precision, recall, and F1 should be accompanied by diagnostics that ask whether retrieved passages truly support the claim, address the correct population, and come from an appropriate evidence source. Verification traces should also be inspected step by step.
The study does not show that retrieval is useless. Rather, it shows that the presence of a citation is not enough. Future systems may need finer claim decomposition, clinically aware evidence matching, explicit uncertainty handling, and evaluation protocols designed for open-ended answers. In high-stakes medicine, identifying why a checker failed can be more valuable than reporting a single impressive score.
Source: Hugging Face Daily Papers
Comments
Checking sign-in status...
Loading comments...