PatchHolmes Uses Listwise Agents to Find the Right Vulnerability Fix
Introduction
A vulnerability record is far more useful when it can be linked to the exact commit that fixed the issue. That link supports patch verification, affected-version tracking, severity analysis, and software supply-chain scanning. Yet roughly 60% to 63% of CVEs in the GitHub Advisory Database and the National Vulnerability Database do not include a patch link. PatchHolmes addresses the task of recovering that missing link: identifying the commit that actually fixed a known vulnerability.
Why patch retrieval is difficult
The search space can be enormous. A repository may contain up to 1.4 million commits, and many vulnerabilities are associated with repositories containing more than 5,000 commits. Candidate diffs are also long, averaging about 15,000 tokens in the reported corpus. Encoders designed around short inputs may therefore miss the relevant change. In addition, vulnerability reports and commit messages often describe the same issue differently. A report may mention a buffer overflow, while the fixing commit refers to an out-of-bounds read.
A two-phase agentic design
PatchHolmes divides the task into two stages:
- Candidate retrieval: A hybrid retriever searches the local Git repository and produces a top-100 list.
- Listwise inspection: Instead of scoring each candidate independently, the second-stage agent sees the entire list. It then uses four budgeted tools to open only a small number of individual commits—between three and ten in the described design—before submitting one final answer.
This distinction matters because independent yes-or-no judgments can fail when several commits touch the same code path. If an agent labels many candidates as plausible, a pointwise system may have no reliable way to distinguish them and can end up selecting whichever candidate appears first. Listwise context gives the model an opportunity to compare intent, code changes, and surrounding evidence across candidates.
Reported results
On GitHubAD, PatchHolmes reaches 59.95% Recall@1, compared with 34.61% for the pointwise classifier Favia and 28.55% for the retrieve-and-reason baseline IRCoT. When the candidate set is held constant, the agent improves top-1 recall by 27.32 percentage points over simply taking the retriever’s first result. Transferred without modification to PatchFinder_top10’s candidate pool, it raises Recall@1 from PatchFinder’s own 24.28% top-1 selection to 39.86%.
The ablations also point to the workflow as the main source of the gain. Changing the language model within the Qwen family shifts Recall@1 by less than one point, while a second model family, gpt-oss, remains well above the no-agent baseline. For each vulnerability, the system uses about 96,000 input tokens, eight tool calls, and reads roughly five commits. It runs with a frozen open-weight model over a local Git clone, without fine-tuning or external search APIs.
Why it matters
PatchHolmes shows how an agent can turn a broad retrieval stage into a focused comparison process under a fixed inspection budget. For vulnerability databases with missing patch links, that could reduce manual triage and improve the foundations of patch tracking, vulnerability response, and supply-chain analysis.
The results do not eliminate the remaining challenges. Top-1 recall is still far from perfect, and performance depends on candidate quality, long-diff understanding, and bridging the vocabulary gap between security reports and code changes. Broader validation across repositories, programming languages, and operational security workflows will be important.
Source: Hugging Face Daily Papers
Comments
Checking sign-in status...
Loading comments...