OpenTumorBoard Tests Whether AI Can Reason With a Cancer Care Team
Introduction
A multidisciplinary tumor board is more than a medical question-and-answer session. Imaging, pathology, treatment history, patient context, and surgical considerations are discussed by specialists who may support, challenge, or refine one another before arriving at a plan. OpenTumorBoard focuses on this collaborative trajectory rather than evaluating a model only on a standalone diagnosis or recommendation.
What the benchmark contains
The authors transcribed 219 publicly available tumor board meetings published on YouTube. The resulting resource includes 611 patient cases, 19,157 discussion turns, and contributions from ten specialist roles. The recordings represent 12,534 minutes of discussion, with 16,215 questions directed to specialists during the meetings. The project also provides a dataset, leaderboard, code, and an automated curation pipeline for future research.
The benchmark has two settings:
- SPECIALIST TURN asks a model to answer a clinically significant question that was actually posed to a specialist during a meeting.
- BOARD SIMULATION provides a case summary and slides, then asks the model to generate a back-and-forth discussion and reach conclusions about therapy, surgery, next steps, or clinical-trial matching.
What the results reveal
The study evaluates 14 general-purpose frontier and medical language models. The strongest model reaches 3.43 out of 5 for clinical equivalence to specialist answers, while alignment with the conclusions recorded by the tumor boards reaches only 2.78 out of 5. These results suggest that models can capture pieces of the medical record without reliably understanding how specialists combine evidence, resolve disagreement, and narrow the space of clinical options.
That distinction matters. A model may produce a plausible answer in a single turn while failing to preserve the longitudinal context of a patient or the division of responsibility across specialties. Board Simulation therefore probes a different capability: sustained, role-aware reasoning that ends in an actionable consensus rather than a fluent paragraph.
The authors also report improvements after supervised fine-tuning and reinforcement learning on a held-out test setting. This indicates that real discussion trajectories may be useful not only for evaluation, but also for adapting models to collaborative clinical workflows.
Significance and limitations
Three M.D. reviewers examined a subset of the benchmark and found high coverage and factuality in the patient cases, along with strong fidelity in the extracted consensus conclusions. That review supports the dataset’s research value, although it does not eliminate important limitations. Publicly recorded meetings may not represent every institution or clinical culture, and transcription, summarization, and consensus extraction can all influence the final benchmark.
OpenTumorBoard should therefore be understood as a research and stress-testing resource, not as a substitute for clinicians or a validated clinical decision system. Its broader contribution is to shift the evaluation question from “What medical facts can a model recall?” to “Can it participate reliably in a complex expert team?” As medical AI moves toward multi-step workflows and agentic systems, benchmarks built from real discussion trajectories could become increasingly important for measuring collaboration, consistency, and safety.
Source: Hugging Face Daily Papers
Comments
Checking sign-in status...
Loading comments...