Back to articles
Evaluation & Benchmarks

OpenTumorBoard Tests Whether AI Can Reason With a Cancer Care Team

3 min read

Introduction

A multidisciplinary tumor board is more than a medical question-and-answer session. Imaging, pathology, treatment history, patient context, and surgical considerations are discussed by specialists who may support, challenge, or refine one another before arriving at a plan. OpenTumorBoard focuses on this collaborative trajectory rather than evaluating a model only on a standalone diagnosis or recommendation.

What the benchmark contains

The authors transcribed 219 publicly available tumor board meetings published on YouTube. The resulting resource includes 611 patient cases, 19,157 discussion turns, and contributions from ten specialist roles. The recordings represent 12,534 minutes of discussion, with 16,215 questions directed to specialists during the meetings. The project also provides a dataset, leaderboard, code, and an automated curation pipeline for future research.

The benchmark has two settings:

  • SPECIALIST TURN asks a model to answer a clinically significant question that was actually posed to a specialist during a meeting.
  • BOARD SIMULATION provides a case summary and slides, then asks the model to generate a back-and-forth discussion and reach conclusions about therapy, surgery, next steps, or clinical-trial matching.

What the results reveal

The study evaluates 14 general-purpose frontier and medical language models. The strongest model reaches 3.43 out of 5 for clinical equivalence to specialist answers, while alignment with the conclusions recorded by the tumor boards reaches only 2.78 out of 5. These results suggest that models can capture pieces of the medical record without reliably understanding how specialists combine evidence, resolve disagreement, and narrow the space of clinical options.

That distinction matters. A model may produce a plausible answer in a single turn while failing to preserve the longitudinal context of a patient or the division of responsibility across specialties. Board Simulation therefore probes a different capability: sustained, role-aware reasoning that ends in an actionable consensus rather than a fluent paragraph.

The authors also report improvements after supervised fine-tuning and reinforcement learning on a held-out test setting. This indicates that real discussion trajectories may be useful not only for evaluation, but also for adapting models to collaborative clinical workflows.

Significance and limitations

Three M.D. reviewers examined a subset of the benchmark and found high coverage and factuality in the patient cases, along with strong fidelity in the extracted consensus conclusions. That review supports the dataset’s research value, although it does not eliminate important limitations. Publicly recorded meetings may not represent every institution or clinical culture, and transcription, summarization, and consensus extraction can all influence the final benchmark.

OpenTumorBoard should therefore be understood as a research and stress-testing resource, not as a substitute for clinicians or a validated clinical decision system. Its broader contribution is to shift the evaluation question from “What medical facts can a model recall?” to “Can it participate reliably in a complex expert team?” As medical AI moves toward multi-step workflows and agentic systems, benchmarks built from real discussion trajectories could become increasingly important for measuring collaboration, consistency, and safety.

Source: Hugging Face Daily Papers

Comments

Checking sign-in status...

Loading comments...

Related articles

CCTest · Blog
PhysVista Tests VLM Physical Intelligence Through a Perception–Reasoning–Assessment Loop
Evaluation & Benchmarks
cctest.ai

PhysVista Tests VLM Physical Intelligence Through a Perception–Reasoning–Assessment Loop

PhysVista introduces a benchmark that evaluates whether vision-language models understand physical consistency rather than merely recognizing visual content. It combines physical state perception, dynamics reasoning, and plausibility assessment across real-world and AI-generated videos.

Read more