Daily Paper

Evaluating Multi-Turn Multimodal Diagnostic Reasoning on Challenging Real-World Clinical Cases

The authors introduce ClinMM-Bench, a multi-turn multimodal benchmark of 1,089 clinical cases, to evaluate diagnostic accuracy and reasoning quality across 15 multimodal large language models.

arXiv:2607.25933 Empirical Study

Rui Yang, Weihao Xuan, Yi Lin, Zhuhan Bao et al.

multimodal-evaluationclinical-reasoningmulti-turn-interactionfailure-mode-analysisvisual-hallucination

1. Introduction: Moving Beyond Static Question-Answering

In a comprehensive evaluation of 15 Multimodal Large Language Models (MLLMs), Yang et al. identify a fundamental misalignment between current AI benchmarking and the dynamic requirements of clinical medicine. Most existing evaluations rely on static, single-turn question-answering formats that provide complete information at once. In contrast, the authors argue that “Clinical Reasoning” is inherently iterative, requiring the synthesis of evolving data, the revision of hypotheses, and the continuous refinement of logic as new multimodal evidence emerges.

The paper reports that while MLLMs can often identify “plausible” diagnostic directions, they frequently lack reasoning fidelity. The authors introduce ClinMM-Bench to simulate the progressive disclosure of patient information, revealing that a model’s ability to reach a surface-level diagnosis does not necessarily imply a grounded or valid reasoning chain. This document analyzes the study’s findings regarding diagnostic accuracy, failure mechanisms, and the implications for medical AI safety.

2. ClinMM-Bench: A High-Fidelity Simulation of Medical Practice

The authors developed ClinMM-Bench to move beyond routine medical records by utilizing 1,089 diagnostically challenging cases derived from published case reports. These cases are characterized by complex, multimodal presentations that require the integration of clinical text and 3,760 medical images.

The benchmark’s methodology and evaluation metrics include:

  • Specialty Coverage: The study spans eight medical specialties: Dermatology, Emergency Medicine, Internal Medicine, Nephrology, Neurology, Oncology, Ophthalmology, and Radiology (the largest segment at 62.35%).
  • Multi-Turn Interaction: Information is disclosed over an average of 5.45 dialogue rounds, forcing models to sustain memory and update hypotheses across turns.
  • Dual-LLM Consensus Mechanism: To ensure scoring rigor, the authors used GPT-5-medium and Claude-4.5-Sonnet as independent judges. These models compared predicted diagnoses against ground truth, assigning scores from 0 (incorrect) to 2 (completely correct).
  • Atomic Fact Decomposition: This novel metric breaks reasoning into “indivisible pieces of medical information” to quantify logic.
  • Reasoning Quality Metrics:
    • Fact Recall: Measures completeness by calculating the proportion of reference facts correctly identified.
    • Hallucination: Measures reliability by identifying the proportion of generated facts unsupported by clinical evidence.
    • Fact Density: Measures efficiency by assessing the proportion of valid atomic facts within the model’s total output.

3. The Accuracy-Reasoning Divergence: Key Findings

The results demonstrate a stark performance gap between proprietary and open-weight models. However, even the most capable systems remain far from clinical readiness, as indicated by the maximum possible diagnostic score of 2.0. The authors emphasize that a model may identify a “plausible” disease category (achieving a partial score) while the underlying reasoning chain is flawed or entirely hallucinated.

The following table summarizes the performance of the top three proprietary models (prioritizing completely correct rates as a proxy for precision) and the top three open-weight models.

RankModel CategoryModel NameDiagnostic Accuracy Score (Mean)Completely Correct Rate (%)
1ProprietaryGPT-5-medium1.14033.88%
2ProprietaryClaude-4.5-Sonnet1.00828.65%
3ProprietaryGemini 3 Pro1.03822.41%
4Open-WeightQwen3-VL-32B0.71811.20%
5Open-WeightLLaMA-4-Scout0.7138.82%
6Open-WeightMedGemma-27B0.6066.61%

4. Taxonomy of Error: The Five Failure Modes of MLLMs

The authors categorize diagnostic failures into five distinct modes, illustrating the multi-level capability deficits observed during progressive disclosure.

  • Information Synthesis Failure: The inability to integrate data across turns or modalities. In a case of smoking-related organizing pneumonia, a model over-relied on early imaging suggestive of malignancy and ignored subsequent biopsy findings and lesion regression, leading to a misdiagnosis of invasive adenocarcinoma.
  • Knowledge Mapping Error: Recognition of clinical facts but failure to link them to the correct disease mechanism. In tickborne encephalitis, one model correctly identified the fever and MRI findings but incorrectly mapped them to acute disseminated encephalomyelitis.
  • Perception Error: Misinterpretation of primary imaging findings. In a Chilaiditi syndrome case, a model misinterpreted colonic gas located between the liver and diaphragm as gas within the gallbladder, resulting in a misdiagnosis of emphysematous cholecystitis.
  • Premature Closure: A cognitive bias where the model commits to an early hypothesis and fails to revise it despite contradictory evidence. In a case of onychomatricoma, models persisted in a subungual melanoma diagnosis despite dermoscopic and MRI evidence supporting a benign condition.
  • Visual Hallucination: The fabrication of non-existent imaging findings. In a cerebral cavernoma case, a model hallucinated an “eccentric scolex” and a “dot sign”—specific markers for parasites—leading to a false diagnosis of neurocysticercosis.

5. Impact of Specialization and “Thinking” Modes

The paper’s analysis of model architectures revealed several counter-intuitive findings that challenge assumptions regarding domain adaptation:

  1. Diminishing Returns of Specialization: While MedGemma-4B showed significant improvement over its general-purpose counterpart, this benefit largely vanished at the 27B scale. In specialties like Radiology and Neurology, the 27B medical variant provided no consistent advantage.
  2. The Reasoning Trace Paradox: Extending reasoning traces (e.g., Qwen3-VL-thinking) did not consistently improve accuracy and often decreased fact recall. The authors suggest that without improved cross-turn memory or visual grounding, longer reasoning traces simply provide more space for a model to drift from the clinical facts.
  3. Specialty-Specific Vulnerability: Performance was highly uneven. Neurology (0.949 [95% CI: 0.880, 1.018]) and Ophthalmology (0.931 [95% CI: 0.797, 1.067]) were the highest-scoring specialties. Conversely, Internal Medicine (0.665 [95% CI: 0.490, 0.851]) and Radiology (0.646 [95% CI: 0.613, 0.680]) proved most difficult, likely due to the higher requirement for synthesizing diverse, conflicting evidence.

6. Conclusion: Implications for AI Safety and Clinical Trust

The findings from ClinMM-Bench suggest that current MLLMs lack the “reasoning fidelity” required for autonomous clinical use. The primary bottleneck is not a lack of medical knowledge, but rather a failure in dynamic hypothesis updating and visual-grounding reliability.

Takeaways for the AI Safety Community:

  • Covert Reasoning Failures: Models frequently reach “partially correct” diagnoses using hallucinated or logically disconnected facts. Relying on accuracy alone as a safety metric is insufficient for clinical applications.
  • Cognitive Bias in LLMs: The phenomenon of “premature closure” confirms that MLLMs are susceptible to human-like cognitive biases, specifically the over-weighting of initial information at the expense of corrective, subsequent data.
  • The Grounding Bottleneck: Safety failures in medical AI are often rooted in the initial perception of multimodal data (e.g., hallucinating a “dot sign”) rather than linguistic reasoning flaws.
  • Failure of Hypothesis Revision: For clinical safety, models must develop robust mechanisms to discard an initial theory when new, contradictory data (such as a biopsy or MRI) is presented. Current architectures tend to force-fit new data into old hypotheses.

Read the full paper on arXiv · PDF