Pith. sign in

REVIEW 3 cited by

Towards a Holistic Framework for Multimodal Large Language Models in Three-dimensional Brain CT Report Generation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.02235 v1 pith:SWLARZKC submitted 2024-07-02 cs.CL

classification cs.CL
keywords modelsradiologybraingptreportbraindatasetevaluationlanguage
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Multi-modal large language models (MLLMs) have been given free rein to explore exciting medical applications with a primary focus on radiology report generation. Nevertheless, the preliminary success in 2D radiology captioning is incompetent to reflect the real-world diagnostic challenge in the volumetric 3D anatomy. To mitigate three crucial limitation aspects in the existing literature, including (1) data complexity, (2) model capacity, and (3) evaluation metric fidelity, we collected an 18,885 text-scan pairs 3D-BrainCT dataset and applied clinical visual instruction tuning (CVIT) to train BrainGPT models to generate radiology-adherent 3D brain CT reports. Statistically, our BrainGPT scored BLEU-1 = 44.35, BLEU-4 = 20.38, METEOR = 30.13, ROUGE-L = 47.6, and CIDEr-R = 211.77 during internal testing and demonstrated an accuracy of 0.91 in captioning midline shifts on the external validation CQ500 dataset. By further inspecting the captioned report, we reported that the traditional metrics appeared to measure only the surface text similarity and failed to gauge the information density of the diagnostic purpose. To close this gap, we proposed a novel Feature-Oriented Radiology Task Evaluation (FORTE) to estimate the report's clinical relevance (lesion feature and landmarks). Notably, the BrainGPT model scored an average FORTE F1-score of 0.71 (degree=0.661; landmark=0.706; feature=0.693; impression=0.779). To demonstrate that BrainGPT models possess objective readiness to generate human-like radiology reports, we conducted a Turing test that enrolled 11 physician evaluators, and around 74% of the BrainGPT-generated captions were indistinguishable from those written by humans. Our work embodies a holistic framework that showcased the first-hand experience of curating a 3D brain CT dataset, fine-tuning anatomy-sensible language models, and proposing robust radiology evaluation metrics.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. MAIRA-Seg: Enhancing Radiology Report Generation with Segmentation-Aware Multimodal Large Language Models

    cs.CV 2024-11 conditional novelty 6.0 of 10

    MAIRA-Seg shows that feeding segmentation mask tokens alongside chest X-rays improves mask-relevant clinical metrics for MLLM radiology report generation, but not lexical overlap.

  2. Large Language Model with Region-guided Referring and Grounding for CT Report Generation

    cs.CV 2024-11 conditional novelty 5.0 of 10

    Reg2RG grounds each generated CT report section in a specific anatomical region by combining decoupled local texture and geometry features with global volume features in an LLM, outperforming prior methods on two ches...

  3. Multimodal Large Language Models for Medicine: A Comprehensive Survey

    cs.LG 2025-04 conditional novelty 2.0 of 10

    A comprehensive review cataloging medical MLLMs, their uses in report generation, diagnosis, and treatment, and the challenges of accuracy, hallucination, fairness, privacy, and deployment.

Pith tools