REVIEW 5 cited by
Improving Factual Completeness and Consistency of Image-to-Text Radiology Report Generation
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Neural image-to-text radiology report generation systems offer the potential to improve radiology reporting by reducing the repetitive process of report drafting and identifying possible medical errors. However, existing report generation systems, despite achieving high performances on natural language generation metrics such as CIDEr or BLEU, still suffer from incomplete and inconsistent generations. Here we introduce two new simple rewards to encourage the generation of factually complete and consistent radiology reports: one that encourages the system to generate radiology domain entities consistent with the reference, and one that uses natural language inference to encourage these entities to be described in inferentially consistent ways. We combine these with the novel use of an existing semantic equivalence metric (BERTScore). We further propose a report generation system that optimizes these rewards via reinforcement learning. On two open radiology report datasets, our system substantially improved the F1 score of a clinical information extraction performance by +22.1 (Delta +63.9%). We further show via a human evaluation and a qualitative analysis that our system leads to generations that are more factually complete and consistent compared to the baselines.
Forward citations
Cited by 5 Pith papers
-
Historical Report Guided Bi-modal Concurrent Learning for Pathology Report Generation
BiGen improves pathology report generation by retrieving relevant sentences from a training report bank and jointly compressing visual and textual knowledge with shared cross-attention tokens.
-
Efficient Medical Vision-Language Alignment Through Adapting Masked Vision Models
ALTA adapts a frozen masked-pretrained X-ray encoder to language with 8% trainable parameters and temporal-multiview inputs, improving medical retrieval and zero-shot classification.
-
Automated Radiology Report Generation Based on Topic-Keyword Semantic Guidance
A topic-keyword semantic guidance framework improves automated radiology report generation and reaches state-of-the-art on two public chest X-ray benchmarks.
-
Holistic Artificial Intelligence in Medicine; improved performance and explainability
An extension of the HAIM multimodal framework that uses LLM-based retrieval and summarization to improve clinical prediction AUC from 79.9% to 90.3% and to generate document-grounded explanations.
-
Enhancing Biomedical Multi-modal Representation Learning with Multi-scale Pre-training and Perturbed Report Discrimination
A biomedical vision-language model is pre-trained with a contrastive loss that distinguishes original radiology reports from nine perturbed variants, plus a local attention loss, and is reported to beat ConVIRT and GL...
Discussion (0). Sign in to comment.