REVIEW 9 cited by
GREEN: Generative Radiology Report Evaluation and Error Notation
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Evaluating radiology reports is a challenging problem as factual correctness is extremely important due to the need for accurate medical communication about medical images. Existing automatic evaluation metrics either suffer from failing to consider factual correctness (e.g., BLEU and ROUGE) or are limited in their interpretability (e.g., F1CheXpert and F1RadGraph). In this paper, we introduce GREEN (Generative Radiology Report Evaluation and Error Notation), a radiology report generation metric that leverages the natural language understanding of language models to identify and explain clinically significant errors in candidate reports, both quantitatively and qualitatively. Compared to current metrics, GREEN offers: 1) a score aligned with expert preferences, 2) human interpretable explanations of clinically significant errors, enabling feedback loops with end-users, and 3) a lightweight open-source method that reaches the performance of commercial counterparts. We validate our GREEN metric by comparing it to GPT-4, as well as to error counts of 6 experts and preferences of 2 experts. Our method demonstrates not only higher correlation with expert error counts, but simultaneously higher alignment with expert preferences when compared to previous approaches.
Forward citations
Cited by 9 Pith papers
-
SliceWorld: A Predictive and Controllable World-State Model for CT Report Generation
SliceWorld introduces a world-state model for CT report generation that uses predictive and factor-aware objectives on axial slice sequences.
-
CARE-X: Towards Clinically Useful Radiology VLMs with Auxiliary Supervision, Reward-Aligned Learning, and Tool-Augmented Measurement
A chest X-ray VLM co-trained with classification and grounding heads, tuned with DAPO reinforcement learning, and augmented with deterministic measurement tools outperforms prior radiology VLMs on report generation, V...
-
ORCA: ORgan-Centroid Aggregation for Training-Free 3D CT Visual Token Compression
ORCA compresses 3D CT tokens into organ-guided connected regions with sinusoidal centroid encoding, outperforming grid average and other compressors at matched budgets.
-
Deterministic Hallucination Detection in Medical VQA via Confidence-Evidence Bayesian Gain
A fixed, sampling-free score — per-token log-probability variance times (1 + average |image-vs-text probability shift|) — detects medical-VQA hallucinations better than semantic-entropy baselines in 13 of 16 settings.
-
Exploring the Capabilities of Large Language Model Encoders for Image-Text Retrieval in Chest X-rays
Domain-adapted LLM encoders trained with masked token prediction and supervised contrastive learning improve chest X-ray image-text retrieval and external generalization, reaching GREEN scores of 0.308 on MIMIC-CXR an...
-
Boosting Vision Semantic Density with Anatomy Normality Modeling for Medical Vision-language Pre-training
On abdominal CT, a vision-language pre-training method using organ-level normal/abnormal contrastive learning and a VQ-VAE normality model achieves 84.9% average zero-shot AUC, beating prior methods by 3.6%.
-
MedErr-CT: A Visual Question Answering Benchmark for Identifying and Correcting Errors in CT Reports
The authors introduce a 3D CT-based visual question answering benchmark with six error types and three task levels, and show that current 3D medical MLLMs perform poorly on it.
-
Disease-Centric Vision-Language Pretraining with Hybrid Visual Encoding for 3D Computed Tomography
A hybrid visual encoder plus disease-centric contrastive pretraining yields SOTA AUC scores on CT-RATE and Rad-ChestCT 3D CT disease classification benchmarks.
-
Region-Aware Multimodal Large Language Model via SlowFast Tokenization and Pseudo-Mask Guidance for 3D CT Report Generation
MedRegion-CT integrates region-representative tokens, mask-driven segmentation tokens, and patient-specific attribute prompts into a multimodal LLM, reporting state-of-the-art scores on RadGenome-Chest CT report generation.
Discussion (0). Sign in to comment.