CCS selects the best radiology report from multiple MLLM candidates by measuring clinical consensus with combined text and multimodal embedding utilities, yielding gains over single-path and Best-of-N baselines on clinical metrics across three datasets.
C.; Schwaighofer, A.; Lungren, M
5 Pith papers cite this work. Polarity classification is still indexing.
citation-role summary
citation-polarity summary
roles
baseline 1polarities
baseline 1representative citing papers
Replacing the generic Stable Diffusion VAE with domain-specific MedVAE pretrained on 1.6M medical images improves diffusion-based SR PSNR by 2.91-3.29 dB on knee/brain MRI and chest X-ray, with gains in fine details and VAE quality predicting SR performance (R²=0.67).
RA-RRG extracts key phrases with LLMs, retrieves them via multimodal similarity, and conditions report generation on them to achieve SOTA CheXbert scores and competitive RadGraph F1 on MIMIC-CXR and IU X-ray while supporting multi-view inputs.
M4CXR is a multi-modal large language model that performs multiple tasks in chest X-ray analysis including report generation with claimed SOTA clinical accuracy using chain-of-thought prompting.
Partial adaptation of DINO-family ViTs via unfreezing the last two blocks yields superior patient-level TMJ OA classification on CBCT compared to frozen or supervised baselines.
citing papers explorer
-
CCS: Clinical Consensus Selection for Radiology Report Generation
CCS selects the best radiology report from multiple MLLM candidates by measuring clinical consensus with combined text and multimodal embedding utilities, yielding gains over single-path and Best-of-N baselines on clinical metrics across three datasets.
-
Domain-Specific Latent Representations Improve the Fidelity of Diffusion-Based Medical Image Super-Resolution
Replacing the generic Stable Diffusion VAE with domain-specific MedVAE pretrained on 1.6M medical images improves diffusion-based SR PSNR by 2.91-3.29 dB on knee/brain MRI and chest X-ray, with gains in fine details and VAE quality predicting SR performance (R²=0.67).
-
RA-RRG: Multimodal Retrieval-Augmented Radiology Report Generation with Key Phrase Extraction
RA-RRG extracts key phrases with LLMs, retrieves them via multimodal similarity, and conditions report generation on them to achieve SOTA CheXbert scores and competitive RadGraph F1 on MIMIC-CXR and IU X-ray while supporting multi-view inputs.
-
M4CXR: Exploring Multi-task Potentials of Multi-modal Large Language Models for Chest X-ray Interpretation
M4CXR is a multi-modal large language model that performs multiple tasks in chest X-ray analysis including report generation with claimed SOTA clinical accuracy using chain-of-thought prompting.
-
Self-Supervised Vision Transformers for CBCT-Based Detection of Temporomandibular Joint Osteoarthritis
Partial adaptation of DINO-family ViTs via unfreezing the last two blocks yields superior patient-level TMJ OA classification on CBCT compared to frozen or supervised baselines.