Pith. sign in

REVIEW 11 cited by

RadGenome-Chest CT: A Grounded Vision-Language Dataset for Chest CT Analysis

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2404.16754 v1 pith:DGGV7M3W submitted 2024-04-25 cs.CV

classification cs.CV
keywords segmentationgroundeddatasetmodelsreportschestdatasetsdeveloping
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Developing generalist foundation model has recently attracted tremendous attention among researchers in the field of AI for Medicine (AI4Medicine). A pivotal insight in developing these models is their reliance on dataset scaling, which emphasizes the requirements on developing open-source medical image datasets that incorporate diverse supervision signals across various imaging modalities. In this paper, we introduce RadGenome-Chest CT, a comprehensive, large-scale, region-guided 3D chest CT interpretation dataset based on CT-RATE. Specifically, we leverage the latest powerful universal segmentation and large language models, to extend the original datasets (over 25,692 non-contrast 3D chest CT volume and reports from 20,000 patients) from the following aspects: (i) organ-level segmentation masks covering 197 categories, which provide intermediate reasoning visual clues for interpretation; (ii) 665 K multi-granularity grounded reports, where each sentence of the report is linked to the corresponding anatomical region of CT volume in the form of a segmentation mask; (iii) 1.3 M grounded VQA pairs, where questions and answers are all linked with reference segmentation masks, enabling models to associate visual evidence with textual explanations. All grounded reports and VQA pairs in the validation set have gone through manual verification to ensure dataset quality. We believe that RadGenome-Chest CT can significantly advance the development of multimodal medical foundation models, by training to generate texts based on given segmentation regions, which is unattainable with previous relevant datasets. We will release all segmentation masks, grounded reports, and VQA pairs to facilitate further research and development in this field.

Discussion (0). Sign in to comment.

Forward citations

Cited by 11 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. ClinFusion: A Vision-Centric Multimodal LLM System for Holistic Medical Understanding

    cs.CV 2026-07 conditional novelty 6.5 of 10

    A cascaded multi-encoder medical MLLM with native 3D fusion and RoI-grounded report metrics claims SOTA on most 2D/3D medical benchmarks and highest radiologist report rankings.

  2. Cheap Probes Predict Expensive Training in 3D-CT Vision--Language Models

    cs.CV 2026-07 conditional novelty 6.0 of 10

    Disease-probe AUROC on frozen 3D-CT tokens predicts report-generation clinical micro-F1 across encoder×compression cells at r=0.95, ρ=0.89 (six cells, preliminary, one dataset).

  3. Machine Learning for Network Attacks Classification and Statistical Evaluation of Adversarial Learning Methodologies for Synthetic Data Generation

    cs.CR 2026-03 unverdicted novelty 6.0 of 10

    DiffVP turns scan-to-normal semantic discrepancies into learnable visual prefix tokens that guide an LLM to write more accurate, fine-grained 3D CT reports.

  4. MMRad-22K: A Structured Multimodal Evidence Dataset for Chest X-ray Report Generation

    cs.CV 2026-02 conditional novelty 6.0 of 10

    A 21,994-case chest X-ray dataset with interleaved regional text and image crops helps LVLMs generate more clinically accurate reports than text-only chain-of-thought.

  5. MedVista3D: Vision-Language Modeling for Reducing Diagnostic Errors in 3D CT Disease Detection, Understanding and Reporting

    cs.CV 2025-09 conditional novelty 6.0 of 10

    A 3D CT vision-language pretraining framework with global and organ-level contrastive alignment plus a text retrieval bank achieves state-of-the-art zero-shot disease classification, report retrieval, and medical visu...

  6. M3Ret: Unleashing Zero-shot Multimodal Medical Image Retrieval via Self-Supervision

    cs.CV 2025-09 conditional novelty 6.0 of 10

    One self-supervised encoder trained on unpaired X-ray, ultrasound, endoscopy, and CT data gives competitive zero-shot retrieval and seems to generalize to unseen MRI tasks.

  7. Disorder-induced stress-flow misalignment in soft glassy materials revealed using multi-directional shear

    cond-mat.soft 2025-08 unverdicted novelty 6.0 of 10

    Soft glassy materials show a transient stress response orthogonal to a newly applied shear direction, which a mesoscopic elasto-plastic model attributes to local yield-stress disorder.

  8. MedErr-CT: A Visual Question Answering Benchmark for Identifying and Correcting Errors in CT Reports

    cs.CV 2025-06 conditional novelty 6.0 of 10

    The authors introduce a 3D CT-based visual question answering benchmark with six error types and three task levels, and show that current 3D medical MLLMs perform poorly on it.

  9. Towards Enhancing 3D Spatial Reasoning in Medical Multimodal Large Language Models

    cs.CV 2026-07 reject novelty 5.0 of 10

    A synthetic chain-of-thought dataset generated from CT reports lets a 2D-pretrained medical MLLM improve on 3D CT spatial-reasoning benchmarks.

  10. HSENet: Hybrid Spatial Encoding Network for 3D Medical Vision-Language Understanding

    cs.CV 2025-06 conditional novelty 5.0 of 10

    HSENet improves 3D CT vision-language understanding by combining global and local 3D encoders with a centroid-based spatial token compressor, posting state-of-the-art results on CT-RATE and RadGenome-ChestCT.

  11. Region-Aware Multimodal Large Language Model via SlowFast Tokenization and Pseudo-Mask Guidance for 3D CT Report Generation

    eess.IV 2025-06 conditional novelty 4.0 of 10

    MedRegion-CT integrates region-representative tokens, mask-driven segmentation tokens, and patient-specific attribute prompts into a multimodal LLM, reporting state-of-the-art scores on RadGenome-Chest CT report generation.

Pith tools