Pith. sign in

REVIEW 3 cited by

A new Vietnamese PET/CT-report dataset improves medical VLM report generation and VQA, but clinical F1 scores remain modest.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

arxiv 2509.24739 v4 pith:DHRDSAFL submitted 2025-09-29 cs.CV

Toward a Vision-Language Foundation Model for Medical Data: Multimodal Dataset and Benchmarks for Vietnamese PET/CT Report Generation

classification cs.CV
keywords medicalvlmsdatasetimagingclinicaldatavietnamesedevelopment
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

ViMed-PET collects PET/CT scans and full-length clinical reports from a Vietnamese hospital. The data is split into head-neck, chest, and abdomen-pelvis regions, creating 8,271 image-report pairs from 2,757 studies. The authors also build augmented datasets with GPT-4o: paraphrased reports, question-answer dialogues, and study comparisons. They then fine-tune four vision-language models by adapting a 3D vision encoder, aligning visual and text features, and instruction-tuning with LoRA.

On standard NLP metrics, the fine-tuned models improve dramatically over zero-shot baselines such as LLaVA-Med, M3D, RadFM, and few-shot GPT-4o. On a new clinical test set of 80 lung-cancer patients with 398 lesions, the models reach clinical F1 scores near 50% for the best configuration, well above GPT-4o's 24% but still far from clinically reliable performance.

The paper claims this is the first large-scale PET/CT dataset with paired Vietnamese clinical reports. The main weaknesses are that the dataset URL is not visible in the manuscript, clinical evaluation relies on GPT-4o to extract structured findings while GPT-4o also generated part of the training data, and no error bars are reported. The resource, if released, is useful for low-resource-language medical AI, but the experimental evidence for clinical utility is provisional.

Core claim

The paper's central claim is that ViMed-PET is the first large-scale paired dataset of PET/CT images and corresponding clinical reports in Vietnamese, and that fine-tuning VLMs on it leads to substantial performance improvements: 'fine-tuning VLMs on our proposed ViMed-PET dataset leads to substantial performance improvements across both standard NLP metrics and clinically specific evaluation metrics' (Section 4.1). If correct, the dataset is a genuinely new resource for PET/CT vision-language research and Vietnamese medical AI.

Load-bearing premise

The clinical evaluation pipeline assumes that GPT-4o-based structured extraction is an unbiased measure of clinical accuracy. GPT-4o generated the paraphrased reports, VQA dialogues, and study comparisons used for training, and the same model extracts the Type, Position, and FDG attributes from both generated and ground-truth reports (Sections 2.4, 3.4, B.4). If GPT-4o extraction favors texts that resemble its own generated style, the reported clinical F1 improvements may partly reflect style matching rather than clinical correctness. A second load-bearing premise is that the 20-slice overlap between adjacent body-part segments does not leak information between training and test splits; the paper does not specify a patient-exclusive split.

Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Axiom & Free-Parameter Ledger

4 free parameters · 5 axioms · 0 invented entities

No new physical or mathematical entities are introduced. The dataset and evaluation protocol are the contributions. The listed free parameters are discretionary choices in the evaluation and preprocessing pipeline rather than parameters of a theory; they affect the reported numbers but are not fitted to a target result. The axioms are background assumptions about data provenance, modality semantics, and the reliability of GPT-4o-based attribute extraction.

free parameters (4)
  • Body-part split ratios = head-neck ~20% of body length; chest starts ~15% below the last head-neck slice; abdomen-pelvis is the remainder
    Expert-defined proportions that determine which slices and report sections are paired; directly affects the 8,271 paired samples and the 20-slice overlap.
  • Reconstruction loss weighting lambda = 1e-2
    Weights the L1 and inverted SSIM terms in Eq. (1) for Cosmos Tokenizer fine-tuning; not derived from data.
  • Clinical attribute categories and grouping rules = Type: 12 classes; FDG: 2 classes; Position: 5 classes; plus semantic grouping rules
    Expert-chosen taxonomy used to define F1-T, F1-TP, F1-TF, F1-TPF; different taxonomies or grouping rules would change the reported clinical scores.
  • Fixed slice depth for Cosmos Tokenizer = 120 slices
    Chosen from the dataset distribution; zero-padding or linear interpolation to 120 slices can add or lose anatomical content.
axioms (5)
  • domain assumption Report text is driven primarily by PET information, with minimal reference to CT anatomical details.
    Stated in Section 5; motivates treating PET/CT volumes as the visual input and ignoring CT-specific content. If false, the dataset and benchmark would be misaligned with the reports.
  • domain assumption Informed consent and de-identification were properly obtained and satisfy privacy regulations (Ethics Approval No. 6184/CN-HDDDBV).
    Section 2.2; needed for legal and ethical distribution of a hospital-derived dataset.
  • domain assumption CT-ViT, pretrained on non-contrast chest CT volumes, transfers to whole-body PET/CT after contrastive fine-tuning.
    Section 3.2 Stage 1; stage-1 fine-tuning relies on transfer from a different modality and anatomy to the target PET/CT data.
  • domain assumption GPT-4o few-shot extraction of clinical attributes is reliable after verification by two physicians.
    Appendix B.4; the entire clinical F1 pipeline assumes structured extraction matches physician judgment for both reference and generated reports, including handling of 'other' categories.
  • domain assumption The 2,757 studies come from independent patients.
    Section 2.2 claims independent patients; if patients have repeated scans, the independence claim and the split could be violated.

pith-pipeline@v1.3.0-alltime-deepseek · 28811 in / 10503 out tokens · 85569 ms · 2026-08-04T13:51:39.559206+00:00 · methodology

0 comments
Cite this review

Pith. "Pith review of Toward a Vision-Language Foundation Model for Medical Data: Multimodal Dataset and Benchmarks for Vietnamese PET/CT Report Generation." pith.science (2026). https://pith.science/paper/DHRDSAFL

@misc{pith2026250924739,
  author       = {Pith},
  title        = {Pith review of: Toward a Vision-Language Foundation Model for Medical Data: Multimodal Dataset and Benchmarks for Vietnamese PET/CT Report Generation},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/DHRDSAFL}},
  note         = {Machine review of arXiv:2509.24739}
}
Share X Bluesky LinkedIn Reddit HN
read the original abstract

Vision-Language Foundation Models (VLMs), trained on large-scale multimodal datasets, have driven significant advances in Artificial Intelligence (AI) by enabling rich cross-modal reasoning. Despite their success in general domains, applying these models to medical imaging remains challenging due to the limited availability of diverse imaging modalities and multilingual clinical data. Most existing medical VLMs are trained on a subset of imaging modalities and focus primarily on high-resource languages, thus limiting their generalizability and clinical utility. To address these limitations, we introduce a novel Vietnamese-language multimodal medical dataset consisting of 2,757 whole-body PET/CT volumes from independent patients and their corresponding full-length clinical reports. This dataset is designed to fill two pressing gaps in medical AI development: (1) the lack of PET/CT imaging data in existing VLMs training corpora, which hinders the development of models capable of handling functional imaging tasks; and (2) the underrepresentation of low-resource languages, particularly the Vietnamese language, in medical vision-language research. To the best of our knowledge, this is the first dataset to provide comprehensive PET/CT-report pairs in Vietnamese. We further introduce a training framework to enhance VLMs' learning, including data augmentation and expert-validated test sets. We conduct comprehensive experiments benchmarking state-of-the-art VLMs on downstream tasks. The experimental results show that incorporating our dataset significantly improves the performance of existing VLMs. We believe this dataset and benchmark will serve as a pivotal step in advancing the development of more robust VLMs for medical imaging, especially for low-resource languages and clinical use in Vietnamese healthcare. The source code is available at https://github.com/AIoT-Lab-BKAI/ViPET-ReportGen.

Figures

Figures reproduced from arXiv: 2509.24739 by Dac Thai Nguyen, Hong Son Mai, Huu Tien Nguyen, Huy Hieu Pham, Johan Barthelemy, Minh Quan Tran, Phi Le Nguyen, Quoc Viet Hung Nguyen, Quynh Anh Chau, Thanh Tam Nguyen, Thanh Trung Nguyen, Thao Nguyen Truong, The Minh Duc Nguyen, Trung Thanh Nguyen.

Figure 1
Figure 1. Figure 1: An example from our ViMed-PET dataset. The visual input consists of aligned 3D CT [PITH_FULL_IMAGE:figures/full_fig_p005_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: Overview of the fine-tuning pipeline. Stage 1: Fine-tuning the 3D Vision Encoder with [PITH_FULL_IMAGE:figures/full_fig_p007_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Message used to prompt GPT-4o to generate our medical VQA conversations. Manually [PITH_FULL_IMAGE:figures/full_fig_p022_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: Few-shot examples included in our prompt to construct the VQA conversation dataset. [PITH_FULL_IMAGE:figures/full_fig_p023_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: System message used to prompt GPT-4o for generating the study comparison dataset. The [PITH_FULL_IMAGE:figures/full_fig_p023_5.png] view at source ↗
Figure 6
Figure 6. Figure 6: Few-shot example used in our prompt for generating the study comparison dataset. The [PITH_FULL_IMAGE:figures/full_fig_p024_6.png] view at source ↗
Figure 7
Figure 7. Figure 7: Clinical evaluation pipeline. Phase 1: Experts define structured clinical attributes from [PITH_FULL_IMAGE:figures/full_fig_p026_7.png] view at source ↗
Figure 8
Figure 8. Figure 8: Message used to prompt GPT-4o for structuring VLM-generated reports into JSON format. [PITH_FULL_IMAGE:figures/full_fig_p027_8.png] view at source ↗
Figure 9
Figure 9. Figure 9: Example of a few-shot prompt used to guide GPT-4o in extracting structured JSON data [PITH_FULL_IMAGE:figures/full_fig_p027_9.png] view at source ↗
Figure 10
Figure 10. Figure 10: Visualization of the flattening strategy. Consecutive 2D slices from a 3D medical volume [PITH_FULL_IMAGE:figures/full_fig_p028_10.png] view at source ↗
Figure 11
Figure 11. Figure 11: Prompt template used with GPT-4o to analyze concatenated 2D grid images and generate [PITH_FULL_IMAGE:figures/full_fig_p028_11.png] view at source ↗
Figure 12
Figure 12. Figure 12: Ground truth and generated PET/CT reports for the chest and abdomen-pelvis regions [PITH_FULL_IMAGE:figures/full_fig_p030_12.png] view at source ↗
Figure 13
Figure 13. Figure 13: Short-form VQA interaction in Vietnamese (EN: translated) between a user and the [PITH_FULL_IMAGE:figures/full_fig_p031_13.png] view at source ↗
Figure 14
Figure 14. Figure 14: Long-form VQA interaction in Vietnamese (EN: translated) using the CT-ViT + Mistral [PITH_FULL_IMAGE:figures/full_fig_p032_14.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. MedOpenClaw and MedFlowBench: Auditing Medical Agents in Full-Study Workflows

    cs.CV 2026-03 conditional novelty 8.0

    MedFlowBench evaluates VLM agents on full radiology and pathology studies by requiring both task answers and verifiable evidence like key slices and regions of interest, revealing that answer-only scores overestimate ...

  2. Region-Grounded Report Generation for 3D Medical Imaging: A Fine-Grained Dataset and Graph-Enhanced Framework

    cs.CV 2026-04 unverdicted novelty 7.0

    Introduces VietPET-RoI dataset with fine-grained RoI annotations for Vietnamese 3D PET/CT and HiRRA graph framework that improves report generation by modeling region dependencies, claiming large gains over prior models.

  3. Region-Grounded Report Generation for 3D Medical Imaging: A Fine-Grained Dataset and Graph-Enhanced Framework

    cs.CV 2026-04 conditional novelty 7.0

    Introduces the first large-scale 3D PET/CT dataset with fine-grained RoI annotations for Vietnamese and a graph-enhanced HiRRA framework that achieves SOTA report generation by modeling RoI dependencies.

Reference graph

Works this paper leans on

62 extracted references · 22 linked inside Pith · cited by 2 Pith papers

  1. [1]

    Learning transferable visual models from natural language supervision

    Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning transferable visual models from natural language supervision. InInternational Conference on Machine Learning, pages 8748–8763, 2021

  2. [2]

    The dawn of LMMs: Preliminary explorations with GPT-4V(ision).arXiv preprint arXiv:2309.17421, pages 1–166, 2023

    Zhengyuan Yang, Linjie Li, Kevin Lin, Jianfeng Wang, Chung-Ching Lin, Zicheng Liu, and Lijuan Wang. The dawn of LMMs: Preliminary explorations with GPT-4V(ision).arXiv preprint arXiv:2309.17421, pages 1–166, 2023

  3. [3]

    Claude: An AI Assistant by Anthropic

    Anthropic. Claude: An AI Assistant by Anthropic. https://www.anthropic.com/index/ claude, 2024. Accessed: 2025-05-11

  4. [4]

    Llava-onevision: Easy visual task transfer.arXiv preprint arXiv:2408.03326, 2024

    Bo Li, Yuanhan Zhang, Dong Guo, Renrui Zhang, Feng Li, Hao Zhang, Kaichen Zhang, Peiyuan Zhang, Yanwei Li, Ziwei Liu, et al. Llava-onevision: Easy visual task transfer.arXiv preprint arXiv:2408.03326, 2024. 10

  5. [5]

    Qwen2.5-VL technical report.arXiv preprint arXiv:2502.13923, 2025

    Shuai Bai, Keqin Chen, Xuejing Liu, Jialin Wang, Wenbin Ge, Sibo Song, Kai Dang, Peng Wang, Shijie Wang, Jun Tang, et al. Qwen2.5-VL technical report.arXiv preprint arXiv:2502.13923, 2025

  6. [6]

    Qwen technical report.arXiv preprint arXiv:2309.16609, 2023

    Jinze Bai, Shuai Bai, Yunfei Chu, Zeyu Cui, Kai Dang, Xiaodong Deng, Yang Fan, Wenbin Ge, Yu Han, Fei Huang, et al. Qwen technical report.arXiv preprint arXiv:2309.16609, 2023

  7. [7]

    Visual instruction tuning.Advances in Neural Information Processing Systems, 36:34892–34916, 2023

    Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. Visual instruction tuning.Advances in Neural Information Processing Systems, 36:34892–34916, 2023

  8. [8]

    Ahive: Anatomy-aware hierarchical vision encoding for interactive radiology report retrieval

    Sixing Yan, William K Cheung, Ivor W Tsang, Keith Chiu, Terence M Tong, Ka Chun Cheung, and Simon See. Ahive: Anatomy-aware hierarchical vision encoding for interactive radiology report retrieval. InProceedings of the 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 14324–14333, 2024

  9. [9]

    Fg-cxr: A radiologist-aligned gaze dataset for enhancing interpretability in chest x-ray report generation

    Trong Thang Pham, Ngoc-Vuong Ho, Nhat-Tan Bui, Thinh Phan, Patel Brijesh, Donald Adjeroh, Gianfranco Doretto, Anh Nguyen, Carol C Wu, Hien Nguyen, and Le Ngang. Fg-cxr: A radiologist-aligned gaze dataset for enhancing interpretability in chest x-ray report generation. InProceedings of the 2024 Asian Conference on Computer Vision, pages 941–958, 2024

  10. [10]

    Cplip: zero-shot learning for histopathology with comprehensive vision-language alignment

    Sajid Javed, Arif Mahmood, Iyyakutti Iyappan Ganapathi, Fayaz Ali Dharejo, Naoufel Werghi, and Mohammed Bennamoun. Cplip: zero-shot learning for histopathology with comprehensive vision-language alignment. InProceedings of the IEEE/CVF 2024 Conference on Computer Vision and Pattern Recognition, pages 11450–11459, 2024

  11. [11]

    Flamingo: a visual language model for few-shot learning.arXiv preprint arXiv:2204.14198, 2022

    Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Arthur Mensch, Katie Milln, Matthew Reynolds, Rebecca Ring, Matthew Tancik, Xiuye Zhai, et al. Flamingo: a visual language model for few-shot learning.arXiv preprint arXiv:2204.14198, 2022

  12. [12]

    Gpt-4o: Openai’s multimodal model with vision, audio, and text capabilities

    OpenAI. Gpt-4o: Openai’s multimodal model with vision, audio, and text capabilities. https: //openai.com/index/gpt-4o, 2024. Accessed: 2025-04-30

  13. [13]

    Can generalist foundation models outcompete special-purpose tuning? case study in medicine.arXiv preprint arXiv:2311.16452, 2023

    Harsha Nori, Yin Tat Lee, Sheng Zhang, Dean Carignan, Richard Edgar, Nicolo Fusi, Nicholas King, Jonathan Larson, Yuanzhi Li, Weishung Liu, et al. Can generalist foundation models outcompete special-purpose tuning? case study in medicine.arXiv preprint arXiv:2311.16452, 2023

  14. [14]

    Multi- modal understanding and generation for medical images and text via vision-language pre- training.IEEE Journal of Biomedical and Health Informatics, 26(12):6070–6080, 2022

    Jong Hak Moon, Hyungyung Lee, Woncheol Shin, Young-Hak Kim, and Edward Choi. Multi- modal understanding and generation for medical images and text via vision-language pre- training.IEEE Journal of Biomedical and Health Informatics, 26(12):6070–6080, 2022

  15. [15]

    Towards a better understanding of annotation tools for medical imaging: a survey.Multimedia Tools and Applications, 81(18):25877–25911, 2022

    Manar Aljabri, Manal AlAmir, Manal AlGhamdi, Mohamed Abdel-Mottaleb, and Fernando Collado-Mesa. Towards a better understanding of annotation tools for medical imaging: a survey.Multimedia Tools and Applications, 81(18):25877–25911, 2022

  16. [16]

    Knowledge-augmented language models interpreting structured chest x-ray findings.arXiv preprint arXiv:2505.01711, 2025

    Alexander Davis, Rafael Souza, and Jia-Hao Lim. Knowledge-augmented language models interpreting structured chest x-ray findings.arXiv preprint arXiv:2505.01711, 2025

  17. [17]

    Knowledge matters: Chest radiology report generation with general and specific knowledge.Medical Image Analysis, 80: 102510, 2022

    Shuxin Yang, Xian Wu, Shen Ge, S Kevin Zhou, and Li Xiao. Knowledge matters: Chest radiology report generation with general and specific knowledge.Medical Image Analysis, 80: 102510, 2022

  18. [18]

    Medclip: Contrastive learning from unpaired medical images and text

    Zifeng Wang, Zhenbang Wu, Dinesh Agarwal, and Jimeng Sun. Medclip: Contrastive learning from unpaired medical images and text. InProceedings of the 2022 Conference on Empirical Methods in Natural Language Processing., volume 2022, page 3876, 2022

  19. [19]

    Med-flamingo: a multimodal medical few-shot learner

    Michael Moor, Qian Huang, Shirley Wu, Michihiro Yasunaga, Yash Dalmia, Jure Leskovec, Cyril Zakka, Eduardo Pontes Reis, and Pranav Rajpurkar. Med-flamingo: a multimodal medical few-shot learner. InMachine Learning for Health, pages 353–367. PMLR, 2023

  20. [20]

    Roentgen: vision-language foundation model for chest x-ray generation.arXiv preprint arXiv:2211.12737, 2022

    Pierre Chambon, Christian Bluethgen, Jean-Benoit Delbrouck, Rogier Van der Sluijs, Mał- gorzata Połacin, Juan Manuel Zambrano Chaves, Tanishq Mathew Abraham, Shivanshu Purohit, Curtis P Langlotz, and Akshay Chaudhari. Roentgen: vision-language foundation model for chest x-ray generation.arXiv preprint arXiv:2211.12737, 2022. 11

  21. [21]

    Medvilam: A multi- modal large language model with advanced generalizability and explainability for medical data understanding and generation.arXiv preprint arXiv:2409.19684, 2024

    Lijian Xu, Hao Sun, Ziyu Ni, Hongsheng Li, and Shaoting Zhang. Medvilam: A multi- modal large language model with advanced generalizability and explainability for medical data understanding and generation.arXiv preprint arXiv:2409.19684, 2024

  22. [22]

    Med3dvlm: An efficient vision- language model for 3d medical image analysis.arXiv preprint arXiv:2503.20047, 2025

    Yu Xin, Gorkem Can Ates, Kuang Gong, and Wei Shao. Med3dvlm: An efficient vision- language model for 3d medical image analysis.arXiv preprint arXiv:2503.20047, 2025

  23. [23]

    Multilingual diversity improves vision- language representations.Advances in Neural Information Processing Systems, 37:91430– 91459, 2024

    Thao Nguyen, Matthew Wallingford, Sebastin Santy, Wei-Chiu Ma, Sewoong Oh, Ludwig Schmidt, Pang Wei W Koh, and Ranjay Krishna. Multilingual diversity improves vision- language representations.Advances in Neural Information Processing Systems, 37:91430– 91459, 2024

  24. [24]

    Pmc-clip: Contrastive language-image pre-training using biomedical documents

    Weixiong Lin, Ziheng Zhao, Xiaoman Zhang, Chaoyi Wu, Ya Zhang, Yanfeng Wang, and Weidi Xie. Pmc-clip: Contrastive language-image pre-training using biomedical documents. In International Conference on Medical Image Computing and Computer-Assisted Intervention, pages 525–536. Springer, 2023

  25. [25]

    Seco de Herrera, et al

    Johannes Rückert, Louise Bloch, Raphael Brüngel, Ahmad Idrissi-Yaghir, Henning Schäfer, Cynthia S Schmidt, Sven Koitka, Obioma Pelka, Asma Ben Abacha, Alba G. Seco de Herrera, et al. Rocov2: Radiology objects in context version 2, an updated multimodal image dataset. Scientific Data, 11(1):688, 2024

  26. [26]

    Mimic-cxr, a de-identified publicly available database of chest radiographs with free-text reports.Scientific data, 6(1):317, 2019

    Alistair EW Johnson, Tom J Pollard, Seth J Berkowitz, Nathaniel R Greenbaum, Matthew P Lungren, Chih-ying Deng, Roger G Mark, and Steven Horng. Mimic-cxr, a de-identified publicly available database of chest radiographs with free-text reports.Scientific data, 6(1):317, 2019

  27. [27]

    Llava-med: Training a large language-and-vision assistant for biomedicine in one day.Advances in Neural Information Processing Systems, 36: 28541–28564, 2023

    Chunyuan Li, Cliff Wong, Sheng Zhang, Naoto Usuyama, Haotian Liu, Jianwei Yang, Tristan Naumann, Hoifung Poon, and Jianfeng Gao. Llava-med: Training a large language-and-vision assistant for biomedicine in one day.Advances in Neural Information Processing Systems, 36: 28541–28564, 2023

  28. [28]

    M3d: Advancing 3d medical image analysis with multi-modal large language models.arXiv preprint arXiv:2404.00578, 2024

    Fan Bai, Yuxin Du, Tiejun Huang, Max Q-H Meng, and Bo Zhao. M3d: Advancing 3d medical image analysis with multi-modal large language models.arXiv preprint arXiv:2404.00578, 2024

  29. [29]

    Towards generalist foundation model for radiology by leveraging web-scale 2d&3d medical data, 2023

    Chaoyi Wu, Xiaoman Zhang, Ya Zhang, Yanfeng Wang, and Weidi Xie. Towards generalist foundation model for radiology by leveraging web-scale 2d&3d medical data, 2023

  30. [30]

    Non-invasive metabolic imaging of brain tumours in the era of precision medicine.Nature Reviews Clinical Oncology, 13(12):725–739, 2016

    Michelle M Kim, Abhijit Parolia, Mark P Dunphy, and Sriram Venneti. Non-invasive metabolic imaging of brain tumours in the era of precision medicine.Nature Reviews Clinical Oncology, 13(12):725–739, 2016

  31. [31]

    Imaging tumor metabolism using positron emission tomography.The Cancer Journal, 21(2):129–136, 2015

    David Y Lewis, Dmitry Soloviev, and Kevin M Brindle. Imaging tumor metabolism using positron emission tomography.The Cancer Journal, 21(2):129–136, 2015

  32. [32]

    Molecular imaging of cancer with positron emission tomography.Nature Reviews Cancer, 2(9):683–693, 2002

    Sanjiv Sam Gambhir. Molecular imaging of cancer with positron emission tomography.Nature Reviews Cancer, 2(9):683–693, 2002

  33. [33]

    Advances in pet imaging of cancer.Nature Reviews Cancer, 23(7):474–490, 2023

    Johannes Schwenck, Dominik Sonanini, Jonathan M Cotton, Hans-Georg Rammensee, Christian la Fougère, Lars Zender, and Bernd J Pichler. Advances in pet imaging of cancer.Nature Reviews Cancer, 23(7):474–490, 2023

  34. [34]

    Is long–axial-field-of-view pet/ct cost-effective? an international health–economic analysis.Journal of Nuclear Medicine, 2025

    Ian Alberts, Stuart More, Karen Knapp, Riccardo Mei, Stefano Fanti, Clemens Mingels, Lorenzo Nardo, Nii Boye Hammond, Harish Nagaraj, Axel Rominger, et al. Is long–axial-field-of-view pet/ct cost-effective? an international health–economic analysis.Journal of Nuclear Medicine, 2025

  35. [35]

    Mohammad Naghavi-Behzad, Oke Gerke, Annette Raskov Kodahl, Marianne V ogsen, Jon Thor Asmussen, Wolfgang Weber, Malene Grubbe Hildebrandt, and Kristian Kidholm. Cost- effectiveness of 2-[18f] fdg-pet/ct versus ce-ct for response monitoring in patients with metastatic breast cancer: a register-based comparative study.Scientific Reports, 13(1):16315, 2023. 12

  36. [36]

    Explore the world population through data (2025)

    World Population Review. Explore the world population through data (2025). https:// worldpopulationreview.com/, 2025. Accessed: 2025-05-12

  37. [37]

    CT2REP: Automated radiology report generation for 3d medical imaging

    Ibrahim Ethem Hamamci, Sezgin Er, and Bjoern Menze. CT2REP: Automated radiology report generation for 3d medical imaging. InProceedings of the 2024 International Conference on Medical Image Computing and Computer-Assisted Intervention, pages 476–486. Springer, 2024

  38. [38]

    Rider lung pet-ct: Data for quantitative imaging biomarker evaluation.The Cancer Imaging Archive, 2015

    Peter Muzi, Matthew Wanner, and Paul Kinahan. Rider lung pet-ct: Data for quantitative imaging biomarker evaluation.The Cancer Imaging Archive, 2015

  39. [39]

    Data from head- neck-pet-ct.The Cancer Imaging Archive, 2017

    Martin Vallières, Emily Kay-Rivest, Léo Jean Perrin, Xavier Liem, Christophe Furstoss, Nader Khaouam, Phuc Félix Nguyen-Tan, Chang-Shu Wang, and Khalil Sultanem. Data from head- neck-pet-ct.The Cancer Imaging Archive, 2017

  40. [40]

    P. Li, S. Wang, T. Li, J. Lu, Y . HuangFu, and D. Wang. A large-scale ct and pet/ct dataset for lung cancer diagnosis (lung-pet-ct-dx).The Cancer Imaging Archive, 2020

  41. [41]

    Gatidis and T

    S. Gatidis and T. Kuestner. A whole-body FDG-PET/CT dataset with manually annotated tumor lesions (FDG-PET-CT-Lesions).The Cancer Imaging Archive, 2022

  42. [42]

    Generatect: Text-conditional generation of 3d chest ct volumes

    Ibrahim Ethem Hamamci, Sezgin Er, Anjany Sekuboyina, Enis Simsar, Alperen Tezcan, Ayse Gulnihan Simsek, Sevval Nil Esirgun, Furkan Almas, Irem Do˘gan, Muhammed Furkan Dasdelen, et al. Generatect: Text-conditional generation of 3d chest ct volumes. InEuropean Conference on Computer Vision, pages 126–143. Springer, 2024

  43. [43]

    Cosmos world foundation model platform for physical ai.arXiv preprint arXiv:2501.03575, 2025

    Niket Agarwal, Arslan Ali, Maciej Bala, Yogesh Balaji, Erik Barker, Tiffany Cai, Prithvijit Chattopadhyay, Yongxin Chen, Yin Cui, Yifan Ding, et al. Cosmos world foundation model platform for physical ai.arXiv preprint arXiv:2501.03575, 2025

  44. [44]

    Albert Q. Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, Lélio Renard Lavaud, Marie-Anne Lachaux, Pierre Stock, Teven Le Scao, Thibaut Lavril, Thomas Wang, Timothée Lacroix, and William El Sayed. Mistral 7b.arXiv preprint arXiv:23...

  45. [45]

    Llama 2: Open foundation and fine-tuned chat models.arXiv preprint arXiv:2307.09288, 2023

    Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. Llama 2: Open foundation and fine-tuned chat models.arXiv preprint arXiv:2307.09288, 2023

  46. [46]

    Lora: Low-rank adaptation of large language models.Interna- tional Conference on Learning Representations, 1(2):3, 2022

    Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, Weizhu Chen, et al. Lora: Low-rank adaptation of large language models.Interna- tional Conference on Learning Representations, 1(2):3, 2022

  47. [47]

    Bleu: a method for automatic evaluation of machine translation

    Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. Bleu: a method for automatic evaluation of machine translation. InProceedings of the 40th Annual Meeting of the Association for Computational Linguistics, pages 311–318, 2002

  48. [48]

    Rouge: A package for automatic evaluation of summaries

    Chin-Yew Lin. Rouge: A package for automatic evaluation of summaries. InText Summarization Branches Out, pages 74–81, 2004

  49. [49]

    Bertscore: Evaluating text generation with bert.arXiv preprint arXiv:1904.09675, 2019

    Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q Weinberger, and Yoav Artzi. Bertscore: Evaluating text generation with bert.arXiv preprint arXiv:1904.09675, 2019

  50. [50]

    Plip: Language-image pre-training for person representation learning

    Jialong Zuo, Jiahao Hong, Feng Zhang, Changqian Yu, Hanyu Zhou, Changxin Gao, Nong Sang, and Jingdong Wang. Plip: Language-image pre-training for person representation learning. Advances in Neural Information Processing Systems, 37:45666–45702, 2024

  51. [51]

    Biomedclip: a multimodal biomedical foundation model pretrained from fifteen million scientific image-text pairs.arXiv preprint arXiv:2303.00915, 2023

    Sheng Zhang, Yanbo Xu, Naoto Usuyama, Hanwen Xu, Jaspreet Bagga, Robert Tinn, Sam Preston, Rajesh Rao, Mu Wei, Naveen Valluri, et al. Biomedclip: a multimodal biomedical foundation model pretrained from fifteen million scientific image-text pairs.arXiv preprint arXiv:2303.00915, 2023. 13

  52. [52]

    Merlin: A vision language foundation model for 3d computed tomography

    Louis Blankemeier, Joseph Paul Cohen, Ashwin Kumar, Dave Van Veen, Syed Jamal Safdar Gardezi, Magdalini Paschali, Zhihong Chen, Jean-Benoit Delbrouck, Eduardo Reis, Cesar Truyts, et al. Merlin: A vision language foundation model for 3d computed tomography. Research Square, pages rs–3, 2024

  53. [53]

    Xraygpt: Chest radiographs summarization using large medical vision-language models

    Omkar Chakradhar Thawakar, Abdelrahman M Shaker, Sahal Shaji Mullappilly, Hisham Cholakkal, Rao Muhammad Anwer, Salman Khan, Jorma Laaksonen, and Fahad Khan. Xraygpt: Chest radiographs summarization using large medical vision-language models. InProceedings of the 23rd Workshop on Biomedical Natural Language Processing, pages 440–448, 2024

  54. [54]

    Elixr: Towards a general purpose x-ray artificial intelligence system through alignment of large language models and radiology vision encoders.arXiv preprint arXiv:2308.01317, 2023

    Shawn Xu, Lin Yang, Christopher Kelly, Marcin Sieniek, Timo Kohlberger, Martin Ma, Wei- Hung Weng, Atilla Kiraly, Sahar Kazemzadeh, Zakkai Melamed, et al. Elixr: Towards a general purpose x-ray artificial intelligence system through alignment of large language models and radiology vision encoders.arXiv preprint arXiv:2308.01317, 2023

  55. [55]

    Chexagent: Towards a foundation model for chest x-ray interpretation.arXiv preprint arXiv:2401.12208, 2024

    Zhihong Chen, Maya Varma, Jean-Benoit Delbrouck, Magdalini Paschali, Louis Blankemeier, Dave Van Veen, Jeya Maria Jose Valanarasu, Alaa Youssef, Joseph Paul Cohen, Eduardo Pontes Reis, et al. Chexagent: Towards a foundation model for chest x-ray interpretation.arXiv preprint arXiv:2401.12208, 2024

  56. [56]

    Qilin- med-vl: Towards chinese large vision-language model for general healthcare.arXiv preprint arXiv:2310.17956, 2023

    Junling Liu, Ziming Wang, Qichen Ye, Dading Chong, Peilin Zhou, and Yining Hua. Qilin- med-vl: Towards chinese large vision-language model for general healthcare.arXiv preprint arXiv:2310.17956, 2023

  57. [57]

    Huatuogpt, towards taming language model to be a doctor.arXiv preprint arXiv:2305.15075, 2023

    Hongbo Zhang, Junying Chen, Feng Jiang, Fei Yu, Zhihong Chen, Jianquan Li, Guiming Chen, Xiangbo Wu, Zhiyi Zhang, Qingying Xiao, et al. Huatuogpt, towards taming language model to be a doctor.arXiv preprint arXiv:2305.15075, 2023

  58. [58]

    Decoupled weight decay regularization.arXiv preprint arXiv:1711.05101, 2017

    Ilya Loshchilov and Frank Hutter. Decoupled weight decay regularization.arXiv preprint arXiv:1711.05101, 2017

  59. [59]

    Phobert: Pre-trained language models for vietnamese

    Dat Quoc Nguyen and Anh Tuan Nguyen. Phobert: Pre-trained language models for vietnamese. arXiv preprint arXiv:2003.00744, 2020

  60. [60]

    Sgdr: Stochastic gradient descent with warm restarts.arXiv preprint arXiv:1608.03983, 2016

    Ilya Loshchilov and Frank Hutter. Sgdr: Stochastic gradient descent with warm restarts.arXiv preprint arXiv:1608.03983, 2016

  61. [61]

    Limitations

    Naoki Wake, Atsushi Kanehira, Kazuhiro Sasabuchi, Jun Takamatsu, and Katsushi Ikeuchi. Open-vocabulary action localization with iterative visual prompting.IEEE Access, 2025. 14 NeurIPS Paper Checklist 1.Claims Question: Do the main claims made in the abstract and introduction accurately reflect the paper’s contributions and scope? Answer: [Yes] Justificat...

  62. [62]

    role":"system

    Institutional review board (IRB) approvals or equivalent for research with human subjects Question: Does the paper describe potential risks incurred by study participants, whether such risks were disclosed to the subjects, and whether Institutional Review Board (IRB) approvals (or an equivalent approval/review based on the requirements of your country or ...