Pith. sign in

REVIEW 4 major objections 4 minor 51 references

Gla-AI4BioMed at RRG24: Visual Instruction-tuned Adaptation for Radiology Report Generation

T0 review · 4 major / 4 minor · reviewed 2026-08-11 · deepseek-v4-flash

Pith's one-line read Two-stage fine-tuned vision-language model generates competitive radiology reports from stitched chest X-rays.

desk verdict Competent shared-task system paper that deserves referee time, but the identical public/hidden F1-RadGraph score needs verification before trusting the 4th-place claim. read the letter →

arxiv 2412.04954 v1 pith:IE6E4GY6 submitted 2024-12-06 cs.CV cs.CLcs.LG

classification cs.CVcs.CLcs.LG
keywords radiologyreportgenerationvisualinstructiontuningchestX-rayvision-languagemodelLoRACLIPVicuna-7BF1-RadGraph
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper claims that a general-purpose vision-language model becomes competitive at radiology report generation when it is fine-tuned in two stages on the RRG24 shared-task dataset: an MLP adapter is first trained to align CLIP chest X-ray features with the Vicuna-7B language model, then the LLM is LoRA-tuned for report generation with the encoder and adapter frozen. The authors report hidden-test F1-RadGraph scores of 24.13 for Findings and 22.10 for Impressions, placing 4th on the leaderboard at submission. They also claim that horizontally stitching up to four X-ray images into one input lets a single image encoder process multi-image studies without dedicated multi-image encoding. A sympathetic reader would care because the recipe suggests that modest reuse of open-source components, not a new architecture, can produce useful clinical text.

What carries the argument

The mechanism is a two-stage visual instruction-tuning pipeline. Stage 1 freezes both the CLIP image encoder and the Vicuna-7B LLM and trains only a GELU-activated MLP adapter with hidden size 1024, aligning chest X-ray features to the LLM's text embedding space. Stage 2 keeps the encoder and adapter frozen and applies Low-Rank Adaptation (LoRA) to the LLM for three epochs on the section-generation task. Multi-image studies are handled by taking the first up to four images and horizontally concatenating them into a single input image before the encoder, so the model consumes a fixed-size stitched input instead of attending over separate images. This machinery is what lets a single-image encoder and a text-only LLM produce the reported results.

What would settle it

Take a cohort of chest X-ray studies with five or more images whose reference report describes a finding visible only in the fifth or later image; run the Med-CXRGen-F model on the first four images stitched as in the paper, and compare its F1-RadGraph with the same model given all images or the same five images reordered to put the finding-bearing image first. If the truncated-input score drops materially while the all-images or reordered score holds, the four-image truncation assumption is falsified.

Watch

Extended reading notes

Core claim

The central discovery claimed is that visual instruction tuning, in the LLaVA-1.5 style, transfers to radiology report generation when applied in two stages on the provided dataset: in Stage 1 the CLIP encoder and Vicuna LLM stay frozen while the MLP adapter learns to align chest X-ray features with the LLM, and in Stage 2 the adapter and encoder remain frozen while LoRA updates the LLM for three epochs on the report-generation task. Separate models are trained for Findings and Impressions, each decoding up to 150 tokens from a prompt requesting a description of that section. With the first up to four study images stitched horizontally into a single input, the models score 24.13 and 22.10 F1-RadGraph on the hidden Findings and Impressions test sections, which the paper reports as 4th place among RRG24 submissions at the time of writing. The authors present this as demonstrating that a domain-specific fine-tuned VLM can handle multiple images and generate clinically relevant report sections, rather than as a new architectural contribution.

Load-bearing premise

The model sees only the first up to four images of a study, stitched side by side, and the paper offers no evidence that this truncation and ordering preserves the clinically relevant content of the full study.

Editorial extensions

If this is right

  • If the recipe is correct, a competitive radiology report generator can be assembled from an open-source VLM and public dataset without a new multimodal architecture, lowering the entry barrier for clinical NLP teams.
  • The horizontal-stitching trick implies that a single-image encoder suffices for studies with a small bounded number of images, which is a practical low-cost option for resource-limited settings.
  • Separate models per report section indicate that Findings and Impressions benefit from section-specific fine-tuning, so the design choice of one model per section is a safe default for similar shared tasks.
  • The reported hidden-test scores give later RRG24 participants a concrete baseline: with the same metric, matching or exceeding F1-RadGraph 24.13 on Findings would establish improvement over this system.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • My inference: because no ablation varies the number or order of stitched images, the 'first four images' truncation probably loses findings that appear only in later images; on such studies the true score would be lower than the leaderboard number.
  • My inference: stitching discards spatial and temporal ordering among views; reordering-invariant or attention-based multi-image fusion would be a natural next step that this paper's setup cannot capture.
  • My inference: the authors' own discussion concedes that performance suffers when superfluous images are included, so an image-selection or relevance-weighting mechanism is a direct, testable follow-up.
  • My inference: because evaluation rests on F1-RadGraph, the clinical safety of the generated text is unmeasured; an error analysis on missed findings would be needed to judge whether the reports are usable rather than just score-competitive.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 4 minor

Summary. The paper presents a radiology report generation system developed for the RRG24 shared task. The authors fine-tune Vicuna-7B with a CLIP vision encoder and an MLP adapter in two stages: first aligning chest X-ray features with the language model, then applying LoRA fine-tuning for radiology report generation. Separate models are trained for Findings and Impressions sections, and up to four input images are horizontally stitched into a single image. The system is evaluated on the RRG24 validation, public test, and hidden test splits across five metrics. The central empirical claim is that the hidden-test F1-RadGraph scores are 24.13 for Findings and 22.10 for Impressions, placing the system 4th on the leaderboard at submission time.

Significance. If the reported hidden-test scores are accurate, the paper demonstrates a competitive recipe for adapting a general-purpose visual language model to radiology report generation with modest computational resources, and it provides a useful data point for the RRG24 benchmark. The manuscript is commendably transparent about its limitations and links to public code. However, the technical novelty is limited: the two-stage adaptation follows LLaVA-1.5 and LLaVA-Med closely, and the multi-image handling is a simple truncation-and-stitch heuristic with no supporting ablation. The main load-bearing evidence is a small set of leaderboard numbers, and at least one of those numbers appears to be internally inconsistent, which is why the result cannot be accepted without verification.

major comments (4)
  1. [Section 5, Table 5] The hidden-test Findings F1-RadGraph is reported as exactly 24.13, identical to the public-test Findings F1-RadGraph, while every other metric in Table 5 differs between the two test sets (e.g., BLEU4 changes from 8.07 to 7.65 and ROUGEL from 24.90 to 24.35). An exact match to two decimal places on two independent test sets is highly implausible and is more likely explained by a copy-paste or logging error. Because the headline claim of a 4th-place finish rests on the hidden-test value, the authors must verify this number against the official RRG24 leaderboard and correct Table 5 and the text accordingly. If the value is confirmed, please provide evidence; otherwise the central empirical claim is unsupported.
  2. [Section 4.1 vs. Section 4.3] Section 4.1 states that the maximum length is 1024 for both text input and inference output, but Section 4.3 states that inference decodes up to 150 tokens, consistent with the baseline. These two statements contradict each other. This is not a cosmetic issue: Table 3 reports an average Findings word count of 380 on test-public, so a 150-token cap could truncate a large fraction of generated reports and directly affect all reported metrics. The authors must clarify the actual decoding limit and, if 150 tokens was used, discuss the effect of truncation on the reported scores.
  3. [Section 4.1, Section 6] The preprocessing choice of taking only the first four images and horizontally stitching them is not validated by any ablation or quantitative analysis. The paper claims this strategy is 'proven to be robust in our experiments' in Section 4.1, but no such experiment is reported, and Section 6 later concedes that performance 'may be compromised in multi-image inference scenarios where it does not account for superfluous images.' The authors should provide at least a comparison of stitching versus separate encoding, or an analysis of image-order sensitivity and the effect of discarding images beyond the fourth, to support the claim that the method preserves clinically relevant information.
  4. [Section 3, Section 3.3] There is an internal contradiction in the training description. Section 3 says the authors follow LLaVA-1.5 protocols 'including a joint tuning phase for the LLM and adapter,' but Section 3.3, Stage 2, states that 'the visual encoder weights and adapter are kept frozen while continuing to update the pre-trained LLM weights using LoRA.' These descriptions cannot both be true. Since the exact parameter-update scheme is essential for reproducing the method, the authors must state unambiguously whether the MLP adapter is updated in Stage 2 or frozen.
minor comments (4)
  1. [Abstract, Section 3.1, Section 3.3] There are several language errors, including 'the ability of model' in the Abstract, 'model ability of to mimic' in Section 3.1, and 'visual instrumental tuning' in Section 3.3, which should read 'visual instruction tuning.'
  2. [Table 5] The model Med-CXRGen-I achieves a much higher validation F1-RadGraph (26.65) than test-public (22.79) or test-hidden (22.10), but this drop is not discussed; a brief comment on possible distribution shift or overfitting would help the reader interpret the results.
  3. [Section 4.1] The dataset statistics in Table 2 are presented without a citation to the original source of the test-hidden split; please clarify whether the hidden split is the official RRG24 hidden split and how it relates to the MIMIC-CXR, CheXpert, PadChest, BIMCV-COVID19, and OpenI collections.
  4. [Table 5] No baseline or comparison system is included in the results table. Since the shared task provides a ViLMedic baseline, the authors should report that baseline (or at least the leaderboard range) so that the reader can assess whether the reported scores are strong or merely moderate.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: the system is trained on the RRG24 training split and scored on organizer-run public and hidden test sets.

full rationale

The paper's central claim is an empirical leaderboard result: the two-stage adapter-pretraining plus LoRA recipe, with horizontally stitched multi-image inputs, attains F1-RadGraph scores of 24.13 (Findings) and 22.10 (Impressions) on the hidden test set. Nothing in the paper defines a model component in terms of these scores, fits a parameter to the reported test values, or derives the evaluation numbers from the training objective by construction. The training loss is standard cross-entropy; the reported metrics (BLEU4, ROUGEL, BERTscore, F1-cheXbert, F1-RadGraph) are computed externally by the shared-task evaluation pipeline on held-out data, not optimized directly during training. The architectural choices (CLIP encoder, Vicuna-7B, MLP adapter, LoRA, four-image horizontal stitching) are empirical design decisions rather than quantities solved for from the target outputs. The only citation involving an author of this paper is Long et al. (2024) for cross-entropy loss, which is not load-bearing. The identical Findings F1-RadGraph value on public and hidden sets (24.13) is a possible data-reporting inconsistency and should be verified against the official leaderboard, but that is a correctness and data-integrity concern, not circularity: the hidden-test score is not definitionally equal to any training input. The limitations section openly acknowledges factors such as disease prevalence and multi-image inference weaknesses, which further supports that the authors are not claiming a derivation. No circular step can be exhibited, so the score is 0.

Assumptions & free parameters 5 free parameters · 5 assumptions · 0 invented entities

All quantities are engineering hyperparameters chosen by hand or copied from LLaVA recipes. No new theoretical entities are introduced. The main assumptions are that the workshop dataset is reliable, that lexical and factual metrics reflect clinical correctness, that frozen CLIP features carry enough X-ray information, and that image stitching preserves clinically relevant content; the last one is ad hoc and untested.

free parameters (5)
  • max_input_images = 4
    Chosen by hand in Section 4.1; no ablation showing first-four truncation preserves report-relevant findings.
  • max_sequence_length = 1024
    Set in Section 4.1 to minimize computation, based on word count distribution; affects training and inference.
  • max_decode_tokens = 150
    Inference decoding length set in Section 4.3 to match the leaderboard baseline; affects output completeness.
  • learning_rate = 1e-5
    Training hyperparameter in Section 4.3, chosen with cosine schedule and warmup; standard but not derived.
  • training_epochs = 1 adapter epoch + 3 LoRA epochs
    Section 3.3 follows the LLaVA-1.5 protocol rather than being optimized or justified by experiments in this paper.
assumptions (5)
  • domain assumption The RRG24 dataset text labels and image-report alignments are correct and the official split is not leaked into training.
    Section 4.1; all training and evaluation uses the workshop-provided dataset; if labels are noisy or split leaks, scores are unreliable.
  • domain assumption F1-RadGraph and F1-CheXbert are valid proxies for clinical correctness of generated reports.
    Section 4.2; the paper adopts these metrics from RRG24 guidelines without validating them on this dataset.
  • domain assumption CLIP visual features contain enough chest X-ray information for report generation after adapter and LoRA tuning.
    Section 3.2; the vision encoder is frozen CLIP pretrained on natural images, not on X-rays; no X-ray-specific encoder comparison is made.
  • ad hoc to paper Horizontal stitching and first-four truncation preserve the clinically relevant image content.
    Section 4.1; asserted as 'robust' in experiments but no ablation is provided; this is the weakest assumption.
  • standard math Cross-entropy loss with standard backpropagation is an appropriate training objective for this generation task.
    Section 3; this is a standard machine-learning assumption, not specific to radiology.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Gla-AI4BioMed at RRG24: Visual Instruction-tuned Adaptation for Radiology Report Generation." pith.science (2026). https://pith.science/paper/IE6E4GY6

@misc{pith2026241204954,
  author       = {Pith},
  title        = {Pith review of: Gla-AI4BioMed at RRG24: Visual Instruction-tuned Adaptation for Radiology Report Generation},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/IE6E4GY6}},
  note         = {Machine review of arXiv:2412.04954}
}
read the original abstract

We introduce a radiology-focused visual language model designed to generate radiology reports from chest X-rays. Building on previous findings that large language models (LLMs) can acquire multimodal capabilities when aligned with pretrained vision encoders, we demonstrate similar potential with chest X-ray images. This integration enhances the ability of model to understand and describe chest X-ray images. Our model combines an image encoder with a fine-tuned LLM based on the Vicuna-7B architecture, enabling it to generate different sections of a radiology report with notable accuracy. The training process involves a two-stage approach: (i) initial alignment of chest X-ray features with the LLM (ii) followed by fine-tuning for radiology report generation.

Figures

Figures reproduced from arXiv: 2412.04954 by the authors.

Figure 1
Figure 1. Our two-stage training framework. In the first stage, visual features are aligned with LLM. In the second [PITH_FULL_IMAGE:figures/full_fig_p004_1.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

51 extracted references · 10 canonical work pages

  1. [1]

    Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katie Millican, Malcolm Reynolds, Roman Ring, Eliza Rutherford, Serkan Cabi, Tengda Han, Zhitao Gong, Sina Samangooei, Marianne Monteiro, Jacob Menick, Sebastian Borgeaud, Andrew Brock, Aida Nematzadeh, Sahand Sharifzadeh, Mikolaj Binkowski,...

  2. [2]

    Dai, Orhan Firat, Melvin Johnson, Dmitry Lepikhin, Alexandre Passos, Siamak Shakeri, Emanuel Taropa, Paige Bailey, Zhifeng Chen, Eric Chu, Jonathan H

    Rohan Anil, Andrew M. Dai, Orhan Firat, Melvin Johnson, Dmitry Lepikhin, Alexandre Passos, Siamak Shakeri, Emanuel Taropa, Paige Bailey, Zhifeng Chen, Eric Chu, Jonathan H. Clark, Laurent El Shafey, Yanping Huang, Kathy Meier-Hellstern, Gaurav Mishra, Erica Moreira, Mark Omernick, Kevin Robinson, Sebastian Ruder, Yi Tay, Kefan Xiao, Yuanzhong Xu, Yujing Z...

  3. [3]

    Anas Awadalla, Irena Gao, Josh Gardner, Jack Hessel, Yusuf Hanafy, Wanrong Zhu, Kalyani Marathe, Yonatan Bitton, Samir Gadre, Shiori Sagawa, Jenia Jitsev, Simon Kornblith, Pang Wei Koh, Gabriel Ilharco, Mitchell Wortsman, and Ludwig Schmidt. 2023. https://arxiv.org/abs/2308.01390 Openflamingo: An open-source framework for training large autoregressive vis...

  4. [4]

    Castro, Benedikt Boecking, Harshita Sharma, Kenza Bouzid, Anja Thieme, Anton Schwaighofer, Maria Wetscherek, Matthew P

    Shruthi Bannur, Stephanie Hyland, Qianchu Liu, Fernando Pérez-García, Maximilian Ilse, Daniel C. Castro, Benedikt Boecking, Harshita Sharma, Kenza Bouzid, Anja Thieme, Anton Schwaighofer, Maria Wetscherek, Matthew P. Lungren, Aditya Nori, Javier Alvarez-Valle, and Ozan Oktay. 2023. https://arxiv.org/abs/2301.04558 Learning to exploit temporal structure fo...

  5. [5]

    Aurelia Bustos, Antonio Pertusa, Jose-Maria Salinas, and Maria de la Iglesia-Vayá. 2020. https://doi.org/10.1016/j.media.2020.101797 Padchest: A large chest x-ray image dataset with multi-label annotated reports . Medical Image Analysis, 66:101797

  6. [6]

    Langlotz

    Pierre Chambon, Jean-Benoit Delbrouck, Thomas Sounack, Shih-Cheng Huang, Zhihong Chen, Maya Varma, Steven QH Truong, Chu The Chuong, and Curtis P. Langlotz. 2024. https://arxiv.org/abs/2405.19538 Chexpert plus: Augmenting a large chest x-ray dataset with text radiology reports, patient demographics and additional image formats . Preprint, arXiv:2405.19538

  7. [7]

    Yu, Qiang Yang, and Xing Xie

    Yupeng Chang, Xu Wang, Jindong Wang, Yuan Wu, Linyi Yang, Kaijie Zhu, Hao Chen, Xiaoyuan Yi, Cunxiang Wang, Yidong Wang, Wei Ye, Yue Zhang, Yi Chang, Philip S. Yu, Qiang Yang, and Xing Xie. 2023. https://arxiv.org/abs/2307.03109 A survey on evaluation of large language models . Preprint, arXiv:2307.03109

  8. [8]

    Gonzalez, Ion Stoica, and Eric P

    Wei-Lin Chiang, Zhuohan Li, Zi Lin, Ying Sheng, Zhanghao Wu, Hao Zhang, Lianmin Zheng, Siyuan Zhuang, Yonghao Zhuang, Joseph E. Gonzalez, Ion Stoica, and Eric P. Xing. 2023. https://lmsys.org/blog/2023-03-30-vicuna/ Vicuna: An open-source chatbot impressing gpt-4 with 90\

Show all 51 references
  1. [9]

    Wenliang Dai, Junnan Li, Dongxu Li, Anthony Tiong, Junqi Zhao, Weisheng Wang, Boyang Li, Pascale Fung, and Steven Hoi. 2023. https://openreview.net/forum?id=vvoWPYqZJA Instruct BLIP : Towards general-purpose vision-language models with instruction tuning . In Thirty-seventh Co...

  2. [10]

    Maria de la Iglesia Vayá, Jose Manuel Saborit, Joaquim Angel Montell, Antonio Pertusa, Aurelia Bustos, Miguel Cazorla, Joaquin Galant, Xavier Barber, Domingo Orozco-Beltrán, Francisco García-García, Marisa Caparrós, Germán González, and Jose María Salinas. 2020. https://arxiv....

  3. [11]

    Jean-Benoit Delbrouck, Pierre Chambon, Christian Bluethgen, Emily Tsai, Omar Almusa, and Curtis Langlotz. 2022 a . https://doi.org/10.18653/v1/2022.findings-emnlp.319 Improving the factual correctness of radiology report generation with semantic rewards . In Findings of the As...

  4. [12]

    Jean-benoit Delbrouck, Khaled Saab, Maya Varma, Sabri Eyuboglu, Pierre Chambon, Jared Dunnmon, Juan Zambrano, Akshay Chaudhari, and Curtis Langlotz. 2022 b . https://doi.org/10.18653/v1/2022.acl-demo.3 V i LM edic: a framework for research at the intersection of vision and lan...

  5. [13]

    Danny Driess, Fei Xia, Mehdi S. M. Sajjadi, Corey Lynch, Aakanksha Chowdhery, Brian Ichter, Ayzaan Wahid, Jonathan Tompson, Quan Vuong, Tianhe Yu, Wenlong Huang, Yevgen Chebotar, Pierre Sermanet, Daniel Duckworth, Sergey Levine, Vincent Vanhoucke, Karol Hausman, Marc Toussaint...

  6. [14]

    Daisuke Endo, Ryota Kobayashi, Ramon Bartolo, Bruno B Averbeck, Yasuko Sugase-Miyamoto, Kazuko Hayashi, Kenji Kawano, Barry J Richmond, and Shigeru Shinomoto. 2021. https://doi.org/10.1038/s41598-021-91244-w A convolutional neural network for estimating synaptic connectivity f...

  7. [15]

    Ryann L Engle, David C Mohr, Sally K Holmes, Marjorie Nealon Seibert, Melissa Afable, Jenniffer Leyson, and Mark Meterko. 2021. https://doi.org/10.1097/HMR.0000000000000254 Evidence-based practice and patient-centered care: doing both well . Health care management review, 46(3...

  8. [16]

    Alex Graves. 2014. https://arxiv.org/abs/1308.0850 Generating sequences with recurrent neural networks . Preprint, arXiv:1308.0850

  9. [17]

    Dan Hendrycks and Kevin Gimpel. 2023. https://arxiv.org/abs/1606.08415 Gaussian error linear units (gelus) . Preprint, arXiv:1606.08415

  10. [18]

    Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen

    Edward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. 2021. https://arxiv.org/abs/2106.09685 Lora: Low-rank adaptation of large language models . Preprint, arXiv:2106.09685

  11. [19]

    Daniel T Huff, Amy J Weisman, and Robert Jeraj. 2021. https://doi.org/10.1088/1361-6560/abcd17 Interpretation and visualization techniques for deep learning models in medical imaging . Physics in Medicine & Biology, 66(4):04TR01

  12. [20]

    Hyland, Shruthi Bannur, Kenza Bouzid, Daniel C

    Stephanie L. Hyland, Shruthi Bannur, Kenza Bouzid, Daniel C. Castro, Mercy Ranjit, Anton Schwaighofer, Fernando Pérez-García, Valentina Salvatelli, Shaury Srivastav, Anja Thieme, Noel Codella, Matthew P. Lungren, Maria Teodora Wetscherek, Ozan Oktay, and Javier Alvarez-Valle. ...

  13. [21]

    Jaehwan Jeong, Katherine Tian, Andrew Li, Sina Hartung, Fardad Behzadi, Juan Calle, David Osayande, Michael Pohlen, Subathra Adithan, and Pranav Rajpurkar. 2023. https://arxiv.org/abs/2303.17579 Multimodal image-text matching improves retrieval-based chest x-ray report generat...

  14. [22]

    Haibo Jin, Haoxuan Che, Yi Lin, and Hao Chen. 2024. https://arxiv.org/abs/2308.12604 Promptmrg: Diagnosis-driven prompts for medical report generation . Preprint, arXiv:2308.12604

  15. [23]

    Alistair EW Johnson, Tom J Pollard, Seth J Berkowitz, Nathaniel R Greenbaum, Matthew P Lungren, Chih-ying Deng, Roger G Mark, and Steven Horng. 2019. https://doi.org/10.1038/s41597-019-0322-0 Mimic-cxr, a de-identified publicly available database of chest radiographs with free...

  16. [24]

    Kahn, Curtis P

    Charles E. Kahn, Curtis P. Langlotz, Elizabeth S. Burnside, John A. Carrino, David S. Channin, David M. Hovsepian, and Daniel L. Rubin. 2009. https://doi.org/10.1148/radiol.2523081992 Toward best practices in radiology reporting . RADIOLOGY, 252(3):852--856

  17. [25]

    Chunyuan Li, Cliff Wong, Sheng Zhang, Naoto Usuyama, Haotian Liu, Jianwei Yang, Tristan Naumann, Hoifung Poon, and Jianfeng Gao. 2023. https://arxiv.org/abs/2306.00890 Llava-med: Training a large language-and-vision assistant for biomedicine in one day . Preprint, arXiv:2306.00890

  18. [26]

    Chin-Yew Lin. 2004. https://aclanthology.org/W04-1013 ROUGE : A package for automatic evaluation of summaries . In Text Summarization Branches Out, pages 74--81, Barcelona, Spain. Association for Computational Linguistics

  19. [27]

    Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. 2023. https://arxiv.org/abs/2304.08485 Visual instruction tuning . Preprint, arXiv:2304.08485

  20. [28]

    Zhengliang Liu, Aoxiao Zhong, Yiwei Li, Longtao Yang, Chao Ju, Zihao Wu, Chong Ma, Peng Shu, Cheng Chen, Sekeun Kim, Haixing Dai, Lin Zhao, Lichao Sun, Dajiang Zhu, Jun Liu, Wei Liu, Dinggang Shen, Xiang Li, Quanzheng Li, and Tianming Liu. 2024. https://arxiv.org/abs/2306.0866...

  21. [29]

    Zijun Long, George Killick, Lipeng Zhuang, Gerardo Aragon-Camarasa, Zaiqiao Meng, and Richard Mccreadie. 2024. https://arxiv.org/abs/2402.14551 Clce: An approach to refining cross-entropy and contrastive learning for optimized learning fusion . Preprint, arXiv:2402.14551

  22. [30]

    Yasuhide Miura, Yuhao Zhang, Emily Tsai, Curtis Langlotz, and Dan Jurafsky. 2021. https://doi.org/10.18653/v1/2021.naacl-main.416 Improving factual completeness and consistency of image-to-text radiology report generation . In Proceedings of the 2021 Conference of the North Am...

  23. [31]

    Maram Mahmoud A Monshi, Josiah Poon, and Vera Chung. 2020. https://doi.org/10.1016/j.artmed.2020.101878 Deep learning in generating radiology reports: A survey . Artificial Intelligence in Medicine, 106:101878

  24. [32]

    Michael Moor, Qian Huang, Shirley Wu, Michihiro Yasunaga, Cyril Zakka, Yash Dalmia, Eduardo Pontes Reis, Pranav Rajpurkar, and Jure Leskovec. 2023. https://arxiv.org/abs/2307.15189 Med-flamingo: a multimodal medical few-shot learner . Preprint, arXiv:2307.15189

  25. [33]

    Aaron Nicolson, Jason Dowling, and Bevan Koopman. 2023. https://doi.org/10.1016/j.artmed.2023.102633 Improving chest x-ray report generation by leveraging warm starting . Artificial Intelligence in Medicine, 144:102633

  26. [34]

    OpenAI, Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, Red Avila, Igor Babuschkin, Suchir Balaji, Valerie Balcom, Paul Baltescu, Haiming Bao, Mohammad Bavarian, Jeff ...

  27. [35]

    Panayides, Amir Amini, Nenad D

    Andreas S. Panayides, Amir Amini, Nenad D. Filipovic, Ashish Sharma, Sotirios A. Tsaftaris, Alistair Young, David Foran, Nhan Do, Spyretta Golemati, Tahsin Kurc, Kun Huang, Konstantina S. Nikita, Ben P. Veasey, Michalis Zervakis, Joel H. Saltz, and Constantinos S. Pattichis. 2...

  28. [36]

    Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002. https://doi.org/10.3115/1073083.1073135 Bleu: a method for automatic evaluation of machine translation . In Proceedings of the 40th Annual Meeting on Association for Computational Linguistics, ACL '02, page 311...

  29. [37]

    Sang-Min Park and Young-Gab Kim. 2023. https://doi.org/10.1016/j.cosrev.2023.100548 Visual language integration: A survey and open challenges . Computer Science Review, 48:100548

  30. [38]

    Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. 2021. https://arxiv.org/abs/2103.00020 Learning transferable visual models from natural lan...

  31. [39]

    Samyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, and Yuxiong He. 2020. https://arxiv.org/abs/1910.02054 Zero: Memory optimizations toward training trillion parameter models . Preprint, arXiv:1910.02054

  32. [40]

    Andrew B. Sellergren, Christina Chen, Zaid Nabulsi, Yuanzhen Li, Aaron Maschinot, Aaron Sarna, Jenny Huang, Charles Lau, Sreenivasa Raju Kalidindi, Mozziyar Etemadi, Florencia Garcia-Vicente, David Melnick, Yun Liu, Krish Eswaran, Daniel Tse, Neeral Beladia, Dilip Krishnan, an...

  33. [41]

    Ng, and Matthew P

    Akshay Smit, Saahil Jain, Pranav Rajpurkar, Anuj Pareek, Andrew Y. Ng, and Matthew P. Lungren. 2020. https://arxiv.org/abs/2004.09167 Chexbert: Combining automatic labelers and expert annotations for accurate radiology report labeling using bert . Preprint, arXiv:2004.09167

  34. [42]

    Tim Tanida, Philip Müller, Georgios Kaissis, and Daniel Rueckert. 2023. https://doi.org/10.1109/cvpr52729.2023.00718 Interactive and explainable region-guided radiology report generation . In 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE

  35. [43]

    Rohan Taori, Ishaan Gulrajani, Tianyi Zhang, Yann Dubois, Xuechen Li, Carlos Guestrin, Percy Liang, and Tatsunori B Hashimoto. 2023. https://crfm. stanford. edu/2023/03/13/alpaca. html Alpaca: A strong, replicable instruction-following model . Stanford Center for Research on F...

  36. [44]

    Corrado, Yossi Matias, Karan Singhal, Pete Florence, Alan Karthikesalingam, and Vivek Natarajan

    Tao Tu, Shekoofeh Azizi, Danny Driess, Mike Schaekermann, Mohamed Amin, Pi-Chuan Chang, Andrew Carroll, Chuck Lau, Ryutaro Tanno, Ira Ktena, Basil Mustafa, Aakanksha Chowdhery, Yun Liu, Simon Kornblith, David Fleet, Philip Mansfield, Sushant Prakash, Renee Wong, Sunny Virmani,...

  37. [45]

    Collins, Ankit Modi, Robert Lloyd, Benjamin Hopkins, Curtis Langlotz, and Jean-Benoit Delbrouck

    Justin Xu, Zhihong Chen, Andrew Johnston, Louis Blankemeier, Maya Varma, Jason Hom, William J. Collins, Ankit Modi, Robert Lloyd, Benjamin Hopkins, Curtis Langlotz, and Jean-Benoit Delbrouck. 2024. Overview of the first shared task on clinical text generation: Rrg24 and `` dis...

  38. [46]

    Corrado, Shravya Shetty, Daniel Tse, Shruthi Prabhakara, Daniel Golden, Rory Pilgrim, Krish Eswaran, and Andrew Sellergren

    Shawn Xu, Lin Yang, Christopher Kelly, Marcin Sieniek, Timo Kohlberger, Martin Ma, Wei-Hung Weng, Atilla Kiraly, Sahar Kazemzadeh, Zakkai Melamed, Jungyeon Park, Patricia Strachan, Yun Liu, Chuck Lau, Preeti Singh, Christina Chen, Mozziyar Etemadi, Sreenivasa Raju Kalidindi, Y...

  39. [47]

    Kuo, Subathra Adithan, Eduardo Pontes Reis, Stephen Kwak, Vasantha Kumar Venugopal, Chloe P

    Benjamin Yan, Ruochen Liu, David E. Kuo, Subathra Adithan, Eduardo Pontes Reis, Stephen Kwak, Vasantha Kumar Venugopal, Chloe P. O'Connell, Agustina Saenz, Pranav Rajpurkar, and Michael Moor. 2023. https://arxiv.org/abs/2310.17811 Style-aware radiology report generation with r...

  40. [48]

    Feiyang Yu, Mark Endo, Rayan Krishnan, Ian Pan, Andy Tsai, Eduardo Pontes Reis, Eduardo Kaiser Ururahy Nunes Fonseca, Henrique Min Ho Lee, Zahra Shakeri Hossein Abad, Andrew Y Ng, et al. 2023. https://doi.org/10.1016/j.patter.2023.100802 Evaluating progress in automatic chest ...

  41. [49]

    Weinberger, and Yoav Artzi

    Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger, and Yoav Artzi. 2020. https://arxiv.org/abs/1904.09675 Bertscore: Evaluating text generation with bert . Preprint, arXiv:1904.09675

  42. [50]

    online" 'onlinestring :=

    ENTRY address archivePrefix author booktitle chapter edition editor eid eprint eprinttype howpublished institution journal key month note number organization pages publisher school series title type volume year doi pubmed url lastchecked label extra.label sort.label short.list...

  43. [51]

    write newline

    " write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 gl...

Pith tools

Reviewed August 11, 2026 · model on record in the stance chip above.