Pith. sign in

REVIEW 4 major objections 6 minor 1 cited by

Natural Language Generation in Healthcare: A Review of Methods and Applications

T0 review · 4 major / 6 minor · reviewed 2026-08-15 · deepseek-v4-flash

Pith's one-line read This review claims the first systematic, multi-modality map of natural language generation in healthcare, covering 113 studies.

desk verdict A useful PRISMA-structured review of 113 medical NLG papers whose modality taxonomy has a genuine internal contradiction and some counting inconsistencies — fixable, not fatal. read the letter →

arxiv 2505.04073 v1 pith:XMYF2BXX submitted 2025-05-07 cs.CL

classification cs.CL
keywords naturallanguagegenerationhealthcarelargemodelssystematicreviewclinicaldocumentationradiologyreportmedicaldialoguesystemsevaluationmetrics
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper tries to establish a comprehensive, PRISMA-based map of natural language generation (NLG) in healthcare, claiming to be the first review to cover multiple medical data modalities and applications together. From 3,988 candidate articles, the authors screened down to 113 studies and grouped them by input type (text-to-text, image-to-text, multimodal-to-text), model architecture, clinical application, and evaluation method. The central result is a set of prevalence statistics: transformer encoder-decoder models dominate text-to-text generation, CNN+Transformer hybrids lead image-to-text work, and ROUGE, BLEU, and Likert-scale human ratings are the most common evaluation tools. A sympathetic reader would care because the paper turns a scattered field into reference numbers and a structured research agenda for clinical generative AI.

What carries the argument

The organizing machinery is a three-way taxonomy of input modality (text-to-text, image-to-text, multimodal-to-text) crossed with model architecture (encoder-decoder, decoder-only, and GAN-based), clinical application (summarization, documentation, dialogue, data augmentation), and evaluation method (n-gram, embedding-based, and human). This taxonomy is what turns 113 individual papers into aggregate prevalence statistics, so the whole review's descriptive claims rest on the categories being coherent and exhaustive.

What would settle it

Take the 113 included studies and re-code each one by input type into four categories—free text, structured/tabular EHR data, images, and multimodal—then compare the resulting prevalence numbers with Table 1. If a substantial share of the text-to-text count moves into a data-to-text category (for example, studies that generate notes from diagnosis codes or lab values), the review's central taxonomic claim is contradicted by its own data.

Watch

Extended reading notes

Core claim

On its own terms, the paper's discovery is that NLG in healthcare has consolidated around transformer-based architectures across three input-modality classes, and that these classes map onto four clinical application families: summarization, automated documentation, dialogue generation, and data augmentation. The paper reports that 61.64% of text-to-text studies use encoder-decoder transformers, decoder-only transformers account for 27.40%, and image-to-text work relies heavily on CNN+Transformer hybrids, while evaluation is dominated by surface-level metrics (ROUGE in 74 studies, BLEU in 60) supplemented by Likert-scale human ratings in 36 studies. It further claims that this is the first systematic review of NLG that spans multiple medical data modalities and multiple healthcare applications, and it identifies synthetic clinical text generation as a promising route to privacy-preserving data scaling.

Load-bearing premise

The whole classification depends on the three input-modality categories being exhaustive and on structured EHR data counting as text-to-text input, even though the introduction defines generation from non-linguistic input as data-to-text; if that labeling is inconsistent, the prevalence table and the grouped summaries stop being a reliable map of the field.

Editorial extensions

If this is right

  • If the taxonomy holds, Table 1's percentages become the field's reference statistics for where NLG research effort concentrates.
  • The dominance of ROUGE and BLEU suggests reported quality may miss clinical relevance and factual accuracy, motivating better evaluation metrics.
  • Synthetic clinical text generation, positioned as privacy-preserving data augmentation, is likely to grow as a way to scale training data without sharing patient records.
  • Multimodal-to-text generation, especially image-plus-structured-data fusion, is a rising pattern that could define the next phase of medical report generation.
  • The application map gives researchers and clinicians a structured menu for matching a clinical documentation problem to an existing architecture and evaluation setup.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • My inference: the text-to-text category probably absorbs data-to-text systems (structured EHR-to-narrative), because the paper's own introduction defines non-linguistic input as data-to-text; a cleaner four-way split would likely shift some prevalence numbers.
  • My inference: because most studies rely on n-gram overlap metrics, the review likely overstates how close the field is to clinically usable generation; a re-analysis correlating automatic scores with expert clinical judgment could test this.
  • My inference: the exclusion of non-English studies and of ethics/safety-only papers means the map is strongest for English-language, application-focused work; low-resource language NLG is a visible gap rather than a surveyed area.
  • My inference: the PRISMA-style pipeline could be rerun periodically to track how fast each modality and application category is growing, making this a baseline rather than a one-time snapshot.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 6 minor

Summary. This manuscript presents a PRISMA-guided systematic review of natural language generation in healthcare. From 3,988 records screened through title/abstract and full-text stages, 113 studies are included and analyzed along four dimensions: input data modality, model architecture, clinical application, and evaluation metrics. The review reports that text-to-text, image-to-text, and multimodal-to-text are the three input-modality categories, identifies encoder-decoder and decoder-only transformers as dominant architectures, lists ROUGE/BLEU as the most common automatic metrics, and organizes applications into summarization, automated documentation, data augmentation, and medical dialogue. The central claim is that this is the first systematic review of NLG across multiple medical data modalities and applications.

Significance. The review is potentially useful as a structured inventory for researchers entering clinical NLG: the search protocol is described in unusual detail (seven databases, two-stage screening with two reviewers per article, training sessions, standardized extraction), and the paper ships concrete prevalence tables and a PRISMA flow diagram. If the modality taxonomy were coherent, Table 1 and Table 2 could serve as reference statistics. The claim to be the first multi-modality systematic review is plausible but not independently verifiable from the manuscript, and it is currently compromised by internal inconsistencies in the modality categories and in the counts, so the central contribution needs repair before the significance can be realized.

major comments (4)
  1. [Introduction and 'Medical data modalities and text generation applications'] The modality taxonomy is internally inconsistent. The Introduction (p. 4) defines 'data-to-text' generation as generating text from non-linguistic input and explicitly lists structured EHR elements in database tables as a medical data modality; the Results section then groups the 113 studies into text-to-text, image-to-text, and multimodal-to-text only, and defines text-to-text as 'generating narrative text from textual input, including structured data (e.g., diagnosis codes, medications), semi-structured data (e.g., templates or tabular EHR entries)'. This classifies studies whose inputs are ICD codes, lab values, or structured EHR tables as text-to-text, even though the Introduction's own definition makes them data-to-text. Examples in the text are Lee et al. (ref 25, FNN-to-LSTM from structured EHR codes), J Kurisinkel & Chen (ref 34, ICD codes to discharge instructions), and Soni & Dina (ref 92, structured/tabular EHR to progress notes). Because the review's stated contribution is a modality-based map, and structured EHRs are named as one of the main medical data modalities, the three-way taxonomy omits the data-to-text category and the Table 1 prevalence counts do not cleanly support the claimed comprehensive overview. The authors should either reintroduce data-to-text as a separate category or explicitly redefine text-to-text to include non-linguistic structured inputs and justify that choice.
  2. [Table 1 and text after Table 1] The prevalence numbers do not add up. The text states '61.64% (45/73) of studies used encoder-decoder transformer models, and 27.40% (20/73) of studies used decoder-only transformer models' in text-to-text generation, but Table 1 lists 74 text-to-text rows (1 GAN + 1 FNN/LSTM + 1 RNN/RNN + 3 GRU/GRU + 4 LSTM/LSTM + 1 LSTM/Transformer + 42 Transformer/Transformer + 1 no-encoder/LSTM + 20 decoder-only Transformer = 74), and no obvious reading yields 45 encoder-decoder transformer studies. The denominator should be a single consistent count, and all percentages should be recomputed from the corrected table.
  3. [Table 2 and 'Image-to-Text Related Metrics'] Table 2 is incomplete relative to the text. The text states that CIDEr is 'widely used in image-to-text generation evaluation' and that pairwise comparison is 'another widely used human evaluation metric', yet neither CIDEr nor pairwise comparison appears in Table 2's counts. Since evaluating metrics is one of the four stated review dimensions, Table 2 should either include these metrics with counts or the text should be revised to avoid claiming they are widely used.
  4. [Definition of multimodal-to-text in Results] The three categories are not mutually exclusive under the given definitions. If structured EHR entries count as text, then a system taking both an image and structured clinical variables is simultaneously text-to-text and multimodal-to-text; if structured entries do not count as text, then the examples in the text-to-text paragraph contradict the definition. The taxonomy should specify whether modality labels are based on the input's linguistic status, its source, or the number of input channels, and Table 1's assignments should follow that rule.
minor comments (6)
  1. [Reference list] Reference list entries 79–81, 84–89, 92, 105, 106, 108, and 111 are incomplete (missing authors, venues, or titles); several appear as bare shared-task names. Please restore full bibliographic details.
  2. [Fig. 3] Figure 3's 'Text-to-Text Generation' panel includes a 'Structural EHR' table with PT_ID, TEST, ENC_DATE, VALUE; once the taxonomy is revised, the figure should be updated to match the corrected modality definitions.
  3. [Limitation section] The sentence 'studies that about the potential bias, security, risks, interpretability, and AI ethics, but without concrete applications were not included' is ungrammatical and should be rewritten for clarity.
  4. [Discussion] The sentence 'We have almost exhausted all the electronic data generated on this planet' is an unsupported overstatement and should be removed or supported with a citation.
  5. [Table 1] In Table 1, the row 'CNN +GAN Transformer' is ambiguous; clarify whether the encoder is CNN+GAN and the decoder is Transformer, and consider separating the entries.
  6. [Evaluation metrics text] The sentence 'Both metrics have been utilized in our reviewed studies62–65' lists four references for two metrics; verify the citation grouping so each metric is supported by the intended studies.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the review's synthesis is independent of the authors' prior work; the one self-citation is not load-bearing.

full rationale

This paper is a PRISMA-based systematic review, not a derivation chain. It fits no parameters, makes no predictions, and proves no theorems; its claims are prevalence counts and qualitative categorizations of 113 external studies. The central claim (first systematic review of NLG across multiple medical data modalities and healthcare applications) is supported by the documented search and screening process (3,988 articles identified, 2,903 excluded at title/abstract, 211 excluded at full-text review, six reviewers with consensus resolution), not by any author-derived result. The only in-house reference is Peng et al. (ref 7), cited descriptively as one included study for decoder-only transformers, synthetic clinical data generation, and a Turing test example; the review's conclusions do not depend on this citation. The modality taxonomy inconsistency noted in the manuscript (the Introduction defines data-to-text as generation from non-linguistic input, while the 'Medical data modalities and text generation applications' section folds structured EHR data into text-to-text) is a substantive correctness and validity concern, but it is not circularity under the rubric: it is an internal classification contradiction, not an input recycled as a prediction or a conclusion forced by self-citation. Because no specific reduction can be exhibited, no circular step is recorded and the score is 0.

Assumptions & free parameters 0 free parameters · 4 assumptions · 0 invented entities

The central claim is a literature synthesis, so its axioms are methodological assumptions about search completeness, screening validity, taxonomy coherence, and extraction accuracy. There are no free parameters or invented entities because the paper contains no fitted model or new theoretical construct.

assumptions (4)
  • domain assumption The seven searched databases and keyword strings identify a representative sample of NLG-in-healthcare research published 2018-2024.
    The review's claim of comprehensiveness rests on the search being complete; coverage of major databases and the ACL Anthology is plausible but not formally demonstrated to be exhaustive.
  • domain assumption Screening criteria correctly separate studies with concrete NLG methods or applications from those without.
    Two reviewers with third-party adjudication were used, but no inter-rater agreement or audit trail is reported, and the excluded 211 full-text papers are not characterized.
  • domain assumption The taxonomy (text-to-text, image-to-text, multimodal-to-text; four application categories) is exhaustive and externally meaningful.
    The categories organize the field, but the paper itself blurs text-to-text and data-to-text, so the taxonomy's coherence is an assumption rather than a demonstrated property.
  • domain assumption The extracted study characteristics from the 113 included papers are accurate.
    Data extraction was done by two reviewers, but Supplementary Table S1 with the full extraction is not present in the preprint, so accuracy cannot be checked by the reader.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Natural Language Generation in Healthcare: A Review of Methods and Applications." pith.science (2026). https://pith.science/paper/XMYF2BXX

@misc{pith2026250504073,
  author       = {Pith},
  title        = {Pith review of: Natural Language Generation in Healthcare: A Review of Methods and Applications},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/XMYF2BXX}},
  note         = {Machine review of arXiv:2505.04073}
}
read the original abstract

Natural language generation (NLG) is the key technology to achieve generative artificial intelligence (AI). With the breakthroughs in large language models (LLMs), NLG has been widely used in various medical applications, demonstrating the potential to enhance clinical workflows, support clinical decision-making, and improve clinical documentation. Heterogeneous and diverse medical data modalities, such as medical text, images, and knowledge bases, are utilized in NLG. Researchers have proposed many generative models and applied them in a number of healthcare applications. There is a need for a comprehensive review of NLG methods and applications in the medical domain. In this study, we systematically reviewed 113 scientific publications from a total of 3,988 NLG-related articles identified using a literature search, focusing on data modality, model architecture, clinical applications, and evaluation methods. Following PRISMA (Preferred Reporting Items for Systematic reviews and Meta-Analyses) guidelines, we categorize key methods, identify clinical applications, and assess their capabilities, limitations, and emerging challenges. This timely review covers the key NLG technologies and medical applications and provides valuable insights for future studies to leverage NLG to transform medical discovery and healthcare.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Clinical Communication Processing with Models Trained on LLM-Generated Synthetic Data: A Structured Survey and Novel Application Case Studies

    cs.CL 2026-08 conditional novelty 6.0 of 10

    Synthetic clinical communication generated by LLMs can train clinical NLP models in thirteen case studies, but only one is tested on real patient text, leaving transfer to authentic communication unproven.

Reference graph

Works this paper leans on

12 extracted references · 6 canonical work pages · cited by 1 Pith paper

  1. [11]

    Sutskever, I., Vinyals, O. & Le, Q. V. Sequence to sequence learning with Neural Networks. arXiv [cs.CL] (2014). 12. Vaswani, A. et al. Attention is all you need. arXiv [cs.CL] (2017). 13. Lin, T., Wang, Y., Liu, X. & Qiu, X. A survey of transformers. AI Open 3, 111–132 (2022). 14. Devlin, J., Chang, M.-W., Lee, K. & Toutanova, K. BERT: Pre-training of De...

  2. [21]

    Zhang, K. et al. A generalist vision-language foundation model for diverse biomedical tasks. Nat. Med. 30, 3129–3141 (2024). 22. Moor, M. et al. Foundation models for generalist medical artificial intelligence. Nature 616, 259–265 (2023). 23. Truhn, D., Eckardt, J.-N., Ferber, D. & Kather, J. N. Large language models and multimodal foundation models for p...

  3. [31]

    Fu, S. et al. Clinical concept extraction: A methodology review. J. Biomed. Inform. 109, 103526 (2020). 32. A Systematic Review of Deep Learning-based Research on Radiology Report Generation. A Systematic Review of Deep Learning-based Research on Radiology Report Generation. 33. Melamud, O. & Shivade, C. Towards Automatic Generation of Shareable Synthetic...

  4. [40]

    Harnessing the Power of Pre-trained Vision-Language Models for Efficient Medical Report Generation

    Li, Q. Harnessing the Power of Pre-trained Vision-Language Models for Efficient Medical Report Generation. Proceedings of the 32nd ACM International Conference on Information and Knowledge Management 1308–1317 (2023) doi:10.1145/3583780.3614961. 41. Delbrouck, J.-B., Zhang, C. & Rubin, D. QIAI at MEDIQA 2021: Multimodal Radiology Report Summarization. (As...

  5. [56]

    A Guided Clinical Abstractive Summarization Model for Generating Medical Reports from Patient-Doctor Conversations. 57. Guided Clinical Abstractive Summarization Model for Generating Medical Reports from Patient-Doctor Conversations. 58. Karn, S. K., Liu, N., Schuetze, H. & Farri, O. Differentiable Multi-Agent Actor-Critic for Multi-Step Radiology Report ...

  6. [65]

    & Rios, A

    Wang, T., Zhao, X. & Rios, A. UTSA-NLP at RadSum23: Multi-Modal Retrieval-Based Chest X-Ray Report Summarization. (Association for Computational Linguistics, 2023). doi:10.18653/v1/2023.bionlp-1.58. 66. Zhou, Y., Ringeval, F. & Portet, F. A survey of evaluation methods of generated medical textual reports. in Proceedings of the 5th Clinical Natural Langua...

  7. [72]

    & Wan, X

    Chen, Z., Song, Y., Chang, T.-H. & Wan, X. Generating Radiology Reports via Memory-Driven Transformer. (Association for Computational Linguistics, 2020). doi:10.18653/v1/2020.emnlp-main.112. 73. Kaur, N. & Mittal, A. CADxReport: Chest x-ray report generation using co-attention mechanism and reinforcement learning. Comput. Biol. Med. 145, 105498 (2022). 74...

  8. [82]

    : Evaluating Constrained Generation of Discharge Summaries with Unstructured and Structured Information. 89. At, U.-H. Discharge Me!

    Sharma, B. et al. Multi-Task Training with In-Domain Language Models for Diagnostic Reasoning. (Association for Computational Linguistics, 2023). doi:10.18653/v1/2023.clinicalnlp-1.10. 83. Ben Abacha, A., Yim, W.-W., Adams, G., Snider, N. & Yetisgen, M. Overview of the MEDIQA-Chat 2023 Shared Tasks on the Summarization & Generation of Doctor-Patient Conve...

Show all 12 references
  1. [91]

    & Jarrahi, M

    Willis, M. & Jarrahi, M. H. Automating documentation: A critical perspective into the role of artificial intelligence in clinical documentation. in Information in Contemporary Society 200–209 (Springer International Publishing, Cham, 2019). doi:10.1007/978-3-030-15742-5_19. 92...

  2. [100]

    Mondal, C. et al. EfficienTransNet: An Automated Chest X-ray Report Generation Paradigm. Proceedings of the 1st International Workshop on Multimodal and Responsible Affective Computing 59–66 (2023) doi:10.1145/3607865.3616174. 101. Yang, S. et al. Radiology report generation w...

  3. [109]

    Hu, Z., Zhao, H., Zhao, Y., Xu, S. & Xu, B. T-agent: A term-aware agent for medical dialogue generation. in 2024 International Joint Conference on Neural Networks (IJCNN) 1–8 (IEEE, 2024). doi:10.1109/ijcnn60899.2024.10650649. 110. Srivastava, A., Pandey, I., Akhtar, M. S. & C...

  4. [116]

    & Deng, C

    Yu, P., Xu, H., Hu, X. & Deng, C. Leveraging generative AI and large language models: A comprehensive roadmap for healthcare integration. Healthcare (Basel) 11, 2776 (2023). 117. Baumann, L. A., Baker, J. & Elshaug, A. G. The impact of electronic health record systems on clini...

Pith tools

Reviewed August 15, 2026 · model on record in the stance chip above.