Pith. sign in

REVIEW 2 major objections 2 minor 140 references

Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation

T0 review · 2 major / 2 minor · reviewed 2026-06-27 · grok-4.3

Pith's one-line read Human evaluation protocols for long-form text generation in recent NLP conference papers are often incompletely reported, creating ambiguity about what was measured and by whom.

desk verdict The paper gives the first big numbers on under-reporting in human eval protocols for long-form generation, but the 20 criteria are unvalidated so the strength of the 'widespread' claim is still open. read the letter →

arxiv 2606.07936 v2 pith:OWOLVEFR submitted 2026-06-06 cs.CL cs.AI

classification cs.CLcs.AI
keywords humanevaluationreproducibilitylong-formtextgenerationNLPconferencesreportingpracticesprotocolsunder-reportingstudydesign
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper examines reporting practices for human evaluations in papers on long-form text generation from *CL conferences between 2023 and 2025. It applies a defined set of 20 criteria for reproducibility to a manual review of 284 papers plus LLM-assisted analysis of more than 1800 additional papers. The analysis shows frequent omission of details on study design, participant information, judgment processes, and interpretation guidelines. This pattern leaves readers uncertain about the reliability of the evaluations and how to compare results across papers. The authors propose concrete steps to improve documentation in future work.

What carries the argument

A set of 20 reportable criteria related to reproducibility of human evaluation studies, applied to check what details papers include about design, participants, and judgment processes.

What would settle it

A re-analysis of the same papers that applies a different or expanded list of criteria and finds high rates of complete reporting on the missing items.

Watch

Extended reading notes

Core claim

A systematic review of human evaluation protocols in *CL publications reveals widespread under-reporting of important aspects of study design, who contributed judgments, and how judgments should be interpreted, which produces ambiguity about what was actually measured and how the results should be understood.

Load-bearing premise

The authors' chosen set of 20 criteria is enough to determine whether a human evaluation study is reproducible and interpretable.

Editorial extensions

If this is right

  • Adopting the 20 criteria would make it easier to interpret and compare human evaluation results across different papers.
  • Papers would need to document participant recruitment, training, and agreement measures more consistently.
  • Ambiguity in current evaluations would decrease if journals and conferences required explicit reporting on these points.
  • Future comparisons of generation systems could rest on clearer evidence of evaluation quality.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The same under-reporting pattern likely appears in evaluations of short-form or other generation tasks outside the long-form focus.
  • Conferences could reduce ambiguity by adding checklist items based on the 20 criteria during submission.
  • Greater transparency might shift research incentives toward more careful study design rather than just reporting results.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 2 minor

Summary. The paper conducts a large-scale observational analysis of human evaluation protocols for long-form text generation in *CL conference papers (2023–2025). It performs a full manual review of 284 papers plus LLM-assisted analysis of an additional 1.8k+ papers against a fixed set of 20 reportable criteria for reproducibility. The central claim is that widespread under-reporting of study-design details creates ambiguity about what was measured, who provided judgments, and how results should be interpreted; the authors provide recommendations and release code plus an annotated dataset.

Significance. If the central observational claim holds, the work is significant because human evaluation remains a primary method for assessing long-form generation quality, and documented under-reporting directly affects reproducibility and interpretability in the field. The explicit release of analysis code and the annotated dataset is a clear strength that supports verification and follow-up studies. The findings could usefully inform community guidelines, provided the 20 criteria are shown to be well-aligned with existing reporting standards.

major comments (2)
  1. [Section defining the 20 criteria] Section defining the 20 criteria: the manuscript presents these criteria as capturing 'important aspects' necessary for reproducibility and interpretability, yet provides no external validation (e.g., expert survey, comparison against ACL or prior meta-study reporting guidelines, or inter-rater agreement on criterion importance). Because the prevalence statistics and the downstream claim of 'widespread under-reporting of important aspects' rest directly on this author-defined set, the absence of such validation makes the quantitative conclusions sensitive to the particular framing chosen.
  2. [LLM-assisted analysis section] LLM-assisted analysis section (1.8k+ papers): the extension from the 284 manually reviewed papers to the larger corpus is load-bearing for the 'widespread' claim, but the manuscript does not report prompt details, few-shot examples, or measured agreement/error rates between the LLM outputs and the manual annotations. Without these, systematic biases in the automated labeling could materially affect the reported under-reporting rates.
minor comments (2)
  1. [Table 1] Table 1 (or equivalent summary table of criteria): the mapping from each criterion to the specific ambiguity it addresses (measurement, contributors, or interpretation) could be made more explicit to help readers trace how missing items produce the claimed ambiguities.
  2. [Data and code release] The GitHub link is provided, but the README should include a clear description of how the 284 manual annotations were performed (annotator background, resolution process) to strengthen reproducibility of the core dataset.

Simulated Author's Rebuttal

2 responses · 0 unresolved

We thank the referee for the constructive feedback. We address each major comment below and outline planned revisions to improve the manuscript.

read point-by-point responses
  1. Referee: [Section defining the 20 criteria] Section defining the 20 criteria: the manuscript presents these criteria as capturing 'important aspects' necessary for reproducibility and interpretability, yet provides no external validation (e.g., expert survey, comparison against ACL or prior meta-study reporting guidelines, or inter-rater agreement on criterion importance). Because the prevalence statistics and the downstream claim of 'widespread under-reporting of important aspects' rest directly on this author-defined set, the absence of such validation makes the quantitative conclusions sensitive to the particular framing chosen.

    Authors: We agree that additional justification for the criteria would strengthen the work. The 20 criteria were synthesized from recurring elements in prior NLP literature on human evaluation reproducibility and reporting standards. In revision we will add a new subsection explicitly mapping each criterion to relevant ACL guidelines and earlier meta-studies, together with a brief rationale for inclusion. While we did not conduct a new expert survey, this explicit alignment will reduce sensitivity to the chosen framing and make the prevalence claims more robust. revision: partial

  2. Referee: [LLM-assisted analysis section] LLM-assisted analysis section (1.8k+ papers): the extension from the 284 manually reviewed papers to the larger corpus is load-bearing for the 'widespread' claim, but the manuscript does not report prompt details, few-shot examples, or measured agreement/error rates between the LLM outputs and the manual annotations. Without these, systematic biases in the automated labeling could materially affect the reported under-reporting rates.

    Authors: We concur that full transparency on the LLM-assisted labeling is required. The original submission omitted these details for brevity. The revised manuscript will include the complete prompts, few-shot examples, and a dedicated error-analysis subsection reporting agreement rates (and disagreement categories) between the LLM and the manual annotations on a held-out validation set. This addition will allow readers to evaluate potential biases directly. revision: yes

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity in observational literature analysis

full rationale

The paper performs a manual and LLM-assisted review of reporting practices in other papers by first defining an explicit set of 20 criteria and then counting their presence or absence. No equations, fitted parameters, predictions, or self-citation chains exist that reduce any central claim to its own inputs by construction. The analysis is self-contained empirical observation against transparently stated criteria; the absence of any enumerated circularity pattern (self-definitional, fitted-input prediction, load-bearing self-citation, etc.) yields a score of 0.

Assumptions & free parameters 0 free parameters · 1 assumptions · 0 invented entities

The analysis depends on the domain assumption that the 20 criteria are adequate for assessing reproducibility; no free parameters or invented entities are introduced.

assumptions (1)
  • domain assumption The defined set of 20 reportable criteria adequately captures key aspects of human evaluation reproducibility.
    Criteria are introduced as the basis for the systematic examination of papers.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation." pith.science (2026). https://pith.science/paper/OWOLVEFR

@misc{pith2026260607936,
  author       = {Pith},
  title        = {Pith review of: Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/OWOLVEFR}},
  note         = {Machine review of arXiv:2606.07936}
}
read the original abstract

Human evaluation plays a critical role in assessing the quality of generated text. However, the reliability and reproducibility of these evaluations depend on transparent and well-documented protocols -- details that are frequently missing in current practice. In this work, we conduct a large-scale analysis of human evaluation protocols for evaluating long-form generation tasks in *CL conference publications from 2023--2025, including a full manual review of 284 papers and LLM-assisted analysis for another 1.8k+ papers. We define a set of 20 reportable criteria related to reproducibility of human evaluation studies, and apply these criteria to systematically examine reporting norms and practices within the community. We find widespread under-reporting of important aspects of human evaluation study design, leading to ambiguity about what was measured and how, who contributed judgments, and how judgments should be interpreted. Based on these findings, we outline actionable recommendations to support more transparent and reproducible reporting in future research. Our analysis code and annotated dataset can be found at: https://github.com/larchlab/Illusions-of-the-Gold-Standard

Figures

Figures reproduced from arXiv: 2606.07936 by the authors.

Figure 1
Figure 1. Average proportion of *CL papers reporting each of 20 core criteria related to the reproducibility of [PITH_FULL_IMAGE:figures/full_fig_p003_1.png] view at source ↗
Figure 2
Figure 2. Distribution of total criteria reported; over [PITH_FULL_IMAGE:figures/full_fig_p007_2.png] view at source ↗
Figure 3
Figure 3. Distributions of annotator and sample counts [PITH_FULL_IMAGE:figures/full_fig_p007_3.png] view at source ↗
Figures from the paper (11 more)
Figure 4
Figure 4. Figure 4: Temporal trends for 2023–2025 *CL con￾ferences. While papers studying long-form generation have increased in the last year, the proportional use of human evaluation for these tasks has decreased. Annotation quality control is rarely employed. We track whether researche…
Figure 5
Figure 5. Figure 5: Prompt used for LLM-based filtering to identify papers studying long-form generation tasks and which [PITH_FULL_IMAGE:figures/full_fig_p015_5.png]
Figure 6
Figure 6. Figure 6: Partial screenshot of our annotation interface in Google Sheets showing questions pertaining to documen [PITH_FULL_IMAGE:figures/full_fig_p017_6.png]
Figure 7
Figure 7. Figure 7: Prompt for LLM-assisted annotation: input prompt structure for each LLM call. [PITH_FULL_IMAGE:figures/full_fig_p017_7.png]
Figure 8
Figure 8. Figure 8: Prompt for LLM-assisted annotation: chunk structure for codebook questions. [PITH_FULL_IMAGE:figures/full_fig_p017_8.png]
Figure 9
Figure 9. Figure 9: Prompt for LLM-assisted annotation: full question schema used for LLM annotation. [PITH_FULL_IMAGE:figures/full_fig_p018_9.png]
Figure 10
Figure 10. Figure 10: Distribution of stemmed evaluation dimensions across all papers (Overall) in the manually annotated set [PITH_FULL_IMAGE:figures/full_fig_p020_10.png]
Figure 11
Figure 11. Figure 11: Temporal trends in reporting: across all *CL papers (2023-2025) with human evaluation and long-form [PITH_FULL_IMAGE:figures/full_fig_p022_11.png]
Figure 12
Figure 12. Figure 12: Frequency of Reporting Criteria for Common NLP Tasks [PITH_FULL_IMAGE:figures/full_fig_p022_12.png]
Figure 13
Figure 13. Figure 13: Frequency of disagreement resolution method reported in manually-annotated sample: Most papers tend not to report how they address disagreement among annotators (n=202). Among the ones that report this criteria, majority vote (n=31) is the most common approach for add…
Figure 14
Figure 14. Figure 14: Distribution of IAA strength reported in [PITH_FULL_IMAGE:figures/full_fig_p023_14.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

140 extracted references · 60 canonical work pages

  1. [1]

    Scientific reports , volume=

    Expert evaluation of large language models for clinical dialogue summarization , author=. Scientific reports , volume=. 2025 , publisher=

  2. [2]

    A Critical Evaluation of Evaluations for Long-form Question Answering

    Xu, Fangyuan and Song, Yixiao and Iyyer, Mohit and Choi, Eunsol. A Critical Evaluation of Evaluations for Long-form Question Answering. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. doi:10.18653/v1/2023.acl-long.181

  3. [3]

    Responsible AI Considerations in Text Summarization Research: A Review of Current Practices

    Liu, Yu Lu and Cao, Meng and Blodgett, Su Lin and Cheung, Jackie Chi Kit and Olteanu, Alexandra and Trischler, Adam. Responsible AI Considerations in Text Summarization Research: A Review of Current Practices. Findings of the Association for Computational Linguistics: EMNLP 2023. 2023. doi:10.18653/v1/2023.findings-emnlp.413

  4. [4]

    Automated Metrics for Medical Multi-Document Summarization Disagree with Human Evaluations

    Wang, Lucy Lu and Otmakhova, Yulia and DeYoung, Jay and Truong, Thinh Hung and Kuehl, Bailey and Bransom, Erin and Wallace, Byron C. Automated Metrics for Medical Multi-Document Summarization Disagree with Human Evaluations. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. doi:10.18653/...

  5. [5]

    O pen R eviewer: A Specialized Large Language Model for Generating Critical Scientific Paper Reviews

    Idahl, Maximilian and Ahmadi, Zahra. O pen R eviewer: A Specialized Large Language Model for Generating Critical Scientific Paper Reviews. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (System Demonstrations). 2025. doi:10.18653/v1/2025.naacl-demo.44

  6. [6]

    First Conference on Language Modeling , year=

    Fine-grained hallucination detection and editing for language models , author=. First Conference on Language Modeling , year=

  7. [7]

    Proceedings of the Fourth Workshop on Generation, Evaluation and Metrics (GEM ^2 ) , pages=

    Heds 3.0: The human evaluation data sheet version 3.0 , author=. Proceedings of the Fourth Workshop on Generation, Evaluation and Metrics (GEM ^2 ) , pages=

  8. [8]

    NPJ digital medicine , volume=

    A framework for human evaluation of large language models in healthcare derived from literature review , author=. NPJ digital medicine , volume=. 2024 , publisher=

Show all 140 references
  1. [9]

    Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing: Tutorial Abstracts , pages=

    Human-centered evaluation of language technologies , author=. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing: Tutorial Abstracts , pages=

  2. [10]

    Proceedings of the 2nd Workshop on Human Evaluation of NLP Systems (HumEval) , pages=

    The human evaluation datasheet: A template for recording details of human evaluation experiments in NLP , author=. Proceedings of the 2nd Workshop on Human Evaluation of NLP Systems (HumEval) , pages=

  3. [11]

    Proceedings of the Fourth Workshop on Insights from Negative Results in NLP , pages=

    Missing information, unresponsive authors, experimental flaws: The impossibility of assessing the reproducibility of previous human evaluations in NLP , author=. Proceedings of the Fourth Workshop on Insights from Negative Results in NLP , pages=

  4. [12]

    Computational Linguistics , volume=

    Common flaws in running human evaluation experiments in NLP , author=. Computational Linguistics , volume=. 2024 , publisher=

  5. [13]

    University of Chicago Coase-Sandor Institute for Law & Economics Research Paper , number=

    Judge AI: Assessing large language models in judicial decision-making , author=. University of Chicago Coase-Sandor Institute for Law & Economics Research Paper , number=

  6. [14]

    Proceedings of the Fourth Workshop on Generation, Evaluation and Metrics (GEM ^2 ) , pages=

    Judging the judges: Evaluating alignment and vulnerabilities in llms-as-judges , author=. Proceedings of the Fourth Workshop on Generation, Evaluation and Metrics (GEM ^2 ) , pages=

  7. [15]

    LLM s instead of Human Judges? A Large Scale Empirical Study across 20 NLP Evaluation Tasks

    Bavaresco, Anna and Bernardi, Raffaella and Bertolazzi, Leonardo and Elliott, Desmond and Fern \'a ndez, Raquel and Gatt, Albert and Ghaleb, Esam and Giulianelli, Mario and Hanna, Michael and Koller, Alexander and Martins, Andre and Mondorf, Philipp and Neplenbroek, Vera and P...

  8. [16]

    Advances in Neural Information Processing Systems , volume=

    Self-refine: Iterative refinement with self-feedback , author=. Advances in Neural Information Processing Systems , volume=

  9. [17]

    npj Health Systems , volume=

    Human evaluation of large language models in healthcare: gaps, challenges, and the need for standardization , author=. npj Health Systems , volume=. 2025 , publisher=

  10. [18]

    Proceedings of the 2024 ACM Conference on Fairness, Accountability, and Transparency , pages=

    Escalation risks from language models in military and diplomatic decision-making , author=. Proceedings of the 2024 ACM Conference on Fairness, Accountability, and Transparency , pages=

  11. [19]

    Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers) , pages=

    Can language model moderators improve the health of online discourse? , author=. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers) , pages=

  12. [20]

    Proceedings of the 31st International Conference on Computational Linguistics , pages=

    Internlm-law: An open-sourced chinese legal large language model , author=. Proceedings of the 31st International Conference on Computational Linguistics , pages=

  13. [21]

    arXiv preprint arXiv:2405.05860 , year=

    The perspectivist paradigm shift: Assumptions and challenges of capturing human labels , author=. arXiv preprint arXiv:2405.05860 , year=

  14. [22]

    Nature human behaviour , volume=

    A manifesto for reproducible science , author=. Nature human behaviour , volume=. 2017 , publisher=

  15. [23]

    2022 , url =

    Shaurya Rohatgi , title =. 2022 , url =

  16. [24]

    On Context Utilization in Summarization with Large Language Models

    Ravaut, Mathieu and Sun, Aixin and Chen, Nancy and Joty, Shafiq. On Context Utilization in Summarization with Large Language Models. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. doi:10.18653/v1/2024.acl-long.153

  17. [25]

    ArXiv , year=

    How Reliable Are Automatic Evaluation Methods for Instruction-Tuned LLMs? , author=. ArXiv , year=

  18. [26]

    arXiv preprint arXiv:2312.07559 , year=

    Paperqa: Retrieval-augmented generative agent for scientific research , author=. arXiv preprint arXiv:2312.07559 , year=

  19. [27]

    arXiv e-prints , pages=

    A foundation model for human-AI collaboration in medical literature mining , author=. arXiv e-prints , pages=

  20. [28]

    arXiv preprint arXiv:2411.14199 , year=

    Openscholar: Synthesizing scientific literature with retrieval-augmented lms , author=. arXiv preprint arXiv:2411.14199 , year=

  21. [29]

    Text summarization branches out , pages=

    Rouge: A package for automatic evaluation of summaries , author=. Text summarization branches out , pages=

  22. [30]

    arXiv preprint arXiv:2310.06825 , year=

    Mistral 7B , author=. arXiv preprint arXiv:2310.06825 , year=

  23. [31]

    arXiv preprint arXiv:2303.08774 , year=

    Gpt-4 technical report , author=. arXiv preprint arXiv:2303.08774 , year=

  24. [32]

    Clinical Natural Language Processing Workshop , year=

    Generating medically-accurate summaries of patient-provider dialogue: A multi-stage approach using large language models , author=. Clinical Natural Language Processing Workshop , year=

  25. [33]

    Conference on Empirical Methods in Natural Language Processing , year=

    Hierarchical Catalogue Generation for Literature Review: A Benchmark , author=. Conference on Empirical Methods in Natural Language Processing , year=

  26. [34]

    Annual Meeting of the Association for Computational Linguistics , year=

    Multi-News: A Large-Scale Multi-Document Summarization Dataset and Abstractive Hierarchical Model , author=. Annual Meeting of the Association for Computational Linguistics , year=

  27. [35]

    Annual Meeting of the Association for Computational Linguistics , year=

    Hierarchical Transformers for Multi-Document Summarization , author=. Annual Meeting of the Association for Computational Linguistics , year=

  28. [36]

    Annual Meeting of the Association for Computational Linguistics , year=

    Summ ^N : A Multi-Stage Summarization Framework for Long Input Dialogues and Documents , author=. Annual Meeting of the Association for Computational Linguistics , year=

  29. [37]

    ArXiv , year=

    Leveraging Long-Context Large Language Models for Multi-Document Understanding and Summarization in Enterprise Applications , author=. ArXiv , year=

  30. [38]

    , author=

    A Hierarchical Decoder with Three-level Hierarchical Attention to Generate Abstractive Summaries of Interleaved Texts. , author=. arXiv: Computation and Language , year=

  31. [39]

    ArXiv , year=

    uMedSum: A Unified Framework for Advancing Medical Abstractive Summarization , author=. ArXiv , year=

  32. [40]

    , author=

    Roles of Document Structure, Cognitive Strategy, and Awareness in Searching for Information. , author=. Reading Research Quarterly , year=

  33. [41]

    , author=

    The Effects of Text Structure Instruction on Middle-Grade Students' Comprehension and Production of Expository Text. , author=. Reading Research Quarterly , year=

  34. [42]

    and Yang, Qiang and Xie, Xing , title =

    Chang, Yupeng and Wang, Xu and Wang, Jindong and Wu, Yuan and Yang, Linyi and Zhu, Kaijie and Chen, Hao and Yi, Xiaoyuan and Wang, Cunxiang and Wang, Yidong and Ye, Wei and Zhang, Yue and Chang, Yi and Yu, Philip S. and Yang, Qiang and Xie, Xing , title =. ACM Trans. Intell. S...

  35. [43]

    Improving Factuality in Clinical Abstractive Multi-Document Summarization by Guided Continued Pre-training

    Elhady, Ahmed and Elsayed, Khaled and Agirre, Eneko and Artetxe, Mikel. Improving Factuality in Clinical Abstractive Multi-Document Summarization by Guided Continued Pre-training. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computati...

  36. [44]

    arXiv preprint arXiv:2305.14251 , year=

    Factscore: Fine-grained atomic evaluation of factual precision in long form text generation , author=. arXiv preprint arXiv:2305.14251 , year=

  37. [45]

    arXiv preprint arXiv:2501.03545 , year=

    Beyond Factual Accuracy: Evaluating Coverage of Diverse Factual Information in Long-form Text Generation , author=. arXiv preprint arXiv:2501.03545 , year=

  38. [46]

    Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers) , pages=

    Embrace Divergence for Richer Insights: A Multi-document Summarization Benchmark and a Case Study on Summarizing Diverse Information from News Articles , author=. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics:...

  39. [47]

    medRxiv , pages=

    Synthetic Data Distillation Enables the Extraction of Clinical Information at Scale , author=. medRxiv , pages=. 2024 , publisher=

  40. [48]

    and Villarroel, Mauricio and Clifford, Gari D

    Lee, Joon and Scott, Daniel J. and Villarroel, Mauricio and Clifford, Gari D. and Saeed, Mohammed and Mark, Roger G. , booktitle=. Open-access MIMIC-II database for intensive care research , year=

  41. [49]

    ArXiv , year=

    ClinicalBERT: Modeling Clinical Notes and Predicting Hospital Readmission , author=. ArXiv , year=

  42. [50]

    Bioinformatics , volume =

    Lee, Jinhyuk and Yoon, Wonjin and Kim, Sungdong and Kim, Donghyeon and Kim, Sunkyu and So, Chan Ho and Kang, Jaewoo , title =. Bioinformatics , volume =. 2019 , month =. doi:10.1093/bioinformatics/btz682 , url =

  43. [51]

    TOPICAL : TOPIC Pages A utomagica L ly

    Giorgi, John and Singh, Amanpreet and Downey, Doug and Feldman, Sergey and Wang, Lucy. TOPICAL : TOPIC Pages A utomagica L ly. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume...

  44. [52]

    What`s in a Summary? Laying the Groundwork for Advances in Hospital-Course Summarization

    Adams, Griffin and Alsentzer, Emily and Ketenci, Mert and Zucker, Jason and Elhadad, No \'e mie. What`s in a Summary? Laying the Groundwork for Advances in Hospital-Course Summarization. Proceedings of the 2021 Conference of the North American Chapter of the Association for Co...

  45. [53]

    2015 26th international workshop on database and expert systems applications (dexa) , pages=

    Clinical decision support systems: a survey of NLP-based approaches from unstructured data , author=. 2015 26th international workshop on database and expert systems applications (dexa) , pages=. 2015 , organization=

  46. [54]

    Journal of Intelligent Connectivity and Emerging Technologies , volume=

    Natural language processing for clinical decision support systems: A review of recent advances in healthcare , author=. Journal of Intelligent Connectivity and Emerging Technologies , volume=

  47. [55]

    Journal of biomedical informatics , volume=

    What can natural language processing do for clinical decision support? , author=. Journal of biomedical informatics , volume=. 2009 , publisher=

  48. [56]

    A Novel System for Extractive Clinical Note Summarization using EHR Data

    Liang, Jennifer and Tsou, Ching-Huei and Poddar, Ananya. A Novel System for Extractive Clinical Note Summarization using EHR Data. Proceedings of the 2nd Clinical Natural Language Processing Workshop. 2019. doi:10.18653/v1/W19-1906

  49. [57]

    Generating SOAP Notes from Doctor-Patient Conversations Using Modular Summarization Techniques

    Krishna, Kundan and Khosla, Sopan and Bigham, Jeffrey and Lipton, Zachary C. Generating SOAP Notes from Doctor-Patient Conversations Using Modular Summarization Techniques. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th Int...

  50. [58]

    Towards Automating Medical Scribing : Clinic Visit D ialogue2 N ote Sentence Alignment and Snippet Summarization

    Yim, Wen-wai and Yetisgen, Meliha. Towards Automating Medical Scribing : Clinic Visit D ialogue2 N ote Sentence Alignment and Snippet Summarization. Proceedings of the Second Workshop on Natural Language Processing for Medical Conversations. 2021. doi:10.18653/v1/2021.nlpmc-1.2

  51. [59]

    DERA : Enhancing Large Language Model Completions with Dialog-Enabled Resolving Agents

    Nair, Varun and Schumacher, Elliot and Tso, Geoffrey and Kannan, Anitha. DERA : Enhancing Large Language Model Completions with Dialog-Enabled Resolving Agents. Proceedings of the 6th Clinical Natural Language Processing Workshop. 2024. doi:10.18653/v1/2024.clinicalnlp-1.12

  52. [60]

    JMIR medical education , volume=

    How does ChatGPT perform on the United States Medical Licensing Examination (USMLE)? The implications of large language models for medical education and knowledge assessment , author=. JMIR medical education , volume=. 2023 , publisher=

  53. [61]

    Cureus , volume=

    Overview of early ChatGPT’s presence in medical literature: insights from a hybrid literature review by ChatGPT and human experts , author=. Cureus , volume=. 2023 , publisher=

  54. [62]

    ArXiv , year=

    Bio-SIEVE: Exploring Instruction Tuning Large Language Models for Systematic Review Automation , author=. ArXiv , year=

  55. [63]

    CoRR , volume =

    Athanasios Lagopoulos and Grigorios Tsoumakas , title =. CoRR , volume =. 2020 , url =. 2011.09752 , timestamp =

  56. [64]

    Benchmarking Large Language Models for News Summarization

    Zhang, Tianyi and Ladhak, Faisal and Durmus, Esin and Liang, Percy and McKeown, Kathleen and Hashimoto, Tatsunori B. Benchmarking Large Language Models for News Summarization. Transactions of the Association for Computational Linguistics. 2024. doi:10.1162/tacl_a_00632

  57. [65]

    On Learning to Summarize with Large Language Models as References

    Liu, Yixin and Shi, Kejian and He, Katherine and Ye, Longtian and Fabbri, Alexander and Liu, Pengfei and Radev, Dragomir and Cohan, Arman. On Learning to Summarize with Large Language Models as References. Proceedings of the 2024 Conference of the North American Chapter of the...

  58. [66]

    Summarizing, Simplifying, and Synthesizing Medical Evidence using GPT -3 (with Varying Success)

    Shaib, Chantal and Li, Millicent and Joseph, Sebastian and Marshall, Iain and Li, Junyi Jessy and Wallace, Byron. Summarizing, Simplifying, and Synthesizing Medical Evidence using GPT -3 (with Varying Success). Proceedings of the 61st Annual Meeting of the Association for Comp...

  59. [67]

    ACM Comput

    Koh, Huan Yee and Ju, Jiaxin and Liu, Ming and Pan, Shirui , title =. ACM Comput. Surv. , month = dec, articleno =. 2022 , issue_date =. doi:10.1145/3545176 , abstract =

  60. [68]

    Conference on Empirical Methods in Natural Language Processing , year=

    DocAsRef: An Empirical Study on Repurposing Reference-based Summary Quality Metrics as Reference-free Metrics , author=. Conference on Empirical Methods in Natural Language Processing , year=

  61. [69]

    ACM Trans

    Nenkova, Ani and Passonneau, Rebecca and McKeown, Kathleen , title =. ACM Trans. Speech Lang. Process. , month = may, pages =. 2007 , issue_date =. doi:10.1145/1233912.1233913 , abstract =

  62. [70]

    ACM Comput

    Jangra, Anubhav and Mukherjee, Sourajit and Jatowt, Adam and Saha, Sriparna and Hasanuzzaman, Mohammad , title =. ACM Comput. Surv. , month = jul, articleno =. 2023 , issue_date =. doi:10.1145/3584700 , abstract =

  63. [71]

    Annual Meeting of the Association for Computational Linguistics , year=

    A Simple Theoretical Model of Importance for Summarization , author=. Annual Meeting of the Association for Computational Linguistics , year=

  64. [72]

    ArXiv , year=

    Earlier Isn’t Always Better: Sub-aspect Analysis on Corpus and System Biases in Summarization , author=. ArXiv , year=

  65. [73]

    Journal of Medical Internet Research , year=

    Potential Roles of Large Language Models in the Production of Systematic Reviews and Meta-Analyses , author=. Journal of Medical Internet Research , year=

  66. [74]

    BMJ Open , year=

    Analysis of the time and workers needed to conduct systematic reviews of medical interventions using data from the PROSPERO registry , author=. BMJ Open , year=

  67. [75]

    Energies , year=

    The Resilience of Critical Infrastructure Systems: A Systematic Literature Review , author=. Energies , year=

  68. [76]

    Proceedings of the 52nd annual meeting of the association for computational linguistics (volume 1: Long papers) , pages=

    Hierarchical summarization: Scaling up multi-document summarization , author=. Proceedings of the 52nd annual meeting of the association for computational linguistics (volume 1: Long papers) , pages=

  69. [77]

    Generating (Factual?) Narrative Summaries of RCTs: Experiments with Neural Multi-Document Summarization , author=. AMIA ... Annual Symposium proceedings. AMIA Symposium , year=

  70. [78]

    Proceedings of the conference

    What’s in a Summary? Laying the Groundwork for Advances in Hospital-Course Summarization , author=. Proceedings of the conference. Association for Computational Linguistics. North American Chapter. Meeting , year=

  71. [79]

    Trends in cognitive sciences , volume=

    Hierarchical process memory: memory as an integral component of information processing , author=. Trends in cognitive sciences , volume=. 2015 , publisher=

  72. [80]

    Craik , abstract =

    F.I.M. Craik , abstract =. Memory: Levels of Processing , editor =. International Encyclopedia of the Social & Behavioral Sciences , publisher =. 2001 , isbn =. doi:https://doi.org/10.1016/B0-08-043076-7/01508-4 , url =

  73. [81]

    Conference on Empirical Methods in Natural Language Processing , year=

    Summarizing Multiple Documents with Conversational Structure for Meta-Review Generation , author=. Conference on Empirical Methods in Natural Language Processing , year=

  74. [82]

    ArXiv , year=

    Nutri-bullets: Summarizing Health Studies by Composing Segments , author=. ArXiv , year=

  75. [83]

    ArXiv , year=

    Multi-Document Scientific Summarization from a Knowledge Graph-Centric View , author=. ArXiv , year=

  76. [84]

    Annual Meeting of the Association for Computational Linguistics , year=

    HIBRIDS: Attention with Hierarchical Biases for Structure-aware Long Document Summarization , author=. Annual Meeting of the Association for Computational Linguistics , year=

  77. [85]

    Proceedings of the 45th International ACM SIGIR Conference on Research and Development in Information Retrieval , year=

    Target-aware Abstractive Related Work Generation with Contrastive Learning , author=. Proceedings of the 45th International ACM SIGIR Conference on Research and Development in Information Retrieval , year=

  78. [86]

    ArXiv , year=

    Disentangling Instructive Information from Ranked Multiple Candidates for Multi-Document Scientific Summarization , author=. ArXiv , year=

  79. [87]

    H i S truct+: Improving Extractive Text Summarization with Hierarchical Structure Information

    Ruan, Qian and Ostendorff, Malte and Rehm, Georg. H i S truct+: Improving Extractive Text Summarization with Hierarchical Structure Information. Findings of the Association for Computational Linguistics: ACL 2022. 2022. doi:10.18653/v1/2022.findings-acl.102

  80. [88]

    Evaluation Metrics in the Era of GPT -4: Reliably Evaluating Large Language Models on Sequence to Sequence Tasks

    Sottana, Andrea and Liang, Bin and Zou, Kai and Yuan, Zheng. Evaluation Metrics in the Era of GPT -4: Reliably Evaluating Large Language Models on Sequence to Sequence Tasks. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023. doi:10.1...

  81. [89]

    and Belz, Anya and Du s ek, Ond r ej and Mille, Simon and van der Lee, Chris and Reiter, Ehud and Santhanam, Shivani and Thomson, Craig

    Howcroft, David M. and Belz, Anya and Du s ek, Ond r ej and Mille, Simon and van der Lee, Chris and Reiter, Ehud and Santhanam, Shivani and Thomson, Craig. A Survey on the Evaluation of Natural Language Generation: Past, Present, and Future. Proceedings of the 17th Internation...

  82. [90]

    G-Eval : NLG Evaluation using GPT-4 with Better Human Alignment

    Liu, Yixuan and Bansal, Mohit. G-Eval : NLG Evaluation using GPT-4 with Better Human Alignment. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing (EMNLP). 2023

  83. [91]

    HD -Eval: Aligning Large Language Model Evaluators Through Hierarchical Criteria Decomposition

    Liu, Yuxuan and Yang, Tianchi and Huang, Shaohan and Zhang, Zihan and Huang, Haizhen and Wei, Furu and Deng, Weiwei and Sun, Feng and Zhang, Qi. HD -Eval: Aligning Large Language Model Evaluators Through Hierarchical Criteria Decomposition. Proceedings of the 62nd Annual Meeti...

  84. [92]

    arXiv preprint arXiv:2402.15754 , year=

    HD-Eval: Aligning Large Language Model Evaluators Through Human Demonstration , author=. arXiv preprint arXiv:2402.15754 , year=

  85. [93]

    Findings of the Association for Computational Linguistics: EMNLP 2023 , year=

    GPT-4 as an Effective Zero-Shot Evaluator for Scientific Figure Captioning , author=. Findings of the Association for Computational Linguistics: EMNLP 2023 , year=

  86. [94]

    Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing (EMNLP) , year=

    G-Eval: NLG Evaluation using GPT-4 with Better Human Alignment , author=. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing (EMNLP) , year=

  87. [95]

    Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers)

    Qiu, Haoyi and Huang, Kung-Hsiang and Qu, Jingnong and Peng, Nanyun. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. doi:10.18653/v1/2024.naacl-long.33

  88. [96]

    Transactions of the Association for Computational Linguistics , volume=

    SummEval: Re-evaluating Summarization Evaluation , author=. Transactions of the Association for Computational Linguistics , volume=. 2021 , publisher=

  89. [97]

    arXiv preprint arXiv:2011.13662 , year=

    FFCI: A Framework for Interpretable Automatic Evaluation of Summarization , author=. arXiv preprint arXiv:2011.13662 , year=

  90. [98]

    Journal of the Royal Society of Medicine , volume =

    Khalid S Khan and Regina Kunz and Jos Kleijnen and Gerd Antes , title =. Journal of the Royal Society of Medicine , volume =. 2003 , doi =

  91. [99]

    The significant cost of systematic reviews and meta-analyses: A call for greater involvement of machine learning to assess the promise of clinical trials , journal =

    Matthew Michelson and Katja Reuter , keywords =. The significant cost of systematic reviews and meta-analyses: A call for greater involvement of machine learning to assess the promise of clinical trials , journal =. 2019 , issn =. doi:https://doi.org/10.1016/j.conctc.2019.1004...

  92. [100]

    npj Digital Medicine , volume=

    Closing the gap between open source and commercial large language models for medical evidence summarization , author=. npj Digital Medicine , volume=. 2024 , publisher=

  93. [101]

    arXiv preprint arXiv:2104.06486 , year=

    Ms2: Multi-document summarization of medical studies , author=. arXiv preprint arXiv:2104.06486 , year=

  94. [102]

    arXiv preprint arXiv:2305.13693 , year=

    Automated metrics for medical multi-document summarization disagree with human evaluations , author=. arXiv preprint arXiv:2305.13693 , year=

  95. [103]

    arXiv preprint arXiv:2402.02420 , year=

    Factuality of large language models in the year 2024 , author=. arXiv preprint arXiv:2402.02420 , year=

  96. [104]

    arXiv preprint arXiv:2307.13528 , year=

    FacTool: Factuality Detection in Generative AI--A Tool Augmented Framework for Multi-Task and Multi-Domain Scenarios , author=. arXiv preprint arXiv:2307.13528 , year=

  97. [105]

    Findings of the Association for Computational Linguistics: ACL 2023 , pages=

    NonFactS: NonFactual summary generation for factuality evaluation in document summarization , author=. Findings of the Association for Computational Linguistics: ACL 2023 , pages=

  98. [106]

    Applied Cognitive Psychology: The Official Journal of the Society for Applied Research in Memory and Cognition , volume=

    Consequences of erudite vernacular utilized irrespective of necessity: Problems with using long words needlessly , author=. Applied Cognitive Psychology: The Official Journal of the Society for Applied Research in Memory and Cognition , volume=. 2006 , publisher=

  99. [107]

    and Maryam Kouchaki and Hancock, Jeffrey T

    Markowitz, David M. and Maryam Kouchaki and Hancock, Jeffrey T. and Francesca Gino. The Deception Spiral: Corporate Obfuscation Leads to Perceptions of Immorality and Cheating Behavior. Journal of Language and Social Psychology. 2021. doi:10.1177/0261927X20949594

  100. [108]

    PNAS Nexus , volume =

    Markowitz, David M , title =. PNAS Nexus , volume =. 2024 , month =. doi:10.1093/pnasnexus/pgae387 , url =

  101. [109]

    CHIME : LLM -Assisted Hierarchical Organization of Scientific Studies for Literature Review Support

    Hsu, Chao-Chun and Bransom, Erin and Sparks, Jenna and Kuehl, Bailey and Tan, Chenhao and Wadden, David and Wang, Lucy and Naik, Aakanksha. CHIME : LLM -Assisted Hierarchical Organization of Scientific Studies for Literature Review Support. Findings of the Association for Comp...

  102. [110]

    FIZZ : Factual Inconsistency Detection by Zoom-in Summary and Zoom-out Document

    Yang, Joonho and Yoon, Seunghyun and Kim, ByeongJeong and Lee, Hwanhee. FIZZ : Factual Inconsistency Detection by Zoom-in Summary and Zoom-out Document. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. doi:10.18653/v1/2024.emnlp-main.3

  103. [111]

    On the Role of Summary Content Units in Text Summarization Evaluation

    Nawrath, Marcel and Nowak, Agnieszka and Ratz, Tristan and Walenta, Danilo and Opitz, Juri and Ribeiro, Leonardo and Sedoc, Jo \ a o and Deutsch, Daniel and Mille, Simon and Liu, Yixin and Gehrmann, Sebastian and Zhang, Lining and Mahamood, Saad and Clinciu, Miruna and Chandu,...

  104. [112]

    arXiv preprint arXiv:2109.11503 , year=

    Finding a balanced degree of automation for summary evaluation , author=. arXiv preprint arXiv:2109.11503 , year=

  105. [113]

    Weinberger and Yoav Artzi , title =

    Tianyi Zhang and Varsha Kishore and Felix Wu and Kilian Q. Weinberger and Yoav Artzi , title =. CoRR , volume =. 2019 , url =. 1904.09675 , timestamp =

  106. [114]

    arXiv preprint arXiv:2304.02554 , year=

    Human-like summarization evaluation with chatgpt , author=. arXiv preprint arXiv:2304.02554 , year=

  107. [115]

    ACM Computing Surveys (CSUR) , volume=

    A survey of evaluation metrics used for NLG systems , author=. ACM Computing Surveys (CSUR) , volume=. 2022 , publisher=

  108. [116]

    Evolutionary Applications , volume=

    Next-generation metrics for monitoring genetic erosion within populations of conservation concern , author=. Evolutionary Applications , volume=. 2018 , publisher=

  109. [117]

    Journal of the Association for Information Science and Technology , volume=

    Measuring text difficulty using parse-tree frequency , author=. Journal of the Association for Information Science and Technology , volume=. 2017 , publisher=

  110. [118]

    APPLS : Evaluating Evaluation Metrics for Plain Language Summarization

    Guo, Yue and August, Tal and Leroy, Gondy and Cohen, Trevor and Wang, Lucy Lu. APPLS : Evaluating Evaluation Metrics for Plain Language Summarization. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. doi:10.18653/v1/2024.emnlp-main.519

  111. [119]

    Aho and Jeffrey D

    Alfred V. Aho and Jeffrey D. Ullman , title =. 1972

  112. [120]

    Publications Manual , year = "1983", publisher =

  113. [121]

    Chandra and Dexter C

    Ashok K. Chandra and Dexter C. Kozen and Larry J. Stockmeyer , year = "1981", title =. doi:10.1145/322234.322243

  114. [122]

    Scalable training of

    Andrew, Galen and Gao, Jianfeng , booktitle=. Scalable training of

  115. [123]

    Dan Gusfield , title =. 1997

  116. [124]

    Tetreault , title =

    Mohammad Sadegh Rasooli and Joel R. Tetreault , title =. Computing Research Repository , volume =. 2015 , url =

  117. [125]

    A Framework for Learning Predictive Structures from Multiple Tasks and Unlabeled Data , Volume =

    Ando, Rie Kubota and Zhang, Tong , Issn =. A Framework for Learning Predictive Structures from Multiple Tasks and Unlabeled Data , Volume =. Journal of Machine Learning Research , Month = dec, Numpages =

  118. [126]

    arXiv preprint arXiv:2005.13213 , year=

    Give me convenience and give her death: Who should decide what uses of NLP are appropriate, and on what basis? , author=. arXiv preprint arXiv:2005.13213 , year=

  119. [127]

    Impact of Annotator Demographics on Sentiment Dataset Labeling , year =

    Ding, Yi and You, Jacob and Machulla, Tonja-Katrin and Jacobs, Jennifer and Sen, Pradeep and H\". Impact of Annotator Demographics on Sentiment Dataset Labeling , year =. Proc. ACM Hum.-Comput. Interact. , month = nov, articleno =. doi:10.1145/3555632 , abstract =

  120. [128]

    Identifying and Measuring Annotator Bias Based on Annotators' Demographic Characteristics

    Al Kuwatly, Hala and Wich, Maximilian and Groh, Georg. Identifying and Measuring Annotator Bias Based on Annotators' Demographic Characteristics. Proceedings of the Fourth Workshop on Online Abuse and Harms. 2020. doi:10.18653/v1/2020.alw-1.21

  121. [129]

    Proceedings of the 57th annual meeting of the association for computational linguistics , pages=

    The risk of racial bias in hate speech detection , author=. Proceedings of the 57th annual meeting of the association for computational linguistics , pages=

  122. [130]

    International journal of qualitative methods , volume=

    Demonstrating rigor using thematic analysis: A hybrid approach of inductive and deductive coding and theme development , author=. International journal of qualitative methods , volume=. 2006 , publisher=

  123. [131]

    Transactions of the Association for Computational Linguistics , volume=

    Bridging the gap: A survey on integrating (human) feedback for natural language generation , author=. Transactions of the Association for Computational Linguistics , volume=. 2023 , publisher=

  124. [132]

    arXiv preprint arXiv:2006.14799 , year=

    Evaluation of text generation: A survey , author=. arXiv preprint arXiv:2006.14799 , year=

  125. [133]

    Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=

    Considers-the-human evaluation framework: Rethinking human evaluation for generative large language models , author=. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=

  126. [134]

    Proceedings of the conference on fairness, accountability, and transparency , pages=

    Model cards for model reporting , author=. Proceedings of the conference on fairness, accountability, and transparency , pages=

  127. [135]

    arXiv preprint arXiv:2109.06598 , year=

    Just what do you think you're doing, Dave?'a checklist for responsible data use in NLP , author=. arXiv preprint arXiv:2109.06598 , year=

  128. [136]

    arXiv preprint arXiv:1909.03004 , year=

    Show your work: Improved reporting of experimental results , author=. arXiv preprint arXiv:1909.03004 , year=

  129. [137]

    Proceedings of the 13th international conference on natural language generation , pages=

    Twenty years of confusion in human evaluation: NLG needs evaluation sheets and standardised definitions , author=. Proceedings of the 13th international conference on natural language generation , pages=

  130. [138]

    The Perspectivist Paradigm Shift: Assumptions and Challenges of Capturing Human Labels

    Fleisig, Eve and Blodgett, Su Lin and Klein, Dan and Talat, Zeerak. The Perspectivist Paradigm Shift: Assumptions and Challenges of Capturing Human Labels. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human ...

  131. [139]

    `Just What do You Think You ' re Doing, Dave?' A Checklist for Responsible Data Use in NLP

    Rogers, Anna and Baldwin, Timothy and Leins, Kobi. `Just What do You Think You ' re Doing, Dave?' A Checklist for Responsible Data Use in NLP. Findings of the Association for Computational Linguistics: EMNLP 2021. 2021. doi:10.18653/v1/2021.findings-emnlp.414

  132. [140]

    Show Your Work: Improved Reporting of Experimental Results

    Dodge, Jesse and Gururangan, Suchin and Card, Dallas and Schwartz, Roy and Smith, Noah A. Show Your Work: Improved Reporting of Experimental Results. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conferen...

Pith tools

Reviewed June 27, 2026 · model on record in the stance chip above.