REVIEW 2 major objections 2 minor 140 references
Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation
T0 review · 2 major / 2 minor · reviewed 2026-06-27 · grok-4.3
Pith's one-line read Human evaluation protocols for long-form text generation in recent NLP conference papers are often incompletely reported, creating ambiguity about what was measured and by whom.
desk verdict The paper gives the first big numbers on under-reporting in human eval protocols for long-form generation, but the 20 criteria are unvalidated so the strength of the 'widespread' claim is still open. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
A set of 20 reportable criteria related to reproducibility of human evaluation studies, applied to check what details papers include about design, participants, and judgment processes.
What would settle it
A re-analysis of the same papers that applies a different or expanded list of criteria and finds high rates of complete reporting on the missing items.
Extended reading notes
Core claim
A systematic review of human evaluation protocols in *CL publications reveals widespread under-reporting of important aspects of study design, who contributed judgments, and how judgments should be interpreted, which produces ambiguity about what was actually measured and how the results should be understood.
Load-bearing premise
The authors' chosen set of 20 criteria is enough to determine whether a human evaluation study is reproducible and interpretable.
Editorial extensions
If this is right
- Adopting the 20 criteria would make it easier to interpret and compare human evaluation results across different papers.
- Papers would need to document participant recruitment, training, and agreement measures more consistently.
- Ambiguity in current evaluations would decrease if journals and conferences required explicit reporting on these points.
- Future comparisons of generation systems could rest on clearer evidence of evaluation quality.
Reading between the lines
- The same under-reporting pattern likely appears in evaluations of short-form or other generation tasks outside the long-form focus.
- Conferences could reduce ambiguity by adding checklist items based on the 20 criteria during submission.
- Greater transparency might shift research incentives toward more careful study design rather than just reporting results.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper conducts a large-scale observational analysis of human evaluation protocols for long-form text generation in *CL conference papers (2023–2025). It performs a full manual review of 284 papers plus LLM-assisted analysis of an additional 1.8k+ papers against a fixed set of 20 reportable criteria for reproducibility. The central claim is that widespread under-reporting of study-design details creates ambiguity about what was measured, who provided judgments, and how results should be interpreted; the authors provide recommendations and release code plus an annotated dataset.
Significance. If the central observational claim holds, the work is significant because human evaluation remains a primary method for assessing long-form generation quality, and documented under-reporting directly affects reproducibility and interpretability in the field. The explicit release of analysis code and the annotated dataset is a clear strength that supports verification and follow-up studies. The findings could usefully inform community guidelines, provided the 20 criteria are shown to be well-aligned with existing reporting standards.
major comments (2)
- [Section defining the 20 criteria] Section defining the 20 criteria: the manuscript presents these criteria as capturing 'important aspects' necessary for reproducibility and interpretability, yet provides no external validation (e.g., expert survey, comparison against ACL or prior meta-study reporting guidelines, or inter-rater agreement on criterion importance). Because the prevalence statistics and the downstream claim of 'widespread under-reporting of important aspects' rest directly on this author-defined set, the absence of such validation makes the quantitative conclusions sensitive to the particular framing chosen.
- [LLM-assisted analysis section] LLM-assisted analysis section (1.8k+ papers): the extension from the 284 manually reviewed papers to the larger corpus is load-bearing for the 'widespread' claim, but the manuscript does not report prompt details, few-shot examples, or measured agreement/error rates between the LLM outputs and the manual annotations. Without these, systematic biases in the automated labeling could materially affect the reported under-reporting rates.
minor comments (2)
- [Table 1] Table 1 (or equivalent summary table of criteria): the mapping from each criterion to the specific ambiguity it addresses (measurement, contributors, or interpretation) could be made more explicit to help readers trace how missing items produce the claimed ambiguities.
- [Data and code release] The GitHub link is provided, but the README should include a clear description of how the 284 manual annotations were performed (annotator background, resolution process) to strengthen reproducibility of the core dataset.
Simulated Author's Rebuttal
We thank the referee for the constructive feedback. We address each major comment below and outline planned revisions to improve the manuscript.
read point-by-point responses
-
Referee: [Section defining the 20 criteria] Section defining the 20 criteria: the manuscript presents these criteria as capturing 'important aspects' necessary for reproducibility and interpretability, yet provides no external validation (e.g., expert survey, comparison against ACL or prior meta-study reporting guidelines, or inter-rater agreement on criterion importance). Because the prevalence statistics and the downstream claim of 'widespread under-reporting of important aspects' rest directly on this author-defined set, the absence of such validation makes the quantitative conclusions sensitive to the particular framing chosen.
Authors: We agree that additional justification for the criteria would strengthen the work. The 20 criteria were synthesized from recurring elements in prior NLP literature on human evaluation reproducibility and reporting standards. In revision we will add a new subsection explicitly mapping each criterion to relevant ACL guidelines and earlier meta-studies, together with a brief rationale for inclusion. While we did not conduct a new expert survey, this explicit alignment will reduce sensitivity to the chosen framing and make the prevalence claims more robust. revision: partial
-
Referee: [LLM-assisted analysis section] LLM-assisted analysis section (1.8k+ papers): the extension from the 284 manually reviewed papers to the larger corpus is load-bearing for the 'widespread' claim, but the manuscript does not report prompt details, few-shot examples, or measured agreement/error rates between the LLM outputs and the manual annotations. Without these, systematic biases in the automated labeling could materially affect the reported under-reporting rates.
Authors: We concur that full transparency on the LLM-assisted labeling is required. The original submission omitted these details for brevity. The revised manuscript will include the complete prompts, few-shot examples, and a dedicated error-analysis subsection reporting agreement rates (and disagreement categories) between the LLM and the manual annotations on a held-out validation set. This addition will allow readers to evaluate potential biases directly. revision: yes
Circularity Check
No significant circularity in observational literature analysis
full rationale
The paper performs a manual and LLM-assisted review of reporting practices in other papers by first defining an explicit set of 20 criteria and then counting their presence or absence. No equations, fitted parameters, predictions, or self-citation chains exist that reduce any central claim to its own inputs by construction. The analysis is self-contained empirical observation against transparently stated criteria; the absence of any enumerated circularity pattern (self-definitional, fitted-input prediction, load-bearing self-citation, etc.) yields a score of 0.
Assumptions & free parameters
assumptions (1)
- domain assumption The defined set of 20 reportable criteria adequately captures key aspects of human evaluation reproducibility.
Cite this review
Pith. "Pith review of Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation." pith.science (2026). https://pith.science/paper/OWOLVEFR
@misc{pith2026260607936,
author = {Pith},
title = {Pith review of: Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation},
year = {2026},
howpublished = {\url{https://pith.science/paper/OWOLVEFR}},
note = {Machine review of arXiv:2606.07936}
}
read the original abstract
Human evaluation plays a critical role in assessing the quality of generated text. However, the reliability and reproducibility of these evaluations depend on transparent and well-documented protocols -- details that are frequently missing in current practice. In this work, we conduct a large-scale analysis of human evaluation protocols for evaluating long-form generation tasks in *CL conference publications from 2023--2025, including a full manual review of 284 papers and LLM-assisted analysis for another 1.8k+ papers. We define a set of 20 reportable criteria related to reproducibility of human evaluation studies, and apply these criteria to systematically examine reporting norms and practices within the community. We find widespread under-reporting of important aspects of human evaluation study design, leading to ambiguity about what was measured and how, who contributed judgments, and how judgments should be interpreted. Based on these findings, we outline actionable recommendations to support more transparent and reproducible reporting in future research. Our analysis code and annotated dataset can be found at: https://github.com/larchlab/Illusions-of-the-Gold-Standard
Figures
Figures from the paper (11 more)
Reference graph
Works this paper leans on
-
[1]
Scientific reports , volume=
Expert evaluation of large language models for clinical dialogue summarization , author=. Scientific reports , volume=. 2025 , publisher=
2025
-
[2]
A Critical Evaluation of Evaluations for Long-form Question Answering
Xu, Fangyuan and Song, Yixiao and Iyyer, Mohit and Choi, Eunsol. A Critical Evaluation of Evaluations for Long-form Question Answering. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. doi:10.18653/v1/2023.acl-long.181
-
[3]
Responsible AI Considerations in Text Summarization Research: A Review of Current Practices
Liu, Yu Lu and Cao, Meng and Blodgett, Su Lin and Cheung, Jackie Chi Kit and Olteanu, Alexandra and Trischler, Adam. Responsible AI Considerations in Text Summarization Research: A Review of Current Practices. Findings of the Association for Computational Linguistics: EMNLP 2023. 2023. doi:10.18653/v1/2023.findings-emnlp.413
-
[4]
Automated Metrics for Medical Multi-Document Summarization Disagree with Human Evaluations
Wang, Lucy Lu and Otmakhova, Yulia and DeYoung, Jay and Truong, Thinh Hung and Kuehl, Bailey and Bransom, Erin and Wallace, Byron C. Automated Metrics for Medical Multi-Document Summarization Disagree with Human Evaluations. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. doi:10.18653/...
-
[5]
O pen R eviewer: A Specialized Large Language Model for Generating Critical Scientific Paper Reviews
Idahl, Maximilian and Ahmadi, Zahra. O pen R eviewer: A Specialized Large Language Model for Generating Critical Scientific Paper Reviews. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (System Demonstrations). 2025. doi:10.18653/v1/2025.naacl-demo.44
-
[6]
First Conference on Language Modeling , year=
Fine-grained hallucination detection and editing for language models , author=. First Conference on Language Modeling , year=
-
[7]
Proceedings of the Fourth Workshop on Generation, Evaluation and Metrics (GEM ^2 ) , pages=
Heds 3.0: The human evaluation data sheet version 3.0 , author=. Proceedings of the Fourth Workshop on Generation, Evaluation and Metrics (GEM ^2 ) , pages=
-
[8]
NPJ digital medicine , volume=
A framework for human evaluation of large language models in healthcare derived from literature review , author=. NPJ digital medicine , volume=. 2024 , publisher=
2024
Show all 140 references
-
[9]
Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing: Tutorial Abstracts , pages=
Human-centered evaluation of language technologies , author=. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing: Tutorial Abstracts , pages=
2024
-
[10]
Proceedings of the 2nd Workshop on Human Evaluation of NLP Systems (HumEval) , pages=
The human evaluation datasheet: A template for recording details of human evaluation experiments in NLP , author=. Proceedings of the 2nd Workshop on Human Evaluation of NLP Systems (HumEval) , pages=
-
[11]
Proceedings of the Fourth Workshop on Insights from Negative Results in NLP , pages=
Missing information, unresponsive authors, experimental flaws: The impossibility of assessing the reproducibility of previous human evaluations in NLP , author=. Proceedings of the Fourth Workshop on Insights from Negative Results in NLP , pages=
-
[12]
Computational Linguistics , volume=
Common flaws in running human evaluation experiments in NLP , author=. Computational Linguistics , volume=. 2024 , publisher=
2024
-
[13]
University of Chicago Coase-Sandor Institute for Law & Economics Research Paper , number=
Judge AI: Assessing large language models in judicial decision-making , author=. University of Chicago Coase-Sandor Institute for Law & Economics Research Paper , number=
-
[14]
Proceedings of the Fourth Workshop on Generation, Evaluation and Metrics (GEM ^2 ) , pages=
Judging the judges: Evaluating alignment and vulnerabilities in llms-as-judges , author=. Proceedings of the Fourth Workshop on Generation, Evaluation and Metrics (GEM ^2 ) , pages=
-
[15]
LLM s instead of Human Judges? A Large Scale Empirical Study across 20 NLP Evaluation Tasks
Bavaresco, Anna and Bernardi, Raffaella and Bertolazzi, Leonardo and Elliott, Desmond and Fern \'a ndez, Raquel and Gatt, Albert and Ghaleb, Esam and Giulianelli, Mario and Hanna, Michael and Koller, Alexander and Martins, Andre and Mondorf, Philipp and Neplenbroek, Vera and P...
2025 doi
-
[16]
Advances in Neural Information Processing Systems , volume=
Self-refine: Iterative refinement with self-feedback , author=. Advances in Neural Information Processing Systems , volume=
-
[17]
npj Health Systems , volume=
Human evaluation of large language models in healthcare: gaps, challenges, and the need for standardization , author=. npj Health Systems , volume=. 2025 , publisher=
2025
-
[18]
Proceedings of the 2024 ACM Conference on Fairness, Accountability, and Transparency , pages=
Escalation risks from language models in military and diplomatic decision-making , author=. Proceedings of the 2024 ACM Conference on Fairness, Accountability, and Transparency , pages=
2024
-
[19]
Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers) , pages=
Can language model moderators improve the health of online discourse? , author=. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers) , pages=
2024
-
[20]
Proceedings of the 31st International Conference on Computational Linguistics , pages=
Internlm-law: An open-sourced chinese legal large language model , author=. Proceedings of the 31st International Conference on Computational Linguistics , pages=
-
[21]
arXiv preprint arXiv:2405.05860 , year=
The perspectivist paradigm shift: Assumptions and challenges of capturing human labels , author=. arXiv preprint arXiv:2405.05860 , year=
-
[22]
Nature human behaviour , volume=
A manifesto for reproducible science , author=. Nature human behaviour , volume=. 2017 , publisher=
2017
-
[23]
2022 , url =
Shaurya Rohatgi , title =. 2022 , url =
2022
-
[24]
On Context Utilization in Summarization with Large Language Models
Ravaut, Mathieu and Sun, Aixin and Chen, Nancy and Joty, Shafiq. On Context Utilization in Summarization with Large Language Models. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. doi:10.18653/v1/2024.acl-long.153
2024 doi
-
[25]
ArXiv , year=
How Reliable Are Automatic Evaluation Methods for Instruction-Tuned LLMs? , author=. ArXiv , year=
-
[26]
arXiv preprint arXiv:2312.07559 , year=
Paperqa: Retrieval-augmented generative agent for scientific research , author=. arXiv preprint arXiv:2312.07559 , year=
-
[27]
arXiv e-prints , pages=
A foundation model for human-AI collaboration in medical literature mining , author=. arXiv e-prints , pages=
-
[28]
arXiv preprint arXiv:2411.14199 , year=
Openscholar: Synthesizing scientific literature with retrieval-augmented lms , author=. arXiv preprint arXiv:2411.14199 , year=
-
[29]
Text summarization branches out , pages=
Rouge: A package for automatic evaluation of summaries , author=. Text summarization branches out , pages=
-
[30]
arXiv preprint arXiv:2310.06825 , year=
Mistral 7B , author=. arXiv preprint arXiv:2310.06825 , year=
-
[31]
arXiv preprint arXiv:2303.08774 , year=
Gpt-4 technical report , author=. arXiv preprint arXiv:2303.08774 , year=
-
[32]
Clinical Natural Language Processing Workshop , year=
Generating medically-accurate summaries of patient-provider dialogue: A multi-stage approach using large language models , author=. Clinical Natural Language Processing Workshop , year=
-
[33]
Conference on Empirical Methods in Natural Language Processing , year=
Hierarchical Catalogue Generation for Literature Review: A Benchmark , author=. Conference on Empirical Methods in Natural Language Processing , year=
-
[34]
Annual Meeting of the Association for Computational Linguistics , year=
Multi-News: A Large-Scale Multi-Document Summarization Dataset and Abstractive Hierarchical Model , author=. Annual Meeting of the Association for Computational Linguistics , year=
-
[35]
Annual Meeting of the Association for Computational Linguistics , year=
Hierarchical Transformers for Multi-Document Summarization , author=. Annual Meeting of the Association for Computational Linguistics , year=
-
[36]
Annual Meeting of the Association for Computational Linguistics , year=
Summ ^N : A Multi-Stage Summarization Framework for Long Input Dialogues and Documents , author=. Annual Meeting of the Association for Computational Linguistics , year=
-
[37]
ArXiv , year=
Leveraging Long-Context Large Language Models for Multi-Document Understanding and Summarization in Enterprise Applications , author=. ArXiv , year=
-
[38]
, author=
A Hierarchical Decoder with Three-level Hierarchical Attention to Generate Abstractive Summaries of Interleaved Texts. , author=. arXiv: Computation and Language , year=
-
[39]
ArXiv , year=
uMedSum: A Unified Framework for Advancing Medical Abstractive Summarization , author=. ArXiv , year=
-
[40]
, author=
Roles of Document Structure, Cognitive Strategy, and Awareness in Searching for Information. , author=. Reading Research Quarterly , year=
-
[41]
, author=
The Effects of Text Structure Instruction on Middle-Grade Students' Comprehension and Production of Expository Text. , author=. Reading Research Quarterly , year=
-
[42]
and Yang, Qiang and Xie, Xing , title =
Chang, Yupeng and Wang, Xu and Wang, Jindong and Wu, Yuan and Yang, Linyi and Zhu, Kaijie and Chen, Hao and Yi, Xiaoyuan and Wang, Cunxiang and Wang, Yidong and Ye, Wei and Zhang, Yue and Chang, Yi and Yu, Philip S. and Yang, Qiang and Xie, Xing , title =. ACM Trans. Intell. S...
2024 doi
-
[43]
Improving Factuality in Clinical Abstractive Multi-Document Summarization by Guided Continued Pre-training
Elhady, Ahmed and Elsayed, Khaled and Agirre, Eneko and Artetxe, Mikel. Improving Factuality in Clinical Abstractive Multi-Document Summarization by Guided Continued Pre-training. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computati...
2024 doi
-
[44]
arXiv preprint arXiv:2305.14251 , year=
Factscore: Fine-grained atomic evaluation of factual precision in long form text generation , author=. arXiv preprint arXiv:2305.14251 , year=
-
[45]
arXiv preprint arXiv:2501.03545 , year=
Beyond Factual Accuracy: Evaluating Coverage of Diverse Factual Information in Long-form Text Generation , author=. arXiv preprint arXiv:2501.03545 , year=
-
[46]
Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers) , pages=
Embrace Divergence for Richer Insights: A Multi-document Summarization Benchmark and a Case Study on Summarizing Diverse Information from News Articles , author=. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics:...
2024
-
[47]
medRxiv , pages=
Synthetic Data Distillation Enables the Extraction of Clinical Information at Scale , author=. medRxiv , pages=. 2024 , publisher=
2024
-
[48]
and Villarroel, Mauricio and Clifford, Gari D
Lee, Joon and Scott, Daniel J. and Villarroel, Mauricio and Clifford, Gari D. and Saeed, Mohammed and Mark, Roger G. , booktitle=. Open-access MIMIC-II database for intensive care research , year=
-
[49]
ArXiv , year=
ClinicalBERT: Modeling Clinical Notes and Predicting Hospital Readmission , author=. ArXiv , year=
-
[50]
Bioinformatics , volume =
Lee, Jinhyuk and Yoon, Wonjin and Kim, Sungdong and Kim, Donghyeon and Kim, Sunkyu and So, Chan Ho and Kang, Jaewoo , title =. Bioinformatics , volume =. 2019 , month =. doi:10.1093/bioinformatics/btz682 , url =
2019 doi
-
[51]
TOPICAL : TOPIC Pages A utomagica L ly
Giorgi, John and Singh, Amanpreet and Downey, Doug and Feldman, Sergey and Wang, Lucy. TOPICAL : TOPIC Pages A utomagica L ly. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume...
2024 doi
-
[52]
What`s in a Summary? Laying the Groundwork for Advances in Hospital-Course Summarization
Adams, Griffin and Alsentzer, Emily and Ketenci, Mert and Zucker, Jason and Elhadad, No \'e mie. What`s in a Summary? Laying the Groundwork for Advances in Hospital-Course Summarization. Proceedings of the 2021 Conference of the North American Chapter of the Association for Co...
2021 doi
-
[53]
2015 26th international workshop on database and expert systems applications (dexa) , pages=
Clinical decision support systems: a survey of NLP-based approaches from unstructured data , author=. 2015 26th international workshop on database and expert systems applications (dexa) , pages=. 2015 , organization=
2015
-
[54]
Journal of Intelligent Connectivity and Emerging Technologies , volume=
Natural language processing for clinical decision support systems: A review of recent advances in healthcare , author=. Journal of Intelligent Connectivity and Emerging Technologies , volume=
-
[55]
Journal of biomedical informatics , volume=
What can natural language processing do for clinical decision support? , author=. Journal of biomedical informatics , volume=. 2009 , publisher=
2009
-
[56]
A Novel System for Extractive Clinical Note Summarization using EHR Data
Liang, Jennifer and Tsou, Ching-Huei and Poddar, Ananya. A Novel System for Extractive Clinical Note Summarization using EHR Data. Proceedings of the 2nd Clinical Natural Language Processing Workshop. 2019. doi:10.18653/v1/W19-1906
2019 doi
-
[57]
Generating SOAP Notes from Doctor-Patient Conversations Using Modular Summarization Techniques
Krishna, Kundan and Khosla, Sopan and Bigham, Jeffrey and Lipton, Zachary C. Generating SOAP Notes from Doctor-Patient Conversations Using Modular Summarization Techniques. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th Int...
2021 doi
-
[58]
Towards Automating Medical Scribing : Clinic Visit D ialogue2 N ote Sentence Alignment and Snippet Summarization
Yim, Wen-wai and Yetisgen, Meliha. Towards Automating Medical Scribing : Clinic Visit D ialogue2 N ote Sentence Alignment and Snippet Summarization. Proceedings of the Second Workshop on Natural Language Processing for Medical Conversations. 2021. doi:10.18653/v1/2021.nlpmc-1.2
2021 doi
-
[59]
DERA : Enhancing Large Language Model Completions with Dialog-Enabled Resolving Agents
Nair, Varun and Schumacher, Elliot and Tso, Geoffrey and Kannan, Anitha. DERA : Enhancing Large Language Model Completions with Dialog-Enabled Resolving Agents. Proceedings of the 6th Clinical Natural Language Processing Workshop. 2024. doi:10.18653/v1/2024.clinicalnlp-1.12
2024 doi
-
[60]
JMIR medical education , volume=
How does ChatGPT perform on the United States Medical Licensing Examination (USMLE)? The implications of large language models for medical education and knowledge assessment , author=. JMIR medical education , volume=. 2023 , publisher=
2023
-
[61]
Cureus , volume=
Overview of early ChatGPT’s presence in medical literature: insights from a hybrid literature review by ChatGPT and human experts , author=. Cureus , volume=. 2023 , publisher=
2023
-
[62]
ArXiv , year=
Bio-SIEVE: Exploring Instruction Tuning Large Language Models for Systematic Review Automation , author=. ArXiv , year=
-
[63]
CoRR , volume =
Athanasios Lagopoulos and Grigorios Tsoumakas , title =. CoRR , volume =. 2020 , url =. 2011.09752 , timestamp =
2020
-
[64]
Benchmarking Large Language Models for News Summarization
Zhang, Tianyi and Ladhak, Faisal and Durmus, Esin and Liang, Percy and McKeown, Kathleen and Hashimoto, Tatsunori B. Benchmarking Large Language Models for News Summarization. Transactions of the Association for Computational Linguistics. 2024. doi:10.1162/tacl_a_00632
2024 doi
-
[65]
On Learning to Summarize with Large Language Models as References
Liu, Yixin and Shi, Kejian and He, Katherine and Ye, Longtian and Fabbri, Alexander and Liu, Pengfei and Radev, Dragomir and Cohan, Arman. On Learning to Summarize with Large Language Models as References. Proceedings of the 2024 Conference of the North American Chapter of the...
2024 doi
-
[66]
Summarizing, Simplifying, and Synthesizing Medical Evidence using GPT -3 (with Varying Success)
Shaib, Chantal and Li, Millicent and Joseph, Sebastian and Marshall, Iain and Li, Junyi Jessy and Wallace, Byron. Summarizing, Simplifying, and Synthesizing Medical Evidence using GPT -3 (with Varying Success). Proceedings of the 61st Annual Meeting of the Association for Comp...
2023 doi
-
[67]
ACM Comput
Koh, Huan Yee and Ju, Jiaxin and Liu, Ming and Pan, Shirui , title =. ACM Comput. Surv. , month = dec, articleno =. 2022 , issue_date =. doi:10.1145/3545176 , abstract =
2022 doi
-
[68]
Conference on Empirical Methods in Natural Language Processing , year=
DocAsRef: An Empirical Study on Repurposing Reference-based Summary Quality Metrics as Reference-free Metrics , author=. Conference on Empirical Methods in Natural Language Processing , year=
-
[69]
ACM Trans
Nenkova, Ani and Passonneau, Rebecca and McKeown, Kathleen , title =. ACM Trans. Speech Lang. Process. , month = may, pages =. 2007 , issue_date =. doi:10.1145/1233912.1233913 , abstract =
2007 doi
-
[70]
ACM Comput
Jangra, Anubhav and Mukherjee, Sourajit and Jatowt, Adam and Saha, Sriparna and Hasanuzzaman, Mohammad , title =. ACM Comput. Surv. , month = jul, articleno =. 2023 , issue_date =. doi:10.1145/3584700 , abstract =
2023 doi
-
[71]
Annual Meeting of the Association for Computational Linguistics , year=
A Simple Theoretical Model of Importance for Summarization , author=. Annual Meeting of the Association for Computational Linguistics , year=
-
[72]
ArXiv , year=
Earlier Isn’t Always Better: Sub-aspect Analysis on Corpus and System Biases in Summarization , author=. ArXiv , year=
-
[73]
Journal of Medical Internet Research , year=
Potential Roles of Large Language Models in the Production of Systematic Reviews and Meta-Analyses , author=. Journal of Medical Internet Research , year=
-
[74]
BMJ Open , year=
Analysis of the time and workers needed to conduct systematic reviews of medical interventions using data from the PROSPERO registry , author=. BMJ Open , year=
-
[75]
Energies , year=
The Resilience of Critical Infrastructure Systems: A Systematic Literature Review , author=. Energies , year=
-
[76]
Proceedings of the 52nd annual meeting of the association for computational linguistics (volume 1: Long papers) , pages=
Hierarchical summarization: Scaling up multi-document summarization , author=. Proceedings of the 52nd annual meeting of the association for computational linguistics (volume 1: Long papers) , pages=
-
[77]
Generating (Factual?) Narrative Summaries of RCTs: Experiments with Neural Multi-Document Summarization , author=. AMIA ... Annual Symposium proceedings. AMIA Symposium , year=
-
[78]
Proceedings of the conference
What’s in a Summary? Laying the Groundwork for Advances in Hospital-Course Summarization , author=. Proceedings of the conference. Association for Computational Linguistics. North American Chapter. Meeting , year=
-
[79]
Trends in cognitive sciences , volume=
Hierarchical process memory: memory as an integral component of information processing , author=. Trends in cognitive sciences , volume=. 2015 , publisher=
2015
-
[80]
Craik , abstract =
F.I.M. Craik , abstract =. Memory: Levels of Processing , editor =. International Encyclopedia of the Social & Behavioral Sciences , publisher =. 2001 , isbn =. doi:https://doi.org/10.1016/B0-08-043076-7/01508-4 , url =
2001 doi
-
[81]
Conference on Empirical Methods in Natural Language Processing , year=
Summarizing Multiple Documents with Conversational Structure for Meta-Review Generation , author=. Conference on Empirical Methods in Natural Language Processing , year=
-
[82]
ArXiv , year=
Nutri-bullets: Summarizing Health Studies by Composing Segments , author=. ArXiv , year=
-
[83]
ArXiv , year=
Multi-Document Scientific Summarization from a Knowledge Graph-Centric View , author=. ArXiv , year=
-
[84]
Annual Meeting of the Association for Computational Linguistics , year=
HIBRIDS: Attention with Hierarchical Biases for Structure-aware Long Document Summarization , author=. Annual Meeting of the Association for Computational Linguistics , year=
-
[85]
Proceedings of the 45th International ACM SIGIR Conference on Research and Development in Information Retrieval , year=
Target-aware Abstractive Related Work Generation with Contrastive Learning , author=. Proceedings of the 45th International ACM SIGIR Conference on Research and Development in Information Retrieval , year=
-
[86]
ArXiv , year=
Disentangling Instructive Information from Ranked Multiple Candidates for Multi-Document Scientific Summarization , author=. ArXiv , year=
-
[87]
H i S truct+: Improving Extractive Text Summarization with Hierarchical Structure Information
Ruan, Qian and Ostendorff, Malte and Rehm, Georg. H i S truct+: Improving Extractive Text Summarization with Hierarchical Structure Information. Findings of the Association for Computational Linguistics: ACL 2022. 2022. doi:10.18653/v1/2022.findings-acl.102
2022 doi
-
[88]
Evaluation Metrics in the Era of GPT -4: Reliably Evaluating Large Language Models on Sequence to Sequence Tasks
Sottana, Andrea and Liang, Bin and Zou, Kai and Yuan, Zheng. Evaluation Metrics in the Era of GPT -4: Reliably Evaluating Large Language Models on Sequence to Sequence Tasks. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023. doi:10.1...
2023 doi
-
[89]
and Belz, Anya and Du s ek, Ond r ej and Mille, Simon and van der Lee, Chris and Reiter, Ehud and Santhanam, Shivani and Thomson, Craig
Howcroft, David M. and Belz, Anya and Du s ek, Ond r ej and Mille, Simon and van der Lee, Chris and Reiter, Ehud and Santhanam, Shivani and Thomson, Craig. A Survey on the Evaluation of Natural Language Generation: Past, Present, and Future. Proceedings of the 17th Internation...
2024
-
[90]
G-Eval : NLG Evaluation using GPT-4 with Better Human Alignment
Liu, Yixuan and Bansal, Mohit. G-Eval : NLG Evaluation using GPT-4 with Better Human Alignment. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing (EMNLP). 2023
2023
-
[91]
HD -Eval: Aligning Large Language Model Evaluators Through Hierarchical Criteria Decomposition
Liu, Yuxuan and Yang, Tianchi and Huang, Shaohan and Zhang, Zihan and Huang, Haizhen and Wei, Furu and Deng, Weiwei and Sun, Feng and Zhang, Qi. HD -Eval: Aligning Large Language Model Evaluators Through Hierarchical Criteria Decomposition. Proceedings of the 62nd Annual Meeti...
2024 doi
-
[92]
arXiv preprint arXiv:2402.15754 , year=
HD-Eval: Aligning Large Language Model Evaluators Through Human Demonstration , author=. arXiv preprint arXiv:2402.15754 , year=
-
[93]
Findings of the Association for Computational Linguistics: EMNLP 2023 , year=
GPT-4 as an Effective Zero-Shot Evaluator for Scientific Figure Captioning , author=. Findings of the Association for Computational Linguistics: EMNLP 2023 , year=
2023
-
[94]
Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing (EMNLP) , year=
G-Eval: NLG Evaluation using GPT-4 with Better Human Alignment , author=. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing (EMNLP) , year=
2023
-
[95]
Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers)
Qiu, Haoyi and Huang, Kung-Hsiang and Qu, Jingnong and Peng, Nanyun. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. doi:10.18653/v1/2024.naacl-long.33
2024 doi
-
[96]
Transactions of the Association for Computational Linguistics , volume=
SummEval: Re-evaluating Summarization Evaluation , author=. Transactions of the Association for Computational Linguistics , volume=. 2021 , publisher=
2021
-
[97]
arXiv preprint arXiv:2011.13662 , year=
FFCI: A Framework for Interpretable Automatic Evaluation of Summarization , author=. arXiv preprint arXiv:2011.13662 , year=
2011
-
[98]
Journal of the Royal Society of Medicine , volume =
Khalid S Khan and Regina Kunz and Jos Kleijnen and Gerd Antes , title =. Journal of the Royal Society of Medicine , volume =. 2003 , doi =
2003
-
[99]
The significant cost of systematic reviews and meta-analyses: A call for greater involvement of machine learning to assess the promise of clinical trials , journal =
Matthew Michelson and Katja Reuter , keywords =. The significant cost of systematic reviews and meta-analyses: A call for greater involvement of machine learning to assess the promise of clinical trials , journal =. 2019 , issn =. doi:https://doi.org/10.1016/j.conctc.2019.1004...
2019 doi
-
[100]
npj Digital Medicine , volume=
Closing the gap between open source and commercial large language models for medical evidence summarization , author=. npj Digital Medicine , volume=. 2024 , publisher=
2024
-
[101]
arXiv preprint arXiv:2104.06486 , year=
Ms2: Multi-document summarization of medical studies , author=. arXiv preprint arXiv:2104.06486 , year=
-
[102]
arXiv preprint arXiv:2305.13693 , year=
Automated metrics for medical multi-document summarization disagree with human evaluations , author=. arXiv preprint arXiv:2305.13693 , year=
-
[103]
arXiv preprint arXiv:2402.02420 , year=
Factuality of large language models in the year 2024 , author=. arXiv preprint arXiv:2402.02420 , year=
2024
-
[104]
arXiv preprint arXiv:2307.13528 , year=
FacTool: Factuality Detection in Generative AI--A Tool Augmented Framework for Multi-Task and Multi-Domain Scenarios , author=. arXiv preprint arXiv:2307.13528 , year=
-
[105]
Findings of the Association for Computational Linguistics: ACL 2023 , pages=
NonFactS: NonFactual summary generation for factuality evaluation in document summarization , author=. Findings of the Association for Computational Linguistics: ACL 2023 , pages=
2023
-
[106]
Applied Cognitive Psychology: The Official Journal of the Society for Applied Research in Memory and Cognition , volume=
Consequences of erudite vernacular utilized irrespective of necessity: Problems with using long words needlessly , author=. Applied Cognitive Psychology: The Official Journal of the Society for Applied Research in Memory and Cognition , volume=. 2006 , publisher=
2006
-
[107]
and Maryam Kouchaki and Hancock, Jeffrey T
Markowitz, David M. and Maryam Kouchaki and Hancock, Jeffrey T. and Francesca Gino. The Deception Spiral: Corporate Obfuscation Leads to Perceptions of Immorality and Cheating Behavior. Journal of Language and Social Psychology. 2021. doi:10.1177/0261927X20949594
2021 doi
-
[108]
PNAS Nexus , volume =
Markowitz, David M , title =. PNAS Nexus , volume =. 2024 , month =. doi:10.1093/pnasnexus/pgae387 , url =
2024 doi
-
[109]
CHIME : LLM -Assisted Hierarchical Organization of Scientific Studies for Literature Review Support
Hsu, Chao-Chun and Bransom, Erin and Sparks, Jenna and Kuehl, Bailey and Tan, Chenhao and Wadden, David and Wang, Lucy and Naik, Aakanksha. CHIME : LLM -Assisted Hierarchical Organization of Scientific Studies for Literature Review Support. Findings of the Association for Comp...
2024 doi
-
[110]
FIZZ : Factual Inconsistency Detection by Zoom-in Summary and Zoom-out Document
Yang, Joonho and Yoon, Seunghyun and Kim, ByeongJeong and Lee, Hwanhee. FIZZ : Factual Inconsistency Detection by Zoom-in Summary and Zoom-out Document. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. doi:10.18653/v1/2024.emnlp-main.3
2024 doi
-
[111]
On the Role of Summary Content Units in Text Summarization Evaluation
Nawrath, Marcel and Nowak, Agnieszka and Ratz, Tristan and Walenta, Danilo and Opitz, Juri and Ribeiro, Leonardo and Sedoc, Jo \ a o and Deutsch, Daniel and Mille, Simon and Liu, Yixin and Gehrmann, Sebastian and Zhang, Lining and Mahamood, Saad and Clinciu, Miruna and Chandu,...
2024 doi
-
[112]
arXiv preprint arXiv:2109.11503 , year=
Finding a balanced degree of automation for summary evaluation , author=. arXiv preprint arXiv:2109.11503 , year=
-
[113]
Weinberger and Yoav Artzi , title =
Tianyi Zhang and Varsha Kishore and Felix Wu and Kilian Q. Weinberger and Yoav Artzi , title =. CoRR , volume =. 2019 , url =. 1904.09675 , timestamp =
2019 arXiv
-
[114]
arXiv preprint arXiv:2304.02554 , year=
Human-like summarization evaluation with chatgpt , author=. arXiv preprint arXiv:2304.02554 , year=
-
[115]
ACM Computing Surveys (CSUR) , volume=
A survey of evaluation metrics used for NLG systems , author=. ACM Computing Surveys (CSUR) , volume=. 2022 , publisher=
2022
-
[116]
Evolutionary Applications , volume=
Next-generation metrics for monitoring genetic erosion within populations of conservation concern , author=. Evolutionary Applications , volume=. 2018 , publisher=
2018
-
[117]
Journal of the Association for Information Science and Technology , volume=
Measuring text difficulty using parse-tree frequency , author=. Journal of the Association for Information Science and Technology , volume=. 2017 , publisher=
2017
-
[118]
APPLS : Evaluating Evaluation Metrics for Plain Language Summarization
Guo, Yue and August, Tal and Leroy, Gondy and Cohen, Trevor and Wang, Lucy Lu. APPLS : Evaluating Evaluation Metrics for Plain Language Summarization. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. doi:10.18653/v1/2024.emnlp-main.519
2024 doi
-
[119]
Aho and Jeffrey D
Alfred V. Aho and Jeffrey D. Ullman , title =. 1972
1972
-
[120]
Publications Manual , year = "1983", publisher =
1983
-
[121]
Chandra and Dexter C
Ashok K. Chandra and Dexter C. Kozen and Larry J. Stockmeyer , year = "1981", title =. doi:10.1145/322234.322243
1981 doi
-
[122]
Scalable training of
Andrew, Galen and Gao, Jianfeng , booktitle=. Scalable training of
-
[123]
Dan Gusfield , title =. 1997
1997
-
[124]
Tetreault , title =
Mohammad Sadegh Rasooli and Joel R. Tetreault , title =. Computing Research Repository , volume =. 2015 , url =
2015
-
[125]
A Framework for Learning Predictive Structures from Multiple Tasks and Unlabeled Data , Volume =
Ando, Rie Kubota and Zhang, Tong , Issn =. A Framework for Learning Predictive Structures from Multiple Tasks and Unlabeled Data , Volume =. Journal of Machine Learning Research , Month = dec, Numpages =
-
[126]
arXiv preprint arXiv:2005.13213 , year=
Give me convenience and give her death: Who should decide what uses of NLP are appropriate, and on what basis? , author=. arXiv preprint arXiv:2005.13213 , year=
2005
-
[127]
Impact of Annotator Demographics on Sentiment Dataset Labeling , year =
Ding, Yi and You, Jacob and Machulla, Tonja-Katrin and Jacobs, Jennifer and Sen, Pradeep and H\". Impact of Annotator Demographics on Sentiment Dataset Labeling , year =. Proc. ACM Hum.-Comput. Interact. , month = nov, articleno =. doi:10.1145/3555632 , abstract =
-
[128]
Identifying and Measuring Annotator Bias Based on Annotators' Demographic Characteristics
Al Kuwatly, Hala and Wich, Maximilian and Groh, Georg. Identifying and Measuring Annotator Bias Based on Annotators' Demographic Characteristics. Proceedings of the Fourth Workshop on Online Abuse and Harms. 2020. doi:10.18653/v1/2020.alw-1.21
2020 doi
-
[129]
Proceedings of the 57th annual meeting of the association for computational linguistics , pages=
The risk of racial bias in hate speech detection , author=. Proceedings of the 57th annual meeting of the association for computational linguistics , pages=
-
[130]
International journal of qualitative methods , volume=
Demonstrating rigor using thematic analysis: A hybrid approach of inductive and deductive coding and theme development , author=. International journal of qualitative methods , volume=. 2006 , publisher=
2006
-
[131]
Transactions of the Association for Computational Linguistics , volume=
Bridging the gap: A survey on integrating (human) feedback for natural language generation , author=. Transactions of the Association for Computational Linguistics , volume=. 2023 , publisher=
2023
-
[132]
arXiv preprint arXiv:2006.14799 , year=
Evaluation of text generation: A survey , author=. arXiv preprint arXiv:2006.14799 , year=
2006
-
[133]
Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=
Considers-the-human evaluation framework: Rethinking human evaluation for generative large language models , author=. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=
-
[134]
Proceedings of the conference on fairness, accountability, and transparency , pages=
Model cards for model reporting , author=. Proceedings of the conference on fairness, accountability, and transparency , pages=
-
[135]
arXiv preprint arXiv:2109.06598 , year=
Just what do you think you're doing, Dave?'a checklist for responsible data use in NLP , author=. arXiv preprint arXiv:2109.06598 , year=
-
[136]
arXiv preprint arXiv:1909.03004 , year=
Show your work: Improved reporting of experimental results , author=. arXiv preprint arXiv:1909.03004 , year=
1909
-
[137]
Proceedings of the 13th international conference on natural language generation , pages=
Twenty years of confusion in human evaluation: NLG needs evaluation sheets and standardised definitions , author=. Proceedings of the 13th international conference on natural language generation , pages=
-
[138]
The Perspectivist Paradigm Shift: Assumptions and Challenges of Capturing Human Labels
Fleisig, Eve and Blodgett, Su Lin and Klein, Dan and Talat, Zeerak. The Perspectivist Paradigm Shift: Assumptions and Challenges of Capturing Human Labels. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human ...
2024 doi
-
[139]
`Just What do You Think You ' re Doing, Dave?' A Checklist for Responsible Data Use in NLP
Rogers, Anna and Baldwin, Timothy and Leins, Kobi. `Just What do You Think You ' re Doing, Dave?' A Checklist for Responsible Data Use in NLP. Findings of the Association for Computational Linguistics: EMNLP 2021. 2021. doi:10.18653/v1/2021.findings-emnlp.414
2021 doi
-
[140]
Show Your Work: Improved Reporting of Experimental Results
Dodge, Jesse and Gururangan, Suchin and Card, Dallas and Schwartz, Roy and Smith, Noah A. Show Your Work: Improved Reporting of Experimental Results. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conferen...
2019 doi
Reviewed June 27, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.