REVIEW 20 cited by
Evaluation of Text Generation: A Survey
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
The paper surveys evaluation methods of natural language generation (NLG) systems that have been developed in the last few years. We group NLG evaluation methods into three categories: (1) human-centric evaluation metrics, (2) automatic metrics that require no training, and (3) machine-learned metrics. For each category, we discuss the progress that has been made and the challenges still being faced, with a focus on the evaluation of recently proposed NLG tasks and neural NLG models. We then present two examples for task-specific NLG evaluations for automatic text summarization and long text generation, and conclude the paper by proposing future research directions.
Forward citations
Cited by 20 Pith papers
-
The Reader is the Metric: How Textual Features and Reader Profiles Explain Conflicting Evaluations of AI Creative Writing
Reader evaluations of AI versus human stories split into two measurable preference profiles, surface-focused and holistic, that track reader expertise and explain why prior studies disagree.
-
LLM-as-a-Judge for Evaluating System Responses in Conversational Music Recommendation
LLM judges agree moderately with human experts when scoring conversational music recommendation responses, outperform reference-based metrics, but are not reliable enough to replace human evaluation.
-
grapheme-kit: Grapheme-Level Metrics and Text Processing for Multilingual NLP
A new open-source library computes distance, similarity, and evaluation metrics on grapheme clusters rather than Unicode code points, and corrects ZWJ/ZWNJ segmentation for Tamil and Sinhala.
-
Preconditioned Test-Time Adaptation for Out-of-Distribution Debiasing in Narrative Generation
CAP-TTA triggers context-aware preconditioned LoRA updates on high bias-risk OOD prompts to reduce toxicity in LLM narrative generation while preserving fluency and avoiding catastrophic forgetting.
-
What Are We Measuring in NLG? A Meta-Analysis of Evaluation Trends 2020-2025
Across 3,334 NLG papers from 2020-2025, legacy n-gram metrics persist, LLM-as-a-judge usage outpaced human validation, and reported LLM-human correlations collapse on criteria like fluency.
-
LLMs for Customized Marketing Content Generation and Evaluation at Scale
MarketingFM generates e-commerce ad copy with RAG and an LLM; AutoEval uses LLM-as-a-Judge plus rule checks and self-refines its prompts, with online tests showing significant clicks and impressions lifts but no signi...
-
COGENT: A Curriculum-oriented Framework for Generating Grade-appropriate Educational Content
Adding curriculum decomposition, readability constraints, and wonder-based topics to LLM prompts yields science reading passages with higher curriculum alignment and human-comparable comprehensibility.
-
Evaluating and Mitigating Bias in AI-Based Medical Text Generation
A training method that selects high-loss examples reduces demographic performance gaps in medical text generation by over 30% across models and datasets.
-
ExPerT: Effective and Explainable Evaluation of Personalized Long-Form Text Generation
ExPerT is a reference-based LLM evaluation metric that extracts and matches atomic aspects, scores content and style, and reports 0.74 human alignment on LongLaMP, a 7.2% relative gain over GEMBA and G-Eval.
-
Interactive Information Need Prediction with Intent and Context
User-selected context plus a short partial intent lets language models generate or retrieve the full information need; partial intent mitigates the noise of larger contexts in adapted Inquisitive and MS MARCO settings.
-
Challenges in Trustworthy Human Evaluation of Chatbots
Open chatbot leaderboards like Chatbot Arena can have their model rankings moved by several positions with only 10% low-quality or adversarial votes.
-
I Can Tell What I am Doing: Toward Real-World Natural Language Grounding of Robot Experiences
RONAR is an LLM-based framework that narrates a mobile robot's experiences in natural language, and its user studies show that these narrations help people localize and explain robot failures faster than raw video interfaces.
-
Fidelity-Diversity Metrics for Text
Optimal-transport fidelity and diversity scores on text embeddings disentangle support mismatch from mass coverage and predict GSM8K finetuning accuracy on synthetic math data.
-
From Multimodal Perception to Strategic Reasoning: A Survey on AI-Generated Game Commentary
A survey that organizes AI-generated game commentary research into a taxonomy of three commentator capabilities and three commentary types, with a review of methods, datasets, and metrics.
-
GLIDER: Grading LLM Interactions and Decisions using Explainable Ranking
A 3.8B evaluator model trained on synthetic data across 685 domains and 183 criteria scores text on arbitrary rubrics with highlighted evidence and matches GPT-4o on FLASK at a fraction of the size.
-
FairJudge: An Adaptive, Debiased, and Consistent LLM-as-a-Judge
A three-stage SFT-DPO-GRPO curriculum on a JudgeLM-derived dataset improves LLM judge agreement/F1, but debiasing and consistency claims are evaluated on a benchmark built from the same pipeline.
-
Statistical Multicriteria Evaluation of LLM-Generated Text
Using generalized stochastic dominance, the authors find that human-written text completions are not significantly outperformed by five LLM decoding strategies across mixed cardinal and ordinal quality metrics.
-
CPR: Leveraging LLMs for Topic and Phrase Suggestion to Facilitate Comprehensive Product Reviews
CPR generates product-review phrase suggestions from topic ratings and LLM-generated topic lists for new products, but the headline 12.3% BLEU claim rests on a cherry-picked, likely self-referential evaluation.
-
AltGen: AI-Driven Alt Text Generation for Enhancing EPUB Accessibility
AltGen applies a standard image-captioning pipeline to generate alt text for EPUB images; the paper reports strong quality scores but gives insufficient detail to verify them.
-
Natural Language Understanding and Inference with MLLM in Visual Question Answering: A Survey
A survey that organizes VQA methods from feature extraction through MLLM reasoning, datasets, and metrics, without introducing new experimental results.
Discussion (0). Continue with ORCID to comment.