REVIEW 20 cited by
LLM-based NLG Evaluation: Current Status and Challenges
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Evaluating natural language generation (NLG) is a vital but challenging problem in natural language processing. Traditional evaluation metrics mainly capturing content (e.g. n-gram) overlap between system outputs and references are far from satisfactory, and large language models (LLMs) such as ChatGPT have demonstrated great potential in NLG evaluation in recent years. Various automatic evaluation methods based on LLMs have been proposed, including metrics derived from LLMs, prompting LLMs, fine-tuning LLMs, and human-LLM collaborative evaluation. In this survey, we first give a taxonomy of LLM-based NLG evaluation methods, and discuss their pros and cons, respectively. Lastly, we discuss several open problems in this area and point out future research directions.
Forward citations
Cited by 20 Pith papers
-
MDEval: Evaluating and Enhancing Markdown Awareness in Large Language Models
MDEval scores Markdown Awareness as the normalized edit distance between a model's HTML-tagged output and a GPT-4o rewrite, and reports human-alignment accuracy of 84.1% when ties are excluded.
-
From Execution to Education: A Bloom-Aligned Framework for Measuring Educational Control in LLMs
On 2,520 programming tasks, matched Qwen general and coder models reliably raise Bloom cognitive demand but fail to lower it, so execution skill does not imply educational control.
-
AI Analyst: Framework and Comprehensive Evaluation of Large Language Models for Financial Time Series Report Generation
LLMs such as GPT-4o can generate coherent financial reports from time series data, and a proposed highlighting system categorizes report segments by whether they stem from data, reasoning, or external knowledge.
-
BriefMe: A Legal NLP Benchmark for Assisting with Legal Briefs
BriefMe introduces a legal brief benchmark with argument summarization, argument completion, and case retrieval, and shows LLMs beat human headings on the first two but struggle on the latter two.
-
LecEval: An Automated Metric for Multimodal Knowledge Acquisition in Multimedia Learning
A fine-tuned multimodal LLM can rate slide-based lecture quality in line with human judges, outperforming generic metrics and prompted LLMs on a new annotated dataset.
-
Can LLMs Replace Human Evaluators? An Empirical Study of LLM-as-a-Judge in Software Engineering
Output-based LLM-as-a-judge methods using large LLMs achieve near-human correlation with human scores for code translation and generation, but not for code summarization or pairwise comparisons.
-
Optimization is Better than Generation: Optimizing Commit Message Leveraging Human-written Commit Message
A commit-message optimization method that starts from human-written messages and uses GPT-4 plus automated evaluators outperforms message generation and completion methods on three of four quality metrics.
-
Natural Language Reinforcement Learning
NLRL replaces scalar RL values with LLM-generated language narratives, trains language critics with language MC/TD, and improves policies via LLM-based policy iteration, outperforming PPO on four small agentic tasks.
-
FlashDP: Private Training Large Language Models with Efficient DP-SGD
FlashDP fuses per-sample gradient computation, norm calculation, clipping, and noise addition into a cache-friendly block-wise all-reduce workflow that avoids explicit per-sample gradient storage and redundant recomputation.
-
Enterprise Large Language Model Evaluation Benchmark
A 14-task enterprise LLM benchmark built mostly from GPT-4o-generated labels and scored by GPT-4o-as-judge shows open-source models closing the reasoning gap, but the dataset is not public and the evaluation is partly...
-
Enhancing Generalization of Speech Large Language Models with Multi-Task Behavior Imitation and Speech-Text Interleaving
Multi-task behavior imitation with speech-text interleaving improves speech LLM generalization on prompts and zero-shot tasks using only paired speech and transcripts.
-
FactsR: A Safer Method for Producing High Quality Healthcare Documentation
FactsR decomposes clinical note generation into real-time fact extraction and recursive refinement, reporting improved completeness and conciseness over a few-shot baseline on the 57-encounter Primock57 benchmark, tho...
-
Leveraging Chain of Thought towards Empathetic Spoken Dialogue without Corresponding Question-Answering Data
Listen, Perceive, Express (LPE) uses two-stage ASR/SER training plus chain-of-thought prompting to let a frozen LLM generate empathetic responses from speech without question-answer fine-tuning data.
-
Reasoning-Enhanced Self-Training for Long-Form Personalized Text Generation
REST-PG trains LLMs to reason over user profiles and self-train on high-reward outputs, improving personalized long-form generation by 14.5% over SFT on LongLaMP.
-
Evaluate Summarization in Fine-Granularity: Auto Evaluation with LLM
SumAutoEval is an LLM-based entity-level summarization evaluator with four dimensions; its claimed human-correlation advantage is not consistently supported by the experiments.
-
Engineering AI Judge Systems
A constitution-based, search-driven development framework for AI judge systems improves judged accuracy by up to 6.2% on commit message generation, with about 58% of general principles reused across five languages.
-
Empowering Meta-Analysis: Leveraging Large Language Models for Scientific Synthesis
Fine-tuning Llama-2 and Mistral 7B on a purpose-built dataset of meta-analysis abstracts improves the relevance of generated meta-analysis abstracts, but the gains rest on a small human evaluation and a questionable l...
-
Reducing Hallucinations in Complex Question Answering using Simple Graph-based Retrieval-Augmented Generation (long version)
A graph-augmented RAG system with vector and graph query tools halves hallucinations and raises factual correctness scores on the MoNaCo complex QA benchmark.
-
Improve LLM-based Automatic Essay Scoring with Linguistic Features
Adding handcrafted linguistic features to zero-shot LLM prompts modestly improves automatic essay scoring for Mistral-7B on ASAP and ELLIPSE, but GPT-4 shows no benefit on ASAP and no significance tests are provided.
-
The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs
A structured survey of LLM safety evaluation that proposes a why/what/where/how taxonomy and catalogs metrics, datasets, benchmarks, evaluators, and frameworks.
Discussion (0). Continue with ORCID to comment.