Pith. sign in

REVIEW 4 cited by

A Comparative Study of Quality Evaluation Methods for Text Summarization

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.00747 v1 pith:YN4EIJBW submitted 2024-06-30 cs.CL cs.AI

A Comparative Study of Quality Evaluation Methods for Text Summarization

classification cs.CL cs.AI
keywords evaluationsummarizationtextautomaticevaluatinghumanmetricscomparative
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Evaluating text summarization has been a challenging task in natural language processing (NLP). Automatic metrics which heavily rely on reference summaries are not suitable in many situations, while human evaluation is time-consuming and labor-intensive. To bridge this gap, this paper proposes a novel method based on large language models (LLMs) for evaluating text summarization. We also conducts a comparative study on eight automatic metrics, human evaluation, and our proposed LLM-based method. Seven different types of state-of-the-art (SOTA) summarization models were evaluated. We perform extensive experiments and analysis on datasets with patent documents. Our results show that LLMs evaluation aligns closely with human evaluation, while widely-used automatic metrics such as ROUGE-2, BERTScore, and SummaC do not and also lack consistency. Based on the empirical comparison, we propose a LLM-powered framework for automatically evaluating and improving text summarization, which is beneficial and could attract wide attention among the community.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. TF1-EN-3M: Three Million Synthetic Moral Fables for Training Small, Open Language Models

    cs.CL 2025-04 unverdicted novelty 7.0

    The authors generate and publicly release the first large-scale open dataset of three million structured moral fables produced by small open language models together with a reproducible LLM-judge evaluation pipeline.

  2. GigaChat Audio: Time-aware Large Audio Language Model

    eess.AS 2026-07 conditional novelty 6.0

    Interleaving periodic time markers with continuous audio tokens, plus duration-mixture synthetic training, yields stable temporal grounding for an audio LLM on inputs up to 120 minutes.

  3. LLM-ReSum: A Framework for LLM Reflective Summarization through Self-Evaluation

    cs.CL 2026-04 unverdicted novelty 6.0

    LLM-ReSum uses LLM self-evaluation in a closed feedback loop to refine summaries, improving factual accuracy by up to 33% and coverage by 39% with 89% human preference.

  4. LongSumEval: Question-Answering Based Evaluation and Feedback-Driven Refinement for Long Document Summarization

    cs.CL 2026-04 unverdicted novelty 6.0

    LongSumEval evaluates long-document summaries via answerability and factual alignment of generated QA pairs, yielding stronger human correlation than prior metrics and enabling iterative self-improvement.