Pith. sign in

REVIEW 23 cited by

News Summarization and Evaluation in the Era of GPT-3

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2209.12356 v2 pith:SM5S6LWJ submitted 2022-09-26 cs.CL

News Summarization and Evaluation in the Era of GPT-3

classification cs.CL
keywords summarizationgpt-3modelssummariesevaluateevaluationfine-tunedkeyword-based
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

The recent success of prompting large language models like GPT-3 has led to a paradigm shift in NLP research. In this paper, we study its impact on text summarization, focusing on the classic benchmark domain of news summarization. First, we investigate how GPT-3 compares against fine-tuned models trained on large summarization datasets. We show that not only do humans overwhelmingly prefer GPT-3 summaries, prompted using only a task description, but these also do not suffer from common dataset-specific issues such as poor factuality. Next, we study what this means for evaluation, particularly the role of gold standard test sets. Our experiments show that both reference-based and reference-free automatic metrics cannot reliably evaluate GPT-3 summaries. Finally, we evaluate models on a setting beyond generic summarization, specifically keyword-based summarization, and show how dominant fine-tuning approaches compare to prompting. To support further research, we release: (a) a corpus of 10K generated summaries from fine-tuned and prompt-based models across 4 standard summarization benchmarks, (b) 1K human preference judgments comparing different systems for generic- and keyword-based summarization.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 23 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. MemMorph: Tool Hijacking in LLM Agents via Memory Poisoning

    cs.CR 2026-05 unverdicted novelty 8.0

    MemMorph poisons LLM agent long-term memory with three crafted records disguised as facts or policies to hijack tool selection, reaching 85.9% success rate across 10 backbones and outperforming baselines while resisti...

  2. Faithful by Construction: Claim-Anchored Attribution for Multi-Document Summarization

    cs.CL 2026-06 unverdicted novelty 7.0

    CAMS framework extracts, clusters, selects, and rewrites atomic claims to produce multi-document summaries with fine-grained, multi-source traceability by construction.

  3. From Text to Voice: A Reproducible and Verifiable Framework for Evaluating Tool Calling LLM Agents

    cs.CL 2026-05 unverdicted novelty 7.0

    A dataset-agnostic framework converts text tool-calling benchmarks to paired audio versions via TTS and noise, showing model-dependent performance with small text-to-voice gaps of 1.8-4.8 points on Confetti and When2Call.

  4. A Recipe for Long-Context Reasoning in Large Language Models via On-Policy Optimization and Distillation

    cs.CL 2026-05 unverdicted novelty 7.0

    dGRPO merges outcome-based policy optimization with dense teacher guidance from on-policy distillation, yielding more stable long-context reasoning on the new LongBlocks synthetic dataset.

  5. Faithful by Construction: Claim-Anchored Attribution for Multi-Document Summarization

    cs.CL 2026-06 unverdicted novelty 6.0

    CAMS is a claim-anchored Extract-Select-Rewrite framework for multi-document summarization that anchors every summary sentence to support-checked atomic claims with multi-source traceability.

  6. M\"OVE: A Holistic LLM Benchmark for the German Public Sector

    cs.CL 2026-06 unverdicted novelty 6.0

    MÖVE presents a new German-language benchmark evaluating 39 LLMs on performance and governance criteria using ten public-administration datasets.

  7. Understanding LLM Behavior in Multi-Target Cross-Lingual Summarization

    cs.CL 2026-05 unverdicted novelty 6.0

    Introduces the MEA benchmark for multi-target cross-lingual summarization across 24 languages and demonstrates that activation steering from English summarization representations improves performance.

  8. From Text to Voice: A Reproducible and Verifiable Framework for Evaluating Tool Calling LLM Agents

    cs.CL 2026-05 unverdicted novelty 6.0

    A dataset-agnostic framework converts text tool-calling benchmarks to paired audio evaluations via TTS, speaker variation and noise, then evaluates seven omni-modal models showing model- and task-dependent performance...

  9. A Meta Reinforcement Learning Approach to Goals-Based Wealth Management

    cs.LG 2026-05 unverdicted novelty 6.0

    MetaRL pre-trained on GBWM problems delivers near-optimal dynamic strategies in 0.01s achieving 97.8% of DP optimal utility and handles larger problems where DP fails.

  10. LLM-ReSum: A Framework for LLM Reflective Summarization through Self-Evaluation

    cs.CL 2026-04 unverdicted novelty 6.0

    LLM-ReSum uses LLM self-evaluation in a closed feedback loop to refine summaries, improving factual accuracy by up to 33% and coverage by 39% with 89% human preference.

  11. Why is "Chicago" Predictive of Deceptive Reviews? Using LLMs to Discover Language Phenomena from Lexical Cues

    cs.CL 2025-11 unverdicted novelty 6.0

    A conjecture-then-validate method lets LLMs convert opaque lexical cues from deceptive-review classifiers into interpretable language phenomena that are empirically grounded and more predictive than direct LLM outputs.

  12. Beyond Math: Stories as a Testbed for Memorization-Constrained Reasoning in LLMs

    cs.CL 2024-12 unverdicted novelty 6.0

    Introduces a mitigation technique that drops LLM accuracy on popular fiction character tasks from 96% to 72% by limiting verbatim memorization while retaining gist cues.

  13. Training Language Models to Self-Correct via Reinforcement Learning

    cs.LG 2024-09 unverdicted novelty 6.0

    SCoRe uses multi-turn online RL with regularization on self-generated traces to improve LLM self-correction, achieving 15.6% and 9.1% gains on MATH and HumanEval for Gemini models.

  14. LaMSUM: Amplifying Voices Against Harassment through LLM Guided Extractive Summarization of User Incident Reports

    cs.CL 2024-06 unverdicted novelty 6.0

    LaMSUM is a novel multi-level LLM framework with voting methods for extractive summarization of large incident report collections that outperforms prior extractive methods.

  15. RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback

    cs.CL 2023-09 conditional novelty 6.0

    RLAIF matches RLHF on summarization and dialogue tasks, with a direct-RLAIF variant achieving superior results by using LLM rewards directly during training.

  16. BloombergGPT: A Large Language Model for Finance

    cs.LG 2023-03 conditional novelty 6.0

    BloombergGPT is a 50B parameter LLM trained on a 708B token mixed financial and general dataset that outperforms prior models on financial benchmarks while preserving general LLM performance.

  17. A Multitask, Multilingual, Multimodal Evaluation of ChatGPT on Reasoning, Hallucination, and Interactivity

    cs.CL 2023-02 accept novelty 6.0

    ChatGPT outperforms zero-shot LLMs on most tasks and improves with interaction but scores only 63.41 percent on reasoning categories and generates extrinsic hallucinations from its training data.

  18. A Recipe for Long-Context Reasoning in Large Language Models via On-Policy Optimization and Distillation

    cs.CL 2026-05 unverdicted novelty 5.0

    Combines GRPO with teacher-guided on-policy distillation and introduces LongBlocks dataset to yield more stable long-context reasoning than either method alone.

  19. Towards Designing for Resilience: Community-Centered Deployment of an AI Business Planning Tool in a Small Business Center

    cs.HC 2026-04 unverdicted novelty 5.0

    An AI business planning tool deployed in community workshops lowers barriers to capital access but requires peer support to preserve sensemaking and build collective resilience in AI use.

  20. Recommendations for Efficient and Responsible LLM Adoption within Industrial Software Development

    cs.SE 2026-04 conditional novelty 4.0

    A multi-case study plus survey produces seven actionable recommendations for efficient and responsible LLM use in industrial software engineering.

  21. A Multi-Model Metric-based Selection Framework for Abstractive Text summarization

    cs.CL 2026-06 unverdicted novelty 3.0

    MASF runs several fine-tuned transformer models to produce candidate summaries, selects the best via lexical and semantic metrics, and reports 88.63% BERTScore on CNN/DailyMail while outperforming listed LLMs.

  22. A Multi-Model Metric-based Selection Framework for Abstractive Text summarization

    cs.CL 2026-06 conditional novelty 3.0

    Selecting among T5, PEGASUS and LED summaries by averaging ROUGE-L + BLEU + BERTScore yields 88.63 % BERTScore on CNN/DailyMail, beating the individual models and several reported LLMs.

  23. A Multi-Model Metric-based Selection Framework for Abstractive Text summarization

    cs.CL 2026-06 unverdicted novelty 3.0

    MASF runs several fine-tuned transformer models to produce candidate summaries for each article, scores them with lexical and semantic metrics, and outputs the highest-scoring one, reporting 88.63% BERTScore on CNN/Da...