REVIEW 23 cited by
News Summarization and Evaluation in the Era of GPT-3
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
News Summarization and Evaluation in the Era of GPT-3
read the original abstract
The recent success of prompting large language models like GPT-3 has led to a paradigm shift in NLP research. In this paper, we study its impact on text summarization, focusing on the classic benchmark domain of news summarization. First, we investigate how GPT-3 compares against fine-tuned models trained on large summarization datasets. We show that not only do humans overwhelmingly prefer GPT-3 summaries, prompted using only a task description, but these also do not suffer from common dataset-specific issues such as poor factuality. Next, we study what this means for evaluation, particularly the role of gold standard test sets. Our experiments show that both reference-based and reference-free automatic metrics cannot reliably evaluate GPT-3 summaries. Finally, we evaluate models on a setting beyond generic summarization, specifically keyword-based summarization, and show how dominant fine-tuning approaches compare to prompting. To support further research, we release: (a) a corpus of 10K generated summaries from fine-tuned and prompt-based models across 4 standard summarization benchmarks, (b) 1K human preference judgments comparing different systems for generic- and keyword-based summarization.
Forward citations
Cited by 23 Pith papers
-
MemMorph: Tool Hijacking in LLM Agents via Memory Poisoning
MemMorph poisons LLM agent long-term memory with three crafted records disguised as facts or policies to hijack tool selection, reaching 85.9% success rate across 10 backbones and outperforming baselines while resisti...
-
Faithful by Construction: Claim-Anchored Attribution for Multi-Document Summarization
CAMS framework extracts, clusters, selects, and rewrites atomic claims to produce multi-document summaries with fine-grained, multi-source traceability by construction.
-
From Text to Voice: A Reproducible and Verifiable Framework for Evaluating Tool Calling LLM Agents
A dataset-agnostic framework converts text tool-calling benchmarks to paired audio versions via TTS and noise, showing model-dependent performance with small text-to-voice gaps of 1.8-4.8 points on Confetti and When2Call.
-
A Recipe for Long-Context Reasoning in Large Language Models via On-Policy Optimization and Distillation
dGRPO merges outcome-based policy optimization with dense teacher guidance from on-policy distillation, yielding more stable long-context reasoning on the new LongBlocks synthetic dataset.
-
Faithful by Construction: Claim-Anchored Attribution for Multi-Document Summarization
CAMS is a claim-anchored Extract-Select-Rewrite framework for multi-document summarization that anchors every summary sentence to support-checked atomic claims with multi-source traceability.
-
M\"OVE: A Holistic LLM Benchmark for the German Public Sector
MÖVE presents a new German-language benchmark evaluating 39 LLMs on performance and governance criteria using ten public-administration datasets.
-
Understanding LLM Behavior in Multi-Target Cross-Lingual Summarization
Introduces the MEA benchmark for multi-target cross-lingual summarization across 24 languages and demonstrates that activation steering from English summarization representations improves performance.
-
From Text to Voice: A Reproducible and Verifiable Framework for Evaluating Tool Calling LLM Agents
A dataset-agnostic framework converts text tool-calling benchmarks to paired audio evaluations via TTS, speaker variation and noise, then evaluates seven omni-modal models showing model- and task-dependent performance...
-
A Meta Reinforcement Learning Approach to Goals-Based Wealth Management
MetaRL pre-trained on GBWM problems delivers near-optimal dynamic strategies in 0.01s achieving 97.8% of DP optimal utility and handles larger problems where DP fails.
-
LLM-ReSum: A Framework for LLM Reflective Summarization through Self-Evaluation
LLM-ReSum uses LLM self-evaluation in a closed feedback loop to refine summaries, improving factual accuracy by up to 33% and coverage by 39% with 89% human preference.
-
Why is "Chicago" Predictive of Deceptive Reviews? Using LLMs to Discover Language Phenomena from Lexical Cues
A conjecture-then-validate method lets LLMs convert opaque lexical cues from deceptive-review classifiers into interpretable language phenomena that are empirically grounded and more predictive than direct LLM outputs.
-
Beyond Math: Stories as a Testbed for Memorization-Constrained Reasoning in LLMs
Introduces a mitigation technique that drops LLM accuracy on popular fiction character tasks from 96% to 72% by limiting verbatim memorization while retaining gist cues.
-
Training Language Models to Self-Correct via Reinforcement Learning
SCoRe uses multi-turn online RL with regularization on self-generated traces to improve LLM self-correction, achieving 15.6% and 9.1% gains on MATH and HumanEval for Gemini models.
-
LaMSUM: Amplifying Voices Against Harassment through LLM Guided Extractive Summarization of User Incident Reports
LaMSUM is a novel multi-level LLM framework with voting methods for extractive summarization of large incident report collections that outperforms prior extractive methods.
-
RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback
RLAIF matches RLHF on summarization and dialogue tasks, with a direct-RLAIF variant achieving superior results by using LLM rewards directly during training.
-
BloombergGPT: A Large Language Model for Finance
BloombergGPT is a 50B parameter LLM trained on a 708B token mixed financial and general dataset that outperforms prior models on financial benchmarks while preserving general LLM performance.
-
A Multitask, Multilingual, Multimodal Evaluation of ChatGPT on Reasoning, Hallucination, and Interactivity
ChatGPT outperforms zero-shot LLMs on most tasks and improves with interaction but scores only 63.41 percent on reasoning categories and generates extrinsic hallucinations from its training data.
-
A Recipe for Long-Context Reasoning in Large Language Models via On-Policy Optimization and Distillation
Combines GRPO with teacher-guided on-policy distillation and introduces LongBlocks dataset to yield more stable long-context reasoning than either method alone.
-
Towards Designing for Resilience: Community-Centered Deployment of an AI Business Planning Tool in a Small Business Center
An AI business planning tool deployed in community workshops lowers barriers to capital access but requires peer support to preserve sensemaking and build collective resilience in AI use.
-
Recommendations for Efficient and Responsible LLM Adoption within Industrial Software Development
A multi-case study plus survey produces seven actionable recommendations for efficient and responsible LLM use in industrial software engineering.
-
A Multi-Model Metric-based Selection Framework for Abstractive Text summarization
MASF runs several fine-tuned transformer models to produce candidate summaries, selects the best via lexical and semantic metrics, and reports 88.63% BERTScore on CNN/DailyMail while outperforming listed LLMs.
-
A Multi-Model Metric-based Selection Framework for Abstractive Text summarization
Selecting among T5, PEGASUS and LED summaries by averaging ROUGE-L + BLEU + BERTScore yields 88.63 % BERTScore on CNN/DailyMail, beating the individual models and several reported LLMs.
-
A Multi-Model Metric-based Selection Framework for Abstractive Text summarization
MASF runs several fine-tuned transformer models to produce candidate summaries for each article, scores them with lexical and semantic metrics, and outputs the highest-scoring one, reporting 88.63% BERTScore on CNN/Da...
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.