REVIEW 18 cited by
Summarization is (Almost) Dead
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
How well can large language models (LLMs) generate summaries? We develop new datasets and conduct human evaluation experiments to evaluate the zero-shot generation capability of LLMs across five distinct summarization tasks. Our findings indicate a clear preference among human evaluators for LLM-generated summaries over human-written summaries and summaries generated by fine-tuned models. Specifically, LLM-generated summaries exhibit better factual consistency and fewer instances of extrinsic hallucinations. Due to the satisfactory performance of LLMs in summarization tasks (even surpassing the benchmark of reference summaries), we believe that most conventional works in the field of text summarization are no longer necessary in the era of LLMs. However, we recognize that there are still some directions worth exploring, such as the creation of novel datasets with higher quality and more reliable evaluation methods.
Forward citations
Cited by 18 Pith papers
-
VA-Blueprint: Uncovering Building Blocks for Visual Analytics System Design
A semi-automated methodology and public knowledge base catalog the building blocks of 101 urban visual analytics systems, using GPT-4 for extraction with human-in-the-loop correction and expert validation.
-
From Alerts to Intelligence: A Novel LLM-Aided Framework for Host-based Intrusion Detection
SHIELD, an LLM-aided pipeline combining a masked autoencoder, deterministic data augmentation, and multi-level prompting, detects host-based attacks with high precision on three public datasets.
-
Fair Document Valuation in LLM Summaries via Shapley Values
Cluster Shapley groups semantically similar documents via embeddings and computes cluster-level Shapley values, claiming better efficiency-accuracy trade-offs than Monte Carlo and Kernel SHAP on Amazon review summarization.
-
MIH-TCCT: Mitigating Inconsistent Hallucinations in LLMs via Event-Driven Text-Code Cyclic Training
MIH-TCCT reduces inconsistent hallucinations by cyclically training LLMs to translate event-based text into structured code and back, without task-specific fine-tuning.
-
Evaluating Small Language Models for News Summarization: Implications and Factors Influencing Performance
A benchmark of 19 small language models on 2,000 news articles shows that the best small models nearly match 70B-parameter LLMs in summary quality while producing shorter summaries.
-
Learn-by-interact: A Data-Centric Framework for Self-Adaptive Agents in Realistic Environments
A data-centric method that relabels agent trajectories with new instructions, called backward construction, improves LLM agent performance on web, code, desktop, and data-science tasks without human labeling.
-
MageBench: Bridging Large Multimodal Models to Agents
MageBench introduces a 483-scenario benchmark showing current large multimodal models are far weaker than humans at agent tasks requiring continuous visual feedback and planning.
-
MIMDE: Exploring the Use of Synthetic vs Human Data for Evaluating Multi-Insight Multi-Document Extraction Tasks
LLMs rank similarly on human and synthetic data for extracting insights, but synthetic data does not predict how well models map insights back to source documents.
-
An Empirical Study of Many-Shot In-Context Learning for Machine Translation of Low-Resource Languages
BM25-retrieved many-shot examples match much larger random sets for translating English into ten truly low-resource languages, and ICL still helps after fine-tuning.
-
Rethinking Hate Speech Detection on Social Media: Can LLMs Replace Traditional Models?
On three hate speech datasets, including a new code-mixed IndoHateMix benchmark, fine-tuned LLMs such as LLaMA-3.1 beat multilingual BERT models, with the largest gains on code-mixed Indian text.
-
A Framework for Generating Conversational Recommendation Datasets from Behavioral Interactions
ConvRecStudio generates roughly 38K synthetic multi-turn recommendation dialogs across three domains from historical interactions, and a cross-attention transformer fusing history with dialog beats dialog-only and his...
-
Summarization for Generative Relation Extraction in the Microbiome Domain
Using LLM-generated summaries as input improves instruction-tuned generative relation extraction in the low-resource microbiome domain, though fine-tuned BERT models remain more accurate.
-
Enhancing Semantic Consistency of Large Language Models through Model Editing: An Interpretability-Oriented Approach
Steering a handful of attention heads with a bias vector derived from consistent prompt pairs improves semantic consistency of LLaMA-2-7B on paraphrased NLU and NLG tasks.
-
(Towards) Scalable Reliable Automated Evaluation with Large Language Models
Multi-LLM pairwise Elo ranking with adjustable consensus thresholds produces rankings of competency profiles that average Spearman ρ≈0.83 with expert judgments.
-
A Research Vision for Web Search on Emerging Topics
The paper lays out three research questions to guide the study and redesign of web search for emerging topics, focused on user knowledge, dynamic topic awareness, and responsible opinion formation.
-
Multiple Abstraction Level Retrieve Augment Generation
MAL-RAG retrieves document, section, paragraph, and multi-sentence chunks together and claims a 25.7% improvement in AI-judged answer correctness on glycoscience questions over single-level RAG.
-
Zero-Shot Strategies for Length-Controllable Summarization
Simple zero-shot tricks, target remapping, best-of-N sampling, and iterative self-revision, substantially improve length compliance in LLaMA-3 summarization without fine-tuning.
-
Multi-LLM Text Summarization
Using multiple LLMs to generate and select summaries raises ROUGE and BLEU scores on ArXiv and GovReport compared with single-LLM chunk-and-concatenate baselines.
Discussion (0). Continue with ORCID to comment.