Instruction-tuned open LLMs beat zero-shot GPT-4 on automatic many-to-many summarization scores without lowering MMLU, but human evaluation shows instruction tuning can increase factual errors.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
An Empirical Study of Many-to-Many Summarization with Large Language Models
Instruction-tuned open LLMs beat zero-shot GPT-4 on automatic many-to-many summarization scores without lowering MMLU, but human evaluation shows instruction tuning can increase factual errors.