A meta-analysis of keyphrase generation shows benchmark datasets are highly correlated, evaluation protocols inflate scores, and a released BART-large baseline provides a stronger reference point.
ChatGPT vs State-of-the-Art Models: A Benchmarking Study in Keyphrase Generation Task
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
Transformer-based language models, including ChatGPT, have demonstrated exceptional performance in various natural language generation tasks. However, there has been limited research evaluating ChatGPT's keyphrase generation ability, which involves identifying informative phrases that accurately reflect a document's content. This study seeks to address this gap by comparing ChatGPT's keyphrase generation performance with state-of-the-art models, while also testing its potential as a solution for two significant challenges in the field: domain adaptation and keyphrase generation from long documents. We conducted experiments on six publicly available datasets from scientific articles and news domains, analyzing performance on both short and long documents. Our results show that ChatGPT outperforms current state-of-the-art models in all tested datasets and environments, generating high-quality keyphrases that adapt well to diverse domains and document lengths.
fields
cs.IR 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
An Analysis of Datasets, Metrics and Models in Keyphrase Generation
A meta-analysis of keyphrase generation shows benchmark datasets are highly correlated, evaluation protocols inflate scores, and a released BART-large baseline provides a stronger reference point.