Pith. sign in

REVIEW 3 major objections 6 minor 101 references

An Analysis of Datasets, Metrics and Models in Keyphrase Generation

T0 review · 3 major / 6 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read Keyphrase generation benchmarks are so similar that evaluating on more than one adds no value, and inconsistent metric computation inflates reported performance.

desk verdict A useful, candid meta-analysis that gives the field a needed look in the mirror; the dataset-redundancy claim is plausible but over-strong given the correlation method. read the letter →

arxiv 2506.10346 v1 pith:LQALZVZA submitted 2025-06-12 cs.IR cs.CL

classification cs.IRcs.CL
keywords keyphrasegenerationbenchmarkdatasetsdatasetredundancyevaluationmetricsreproducibilitypre-trainedlanguagemodelsKP20kF1@k
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper claims that the field's standard evaluation practice for keyphrase generation is broken in two specific ways. The five most-used benchmark datasets (KP20k, SemEval-2010, Inspec, Krapivin, NUS) are so similar both in content and in the model rankings they produce that reporting results on more than one of them adds no information. At the same time, inconsistent metric computation, above all a ground-truth normalization recipe inherited from the KP20k work and a convention of padding short outputs with dummy phrases, inflates reported scores and makes cross-paper comparisons unreliable. If these findings are right, the community should consolidate benchmarks, standardize metric calculation, and treat many published performance gaps as artifacts rather than real progress. The paper also releases a fine-tuned BART-large baseline trained without normalization, intended as a solid reference point for future work.

What carries the argument

Three instruments carry the analysis. The first is a manually assembled collection of 826 score triples from 50 models across 26 datasets, recording each model's best reported score per dataset and metric along with contribution type, architecture, significance-testing practice, and code and weight availability. The second is a correlation matrix computed over those best scores across the five dominant datasets; correlations above 0.9 are the quantitative basis for the dataset-redundancy claim. The third is a replication study that recomputes F1 for three published models (catSeqTG-2RF1, ExHiRD-h, SetTrans) under a fixed protocol, toggling the normalization on and off, which isolates the metric-inflation effect. The released object is a BART-large baseline fine-tuned on KP20k in the ONE2MANY format, in which keyphrases are generated as a single delimiter-separated sequence; it is trained without preprocessing, selected by validation F1, and evaluated with greedy and beam-search decoding.

What would settle it

Recompute the five-dataset correlation matrix under one standardized evaluation pipeline, with identical normalization, identical F1@k padding rules, identical present/absent matching, and the same released model outputs on all five datasets, and check whether the correlations remain above 0.9.

Watch

Extended reading notes

Core claim

Analyzing 52 papers published after the first neural keyphrase generation work, the paper establishes three claims. First, model scores on the five dominant benchmark datasets are almost perfectly correlated (Pearson $\rho > 0.9$, p-value < 0.01 for the pairwise correlations), and several of these datasets share source documents from the same digital library, so using more than one of them contributes no additional signal. Second, two inconsistencies inflate results: some authors compute F1@k after padding short predictions with dummy phrases, and at least 60% of papers apply a normalization that strips abbreviations, tokenizes on non-letter characters, and replaces digits with a placeholder before scoring. The paper's replication experiments show this normalization adds up to 3.5 F1@5 points for present keyphrases. Third, although model architectures have moved from RNNs to Transformers to fine-tuned pre-trained language models, absolute progress is limited: only about 3.1 F1@M points separate the best 2019 and best 2024 models on present keyphrases, and absent keyphrase F1@M hovers near 11%. In response, the paper releases a BART-large model fine-tuned on KP20k in the ONE2MANY format without normalization, which beats most prior models and reaches state-of-the-art absent keyphrase F1@5.

Load-bearing premise

The redundancy claim assumes that correlations above 0.9 among best reported model scores reflect genuine dataset similarity, and would collapse if those correlations mostly come from shared evaluation quirks, small sample sizes, or the general tendency of better models to score higher on every benchmark.

Editorial extensions

If this is right

  • The five standard scientific-abstract datasets can be treated as one evaluation signal; papers should pair one of them with a different-domain set such as KPTimes instead of reporting all five.
  • Results computed with ground-truth normalization are inflated relative to unnormalized results, so standardizing evaluation will lower some reported scores and can reorder model comparisons.
  • State-of-the-art keyphrase generation is now dominated by fine-tuned pre-trained language models, yet the net gain over early sequence-to-sequence models is small, around 3.1 F1@M points for present keyphrases from 2019 to 2024.
  • Absent keyphrase generation remains far behind present keyphrase generation, with best F1@M near 11%, and strict single-ground-truth matching makes those scores especially unreliable.
  • Only 8 of the 50 model papers release model weights; without weights, statistical significance testing and fair comparison remain rare, so releasing weights is a necessary condition for credible progress claims.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A direct extension of the redundancy claim is that new datasets should be chosen for domain diversity rather than size, because one scientific-abstract benchmark plus one news or social media benchmark would carry more signal than five overlapping abstract sets.
  • A testable extension of the metric-inflation argument is that retroactively applying a single standardized evaluation protocol to released model outputs would reorder published leaderboards, since normalization and padding affect different models by different amounts.
  • The paper's correlation table shows KPTimes and DUC2001 behaving markedly differently from the five scientific datasets, suggesting that cross-domain evaluation is where new benchmarks can still add information.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. The paper presents a meta-analysis of 52 keyphrase-generation papers published after Meng et al. (2017). The authors manually extracted 826 best-score triples across 50 models, analyzed contribution types, benchmark-dataset usage, evaluation metrics, and model architectures, and released a fine-tuned BART-large baseline along with an evaluation framework. The three headline findings are that the five commonly used benchmark datasets are so correlated that reporting results on more than one adds no value, that evaluation-protocol inconsistencies such as keyphrase normalization overestimate model performance, and that despite a shift to PLM-based models, overall progress has been modest. The paper also includes replication experiments measuring the effect of the Meng et al. (2017) normalization procedure on three published models.

Significance. If the findings were fully established, they would give the field a clear, actionable message: consolidate redundant benchmarks, standardize metric computation, and treat many published performance gaps with caution. The paper provides a useful service by making the collected metadata, replication scripts, and model weights publicly available, and the replication experiments are a concrete step beyond a standard survey. The main significance currently rests on two empirical claims that need stronger support: the dataset-redundancy claim in Section 3.2 and the overestimation claim in Section 3.3. I do not see a circularity problem: the analysis compares externally reported results and evaluates its released baseline under a standard benchmark, rather than deriving its conclusions from assumptions secretly built into the data.

major comments (3)
  1. [Section 3.2, Figure 2] The dataset-redundancy claim is not established by the reported correlation analysis. The figure is computed from "best scores" on each dataset, but the paper does not specify which of the 42 extracted metrics was used to obtain one score per model and dataset, how many models have scores on each dataset pair, or whether the compared scores were produced under the same evaluation protocol. This matters because, under the paper's own Section 3.3, normalization changes F1 by 2.2 to 3.5 absolute points and two incompatible F1@k padding rules coexist. A paper that uses one protocol for all five datasets will have its scores shifted on all five datasets, inflating cross-dataset Pearson correlations even if the datasets share no documents; a shared latent "model quality" factor will do the same. The statement that "reporting results on more than one adds no value" therefore needs a matched-protocol analysis, such as restricting correlations to models evaluated with identical normalization and F1@k conventions, reporting pairwise sample sizes, and correcting for the 28 pairwise comparisons shown in Figure 2.
  2. [Section 3.3, Figure 4] The replication evidence as printed is internally inconsistent and should be repaired before it can support the overestimation claim. The figure appears to show that "Ours w/o norm" exceeds "Ours" for present F1@M on ExHiRD and SetTrans, and exceeds "Ours" for all three models on absent F1@5 (e.g., catSeqTG absent F1@5: 5.8 vs 1.6 in the printed bars). The text says normalization "significantly increases the scores for the majority of the evaluation metrics" and gives +2.2 (F1@M) and +3.5 (F1@5) for present keyphrases, but these numbers are not derived from the figure in an obvious way. The authors should report the per-model, per-metric values in a table, define which pairwise comparison yields the claimed averages, and explain the sign of the effect for absent keyphrases; otherwise the central claim that normalization overestimates performance is not quantitatively supported.
  3. [Section 3.5, Figure 6] The state-of-the-art-over-time plot and the statement that only 3.1% present F1@M separates Chan et al. (2019) from Thomas and Vajjala (2024a) inherit the protocol-mixing problem identified in Sections 3.2 and 3.3. The plotted scores are the best scores extracted from 50 papers under different normalization rules, different F1@k padding conventions, and different present/absent matching methods; Section 3.3 shows these choices alone move scores by several points. Restricting the SOTA lines to a single protocol, or clearly labeling them as illustrative rather than directly comparable, is necessary for the claim about limited overall progress.
minor comments (6)
  1. [Section 2] The selection section says the sample includes papers "published at major NLP venues in the last seven years," but the same paragraph states that AAAI, SIGIR, and CIKM papers are included; the sentence should be reworded to acknowledge that non-ACL venues are also represented.
  2. [Section 3.2, Figure 2 caption] The caption for Figure 2 should state the exact metric used for the correlation, the number of models per pair, and the method used to aggregate multiple reported metrics per dataset.
  3. [Section 3.3] Footnote 3 notes that normalization information is often located only in source code; the sentence in the main text says normalization was observed "in at least 30 out of 50 papers" but does not say how much uncertainty remains or how the count was determined; a short note would make the statistic reproducible.
  4. [Section 4] There is a typo "ONE2M ANY" in the baseline description; elsewhere the paper writes "ONE2MANY" and "ONE2SET"; these should be made consistent.
  5. [Appendix A.3] The appendix says keyphrases are lowercased, stemmed, and deduplicated before score calculation, but the main analysis does not state whether the 826 extracted triples were all normalized in this way; adding this detail would help readers interpret the correlation matrix.
  6. [References] The Jiang et al. (2023a) and Jiang et al. (2023b) references appear to point to the same paper; this duplicate should be resolved, and the in-text citations should match the reference list.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the redundancy and metric-inconsistency findings are empirical meta-analyses with independent replication, and the baseline is evaluated on a held-out test set.

full rationale

The paper's central claims are meta-analytic rather than derived from a fitted model. The dataset-redundancy claim is inferred from Pearson correlations over best-reported scores extracted from 52 external papers; the high correlations are an empirical observation about how models rank across datasets, not an identity or a definitional consequence. The evaluation-protocol critique is supported by independent replication experiments in Section 3.3, where the authors recompute F1 scores for three external models under a stated protocol (dummy padding, stricter subsequence matching), so the discrepancy between original and replicated scores is measured rather than assumed. The released baseline in Section 4 is fine-tuned on KP20k, checkpoint-selected on the KP20k validation set using F1@{M,5,10}, and then evaluated on the KP20k test set; no fitted quantity is renamed as a prediction and no held-out result is constructed from the training objective. Self-citations appear (e.g., Boudin et al. 2020; Boudin and Gallina 2021; Boudin and Aizawa 2024), but they support peripheral suggestions or the authors' own released resources, not the load-bearing redundancy or overestimation findings. The weakest point, that protocol differences and a latent model-quality factor could inflate cross-dataset correlations, is a threat to validity rather than a circular derivation: the conclusion would be false if those confounds dominated, so it is not forced by construction. No equation is reused as its own output, and no fitted parameter is relabeled as a finding. Therefore no circularity step meets the evidence bar.

Assumptions & free parameters 4 free parameters · 5 assumptions · 0 invented entities

The main claims are empirical meta-analytic claims, not derivations, so the ledger contains few theoretical free parameters. The released baseline's hyperparameters are the main hand-chosen numbers, and the strongest interpretive assumption is the correlation-to-redundancy inference, which is the load-bearing premise.

free parameters (4)
  • Number of fine-tuning epochs for released baseline = 9
    The checkpoint with the highest validation F1 on KP20k is selected, so the reported baseline scores depend on this model-selection choice; it is a standard hyperparameter, not a theoretical constant.
  • Learning rate for baseline fine-tuning = 1e-5
    Hand-chosen optimization hyperparameter for the released baseline; it affects the model scores but not the survey's dataset and metric findings.
  • Batch size for baseline fine-tuning = 4
    Hand-chosen training hyperparameter dictated by GPU memory on two RTX 2080 cards; it affects the reported baseline scores.
  • Beam search width for baseline inference = 20
    Hand-chosen decoding parameter used to assemble top-k keyphrases at test time; it determines the final reported scores.
assumptions (5)
  • domain assumption The 52-paper sample restricted to ACL Anthology plus six non-ACL papers is representative of keyphrase generation research.
    Used to generalize findings across the field; the authors acknowledge preprints and non-ACL journals are excluded (Section 2, Limitations).
  • domain assumption Manually extracted best reported scores (826 triples) faithfully represent model performance.
    The correlation matrix and SOTA-over-time lines depend on these extractions; the authors note typos, ambiguities, and a disambiguation rule that may produce suboptimal scores for non-KP20k datasets (Limitations).
  • ad hoc to paper High Pearson correlation between best scores on two datasets implies the datasets are redundant for evaluation.
    This inference drives finding 1 in the Introduction and Section 3.2; it is an interpretive leap not supported by a formal model of evaluation redundancy.
  • domain assumption Standard evaluation preprocessing (lowercasing, stemming, duplicate removal, dummy-phrase padding) is the correct basis for comparing models.
    Adopted for the replication experiments and baseline evaluation (Appendix A.3); while standard, it is a contested choice in the same literature.
  • standard math Pearson correlation and Student's paired t-test are valid for the compared model score distributions.
    Underlying assumptions of linearity and paired observations are unstated (Section 3.2 and Figure 4).

how reviews work

0 comments
Cite this review

Pith. "Pith review of An Analysis of Datasets, Metrics and Models in Keyphrase Generation." pith.science (2026). https://pith.science/paper/LQALZVZA

@misc{pith2026250610346,
  author       = {Pith},
  title        = {Pith review of: An Analysis of Datasets, Metrics and Models in Keyphrase Generation},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/LQALZVZA}},
  note         = {Machine review of arXiv:2506.10346}
}
read the original abstract

Keyphrase generation refers to the task of producing a set of words or phrases that summarises the content of a document. Continuous efforts have been dedicated to this task over the past few years, spreading across multiple lines of research, such as model architectures, data resources, and use-case scenarios. Yet, the current state of keyphrase generation remains unknown as there has been no attempt to review and analyse previous work. In this paper, we bridge this gap by presenting an analysis of over 50 research papers on keyphrase generation, offering a comprehensive overview of recent progress, limitations, and open challenges. Our findings highlight several critical issues in current evaluation practices, such as the concerning similarity among commonly-used benchmark datasets and inconsistencies in metric calculations leading to overestimated performances. Additionally, we address the limited availability of pre-trained models by releasing a strong PLM-based model for keyphrase generation as an effort to facilitate future research.

Figures

Figures reproduced from arXiv: 2506.10346 by the authors.

Figure 1
Figure 1. Number of papers utilizing each dataset. [PITH_FULL_IMAGE:figures/full_fig_p003_1.png] view at source ↗
Figure 2
Figure 2. Pearson’s correlation coefficient ρ computed between the model scores across datasets [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗
Figure 3
Figure 3. Number of papers employing each evaluation [PITH_FULL_IMAGE:figures/full_fig_p004_3.png] view at source ↗
Figures from the paper (5 more)
Figure 4
Figure 4. Figure 4: Replicated evaluation results on the KP20k [PITH_FULL_IMAGE:figures/full_fig_p004_4.png]
Figure 5
Figure 5. Figure 5: Evolutionary tree of the keyphrase generation [PITH_FULL_IMAGE:figures/full_fig_p005_5.png]
Figure 6
Figure 6. Figure 6: Best scores achieved by each model in terms of [PITH_FULL_IMAGE:figures/full_fig_p006_6.png]
Figure 7
Figure 7. Figure 7: Architectures of the proposed keyphrase gen [PITH_FULL_IMAGE:figures/full_fig_p006_7.png]
Figure 8
Figure 8. Figure 8: Performance of our baseline model on the [PITH_FULL_IMAGE:figures/full_fig_p007_8.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

101 extracted references · 55 canonical work pages

  1. [1]

    URL: " 'urlintro :=

    ENTRY address author booktitle chapter edition editor howpublished institution journal key month note number organization pages publisher school series title type volume year eprint doi pubmed url lastchecked label extra.label sort.label short.list INTEGERS output.state before.all mid.sentence after.sentence after.block STRINGS urlintro eprinturl eprintpr...

  2. [2]

    write newline

    " write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION word.in bbl.in capitalize " " * FUNCT...

  3. [3]

    Wasi Ahmad, Xiao Bai, Soomin Lee, and Kai-Wei Chang. 2021. https://doi.org/10.18653/v1/2021.acl-long.111 Select, extract and generate: Neural keyphrase generation with layer-wise coverage attention . In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Pr...

  4. [4]

    Mohammad Arvan, Lu \' s Pina, and Natalie Parde. 2022. https://doi.org/10.18653/v1/2022.emnlp-main.150 Reproducibility in computational linguistics: Is source code enough? In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 2350--2361, Abu Dhabi, United Arab Emirates. Association for Computational Linguistics

  5. [5]

    Hareesh Bahuleyan and Layla El Asri. 2020. https://doi.org/10.18653/v1/2020.coling-main.462 Diverse keyphrase generation with neural unlikelihood training . In Proceedings of the 28th International Conference on Computational Linguistics, pages 5271--5287, Barcelona, Spain (Online). International Committee on Computational Linguistics

  6. [6]

    Xiao Bai, Xue Wu, Ivan Stojkovic, and Kostas Tsioutsiouliklis. 2024. https://doi.org/10.1145/3627673.3680093 Leveraging large language models for improving keyphrase generation for contextual targeting . In Proceedings of the 33rd ACM International Conference on Information and Knowledge Management, CIKM '24, page 4349–4357, New York, NY, USA. Association...

  7. [7]

    Florian Boudin and Akiko Aizawa. 2024. https://doi.org/10.18653/v1/2024.findings-emnlp.33 Unsupervised domain adaptation for keyphrase generation using citation contexts . In Findings of the Association for Computational Linguistics: EMNLP 2024, pages 598--614, Miami, Florida, USA. Association for Computational Linguistics

  8. [8]

    Florian Boudin and Ygor Gallina. 2021. https://doi.org/10.18653/v1/2021.naacl-main.330 Redefining absent keyphrases and their effect on retrieval effectiveness . In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 4185--4193, Online. Association for Comput...

Show all 101 references
  1. [9]

    Florian Boudin, Ygor Gallina, and Akiko Aizawa. 2020. https://doi.org/10.18653/v1/2020.acl-main.105 Keyphrase generation for scientific document retrieval . In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 1118--1126, Online. As...

  2. [10]

    Erion C ano and Ond r ej Bojar. 2019. https://doi.org/10.18653/v1/N19-1070 Keyphrase generation: A text summarization struggle . In Proceedings of the 2019 Conference of the North A merican Chapter of the Association for Computational Linguistics: Human Language Technologies, ...

  3. [11]

    Hou Pong Chan, Wang Chen, Lu Wang, and Irwin King. 2019. https://doi.org/10.18653/v1/P19-1208 Neural keyphrase generation via reinforcement learning with adaptive rewards . In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 2163--...

  4. [12]

    Jun Chen, Xiaoming Zhang, Yu Wu, Zhao Yan, and Zhoujun Li. 2018. https://doi.org/10.18653/v1/D18-1439 Keyphrase generation with correlation constraints . In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 4057--4066, Brussels, Belg...

  5. [13]

    Wang Chen, Hou Pong Chan, Piji Li, Lidong Bing, and Irwin King. 2019 a . https://doi.org/10.18653/v1/N19-1292 An integrated approach for keyphrase generation via exploring the power of retrieval and extraction . In Proceedings of the 2019 Conference of the North A merican Chap...

  6. [14]

    Wang Chen, Hou Pong Chan, Piji Li, and Irwin King. 2020. https://doi.org/10.18653/v1/2020.acl-main.103 Exclusive hierarchical decoding for deep keyphrase generation . In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 1095--1105, ...

  7. [15]

    Wang Chen, Yifan Gao, Jiani Zhang, Irwin King, and Michael R. Lyu. 2019 b . https://doi.org/10.1609/aaai.v33i01.33016268 Title-guided encoding for keyphrase generation . Proceedings of the AAAI Conference on Artificial Intelligence, 33(01):6268--6275

  8. [16]

    Qikai Cheng, Jiamin Wang, Wei Lu, Yong Huang, and Yi Bu. 2020. https://doi.org/10.1007/s11192-020-03576-5 Keyword-citation-keyword network: a new perspective of discipline knowledge structure analysis . Scientometrics, 124(3):1923–1943

  9. [17]

    Chi, Michelle Gumbrecht, and Lichan Hong

    Ed H. Chi, Michelle Gumbrecht, and Lichan Hong. 2007. Visual foraging of highlighted text: an eye-tracking study. In Proceedings of the 12th International Conference on Human-Computer Interaction: Intelligent Multimodal Interaction Environments, HCI'07, page 589–598, Berlin, H...

  10. [18]

    Cheng-Han Chiang and Hung-yi Lee. 2023. https://doi.org/10.18653/v1/2023.findings-emnlp.599 A closer look into using large language models for automatic evaluation . In Findings of the Association for Computational Linguistics: EMNLP 2023, pages 8928--8942, Singapore. Associat...

  11. [19]

    Kyunghyun Cho, Bart van Merri \"e nboer, Caglar Gulcehre, Dzmitry Bahdanau, Fethi Bougares, Holger Schwenk, and Yoshua Bengio. 2014. https://doi.org/10.3115/v1/D14-1179 Learning phrase representations using RNN encoder -- decoder for statistical machine translation . In Procee...

  12. [20]

    Minseok Choi, Chaeheon Gwak, Seho Kim, Si Kim, and Jaegul Choo. 2023. https://doi.org/10.18653/v1/2023.findings-emnlp.199 S im CKP : Simple contrastive learning of keyphrase representations . In Findings of the Association for Computational Linguistics: EMNLP 2023, pages 3003-...

  13. [21]

    Lam Do, Pritom Saha Akash, and Kevin Chen-Chuan Chang. 2023. https://doi.org/10.18653/v1/2023.acl-long.592 Unsupervised open-domain keyphrase generation . In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages ...

  14. [22]

    Rotem Dror, Gili Baumer, Segev Shlomov, and Roi Reichart. 2018. https://doi.org/10.18653/v1/P18-1128 The hitchhiker ' s guide to testing statistical significance in natural language processing . In Proceedings of the 56th Annual Meeting of the Association for Computational Lin...

  15. [23]

    J. Fagan. 1987. https://doi.org/10.1145/42005.42016 Automatic phrase indexing for document retrieval . In Proceedings of the 10th Annual International ACM SIGIR Conference on Research and Development in Information Retrieval, SIGIR '87, page 91–101, New York, NY, USA. Associat...

  16. [24]

    Nazanin Firoozeh, Adeline Nazarenko, Fabrice Alizon, and Béatrice Daille. 2020. https://doi.org/10.1017/S1351324919000457 Keyword extraction: Issues and methods . Natural Language Engineering, 26(3):259–291

  17. [25]

    Ygor Gallina, Florian Boudin, and Beatrice Daille. 2019. https://doi.org/10.18653/v1/W19-8617 KPT imes: A large-scale dataset for keyphrase generation on news documents . In Proceedings of the 12th International Conference on Natural Language Generation, pages 130--135, Tokyo,...

  18. [26]

    Yifan Gao, Qingyu Yin, Zheng Li, Rui Meng, Tong Zhao, Bing Yin, Irwin King, and Michael Lyu. 2022. https://doi.org/10.18653/v1/2022.findings-naacl.92 Retrieval-augmented multilingual keyphrase generation with retriever-generator iterative training . In Findings of the Associat...

  19. [27]

    Krishna Garg, Jishnu Ray Chowdhury, and Cornelia Caragea. 2022. https://doi.org/10.18653/v1/2022.findings-emnlp.427 Keyphrase generation beyond the boundaries of title and abstract . In Findings of the Association for Computational Linguistics: EMNLP 2022, pages 5809--5821, Ab...

  20. [28]

    Krishna Garg, Jishnu Ray Chowdhury, and Cornelia Caragea. 2023. https://doi.org/10.18653/v1/2023.findings-acl.534 Data augmentation for low-resource keyphrase generation . In Findings of the Association for Computational Linguistics: ACL 2023, pages 8442--8455, Toronto, Canada...

  21. [29]

    Aaron Grattafiori, Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Alex Vaughan, Amy Yang, Angela Fan, Anirudh Goyal, Anthony Hartshorn, Aobo Yang, Archi Mitra, Archie Sravankumar, Artem Korenev, Art...

  22. [30]

    Jiatao Gu, Zhengdong Lu, Hang Li, and Victor O.K. Li. 2016. https://doi.org/10.18653/v1/P16-1154 Incorporating copying mechanism in sequence-to-sequence learning . In Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers...

  23. [31]

    Carl Gutwin, Gordon Paynter, Ian Witten, Craig Nevill-Manning, and Eibe Frank. 1999. https://doi.org/https://doi.org/10.1016/S0167-9236(99)00038-X Improving browsing in digital libraries with keyphrase indexes . Decision Support Systems, 27(1):81--104

  24. [32]

    Kazi Saidul Hasan and Vincent Ng. 2014. https://doi.org/10.3115/v1/P14-1119 Automatic keyphrase extraction: A survey of the state of the art . In Proceedings of the 52nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1262--1273, ...

  25. [33]

    Ma \"e l Houbre, Florian Boudin, and Beatrice Daille. 2022. https://doi.org/10.18653/v1/2022.louhi-1.6 A large-scale dataset for biomedical keyphrase generation . In Proceedings of the 13th International Workshop on Health Text Mining and Information Analysis (LOUHI), pages 47...

  26. [34]

    Kai Hu, Qing Luo, Kunlun Qi, Siluo Yang, Jin Mao, Xiaokang Fu, Jie Zheng, Huayi Wu, Ya Guo, and Qibing Zhu. 2019. https://doi.org/https://doi.org/10.1016/j.ipm.2019.02.014 Understanding the topic evolution of scientific literatures like an evolving city: Using google word2vec ...

  27. [35]

    Xiaoli Huang, Tongge Xu, Lvan Jiao, Yueran Zu, and Youmin Zhang. 2021. https://doi.org/10.1609/aaai.v35i14.17546 Adaptive beam search decoding for discrete keyphrase generation . Proceedings of the AAAI Conference on Artificial Intelligence, 35(14):13082--13089

  28. [36]

    Anette Hulth. 2003. https://aclanthology.org/W03-1028 Improved automatic keyword extraction given more linguistic knowledge . In Proceedings of the 2003 Conference on Empirical Methods in Natural Language Processing, pages 216--223

  29. [38]

    Yi Jiang, Rui Meng, Yong Huang, Wei Lu, and Jiawei Liu. 2023 b . https://doi.org/https://doi.org/10.1002/asi.24749 Generating keyphrases for readers: A controllable keyphrase generation framework . Journal of the Association for Information Science and Technology, 74(7):759--774

  30. [39]

    Staveley

    Steve Jones and Mark S. Staveley. 1999. https://doi.org/10.1145/312624.312671 Phrasier: A system for interactive document retrieval using keyphrases . In Proceedings of the 22nd Annual International ACM SIGIR Conference on Research and Development in Information Retrieval, SIG...

  31. [40]

    Byungha Kang and Youhyun Shin. 2024. https://aclanthology.org/2024.lrec-main.775 Improving low-resource keyphrase generation through unsupervised title phrase generation . In Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resource...

  32. [41]

    Jihyuk Kim, Myeongho Jeong, Seungtaek Choi, and Seung-won Hwang. 2021. https://doi.org/10.18653/v1/2021.emnlp-main.209 Structure-augmented keyphrase generation . In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 2657--2667, Online...

  33. [42]

    Su Nam Kim, Olena Medelyan, Min-Yen Kan, and Timothy Baldwin. 2010. https://aclanthology.org/S10-1004 S em E val-2010 task 5 : Automatic keyphrase extraction from scientific articles . In Proceedings of the 5th International Workshop on Semantic Evaluation, pages 21--26, Uppsa...

  34. [43]

    Fajri Koto, Timothy Baldwin, and Jey Han Lau. 2022. https://aclanthology.org/2022.coling-1.303 L ip K ey: A large-scale news dataset for absent keyphrases generation and abstractive summarization . In Proceedings of the 29th International Conference on Computational Linguistic...

  35. [44]

    Mikalai Krapivin, Aliaksandr Autaeu, Maurizio Marchese, et al. 2009. Large dataset for keyphrases extraction. Technical report, University of Trento-Dipartimento di Ingegneria e Scienza dell'Informazione

  36. [45]

    Mayank Kulkarni, Debanjan Mahata, Ravneet Arora, and Rajarshi Bhowmik. 2022. https://doi.org/10.18653/v1/2022.findings-naacl.67 Learning rich representation of keyphrases from text . In Findings of the Association for Computational Linguistics: NAACL 2022, pages 891--906, Seat...

  37. [46]

    Giuseppe Lancioni, Saida S.Mohamed, Beatrice Portelli, Giuseppe Serra, and Carlo Tasso. 2020. https://doi.org/10.18653/v1/2020.sustainlp-1.12 Keyphrase generation with GAN s in low-resources scenarios . In Proceedings of SustaiNLP: Workshop on Simple and Efficient Natural Lang...

  38. [47]

    Hwanhee Lee, Seunghyun Yoon, Franck Dernoncourt, Doo Soon Kim, Trung Bui, Joongbo Shin, and Kyomin Jung. 2021. https://doi.org/10.18653/v1/2021.naacl-main.170 KPQA : A metric for generative question answering using keyphrase weights . In Proceedings of the 2021 Conference of t...

  39. [48]

    Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Veselin Stoyanov, and Luke Zettlemoyer. 2020. https://doi.org/10.18653/v1/2020.acl-main.703 BART : Denoising sequence-to-sequence pre-training for natural language generation, translatio...

  40. [49]

    Jiahao Liu, Qifan Wang, Jingang Wang, and Xunliang Cai. 2024. https://doi.org/10.18653/v1/2024.findings-acl.179 Speculative decoding via early-exiting for faster LLM inference with T hompson sampling control mechanism . In Findings of the Association for Computational Linguist...

  41. [50]

    Yizhu Liu, Qi Jia, and Kenny Zhu. 2021. https://doi.org/10.1145/3442381.3449906 Keyword-aware abstractive summarization by extracting set-level intermediate summaries . In Proceedings of the Web Conference 2021, WWW '21, page 3042–3054, New York, NY, USA. Association for Compu...

  42. [51]

    Zhiyuan Liu, Xinxiong Chen, Yabin Zheng, and Maosong Sun. 2011. https://aclanthology.org/W11-0316 Automatic keyphrase extraction by bridging vocabulary gap . In Proceedings of the Fifteenth Conference on Computational Natural Language Learning, pages 135--144, Portland, Oregon...

  43. [52]

    Wei Lu, Shengzhi Huang, Jinqing Yang, Yi Bu, Qikai Cheng, and Yong Huang. 2021. https://doi.org/https://doi.org/10.1016/j.ipm.2021.102594 Detecting research topic trends by author-defined keyword frequency . Information Processing & Management, 58(4):102594

  44. [53]

    Yichao Luo, Yige Xu, Jiacheng Ye, Xipeng Qiu, and Qi Zhang. 2021. https://doi.org/10.18653/v1/2021.findings-emnlp.45 Keyphrase generation with fine-grained evaluation-guided reinforcement learning . In Findings of the Association for Computational Linguistics: EMNLP 2021, page...

  45. [54]

    Debanjan Mahata, Navneet Agarwal, Dibya Gautam, Amardeep Kumar, Swapnil Parekh, Yaman Kumar Singla, Anish Acharya, and Rajiv Ratn Shah. 2022. https://ceur-ws.org/Vol-3317/Paper9.pdf LDKP - A dataset for identifying keyphrases from long scientific documents . In Proceedings of ...

  46. [55]

    López-López, and José Portela

    Roberto Martínez-Cruz, Alvaro J. López-López, and José Portela. 2023. http://arxiv.org/abs/2304.14177 Chatgpt vs state-of-the-art models: A benchmarking study in keyphrase generation task

  47. [56]

    Rui Meng, Tong Wang, Xingdi Yuan, Yingbo Zhou, and Daqing He. 2023. https://doi.org/10.18653/v1/2023.findings-acl.102 General-to-specific transfer labeling for domain adaptable keyphrase generation . In Findings of the Association for Computational Linguistics: ACL 2023, pages...

  48. [57]

    Rui Meng, Xingdi Yuan, Tong Wang, Sanqiang Zhao, Adam Trischler, and Daqing He. 2021. https://doi.org/10.18653/v1/2021.naacl-main.396 An empirical study on neural keyphrase generation . In Proceedings of the 2021 Conference of the North American Chapter of the Association for ...

  49. [58]

    Rui Meng, Sanqiang Zhao, Shuguang Han, Daqing He, Peter Brusilovsky, and Yu Chi. 2017. https://doi.org/10.18653/v1/P17-1054 Deep keyphrase generation . In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 582...

  50. [59]

    Thuy Dung Nguyen and Min-Yen Kan. 2007. Keyphrase extraction in scientific publications. In Asian Digital Libraries. Looking Back 10 Years and Forging New Frontiers, pages 317--326, Berlin, Heidelberg. Springer Berlin Heidelberg

  51. [60]

    Madhur Panwar, Shashank Shailabh, Milan Aggarwal, and Balaji Krishnamurthy. 2021. https://doi.org/10.18653/v1/2021.acl-long.299 TAN - NTM : T opic attention networks for neural topic modeling . In Proceedings of the 59th Annual Meeting of the Association for Computational Ling...

  52. [61]

    Eirini Papagiannopoulou and Grigorios Tsoumakas. 2020. https://doi.org/https://doi.org/10.1002/widm.1339 A review of keyphrase extraction . WIREs Data Mining and Knowledge Discovery, 10(2):e1339

  53. [62]

    Fr\' e d\' e ric Piedboeuf and Philippe Langlais. 2022. https://proceedings.neurips.cc/paper_files/paper/2022/file/f88709551258331f9ab31b33c71021a4-Paper-Datasets_and_Benchmarks.pdf A new dataset for multilingual keyphrase generation . In Advances in Neural Information Process...

  54. [63]

    M. F. Porter. 1997. An algorithm for suffix stripping, page 313–316. Morgan Kaufmann Publishers Inc., San Francisco, CA, USA

  55. [64]

    Jishnu Ray Chowdhury, Seo Yeon Park, Tuhin Kundu, and Cornelia Caragea. 2022. https://doi.org/10.18653/v1/2022.findings-emnlp.357 KPDROP : Improving absent keyphrase generation . In Findings of the Association for Computational Linguistics: EMNLP 2022, pages 4853--4870, Abu Dh...

  56. [65]

    Anna Rogers, Marzena Karpinska, Jordan Boyd-Graber, and Naoaki Okazaki. 2023. https://aclanthology.org/2023.acl-long.report Program chairs ' report on peer review at acl 2023 . In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1...

  57. [66]

    Tokala Yaswanth Sri Sai Santosh, Nikhil Reddy Varimalla, Anoop Vallabhajosyula, Debarshi Kumar Sanyal, and Partha Pratim Das. 2021. https://doi.org/10.1145/3459637.3482119 Hicova: Hierarchical conditional variational autoencoder for keyphrase generation . In Proceedings of the...

  58. [67]

    Liangying Shao, Liang Zhang, Minlong Peng, Guoqi Ma, Hao Yue, Mingming Sun, and Jinsong Su. 2024. https://doi.org/10.18653/v1/2024.emnlp-main.624 O ne2 S et + large language model: Best partners for keyphrase generation . In Proceedings of the 2024 Conference on Empirical Meth...

  59. [68]

    Xianjie Shen, Yinghan Wang, Rui Meng, and Jingbo Shang. 2022. https://doi.org/10.1609/aaai.v36i10.21381 Unsupervised deep keyphrase generation . Proceedings of the AAAI Conference on Artificial Intelligence, 36(10):11303--11311

  60. [69]

    Mingyang Song, Yi Feng, and Liping Jing. 2023 a . https://doi.org/10.18653/v1/2023.findings-eacl.161 A survey on recent advances in keyphrase extraction from pre-trained language models . In Findings of the Association for Computational Linguistics: EACL 2023, pages 2153--2164...

  61. [70]

    Mingyang Song, Haiyun Jiang, Shuming Shi, Songfang Yao, Shilong Lu, Yi Feng, Huafeng Liu, and Liping Jing. 2023 b . http://arxiv.org/abs/2303.13001 Is chatgpt a good keyphrase generator? a preliminary study

  62. [71]

    Sandeep Subramanian, Tong Wang, Xingdi Yuan, Saizheng Zhang, Adam Trischler, and Yoshua Bengio. 2018. https://doi.org/10.18653/v1/W18-2609 Neural models for key phrase extraction and question generation . In Proceedings of the Workshop on Machine Reading for Question Answering...

  63. [72]

    Ilya Sutskever, Oriol Vinyals, and Quoc V Le. 2014. https://proceedings.neurips.cc/paper_files/paper/2014/file/a14ac55a4f27472c5d894ec1c3c743d2-Paper.pdf Sequence to sequence learning with neural networks . In Advances in Neural Information Processing Systems, volume 27. Curra...

  64. [73]

    Avinash Swaminathan, Haimin Zhang, Debanjan Mahata, Rakesh Gosangi, Rajiv Ratn Shah, and Amanda Stent. 2020. https://doi.org/10.18653/v1/2020.emnlp-main.645 A preliminary exploration of GAN s for keyphrase generation . In Proceedings of the 2020 Conference on Empirical Methods...

  65. [74]

    Edwin Thomas and Sowmya Vajjala. 2024 a . https://doi.org/10.18653/v1/2024.findings-naacl.102 Improving absent keyphrase generation with diversity heads . In Findings of the Association for Computational Linguistics: NAACL 2024, pages 1568--1584, Mexico City, Mexico. Associati...

  66. [75]

    Edwin Thomas and Sowmya Vajjala. 2024 b . https://aclanthology.org/2024.lrec-main.849 Keyphrase generation: Lessons from a reproducibility study . In Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-CO...

  67. [76]

    Xiaojun Wan and Jianguo Xiao. 2008. Single document keyphrase extraction using neighborhood knowledge. In Proceedings of the 23rd National Conference on Artificial Intelligence - Volume 2, AAAI'08, page 855–860. AAAI Press

  68. [77]

    Xiaojun Wan, Jianwu Yang, and Jianguo Xiao. 2007. https://aclanthology.org/P07-1070 Towards an iterative reinforcement approach for simultaneous document summarization and keyword extraction . In Proceedings of the 45th Annual Meeting of the Association of Computational Lingui...

  69. [78]

    Siyu Wang, Jianhui Jiang, Yao Huang, and Yin Wang. 2022. https://aclanthology.org/2022.coling-1.204 Automatic keyphrase generation by incorporating dual copy mechanisms in sequence-to-sequence learning . In Proceedings of the 29th International Conference on Computational Ling...

  70. [79]

    Lyu, and Shuming Shi

    Yue Wang, Jing Li, Hou Pong Chan, Irwin King, Michael R. Lyu, and Shuming Shi. 2019. https://doi.org/10.18653/v1/P19-1240 Topic-aware neural keyphrase generation for social media language . In Proceedings of the 57th Annual Meeting of the Association for Computational Linguist...

  71. [80]

    Di Wu, Wasi Ahmad, and Kai-Wei Chang. 2023. https://doi.org/10.18653/v1/2023.emnlp-main.410 Rethinking model selection and decoding for keyphrase generation with pre-trained sequence-to-sequence models . In Proceedings of the 2023 Conference on Empirical Methods in Natural Lan...

  72. [81]

    Di Wu, Wasi Ahmad, and Kai-Wei Chang. 2024 a . https://aclanthology.org/2024.lrec-main.1083 On leveraging encoder-only pre-trained language models for effective keyphrase generation . In Proceedings of the 2024 Joint International Conference on Computational Linguistics, Langu...

  73. [82]

    Di Wu, Wasi Ahmad, Sunipa Dev, and Kai-Wei Chang. 2022 a . https://doi.org/10.18653/v1/2022.findings-emnlp.49 Representation learning for resource-constrained keyphrase generation . In Findings of the Association for Computational Linguistics: EMNLP 2022, pages 700--716, Abu D...

  74. [83]

    Di Wu, Xiaoxian Shen, and Kai-Wei Chang. 2024 b . https://doi.org/10.18653/v1/2024.findings-emnlp.494 M eta KP : On-demand keyphrase generation . In Findings of the Association for Computational Linguistics: EMNLP 2024, pages 8420--8437, Miami, Florida, USA. Association for Co...

  75. [84]

    Di Wu, Da Yin, and Kai-Wei Chang. 2024 c . https://doi.org/10.18653/v1/2024.findings-acl.117 KPE val: Towards fine-grained semantic-based keyphrase evaluation . In Findings of the Association for Computational Linguistics: ACL 2024, pages 1959--1981, Bangkok, Thailand. Associa...

  76. [85]

    Huanqin Wu, Wei Liu, Lei Li, Dan Nie, Tao Chen, Feng Zhang, and Di Wang. 2021. https://doi.org/10.18653/v1/2021.findings-acl.73 U ni K eyphrase: A unified extraction and generation framework for keyphrase prediction . In Findings of the Association for Computational Linguistic...

  77. [86]

    Huanqin Wu, Baijiaxin Ma, Wei Liu, Tao Chen, and Dan Nie. 2022 b . https://doi.org/10.1609/aaai.v36i10.21402 Fast and constrained absent keyphrase generation by prompt-based learning . Proceedings of the AAAI Conference on Artificial Intelligence, 36(10):11495--11503

  78. [87]

    Binbin Xie, Jia Song, Liangying Shao, Suhang Wu, Xiangpeng Wei, Baosong Yang, Huan Lin, Jun Xie, and Jinsong Su. 2023. https://doi.org/https://doi.org/10.1016/j.ipm.2023.103382 From statistical methods to deep learning, automatic keyphrase prediction: A survey . Information Pr...

  79. [88]

    Binbin Xie, Xiangpeng Wei, Baosong Yang, Huan Lin, Jun Xie, Xiaoli Wang, Min Zhang, and Jinsong Su. 2022. https://doi.org/10.18653/v1/2022.emnlp-main.491 WR - O ne2 S et: Towards well-calibrated keyphrase generation . In Proceedings of the 2022 Conference on Empirical Methods ...

  80. [89]

    Jianxin Yang, Wenge Rong, Libin Shi, and Zhang Xiong. 2019. https://doi.org/10.18653/v1/N19-1228 S equential A ttention with K eyword M ask M odel for C ommunity-based Q uestion A nswering . In Proceedings of the 2019 Conference of the North A merican Chapter of the Associatio...

  81. [90]

    Hai Ye and Lu Wang. 2018. https://doi.org/10.18653/v1/D18-1447 Semi-supervised learning for neural keyphrase generation . In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 4142--4153, Brussels, Belgium. Association for Computation...

  82. [91]

    Jiacheng Ye, Ruijian Cai, Tao Gui, and Qi Zhang. 2021 a . https://doi.org/10.18653/v1/2021.emnlp-main.213 Heterogeneous graph neural networks for keyphrase generation . In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 2705--2715,...

  83. [92]

    Jiacheng Ye, Tao Gui, Yichao Luo, Yige Xu, and Qi Zhang. 2021 b . https://doi.org/10.18653/v1/2021.acl-long.354 O ne2 S et: G enerating diverse keyphrases as a set . In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th Interna...

  84. [93]

    Xingdi Yuan, Tong Wang, Rui Meng, Khushboo Thaker, Peter Brusilovsky, Daqing He, and Adam Trischler. 2020. https://doi.org/10.18653/v1/2020.acl-main.710 One size does not fit all: Generating and evaluating variable number of keyphrases . In Proceedings of the 58th Annual Meeti...

  85. [94]

    Hongyuan Zha. 2002. https://doi.org/10.1145/564376.564398 Generic summarization and keyphrase extraction using mutual reinforcement principle and sentence clustering . In Proceedings of the 25th Annual International ACM SIGIR Conference on Research and Development in Informati...

  86. [95]

    Chengxiang Zhai. 1997. https://doi.org/10.3115/974557.974603 Fast statistical parsing of noun phrases for document indexing . In Fifth Conference on Applied Natural Language Processing, pages 312--319, Washington, DC, USA. Association for Computational Linguistics

  87. [96]

    Weichao Zhang, Ruqing Zhang, Jiafeng Guo, Maarten de Rijke, Yixing Fan, and Xueqi Cheng. 2024. https://doi.org/10.18653/v1/2024.emnlp-main.300 Pretraining data detection for large language models: A divergence-based calibration method . In Proceedings of the 2024 Conference on...

  88. [97]

    Yuxiang Zhang, Tao Jiang, Tianyu Yang, Xiaoli Li, and Suge Wang. 2022. https://doi.org/10.1145/3477495.3531990 Htkg: Deep keyphrase generation with neural hierarchical topic guidance . In Proceedings of the 45th International ACM SIGIR Conference on Research and Development in...

  89. [98]

    Guangzhen Zhao, Guoshun Yin, Peng Yang, and Yu Yao. 2022. https://doi.org/10.18653/v1/2022.emnlp-main.529 Keyphrase generation via soft and hard semantic corrections . In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 7757--7768, ...

  90. [99]

    Jing Zhao, Junwei Bao, Yifan Wang, Youzheng Wu, Xiaodong He, and Bowen Zhou. 2021. https://doi.org/10.18653/v1/2021.naacl-main.455 SGG : Learning to select, guide, and generate for keyphrase generation . In Proceedings of the 2021 Conference of the North American Chapter of th...

  91. [100]

    Jing Zhao and Yuxiang Zhang. 2019. https://doi.org/10.18653/v1/P19-1515 Incorporating linguistic constraints into keyphrase generation . In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 5224--5233, Florence, Italy. Association f...

  92. [101]

    Baohang Zhou, Zezhong Wang, Lingzhi Wang, Hongru Wang, Ying Zhang, Kehui Song, Xuhui Sui, and Kam-Fai Wong. 2024. https://doi.org/10.18653/v1/2024.findings-acl.35 DPDLLM : A black-box framework for detecting pre-training data from large language models . In Findings of the Ass...

  93. [102]

    Erion Çano and Ondřej Bojar. 2019. https://doi.org/10.23919/FRUCT48121.2019.8981519 Keyphrase generation: A multi-aspect survey . In 2019 25th Conference of Open Innovations Association (FRUCT), pages 85--94

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.