Pith. sign in

REVIEW 1 cited by

SimCKP: Simple Contrastive Learning of Keyphrase Representations

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2310.08221 v1 pith:HTKBPJDT submitted 2023-10-12 cs.CL cs.AIcs.LG

classification cs.CLcs.AIcs.LG
keywords documentkeyphrasescontrastivekeyphraselearningrepresentationsaimscorresponding
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Keyphrase generation (KG) aims to generate a set of summarizing words or phrases given a source document, while keyphrase extraction (KE) aims to identify them from the text. Because the search space is much smaller in KE, it is often combined with KG to predict keyphrases that may or may not exist in the corresponding document. However, current unified approaches adopt sequence labeling and maximization-based generation that primarily operate at a token level, falling short in observing and scoring keyphrases as a whole. In this work, we propose SimCKP, a simple contrastive learning framework that consists of two stages: 1) An extractor-generator that extracts keyphrases by learning context-aware phrase-level representations in a contrastive manner while also generating keyphrases that do not appear in the document; 2) A reranker that adapts scores for each generated phrase by likewise aligning their representations with the corresponding document. Experimental results on multiple benchmark datasets demonstrate the effectiveness of our proposed approach, which outperforms the state-of-the-art models by a significant margin.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. QExplorer: Large Language Model Based Query Extraction for Toxic Content Exploration

    cs.IR 2025-02 conditional novelty 5.0 of 10

    A two-stage fine-tuned LLM, using SFT followed by DPO with search-engine feedback, extracts queries that find more toxic items on a second-hand marketplace than human auditors do.

Pith tools