Pith. sign in

REVIEW 1 cited by

Speaker-Text Retrieval via Contrastive Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2312.06055 v1 pith:XC4VOBX2 submitted 2023-12-11 cs.SD eess.AS

classification cs.SDeess.AS
keywords speakerlearningretrievalcontrastivetasktextacrossadditional
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In this study, we introduce a novel cross-modal retrieval task involving speaker descriptions and their corresponding audio samples. Utilizing pre-trained speaker and text encoders, we present a simple learning framework based on contrastive learning. Additionally, we explore the impact of incorporating speaker labels into the training process. Our findings establish the effectiveness of linking speaker and text information for the task for both English and Japanese languages, across diverse data configurations. Additional visual analysis unveils potential nuanced associations between speaker clustering and retrieval performance.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. RA-CLAP: Relation-Augmented Emotional Speaking Style Contrastive Language-Audio Pretraining For Speech Retrieval

    cs.SD 2025-05 conditional novelty 6.0 of 10

    Authors propose ESS-CLAP and RA-CLAP, contrastive speech-text models for emotional speaking style retrieval, evaluated on PromptSpeech, TextrolSpeech, and SpeechCraft.

Pith tools