ParaSpeechCLAP: A Dual-Encoder Speech-Text Model for Rich Stylistic Language-Audio Pretraining

Anuj Diwan; David Harwath; Eunsol Choi

arxiv: 2603.28737 · v2 · pith:SHWJ7X7Pnew · submitted 2026-03-30 · 📡 eess.AS · cs.AI· cs.CL· cs.SD

ParaSpeechCLAP: A Dual-Encoder Speech-Text Model for Rich Stylistic Language-Audio Pretraining

Anuj Diwan , Eunsol Choi , David Harwath This is my paper

classification 📡 eess.AS cs.AIcs.CLcs.SD

keywords modelsparaspeechclapmodelstyleclassificationdual-encoderintrinsicrich

0 comments

read the original abstract

We introduce ParaSpeechCLAP, a family of dual-encoder models that map speech and text style captions into a shared embedding space, supporting rich intrinsic (speaker-level) and situational (utterance-level) descriptors, such as pitch, texture, and emotion, beyond the narrow set handled by existing models. We train separate Intrinsic and Situational models alongside a unified Combined model, finding that specialized models are stronger on individual style dimensions while the unified model excels on compositional evaluation. We further show that ParaSpeechCLAP-Intrinsic benefits from an additional classification loss and class-balanced training. We demonstrate performance on style caption retrieval, speech attribute classification, and usability as inference-time reward models for style-prompted TTS. ParaSpeechCLAP models outperform baselines on most metrics across all three applications. Our models and code are released at https://github.com/ajd12342/paraspeechclap .

This paper has not been read by Pith yet.

ParaSpeechCLAP: A Dual-Encoder Speech-Text Model for Rich Stylistic Language-Audio Pretraining

discussion (0)