Pith. sign in

REVIEW 3 cited by

ParaCLAP -- Towards a general language-audio model for computational paralinguistic tasks

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.07203 v1 pith:F7NJ6MML submitted 2024-06-11 cs.SD eess.AS

ParaCLAP -- Towards a general language-audio model for computational paralinguistic tasks

classification cs.SD eess.AS
keywords audiomodelstasksclapcomputationalgeneralparalinguisticqueries
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Contrastive language-audio pretraining (CLAP) has recently emerged as a method for making audio analysis more generalisable. Specifically, CLAP-style models are able to `answer' a diverse set of language queries, extending the capabilities of audio models beyond a closed set of labels. However, CLAP relies on a large set of (audio, query) pairs for pretraining. While such sets are available for general audio tasks, like captioning or sound event detection, there are no datasets with matched audio and text queries for computational paralinguistic (CP) tasks. As a result, the community relies on generic CLAP models trained for general audio with limited success. In the present study, we explore training considerations for ParaCLAP, a CLAP-style model suited to CP, including a novel process for creating audio-language queries. We demonstrate its effectiveness on a set of computational paralinguistic tasks, where it is shown to surpass the performance of open-source state-of-the-art models.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. ParaSpeechCLAP: A Dual-Encoder Speech-Text Model for Rich Stylistic Language-Audio Pretraining

    eess.AS 2026-03 conditional novelty 6.0

    Dual-encoder speech-text models trained on rich intrinsic and situational style captions outperform prior CLAP-style baselines on retrieval, classification, and inference-time TTS style guidance.

  2. AutoSIFT: Automatic Style Sifting for Controllable Speech Generation with Arbitrary Style Infilling

    cs.SD 2026-07 unverdicted novelty 5.0

    AutoSIFT disentangles text-describable style categories from residual speech styles and selectively infills only the categories the user specifies.

  3. AutoSIFT: Automatic Style Sifting for Controllable Speech Generation with Arbitrary Style Infilling

    cs.SD 2026-07 conditional novelty 5.0

    A TTS framework that replaces only text-specified style categories and infills all unspecified and residual style from a reference speech recording.