Pith. sign in

REVIEW 1 cited by

Improving Query-by-Vocal Imitation with Contrastive Learning and Audio Pretraining

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2408.11638 v1 pith:OXZ6ESC3 submitted 2024-08-21 eess.AS

Improving Query-by-Vocal Imitation with Contrastive Learning and Audio Pretraining

classification eess.AS
keywords audiocontrastiveimitationlearningpre-trainedvocalfeatureimitations
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Query-by-Vocal Imitation (QBV) is about searching audio files within databases using vocal imitations created by the user's voice. Since most humans can effectively communicate sound concepts through voice, QBV offers the more intuitive and convenient approach compared to text-based search. To fully leverage QBV, developing robust audio feature representations for both the vocal imitation and the original sound is crucial. In this paper, we present a new system for QBV that utilizes the feature extraction capabilities of Convolutional Neural Networks pre-trained with large-scale general-purpose audio datasets. We integrate these pre-trained models into a dual encoder architecture and fine-tune them end-to-end using contrastive learning. A distinctive aspect of our proposed method is the fine-tuning strategy of pre-trained models using an adapted NT-Xent loss for contrastive learning, creating a shared embedding space for reference recordings and vocal imitations. The proposed system significantly enhances audio retrieval performance, establishing a new state of the art on both coarse- and fine-grained QBV tasks.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Quality Audio Prototyping: a prototype system for unified sound retrieval and procedural generation

    cs.SD 2026-05 unverdicted novelty 4.0

    QuAP is a working prototype combining similarity-based audio retrieval, real-time procedural models, and a perceptually guided parameter assistant, with evaluations showing quality gains in five of six synthesis model...