Pith. sign in

REVIEW 2 cited by

VSVC: Backdoor attack against Keyword Spotting based on Voiceprint Selection and Voice Conversion

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2212.10103 v1 pith:2OZTWF6V submitted 2022-12-20 cs.SD cs.AIcs.CRcs.LGeess.AS

classification cs.SDcs.AIcs.CRcs.LGeess.AS
keywords attacktrainingbackdoordatavoicevsvcconversionkeyword
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Keyword spotting (KWS) based on deep neural networks (DNNs) has achieved massive success in voice control scenarios. However, training of such DNN-based KWS systems often requires significant data and hardware resources. Manufacturers often entrust this process to a third-party platform. This makes the training process uncontrollable, where attackers can implant backdoors in the model by manipulating third-party training data. An effective backdoor attack can force the model to make specified judgments under certain conditions, i.e., triggers. In this paper, we design a backdoor attack scheme based on Voiceprint Selection and Voice Conversion, abbreviated as VSVC. Experimental results demonstrated that VSVC is feasible to achieve an average attack success rate close to 97% in four victim models when poisoning less than 1% of the training data.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. SPBA: Utilizing Speech Large Language Model for Backdoor Attacks on Speech Classification Models

    cs.SD 2025-06 conditional novelty 6.0 of 10

    A speech backdoor attack uses SLLM-generated timbre and emotion triggers with MGDA-balanced training to implant multiple effective backdoors in speech classifiers.

  2. Investigating Vulnerabilities and Defenses Against Audio-Visual Attacks: A Comprehensive Survey Emphasizing Multimodal Models

    cs.CR 2025-06 conditional novelty 4.0 of 10

    A survey that organizes audio and video AI security research into adversarial, backdoor, and jailbreak attacks, with extra attention to multimodal large language models.

Pith tools