Pith. sign in

REVIEW 2 cited by

PBSM: Backdoor attack against Keyword spotting based on pitch boosting and sound masking

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2211.08697 v1 pith:M7OE6M4G submitted 2022-11-16 cs.SD cs.AIcs.CRcs.LGeess.AS

classification cs.SDcs.AIcs.CRcs.LGeess.AS
keywords attackdatatrainingbackdoorpbsmboostingdeepkeyword
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Keyword spotting (KWS) has been widely used in various speech control scenarios. The training of KWS is usually based on deep neural networks and requires a large amount of data. Manufacturers often use third-party data to train KWS. However, deep neural networks are not sufficiently interpretable to manufacturers, and attackers can manipulate third-party training data to plant backdoors during the model training. An effective backdoor attack can force the model to make specified judgments under certain conditions, i.e., triggers. In this paper, we design a backdoor attack scheme based on Pitch Boosting and Sound Masking for KWS, called PBSM. Experimental results demonstrated that PBSM is feasible to achieve an average attack success rate close to 90% in three victim models when poisoning less than 1% of the training data.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. SPBA: Utilizing Speech Large Language Model for Backdoor Attacks on Speech Classification Models

    cs.SD 2025-06 conditional novelty 6.0 of 10

    A speech backdoor attack uses SLLM-generated timbre and emotion triggers with MGDA-balanced training to implant multiple effective backdoors in speech classifiers.

  2. Investigating Vulnerabilities and Defenses Against Audio-Visual Attacks: A Comprehensive Survey Emphasizing Multimodal Models

    cs.CR 2025-06 conditional novelty 4.0 of 10

    A survey that organizes audio and video AI security research into adversarial, backdoor, and jailbreak attacks, with extra attention to multimodal large language models.

Pith tools