REVIEW 3 cited by
Muting Whisper: A Universal Acoustic Adversarial Attack on Speech Foundation Models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
abstract
Recent developments in large speech foundation models like Whisper have led to their widespread use in many automatic speech recognition (ASR) applications. These systems incorporate `special tokens' in their vocabulary, such as $\texttt{<|endoftext|>}$, to guide their language generation process. However, we demonstrate that these tokens can be exploited by adversarial attacks to manipulate the model's behavior. We propose a simple yet effective method to learn a universal acoustic realization of Whisper's $\texttt{<|endoftext|>}$ token, which, when prepended to any speech signal, encourages the model to ignore the speech and only transcribe the special token, effectively `muting' the model. Our experiments demonstrate that the same, universal 0.64-second adversarial audio segment can successfully mute a target Whisper ASR model for over 97\% of speech samples. Moreover, we find that this universal adversarial audio segment often transfers to new datasets and tasks. Overall this work demonstrates the vulnerability of Whisper models to `muting' adversarial attacks, where such attacks can pose both risks and potential benefits in real-world settings: for example the attack can be used to bypass speech moderation systems, or conversely the attack can also be used to protect private speech data.
Forward citations
Cited by 3 Pith papers
-
Generative Testing of Automated Speech Recognition Systems
Phoneme-level latent interpolation in a TTS model yields ~98% black-box ASR failures with higher naturalness than waveform attacks and quality competitive with white-box PGD.
-
Universal Acoustic Adversarial Attacks for Flexible Control of Speech-LLMs
A single learned 3.2-second audio prefix can mute or redirect speech LLMs, and can be trained to selectively mute only targeted genders or languages.
-
Whisper Smarter, not Harder: Adversarial Attack on Partial Suppression
Aiming for partial instead of full suppression can make adversarial audio attacks less noticeable, and low-pass filtering may defend against them.
Discussion (0). Continue with ORCID to comment.