Pith. sign in

REVIEW 4 cited by

Towards Explainable Spoofed Speech Attribution and Detection:a Probabilistic Approach for Characterizing Speech Synthesizer Components

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2502.04049 v3 pith:CSB7EKJU submitted 2025-02-06 eess.AS

classification eess.AS
keywords embeddingsprobabilisticdetectionattributeattributionspeechspoofingtask
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

We propose an explainable probabilistic framework for characterizing spoofed speech by decomposing it into probabilistic attribute embeddings. Unlike raw high-dimensional countermeasure embeddings, which lack interpretability, the proposed probabilistic attribute embeddings aim to detect specific speech synthesizer components, represented through high-level attributes and their corresponding values. We use these probabilistic embeddings with four classifier back-ends to address two downstream tasks: spoofing detection and spoofing attack attribution. The former is the well-known bonafide-spoof detection task, whereas the latter seeks to identify the source method (generator) of a spoofed utterance. We additionally use Shapley values, a widely used technique in machine learning, to quantify the relative contribution of each attribute value to the decision-making process in each task. Results on the ASVspoof2019 dataset demonstrate the substantial role of duration and conversion modeling in spoofing detection; and waveform generation and speaker modeling in spoofing attack attribution. In the detection task, the probabilistic attribute embeddings achieve $99.7\%$ balanced accuracy and $0.22\%$ equal error rate (EER), closely matching the performance of raw embeddings ($99.9\%$ balanced accuracy and $0.22\%$ EER). Similarly, in the attribution task, our embeddings achieve $90.23\%$ balanced accuracy and $2.07\%$ EER, compared to $90.16\%$ and $2.11\%$ with raw embeddings. These results demonstrate that the proposed framework is both inherently explainable by design and capable of achieving performance comparable to raw CM embeddings.

Discussion (0). Sign in to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Towards Neural Audio Codec Source Parsing

    eess.AS 2025-06 conditional novelty 7.0 of 10

    NACSP predicts codec generation parameters from deepfake audio, and the proposed HYDRA hyperbolic model beats Euclidean baselines on most benchmark tasks.

  2. How to Leverage Synthetic Speech for LLM-Based ASR Systems?

    cs.CL 2026-06 unverdicted novelty 6.0 of 10

    Layer-wise pooling plus RIR-augmented synthetic speech matches a 100%-real ASR baseline with only 25% real data (13.6 h) and beats it at higher real fractions.

  3. Open-Set Source Tracing of Audio Deepfake Systems

    eess.AS 2025-07 conditional novelty 5.0 of 10

    Softmax energy, a modified out-of-distribution score, improves open-set source tracing of audio deepfake systems, achieving a 31% relative FPR95 reduction and best FPR95 of 8.3% with augmentation.

  4. Towards Generalized Source Tracing for Codec-Based Deepfake Speech

    cs.SD 2025-06 conditional novelty 5.0 of 10

    SASTNet, which fuses Whisper semantic features with Wav2Vec2 and AudioMAE acoustic features, improves source tracing for codec-based deepfake speech on CodecFake+, while exposing that prior models overfit to silence.

Pith tools