Pith. sign in

REVIEW 5 cited by

Temporal-Channel Modeling in Multi-head Self-Attention for Synthetic Speech Detection

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.17376 v1 pith:KBXNYRXJ submitted 2024-06-25 cs.SD cs.AIeess.AS

classification cs.SDcs.AIeess.AS
keywords speechsynthetictemporal-channelmhsamodelingtemporalimprovementinput
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Recent synthetic speech detectors leveraging the Transformer model have superior performance compared to the convolutional neural network counterparts. This improvement could be due to the powerful modeling ability of the multi-head self-attention (MHSA) in the Transformer model, which learns the temporal relationship of each input token. However, artifacts of synthetic speech can be located in specific regions of both frequency channels and temporal segments, while MHSA neglects this temporal-channel dependency of the input sequence. In this work, we proposed a Temporal-Channel Modeling (TCM) module to enhance MHSA's capability for capturing temporal-channel dependencies. Experimental results on the ASVspoof 2021 show that with only 0.03M additional parameters, the TCM module can outperform the state-of-the-art system by 9.25% in EER. Further ablation study reveals that utilizing both temporal and channel information yields the most improvement for detecting synthetic speech.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Replay Attacks Against Audio Deepfake Detection

    cs.SD 2025-05 accept novelty 7.0 of 10

    Physical replay attacks, playing and re-recording synthetic speech, sharply degrade the accuracy of six open-source audio deepfake detectors, and the new ReplayDF dataset lets the community measure and defend against this.

  2. Leveraging Gradient Reversal Loss and Multitask Learning for Datasets-Aware Audio Deepfake Detection

    eess.AS 2026-07 conditional novelty 5.0 of 10

    Adding dataset identity as an auxiliary task or adversarial label improves aggregate audio-deepfake detection EER on the 2025 Speech DeepFake Arena benchmark.

  3. Bona fide Cross Testing Reveals Weak Spot in Audio Deepfake Detection Systems

    cs.SD 2025-09 reject novelty 5.0 of 10

    A new evaluation protocol exhaustively pairs 164 speech synthesizers with nine bona fide speech types and reports max-pooled EERs, revealing larger failures than pooled averages show.

  4. Teffic-Audio: Tell Fact from Fiction

    cs.SD 2026-07 conditional novelty 4.0 of 10

    A simple Conformer deepfake detector trained with multi-source balanced sampling and diverse augmentation reaches 1.454% pooled EER on Speech-DF-Arena, first among public systems.

  5. Two Views, One Truth: Spectral and Self-Supervised Features Fusion for Robust Speech Deepfake Detection

    cs.SD 2025-07 conditional novelty 4.0 of 10

    Fusing CQCC spectral features with Wav2Vec2.0 embeddings via cross-attention lowers average equal error rate from 10.87% to 6.80% across four speech deepfake benchmarks.

Pith tools