Pith. sign in

REVIEW 4 cited by

End-to-End Spectro-Temporal Graph Attention Networks for Speaker Verification Anti-Spoofing and Speech Deepfake Detection

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2107.12710 v2 pith:3KZKBTTW submitted 2021-07-27 eess.AS cs.SD

classification eess.AScs.SD
keywords graphdetectionfusionmodelspeechtemporalartefactsattention
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Artefacts that serve to distinguish bona fide speech from spoofed or deepfake speech are known to reside in specific subbands and temporal segments. Various approaches can be used to capture and model such artefacts, however, none works well across a spectrum of diverse spoofing attacks. Reliable detection then often depends upon the fusion of multiple detection systems, each tuned to detect different forms of attack. In this paper we show that better performance can be achieved when the fusion is performed within the model itself and when the representation is learned automatically from raw waveform inputs. The principal contribution is a spectro-temporal graph attention network (GAT) which learns the relationship between cues spanning different sub-bands and temporal intervals. Using a model-level graph fusion of spectral (S) and temporal (T) sub-graphs and a graph pooling strategy to improve discrimination, the proposed RawGAT-ST model achieves an equal error rate of 1.06 % for the ASVspoof 2019 logical access database. This is one of the best results reported to date and is reproducible using an open source implementation.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Probing Speaker Identity Sensitivity in Audio Deepfake Detectors

    cs.SD 2026-07 conditional novelty 6.0 of 10

    A new Identity Sensitivity Score flags misclassified audio deepfake detections with AUC up to 0.954, but its claim to isolate speaker-identity behavior from plain confidence is not yet controlled.

  2. Generalizable Audio Spoofing Detection using Non-Semantic Representations

    cs.SD 2025-08 conditional novelty 5.0 of 10

    Frozen non-semantic TRILLson embeddings with a lightweight backend beat prior spoofing detectors on out-of-domain datasets while staying competitive in-domain.

  3. Teffic-Audio: Tell Fact from Fiction

    cs.SD 2026-07 conditional novelty 4.0 of 10

    A simple Conformer deepfake detector trained with multi-source balanced sampling and diverse augmentation reaches 1.454% pooled EER on Speech-DF-Arena, first among public systems.

  4. Trusted Fake Audio Detection Based on Dirichlet Distribution

    cs.SD 2025-06 conditional novelty 4.0 of 10

    Applying Dirichlet-based evidential learning to three fake audio detectors yields modest EER gains and apparently better calibration on ASVspoof, but the calibration comparison is methodologically weak.

Pith tools