Pith. sign in

Graph Attention Networks for Anti-Spoofing

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it
abstract

The cues needed to detect spoofing attacks against automatic speaker verification are often located in specific spectral sub-bands or temporal segments. Previous works show the potential to learn these using either spectral or temporal self-attention mechanisms but not the relationships between neighbouring sub-bands or segments. This paper reports our use of graph attention networks (GATs) to model these relationships and to improve spoofing detection performance. GATs leverage a self-attention mechanism over graph structured data to model the data manifold and the relationships between nodes. Our graph is constructed from representations produced by a ResNet. Nodes in the graph represent information either in specific sub-bands or temporal segments. Experiments performed on the ASVspoof 2019 logical access database show that our GAT-based model with temporal attention outperforms all of our baseline single systems. Furthermore, GAT-based systems are complementary to a set of existing systems. The fusion of GAT-based models with more conventional countermeasures delivers a 47% relative improvement in performance compared to the best performing single GAT system.

citation-role summary

background 1

citation-polarity summary

fields

cs.SD 1

years

2025 1

verdicts

CONDITIONAL 1

roles

background 1

polarities

unclear 1

representative citing papers

Trusted Fake Audio Detection Based on Dirichlet Distribution

cs.SD · 2025-06-03 · conditional · novelty 4.0

Applying Dirichlet-based evidential learning to three fake audio detectors yields modest EER gains and apparently better calibration on ASVspoof, but the calibration comparison is methodologically weak.

citing papers explorer

Showing 1 of 1 citing paper.

  • Trusted Fake Audio Detection Based on Dirichlet Distribution cs.SD · 2025-06-03 · conditional · none · ref 27 · internal anchor

    Applying Dirichlet-based evidential learning to three fake audio detectors yields modest EER gains and apparently better calibration on ASVspoof, but the calibration comparison is methodologically weak.