Pith. sign in

REVIEW 1 cited by

Complex-valued neural networks for voice anti-spoofing

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2308.11800 v1 pith:DPIXBEVD submitted 2023-08-22 cs.SD cs.LGeess.AS

classification cs.SDcs.LGeess.AS
keywords methodsaudioanti-spoofingcomplex-valuedexplainableinformationphaseapproach
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Current anti-spoofing and audio deepfake detection systems use either magnitude spectrogram-based features (such as CQT or Melspectrograms) or raw audio processed through convolution or sinc-layers. Both methods have drawbacks: magnitude spectrograms discard phase information, which affects audio naturalness, and raw-feature-based models cannot use traditional explainable AI methods. This paper proposes a new approach that combines the benefits of both methods by using complex-valued neural networks to process the complex-valued, CQT frequency-domain representation of the input audio. This method retains phase information and allows for explainable AI methods. Results show that this approach outperforms previous methods on the "In-the-Wild" anti-spoofing dataset and enables interpretation of the results through explainable AI. Ablation studies confirm that the model has learned to use phase information to detect voice spoofing.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. What You Read Isn't What You Hear: Linguistic Sensitivity in Deepfake Speech Detection

    cs.LG 2025-05 conditional novelty 6.0 of 10

    Small semantic-preserving changes to transcripts, passed through text-to-speech, significantly reduce the accuracy of both open-source and commercial audio anti-spoofing detectors.

Pith tools