REVIEW 3 major objections 3 minor 1 cited by
Multi-Granularity Adaptive Time-Frequency Attention Framework for Audio Deepfake Detection under Real-World Communication Degradations
T0 review · 3 major / 3 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read Audio deepfake detectors fail on real phone calls; this paper offers a fix
desk verdict The abstract pitches a sensible robustness idea, but the submitted full text is unreadable mojibake, so the performance claims are uncheckable. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The Multi-Granularity Adaptive Attention (MGAA) architecture sits at the core: a set of customizable multi-scale attention heads captures both global and local receptive fields at different time-frequency granularities. A learned adaptive fusion mechanism then weights those attention branches by the saliency of time-frequency regions, dynamically reallocating focus according to the degradation present. This mechanism is what localizes subtle forgery traces that channel distortions would otherwise obscure.
What would settle it
Run the same trained framework on an unseen codec such as Opus at a low bitrate, on tandem-coded audio, or on audio with added background noise and jitter; if accuracy collapses to baseline levels while existing methods hold up, the paper's claim of real-world robustness is not supported.
Extended reading notes
Core claim
The central claim is that a single audio deepfake detection framework can remain accurate across realistic communication degradations instead of collapsing like clean-condition detectors do. The framework accepts multiple types of time-frequency representations and, through MGAA, learns to locate and amplify subtle synthetic-speech artifacts even after compression and packet loss have distorted the signal. The authors report consistent outperformance over state-of-the-art baselines across six speech codecs and five levels of packet loss, and show that MGAA-enhanced features increase the separability of real and fake audio and sharpen the decision boundary.
Load-bearing premise
Everything rests on the assumption that the six tested speech codecs and five packet-loss levels represent real-world communication degradations well enough that performance on them transfers to other distortions such as noise, jitter, or new codecs.
Editorial extensions
If this is right
- A single trained model can handle multiple codecs and packet-loss levels without needing a separate detector per condition.
- Attention reallocation should make detection depend less on distorted content and more on the time-frequency regions that still carry forgery evidence.
- Sharper separation between real and fake feature distributions should translate into lower false-accept rates on degraded voice calls.
- The framework provides a practical path toward deploying deepfake detection in voice communication pipelines rather than only on pristine audio.
- The design's ability to accept multiple time-frequency representations means it can be adapted as new spectrogram-like features become available.
Reading between the lines
- If adaptive time-frequency saliency is the real driver of the gain, the same architecture may partially transfer to unseen codecs whose artifacts are spectrally similar, but that transfer is untested and would need its own evaluation.
- The paper's logic suggests combining MGAA with channel-simulation data augmentation during training could push robustness further, since the model would see an even wider menu of distortions.
- The degradation menu may underrepresent real-world conditions such as background noise, jitter, tandem codecs, and network echo; a testable extension is to measure MGAA against those distortions explicitly.
- Because the framework fuses multiple saliency-weighted branches, it could also serve as a diagnostic tool for understanding which time-frequency regions current deepfake generators fail to reproduce under channel stress.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes MGAA, a multi-granularity adaptive time-frequency attention framework for audio deepfake detection under communication degradations. The abstract claims that MGAA is the first unified framework for robust audio deepfake detection under such degradations, that its adaptive fusion mechanism dynamically reallocates attention according to time-frequency saliency, and that extensive experiments show consistent state-of-the-art outperformance across six speech codecs and five levels of packet loss. The full text as submitted is almost entirely corrupted mojibake, so the architecture description, equations, experimental tables, dataset specifications, and training details cannot be inspected. The central contribution is therefore an empirical claim that is currently unsupported by any readable evidence in the submitted artifact.
Significance. If the empirical claims are verified, the paper addresses a practically important gap: most audio deepfake detectors are evaluated on clean audio, whereas real communication channels introduce codec compression and packet loss. The proposed architecture, combining multi-scale attention heads over time-frequency representations with a saliency-based adaptive fusion mechanism, is conceptually plausible and could offer a useful robustness-oriented design for the field. However, the lack of any inspectable experimental protocol, quantitative results, or statistical analysis means that the claimed consistent outperformance cannot be assessed. The paper does not appear to provide machine-checked proofs, reproducible code, or a parameter-free derivation; the contribution is entirely empirical and depends on the unreadable experimental section.
major comments (3)
- [Full text / Abstract] The central empirical claim of consistent state-of-the-art outperformance across six speech codecs and five packet-loss levels is unsupported by the submitted artifact because the full text is heavily corrupted. The experimental section, tables, equations, and methodology descriptions are unreadable; no dataset identity, train/validation split, evaluation metric, codec implementation details, packet-loss simulation parameters, or baseline versions can be verified. This is a load-bearing issue: the paper's contribution is an empirical result, and without a readable experimental protocol the claim cannot be checked. The authors should resubmit a clean, complete PDF with the experiments section fully intact.
- [Experiments (unreadable tables)] Even if the garbled table-like blocks are intended to contain results, the entries are uninterpretable, and no error bars, confidence intervals, or significance tests are visible. The phrase 'consistently outperforms' in the Abstract requires a per-condition comparison with some measure of variability; unless the tables distinguish statistically meaningful gains from negligible differences, the claim is not supported. Please provide readable numeric tables with per-codec and per-packet-loss results, the metric (e.g., EER or accuracy), and the number of trials or seeds.
- [Abstract] The generalization from 'six speech codecs and five levels of packet losses' to 'real-world communication degradation scenarios' needs either a coverage argument or a held-out evaluation. As written, the evidence covers the specific tested codecs and loss levels, not jitter, tandem codecs, background noise, or future encoders. Without an additional held-out degradation condition or an explicit argument that the tested conditions span the space of interest, the broad real-world claim is stronger than the stated experimental menu supports.
minor comments (3)
- [Abstract] The phrase 'first unified framework' should be justified against prior robust-ADD work; the related-work section is currently unreadable, so the novelty claim cannot be checked. Please make the comparison explicit in the resubmission.
- [Full text / availability] No data or code availability statement is visible; for an empirical paper of this type, a public release or a clear availability statement would substantially aid verification.
- [Methods / degradation simulation] The packet-loss simulation should be specified precisely: random versus burst loss, loss rates, and whether loss is applied before or after codec encoding. The current text does not permit this determination at the submitted level of readability.
Circularity Check
No significant circularity: the paper's central claims are empirical architecture evaluations, and no derivation reduces by construction to its fitted inputs or to self-citations.
full rationale
I walked the available derivation chain, which is mostly the abstract plus a heavily corrupted full text. The central claims are that the proposed MGAA framework is the first unified robust ADD framework and that it consistently outperforms state-of-the-art baselines across six speech codecs and five packet-loss levels. These are empirical claims supported by experiments against external baselines, not claims derived from a first-principles theory whose outputs are built into its inputs. The adaptive fusion weights are supervised model parameters learned during training, which is standard fitting rather than a hidden circular loop. No equation could be exhibited in which a predicted quantity equals an input by construction, and no load-bearing self-citation or imported uniqueness theorem is visible. The corrupted full text prevents full inspection of the experimental protocol, but that is a verifiability and completeness concern, not circularity. Under the hard rule that circularity may only be claimed when a specific reduction can be quoted, no circular step is identifiable, so the score is 0.
Assumptions & free parameters
assumptions (3)
- domain assumption Time-frequency representations retain detectable forgery traces after codec compression and packet loss.
- domain assumption The six codecs and five packet-loss levels span the practically relevant degradation space.
- domain assumption Saliency-based adaptive fusion improves class separability without merely overfitting to degradation artifacts.
Cite this review
Pith. "Pith review of Multi-Granularity Adaptive Time-Frequency Attention Framework for Audio Deepfake Detection under Real-World Communication Degradations." pith.science (2026). https://pith.science/paper/5UWXBRSX
@misc{pith2026250801467,
author = {Pith},
title = {Pith review of: Multi-Granularity Adaptive Time-Frequency Attention Framework for Audio Deepfake Detection under Real-World Communication Degradations},
year = {2026},
howpublished = {\url{https://pith.science/paper/5UWXBRSX}},
note = {Machine review of arXiv:2508.01467}
}
read the original abstract
The rise of highly convincing synthetic speech poses a growing threat to audio communications. Although existing Audio Deepfake Detection (ADD) methods have demonstrated good performance under clean conditions, their effectiveness drops significantly under degradations such as packet losses and speech codec compression in real-world communication environments. In this work, we propose the first unified framework for robust ADD under such degradations, which is designed to effectively accommodate multiple types of Time-Frequency (TF) representations. The core of our framework is a novel Multi-Granularity Adaptive Attention (MGAA) architecture, which employs a set of customizable multi-scale attention heads to capture both global and local receptive fields across varying TF granularities. A novel adaptive fusion mechanism subsequently adjusts and fuses these attention branches based on the saliency of TF regions, allowing the model to dynamically reallocate its focus according to the characteristics of the degradation. This enables the effective localization and amplification of subtle forgery traces. Extensive experiments demonstrate that the proposed framework consistently outperforms state-of-the-art baselines across various real-world communication degradation scenarios, including six speech codecs and five levels of packet losses. In addition, comparative analysis reveals that the MGAA-enhanced features significantly improve separability between real and fake audio classes and sharpen decision boundaries. These results highlight the robustness and practical deployment potential of our framework in real-world communication environments.
Forward citations
Cited by 1 Pith paper
-
Audio Deepfake Detection at the First Greeting: "Hi!"
S-MGAA adds pixel-channel enhancement and frequency compensation modules to improve audio deepfake detection on very short, degraded speech inputs.
Reference graph
Works this paper leans on
-
[1]
, " * write output.state after.block = add.period write newline
ENTRY address archivePrefix author booktitle chapter edition editor eid eprint howpublished institution isbn journal key month note number organization pages publisher school series title type volume year label extra.label sort.label short.list INTEGERS output.state before.all mid.sentence after.sentence after.block FUNCTION init.state.consts #0 'before.a...
-
[2]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION word.in bbl.in capitalize " " * FUNCT...
-
[3]
Astrom, H.; Astrom, H.; Spittka, J.; and Vos, K. 2009. RTP payload format and file storage format for silk speech and audio codec. Technical report, Internet Engineering Task Force
work page 2009
-
[4]
Besacier, L.; Mayorga, P.; Bonastre, J.-F.; Fredouille, C.; and Meignier, S. 2003. Overview of compression and packet loss effects in speech biometrics. IEE Proceedings-Vision, Image and Signal Processing, 150(6): 372--376
work page 2003
-
[5]
Bessette, B.; Salami, R.; Lefebvre, R.; Jelinek, M.; Rotola-Pukkila, J.; Vainio, J.; Mikkola, H.; and Jarvinen, K. 2002. The adaptive multirate wideband speech codec (AMR-WB). IEEE Transactions on Speech and Audio Processing, 10(8): 620--636
work page 2002
-
[6]
Bisogni, C.; Loia, V.; Nappi, M.; and Pero, C. 2024. Acoustic features analysis for explainable machine learning-based audio spoofing detection. Computer Vision and Image Understanding, 249: 104145
work page 2024
-
[7]
Blue, L.; Warren, K.; Abdullah, H.; Gibson, C.; Vargas, L.; O'Dell, J.; Butler, K.; and Traynor, P. 2022. Who are you (I really wanna know)? detecting audio DeepFakes through vocal tract reconstruction. In 31st USENIX Security Symposium (USENIX Security 22), 2691--2708
work page 2022
-
[8]
Borodin, K.; Kudryavtsev, V.; Korzh, D.; Efimenko, A.; Mkrtchian, G.; Gorodnichev, M.; and Rogov, O. Y. 2024. AASIST3: KAN-enhanced AASIST speech deepfake detection using SSL features and additional regularization for the ASVspoof 2024 Challenge. In Proc. ASVspoof 2024, 48--55
work page 2024
Show all 75 references
-
[9]
E.; and Nocedal, J
Bottou, L.; Curtis, F. E.; and Nocedal, J. 2018. Optimization methods for large-scale machine learning. SIAM review, 60(2): 223--311
2018
-
[10]
Brewster, T. 2021. Fraudsters cloned company director’s voice in \ 35 million heist, police find. https://www.forbes.com/sites/thomasbrewster/2021/10/14/huge-bank-fraud-uses-deep-fake-voice-tech-to-steal-millions/. Accessed: 2025-05-12
2021
-
[11]
Bruhn, S.; Pobloth, H.; Schnell, M.; Grill, B.; Gibbs, J.; Miao, L.; J \"a rvinen, K.; Laaksonen, L.; Harada, N.; Naka, N.; et al. 2015. Standardization of the new 3GPP EVS codec. In IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 5703--5707
2015
-
[12]
Chakravarty, N.; and Dua, M. 2024. A lightweight feature extraction technique for deepfake audio detection. Multimedia Tools and Applications, 83(26): 67443--67467
2024
-
[13]
Chettri, B. 2023. The clever hans effect in voice spoofing detection. In 2022 IEEE Spoken Language Technology Workshop (SLT), 577--584. IEEE
2023
-
[14]
Cohen, A.; Rimon, I.; Aflalo, E.; and Permuter, H. H. 2022. A study on data augmentation in voice anti-spoofing. Speech Communication, 141: 56--67
2022
-
[15]
Coldewey, D. 2024. Six million fine for robocaller who used ai to clone biden’s voice. https://techcrunch.com/2024/05/23/6m-fine-for-robocaller-who-used-ai-to-clone-bidens-voice/. Accessed: 2025-05-12
2024
-
[16]
Cox, J. 2023. How i broke into a bank account with an ai-generated voice. https://www.vice.com/en/ article/dy7axa/how-i-broke-into-a-bank-account-with-an-ai-generated-voice. Accessed: 2025-05-12
2023
-
[17]
Doan, T.-P.; Nguyen-Vu, L.; Jung, S.; and Hong, K. 2023. BTS-E: Audio Deepfake Detection Using Breathing-Talking-Silence Encoder. In IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 1--5
2023
-
[18]
ETSI. 2024. LTE; 5G; Codec for Immersive Voice and Audio Services - Detailed Algorithmic Description incl. RTP payload format and SDP parameter definitions. https://www.etsi.org/
2024
-
[19]
Frank, J.; and Sch\" o nherr, L. 2021. WaveFake: A Data Set to Facilitate Audio Deepfake Detection. In Proceedings of the Neural Information Processing Systems Track on Datasets and Benchmarks, volume 1
2021
-
[20]
Gerken, T.; and McMahon, L. 2022. Big tech must deal with disinformation or face fines, says eu. https://www.bbc.co.uk/news/technology-61817647. Accessed: 2025-05-12
2022
-
[21]
Goode, B. 2002. Voice over internet protocol (voip). Proceedings of the IEEE, 90(9): 1495--1517
2002
-
[22]
Guo, Y.; Huang, H.; Chen, X.; Zhao, H.; and Wang, Y. 2024. Audio Deepfake Detection With Self-Supervised Wavlm And Multi-Fusion Attentive Classifier. In IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 12702--12706
2024
-
[23]
Hamza, A.; Javed, A. R. R.; Iqbal, F.; Kryvinska, N.; Almadhor, A. S.; Jalil, Z.; and Borghol, R. 2022. Deepfake Audio Detection via MFCC Features Using Machine Learning. IEEE Access, 10: 134018--134028
2022
-
[24]
Hu, J.; Shen, L.; and Sun, G. 2018. Squeeze-and-excitation networks. In Proceedings of the IEEE conference on computer vision and pattern recognition, 7132--7141
2018
-
[25]
Ito, K.; and Johnson, L. 2017. The LJ Speech Dataset. https://keithito.com/LJ-Speech-Dataset/
2017
-
[26]
Jia, X.; De Brabandere, B.; Tuytelaars, T.; and Gool, L. V. 2016. Dynamic filter networks. Advances in neural information processing systems, 29
2016
-
[27]
S.; Lee, B.-J.; Yu, H.-J.; and Evans, N
Jung, J.-w.; Heo, H.-S.; Tak, H.; Shim, H.-j.; Chung, J. S.; Lee, B.-J.; Yu, H.-J.; and Evans, N. 2022. Aasist: Audio anti-spoofing using integrated spectro-temporal graph attention networks. In IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP),...
2022
-
[28]
M.; Sharaf, M.; and Hassan, H
Kanwal, T.; Mahum, R.; AlSalman, A. M.; Sharaf, M.; and Hassan, H. 2024. Fake speech detection using VGGish with attention block. EURASIP Journal on Audio, Speech, and Music Processing, 2024(1): 35
2024
-
[29]
Knibbs, K. 2024. Researchers say the deepfake biden robocall was likely made with tools from ai startup elevenlabs. https://www.wired.com/story/biden-robocall-deepfake-elevenlabs/. Accessed: 2025-05-12
2024
-
[30]
Kong, J.; Kim, J.; and Bae, J. 2020. Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis. Advances in Neural Information Processing Systems, 33: 17022--17033
2020
-
[31]
Kong, Z.; Ping, W.; Huang, J.; Zhao, K.; and Catanzaro, B. 2021. DiffWave: A Versatile Diffusion Model for Audio Synthesis. In International Conference on Learning Representations
2021
-
[32]
Z.; Sotelo, J.; De Brebisson, A.; Bengio, Y.; and Courville, A
Kumar, K.; Kumar, R.; De Boissiere, T.; Gestin, L.; Teoh, W. Z.; Sotelo, J.; De Brebisson, A.; Bengio, Y.; and Courville, A. C. 2019. Melgan: Generative adversarial networks for conditional waveform synthesis. Advances in Neural Information Processing Systems, 32: 14910 -- 14921
2019
-
[33]
Lavrentyeva, G.; Novoselov, S.; Tseren, A.; Volkova, M.; Gorlanov, A.; and Kozlov, A. 2019. STC Antispoofing Systems for the ASVspoof2019 Challenge. Interspeech 2019
2019
-
[34]
Li, M.; Ahmadiadli, Y.; and Zhang, X.-P. 2022. A comparative study on physical and perceptual features for deepfake audio detection. In Proceedings of the 1st International Workshop on Deepfake Detection for Audio Multimedia, 35--41
2022
-
[35]
Li, X.; Wang, W.; Hu, X.; and Yang, J. 2019. Selective Kernel Networks. In 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE
2019
-
[36]
Lin, T.-Y.; Doll \'a r, P.; Girshick, R.; He, K.; Hariharan, B.; and Belongie, S. 2017. Feature pyramid networks for object detection. In Proceedings of the IEEE conference on computer vision and pattern recognition, 2117--2125
2017
-
[37]
Liu, X.; Wang, X.; Sahidullah, M.; Patino, J.; Delgado, H.; Kinnunen, T.; Todisco, M.; Yamagishi, J.; Evans, N.; Nautsch, A.; et al. 2023. Asvspoof 2021: Towards spoofed and deepfake speech detection in the wild. IEEE/ACM Transactions on Audio, Speech, and Language Processing,...
2023
-
[38]
Loshchilov, I.; and Hutter, F. 2017 a . Decoupled weight decay regularization. arXiv preprint arXiv:1711.05101
2017 arXiv
-
[39]
Loshchilov, I.; and Hutter, F. 2017 b . SGDR : Stochastic Gradient Descent with Warm Restarts. In International Conference on Learning Representations
2017
-
[40]
M, S.; Rajput, A.; and M, S. 2024. Classification of Deep Fake Audio Using MFCC Technique. In IEEE International Conference on Information Technology, Electronics and Intelligent Communication Systems (ICITEICS), 1--6
2024
-
[41]
M-AILABS . 2019. The M-AILABS Speech Dataset. https://github.com/imdatceleste/m-ailabs-dataset
2019
-
[42]
M.; and Álvarez, A
Martín-Doñas, J. M.; and Álvarez, A. 2022. The Vicomtech Audio Deepfake Detection System Based on Wav2vec2 for the 2022 ADD Challenge. In IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 9241--9245
2022
-
[43]
Molisch, A. F. 2012. Wireless communications, volume 34. John Wiley & Sons
2012
-
[44]
u ller, N. M.; Kawa, P.; Choong, W. H.; Casanova, E.; G \
M \"u ller, N. M.; Kawa, P.; Choong, W. H.; Casanova, E.; G \"o lge, E.; M \"u ller, T.; Syga, P.; Sperl, P.; and B \"o ttinger, K. 2024. Mlaad: The multi-language audio anti-spoofing dataset. In 2024 International Joint Conference on Neural Networks (IJCNN), 1--7
2024
-
[45]
Reimao, R.; and Tzerpos, V. 2019. For: A dataset for synthetic speech detection. In 2019 International Conference on Speech Technology and Human-Computer Dialogue (SpeD), 1--10
2019
-
[46]
Ren, Y.; Hu, C.; Tan, X.; Qin, T.; Zhao, S.; Zhao, Z.; and Liu, T.-Y. 2021. FastSpeech 2: Fast and High-Quality End-to-End Text to Speech. In International Conference on Learning Representations
2021
-
[47]
G.; and Kinnunen, T
Sahidullah, M.; Shim, H.-j.; Hautam \"a ki, R. G.; and Kinnunen, T. H. 2025. Shortcut Learning in Binary Classifier Black Boxes: Applications to Voice Anti-Spoofing and Biometrics. IEEE Journal of Selected Topics in Signal Processing
2025
-
[48]
Sesia, S.; Toufik, I.; and Baker, M. 2011. Lte-the umts long term evolution: from theory to practice. Wiley
2011
-
[49]
J.; Schuster, M.; Jaitly, N.; Yang, Z.; Chen, Z.; Zhang, Y.; Wang, Y.; Skerrv-Ryan, R.; et al
Shen, J.; Pang, R.; Weiss, R. J.; Schuster, M.; Jaitly, N.; Yang, Z.; Chen, Z.; Zhang, Y.; Wang, Y.; Skerrv-Ryan, R.; et al. 2018. Natural tts synthesis by conditioning wavenet on mel spectrogram predictions. In IEEE International Conference on Acoustics, Speech and Signal Pro...
2018
-
[50]
Shi, H.; Shi, X.; and Dogan, S. 2024. Speech inpainting based on multi-layer long short-term memory networks. Future Internet, 16(2): 63
2024
-
[51]
Shi, H.; Shi, X.; Dogan, S.; Alzubi, S.; Huang, T.; and Zhang, Y. 2025. Benchmarking Audio Deepfake Detection Robustness in Real-world Communication Scenarios. arXiv preprint arXiv:2504.12423. Accepted by EUSIPCO 2025
2025 arXiv
-
[52]
Shih, T.-H.; Yeh, C.-Y.; and Chen, M.-S. 2024. Does Audio Deepfake Detection Rely on Artifacts? In ICASSP 2024-2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 12446--12450. IEEE
2024
-
[53]
Shim, H.-j.; Gonzalez Hautam \"a ki, R.; Sahidullah, M.; and Kinnunen, T. 2023. How to Construct Perfect and Worse-than-Coin-Flip Spoofing Countermeasures: A Word of Warning on Shortcut Learning. In Proc. Interspeech 2023, 785--789
2023
-
[54]
Tak, H.; Jung, J.-W.; Patino, J.; Kamble, M.; Todisco, M.; and Evans, N. 2021 a . End-to-End Spectro-Temporal Graph Attention Networks for Speaker Verification Anti-Spoofing and Speech Deepfake Detection. In ASVSPOOF 2021, Automatic Speaker Verification and Spoofing Countermea...
2021
-
[55]
Tak, H.; Patino, J.; Todisco, M.; Nautsch, A.; Evans, N.; and Larcher, A. 2021 b . End-to-end anti-spoofing with rawnet2. In IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 6369--6373
2021
-
[56]
Tak, H.; Todisco, M.; Wang, X.; Jung, J.-w.; Yamagishi, J.; and Evans, N. 2022. Automatic speaker verification spoofing and deepfake detection using wav2vec 2.0 and data augmentation. arXiv preprint arXiv:2202.12233
2022 arXiv
-
[57]
Todisco, M.; Delgado, H.; and Evans, N. 2017. Constant Q cepstral coefficients: A spoofing countermeasure for automatic speaker verification. Computer Speech & Language, 45: 516--535
2017
-
[58]
Todisco, M.; Wang, X.; Vestman, V.; Sahidullah, M.; Delgado, H.; Nautsch, A.; Yamagishi, J.; Evans, N.; Kinnunen, T.; and Lee, K. A. 2019. ASVspoof 2019: Future horizons in spoofed and fake audio detection. arXiv preprint arXiv:1904.05441
2019 arXiv
-
[59]
Valenti, M.; Squartini, S.; Diment, A.; Parascandolo, G.; and Virtanen, T. 2017. A convolutional neural network approach for acoustic scene classification. In 2017 International Joint Conference on Neural Networks (IJCNN), 1547--1554. IEEE
2017
-
[60]
Valin, J.-M. 2016. Speex: A free codec for free speech. arXiv preprint arXiv:1602.08668
2016 arXiv
-
[61]
B.; and Vos, K
Valin, J.-M.; Maxwell, G.; Terriberry, T. B.; and Vos, K. 2016. High-quality, low-delay music coding in the opus codec. arXiv preprint arXiv:1602.04845
2016 arXiv
-
[62]
Van Den Oord, A.; Dieleman, S.; Zen, H.; Simonyan, K.; Vinyals, O.; Graves, A.; Kalchbrenner, N.; Senior, A.; Kavukcuoglu, K.; et al. 2016. Wavenet: A generative model for raw audio. arXiv preprint arXiv:1609.03499, 12
2016 arXiv
-
[63]
Van der Maaten, L.; and Hinton, G. 2008. Visualizing data using t-SNE. Journal of Machine Learning Research, 9(11): 2579--2605
2008
-
[64]
Wang, X.; Delgado, H.; Tak, H.; Jung, J.-w.; Shim, H.-j.; Todisco, M.; Kukanov, I.; Liu, X.; Sahidullah, M.; Kinnunen, T.; et al. 2024. ASVspoof 5: Crowdsourced speech data, deepfakes, and adversarial attacks at scale. arXiv preprint arXiv:2408.08739
2024 arXiv
-
[65]
Wang, X.; Girshick, R.; Gupta, A.; and He, K. 2018. Non-local neural networks. In Proceedings of the IEEE conference on computer vision and pattern recognition, 7794--7803
2018
-
[66]
Wang, X.; and Yamagishi, J. 2021. Investigating self-supervised front ends for speech spoofing countermeasures. arXiv preprint arXiv:2111.07725
2021 arXiv
-
[67]
M.; Qadri, S
Wani, T. M.; Qadri, S. A. A.; Comminiello, D.; and Amerini, I. 2024 a . Detecting audio deepfakes: Integrating CNN and BiLSTM with multi-feature concatenation. In Proceedings of the 2024 ACM Workshop on Information Hiding and Multimedia Security, 271--276
2024
-
[68]
M.; Qadri, S
Wani, T. M.; Qadri, S. A. A.; Comminiello, D.; and Amerini, I. 2024 b . Detecting audio deepfakes: Integrating CNN and BiLSTM with multi-feature concatenation. In Proceedings of the 2024 ACM Workshop on Information Hiding and Multimedia Security, 271--276
2024
-
[69]
Woo, S.; Park, J.; Lee, J.-Y.; and Kweon, I.-S. 2018. CBAM: Convolutional Block Attention Module. In European Conference on Computer Vision, 3--19. European Conference on Computer Vision
2018
-
[70]
Yadav, S.; and Rai, A. 2020. Frequency and temporal convolutional attention for text-independent speaker recognition. In ICASSP 2020-2020 IEEE international conference on acoustics, speech and signal processing (ICASSP), 6794--6798. IEEE
2020
-
[71]
Yamamoto, R.; Song, E.; and Kim, J.-M. 2020. Parallel WaveGAN: A fast waveform generation model based on generative adversarial networks with multi-resolution spectrogram. In IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 6199--6203
2020
-
[72]
Yi, J.; Fu, R.; Tao, J.; Nie, S.; Ma, H.; Wang, C.; Wang, T.; Tian, Z.; Bai, Y.; Fan, C.; et al. 2022. Add 2022: the first audio deep synthesis detection challenge. In IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 9216--9220
2022
-
[73]
Yu, N.; Chen, L.; Leng, T.; Chen, Z.; and Yi, X. 2024. An explainable deepfake of speech detection method with spectrograms and waveforms. Journal of Information Security and Applications, 81: 103720
2024
-
[74]
Zhang, Q.; Wen, S.; and Hu, T. 2024. Audio deepfake detection with self-supervised xls-r and sls classifier. In Proceedings of the 32nd ACM International Conference on Multimedia, 6765--6773
2024
-
[75]
Zhu, Y.; Koppisetti, S.; Tran, T.; and Bharaj, G. 2024. Slim: Style-linguistics mismatch model for generalized audio deepfake detection. Advances in Neural Information Processing Systems, 37: 67901--67928
2024
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.