REVIEW 3 major objections 7 minor 43 references
IndieFake Dataset: A Benchmark Dataset for Audio Deepfake Detection
T0 review · 3 major / 7 minor · reviewed 2026-08-15 · deepseek-v4-flash
Pith's one-line read The IndieFake Dataset, 27.17 hours of bonafide and deepfake English speech from 50 Indian speakers, is introduced to show that a small balanced accent-specific dataset can train audio deepfake detectors better than the much larger…
desk verdict A genuinely useful new dataset for Indian-accented audio deepfake detection, but the headline superiority claims outrun the evidence; the dataset itself is worth engaging with. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the IndieFake Dataset (IFD) itself: 27.17 hours of five-second audio clips from 50 English-speaking Indian speakers, with 8,164 bonafide samples drawn from Creative-Commons YouTube speech and 11,396 deepfake samples generated by TTS and voice-cloning services under three scenarios (hypothetical transcripts, transcripts of the same speaker, and transcripts of another listed speaker). The dataset's design—balanced bonafide/deepfake counts, speaker-level metadata, and a subject-independent 80:20 split—is what carries the argument, because it lets the authors attribute differences in detector performance to accent and content diversity rather than speaker overlap or class imbalance.
What would settle it
Train identical baseline models on ASVspoof21 (DF) and IFD with equal epochs, equal compute, or matched convergence criteria, then evaluate on the same held-out test sets; if the IFD-trained models no longer achieve lower equal error rates than the ASVspoof21-trained models, the paper's superior-training-resource claim is not supported.
Extended reading notes
Core claim
The central claim is that IFD outperforms ASVspoof21 (DF) as a training resource and is more challenging than In-The-Wild as a benchmark, despite its smaller scale. Training five baseline detectors (LFCC-LCNN, MFCC-LCNN, LFCC-MesoNet, MFCC-MesoNet, RawNet3) on IFD and testing on ITW gives lower EER in most cases than training the same baselines on ASVspoof21 (DF). Conversely, models trained on ASVspoof21 (DF) show higher EER and lower accuracy on IFD than on ITW, indicating that IFD contains harder or less familiar spoofing conditions. The dataset also contributes speaker-level characterization and a subject-independent train-test split, which the authors say are missing from ASVspoof21 (DF).
Load-bearing premise
The claim that IFD is a better training resource assumes that comparing error rates after 10 training epochs on ASVspoof21 (DF) and 50 epochs on IFD is a fair measure of dataset quality.
Editorial extensions
If this is right
- Audio deepfake detectors can be trained more cheaply on a small balanced dataset with Indian English accents than on a much larger imbalanced set, if the reported EER advantage holds.
- Starting from scratch on IFD appears preferable to initializing from ASVspoof21 (DF) pre-trained weights, since the paper observes performance drops after such fine-tuning.
- IFD can serve as a more stringent evaluation benchmark than In-The-Wild for detectors meant to work in Indian English contexts.
- The subject-independent split means a detector that does well on IFD has not seen the same speaker's voice in training, which is closer to real-world deployment against new voices.
Reading between the lines
- The reported advantage may depend on the unequal training budgets, so a matched-budget comparison would tell whether the dataset itself or the extra epochs drive the result.
- The cross-speaker transcript scenario isolates content-transfer attacks, so IFD could be used to study a class of spoofing that most existing benchmarks do not separate out.
- The same collection recipe—balanced accent-specific speakers plus scenario-driven deepfake generation—could be applied to other under-represented accents, which would show whether accent diversity is the active ingredient in detection performance.
- Because the dataset is public, independent groups can reproduce the baseline table and extend it to newer detectors without generating new deepfake audio themselves.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces the IndieFake Dataset (IFD), a new benchmark for audio deepfake detection containing 27.17 hours of bonafide and deepfake English speech from 50 Indian speakers. The authors describe a subject-independent train-test split, speaker-level characterization, and four generation scenarios. They evaluate five baselines (LFCC-LCNN, MFCC-LCNN, LFCC-MesoNet, MFCC-MesoNet, RawNet3) under multiple training and evaluation configurations across IFD, ASVspoof21 (DF), and In-The-Wild (ITW), and they claim that IFD outperforms ASVspoof21 (DF) as a training resource and is more challenging than ITW as a benchmark.
Significance. If the comparative claims were rigorously established, the dataset would fill a real gap: South-Asian-accented English is underrepresented in existing audio deepfake detection benchmarks, and the speaker-level metadata plus balanced design would be useful for generalization studies. The paper's strengths include the public release with documentation, the subject-independent split, the inclusion of multiple generation scenarios, and the ablation study of normalization variants for raw-audio models. However, the headline claims of superiority over ASVspoof21 (DF) and ITW are not supported by the reported evidence as it stands, because the training budgets are not matched and the result patterns are mixed across baselines.
major comments (3)
- [Section III.C and Table IV] The claim that IFD outperforms ASVspoof21 (DF) as a training resource is not supported by the reported results. Baselines are trained for 10 epochs on ASVspoof21 (DF) and 50 epochs on IFD, so the comparison conflates dataset content with optimization budget; a matched-updates or matched-compute comparison is needed. Moreover, Table IV shows that LFCC-LCNN (0.233 vs. 0.3718) and MFCC-LCNN (0.34 vs. 0.389) achieve lower EER after training on ASVspoof21 (DF), so only 3 of 5 baselines favor IFD. The abstract and conclusion state an unqualified superiority, while Section III.D says 'consistently better ... for most baselines,' which is internally inconsistent. No error bars or significance tests are provided, so the observed differences could reverse under a controlled protocol.
- [Section III.D and Table II, columns E/F] The secondary claim that IFD is more challenging than ITW also rests on a 3-of-5 pattern: MFCC-MesoNet and RawNet3 have lower EER on IFD (0.469 vs. 0.698 and 0.402 vs. 0.497, respectively), and the text's 'in most cases, the EER is higher' is not matched by an unqualified statement in the abstract and conclusion. This claim needs either a matched evaluation protocol and statistical support or a careful rewording to reflect the actual per-baseline results.
- [Section II.B] The dataset is first described as 'a subject dependent dataset' and then the splitting strategy is called 'subject independent splitting approach.' These labels are confusing; if the train and test sets contain disjoint speakers, the design is subject-independent, and the wording should be corrected. The distinction matters because the paper's generalization claims depend on it.
minor comments (7)
- [Abstract] There are grammar and punctuation errors: 'Advancements ... offers benefits' should be 'offer benefits,' and 'worlds population' should be 'world's population.'
- [References] References [9] and [17] are the same Whisper-features paper, and references [10] and [18] are both the Tacotron paper. Duplicate entries should be merged.
- [Section III.C] The phrase 'an average batch size of 64' is odd; batch size is a fixed hyperparameter, so 'a batch size of 64' would be clearer.
- [Section III.D] The text states that IFD is 'one-sixth of the size of ASVspoof21 (DF),' but the sample counts in Section III.B imply a ratio closer to 1/30; please correct this quantitative comparison.
- [Table I] The table entries such as 'Complete Speaker-13' and 'Bonafide Only Speaker-45' are difficult to parse; a clearer layout separating speaker ID, type, and counts would improve readability.
- [Figures 3 and 4] The architecture diagrams for ICDD and ISCSE are not referenced or explained in the body text, and the convolution parameters in the diagrams are unclear; consider adding a textual description or moving the figures to an appendix.
- [Section IV] The conclusion uses 'proves' for empirical findings; 'suggests' or 'is consistent with' would be more appropriate given the mixed per-baseline results.
Circularity Check
No circularity: the paper is an empirical dataset benchmark with no fitted parameters and no derivation chain; the uncontrolled training budget is a fairness concern, not a circularity concern.
full rationale
This is an empirical dataset paper. The central claims are (i) that training on IFD yields lower EER than training on ASVspoof21 (DF) when both are tested on ITW (Table IV), and (ii) that IFD is a more challenging test set than ITW when both are evaluated after training on ASVspoof21 (DF) (Table II, columns E and F). Neither claim is derived from an equation, fitted parameter, or self-citation. The evaluation uses external benchmark test sets (ITW and ASVspoof21) whose labels are not constructed from the paper's own outputs, so the results are not true by definition. The skeptical concern that the comparison is not controlled for training budget (10 epochs for ASVspoof21 versus 50 epochs for IFD) is a legitimate experimental-validity criticism, but it is not circularity: the EER values are measured, not deduced from the input. Likewise, the mixed results across baselines in Table IV concern the strength of the evidence, not whether the claim reduces to its inputs. The paper contains no self-citation chain that forces a conclusion, no ansatz smuggled in by citation, and no renamed known result. Therefore the appropriate circularity score is 0.
Assumptions & free parameters
assumptions (4)
- domain assumption Bonafide audio from YouTube under Creative Commons is genuine human speech.
- domain assumption Commercial TTS services (Amazon Polly, Play.ht, ElevenLabs) generate deepfake audio representative of real-world attacks.
- domain assumption Cross-dataset EER after training on a dataset is a valid measure of that dataset's training quality.
- domain assumption The 50 selected speakers are representative of Indian English accents and the subject-independent split prevents speaker leakage.
Cite this review
Pith. "Pith review of IndieFake Dataset: A Benchmark Dataset for Audio Deepfake Detection." pith.science (2026). https://pith.science/paper/KICBL2IA
@misc{pith2026250619014,
author = {Pith},
title = {Pith review of: IndieFake Dataset: A Benchmark Dataset for Audio Deepfake Detection},
year = {2026},
howpublished = {\url{https://pith.science/paper/KICBL2IA}},
note = {Machine review of arXiv:2506.19014}
}
read the original abstract
Advancements in audio deepfake technology offers benefits like AI assistants, better accessibility for speech impairments, and enhanced entertainment. However, it also poses significant risks to security, privacy, and trust in digital communications. Detecting and mitigating these threats requires comprehensive datasets. Existing datasets lack diverse ethnic accents, making them inadequate for many real-world scenarios. Consequently, models trained on these datasets struggle to detect audio deepfakes in diverse linguistic and cultural contexts such as in South-Asian countries. Ironically, there is a stark lack of South-Asian speaker samples in the existing datasets despite constituting a quarter of the worlds population. This work introduces the IndieFake Dataset (IFD), featuring 27.17 hours of bonafide and deepfake audio from 50 English speaking Indian speakers. IFD offers balanced data distribution and includes speaker-level characterization, absent in datasets like ASVspoof21 (DF). We evaluated various baselines on IFD against existing ASVspoof21 (DF) and In-The-Wild (ITW) datasets. IFD outperforms ASVspoof21 (DF) and proves to be more challenging compared to benchmark ITW dataset. The complete dataset, along with documentation and sample reference clips, is publicly accessible for research use on project website.
Figures
Figures from the paper (1 more)
Reference graph
Works this paper leans on
-
[1]
Mohammadi, Mohsen & Sadegh Mohammadi, H. R.. (2017). Robust features fusion for text independent speaker verification enhancement in noisy environments. 10.1109/IranianCEE.2017.7985357
arXiv 2017
-
[2]
D. A. Reynolds and R. C. Rose, ”Robust text-independent speaker identification using Gaussian mixture speaker models,” in IEEE Transactions on Speech and Audio Processing, vol. 3, no. 1, pp. 72-83, Jan. 1995, doi: 10.1109/89.365379. keywords: Robustness;Telephony;Speech analysis;Speaker recognition;Spectral shape;Databases;Data mining;Loudspeakers;Degradation,
-
[3]
L. R. Rabiner, ”A tutorial on hidden Markov models and selected applications in speech recognition,” in Proceedings of the IEEE, vol. 77, no. 2, pp. 257-286, Feb. 1989, doi: 10.1109/5.18626. keywords: Tutorial;Hidden Markov models;Speech recognition,
doi:10.1109/5.18626 1989
-
[4]
Jung, Jee-Weon & Kim, Youjin & Heo, Hee-Soo & Lee, Bong-Jin & Kwon, Youngki & Chung, Joon Son. (2022). Pushing the limits of raw waveform speaker recognition. 2228-2232. 10.21437/Interspeech.2022- 126
-
[5]
Wu, Xiang & He, Ran & Sun, Zhenan & Tan, Tieniu. (2018). A Light CNN for Deep Face Representation with Noisy Labels. IEEE Transactions on Information Forensics and Security. 13. 1-1. 10.1109/TIFS.2018.2833032
arXiv 2018
-
[6]
D. Afchar, V . Nozick, J. Yamagishi and I. Echizen, ”MesoNet: a Com- pact Facial Video Forgery Detection Network,” 2018 IEEE International Workshop on Information Forensics and Security (WIFS), Hong Kong, China, 2018, pp. 1-7, doi: 10.1109/WIFS.2018.8630761. keywords: Face;Training;Convolutional codes;Forgery;Deep learning;Decoding,
arXiv 2018
-
[7]
M ¨uller, Nicolas & Czempin, Pavel & Diekmann, Franziska & Froghyar, Adam & B ¨ottinger, Konstantin. (2022). Does Audio Deepfake Detection Generalize?. 2783-2787. 10.21437/Interspeech.2022-108
-
[8]
Yamagishi, Junichi & Wang, Xin & Todisco, Massimiliano & Sahidullah, Md & Patino, Jose & Nautsch, Andreas & Liu, Xuechen & Lee, Kong Aik & Kinnunen, Tomi & Evans, Nicholas & Delgado, H ´ector. (2021). ASVspoof 2021: accelerating progress in spoofed and deepfake speech detection. 47-54. 10.21437/ASVSPOOF.2021-8
Show all 43 references
-
[11]
Kameoka, Hirokazu & Kaneko, Takuhiro & Tanaka, Kou & Hojo, Nobukatsu. (2018). StarGAN-VC: non-parallel many-to-many V oice Conversion Using Star Generative Adversarial Networks. 266-273. 10.1109/SLT.2018.8639535
2018
-
[12]
Zen, Heiga & Dang, Viet & Clark, Rob & Zhang, Yu & Weiss, Ron & Jia, Ye & Chen, Zhifeng. (2019). LibriTTS: A Corpus Derived from LibriSpeech for Text-to-Speech. 1526-1530. 10.21437/Interspeech.2019- 2441
2019 doi
-
[13]
Lorenzo-Trueba, Jaime & Yamagishi, Junichi & Toda, Tomoki & Saito, Daisuke & Villavicencio, Fernando & Kinnunen, Tomi & Ling, Zhen- Hua. (2018). The V oice Conversion Challenge 2018: Promoting Devel- opment of Parallel and Nonparallel Methods
2018
-
[14]
The lj speech dataset,
K. Ito and L. Johnson, “The lj speech dataset,” https://keithito.com/LJ- Speech-Dataset/, 2017
2017
-
[15]
Veaux, Christophe; Yamagishi, Junichi; MacDonald, Kirsten. (2017). CSTR VCTK Corpus: English Multi-speaker Corpus for CSTR V oice Cloning Toolkit, [sound]. University of Edinburgh. The Centre for Speech Technology Research (CSTR). https://doi.org/10.7488/ds/1994
2017 doi
-
[16]
Reimao, Ricardo & Tzerpos, Vassilios. (2019). FoR: A Dataset for Synthetic Speech Detection. 1-10. 10.1109/SPED.2019.8906599
2019
-
[17]
Kawa, Piotr & Plata, Marcin & Czuba, Michał & Szyma ´nski, Piotr & Syga, Piotr. (2023). Improved DeepFake Detection Using Whisper Features. 4009-4013. 10.21437/Interspeech.2023-1537
2023 doi
-
[18]
& Stanton, Daisy & Weiss, Ron & Jaitly, Navdeep & Yang, Zongheng & Xiao, Ying & Chen, Zhifeng & Bengio, Samy & Le, Quoc & Agiomyrgiannakis, Yannis & Clark, Rob & Saurous, Rif
Wang, Yuxuan & Skerry-Ryan, R.J. & Stanton, Daisy & Weiss, Ron & Jaitly, Navdeep & Yang, Zongheng & Xiao, Ying & Chen, Zhifeng & Bengio, Samy & Le, Quoc & Agiomyrgiannakis, Yannis & Clark, Rob & Saurous, Rif. (2017). Tacotron: Towards End-to-End Speech Synthesis. 4006-4010. 10...
2017 doi
-
[19]
Shen, Jonathan & Pang, Ruoming & Weiss, Ron & Schuster, Mike & Jaitly, Navdeep & Yang, Zongheng & Chen, Zhifeng & Zhang, Yu & Wang, Yuxuan & Skerrv-Ryan, Rj & Saurous, Rif & Agiomvrgiannakis, Yannis. (2018). Natural TTS Synthesis by Conditioning Wavenet on MEL Spectrogram Pred...
2018
-
[20]
M ¨uller, Nicolas & Sperl, Philip & B ¨ottinger, Konstantin. (2023). Complex-valued neural networks for voice anti-spoofing. 3814-3818. 10.21437/Interspeech.2023-901
2023 doi
-
[21]
J., & Puckette, M
Donahue, C., McAuley, J. J., & Puckette, M. S. (2018). Synthesizing Audio with Generative Adversarial Networks. CoRR, abs/1802.04208. http://arxiv.org/abs/1802.04208
2018 arXiv
-
[23]
& Courville, Aaron
Kumar, Kundan & Kumar, Rithesh & Boissiere, Thibault & Gestin, Lucas & Teoh, Wei & Sotelo, Jos´e & Brebisson, Alexandre & Bengio, Y . & Courville, Aaron. (2019). MelGAN: Generative Adversarial Networks for Conditional Waveform Synthesis
2019
-
[24]
Jaehyeon, Kim & Kong, Jungil & Son, Juhee. (2021). Conditional Variational Autoencoder with Adversarial Learning for End-to-End Text- to-Speech
2021
-
[25]
Miao, Chenfeng & Zhu, Qingying & Chen, Minchuan & Ma, Jun & Wang, Shaojun & Xiao, Jing. (2024). EfficientTTS 2: Variational End- to-End Text-to-Speech Synthesis and V oice Conversion. IEEE/ACM Transactions on Audio, Speech, and Language Processing. PP. 1-13. 10.1109/TASLP.2024.3369528
2024
-
[26]
Qian, Kaizhi & Zhang, Yang & Chang, Shiyu & Yang, Xuesong & Hasegawa-Johnson, Mark. (2019). AutoVC: Zero-Shot V oice Style Transfer with Only Autoencoder Loss
2019
-
[27]
Pariente, Manuel & Cornell, Samuele & Deleforge, Antoine & Vincent, Emmanuel. (2020). Filterbank Design for End-to-end Speech Separation. 6364-6368. 10.1109/ICASSP40776.2020.9053038
2020
-
[28]
Guha Roy, Abhijit & Navab, Nassir & Wachinger, Christian. (2018). Concurrent Spatial and Channel Squeeze & Excitation in Fully Convo- lutional Networks
2018
-
[29]
Wu, Yuxin & He, Kaiming. (2018). Group Normalization
2018
-
[30]
Ulyanov, Dmitry & Vedaldi, Andrea & Lempitsky, Victor. (2016). Instance Normalization: The Missing Ingredient for Fast Stylization
2016
-
[31]
Capes, Tim & Coles, Paul & Conkie, Alistair & Golipour, Ladan & Hadjitarkhani, Abie & Hu, Qiong & Huddleston, Nancy & Hunt, Melvyn & Li, Jiangchuan & Neeracher, Matthias & Prahallad, Kishore & Raitio, Tuomo & Rasipuram, Ramya & Townsend, Greg & Williamson, Becci & Winarsky, Da...
2017 doi
-
[32]
Felix, Shubham & Kumar, Sumer & Veeramuthu, A.. (2018). A Smart Personal AI Assistant for Visually Impaired People. 1245-1250. 10.1109/ICOEI.2018.8553750
2018
-
[33]
Ho, Derek kwun-hong. (2017). V oice-controlled virtual assistants for the older people with visual impairment. Eye. 32. 10.1038/eye.2017.165
2017 doi
-
[34]
Mart ´ın-Do˜nas, Juan & ´Alvarez, Aitor. (2022). The Vicomtech Audio Deepfake Detection System Based on Wav2vec2 for the 2022 ADD Challenge. 9241-9245. 10.1109/ICASSP43922.2022.9747768
2022
-
[35]
Sree, Katamneni & Rattani, Ajita. (2023). MIS-A V oiDD: Modality Invariant and Specific Representation for Audio-Visual Deepfake De- tection. 10.1109/ICMLA58977.2023.00207
2023
-
[36]
oord, Aaron & Dieleman, Sander & Zen, Heiga & Simonyan, Karen & Vinyals, Oriol & Graves, Alex & Kalchbrenner, Nal & Senior, Andrew & Kavukcuoglu, Koray. (2016). WaveNet: A Generative Model for Raw Audio
2016
-
[37]
Wang, Wenfu & Xu, Shuang & Xu, Bo. (2016). First Step Towards End- to-End Parametric TTS Synthesis: Generating Spectral Parameters with Neural Attention. 2243-2247. 10.21437/Interspeech.2016-134
2016 doi
-
[38]
Oloko-oba, Mustapha & T.S, Ibiyemi & Samuel, Osagie. (2016). Text-to- Speech Synthesis Using Concatenative Approach. International Journal of Trend in Research and Development. 3. 559-462
2016
-
[39]
Park, Daniel & Chan, William & Zhang, Yu & Chiu, Chung-Cheng & Zoph, Barret & Cubuk, Ekin & Le, Quoc. (2019). SpecAugment: A Simple Data Augmentation Method for Automatic Speech Recognition. 2613-2617. 10.21437/Interspeech.2019-2680
2019 doi
-
[40]
Ko, Tom & Peddinti, Vijayaditya & Povey, Daniel & Khudanpur, Sanjeev. (2015). Audio augmentation for speech recognition. 3586-3589. 10.21437/Interspeech.2015-711
2015 doi
-
[41]
Verma, Prateek & Smith, Julius. (2018). Neural Style Transfer for Audio Spectograms
2018
-
[42]
Ratnarajah, Anton & Ghosh, Sreyan & Kumar, Sonal & Chiniya, Purva & Manocha, Dinesh. (2024). A V-RIR: Audio-Visual Room Impulse Response Estimation
2024
-
[43]
Kim, Jong & Salamon, Justin & Li, Peter & Bello, Juan. (2018). CREPE: A Convolutional Representation for Pitch Estimation. Acoustics, Speech, and Signal Processing, 1988. ICASSP-88., 1988 International Confer- ence on
2018
-
[44]
Moliner, Eloi & Fierro, Leonardo & Wright, Alec & H ¨am¨al¨ainen, Matti & V ¨alim¨aki, Vesa. (2024). Noise Morphing for Audio Time Stretching. Signal Processing Letters, IEEE. 31. 1144-1148. 10.1109/LSP.2024.3386118
2024
-
[45]
Abedi, Mostafa & Pourmohammad, Ali. (2015). Super-Gaussian non- stationary audio noises sparse representation. 152-155. 10.1109/INNO- V ATIONS.2015.7381531
2015
-
[46]
Available at: https://github.com/AI4Bharat/ NPTEL2020-Indian-English-Speech-Dataset/tree/master
AI4Bharat, AI4BHARAT/NPTEL2020-Indian-English-speech- dataset: NPTEL2020: Speech2Text dataset for Indian-English accent, GitHub. Available at: https://github.com/AI4Bharat/ NPTEL2020-Indian-English-Speech-Dataset/tree/master
Reviewed August 15, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.