{"id":"c37c3301-bba8-41fe-875b-ca7f812734ce","arxiv_id":"2506.19014","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"The IndieFake dataset adds 19,560 Indian-English audio clips (8,164 bonafide, 11,396 deepfake) from 50 speakers, and baseline tests show it is a harder benchmark than In-The-Wild for detection models.","lead":"This paper introduces IndieFake, a new benchmark dataset of 27.17 hours of genuine and AI-generated English speech from 50 Indian speakers. It aims to close a gap in audio deepfake detection for South Asian accents and tests standard detectors against it.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Claim that IFD outperforms ASVspoof21 is supported by only 3/5 baselines in Table IV, and the comparison is not controlled for training budget.","rationale":"The reader's weakest assumption, that comparing 10 epochs on ASVspoof21 with 50 epochs on IFD is unfair, is a real confound. However, the more direct problem is that Table IV, the only head-to-head evidence for the 'better training resource' claim, is mixed even before controlling for epochs: ASVspoof21 yields lower EER on LFCC-LCNN and MFCC-LCNN. The abstract and conclusion overstate this as 'IFD outperforms ASVspoof21 (DF)', while Section III.D admits 'for most baselines' and even says 'consistently better' in the same sentence. A matched-budget rerun with paired significance testing would settle whether the 3-of-5 majority is reliable or an artifact of the uncontrolled protocol. The 'more challenging than ITW' claim is secondary and depends on the same ASVspoof21-trained comparison, where EER is higher on IFD for only 3 of 5 baselines, so it needs the same qualification. The dataset itself appears to be a meaningful contribution with internally consistent counts and a subject-independent split, and the ablation study is plausible, so the paper should not be rejected. The CONDITIONAL verdict already captures the need to qualify the headline; I would not change it. Hence UNCHANGED.","tokens_in":11026,"tokens_out":9425,"duration_ms":84490,"concrete_test":"Recompute Table IV with a matched optimization budget: train each of the five baselines on ASVspoof21 and IFD for the same number of parameter updates (e.g., 773k, IFD's current budget, and 1.25M, ASVspoof21's current budget), using early stopping on a common held-out validation set, and report per-baseline EER with a paired bootstrap or Wilcoxon signed-rank test over test utterances. If ASVspoof21 wins on 2 or more of the 5 baselines at matched budget, or if the IFD advantage is not statistically significant, the abstract's unqualified 'outperforms' should be weakened to 'comparable' or 'better on some baselines'. As a secondary check, repeat the ITW-vs-IFD challenge comparison (Table II columns E/F) after matching input length and score normalization; if the 3/5 EER differences are not robust, soften the 'more challenging' claim as well.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that IFD is a better training resource than ASVspoof21 (DF) rests on Table IV, which compares EER on ITW after training on each dataset. Two problems make this claim insecure. First, the optimization budget is not matched: Section III.C fixes ASVspoof21 at 10 epochs and IFD at 50 epochs, so the comparison conflates dataset content with training effort. Second, even taken at face value, Table IV is mixed: LFCC-LCNN (EER 0.233 vs 0.3718) and MFCC-LCNN (0.34 vs 0.389) favor ASVspoof21, while only LFCC-MesoNet, MFCC-MesoNet, and RawNet3 favor IFD. The abstract and conclusion state an unqualified 'IFD outperforms ASVspoof21', but the body says only 'for most baselines' and even calls the results 'consistently better ... for most baselines', an internal inconsistency. If the comparison were controlled for updates and significance-tested, the claimed advantage could disappear or reverse; the paper provides no error bars or significance test on any EER difference. The secondary 'more challenging than ITW' claim relies on the same uncontrolled protocol (Table II, columns E/F), with EER higher on IFD in only 3 of 5 baselines. Because the paper's headline is exactly this superiority claim, the weakness is load-bearing.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces the IndieFake Dataset (IFD), a new benchmark for audio deepfake detection containing 27.17 hours of bonafide and deepfake English speech from 50 Indian speakers. The authors describe a subject-independent train-test split, speaker-level characterization, and four generation scenarios. They evaluate five baselines (LFCC-LCNN, MFCC-LCNN, LFCC-MesoNet, MFCC-MesoNet, RawNet3) under multiple training and evaluation configurations across IFD, ASVspoof21 (DF), and In-The-Wild (ITW), and they claim that IFD outperforms ASVspoof21 (DF) as a training resource and is more challenging than ITW as a benchmark.","tokens_in":11292,"tokens_out":3796,"duration_ms":35417,"significance":"If the comparative claims were rigorously established, the dataset would fill a real gap: South-Asian-accented English is underrepresented in existing audio deepfake detection benchmarks, and the speaker-level metadata plus balanced design would be useful for generalization studies. The paper's strengths include the public release with documentation, the subject-independent split, the inclusion of multiple generation scenarios, and the ablation study of normalization variants for raw-audio models. However, the headline claims of superiority over ASVspoof21 (DF) and ITW are not supported by the reported evidence as it stands, because the training budgets are not matched and the result patterns are mixed across baselines.","major_comments":[{"comment":"The claim that IFD outperforms ASVspoof21 (DF) as a training resource is not supported by the reported results. Baselines are trained for 10 epochs on ASVspoof21 (DF) and 50 epochs on IFD, so the comparison conflates dataset content with optimization budget; a matched-updates or matched-compute comparison is needed. Moreover, Table IV shows that LFCC-LCNN (0.233 vs. 0.3718) and MFCC-LCNN (0.34 vs. 0.389) achieve lower EER after training on ASVspoof21 (DF), so only 3 of 5 baselines favor IFD. The abstract and conclusion state an unqualified superiority, while Section III.D says 'consistently better ... for most baselines,' which is internally inconsistent. No error bars or significance tests are provided, so the observed differences could reverse under a controlled protocol.","section":"Section III.C and Table IV"},{"comment":"The secondary claim that IFD is more challenging than ITW also rests on a 3-of-5 pattern: MFCC-MesoNet and RawNet3 have lower EER on IFD (0.469 vs. 0.698 and 0.402 vs. 0.497, respectively), and the text's 'in most cases, the EER is higher' is not matched by an unqualified statement in the abstract and conclusion. This claim needs either a matched evaluation protocol and statistical support or a careful rewording to reflect the actual per-baseline results.","section":"Section III.D and Table II, columns E/F"},{"comment":"The dataset is first described as 'a subject dependent dataset' and then the splitting strategy is called 'subject independent splitting approach.' These labels are confusing; if the train and test sets contain disjoint speakers, the design is subject-independent, and the wording should be corrected. The distinction matters because the paper's generalization claims depend on it.","section":"Section II.B"}],"minor_comments":[{"comment":"There are grammar and punctuation errors: 'Advancements ... offers benefits' should be 'offer benefits,' and 'worlds population' should be 'world's population.'","section":"Abstract"},{"comment":"References [9] and [17] are the same Whisper-features paper, and references [10] and [18] are both the Tacotron paper. Duplicate entries should be merged.","section":"References"},{"comment":"The phrase 'an average batch size of 64' is odd; batch size is a fixed hyperparameter, so 'a batch size of 64' would be clearer.","section":"Section III.C"},{"comment":"The text states that IFD is 'one-sixth of the size of ASVspoof21 (DF),' but the sample counts in Section III.B imply a ratio closer to 1/30; please correct this quantitative comparison.","section":"Section III.D"},{"comment":"The table entries such as 'Complete Speaker-13' and 'Bonafide Only Speaker-45' are difficult to parse; a clearer layout separating speaker ID, type, and counts would improve readability.","section":"Table I"},{"comment":"The architecture diagrams for ICDD and ISCSE are not referenced or explained in the body text, and the convolution parameters in the diagrams are unclear; consider adding a textual description or moving the figures to an appendix.","section":"Figures 3 and 4"},{"comment":"The conclusion uses 'proves' for empirical findings; 'suggests' or 'is consistent with' would be more appropriate given the mixed per-baseline results.","section":"Section IV"}],"recommendation":"major_revision","confidential_remarks":"The dataset itself appears to be a useful contribution for the ADD community, and the release with documentation is commendable. The main problem is that the abstract and conclusion make stronger claims than the evidence supports, and the evaluation protocol confounds dataset quality with training budget. I believe these issues are fixable with matched experiments, significance testing, and calibrated wording, so I recommend major revision rather than rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The thing to know: this paper gives the field a new, deliberately constructed dataset of Indian-accented English speech for audio deepfake detection, and that artifact is real and useful. The claims that go with it, however, are too strong. Read the paper for the dataset, not for the comparative verdicts.\n\nWhat is actually new: IFD has 50 Indian English speakers, 27.17 hours of bonafide and deepfake audio, balanced class distribution, a subject-independent split, and multiple generation scenarios including hypothetical transcripts and cross-speaker content transfer. That is a gap worth filling—no existing benchmark focuses on South Asian accents, and the authors document the collection pipeline honestly. They also release the data, which is the part that will help people. The ablation study on normalization is minor but not useless.\n\nWhere it gets soft: the abstract and conclusion say IFD \"outperforms ASVspoof21 (DF)\" and is \"more challenging\" than In-The-Wild. The body itself is more careful—\"consistently better ... for most baselines\"—and the numbers support only \"some baselines.\" In Table IV, training on IFD beats training on ASVspoof21 in only 3 of 5 models when tested on ITW; for LFCC-LCNN and MFCC-LCNN, ASVspoof21-trained models do better. The \"more challenging\" claim rests on the same mixed pattern: higher EER on IFD in 3 of 5 baselines. Worse, the training budgets are unmatched—10 epochs on ASVspoof21 versus 50 on IFD—so the comparison conflates dataset content with optimization effort. No error bars, no significance tests. These flaws are load-bearing for the paper's headline, and the fix is straightforward: report matched training budgets, add repetitions or significance tests, and soften the wording to match the actual results.\n\nNone of this undermines the dataset itself. The internal statistics are consistent, the splits are described clearly, and the authors are not hiding their protocol—the confound is right there in Section III.C. I would not call this a deceptive paper; I would call it an over-eager one.\n\nWho is this for? Anyone working on audio deepfake detection in Indian or South Asian contexts, and anyone who wants a stress test for cross-dataset generalization. It deserves a serious referee, because the dataset is a meaningful contribution even if the comparative claims need rework before publication.","headline":"A genuinely useful new dataset for Indian-accented audio deepfake detection, but the headline superiority claims outrun the evidence; the dataset itself is worth engaging with.","tokens_in":11805,"tokens_out":1019,"would_cite":true,"duration_ms":11646,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The IndieFake Dataset, 27.17 hours of bonafide and deepfake English speech from 50 Indian speakers, is introduced to show that a small balanced accent-specific dataset can train audio deepfake detectors better than the much larger…","keywords":["audio deepfake detection","IndieFake dataset","Indian English speech","text-to-speech deepfakes","voice conversion","benchmark dataset","ASVspoof21","equal error rate"],"falsifier":"Train identical baseline models on ASVspoof21 (DF) and IFD with equal epochs, equal compute, or matched convergence criteria, then evaluate on the same held-out test sets; if the IFD-trained models no longer achieve lower equal error rates than the ASVspoof21-trained models, the paper's superior-training-resource claim is not supported.","tokens_in":10834,"feed_emoji":"🎙️","tokens_out":6219,"duration_ms":57345,"temperature":0.7,"pith_summary":"The paper introduces IndieFake (IFD), a 27.17-hour audio deepfake detection dataset built from 50 English-speaking Indian speakers, with 8,164 bonafide and 11,396 deepfake samples. Its goal is to fill the gap left by existing datasets that lack South-Asian accents, and to show that such a dataset can be a better training resource than ASVspoof21 (DF) and a more demanding evaluation benchmark than In-The-Wild (ITW). The authors report that baselines trained on IFD from scratch consistently achieve lower equal error rates on ITW than the same baselines trained on ASVspoof21 (DF), despite IFD being about one-sixth the size. They also report that models trained on ASVspoof21 (DF) perform worse on IFD than on ITW, which they read as evidence that IFD is the harder benchmark.","feed_headline":"27-hour Indian-English audio set beats ASVspoof21 for deepfake training","feed_subtitle":"A balanced 50-speaker set trains detectors better than a much larger set and tests harder than In-The-Wild.","key_machinery":"The central object is the IndieFake Dataset (IFD) itself: 27.17 hours of five-second audio clips from 50 English-speaking Indian speakers, with 8,164 bonafide samples drawn from Creative-Commons YouTube speech and 11,396 deepfake samples generated by TTS and voice-cloning services under three scenarios (hypothetical transcripts, transcripts of the same speaker, and transcripts of another listed speaker). The dataset's design—balanced bonafide/deepfake counts, speaker-level metadata, and a subject-independent 80:20 split—is what carries the argument, because it lets the authors attribute differences in detector performance to accent and content diversity rather than speaker overlap or class imbalance.","core_discovery":"The central claim is that IFD outperforms ASVspoof21 (DF) as a training resource and is more challenging than In-The-Wild as a benchmark, despite its smaller scale. Training five baseline detectors (LFCC-LCNN, MFCC-LCNN, LFCC-MesoNet, MFCC-MesoNet, RawNet3) on IFD and testing on ITW gives lower EER in most cases than training the same baselines on ASVspoof21 (DF). Conversely, models trained on ASVspoof21 (DF) show higher EER and lower accuracy on IFD than on ITW, indicating that IFD contains harder or less familiar spoofing conditions. The dataset also contributes speaker-level characterization and a subject-independent train-test split, which the authors say are missing from ASVspoof21 (DF).","pith_inferences":["The reported advantage may depend on the unequal training budgets, so a matched-budget comparison would tell whether the dataset itself or the extra epochs drive the result.","The cross-speaker transcript scenario isolates content-transfer attacks, so IFD could be used to study a class of spoofing that most existing benchmarks do not separate out.","The same collection recipe—balanced accent-specific speakers plus scenario-driven deepfake generation—could be applied to other under-represented accents, which would show whether accent diversity is the active ingredient in detection performance.","Because the dataset is public, independent groups can reproduce the baseline table and extend it to newer detectors without generating new deepfake audio themselves."],"forward_implications":["Audio deepfake detectors can be trained more cheaply on a small balanced dataset with Indian English accents than on a much larger imbalanced set, if the reported EER advantage holds.","Starting from scratch on IFD appears preferable to initializing from ASVspoof21 (DF) pre-trained weights, since the paper observes performance drops after such fine-tuning.","IFD can serve as a more stringent evaluation benchmark than In-The-Wild for detectors meant to work in Indian English contexts.","The subject-independent split means a detector that does well on IFD has not seen the same speaker's voice in training, which is closer to real-world deployment against new voices."],"supporting_citations":[{"why":"The ASVspoof21 (DF) dataset is the primary comparison training resource, providing the large imbalanced corpus that IFD is claimed to outperform.","marker":"[8]"},{"why":"The In-The-Wild (ITW) dataset is the baseline benchmark used for the cross-dataset difficulty comparison.","marker":"[7]"},{"why":"RawNet3, the end-to-end raw-waveform baseline, supplies the raw-audio detection architecture and is the best-performing model in the evaluation.","marker":"[4]"},{"why":"LCNN is one of the front-end baseline detectors evaluated on all datasets.","marker":"[5]"},{"why":"MesoNet is the second front-end baseline detector evaluated on all datasets.","marker":"[6]"},{"why":"Provides the pre-trained weights used to initialize baselines in fine-tuning experiments.","marker":"[9]"},{"why":"Contributes miscellaneous bonafide audio samples to the IFD test set.","marker":"[46]"}],"fun_headline_variants":["Indian-English deepfake dataset outperforms ASVspoof21 in training","New Indian-English audio benchmark challenges deepfake detectors","50 Indian speakers, 27 hours: a tougher deepfake test set","IndieFake: 27-hour Indian audio set beats bigger deepfake benchmarks","Indian-English deepfake set more challenging than In-The-Wild"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The claim that IFD is a better training resource assumes that comparing error rates after 10 training epochs on ASVspoof21 (DF) and 50 epochs on IFD is a fair measure of dataset quality.","fun_headline_variants_meta":{"raw":{"variants":["Indian-English deepfake dataset outperforms ASVspoof21 in training","New Indian-English audio benchmark challenges deepfake detectors","50 Indian speakers, 27 hours: a tougher deepfake test set","IndieFake: 27-hour Indian audio set beats bigger deepfake benchmarks","Indian-English deepfake set more challenging than In-The-Wild"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000876,"raw_usage":{"total_tokens":3792,"prompt_tokens":947,"completion_tokens":2845,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":563,"completion_tokens_details":{"reasoning_tokens":2754}},"tokens_in":563,"tokens_out":2845,"duration_ms":19359,"temperature":1.0,"reasoning_tokens":2754,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T18:39:17.727594+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Train identical baseline models on ASVspoof21 (DF) and IFD with equal epochs, equal compute, or matched convergence criteria, then evaluate on the same held-out test sets; if the IFD-trained models no longer achieve lower equal error rates than the ASVspoof21-trained models, the paper's superior-training-resource claim is not supported.","supporting_citations":[{"cited_title":"Available at: https://github.com/AI4Bharat/ NPTEL2020-Indian-English-Speech-Dataset/tree/master","cited_arxiv_id":null,"evidence_quote":"Contributes miscellaneous bonafide audio samples to the IFD test set."}],"review_version":2}