{"id":"08daf953-3a9f-45d6-a983-5e46f6393f0d","arxiv_id":"2509.09204","paper_version":1,"verdict":"REJECT","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"A new evaluation protocol exhaustively pairs 164 speech synthesizers with nine bona fide speech types and reports max-pooled EERs, revealing larger failures than pooled averages show.","lead":"This paper proposes a new way to test audio deepfake detectors: pair each of 164 fake-voice generators with nine types of real speech, then report the worst error rate for each type. It finds that interview-style real speech is a hidden weak spot for current detectors, and that average error rates hide much larger worst-case failures.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Worst-case EERs >0.5 imply score anti-correlation; unless score polarity is verified, the reported 'weak spot' may be a sign-flip artifact—and the 30%/10% reading conflicts with the table's % label.","rationale":"The proposed cross-testing idea is reasonable, and the paper includes a heatmap and table that, if correct, would be valuable. The most load-bearing claim is not the framework's mathematical definition but the magnitude and interpretation of the headline numbers, because the entire 'weak spot' conclusion depends on the max and average EER values in Table 2. The reader's weakest_assumption concerned cross-corpus comparability and score normalization; I think a more immediate, argument-level problem is that the key numbers are inconsistent with their labeling and, under the only reading that supports the 'weak spot', the models are reported as anti-predictive on many subsets. That would be an extraordinary result that needs a polarity check before any conclusion about b6/b9. I agree with the reader's REJECT verdict overall: the placeholder link, the Eq. (8) typo, and the table ambiguity independently make verification impossible. This polarity/numerical ambiguity is the flaw most directly tied to the central empirical claim: if the check shows min(EER, 1−EER) is low, the headline vulnerability is an artifact; if not, the framework would merit a more positive assessment conditional on fixing the presentation errors.","tokens_in":8674,"tokens_out":17042,"duration_ms":203189,"concrete_test":"Once the released score files are available (the paper's link is https://empty.com), take the (k,m) pairs that give the maximum Table 2 EERs for Wav2Vec-Conformer and recompute EER two ways: using scores as-is and using negated scores. Also compute min(EER, 1−EER) for every cell with EER>0.5. If most of these cells have min < 0.1, the high EERs are sign-flip artifacts; if min remains ≥0.3, the weak spot is genuine. Also confirm whether the reproduced EERs match the decimal values in Table 2 or ten times smaller values; this discriminates between the fraction and percentage interpretations.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central evidence is Table 2. The header says 'EER (%)' but the text ('fail to detect more than 30%', 'average EER approximately 10%') only makes sense if the entries are fractions. Under the fraction reading, many max-EER cells (0.73–0.99) exceed 0.5. A binary detector with scores independent of labels has EER=0.5; values above 0.5 mean the score ordering used in Eqs. (1)–(2) is anti-correlated with the class labels for that subset. The paper never checks the polarity of the SSL models' scores across the 164×9 pairs. If the worst pairs simply have inverted scores, the true EER is 1−EER—e.g., 0.98 becomes 0.02—and the claimed 'hidden vulnerabilities' largely disappear. If instead the entries are literal percentages, max EERs are <1% and the 'more than 30%' sentence is unsupported. Either way the headline magnitude is not established.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes 'bona fide cross-testing' for audio deepfake detection: instead of combining synthesizers into a single test set, each of 164 synthesizer subsets is paired with each of 9 bona fide speech corpora, and per-pair EERs are computed and then summarized by maximum and average across synthesizers. The authors argue that this reveals hidden vulnerabilities, especially on the celebrity-interview bona fide types b6 and b9, and that it is more robust and interpretable than traditional spoof cross-testing. They evaluate three SSL-based ADD models and claim to release a dataset, code, and score files.","tokens_in":8946,"tokens_out":4639,"duration_ms":55445,"significance":"The central idea addresses a real limitation of current ADD evaluation: pooling synthesizers in a single EER underweights underrepresented spoof subsets and ignores bona fide speech diversity. The construction of a benchmark with 164 synthesizers and 9 bona fide types is a useful contribution, and the proposal to report max-pooled EER is a reasonable way to surface worst-case performance. If the results were correct, the claim that current detectors are far less robust than average EERs suggest would be an important finding. However, the manuscript as written contains several load-bearing errors—an inconsistent definition of FPR in Eq. (8), an unresolved unit/polarity problem in Table 2, and a placeholder repository URL—that currently prevent the central evidence from being accepted.","major_comments":[{"comment":"The table header says 'EER (%)' but the entries are decimals (e.g., 0.95), and the text interprets them as fractions ('fail to detect more than 30%', 'average EER approximately 10%'). Under the fraction reading, many max-EER values exceed 0.5 (e.g., 0.73, 0.98, 0.99). Equal error rate is bounded at 0.5 for a random or score-independent binary detector; values above 0.5 imply that the score ordering used in Eqs. (1)-(2) is anti-correlated with the class labels for those subsets. The paper never checks the polarity of the models' scores across the 164x9 pairs. If the polarity is reversed, the correct EER would be 1 - EER, and the 'hidden vulnerabilities' may largely disappear. If the entries are instead literal percentages, they are all below 1%, and the 'more than 30%' sentence is unsupported. Either way, the headline magnitude is not established.","section":"Section 4, Table 2"},{"comment":"The definition of P^k_FP is inconsistent with Eq. (1). Eq. (8) uses the spoof-set symbol Lambda^k_P and the condition s_j >= tau, while Eq. (1) defines FPR over the bona fide set Lambda_N with s_j < tau. This is not a mere typo: Eqs. (9)-(10) use P^k_FP to compute EER^{k,m}, so if Eq. (8) is implemented literally, the framework computes a false-negative rate rather than a false-positive rate. The correct definition and the evaluation code need to be fixed and verified.","section":"Section 2.4, Eq. (8)"},{"comment":"The abstract and contribution list promise code, datasets, and score files at https://github.com/cyaaronk/audio_deepfake_eval, but Section 3 says the material is 'available at https://empty.com'. This is a placeholder, not a usable repository. The reproducibility claim cannot be verified, and the reference to a nonexistent URL is a serious omission for a benchmark paper.","section":"Section 3, data release"},{"comment":"The claim that b6 and b9 are harder because they are 'celebrity interview speech... recorded in a noisy public area' is confounded by the dataset-level properties listed in Table 1: different sampling rates (44.1 kHz vs. 16/48 kHz), different durations, different codecs, and different recording setups. EERs computed on different corpora are compared directly without per-dataset score normalization or an analysis of score distribution shifts. It is therefore possible that the observed b6/b9 differences are artifacts of corpus-specific characteristics rather than intrinsic properties of the bona fide speech style. The paper should provide evidence that the effect persists after controlling for these factors.","section":"Section 4, b6/b9 claim"}],"minor_comments":[{"comment":"Typos: 'Methodoglogy' in Section 2, 'nagative' in Section 2.2, 'accross' in Section 2.5, 'availabe' in Section 3, and 'AV-Deefake' in Table 1 should be corrected.","section":"General"},{"comment":"The heading says 'Maximum pooling on spoof cross-testing results', but Eq. (11) pools over synthesizers for each bona fide type. The heading should reflect that this is max pooling over synthesizers within the bona fide cross-testing framework.","section":"Section 2.5"},{"comment":"The same source datasets (FakeAVCeleb and AV-Deepfake-1M) appear as both bona fide subsets (b6, b9) and spoof subsets (s3, s5). While this is a defensible choice, it should be stated explicitly and the implications for the 'cross-testing' terminology should be discussed.","section":"Table 1"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should know about this paper for its evaluation idea, but don't trust its numbers yet. Bona fide cross-testing—varying the genuine-speech side while sweeping 164 synthesizers—is a sensible stress test that the ADD community needs. The scope alone is valuable if the data is real. Max-pooling over synthesizers to reveal worst-case vulnerabilities is a reasonable, attacker-motivated aggregation, and the paper is right that most ADD evals ignore bona fide diversity.\n\nThe problem is that the current manuscript cannot be verified. Eq. (8), which defines the core FPR for the k-th bona fide subset, uses the spoof-set symbol Λ_P and the wrong inequality sign; it should be Λ_N with `<`. That is a load-bearing typo. Table 2 is labeled 'EER (%)' but the entries (0.95, 0.98, etc.) are obviously fractions—the text's 'more than 30%' and 'average ERR about 10%' only make sense as fractions. This isn't cosmetic: many max-pooled EERs exceed 0.5. The stress-test note is right that an EER above 0.5 means the score ordering is anti-correlated with labels for that subset; flipping the decision rule gives 1−EER, often under 0.1. The paper never checks score polarity, so the advertised 'weak spot' on celebrity interviews (b6/b9) may be an artifact. Without that check, the headline magnitude is not established.\n\nAlso, the data link in Section 3 is literally 'https://empty.com', contradicting the abstract's GitHub URL. That makes the benchmark unreproducible as submitted. The 'first such framework' claim is also overstrong given In-The-Wild [25] already studied cross-domain generalization, though the specific sweep over 164 synthesizers and 9 bona fide types is new.\n\nIf the authors fix the equation, relabel the table, add a per-subset polarity check (or justify their convention), and actually release the code/data, this could be a useful resource. The core idea deserves a serious referee; the current presentation doesn't. I'd encourage you to engage with it as a promising draft rather than dismiss it, but don't cite the numbers yet.\n\nFor peer review: send it out. The methodology idea is important enough that a careful referee could help the authors turn it into something solid. But it needs heavy revision.","headline":"Bona fide cross-testing is a genuinely useful evaluation idea, but the paper's own numbers are unreliable until the metric equation, the table units, and the score-polarity question are fixed.","tokens_in":9423,"tokens_out":6492,"would_cite":false,"duration_ms":73160,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that audio deepfake detectors are far less robust than average error-rate scores suggest, and that a new evaluation design—bona fide cross-testing—exposes worst-case equal error rates of 73–99% that standard benchmarks mis","keywords":["audio deepfake detection","bona fide cross-testing","equal error rate","evaluation methodology","spoof detection","maximum pooling","speech anti-spoofing","generalization"],"falsifier":"Look for a dataset-identity artifact: compute the same mEER after per-corpus score normalization (e.g., z-scoring each bona fide subset's scores by its own mean and variance) or after matching duration and channel conditions; if the b6/b9 elevation shrinks toward the other types, the claimed genuine-speech vulnerability is an artifact. Alternatively, retrain a detector on interview-style bona fide data and check whether the max-pooled EER for b6/b9 drops substantially; it should if the paper's explanation is right.","tokens_in":8579,"feed_emoji":"🎙️","tokens_out":4978,"duration_ms":55962,"temperature":0.7,"pith_summary":"The authors propose a new way to evaluate audio deepfake detectors. Instead of melting many synthesizers into one test set and reporting a single error rate, they pair every synthesizer with nine different kinds of genuine speech and compute an equal error rate for each pair. Across 164 synthesizers, the maximum pooled error rate—the worst synthesizer for each genuine speech type—reaches 0.73 to 0.99 for three recent self-supervised detectors, while average error rates stay near 0.10. The paper argues that these worst-case numbers are the realistic ones because attackers will pick the easiest-to-miss synthesizer, and that genuine speech variety, not just synthesizer variety, drives the failures. The authors release their benchmark data and code so the field can use the same protocol.","feed_headline":"Worst-case audio deepfake error hits 99 percent in cross-test","feed_subtitle":"Pairing 164 synthesizers with nine speech types shows average scores hide real-world weak spots such as celebrity interviews.","key_machinery":"The key machinery is bona fide cross-testing combined with maximum pooling. Bona fide cross-testing takes K diverse genuine-speech datasets and M synthesizer datasets, forms M×K test sets, and computes a separate EER for each pair using the threshold that balances false positives and false negatives. Maximum pooling then summarizes the M EERs for each genuine speech type by taking the maximum, yielding mEER_k. This design counters two problems: underrepresented subsets lose influence in a combined dataset, and a single genuine-speech type gives no view of real-world diversity. The reported mEER_k values are the paper's main evidence for hidden vulnerabilities.","core_discovery":"The central discovery is that the vulnerability of current audio deepfake detectors lies partly in the bona fide side of the test, not only in the spoof side. When the same set of 164 synthesizers is crossed with nine bona fide speech types—clean read speech, meetings, noisy accented speech, news, emotion, social media, and celebrity interviews—the equal error rates vary sharply by bona fide type. Two interview corpora (labeled b6 and b9) consistently produce the worst results, which the authors attribute to fast-paced celebrity speech recorded in noisy public areas. Averaging over synthesizers masks this: average EERs cluster near 10%, while maximum-pooled EERs reach 73–99%. The paper concl","pith_inferences":["If bona fide cross-testing becomes standard, detector training should include diverse genuine speech, especially noisy conversational interview audio, rather than only clean read speech; this extension is implied but not developed in the paper.","A possible confound: EERs across corpora are compared without per-dataset score normalization; channel, codec, or speaker differences between b6/b9 and other corpora could contribute to the observed gaps. The paper does not test this.","The manuscript's reproducibility section contains a placeholder URL (https://empty.com) alongside the abstract's GitHub link; the stated benchmark release is not verifiable until that link resolves.","One testable extension: applying the same cross-test protocol to a detector trained with interview-style bona fide augmentation should shrink the b6/b9 gap, which would corroborate the authors' attribution to environmental difficulty."],"forward_implications":["Average EER in the 10% range can coexist with worst-case per-synthesizer EERs above 0.9, so published averages understate risk.","Detectors trained on older benchmarks generalize poorly to newer synthesizers; several 2024 synthesizer families cause large errors.","Bona fide speech type is a first-order factor: celebrity interview speech (b6, b9) is markedly harder than clean read speech (b3, b5) for all three detectors.","Evaluating on combined multi-synthesizer sets can hide failures of underrepresented synthesizers because the global EER threshold is dominated by larger subsets.","Reporting per bona fide type is necessary for deployment decisions; for example, fake-news detection should use the mEER for news-domain bona fide audio."],"fun_headline_variants":["Deepfake detectors hit up to 99% error on celebrity interviews","Cross-testing exposes deepfake detectors' celebrity interview blind spot","Fast-paced celebrity speech exposes deepfake detector flaws","Average EERs mask 99% failure on interviews, cross-test shows"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The framework assumes the nine selected bona fide corpora, and the 600-sample cap per synthesizer–bona fide pair, capture real-world speech diversity and are directly comparable; systematic score offsets from differing channels, codecs, or speakers could make the b6/b9 weakness an artifact of dataset mismatch rather than genuine vulnerability.","fun_headline_variants_meta":{"raw":{"variants":["Deepfake detectors hit up to 99% error on celebrity interviews","Cross-testing exposes deepfake detectors' celebrity interview blind spot","Fast-paced celebrity speech exposes deepfake detector flaws","Average EERs mask 99% failure on interviews, cross-test shows"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001138,"raw_usage":{"total_tokens":4535,"prompt_tokens":693,"completion_tokens":3842,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":437,"completion_tokens_details":{"reasoning_tokens":3770}},"tokens_in":437,"tokens_out":3842,"duration_ms":30553,"temperature":1.0,"reasoning_tokens":3770,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-04T19:29:46.010764+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Look for a dataset-identity artifact: compute the same mEER after per-corpus score normalization (e.g., z-scoring each bona fide subset's scores by its own mean and variance) or after matching duration and channel conditions; if the b6/b9 elevation shrinks toward the other types, the claimed genuine-speech vulnerability is an artifact. Alternatively, retrain a detector on interview-style bona fide data and check whether the max-pooled EER for b6/b9 drops substantially; it should if the paper's explanation is right.","supporting_citations":[],"review_version":1}