{"id":"0d81f7b1-6070-4aa6-b239-4be462ab8d6b","arxiv_id":"2502.00961","paper_version":1,"verdict":"REJECT","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"high","formal_verification":"none","parameter_count":0,"one_line_summary":"In a 44-person test, 66% of audio and 43% of video answers about AI-generated clips were wrong, but those rates include errors on genuine clips and are not clean measures of missed fakes.","lead":"The authors made fake audio and video clips with consumer deepfake tools and asked 44 people to pick which were fake. Most answers were wrong, but the way the results were counted makes the headline numbers unreliable.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The headline 66%/43% figures are overall misclassification rates over all clips, including genuine ones; they do not measure failure to detect AI-created audio/video, so the abstract's central claim is unsupported.","rationale":"The reader's weakest-assumption analysis and mine converge on the same point: the headline rates are computed over all selections, not over AI-created clips, and the statement '66% failed to identify AI created audio as fake' cannot be recovered from the data as reported. This is load-bearing because it is the only quantitative evidence for the paper's central claim. If the overall error rate is mostly false positives on genuine clips, the experiment may even show that participants were biased toward calling things fake, not that they were fooled by deepfakes. The internal inconsistency in Table 2 (e.g., the 25–34 age group shows higher video detection than several younger groups while the prose claims audio was hardest) reinforces but does not carry the objection. I also note the absence of statistical tests and confidence intervals, but the more fundamental problem is not inferential noise; it is that the reported statistic measures a different quantity than the claim. The general observation that deepfakes can support phishing is consistent with prior case reports and is not the contested part; the contested part is the specific 66%/43% quantification, which should be rejected as stated. A revised version that reports fake-only sensitivity, false-positive rates, and uncertainty could be reconsidered, but the current paper's central empirical claim is not supported.","tokens_in":15093,"tokens_out":3308,"duration_ms":32424,"concrete_test":"Recompute the Section 5 results as a confusion matrix per modality from the per-clip raw responses (or rerun the questionnaire with balanced fake/genuine counts): sensitivity = fraction of the 3 fake audio clips and 3 fake video clips mislabeled as genuine, separately from specificity = fraction of genuine clips correctly labeled. If the fake-only miss rate differs materially from 66.2% and 43.1% — for example if it is near the random-guessing baseline while the high overall error is driven by false positives on the more numerous genuine clips — then the abstract's 'failed to identify AI created media' claim is an artifact of pooling the two error types.","verdict_should_be":"REJECT","load_bearing_attack":"The paper's central empirical claim is the abstract statement that 66% of participants failed to identify AI-created audio as fake and 43% failed to identify such videos as fake. That claim is not what Section 5 actually reports. Section 4.4.1 describes a video task with 3 fake and 5 genuine clips and an audio task with 3 fake and 10 genuine clips. Section 5 then states that 'of the audio selections made 66.2% were incorrect' and '43.1% incorrect when selecting video examples.' These are overall error rates across all clips, pooling missed fakes with false positives on genuine clips. Because genuine clips outnumber fakes 10:3 in audio and 5:3 in video, a participant who succumbs to the stated instruction that fakes are present and labels many genuine clips as fake will contribute a high overall error rate even while detecting every actual fake. The abstract's verb 'failed to identify AI created audio/video' can only be sustained by the fake-only miss rate, which is never reported. The conclusion's inference that 'the remaining 66% of attacks would convince a victim' adds a further unjustified step, equating a wrong selection with falling for an attack. The paper itself notes high false positives in Section 5, which is consistent with this conflation. Thus the headline quantitative finding is unsupported by the experiment as analyzed.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper investigates whether deepfake audio and video could enable spearphishing attacks. It reviews reported and hypothetical cases, creates deepfake media using DeepFaceLab, ResembleAI, and speech-driven animation on a consumer desktop, and then presents 44 participants with genuine and fake clips across separate audio, video, and combined conditions. The abstract's headline finding is that 66% of participants failed to identify AI-created audio as fake and 43% failed to identify such videos as fake, leading the authors to conclude that deepfake-enabled spearphishing is a serious and easily accessible threat.","tokens_in":15234,"tokens_out":6418,"duration_ms":64299,"significance":"If the headline numbers were valid, the study would be a useful empirical data point on human detection of deepfake media in a security context. The paper gives credit for several things: it actually creates deepfake media rather than only speculating, it uses a public audio-video dataset, it collects data across age groups and before/after threat perceptions, and it compares audio, video, and combined conditions. However, the central quantitative claim is not supported by the analysis as reported: the 66% and 43% figures are overall incorrect selection rates over a mixture of genuine and fake clips, not failure-to-detect rates for fake clips. Without a corrected analysis, the abstract's main claim and the conclusion's inference about attack success do not follow from the experiment.","major_comments":[{"comment":"The abstract's headline claim is not what Section 5 actually measures. Section 4.4.1 states that the video task contains 3 fake and 5 genuine clips and the audio task contains 3 fake and 10 genuine clips, while Section 5 reports 'of the audio selections made 66.2% were incorrect' and '43.1% incorrect when selecting video examples.' These are overall error rates over all clips, pooling missed fakes with false positives on genuine clips. Because genuine clips heavily outnumber fakes (10:3 in audio, 5:3 in video), a participant who is told fakes are present and therefore mislabels many genuine clips as fake can be counted as 'incorrect' even while detecting every actual fake. The abstract's phrase 'failed to identify AI created audio as fake' requires the miss rate restricted to fake items, which is never reported. Please report a per-condition confusion matrix, the fake-only miss rate, the false-alarm rate on genuine clips, and confidence intervals.","section":"Section 5 and Section 4.4.1"},{"comment":"The inference that 'if only 34% of assumptions were correct when attempting to identify fake audio samples then ... the remaining 66% of attacks would convince a victim into falling for a phishing attack' conflates an incorrect forced-choice classification in a laboratory task with falling victim to a spearphishing attack. The experiment did not measure whether participants would click a link, enter credentials, transfer funds, or otherwise comply. Moreover, Section 4.4.1 tells participants that fake examples are present, which is unlike most real-world encounters and can inflate suspicion and false positives. This conclusion is therefore not supported by the data.","section":"Section 6 (Conclusion)"},{"comment":"The age-disaggregated numbers in Table 2 appear inconsistent with the overall correct-rate figures in Section 5. Using the participant counts in Table 1, a weighted average of the 'Video Detection Rate' column is roughly 36%, whereas Section 5 reports 56.9% correct when selecting video examples. This suggests that Table 2 and Section 5 are using different definitions of 'detection rate' (for example, per-fake detection versus per-selection accuracy), or that the table contains an error. The metric must be defined explicitly for each reported statistic, and raw counts should accompany percentages. In addition, several age cells contain only one or two participants (e.g., 'Below 18' has n=1), so the conclusion in Section 6 that 'the older an individual is, the less likely they are to correctly identify artificial media' is not supported by these data.","section":"Section 5, Table 2"},{"comment":"No inferential statistics, confidence intervals, or chance-level baselines are provided. For the audio task, a participant who simply classified every clip as genuine would be correct on 10 of 13 clips (76.9%), while a participant who classified every clip as fake would be correct on only 3 of 13 (23.1%); the reported 33.8% overall correct rate therefore cannot be interpreted without a chance baseline and a measure of variability. The absence of such analysis is particularly consequential for the small age subgroups where percentages such as 100% and 0% appear in Tables 3 and 4.","section":"Section 5 generally"}],"minor_comments":[{"comment":"The hardware described (Intel i7-8700K, GTX 1080 Ti, 32GB RAM) is perhaps better described as a 'consumer desktop' rather than a 'low-spec computing facility,' which is the phrase used in the introduction and conclusion.","section":"Section 4.1"},{"comment":"The phrase 'a varied combination of real and genuine audio and video' is confusing because 'real' and 'genuine' are synonyms; please clarify the intended contrast.","section":"Section 4.4.1"},{"comment":"Entries of 100.00% and 0.00% in small age cells should be suppressed or accompanied by raw counts, and the 'N/A' cells should be explained (e.g., no participants from Form A in those age groups).","section":"Tables 3 and 4"},{"comment":"The abstract rounds 66.2% and 43.1% to 66% and 43%, respectively; this is acceptable, but the paper should ensure that all quoted percentages are traceable to the same denominator and metric.","section":"Abstract and Section 5"},{"comment":"References [38] and [39] both appear to cite the same arXiv paper by Li and Lyu; one duplicate should be removed, and the in-text citations should be checked.","section":"References"},{"comment":"There are numerous typographical and formatting errors, including 'conﬁdently,' 'Artiﬁcial,' and inconsistent spacing in the references; a careful proofread is needed.","section":"General presentation"}],"recommendation":"major_revision","confidential_remarks":"The central problem is that the headline statistics measure the wrong quantity, and the conclusion draws a behavioral inference the experiment cannot support. If the authors can reanalyze their raw responses to report fakes-only miss rates, false-alarm rates, and confidence intervals, the paper could be salvaged as a modest empirical study; if the raw data are not available or the reanalysis does not support the abstract's claim, the paper would likely need to be rejected. I would also ask the editor to consider whether the journal's standards require a balanced set of fake and genuine stimuli and a deception or behavioral outcome for conclusions about spearphishing susceptibility."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Kemp et al. ask a sensible question: can off-the-shelf deepfake tools on consumer hardware fool average people in a spearphishing scenario? That is worth testing, and the paper does some real work: it generates video and audio deepfakes with DeepFaceLab, Resemble AI, and a lip-sync tool, embeds them among genuine clips from the VidTIMIT dataset, and runs a 44-person evaluation. The literature review also usefully collects the known real-world cases, including the UK energy company CEO voice fraud. I credit the authors for attempting an empirical human-subjects study rather than another opinion piece.\n\nBut the central result, as stated, does not hold up. The abstract says 66% failed to identify AI-created audio and 43% failed to identify AI-created video. The paper does not report those numbers. Section 5 reports that 66.2% of all audio selections were incorrect and 43.1% of all video selections were incorrect. The audio task had 3 fake and 10 genuine clips; the video task had 3 fake and 5 genuine clips. So a participant who correctly identified every fake but also flagged several genuine clips as fake would rack up a high overall error rate. That is a mix of missed fakes and false positives, not a failure-to-detect rate. The authors even acknowledge high false positives in Section 5, which is consistent with this conflation.\n\nThere are further soft spots. Table 2's age-disaggregated detection rates conflict with the prose claim that audio was the hardest to detect; for several age brackets the audio detection rate is higher than the video rate. There are no statistical tests, confidence intervals, or a chance-level baseline, and the conclusion's leap from “wrong selection” to “66% of attacks would convince a victim” is unjustified. The sample is small and convenience-based. These are not minor issues; they undermine the paper's headline contribution.\n\nWhat is genuinely new is the specific detection data for these generated clips, but it is not presented correctly. The paper deserves a serious referee, not because the current analysis is acceptable, but because the underlying question is timely and the data could be reanalyzed. If the authors re-report fake-only miss rates and false-positive rates separately, and add basic inference, the study could make a modest but honest contribution. As it stands, the conclusions should not be taken at face value.","headline":"The paper's headline 66%/43% detection-failure claims are not supported by the reported analysis — those are overall error rates over mixed real/fake sets, not missed-fake rates — but the study is a genuine attempt worth a careful revision rather than a desk reject.","tokens_in":15859,"tokens_out":1837,"would_cite":false,"duration_ms":19887,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Most people, even when told fakes may be present, misjudged 66% of AI-generated audio clips and 43% of AI-generated video clips, and the fakes were built on an ordinary desktop.","keywords":["Techno-progressivism","Deepfakes","Spearphishing","Artificial Intelligence","Social engineering","Audio deepfake detection","Video deepfake detection","Phishing awareness"],"falsifier":"Recompute the audio and video error rates separately for fake items and for genuine items from the raw responses; if most incorrect audio selections were participants labelling genuine clips as fake rather than missing fake clips, the headline claim that 66% of people failed to identify AI-created audio is not supported.","tokens_in":14781,"feed_emoji":"🎭","tokens_out":19045,"duration_ms":166098,"temperature":0.7,"pith_summary":"This paper argues that deepfake-enabled spearphishing is a practical, near-term threat rather than a hypothetical one. The authors used openly available tools and a desktop computer to create synthetic speech and swapped-face video of a known speaker, then hid these fakes among genuine clips and asked 44 participants to pick them out. Their headline finding is that 66% of participants failed to identify the AI-created audio as fake and 43% failed to identify the AI-created video as fake, figures they compute from the wrong selections made over a mix of fake and genuine items. If this is right, the barrier to running convincing voice- and video-based phishing attacks is already low, and the question becomes how to give ordinary people reliable cues for spotting the remaining artifacts.","feed_headline":"66% of AI voice clones fooled listeners in phishing test","feed_subtitle":"Simulated spearphishing test: most volunteers couldn't tell cloned voices from real ones; two in five missed fake video.","key_machinery":"The argument is carried by a detection experiment built around deepfakes—synthetic audio and video that replace a person's likeness or voice—and by the toolchain that produced them. In the questionnaire, the video condition had 8 clips (3 fake, 5 genuine), the audio condition had 13 clips (3 fake, 10 genuine), and the combined condition had 8 lip-synced clips; participants were told that fakes were present and were asked to identify the generated items and give their reasons. The generated media themselves—a face swap, a cloned voice, and a lip-synced combination—are the objects whose realism the experiment measures.","core_discovery":"The paper's central claim is that synthetic audio and video are already convincing enough to support spearphishing attacks, and that the production side is no longer a barrier. Using a public dataset of 43 speakers as source material, the authors generated swapped-face video with an open-source face-swapping tool, cloned speech with a commercial voice-synthesis service, and lip-synced combinations with a speech-driven animation model, all on an ordinary desktop computer. In a questionnaire with 44 respondents, participants were told that some clips were fakes and were asked to identify them; the paper reports 66.2% of audio selections were incorrect and 43.1% of video selections were incorrect, and in the combined condition the audio component stayed about as hard to detect (68.3% incorrect) while the video component became easier (30.2% incorrect). The paper reads these rates as evidence that a large share of such attacks would succeed, that synthetic audio is the most dangerous component, and that people without prior knowledge of deepfakes—along with older adults—are especially vulnerable.","pith_inferences":["Beyond the paper's aggregate numbers, the 66% and 43% figures are rates over selections, not over people; a participant who incorrectly flags several genuine clips as fake counts as an error without having been deceived by a fake. Recomputing the rates for fake-only and genuine-only items would separate true deception from an over-cautious response bias.","Because participants were told that fakes might be present, the experiment may overstate how well people would perform when nothing has primed their suspicion; running the same test without any warning would give a cleaner measure of real-world vulnerability and could plausibly show even higher deception rates.","The paper uses one speaker's voice and actors recorded in identical studio conditions, so the results may not transfer to attacks that impersonate a widely recognized public figure; repeating the protocol with a well-known voice would test whether familiarity helps detection or makes the clone more believable."],"forward_implications":["Because the fakes were produced on a desktop computer with accessible tools, the cost and skill barrier for running deepfake spearphishing attacks is low enough that this is a current threat, not a distant one.","Audio is the riskier medium in this study: it was misidentified more often than video, and in the combined condition the audio portion stayed nearly as hard to detect while the video portion became easier to spot.","Older respondents and respondents with no prior knowledge of deepfakes had lower detection rates, which points to awareness and training as targeted countermeasures.","The poor visual quality of existing lip-syncing tools currently helps defenders, since adding lip-synced video made the video component easier to detect.","The documented real-world case of a CEO's cloned voice gives the lab results a concrete anchor: the authors treat the audio error rate as the share of audio-based attacks that could plausibly succeed."],"supporting_citations":[{"why":"Supplies the video and audio recordings of 43 actors that are the genuine source material for the experiment.","marker":"[1]"},{"why":"The academic publication describing that actor dataset, cited alongside [1] as the data source.","marker":"[48]"},{"why":"The open-source face-swapping tool used by the authors to generate the deepfake video clips.","marker":"[28]"},{"why":"The academic paper describing the face-swapping framework behind that tool.","marker":"[42]"},{"why":"The commercial voice-cloning service used to synthesize the fake audio clips.","marker":"[8]"},{"why":"The lip-syncing application used to build the combined audio-video stimuli.","marker":"[9]"},{"why":"The temporal generative-adversarial method behind the lip-syncing application, cited as the technical basis for the combined clips.","marker":"[52]"},{"why":"The documented real-world fraud in which a CEO's cloned voice convinced a victim to transfer funds, used as evidence that such attacks occur.","marker":"[10]"}],"fun_headline_variants":["66% fell for AI-cloned voice in spearphishing test","AI voice scams fool 66% of testers in phishing trial","Most listeners can't identify AI-generated voice in phishing","Deepfake audio tripped up 66% in spearphishing simulation"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The central claim depends on treating an incorrect answer on a test where participants were told fakes exist as the same thing as being deceived by a fake in a real-world attack.","fun_headline_variants_meta":{"raw":{"variants":["66% fell for AI-cloned voice in spearphishing test","AI voice scams fool 66% of testers in phishing trial","Most listeners can't identify AI-generated voice in phishing","Deepfake audio tripped up 66% in spearphishing simulation"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000487,"raw_usage":{"total_tokens":2385,"prompt_tokens":913,"completion_tokens":1472,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":529,"completion_tokens_details":{"reasoning_tokens":1399}},"tokens_in":529,"tokens_out":1472,"duration_ms":11031,"temperature":1.0,"reasoning_tokens":1399,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-09T17:05:45.724658+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Recompute the audio and video error rates separately for fake items and for genuine items from the raw responses; if most incorrect audio selections were participants labelling genuine clips as fake rather than missing fake clips, the headline claim that 66% of people failed to identify AI-created audio is not supported.","supporting_citations":[{"cited_title":"Vidtimit audio-video dataset","cited_arxiv_id":null,"evidence_quote":"Supplies the video and audio recordings of 43 actors that are the genuine source material for the experiment."},{"cited_title":"Multi-region probabilistic his- tograms for robust and scalable identity inference, in: International conference on biometrics, Springer","cited_arxiv_id":null,"evidence_quote":"The academic publication describing that actor dataset, cited alongside [1] as the data source."},{"cited_title":"iperov/deepfacelab","cited_arxiv_id":null,"evidence_quote":"The open-source face-swapping tool used by the authors to generate the deepfake video clips."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"The commercial voice-cloning service used to synthesize the fake audio clips."},{"cited_title":"Dinoman/speech-driven-animation","cited_arxiv_id":null,"evidence_quote":"The lip-syncing application used to build the combined audio-video stimuli."},{"cited_title":"Fraudsters used ai to mimic ceo’s voice in un- usual cybercrime case","cited_arxiv_id":null,"evidence_quote":"The documented real-world fraud in which a CEO's cloned voice convinced a victim to transfer funds, used as evidence that such attacks occur."}],"review_version":1}