{"id":"41b60b32-cff6-406b-be25-2d08ac3ae741","arxiv_id":"2411.14586","paper_version":1,"verdict":"REJECT","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"high","formal_verification":"none","parameter_count":0,"one_line_summary":"A pretest-posttest study of 129 undergraduates found that a brief training on linguistic cues reduced 'unsure' responses when judging audio deepfakes, but did not significantly improve overall discernment accuracy relative to a reading-based control.","lead":"A brief training on five linguistic cues made college students less likely to say 'unsure' when judging audio deepfakes, but it did not improve how often they got it right. A control group that read an article improved accuracy more than the trained group did.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The abstract's claimed accuracy improvement on initially unsure clips is contradicted by the paper's own Table 2; the only significant experimental effect is reduced unsurety, which is a response-bias effect, not improved discernment.","rationale":"The reader's verdict is REJECT, and my analysis supports that rejection, but I identify a slightly different load-bearing concern than the reader's stated weakest assumption. The reader emphasized the memory/practice confound from reusing the same 20 clips in pre and post tests; that is a serious validity threat. The more direct and decisive problem, however, is that the paper's own tables contradict the abstract's central claim regardless of design validity. Table 2 shows no significant accuracy improvement for the experimental group on the exact subset (initially unsure clips) highlighted in the abstract, and the control group improved more. This internal inconsistency means the claim fails even on the paper's own terms. The significant decrease in unsurety is a real but distinct outcome, and it cannot substitute for accuracy improvement. The proposed concrete test would settle whether any accuracy effect exists after controlling for response bias, but the current manuscript already provides sufficient evidence from its own statistics to reject the abstract's claim. Therefore the verdict remains REJECT, and no adjustment to the reader's decision is needed.","tokens_in":6286,"tokens_out":2980,"duration_ms":28437,"concrete_test":"Reanalyze the pre/post data for the subset of clips each experimental student marked 'unsure' in the pre-test: compute per-student accuracy change on those clips, test with a paired Wilcoxon test and a between-group comparison against the control group (e.g., ANCOVA on post accuracy with pre unsurety as covariate). Also compute signal-detection d' and response-bias c for real vs fake clips pre and post for each group. If the experimental group's accuracy change on initially unsure clips is not significantly greater than zero or than the control group's, the central claim in the abstract is falsified.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim requires that the training improves students' ability to correctly identify audio deepfakes, specifically on clips they were initially unsure about. The paper's own inferential statistics contradict this. In Table 2, among students who were unsure on at least one pre-test clip, the experimental group's mean paired difference in accuracy is 0.02 (not significant), while the control group's is 0.05 (significant). Thus the experimental group did not significantly improve on the exact subset highlighted in the abstract, and the control group improved more. The only statistically significant experimental effect is the decrease in unsurety (Table 1), which measures willingness to commit, not correctness. The supplementary observation that 85% of initially-unsure fake clips and only 20% of initially-unsure real clips were correctly identified in the post-test is not a paired significance test and is compatible with a learned response bias toward labeling clips 'fake': because the stimulus set is 80% fake, such a bias would mechanically raise fake-clip accuracy while lowering real-clip accuracy, leaving overall accuracy unchanged. Therefore the central claim is unsupported by the reported data.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript reports a pre/post intervention study of 129 undergraduate students (from 264 enrolled) at UMBC, in which an experimental group received a 15–20 minute module teaching five sociolinguistic cues (pitch, pause, stop-consonant bursts, breath intake/outtake, and audio quality) and a control group read an article about deepfakes. All participants listened to the same 20 clips (4 real, 16 fake) before and after the intervention and judged each clip as real, fake, or unsure. The authors report that the experimental group significantly reduced unsurety and, in the abstract, claim an improvement in correct identification of clips they were initially unsure about. They also examine whether training effects differ by gender and by computing/non-computing major.","tokens_in":6479,"tokens_out":5699,"duration_ms":54252,"significance":"The topic is timely, and the idea of translating expert sociolinguistic cues from algorithmic deepfake detection into human auditory training is a useful and interdisciplinary contribution. The EDLFs are grounded in prior work rather than fitted to the current outcome data, and the study uses a relatively large classroom sample. However, the paper's central claim is not supported by its own inferential statistics: Table 2 shows no significant experimental-group improvement in accuracy on clips that were initially marked unsure, and the only significant experimental effect, decreased unsurety, is a change in response criterion rather than evidence of improved discernment. The pre/post use of identical clips and the nonrandom assignment to condition further prevent causal interpretation. As submitted, the headline claim is inaccurate, although a repositioned exploratory report on response bias and media literacy could have value.","major_comments":[{"comment":"The abstract's claim that the experimental group showed an improvement in their ability to correctly identify clips they were initially unsure about is contradicted by the paper's own paired tests. In Table 2, among students who were unsure on at least one pre-test clip, the experimental group's mean paired difference in accuracy is 0.02 for all clips, 0.04 for real clips, and 0.01 for fake clips, none of which is marked significant; the control group's corresponding all-clip difference is 0.05 and is significant. No reported experimental accuracy contrast reaches significance, so the central claim of improved discernment is unsupported.","section":"Abstract; §3.2, Table 2"},{"comment":"Because the post-survey uses the same 20 audio clips as the pre-survey, any pre/post accuracy change confounds training with memory and practice effects. This is not a hypothetical concern: the control group also improves significantly on the same measure (Table 2, all clips, 0.05, marked significant), which indicates that repeated exposure alone can produce the outcome attributed to the intervention. The design therefore cannot distinguish training-induced improvement in general discernment from item-specific familiarity.","section":"§2.1 and §2.3"},{"comment":"The only statistically significant experimental effect reported is the decrease in unsurety (Table 1, All Students, experimental group: -0.02, marked significant). A decrease in unsure responses is a change in decision threshold, not evidence of improved perceptual accuracy. The supplementary observation that 85% of initially unsure fake clips but only 20% of initially unsure real clips were correctly identified in the post-test is not a paired significance test; given that the stimulus set is 80% fake, this pattern is compatible with a learned bias toward labeling uncertain clips as fake, which would leave overall accuracy unchanged while shifting real-clip accuracy downward.","section":"§3.2 and Table 1"},{"comment":"Assignment to experimental and control conditions was at the level of course sections and depended on instructors' willingness to participate; students were not randomly assigned, and no baseline equivalence between groups is reported. Differential attrition (264 enrolled but only 129 completing all phases) and the lack of pre-survey accuracy comparisons between groups make it difficult to attribute the observed differences to the training module.","section":"§3.1"},{"comment":"The many subgroup tests in Tables 1 and 2 (gender, English first language, fluency, and major) are performed without correction for multiple comparisons and include very small groups, such as control females with N=7, English-not-first-language control with N=7, and non-computing control with N=8. These significant results are likely to include false positives, and the paper's RQ2 and RQ3 conclusions should be treated as exploratory; additionally, no effect sizes or confidence intervals are reported for any of the paired differences.","section":"§3.2, Tables 1–2"}],"minor_comments":[{"comment":"The abstract says the study evaluated 264 students, but §3.1 reports that only 129 students completed all three phases; this discrepancy should be stated transparently, and calling the sample a representative cross section of all students at UMBC is not supported by the described convenience sample of nine course sections.","section":"Abstract and §3.1"},{"comment":"Please provide a formal definition of the unsurety rate and specify how it is computed in the table captions, since the text only states that unsure responses were counted as incorrect for accuracy.","section":"Tables 1 and 2"},{"comment":"The control-group paragraph reports a significant improvement for clips shorter than 2 seconds, but no analysis, table, or definition of this duration split appears elsewhere in the results section.","section":"§3.2, Control Group paragraph"},{"comment":"The open-ended questions about familiarity with audio deepfakes are described in the pre-survey methodology but are not analyzed or reported in the results; either include that analysis or remove the mention from the methodology.","section":"§2.1 and §3"},{"comment":"The 4:16 real-to-fake imbalance in the stimulus set should be discussed as a design limitation, because it directly affects accuracy rates and the interpretation of changes in unsurety.","section":"§2.1"}],"recommendation":"reject","confidential_remarks":"The paper's central claim is contradicted by its own Table 2, and the reuse of identical test clips makes the pre/post comparison uninterpretable as a measure of generalizable discernment. I do not see how a revision within the scope of the current data could salvage the central claim; a new study with unseen post-test audio and random assignment would be required. The paper may still have value as an exploratory report on response bias and media literacy if substantially rewritten."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The headline you need: this paper is a real pilot study with one clean, small result—sociolinguistic training reduces expressed unsurety—but the abstract overclaims an accuracy improvement that the paper's own statistics do not support. Treat it as evidence about listener confidence, not listener ability.\n\nWhat's actually new here is the intervention itself. Instead of informational training or a game, the authors trained students to listen for five expert-defined phonetic features (pitch, pause, stop bursts, breath, audio quality). That is a concrete, reproducible module, and I don't see a prior human-listener study doing exactly that. The paper also does some things well: a pre/post design with 264 enrolled and 129 complete cases, subgroup analyses by gender and major, and—importantly—a results section that openly states the unsurety decrease \"did not always co-occur\" with an accuracy increase. The authors are not hiding the null accuracy result; the abstract just ignores it.\n\nThe soft spots are structural. The same 20 clips are used pre and post a month apart, so memory and practice effects are baked into every comparison. Assignment to control versus experimental was by intact course sections chosen by instructor willingness, not random. They run many small subgroup tests without any correction, so the handful of significant stars should be read as suggestive. And the central claimed finding—improvement on clips students were initially unsure about—is directly contradicted by Table 2: the experimental group's mean paired difference in accuracy on those clips is 0.02 and not significant, while the control group's is 0.05 and significant. The paper then highlights that 85% of initially-unsure fake clips but only 20% of initially-unsure real clips were correctly identified post-test. That is not a paired significance test, and it is exactly what you would expect from a learned bias toward labeling clips fake in an 80%-fake stimulus set. So the accuracy gain is likely response bias, not discernment.\n\nWho gets value from this? Researchers working on human factors in deepfake detection, and educators designing media literacy interventions. But the value is as a proof-of-concept for a training approach, not as evidence that the training works. The citation pattern is fair; no self-citation abuse beyond their own prior EDLF work.\n\nMy recommendation: send it to peer review. The research question matters, the intervention is novel, and a good referee can force the authors to fix the design and scale the claims back. But it should not be accepted in anything close to this form. If we cite it at all, we cite it for the unsurety effect and the EDLF module, not for improved discernment.","headline":"A modest, honestly reported pilot study whose abstract overclaims an accuracy benefit its own Table 2 contradicts; the real finding is reduced unsurety, not improved discernment.","tokens_in":7002,"tokens_out":1651,"would_cite":false,"duration_ms":18012,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that a short training module teaching listeners to attend to five expert-defined linguistic cues significantly reduces uncertainty when judging audio deepfakes and improves how often previously 'unsure' clips are later…","keywords":["audio deepfake","deepfake detection","sociolinguistics","phonetic features","human perception","training module","undergraduate education","misinformation"],"falsifier":"Run the experiment again with two randomly selected sets of 20 clips (set A for pre-test, set B for post-test) drawn from the same datasets; if the experimental group's accuracy advantage over the control group disappears or reverses, the claim that the training improves discernment rather than just confidence would be falsified.","tokens_in":6073,"feed_emoji":"🎧","tokens_out":7133,"duration_ms":60547,"temperature":0.7,"pith_summary":"This paper tests whether a 15–20 minute instructional module that teaches undergraduate students to listen for five expert-defined linguistic features—pitch, pauses, stop-consonant bursts, audible breath, and overall audio quality—improves their ability to tell real speech from AI-generated audio. In a pre/post design across experimental and control groups, the training produced a statistically significant drop in how often students answered 'unsure' when judging audio clips, and the authors read this as improved confidence and improved identification of clips students had earlier been unsure about. The paper positions this as the first sociolinguistics-informed approach to human audio deepfake discernment, building on earlier work where the same features improved AI detection. The authors argue that because listeners are often the weakest link in defending against audio deepfakes, giving them concrete linguistic cues is a promising complement to algorithmic detection.","feed_headline":"Audio deepfake training makes students surer, study finds","feed_subtitle":"A 20-minute module on pitch, pauses, and breath cues trimmed 'unsure' responses, though accuracy gains were modest.","key_machinery":"The central mechanism is the Expert-Defined Linguistic Features (EDLFs): five phonetic and phonological cues (pitch, pause, word-initial or word-final stop-consonant bursts, audible intake or outtake of breath, and overall audio quality) selected by two sociolinguistics experts after reviewing 344 real and fake English audio samples. The EDLFs are taught in a four-part instructional module and serve as a listening protocol, giving ordinary listeners concrete, named things to attend to while judging whether speech is real or synthetic. The measurement machinery is a pre-post design using the same 20 audio clips in both surveys, with paired differences in accuracy and unsurety as the outcome statistics.","core_discovery":"The paper's central claim is that sociolinguistically informed training improves listeners' audio deepfake discernment, primarily by reducing uncertainty and by sharpening their judgment of clips they were initially unsure about: 85% of previously 'unsure' clips that were actually fake were later correctly labeled as fake after training, while only about 20% of such real clips were. The mechanism is the Expert-Defined Linguistic Features (EDLFs), five phonetic and phonological cues selected by sociolinguists from 344 real and fake samples: pitch anomalies, unexpected pauses, abnormal stop-consonant release bursts, audible breath, and degraded audio quality. The experimental group showed a significant decrease in unsurety, but the paper reports that this decrease did not always co-occur with an increase in discernment accuracy; the control group, which merely read an article about deepfakes, showed a significant gain in overall accuracy, especially on real clips. The authors conclude that the training increases listener confidence, that effects vary by demographics such as gender and first-language status, and that holistic, interdisciplinary training is a worthwhile direction for audio misinformation.","pith_inferences":["The reported statistics indicate that the reliable, significant effect of the training is reduced uncertainty rather than improved overall accuracy; the experimental group's mean accuracy changes in Tables 1 and 2 are not all statistically significant, so the abstract's claim of improved discernment may overstate the data.","A natural next study would use a fresh set of audio clips in the post-survey to rule out memory effects; if the accuracy gains disappear, the mechanism might be familiarity rather than skill acquisition.","The control group's accuracy improvement hints that the active ingredient could be general attentiveness or metacognitive reflection from having taken the survey twice, rather than the specific linguistic cues; a third group receiving generic critical-listening instructions would isolate this.","The training's skew toward labeling uncertain clips as fake suggests a possible societal trade-off: it may protect against fraud but could increase suspicion toward legitimate voices, with consequences for trust in audio content."],"forward_implications":["A short, low-cost training module can make listeners more confident and decisive when judging whether audio is real, which may help in everyday exposure to social media content.","Because the control group improved accuracy from a brief article, even minimal deepfake education appears to sharpen judgment of real audio clips, supporting digital media literacy efforts.","The differential effects by gender and first-language status suggest training materials could be tailored to specific listener populations, such as English-language learners.","The observed bias away from labeling clips as real (85% correct on fake clips vs. 20% on real clips) implies training may induce skepticism toward audio, which could reduce false trust but also increase false alarms."],"supporting_citations":[{"why":"Supplies the five Expert-Defined Linguistic Features and the earlier finding that integrating them improved AI spoofed-audio detection, motivating the human training.","marker":"[5]"},{"why":"Provides prior evidence that humans cannot reliably detect speech deepfakes, the baseline the training aims to beat.","marker":"[3]"},{"why":"Shows AI algorithms outperforming or matching untrained human listeners in an online game, defining the human-perception gap.","marker":"[4]"},{"why":"Provide the standard machine-learning datasets from which the 20 real and fake audio clips in the pre/post survey were randomly selected.","marker":"[6-11]"},{"why":"Is the article given to the control group, serving as the comparison condition for the training module.","marker":"[12]"}],"fun_headline_variants":["Linguistic cue training cuts student uncertainty on audio deepfakes","Audio deepfake training makes students surer, but not more accurate","For undergrads, linguistic training reduces audio deepfake unsurety","Hearing fakes: sociolinguistic cues boost listener confidence","Trained ears: linguistic cues help students spot audio deepfakes they doubted"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The study relies on the same 20 audio clips being presented in both the pre- and post-surveys, so any improvement could reflect memory of the clips or practice effects rather than a genuine gain in discernment ability.","fun_headline_variants_meta":{"raw":{"variants":["Linguistic cue training cuts student uncertainty on audio deepfakes","Audio deepfake training makes students surer, but not more accurate","For undergrads, linguistic training reduces audio deepfake unsurety","Hearing fakes: sociolinguistic cues boost listener confidence","Trained ears: linguistic cues help students spot audio deepfakes they doubted"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000962,"raw_usage":{"total_tokens":4140,"prompt_tokens":1035,"completion_tokens":3105,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":651,"completion_tokens_details":{"reasoning_tokens":3013}},"tokens_in":651,"tokens_out":3105,"duration_ms":23676,"temperature":1.0,"reasoning_tokens":3013,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T15:07:05.515744+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the experiment again with two randomly selected sets of 20 clips (set A for pre-test, set B for post-test) drawn from the same datasets; if the experimental group's accuracy advantage over the control group disappears or reverses, the claim that the training improves discernment rather than just confidence would be falsified.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the five Expert-Defined Linguistic Features and the earlier finding that integrating them improved AI spoofed-audio detection, motivating the human training."},{"cited_title":"Warning: Humans cannot reliably detect speech deepfakes","cited_arxiv_id":null,"evidence_quote":"Provides prior evidence that humans cannot reliably detect speech deepfakes, the baseline the training aims to beat."},{"cited_title":"Human perception of audio deepfakes","cited_arxiv_id":null,"evidence_quote":"Shows AI algorithms outperforming or matching untrained human listeners in an online game, defining the human-perception gap."},{"cited_title":"Who are you (i really wanna know)? detecting audio{DeepFakes} through vocal tract reconstruction","cited_arxiv_id":null,"evidence_quote":"Is the article given to the control group, serving as the comparison condition for the training module."}],"review_version":1}