{"id":"1504d1f7-8776-4c3f-9d25-3360b6490581","arxiv_id":"2507.01081","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"An AI-guided version of the Imagery Competing Task Intervention, with pupillometry, reduced self-reported intrusive memories after analogue trauma in a preregistered randomized experiment.","lead":"Researchers tested an AI chatbot that guided healthy volunteers through a memory-reduction exercise while tracking pupil size. People who played the visuospatial game reported fewer intrusive memories of a traumatic film over the next week than those who listened to a podcast.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"In-session vigilance-intrusion task shows no group difference (p=0.63, direction reversed), undercutting the diary-based central claim unless differential reporting is ruled out.","rationale":"The reader's weakest assumption concerns the validity and equivalence of the self-report diary outcome; this stress-test identifies a concrete internal inconsistency—the in-session vigilance-intrusion task showing no group difference—that directly supports that assumption. The paper is transparent about this null result but does not integrate it into the interpretation of the primary outcome. This does not overturn the preregistered diary finding, which is robust to several sensitivity analyses, but it does reinforce the need for caution and further evidence (e.g., data availability, or a preregistered analysis plan for the vigilance task). The reader's CONDITIONAL verdict remains appropriate; no change to the verdict is needed, but the condition should explicitly require addressing the discrepancy between the diary and the vigilance-intrusion task. I set agreement_with_reader to 'agree' because my concern is the same underlying issue, specified with additional evidence from within the manuscript.","tokens_in":23863,"tokens_out":3572,"duration_ms":39104,"concrete_test":"Perform an ANCOVA on total diary intrusions with condition as the predictor and in-session vigilance-intrusion count as a covariate, and compute within-group Spearman correlations between diary and vigilance counts. Additionally, compute a Bayes factor for the vigilance-task group comparison (intervention vs control) to quantify evidence for the null. If the diary effect remains substantial after adjusting for the vigilance count, or if the vigilance null is underpowered (e.g., BF inconclusive), the concern is weakened; if the diary effect attenuates to null or the BF favors absence of an effect, the central claim would require reinterpretation.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that ANTIDOTE reduces intrusive memories rests on the 7-day self-report electronic diary. The weakest assumption is that diary counts faithfully and equivalently reflect the true frequency of intrusive memories in both conditions. The manuscript itself contains a direct, more controlled test: the vigilance-intrusion task completed in the same session (Supplementary Results). Reported values are intervention #=51.10 [41.90, 60.46], n=40; control #=48.04 [40.08, 66.46], n=48; p=0.63, with the effect direction opposite to the diary outcome. The paper flags this as differing from prior work (Lau-Zhu et al. 2019) but does not resolve the tension. If the intervention genuinely reduced intrusions, one would expect at least a trend on this in-session, less memory-demanding measure; its absence suggests the diary difference could reflect differential reporting or demand characteristics. The intervention group received AI-guided instruction about the intervention rationale and diary importance, while the control group did not receive equivalent engagement during the podcast task, creating an asymmetry in expectations. This is specifically acknowledged in the Discussion ('AI guidance was not used in the podcast listening control task, which may have introduced an imbalance in instructional engagement between conditions'). Given the primary outcome is entirely self-report and the only objective behavioral measure shows no effect, the central claim is not yet secure.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports a preregistered randomized controlled experiment (N=100 healthy volunteers) testing ANTIDOTE, a fully automated AI-guided digital intervention combining the Imagery Competing Task Intervention (ICTI) with a GPT-4 chatbot guide and pupillometry, against an active control (podcast listening) in the trauma film paradigm. The primary outcome is the total number of intrusive memories self-reported in a 7-day electronic diary. The authors report that intervention participants recorded significantly fewer intrusions than controls (11.62 vs. 21.00; one-tailed p=0.007; Cohen's d=0.49), that human and AI ratings confirmed competent AI-guided instruction, and that pupil size increased with game difficulty and predicted diary-reported intrusions in exploratory analyses. The central claim is that ANTIDOTE reduces intrusive memories, with supporting evidence that the AI can deliver the intervention and that pupillometry may track engagement.","tokens_in":24122,"tokens_out":3673,"duration_ms":42506,"significance":"If the central claim holds, this is an important proof-of-concept: it suggests that a fully automated, AI-guided digital intervention with physiological monitoring can reproduce the effects of a human-guided evidence-based intervention, with substantial scalability implications. The study has notable strengths: the primary hypothesis and analysis were preregistered; the primary result is robust to permutation tests, parametric tests, and several sensitivity analyses; AI-guided conversations were independently graded by human raters using a clinical rubric; and the analysis code and data are promised to be openly available. The pupillometry results, especially the difficulty–pupil relationship and its predictive link to diary outcomes, are a novel contribution, albeit exploratory. However, the current evidence is not yet secure because the primary outcome relies entirely on unverified self-report in a naturalistic setting, and the one in-session behavioral measure of intrusions shows no group difference.","major_comments":[{"comment":"The primary outcome is the total number of self-reported intrusive memories entered into a Google Sheets diary on participants' personal devices over 7 days. There is no objective verification of diary compliance or accuracy, and the active control group did not receive AI-guided instruction during the podcast task, an asymmetry the authors acknowledge in the Discussion. This creates a plausible alternative explanation: the group difference may reflect differential motivation, attention, or demand characteristics rather than a genuine reduction in intrusive memories. To make the central claim convincing, the authors should provide additional evidence ruling out reporting bias—for example, diary metadata (timestamps, daily email opens, 'No Intrusions' selections), a per-day compliance analysis, or a comparison of results restricted to participants with objective evidence of engagement. Without such evidence, the primary result is consistent with, but does not establish, a true reduction in intrusion frequency.","section":"Methods, 'Intrusive memory reporting'; Results, 'Reduction of intrusive memories'; Discussion, 'AI Guidance'"},{"comment":"The vigilance-intrusion task, completed in the laboratory at the end of the experimental session, shows no group difference in intrusions (intervention 51.10, control 48.04; p=0.63), with the direction opposite to the diary outcome. This task is a more controlled, less memory-demanding measure of the same construct, and the authors note that prior work has observed group differences on this task. The paper's response—that remote diaries are the 'clinical gold standard'—does not resolve the tension; if the intervention genuinely reduced intrusions, one would expect at least a trend in the in-session measure. The authors should address this discrepancy head-on, for example by reporting the correlation between diary and vigilance-task intrusions within each condition, discussing power or task-sensitivity limitations, or explicitly framing the vigilance-intrusion result as a failed manipulation check with implications for the interpretation of the primary outcome.","section":"Supplementary Results, 'Comparing in-person and remote assessments of intrusions'"},{"comment":"The sentence 'The total number of intrusive memories across participants did not reliably change (Δ=1.15 entries, 95% CIs [0.35, 2.15], n=100; p=0.68)' is internally inconsistent: a 95% CI that excludes zero should correspond to a p-value below 0.05, not p=0.68. This appears to be a reporting error, but it affects a quality-control result that is used to argue the primary analysis is robust. The authors should correct the reported statistic and clarify whether the analysis refers to total counts, per-participant differences, or something else.","section":"Results, 'Evaluating AI guidance', paragraph on quality control"}],"minor_comments":[{"comment":"The color coding for conditions is inconsistent: Figure 1 uses red for the intervention group and blue for the control group, while Figure 2a and 2b use blue for the intervention and red for the control. Please standardize the color scheme across all figures.","section":"Figure 1 and Figure 2 captions"},{"comment":"The text contains a typo: 'muli-level analysis' should be 'multi-level analysis'.","section":"Results, 'Evaluating AI guidance'"},{"comment":"The phrase 'inhibiting the preprotent response' should read 'inhibiting the prepotent response'.","section":"Methods, 'Vigilance-intrusion task'"},{"comment":"The description of the AI guide's blinding states that prompt content was 'identical across groups, except for the task-specific prompt for the intervention phase.' If the prompt differs for the intervention phase, the AI guide is not fully blinded to condition during that conversation; this should be clarified to avoid overstating the blinding procedure.","section":"Methods, 'Evaluating AI guidance'"},{"comment":"The reported mood changes (sadness, depression, hopelessness) use the labels s, d, h; please ensure these abbreviations are defined or not used ambiguously, as 's' is also used for the AI-guidance score.","section":"Supplementary Results, 'Mood induction'"}],"recommendation":"major_revision","confidential_remarks":"The paper addresses a timely and important topic, and the experimental design is more rigorous than typical for a first demonstration of an AI-guided digital intervention. My main concern is that the central claim rests on a self-report diary measure that is vulnerable to differential demand and engagement, and the single in-session behavioral measure does not corroborate the effect. I would be willing to accept a revision that either provides additional objective or quasi-objective evidence (e.g., diary compliance metrics, a pre-registered sensitivity analysis on engaged participants) or substantially tempers the causal claim to reflect the current evidence. The internal inconsistency in the quality-control statistic (Δ vs. p) should also be corrected. Given the potential impact and the otherwise sound methodology, major revision seems appropriate."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is the first test of a fully automated, AI-guided and pupillometry-monitored version of the imagery-competing task (ICTI), and the 7-day diary outcome supports the preregistered hypothesis. But the one objective in-session measure of intrusions showed nothing, and the paper doesn't resolve what that means for the self-report result.\n\nWhat's genuinely new: prior ICTI required a human guide; here the LLM delivered all instructions, and human raters independently scored the conversations as competent. The preregistration is real, the sample meets the target, and the primary permutation test survives robustness checks. The pupil-difficulty relationship within the game is a nice internal validation that the physiological measure tracks something.\n\nThe soft spots are concentrated in the outcome. The electronic diary is the clinical gold standard for intrusive-memory assessment, but it is still self-report, and the two arms differed in how much engaging instruction they received. The control group listened to a podcast with static instructions; the intervention group got interactive AI guidance. That imbalance alone could explain differential reporting or attention. More telling is the supplementary vigilance-intrusion task: no group difference (p=0.63), with the effect going the wrong way. The paper flags this but moves on. It's not dispositive, since the task is brief and the diary is the preregistered outcome, but a reader should not assume the effect is robust until this is squared away.\n\nA smaller issue: the conclusion says the AI intervention \"replicates\" human-guided ICTI, but no human-guided arm was run here; the comparison is just to prior studies. Also, data and code are not yet public, so the analyses cannot be independently checked.\n\nBottom line: this is a worthwhile proof-of-concept for AI-guided delivery, and the physiological monitoring is a promising direction. It deserves peer review, but the claims in the Discussion should be tempered until the reporting-bias concern is addressed, ideally with a matched control condition or a demonstration that diary compliance and entry validity did not differ by condition. I would send it to a serious referee, expecting revision on interpretation rather than on the basic experimental execution.","headline":"First fully automated AI-guided ICTI study with a promising diary result, but the null on the in-session vigilance task leaves the central claim less secure than the authors suggest.","tokens_in":24604,"tokens_out":2572,"would_cite":true,"duration_ms":137758,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that a fully automated AI guide paired with pupillometry can deliver an evidence-based trauma intervention and reproduce its reduction of intrusive memories in a lab model.","keywords":["intrusive memories","generative AI","large language models","pupillometry","digital therapeutic","imagery competing task intervention","trauma film paradigm","post-traumatic stress disorder"],"falsifier":"The claim would be falsified if a preregistered replication that adds objective diary compliance tracing, such as time-stamped entry metadata or passive recording of diary use, and an equally interactive AI-guided control task found no intrusion difference between groups, or found that the diary difference vanished once compliance was statistically equated.","tokens_in":23698,"feed_emoji":"🧠","tokens_out":7282,"duration_ms":80395,"temperature":0.7,"pith_summary":"This paper tests whether a fully automated system can replace the human guide in a brief psychological intervention for intrusive memories. One hundred healthy volunteers watched traumatic video clips, then either played a mental-rotation block puzzle game or listened to a neutral podcast, with an AI chatbot delivering instructions and an eye tracker recording pupil size throughout. The preregistered result is that intervention participants reported fewer intrusive memories over the following week, 11.62 versus 21.00 in the active control (one-tailed p = 0.007, Cohen's d = 0.49), a reduction of roughly 45 percent. The authors also report that AI-delivered instruction earned competent ratings on a clinical rubric and that pupil size during gameplay tracked task difficulty and predicted fewer later intrusions. If the finding holds, generative AI combined with physiological monitoring could let an evidence-based trauma intervention be delivered without a trained human guide.","feed_headline":"AI guide plus eye tracking cuts trauma-film intrusions 45%","feed_subtitle":"A chatbot-led visuospatial task cut intrusive memories by 45 percent versus a podcast control in 100 volunteers.","key_machinery":"The load-bearing object is the Imagery Competing Task Intervention (ICTI), a two-part procedure in which a brief memory reminder reactivates the trauma memory and a visuospatial task, here a falling-block puzzle game demanding mental rotation and planning, disrupts reconsolidation. ANTIDOTE automates the two roles of the human guide with a structured large-language-model chatbot that teaches each protocol step and checks comprehension through summarization and corrective feedback, and with a screen-mounted eye tracker that records pupil size as an online index of cognitive effort. The pupil signal does double duty: it provides evidence that participants engaged the intended cognitive strategies, and it predicts how many intrusions each participant later reports.","core_discovery":"The central claim is that a fully automated system can deliver the Imagery Competing Task Intervention (ICTI) and reproduce its known effect on intrusive memories in an experimental analogue of trauma. In a preregistered between-subjects study of 100 healthy adults, participants who received the AI-guided visuospatial gameplay intervention recorded significantly fewer intrusive memories in a 7-day electronic diary than participants who listened to a neutral podcast, and this difference was robust to the paper's quality-control and outlier exclusions. The paper further claims that this automated delivery was not merely behavioral: the AI guide's instruction was rated competent on a rubric used to train human therapists, human and AI grading converged, and pupil size rose with game difficulty and predicted outcome in the expected direction, suggesting the system can monitor the internal cognitive engagement that a human guide would normally observe.","pith_inferences":["I would be cautious about generalizing beyond the laboratory: the paper's active control differs from the intervention in both modality and interactivity, so demand characteristics or differential engagement could account for part of the diary difference.","A natural extension the paper does not test is a closed-loop system that uses pupil size in real time to adjust game difficulty or end memory reactivation when effort drops; such a design could be compared directly with fixed ANTIDOTE.","The absence of a group difference in the in-session vigilance-intrusion task while the week-long diary showed a difference suggests the effect may operate on later consolidation or on reporting behavior rather than immediate intrusion production; time-stamped diary metadata could separate these.","Replacing the passive podcast control with an equally interactive AI-guided non-visuospatial control would isolate the causal contribution of visuospatial disruption from general engagement with the AI."],"forward_implications":["A version of the intervention could be delivered remotely with no clinician in the loop, removing the trained human guide as the scaling bottleneck.","Pupil size can serve as a compliance and engagement check in place of a human observer, flagging whether a participant is actually performing mental rotation during gameplay.","AI-based grading of instructional conversations tracked human grading, so treatment fidelity could be audited at scale without manual review of every chat.","The effect size closely parallels human-guided ICTI laboratory studies, suggesting automation does not obviously trade away the intervention's known benefit.","The reduction was consistent across all seven diary days, so the effect appears to be a sustained change rather than a single-day spike."],"supporting_citations":[{"why":"Establishes the trauma film paradigm and shows Tetris-style gameplay reduces intrusive memories via reconsolidation-update mechanisms; this is the effect ANTIDOTE aims to automate.","marker":"James et al. 2015"},{"why":"Defines the experimental analogue-trauma model that the present protocol uses to evaluate interventions.","marker":"James et al. 2016"},{"why":"Supplies the scoring rubric and clinical precedent for human-guided ICTI used to grade the AI guide's conversations.","marker":"Kanstrup et al. 2024"},{"why":"Proof-of-concept randomized controlled trial in emergency department patients showing human-guided ICTI reduces intrusions after real trauma, the clinical benchmark.","marker":"L. Iyadurai et al. 2018"},{"why":"Randomized controlled trial in COVID-19 intensive care staff supporting ICTI effectiveness and the clinical expectation ANTIDOTE is compared against.","marker":"Lalitha Iyadurai et al. 2023"},{"why":"Bayesian adaptive trial with healthcare workers that shapes the memory reminder design and clinical implementation of ICTI.","marker":"Ramineni et al. 2023"},{"why":"Prior trauma-film study whose intrusion counts provide the comparison for the effect size reported here.","marker":"Lau-Zhu, Henson, and Holmes 2021"},{"why":"Reviews pupil dilation as a cognitive-effort index, the interpretive foundation for the pupillometry analyses.","marker":"van der Wel and van Steenbergen 2018"},{"why":"Proposed the original cognitive-science rationale that playing Tetris after trauma reduces flashback build-up.","marker":"Holmes et al. 2009"}],"fun_headline_variants":["AI and pupil tracking cut trauma-film intrusions 45%","Automated AI guide plus pupil check cuts intrusive memories 45%","AI-delivered therapy with eye tracking reduces trauma intrusions 45%","Bot-led ICTI with pupillometry lowers intrusive memories 45%"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The primary outcome rests on participants' own electronic diary entries made on personal devices over 7 days, with no objective verification that both groups completed the diaries equally diligently or accurately, so a difference in motivation or in what participants think the study wants could produce the reported difference.","fun_headline_variants_meta":{"raw":{"variants":["AI and pupil tracking cut trauma-film intrusions 45%","Automated AI guide plus pupil check cuts intrusive memories 45%","AI-delivered therapy with eye tracking reduces trauma intrusions 45%","Bot-led ICTI with pupillometry lowers intrusive memories 45%"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000871,"raw_usage":{"total_tokens":3731,"prompt_tokens":864,"completion_tokens":2867,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":480,"completion_tokens_details":{"reasoning_tokens":2791}},"tokens_in":480,"tokens_out":2867,"duration_ms":21986,"temperature":1.0,"reasoning_tokens":2791,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T20:59:45.683925+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"The claim would be falsified if a preregistered replication that adds objective diary compliance tracing, such as time-stamped entry metadata or passive recording of diary use, and an equally interactive AI-guided control task found no intrusion difference between groups, or found that the diary difference vanished once compliance was statistically equated.","supporting_citations":[],"review_version":1}