{"id":"0d17f236-ab4a-4961-aa86-f066d7415eb1","arxiv_id":"2608.09719","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Seeing a video of yourself as the narrator made history feel closer and more engaging, but it did not help and may have hurt immediate quiz performance.","lead":"This study had 36 learners watch history videos narrated by an AI-generated avatar of their own face and voice, and compared that with a generic narrator. The self avatar boosted feelings of immersion and closeness, but quiz scores were actually lower, revealing a trade-off between engagement and immediate learning.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Stimulus-quality confound is unaddressed: self videos were made from casual photos and 10–15 s recordings via face/voice cloning, while non-self videos used an expert-designed character, so higher uncanniness and lower quiz scores may reflect production artifacts rather than self-reference.","rationale":"The reader's conditional verdict is appropriate, and my stress-test converges on the same weakest link: stimulus-quality equivalence. The central claim has two parts: (1) the Digital Self agent improves experiential measures, and (2) it does not improve and may reduce short-term retention. Both parts are between-condition comparisons, and both are interpretable as self-reference effects only if the videos are matched on everything except whose face and voice are used. The threat is concrete: the self condition necessarily uses lower-bound source material (a casual photo and 10–15 s recording) through a cloning pipeline, whereas the non-self condition uses an expert-designed character. The paper states both agents were generated through the same pipeline, but this controls software, not output quality; cloning quality depends heavily on source input, and no objective or subjective quality check or manipulation check is reported. This is not an external-consensus disagreement; it is an internal control gap. A blinded quality-rating study would settle it. Multiple-comparison control is a secondary worry (ten tests at α = .05 without correction, and several p-values are modest), but the stimulus confound is more load-bearing because it threatens construct validity, not just inference. I would keep the reader's CONDITIONAL verdict rather than escalating to REJECT: the qualitative data and large effect sizes make the phenomenon worth reporting, but the current evidence cannot cleanly isolate self-reference from video quality. No change to the reader's verdict is needed.","tokens_in":7504,"tokens_out":4509,"duration_ms":39486,"concrete_test":"Run a blinded perceptual-quality rating study on the exact stimulus videos (or representative 20 s clips from each condition for all 36 participants). Independent raters who are naive to the manipulation and do not know the participants rate each clip on lip-sync precision, voice naturalness, visible artifacts, and overall production quality, with audio-only and muted-video passes to separate voice from visual quality. If the Digital Self clips score significantly worse on any quality dimension, the higher UVS and lower quiz scores in Table 1 cannot be cleanly attributed to identity mirroring; if the clips are rated equivalent, the confound is substantially weakened.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central empirical claim—that self-mirroring increases experiential engagement but lowers immediate retention—requires that the only systematic difference between conditions is identity mirroring. Section 3.2 does not establish this. The Digital Self stimulus is generated by face transfer and voice cloning from the participant's casual frontal photo and 10–15 s recording; the non-self agent is an expert-designed standardized female character with confirmed Lingnan features. Stating that both agents were generated through the same pipeline does not equate their output quality, because input fidelity differs: a casual selfie and short recording are more likely to yield lip-sync errors, voice artifacts, and visual uncanniness than professionally designed character assets. No manipulation check, video-quality rating, or artifact-absence check is reported in Section 4. Table 1 shows UVS eeriness is dramatically higher in the self condition (d = .957) and History Retention is lower (d = −.413); both effects are exactly what a technical-quality confound would predict. If the self videos are merely lower-fidelity, the paper's design implications (adaptation phase, selective presence, stylized abstraction) may target the wrong cause, and the conclusion that self-identity decouples experience from learning is unsupported. The qualitative theme of attention being drawn to the agent (P22, P28) is consistent with artifact-driven distraction as well as with identity novelty.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces an 'Ancestral Digital Self': an AI-generated pedagogical agent, presented in prerecorded video, that mirrors a learner's facial features and vocal timbre within a historical narrative. In a within-subjects study with 36 participants, the authors compare this Digital Self to a non-self pedagogical agent narrating Wubeiling-culture content. Quantitative results show that the Digital Self condition increases several experiential measures (self-other inclusion, agent perception, narrative transportation, relatedness) but also raises uncanny-valley eeriness, while immediate history-retention quiz scores are lower; Remember/Know judgments do not differ. Interview data suggest that self-similarity increases familiarity and motivation, but that novelty and uncanniness can draw attention away from content. The paper concludes that self-mirroring decouples experiential engagement from short-term learning outcomes and offers design implications (adaptation phase, selective presence, stylized abstraction).","tokens_in":88,"tokens_out":5760,"duration_ms":305010,"significance":"If the reported effects are causal, the paper makes a useful contribution to personalized pedagogical-agent design: it offers a reproducible AI-generation workflow, uses an appropriate within-subjects design with counterbalancing, and combines quantitative and qualitative evidence. The core claims are not circular: the experience and outcome measures are independent instruments with no fitted parameters. The potential significance is real for HCI and learning-technology audiences. However, the causal interpretation is currently threatened by stimulus-fidelity and gender confounds and by uncorrected multiple testing, so the significance depends on whether those concerns can be addressed.","major_comments":[{"comment":"The crucial comparison assumes that the only systematic difference between conditions is identity mirroring, but §3.2 does not establish this. The Digital Self was generated by face transfer and voice cloning from the participant's casual frontal photograph and 10–15 s recording, while the non-self agent was an expert-designed standardized character. Stating that both agents were generated through the same pipeline does not equate their output fidelity: a casual selfie and a short recording are more prone to lip-sync errors, voice artifacts, and visual uncanniness than a professionally designed character. Table 1 in §4.1 shows exactly the pattern such a confound would predict (UVS d = .957, higher eeriness; History Retention d = −.413, lower performance), and no manipulation check, video-quality rating, or artifact-absence check is reported. The qualitative theme of attention being drawn to the agent (P22, P28) is also consistent with artifact-driven distraction. The paper should report per-video quality ratings, an artifact check, or a matched-fidelity control before attributing these outcomes to self-reference; at minimum, the causal framing in §5 must be softened.","section":"§3.2, Table 1"},{"comment":"The non-self condition is a standardized female character, whereas the Digital Self mirrors each participant's own gender and voice. Because 21 of the 36 participants were male, narrator gender is confounded with condition for the majority of the sample, so the IOS, API, IMI, and UVS differences could reflect gender/voice matching rather than self-identity. No subgroup analysis by participant gender is reported. The authors should use a gender-matched non-self agent, test for condition-by-gender interactions, or explicitly justify why gender mismatch is not an alternative explanation for the Table 1 effects.","section":"§3.2, §3.1"},{"comment":"The analysis reports ten outcome tests with no multiple-comparison correction and no pre-specified primary endpoints. Under a simple Bonferroni correction (α = .005), History Retention (p = .018), API (p = .027), and NT (p = .012) are no longer significant, so the central claims of lower retention and of several experiential benefits rest on uncorrected p-values. This is load-bearing because the 'decoupling' conclusion in §5.2 depends on the retention difference. The authors should apply a family-wise or FDR correction, pre-register primary outcomes, or explicitly label the uncorrected tests as exploratory.","section":"§4.1, Table 1"},{"comment":"There is no manipulation check for successful self-recognition. Although participants saw their own face and heard their own voice, the identity manipulation is never verified quantitatively (e.g., with a recognition or self-identification item), and the only supporting evidence is retrospective interview quotes. A closed-ended manipulation check would substantially strengthen the claim that the observed effects are due to perceived self-mirroring rather than to novelty or to incidental features of the stimuli.","section":"§4.1, §3.4"}],"minor_comments":[{"comment":"Effect sizes are reported without confidence intervals; adding 95% confidence intervals for Cohen's d and r would improve interpretability, especially for the null R/K results.","section":"Table 1"},{"comment":"The within-subjects design is counterbalanced, but no order effects or content-set effects are reported; with two consecutive rounds, fatigue or practice could influence history-retention scores.","section":"§3.3"},{"comment":"The quizzes are said to have been piloted with Wubeiling-unfamiliar individuals to ensure comparable difficulty, but no pilot details, sample size, or equivalence statistics are reported.","section":"§3.4"},{"comment":"The Remember/Know/Guess procedure is not described in enough detail to be reproduced; the authors should specify the instructions, the guess option, and the scoring rule.","section":"§3.4"},{"comment":"Reference [22] appears to be about familiarity enhancing memory when novelty does not, which is a poor fit for the claim that a novelty effect draws attention to the avatar; please verify the citation or replace it.","section":"§5.2"},{"comment":"The workflow in Figure 2 would be more reproducible if the specific face-transfer, voice-cloning, and video-composition tools or parameter settings were identified.","section":"§3.2"}],"recommendation":"major_revision","confidential_remarks":"I concur with the reader's conditional assessment: the stress-test concern about stimulus quality is valid and is the main obstacle to accepting the paper's causal conclusions. The manuscript is within scope for a UbiComp companion paper, and the topic is timely. If the authors can add fidelity checks, address the gender confound, and reanalyze the outcomes with appropriate correction or explicit exploratory labeling, the central claim could become defensible; in its current form, the evidence does not yet rule out plausible artifact explanations."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a clean, honest 5-page study of a genuinely new application—using a learner's own face and voice in an AI narrator for history videos. The key empirical outcome, higher self-reported engagement but lower immediate quiz scores, is the kind of result people will cite. But the paper does not yet control for a plausible confound that I think the stress-test note nailed: the self videos are produced from casual photos and 10–15-second recordings via face/voice cloning, while the control agent is a polished, expert-validated character. If the self simply looks and sounds more artificial, the higher eeriness and the lower quiz scores are exactly what you'd expect without any identity effect. There's no manipulation check, no video-quality rating, no artifact audit. That's the main soft spot; it's not fatal, but it blocks strong conclusions.\n\nWhat it does well: within-subjects design with counterbalancing of both content and order; standard, published scales (IOS, API, NT, IMI, UVS); a sensible R/K paradigm; and a qualitative analysis that doesn't overclaim. The Discussion is appropriately cautious, explicitly noting that historical empathy wasn't measured and that novelty/uncanniness may distract. The proposed design strategies (adaptation phase, selective presence, stylized abstraction) are reasonable.\n\nThe multiple-comparison issue is real but minor; the pattern of effects is consistent, and a Bonferroni correction wouldn't change the main story at the 0.05 level for the largest effects. The N=36 is small but typical for this subfield.\n\nWho should read it: researchers working on pedagogical agents, self-avatars, and narrative learning. It's a useful existence proof that self-mirroring can boost experiential measures without improving—and perhaps slightly hurting—immediate retention. It deserves a serious referee: the question is timely, the method is a reasonable first step, and the confounds are addressable. I'd send it to review with a request for a manipulation check, perceived-video-quality ratings, and a direct comparison of the two agents' fidelity.","headline":"A plausible, honestly-reported small study of a genuinely new application; the central decoupling claim is weakened by an unaddressed stimulus-quality confound, so the paper needs a manipulation check and fidelity ratings before its design implications are trusted.","tokens_in":8244,"tokens_out":3112,"would_cite":false,"duration_ms":25540,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"An AI-generated narrator that mirrors a learner's face and voice increases immersion and emotional connection to history, but lowers immediate quiz scores.","keywords":["digital self","pedagogical agents","history learning","generative AI","self-reference effect","uncanny valley","narrative transportation","within-subjects study"],"falsifier":"Run the same comparison with a manipulation check of perceived video quality and uncanniness, or swap the faces and voices while holding the underlying animation and audio pipeline fixed; if quiz performance becomes equal when perceived video quality is matched, the lower retention score is an artifact of video quality rather than self-reference.","tokens_in":7350,"feed_emoji":"🪞","tokens_out":4599,"duration_ms":36449,"temperature":0.7,"pith_summary":"This paper proposes the Ancestral Digital Self, an AI-generated pedagogical narrator in prerecorded videos that mirrors a learner's face and voice, and tests whether this self-referential agent helps people learn history. In a within-subjects study of 36 adults, the digital-self agent increased narrative transportation, perceived relatedness to the historical culture, self-other inclusion, and positive agent perception, compared with a standard non-self narrator. However, it did not improve learning: history quiz scores were lower in the digital-self condition, and Remember/Know memory-state judgments showed no reliable difference. The authors conclude that identity-mirroring agents can enhance the experiential side of history learning while failing to improve, and possibly harming, immediate knowledge retention, a trade-off they attribute to novelty and uncanniness drawing attention away from content.","feed_headline":"Seeing your own face in a history video lowers quiz scores","feed_subtitle":"An AI avatar mirroring the learner boosts immersion, but immediate retention drops in first exposure.","key_machinery":"The load-bearing object is the Ancestral Digital Self: a prerecorded AI-generated video narrator built by transferring the learner's face onto a historical presenter and cloning the learner's voice, presented as a historically situated version of the self. This mechanism operationalizes the Self-Reference Effect, the idea that self-related cues serve as salient cognitive anchors for deeper processing, inside a narrative-centered learning video. The counterbalanced within-subjects design and established scales (IOS, API, NT, IMI, UVS, and Remember/Know) carry the measurement, while the face- and voice-transfer pipeline carries the argument because it makes the experimental contrast about identity mirroring rather than content.","core_discovery":"The central discovery is that embedding a learner's own facial features and vocal timbre into a historically situated pedagogical agent, called the Ancestral Digital Self, produces a decoupling between subjective learning experience and objective short-term learning outcome. Relative to a neutral pedagogical agent created through the same pipeline, the digital self significantly increased narrative transportation, perceived relatedness, self-other inclusion, and agent persona ratings, with medium-to-large effect sizes. Yet participants scored significantly lower on the content quiz after watching the digital-self video, and Remember/Know judgments did not differ. The authors interpret this as evidence that self-similarity acts as an identity mediator that narrows psychological distance to the past, but that novelty and uncanny eeriness can capture attention at the expense of the historical content in a single-session setting.","pith_inferences":["A longitudinal or repeated-exposure version of this study could reveal whether the retention deficit is a first-session novelty artifact that reverses once the avatar becomes familiar.","If stylized abstraction (non-photorealistic representation) removes the uncanny response while preserving self-recognition, it might keep the experiential gains and eliminate the quiz deficit; this is directly testable with the same workflow.","The decoupling result suggests that self-relevance manipulations in other instructional media, not just video agents, may boost engagement metrics while leaving or lowering immediate recall, so outcome measures should accompany engagement measures in evaluation."],"forward_implications":["When the instructional goal is to draw learners into a narrative and build emotional connection, identity-mirroring agents are an effective design lever.","Identity salience should be calibrated to instructional goals rather than maximized; the paper proposes adaptation phases, selective presence, and stylized abstraction as mitigations.","Short-term engagement gains from self-reference do not automatically translate into short-term knowledge gains, and may reduce them in a first exposure.","The approach could extend to other high-psychological-distance domains such as cross-cultural learning and social-issue documentaries, though the authors call for further research."],"supporting_citations":[{"why":"Supplies the Self-Reference Effect that grounds the prediction that self-relevant cues anchor deeper processing.","marker":"[23]"},{"why":"Provides the Video Self-Modeling lineage that the digital self extends to history learning.","marker":"[6]"},{"why":"Establishes psychological distance in heritage experience, the gap the paper tries to bridge.","marker":"[16]"},{"why":"Provides the Narrative Transportation scale used to measure immersion.","marker":"[10]"},{"why":"Provides the Agent Persona Instrument used to assess agent perception.","marker":"[4]"},{"why":"Supplies the Uncanny Valley scale used to measure eeriness.","marker":"[13]"},{"why":"Supplies the uncanny valley concept used to explain the retention cost.","marker":"[20]"},{"why":"Supplies the novelty-effect account used to explain attention diversion.","marker":"[22]"},{"why":"Provides the historical-empathy model that motivates the experiential outcome.","marker":"[7]"}],"fun_headline_variants":["Self-mirroring history avatar: immersive but worse for quizzes","Seeing yourself in history video cuts test scores despite engagement","Ancestral self-avatar boosts immersion, yet hurts immediate recall","Your digital twin in history lesson: more engaged, less learned","Mirror avatar in history video: high immersion, low quiz scores"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The self and non-self videos are assumed to be matched in quality beyond the identity manipulation, since the self videos come from face transfer and voice cloning of a casual photo and a short recording, and no manipulation check or video-quality rating is reported.","fun_headline_variants_meta":{"raw":{"variants":["Self-mirroring history avatar: immersive but worse for quizzes","Seeing yourself in history video cuts test scores despite engagement","Ancestral self-avatar boosts immersion, yet hurts immediate recall","Your digital twin in history lesson: more engaged, less learned","Mirror avatar in history video: high immersion, low quiz scores"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000165,"raw_usage":{"total_tokens":1214,"prompt_tokens":871,"completion_tokens":343,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":487,"completion_tokens_details":{"reasoning_tokens":255}},"tokens_in":487,"tokens_out":343,"duration_ms":3719,"temperature":1.0,"reasoning_tokens":255,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T12:02:18.851893+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same comparison with a manipulation check of perceived video quality and uncanniness, or swap the faces and voices while holding the underlying animation and audio pipeline fixed; if quiz performance becomes equal when perceived video quality is matched, the lower retention score is an artifact of video quality rather than self-reference.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the Self-Reference Effect that grounds the prediction that self-relevant cues anchor deeper processing."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the Video Self-Modeling lineage that the digital self extends to history learning."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Establishes psychological distance in heritage experience, the gap the paper tries to bridge."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the Narrative Transportation scale used to measure immersion."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the Agent Persona Instrument used to assess agent perception."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the Uncanny Valley scale used to measure eeriness."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the historical-empathy model that motivates the experiential outcome."}],"review_version":1}