{"id":"e7a0458d-1dd0-48b0-89e7-7489c6558da4","arxiv_id":"2501.17819","paper_version":2,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A system that generates social-emotional learning activities from children's videos increased emotion-word use in 5-8 year olds' story retellings, and parents saw it as helping family conversations.","lead":"This paper introduces eaSEL, a system that uses AI to spot social-emotional moments in children's videos and generates reflection activities for kids plus discussion prompts for parents. In a 20-family study, children used more emotion words when retelling a video after doing an eaSEL activity than after just watching.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"No active control in the user study (Section 5.1) means the reported affect-word increase (p=0.02, d=0.27) could result from any post-viewing activity, not specifically from eaSEL's SEL content.","rationale":"The only quantitative evidence for the paper's central claim about children's SEL reflection is the affect-word finding in Section 6.1. That finding compares an activity to doing nothing, so it cannot distinguish the system's SEL-specific mechanism from a general effect of engaging children in any reflective task after viewing. This is precisely the reader's weakest assumption, phrased as 'any interactive post-viewing task.' I agree, and the proposed three-arm study would settle it. The issue is serious but not disqualifying: the paper is transparent about its single-session, small-sample design, and its technical evaluations provide supporting evidence for the system components. Thus the existing CONDITIONAL verdict is appropriate; I would not change it on the basis of this concern alone.","tokens_in":34583,"tokens_out":5711,"duration_ms":54632,"concrete_test":"Run a three-arm within-subjects study with a new sample (N≥30), counterbalancing two comparable episodes and conditions: (1) No Activity, (2) eaSEL Activity, (3) non-SEL active control, e.g., 'Draw a picture of your favorite part of the video and then tell me about it,' matched for modality, duration, and prompt length, with no emotion-related language. Use the same retelling prompt ('Please tell me what happened in the video you just watched') and the same LIWC15 affect-word proportion as the outcome. Pre-register the contrast: if both activity arms significantly exceed No Activity and do not differ from each other, the SEL-specific claim is not supported; if only the eaSEL arm is significantly higher, the concern is resolved. As a secondary check, have condition-blinded annotators rate the emotional content of the retellings to validate LIWC output for this age group.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The one-way within-subjects design in Section 5.1 uses a single factor, Activity, with levels No Activity and eaSEL Activity. Section 6.1 reports that children use a higher proportion of affect words after the eaSEL activity. This contrast does not isolate the SEL component of the intervention: any interactive post-viewing task—e.g., drawing a favorite character, retelling the plot to a puppet, or answering generic comprehension questions—could prompt more talk and more emotion-related vocabulary simply by increasing reflection about the video. The paper's central contribution is the AI-mediated SEL activity, not post-video activities in general, so H1 should have been tested against an active control condition matched for modality, duration, and adult attention. The absence of such a condition is the most load-bearing weakness because it underdetermines the causal interpretation of the paper's key positive result. Additionally, the LIWC15 affect-word lexicon (Section 5.4) is derived from adult text and has not been validated for 5–8-year-old children, so the measured difference may also be affected by differential lexical coverage across conditions, though the active-control issue is more fundamental.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents eaSEL, an LLM-based pipeline that (1) detects social-emotional learning (SEL) moments in children's video transcripts, (2) generates child-facing reflection activities tied to those moments, and (3) produces parent-facing conversation starters intended to scaffold parent-child discussion without co-viewing. The technical evaluation uses human gold labels from two annotators (Krippendorff's alpha 0.64, 0.88 overall agreement) and ratings from five parent annotators on 59 generated child activities and 59 conversation starters; the authors report high relevance and reflection-promotion scores for activities, with weaker child-appropriateness results, and high relevance for conversation starters. The user study is a one-way within-subjects experiment with N=20 dyads comparing an eaSEL Activity condition with a No Activity condition on children's use of affect words in story retellings, plus qualitative analysis of child artifacts and parent interviews. The paper reports a significant increase in general affect-word proportion (p=0.02, Cliff's d=0.27) and positive parent perceptions. The central claimed contribution is that eaSEL promotes children's SEL reflection during independent media consumption and scaffolds parent-child engagement.","tokens_in":34755,"tokens_out":3405,"duration_ms":41008,"significance":"If the central claim is supported, the contribution is useful and timely: it addresses independent media consumption, a realistic context that joint-media-engagement systems do not cover, and it connects LLM-generated activities to established SEL curricula. Strengths of the manuscript include the transparency of the pipeline (full prompts in the appendix), the use of external human judgments for the technical evaluation, the explicit analysis of SEL detection failure modes (e.g., the R2 social-skills category), and the qualitative coding of child artifacts and parent interviews. The work is not circular: system outputs are judged against human gold labels and the user-study hypothesis was not used to set system parameters. However, the user study's design does not support the specific causal claim that the SEL content of eaSEL activities, rather than any interactive post-viewing activity or episode differences, drove the observed affect-word increase. The main quantitative result is therefore underdetermined, and the paper's headline interpretation is stronger than the evidence.","major_comments":[{"comment":"The user study does not include an active control condition. The single factor, Activity, has only two levels: No Activity and eaSEL Activity. Any interactive post-viewing task—generic comprehension questions, retelling to a puppet, or drawing a favorite character—could increase reflective talk and affect-word use simply by prompting additional engagement with the video. The observed p=0.02, Cliff's d=0.27 for general affect words therefore does not isolate the SEL-specific component of eaSEL, which is the paper's central contribution. In addition, each child watched one of two fixed episodes per condition; counterbalancing order across participants does not remove the episode-activity confound because episode content (including emotional salience) is not matched or varied independently of the activity factor. Section 7.4 lists limitations but does not mention this active-control gap. To support H1 as stated, the authors should add an active control condition matched for modality, duration, and adult attention, or should substantially weaken the causal claims in the abstract and discussion.","section":"Sections 5.1, 6.1, and 7.4"},{"comment":"The outcome measure itself needs justification for this age group and analysis approach. LIWC15 is an adult-text lexicon and has not been validated for 5-8-year-old children; words such as 'kind' and 'hugged' are counted as affect words, and coverage may differ across the two episodes and across activity conditions in ways unrelated to SEL reflection. The manuscript also tests three LIWC outcome families (general affect, positive emotion, negative emotion) without multiple-comparison correction; only the general affect proportion is significant at p=0.02. The reported effect size is small (Cliff's d=0.27). At minimum, the authors should report adjusted p-values or pre-specify a single primary outcome, and should provide evidence or a reasoned argument that LIWC15's affect lexicon is appropriate for child retellings.","section":"Sections 5.4 and 6.1"},{"comment":"The metric used for retellings is described as 'affect word proportions (number of unique emotion words / unique words in full re-telling)' but Figure 6 is labeled 'Emotion word counts' and the text elsewhere refers to 'frequency of positive or negative emotion words.' This ambiguity matters because the hypothesis concerns 'more emotional language,' which could mean token counts, unique types, or proportions. Please clarify the exact metric used for each comparison and ensure the figure and text are consistent. If the analysis used unique-word proportions, the interpretation in terms of 'using more emotional language' requires an assumption that lexical diversity is the relevant construct; that assumption should be stated and justified.","section":"Section 5.4"}],"minor_comments":[{"comment":"The sample is described as ages 5-8, but the participant table shows only one 5-year-old; the effective age range is mostly 6-8. This should be acknowledged when discussing generalizability to the full stated age range.","section":"Section 5.2"},{"comment":"The two video episodes are not described in terms of their emotional content, length, or narrative complexity. Since episode and condition are confounded, the manuscript should at least report basic properties of the two episodes and ideally provide the episode titles or scripts for reproducibility.","section":"Section 5.3"},{"comment":"The statement that 74% of activities 'do not contain yes or no questions' means 26% do; the paper then says this was 'not harmful,' but no data are presented to support that conclusion. Either provide evidence or soften the claim.","section":"Section 4.3.2"},{"comment":"The average Sentence-BERT cosine similarity of 0.36 is interpreted as 'weak similarity' but the qualitative examples show semantically similar explanations. The authors should note that cosine similarity on this embedding space is not a calibrated measure of semantic equivalence for short, explanatory texts.","section":"Section 4.2.2"},{"comment":"The 'fail-safes' scenario ('if a child chooses to skip or skimp on an activity, a parent might notice') is speculative; no data in the study address skipping behavior. Consider labeling this as a design vision rather than a finding.","section":"Section 7.3"}],"recommendation":"major_revision","confidential_remarks":"The paper has a solid technical evaluation and a well-documented system, but the user study's missing active control is a load-bearing weakness for the main positive claim. A revision that adds an active control condition or reframes the H1 claim as 'any interactive post-viewing activity' rather than 'SEL-specific reflection' would substantially strengthen the paper. The LIWC validity and multiple-comparison issues reinforce the need for a more cautious interpretation. I see no novelty-disclosure or citation-pattern concerns; the related work is appropriately situated."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Read this one. It's a good system paper: eaSEL takes LLMs and applies them to a real niche—turning solo video watching into scaffolded SEL practice for 5-8 year olds, plus parent conversation starters that don't require co-viewing. The technical evaluation is honest and solid: two annotators reached Krippendorff's alpha 0.64 on SEL moment detection, and five parent annotators rated generated activities as relevant, with clear note of where child-language appropriateness falls short. No circularity: the detection is judged against human gold labels, and the user study hypothesis doesn't set any constants. That part earns credit.\n\nThe user study itself is competently executed for an initial system test: N=20 dyads, counterbalanced within-subjects design, and they collect retellings, artifacts, and parent interviews. But the main quantitative result—more affect words after eaSEL activities (p=0.02, Cliff's d=0.27)—comes from a one-way contrast between activity and no activity. There's no active control. So the result shows the whole package works better than nothing, but it doesn't show the SEL-specific content is the active ingredient. Any post-video reflective task could plausibly produce the same increase. The paper claims a bit more than that in the abstract and discussion. This is a real limitation, not a fatal one, since the system as a whole is the contribution, but it should be acknowledged and ideally addressed in a follow-up.\n\nThe LIWC15 measure for children is also unvalidated, but it's the same lexicon across conditions, so it's less concerning than the active-control issue. Sample is small and homogeneous, which the paper openly admits. They also mention they didn't observe parent-child conversations, so the parent-side claims are based on interview reports, and the paper says that.\n\nOverall: not a load-bearing flaw, but a gap between the evidence and the strength of the causal language. I'd want the revise to add an active control or soften the conclusions. This deserves peer review, and a serious referee would catch the active-control point and ask for one.\n\nFor a reader: useful for anyone working on child-AI interaction, family media, or LLM-based learning tools. I'd bring it to a reading group.","headline":"A well-put-together system paper whose user study supports eaSEL as a whole, but does not isolate the SEL component; worth a serious referee.","tokens_in":35307,"tokens_out":3277,"would_cite":true,"duration_ms":34241,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that AI-generated reflection activities inserted into children's video watching make children retell stories with more emotion words, and that parents see the generated prompts as scaffolding for deeper conversations.","keywords":["social-emotional learning","parent-child interaction","children's media","large language models","video reflection activities","conversation starters","user study","affect language"],"falsifier":"A direct test would run the same two episodes with three conditions: no activity, an eaSEL activity, and an equally engaging non-SEL activity (for example, drawing a favorite scene or retelling the plot). If the non-SEL condition produces the same emotion-word increase as eaSEL, the effect is not specific to social-emotional content; if a different child vocabulary instrument fails to reproduce the difference, the measure may be driving the result.","tokens_in":1559,"feed_emoji":"🎬","tokens_out":2061,"duration_ms":92289,"temperature":0.7,"pith_summary":"Children aged 5-8 spend much of their screen time watching videos on their own, and parents want that time to support learning. The paper introduces eaSEL, a tablet system that watches along with a child's video, uses a large language model to spot social-emotional moments in the transcript, and then offers the child a short reflection activity such as drawing a related memory or role-playing a scene. After the video it gives the parent a summary, the child's artifact, and a conversation starter tied to the same social-emotional skill. The paper's central empirical claim is that doing one of these activities changes what children take away from the show: in retellings, children used proportionally more emotion words after an eaSEL activity than after watching without an activity (p=0.02, Cliff's d=0.27). Parents in interviews said the tool made otherwise passive viewing feel constructive and gave them entry points for talking about feelings and lessons.","feed_headline":"AI reflection prompts make kids retell stories with more feeling words","feed_subtitle":"A 20-family study found the effect, and parents said the conversation starters could deepen talks.","key_machinery":"The machinery is a pipelined prompting chain built on a large language model using in-context learning. Automatic speech transcription converts each episode into a transcript; the model then rates whether each of ten skills drawn from the standard five-competency social-emotional learning framework appears in a central plot moment, with positive and negative examples for each skill. Next, conditioned on the detected skill and moment, the model generates one of four activity types for the child (drawing, creative story play, personal storytelling, role play) and a separate parent-child conversation starter that asks the parent to share a personal experience tied to the same skill. The quantitative outcome that carries the user-study claim is the proportion of affect words in children's story retellings, measured with an automated emotion-word lexicon and compared across conditions with a Wilcoxon signed-rank test.","core_discovery":"On the paper's own terms, the discovery is that an LLM-driven pipeline can turn a children's video transcript into a teachable social-emotional moment and a developmentally plausible activity, and that completing such an activity measurably shifts a child's own language about the story. In a within-subjects study with 20 parent-child dyads, children's retellings contained a higher proportion of affect words (e.g., 'feel,' 'kind,' 'angry') after the eaSEL condition than after the no-activity condition, and this difference was significant by a Wilcoxon signed-rank test. Child-produced drawings and recordings showed that 16 of 20 children gave substantive, reflective answers, often by linking the story to personal experience or by adopting a character's perspective. Parents reported that the child activity encouraged active reflection and regular practice of emotional vocabulary, and that the parent-facing summary, artifact, and conversation starter could scaffold conversations they would not otherwise have had. The technical evaluation found strong relevance of generated activities to the detected SEL skill and moment, with the weakest areas being detection of 'social skills' moments and use of child-appropriate language.","pith_inferences":["Because the Activity factor only contrasts no activity with eaSEL, the design cannot rule out that any interactive post-viewing task—not the SEL content—drives the affect-word increase; a non-SEL active control would isolate the mechanism.","If the emotion-word measure is a valid proxy, a longer-term deployment could test whether repeated eaSEL sessions grow children's emotion vocabulary beyond the study session and whether parent-child conversations triggered by the starters actually happen and deepen.","The same detect-and-generate pipeline could be extended to audiobooks, games, or other passively consumed children's media, where a parent-facing artifact might be even more valuable because there is no visual record of the experience.","The paper's out-of-scope note about inappropriate source content implies a deployment decision: the system should probably refuse to generate activities for content parents would find objectionable, or let parents set boundaries on which lessons are acceptable."],"forward_implications":["If the effect holds, video-watching time can be converted into a low-cost, at-home SEL practice without requiring parents to watch the same content.","The pipeline's transcript-based detection means the approach could attach reflection activities to any children's show, not just a curated set of episodes.","Parents can receive conversation starters and artifacts without co-viewing, which addresses the time constraint that makes joint media engagement impractical for many families.","The technical evaluation identifies two bottlenecks for deployment: social-skill moments are often missed, and generated language is sometimes too advanced for 5-8 year olds, so content selection and child-appropriate rewrites are needed.","The paper's own limitation list says the single-session, homogeneous sample means longitudinal learning outcomes and actual parent-child conversations remain untested."],"supporting_citations":[{"why":"Supplies the emotion-word lexicon used to score children's retellings, the outcome measure behind the central H1 result.","marker":"[53]"},{"why":"Provides the evaluation rubric for generated questions and is the prior system whose SEL-focused question design eaSEL extends to video.","marker":"[76]"},{"why":"Identifies SEL classroom activities and at-home practice needs that define the four child activity types.","marker":"[73]"},{"why":"Defines the five-competency SEL framework from which the ten detected skills are derived.","marker":"[23]"},{"why":"Establishes the in-context learning prompting approach used for detection and generation.","marker":"[7]"},{"why":"Automatic transcription component that converts episodes into the transcripts the pipeline operates on.","marker":"[56]"}],"fun_headline_variants":["AI activity boosts kids' emotion words in story retells","Parent-child study: AI prompts deepen kids' emotional reflection","eaSEL turns video time into SEL moments for families","After AI activity, children use more feeling words in retellings","SEL tool: AI prompts lead to more affect words in kids' stories"],"cache_read_input_tokens":37504,"weakest_assumption_plain":"The causal reading of the user study rests on the assumptions that the two video episodes are emotionally comparable and that a higher proportion of emotion words in a 5-8 year old's retelling actually reflects social-emotional reflection rather than, say, general talkativeness or a task-related priming effect.","fun_headline_variants_meta":{"raw":{"variants":["AI activity boosts kids' emotion words in story retells","Parent-child study: AI prompts deepen kids' emotional reflection","eaSEL turns video time into SEL moments for families","After AI activity, children use more feeling words in retellings","SEL tool: AI prompts lead to more affect words in kids' stories"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000288,"raw_usage":{"total_tokens":1688,"prompt_tokens":941,"completion_tokens":747,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":557,"completion_tokens_details":{"reasoning_tokens":660}},"tokens_in":557,"tokens_out":747,"duration_ms":7668,"temperature":1.0,"reasoning_tokens":660,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T04:31:54.127617+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A direct test would run the same two episodes with three conditions: no activity, an eaSEL activity, and an equally engaging non-SEL activity (for example, drawing a favorite scene or retelling the plot). If the non-SEL condition produces the same emotion-word increase as eaSEL, the effect is not specific to social-emotional content; if a different child vocabulary instrument fails to reproduce the difference, the measure may be driving the result.","supporting_citations":[{"cited_title":"ContextQ: Generated Questions to Support Meaningful Parent-Child Dialogue While Co-Reading","cited_arxiv_id":"2405.03889","evidence_quote":"Provides the evaluation rubric for generated questions and is the prior system whose SEL-focused question design eaSEL extends to video."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Identifies SEL classroom activities and at-home practice needs that define the four child activity types."}],"review_version":1}