{"id":"d8811bda-bb1c-4141-bbf9-5270d8a88709","arxiv_id":"2502.03862","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Video-based reflection nudges improved several deliberation-quality measures over text, image, and audio formats, despite text being the subjectively preferred modality for persona-style prompts.","lead":"This paper reports two user studies on how the format of reflection prompts (text, image, video, or audio) affects the quality of people's written opinions in online deliberation. Video prompts generally led to higher scores on several quality measures, even though participants said they preferred text for persona-style prompts.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 'strictly same content' premise in Section 3.3 is violated: image stimuli omit the persona/story script and video adds narration plus visuals, so the reported modality effects are confounded with information content.","rationale":"The reader's weakest-assumption analysis identifies the modality-content confound as the central threat. I agree. The paper's own generated-stimulus constraints (image no wordings; video with narration and background visuals) make it impossible to hold content fixed across modalities, so the causal claim that modality drives deliberativeness is not yet supported. However, this is a fixable design flaw rather than evidence against the existence of modality effects; the empirical pattern may survive with proper controls. Therefore the condition on the verdict remains appropriate. The proposed two-condition follow-up isolates the contribution of the video format from the contribution of simultaneous verbal+visual information, testing the load-bearing premise directly. No ad hominem or external-consensus issue is involved; the concern is internal to the design.","tokens_in":35923,"tokens_out":5808,"duration_ms":64968,"concrete_test":"Run a follow-up experiment for the persona nudge with two additional between-subject conditions matched to the existing Video condition: (1) Text+Image, where the full persona text appears as readable captions on the same still image used in the video, and (2) Audio+Static Image, where the same narration is played over that still image with no motion. Use the same N=25 per cell, covariates, and dependent measures. If the argument-repertoire mean for Text+Image is within a pre-specified non-inferiority margin (e.g., 0.5 arguments) of the Video mean, and Audio+Static Image is also non-inferior, then the video advantage is attributable to combined verbal+visual content, not to the video modality; the central claim should be weakened. If Video still significantly beats both controls, the concern is resolved.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim—video nudges cause higher deliberativeness—rests on Section 3.3's assertion that 'the content across all modalities is strictly the same, just presented differently by their own modality.' The implementation contradicts this. For the image condition, Section 3.3 and Appendix Table 4 impose 'The image has no wordings'; the persona's perspective text and the story text are therefore absent from the image stimulus. Section 3.2's earlier claim that the image modality presents 'the same text content, with additional images' is thus internally inconsistent. For video, the generation prompt requires the full script to be voiced as narration ('follow the script exactly as provided') and adds 'a background relevant to the persona's occupation' with 'photo-realistic' moving visuals. Consequently, the video condition delivers the complete verbal script plus extra visual scene information; the text condition delivers only the script; the audio condition delivers only the narration; the image condition delivers only an illustration. The observed advantages of video (e.g., argument repertoire for persona, opinion expression for storytelling) could therefore be produced by the larger amount of information presented (verbal script plus visual context), not by the video format as such. This is a confound, not just an alternative interpretation, because the paper's stated design goal was to hold content fixed.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper investigates whether presenting reflective nudges in different modalities (text, image, video, audio) affects the quality of online deliberation. It reports two user studies: Study 1 (N=20) measures subjective modality preference for two nudge types, persona (direct) and storytelling (indirect), finding text most preferred for persona and video for storytelling. Study 2 (N=200, between-subjects, 25 per condition) measures five deliberativeness dimensions under each modality and nudge type, reporting that video yields higher argument repertoire for persona and higher argument diversity, opinion expression, and justification level for storytelling. The paper concludes that video is generally the most effective modality and advocates matching modality to nudge type. The central causal claim rests on the assertion in Section 3.3 that content is 'strictly the same' across modalities, so that observed differences are attributable to modality alone.","tokens_in":36167,"tokens_out":3808,"duration_ms":38778,"significance":"If the causal claim held, the paper would make a useful empirical contribution to online deliberation design, extending prior text-only nudging work to multimodal presentation and offering practical guidance. The study has genuine strengths: a pre-specified power analysis, inter-coder reliability above 0.83 for all dependent variables, use of established deliberativeness measures, and a reasonably powered between-subject design in Study 2. However, the central attribution of effects to modality is undermined by a content-equivalence confound that is both internally contradicted and unvalidated. The paper's own Limitations section does not mention this confound. With the finding reframed as an effect of multimodal presentation formats (which differ in information content), the study would still be of interest, but the causal language in the abstract and conclusion is not currently supported.","major_comments":[{"comment":"The central premise that 'the content across all modalities is strictly the same, just presented differently by their own modality' is contradicted by the implementation. The image prompt explicitly requires 'The image has no wordings,' so the persona perspective or story text is absent from image stimuli, while the video prompt requires the voice narration to 'follow the script exactly as provided' and adds 'a background relevant to the persona's occupation' and photo-realistic moving visuals. Thus the video condition presents the full verbal script plus additional scene information, the audio condition presents only the narration, the text condition only the script, and the image condition only an illustration. These differences in information content, not modality alone, could explain the observed video advantages in argument repertoire (Section 6.1.1) and argument diversity/opinion expression/justification (Sections 6.1.2-6.1.4). This is a load-bearing confound because the paper's stated design goal was to hold content fixed in order to attribute effects to modality.","section":"Section 3.3, Appendix Table 4"},{"comment":"No manipulation check or content-equivalence validation is reported. Given the documented differences in stimulus composition, the authors need to provide evidence that participants perceived the same underlying content across modalities, or substantially temper the causal interpretation. The Limitations section (Section 8) lists several limitations but does not acknowledge this confound, which is directly relevant to the paper's main claim.","section":"Section 3.3 / Section 8"},{"comment":"The summary of Study 2 overstates the persona result. For the direct (persona) nudge, only Argument Repertoire reached significance (Table 2, p<.01); Argument Diversity, Opinion Expression, Justification Level, and Constructiveness were all non-significant. The statement 'Video generally emerges as the most effective modality for the direct reflective nudge' is therefore too strong, and the Introduction's claim that 'video leads to higher level of deliberative quality for persona' should be corrected to specify the single significant dimension.","section":"Section 6.1.6, Table 2, Introduction"}],"minor_comments":[{"comment":"The text says post-hoc comparisons were applied to 'the eight measures of deliberativeness,' but Section 5.2 defines only five dependent variables; this should be corrected to 'five measures' or otherwise clarified.","section":"Section 6.1"},{"comment":"The caption contains a typo: 'ANCOV A' should be 'ANCOVA'.","section":"Figure 4 caption"},{"comment":"The education category 'Post-Graudate Degree' is misspelled; it should read 'Post-Graduate Degree.'","section":"Table 6"},{"comment":"Study 1 required participants to write at least 30 words, but Study 2 does not state a minimum word count; please clarify whether the same instruction applied in Study 2, since length constraints could interact with deliberativeness measures.","section":"Section 4.5 / Section 5.5"},{"comment":"The ANCOVA in Study 1 is conducted on only 20 participants with four covariates; the authors should acknowledge in Section 8 that this small sample size limits the reliability of the preference ranking analysis, which is currently described without such a caveat.","section":"Section 4.7"}],"recommendation":"major_revision","confidential_remarks":"The content-equivalence confound is the key blocking point. I would look favorably on a revision that either (a) re-runs the study with stimuli that truly hold information content constant (e.g., image with accompanying text, video without added narration or with only the same verbal content), or (b) reframes the contribution as an exploration of multimodal presentation formats that naturally differ in information content, removing the causal language of modality effects. The paper's empirical measures and inter-coder reliability are solid, and the topic is within the scope of the venue. The manuscript also needs to correct the overstated persona result in the Introduction and Section 6.1.6."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a well-run pair of studies and a genuinely new comparison, but the headline conclusion—video nudges cause higher deliberativeness—is undermined by a content confound the authors built into their own design. Their stated goal was that \"content across all modalities is strictly the same\" (Section 3.3), but that isn't what happened. The image condition omits the persona/story text entirely (the generation prompt in Appendix Table 4 requires \"no wordings\"), while the video condition voices the full script as narration and adds a visual background. So video presents more information than text or audio, and much more than image. The reported advantages of video (argument repertoire for persona; argument diversity, opinion expression, justification for storytelling) could easily come from having the full script read aloud plus a vivid scene, rather than from the video format as such.\n\nWhat is genuinely new: this is the first head-to-head comparison I know of for text, image, video, and audio reflection nudges in deliberation. Study 1's preference results (text for persona, video for storytelling) are interesting and the qualitative analysis is thorough. Study 2 uses a sensible 2x4 between-subjects design with a power analysis, inter-coder reliability (kappas 0.83–0.92), and covariate control. The dual-systems framing is used post hoc to interpret comments, not to predict outcomes, so there's no circularity problem.\n\nSoft spots beyond the confound: there is no no-nudge control, so we can't tell whether any modality actually improves deliberativeness relative to reflecting without a nudge. No effect sizes are reported, only p-values. The demographic tables show some imbalance (e.g., some cells skew male), but that's minor.\n\nThe confound is load-bearing. The paper's central claim—that modality itself drives the differences—doesn't hold up. The authors could re-frame the work as a comparison of multimodal stimuli rather than pure modality, or add a text-plus-image condition to disentangle content from format, or at minimum discuss the confound prominently and soften the causal language. As written, the causal conclusion is not established.\n\nWho this is for: researchers designing deliberation interfaces and anyone studying modality effects in nudging. A serious referee should engage—this is not a desk reject. With a major revision that addresses the content/modality distinction, it could be a solid contribution. I'd accept it for review.","headline":"The 'video works best' result is real in the data but can't be cleanly attributed to modality: the image condition drops the text and the video condition adds narration plus visuals.","tokens_in":36678,"tokens_out":2241,"would_cite":false,"duration_ms":22524,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Presenting reflection prompts as videos improves deliberative quality more than text, image, or audio, even when users say they prefer text.","keywords":["deliberation","deliberativeness","reflective nudges","multimodality","online deliberation","self-reflection","large language models","user study"],"falsifier":"Run the storytelling comparison again with a matched audio script: the same narration presented as audio-only, as narration with a static image, and as narration with moving visuals. If the three produce equal deliberativeness scores, the paper's claim that video modality itself drives the effect is wrong.","tokens_in":35745,"feed_emoji":"🎥","tokens_out":7528,"duration_ms":66380,"temperature":0.7,"pith_summary":"The paper asks whether the sensory format of a reflection prompt—text, image, video, or audio—changes how thoughtfully a person writes before joining an online discussion. It reports that format matters: for persona-style direct nudges, video produced significantly more non-redundant arguments than text, image, or audio, even though text was the most preferred format. For storytelling-style indirect nudges, video again led on argument diversity, opinion expression, and justification level, while text scored worst on opinion expression. The authors conclude that reflective nudge design should match modality to nudge type, and that subjective preference is not a reliable guide to deliberative quality. This matters because online deliberation platforms currently rely on text by default, potentially under-servicing complex reflective tasks.","feed_headline":"Video reflection nudges beat text for deeper online deliberation","feed_subtitle":"A 200-person experiment shows video boosts argument diversity and justification for storytelling prompts, even where text is preferred.","key_machinery":"The carrying mechanism is the multimodal reflective nudge—a short prompt, generated by a large language model and rendering tools, that asks a user to consider another person's perspective (persona) or a narrative (storytelling) before writing. It is presented inside the Reflect interface in one of four modalities: text, image, video, or audio. The experiment's engine is the 2x4 between-subjects comparison of these nudges against five coded measures of deliberativeness: argument repertoire, argument diversity, opinion expression, justification level, and constructiveness. Content equivalence across modalities is maintained by using the same generated script as the source for every modality.","core_discovery":"The central claim is that modality changes deliberativeness, and video is the most consistently effective modality. In a between-subjects experiment with 200 participants, the persona-type nudge delivered as video yielded an average of 3.60 non-redundant arguments, significantly more than text (2.32), image (2.56), or audio (2.60). For the storytelling nudge, video significantly outperformed image on argument diversity (5.52 vs 3.72 themes) and justification level (2.32 vs 1.44), and text was significantly worse than image, video, and audio for opinion expression (0.48 vs 0.76, 0.84, 0.84). The paper interprets these effects through dual-system thinking: text favours fast autonomous reading suited to short direct prompts, while video's combined visual and auditory channels sustain engagement for longer narrative prompts. The conclusion is that the optimal modality depends on the type of reflective nudge, and platforms should offer video rather than defaulting to text.","pith_inferences":["The paper's content-equivalence assumption is incomplete: image prompts were required to contain no wording and video added voice narration and motion, so the video condition carries more information than the image condition; part of video's measured advantage may therefore be an information-richness effect rather than a pure modality effect.","A direct replication could isolate the mechanism by using the identical narration in audio-only and video conditions; if the visual channel adds nothing beyond the voice, the video advantage should shrink.","The results suggest a testable personalization rule: users with low prior topic knowledge or non-native language backgrounds may benefit more from image or video nudges, since those modalities convey context without requiring dense reading.","Because storytelling videos ran 90 to 120 seconds while persona videos ran 9 to 12 seconds, the duration difference is confounded with nudge type; shorter narrative videos might not reproduce the storytelling benefit."],"forward_implications":["Online deliberation platforms can raise individual opinion quality by offering video-based reflection nudges, especially when the nudge is a narrative.","Text-based nudges remain defensible for short persona prompts, but platforms should avoid long text-only storytelling prompts because text scored lowest on opinion expression.","Subjective preference data is not sufficient for choosing nudge formats; objective deliberation metrics should drive design decisions.","LLM-backed generation makes it feasible to render the same reflective script across modalities, letting platforms adapt to different tasks and users without manual content authoring."],"supporting_citations":[{"why":"Provides the prior result that textual reflective nudges improve deliberativeness, plus the persona/storytelling nudge designs and the five deliberativeness measures reused here.","marker":"[197]"},{"why":"Shows reflective question prompts improve opinion quality and expression, grounding the direct-nudge manipulation.","marker":"[199]"},{"why":"Supplies multimedia learning theory and the principle that words plus pictures outperform text alone, motivating the hypothesis that modality changes reflection.","marker":"[126]"},{"why":"Supplies the visual-auditory-reading/writing-kinesthetic taxonomy used to measure participants' inherent reflecting styles as a covariate.","marker":"[62]"},{"why":"Provides the dual-system account of fast intuitive and slow reflective thought used to interpret why modalities differ.","marker":"[100]"},{"why":"Defines the deliberativeness dimensions—rationality and constructiveness—on which the five dependent measures are built.","marker":"[186]"},{"why":"Establishes that disagreement and opinion quality can be measured through argumentation, supporting the argument-repertoire approach.","marker":"[160]"}],"fun_headline_variants":["Video nudges boost argument diversity in online deliberation","For storytelling prompts, video outperforms text in deliberation quality","Optimal nudge modality depends on prompt type: video wins for storytelling","Multimodal nudges: video boosts deliberation, but type matters"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The central assumption is that the four modality versions contain strictly the same content, so any quality difference must come from the presentation format; but the image version omits the text and the video version adds narration and motion, so the content is not actually identical.","fun_headline_variants_meta":{"raw":{"variants":["Video nudges boost argument diversity in online deliberation","For storytelling prompts, video outperforms text in deliberation quality","Optimal nudge modality depends on prompt type: video wins for storytelling","Multimodal nudges: video boosts deliberation, but type matters"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000274,"raw_usage":{"total_tokens":1618,"prompt_tokens":905,"completion_tokens":713,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":521,"completion_tokens_details":{"reasoning_tokens":643}},"tokens_in":521,"tokens_out":713,"duration_ms":6838,"temperature":1.0,"reasoning_tokens":643,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-09T00:26:28.417886+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the storytelling comparison again with a matched audio script: the same narration presented as audio-only, as narration with a static image, and as narration with moving visuals. If the three produce equal deliberativeness scores, the paper's claim that video modality itself drives the effect is wrong.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the prior result that textual reflective nudges improve deliberativeness, plus the persona/storytelling nudge designs and the five deliberativeness measures reused here."},{"cited_title":"system role","cited_arxiv_id":null,"evidence_quote":"Shows reflective question prompts improve opinion quality and expression, grounding the direct-nudge manipulation."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies multimedia learning theory and the principle that words plus pictures outperform text alone, motivating the hypothesis that modality changes reflection."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines the deliberativeness dimensions—rationality and constructiveness—on which the five dependent measures are built."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Establishes that disagreement and opinion quality can be measured through argumentation, supporting the argument-repertoire approach."}],"review_version":1}