{"id":"c8106c4b-83c3-4f01-9df4-64c203daf201","arxiv_id":"2505.03164","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"InfoVids, which co-locate presenter and visualization in one space, made viewers focus more on the presenter and feel more engaged than an equivalent slide-based video.","lead":"This paper introduces InfoVids, videos that place the presenter and the data visualization in the same three-dimensional space instead of a separate slide. In a 30-person comparison, viewers reported less attention splitting, more focus on the presenter, and more natural full-body performances.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The baseline is not a fair comparison: the same take was constrained by a scripted 'common body language' that the authors admit detracted from InfoVid benefits, so measured differences may be artifacts of baseline construction.","rationale":"The most load-bearing link in the argument is the validity of the baseline condition, because every quantitative comparison and the qualitative interpretations are relative to it. The reader identified the same point, and the manuscript itself is transparent about the damage: Appendix B.1 admits that the scripted common body language 'detracted from InfoVid benefits.' That admission means the control was deliberately degraded on outcome measures that the paper then claims InfoVids improve. Such a control cannot support the causal phrasing in the abstract. I do not think this is fatal if the paper is treated as an exploratory technology probe; the qualitative insights and design lessons stand on their own. However, the statistical layer claims more than this design can support, especially with 36 uncorrected binomial tests and no released stimuli or data. A single rerun with an ecologically valid baseline is the decisive check. Because the reader's verdict is already CONDITIONAL and my concern does not change that verdict, I recommend UNCHANGED.","tokens_in":19824,"tokens_out":3401,"duration_ms":34957,"concrete_test":"Rerun the comparison with an independently authored baseline: same presenter, script, content, duration, and resolution, but produced as a normal slides/videoconferencing video (natural gaze toward the slides, ordinary slide-specific gestures, standard upper-body picture-in-picture, no common-body-language constraint, no post-hoc animation synchronization), and collect Q1–Q9 from a fresh N=30 sample. The original finding is supported only if the InfoVid advantages on natural body movement, enjoyability, and attention allocation (Q4/Q5/Q8–Q9) persist; if they shrink or reverse, the reported differences were artifacts of baseline construction.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The causal claim in the abstract—that InfoVids 'reduced viewer attention splitting, shifted the focus from the visualization to the presenter, and led to more interactive, natural, and engaging full-body data performances'—rests on the baseline being a representative slides/videoconferencing condition. Appendix B.1 describes how both conditions were recorded from the same take and then choreographed so that one performance would work in both formats: the actor was instructed to use large, open hand movements and to face the camera, and the authors state these modifications 'detracted from InfoVid benefits' and 'made the presenter look less immersed.' The baseline additionally required upper-body-only framing, rule-of-thirds post-hoc zoom/pan, and animations post-hoc synced to presenter movements (B.2, B.3). Thus every metric that distinguishes the conditions—immersion, engagement, natural body movement, attention allocation—is confounded with these design choices. The observed shift in attention and perceived naturalness could be an artifact of a non-representative baseline deliberately optimized to make the InfoVid look good, rather than evidence that co-locating presenter and visualization changes the viewing experience.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces InfoVids, a presentation paradigm in which the presenter and an AR visualization occupy a shared 3D frame, with the stated goal of creating a more equitable relationship between presenter and visualization. The authors describe an iterative, autobiographical design process (nine months), a prototype implementation called the Body Object Model (BOM), and four technology probes (AirplaneVis, NapoleonVis, InjuryVis, WalmartVis) that vary spatial layout, body-visualization attachment, and interaction. They then report a comparative study with 30 public participants who viewed each InfoVid and a corresponding 'slides' baseline and answered nine forced-choice-style questions, followed by semi-structured interviews. The central claim is that InfoVids reduce viewer attention splitting, shift focus from the visualization to the presenter, and produce more interactive, natural, and engaging full-body performances.","tokens_in":20036,"tokens_out":4336,"duration_ms":45295,"significance":"If the empirical claims were supported, the paper would make a useful contribution to the emerging design space of AR-mediated data presentation and performance, complementing work on data videos and shorts. The four probes are thoughtfully chosen, the qualitative analysis is grounded in participant quotes, and the autobiographical reflection provides practical design lessons. The authors also took several pains to make the comparison more controlled: using an external actor, recording from the same take, randomizing order, and reporting the limitations of the baseline construction. However, the strength of the central empirical claims is not justified by the current evidence, primarily because the baseline condition is not a representative or neutral comparison and because the statistical analysis is under-powered and over-tested.","major_comments":[{"comment":"The fairness of the baseline is load-bearing for every causal claim in the abstract, and the baseline is not a neutral representation of a conventional slide presentation. To record both versions from the same take, the actor was scripted to use large open hand movements and to face the camera; the authors themselves write that these modifications 'detracted from InfoVid benefits and made the presenter look less immersed' (B.1). In addition, the baseline was framed to show only the upper torso (B.2), slide animations were post-hoc synchronized to the performer's movements (B.3), and rule-of-thirds reframing was applied. Consequently, the observed differences in immersion, natural body movement, engagement, and attention (Figures 7–8) may reflect the intentionally constrained baseline rather than the co-located presentation relationship itself. The abstract states that InfoVids 'reduced viewer attention splitting, shifted the focus from the visualization to the presenter, and led to more interactive, natural, and engaging full-body data performances'; such causal statements are not supported by a comparison against a baseline that suppresses the very features being tested. The claims should be explicitly scoped to 'this particular baseline construction,' or a more representative baseline should be devised.","section":"Appendix B.1–B.3, §3.3, Abstract"},{"comment":"The statistical layer does not support the number of significance claims. The paper reports 36 binomial tests (9 questions × 4 visualizations) at α = 0.05 with no multiple-comparison correction, so roughly 1.8 false positives are expected by chance alone. The stars in Figure 7 should therefore be interpreted with caution. Additionally, the 6-point Likert responses are binarized into a 2AFC for analysis, discarding the strength of preference that the scale was designed to capture. The analysis should be re-run with appropriate multiplicity control (e.g., Bonferroni or FDR), and effect sizes or confidence intervals should be reported so that readers can judge the magnitude and precision of the claimed effects.","section":"§4.2, Figure 7"},{"comment":"The attention-switch analysis in Figure 8 uses Fisher's exact test on paired contingency tables, but Fisher's exact test assumes independent (unpaired) observations. Because the same 30 participants answered the attention question for both the InfoVid and the baseline version, a paired test such as McNemar's test should be used. The reported p-values for the attention comparisons are therefore not valid as computed. This issue directly affects the claim that 'more than half the participants switched from focusing on the visualizations to the presenter with InfoVids.'","section":"§4.2, Figure 8"},{"comment":"The paper's claim that InfoVids 'reduced viewer attention splitting' is not directly measured. Survey questions 8 and 9 ask which stimulus—presenter or visualization—the participant viewed more, which measures relative attention allocation, not the degree to which attention was split, divided, or conflicted. The qualitative data include statements like 'two videos were fighting for their attention,' but no systematic measure of split attention (e.g., eye-tracking, self-reported cognitive load, or a dedicated split-attention scale) is presented. The abstract-level claim of reduced attention splitting should be softened to 'reported less attention conflict' or supported by an appropriate measurement.","section":"§5, 'InfoVids Reduce Split Viewer Attention...'"}],"minor_comments":[{"comment":"There are several typographical errors that should be corrected in a revision: 'presnter' in §2.2, 'showcae' in §5, 'inseprable' in §3.3, 'T orso' in Figure 5, and inconsistent capitalization of 'Vishandler' in Appendix A. The table header 'W almartVis' should be formatted consistently.","section":"Throughout"},{"comment":"The design duration is inconsistent: the Introduction says InfoVids were 'iteratively designed over the span of four months,' while §3 says the iterative design process took place 'over the course of nine months' and §6 refers to a 'nine-month experience.' Please reconcile these numbers.","section":"Introduction vs. §3 vs. §6"},{"comment":"Figure 7 shows significance stars but no effect sizes, confidence intervals, or exact p-values. The caption says significance is from a binomial test at α = 0.05, but the reader cannot tell which bars are significant or how large the underlying preferences are; adding numeric annotations would improve transparency.","section":"Figure 7"},{"comment":"The caption of Figure 6 refers to an 'orange overlay' to illustrate rule-of-thirds, but the figure (as described in the text) does not clearly label the overlay. Please make the overlay explicit in the figure or describe it in the caption.","section":"Figure 6"}],"recommendation":"major_revision","confidential_remarks":"The paper's candid appendix is both a strength and a liability: the authors acknowledge that their baseline construction detracted from the InfoVid condition, but they do not seem to fully recognize that this undermines the causal framing of their abstract. The paper may still be publishable if the claims are substantially reframed as exploratory and condition-specific, and if the statistical analyses are corrected or de-emphasized. The quantitative layer currently gives a false impression of rigor relative to the strength of the conclusions."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should know two things about this paper. First, it opens up a concrete new design space for AR-based data presentation: co-locating a full-body presenter and visualization in one 3D frame, with the Body Object Model as an authoring concept. Second, the paper's central quantitative comparison rests on a baseline that the authors openly modified in ways that disadvantage the InfoVid format, so the numbers should be read as indicative, not causal.\n\nThe four InfoVid probes are well chosen and thoughtfully designed. The autobiographical design process is candid, the external actor and same-take control are good instincts, and the qualitative themes match the vote tallies. The attention-shift result—viewers moving from visualization to presenter in the InfoVid condition—is genuinely interesting, as is the WalmartVis counterexample showing when body-binding hurts. Those parts are real contributions.\n\nThe soft spot is the baseline. As the authors describe in Appendix B.1, they scripted a common body language for both conditions—large open hand movements, facing the camera—because the original InfoVid performance would look awkward in a slide format. They admit this \"detracted from InfoVid benefits\" and made the presenter look less immersed. That means the measured differences between conditions are likely driven in part by the baseline being an artificial, constrained performance, not just by the spatial co-location. The same applies to the upper-body-only framing and the post-hoc synced animations. None of this invalidates the qualitative findings, but it does undercut the abstract's causal phrasing.\n\nThe statistical layer is also fragile: 36 binomial tests at alpha 0.05 with no correction, 6-point Likert responses collapsed to 2AFC, and no confidence intervals. That's fixable with a more conservative analysis. The authors also didn't release stimuli or data, which is a miss for a study this dependent on the visual materials.\n\nWho should read this? HCI and visualization researchers working on AR storytelling or data presentation will get useful design vocabulary and a clear set of open problems. It deserves a serious peer review, but the authors should be pushed to reframe the claims as exploratory and to address the baseline confound directly—ideally with a more natural slide condition using the same full-body performance, or at least with a sensitivity analysis. I'd take it to a reading group but would not cite the quantitative results as evidence until those issues are resolved.","headline":"A genuinely new design space with an honest, exploratory study, but the headline quantitative claims are undercut by a baseline the authors themselves admit was artificially constrained.","tokens_in":20542,"tokens_out":1851,"would_cite":false,"duration_ms":20047,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Co-locating the presenter and the visualization in one shared 3D frame reduces viewers' attention splitting and shifts their focus from the visualization to the presenter.","keywords":["InfoVids","augmented reality presentations","data storytelling","presenter-visualization relationship","full-body data performance","viewer attention","body-vis attachment","technology probes"],"falsifier":"Run an eye-tracking study comparing an InfoVid with a baseline where the presenter is allowed to look at and point to the slide content naturally, instead of using the scripted common body language; if switching counts and gaze dwell times are equal, the claimed attention-splitting reduction is an artifact of the baseline style, not of co-location.","tokens_in":19670,"feed_emoji":"🎬","tokens_out":8007,"duration_ms":77320,"temperature":0.7,"pith_summary":"The paper tries to establish that a presentation format can treat the presenter and the data visualization as equal partners inside one shared three-dimensional space, rather than confining the presenter to a small box next to slides. It designs four 'InfoVids'—short information videos that place augmented visualizations around a full-body presenter—and compares each with a 2D slide-style baseline using 30 viewers and nine forced-choice metrics. The authors report that co-location reduced the sense of splitting attention between two competing areas, moved viewers' self-reported focus from the visualization toward the presenter, and made the presenter's body language feel more natural, engaged, and story-relevant. They also show the boundary condition: when a visualization is attached to the body with no narrative purpose, the same format is rated worse than slides. If the effect is real, it gives video and AR presentation designers a concrete design lever: where the presenter stands relative to the data can itself tell the story.","feed_headline":"Putting presenter and chart in one frame cuts split attention","feed_subtitle":"In a 30-person comparison, co-located augmented-reality videos pulled focus from charts to the presenter","key_machinery":"The machinery that carries the argument is the Body Object Model (BOM), a phone-based augmented-reality authoring and performance system that treats the presenter's face, joints, and hands as 'BodyAnchors' onto which 'VisNodes' (containers for 3D visualizations) can be nested, with a 'VisHandler' scripting gestures and actions. This nesting lets a designer create 'body-vis attachments' in which the visualization and the presenter move each other simultaneously, rather than one simply following the other. The system is used to produce four InfoVids implementing three design conditions: use of 3D physical space (C1), body-vis attachments for bidirectional interaction (C2), and unilateral body-vis interactions by the presenter (C3). The experimental contrast pairs each InfoVid with a slide-style baseline recorded from the same performance take, so the reported differences are attributed to the spatial relationship rather than to the content or the actor's delivery.","core_discovery":"The paper's central claim is that merging the spatial world of the presenter with the world of the visualization changes the viewing experience in measurable ways: viewers switch less between presenter and chart, pay more attention to the presenter, and read the presenter's body as part of the narrative. The evidence comes from four paired probes—AirplaneVis, NapoleonVis, InjuryVis, and WalmartVis—which differ in whether the visualization uses 3D space, is attached to the body, or is controlled by gestures. Across all four, InfoVids beat the slide baseline on perceived presenter immersion, engagement, and co-presence in the same room as the visualization; for three of the four, viewers also rated the presenter's body movement as more natural and the storytelling as stronger, and the majority of the 30 participants switched from focusing on the chart in the baseline to focusing on the presenter in the InfoVid. The counterexample, WalmartVis, shows the same co-location can hurt: when a map is strapped to the chest with no narrative benefit, viewers find the movement unnatural and prefer the baseline for enjoyability, storytelling, and information understanding. The authors conclude that the relationship between presenter and visualization, not the mere fact of co-location, is what drives the experience.","pith_inferences":["One testable extension is a full factorial design that crosses 3D depth (C1) with body binding (C2/C3) to separate the effect of depth from the effect of attachment; the current four probes vary both together.","Self-reported focus could be checked with eye tracking or saliency maps; if gaze data confirm the attention shift, the finding becomes a stronger design principle rather than a preference.","The co-location principle should transfer to any single-frame medium, even a carefully staged 2D video where the presenter's hands directly control on-screen graphics; a comparison with such a baseline would isolate the role of true 3D depth.","Building on the paper's self-spectating observation, future performance tools may need separate private cues for the presenter and public visuals for the viewer, letting the presenter interact without visible awkwardness."],"forward_implications":["Video-conferencing and slide tools could offer a co-located mode that shows the presenter full-body inside the same frame as the visualization, rather than boxing the presenter into a corner.","Full-body visibility alone can improve perceived naturalness and storytelling, since the simplest InfoVid, AirplaneVis, used no body attachment and still beat its baseline on those metrics.","Body-vis attachment should be reserved for moments where the body carries narrative meaning; a purposeless attachment, as in WalmartVis, reverses the benefits and makes the performance feel awkward.","Presentation designers should expect some viewer resistance when the format contradicts the mental model of a formal academic talk; acceptance may depend on the viewing context and the presenter's identity.","Authoring tools should reduce the presenter's memorization burden, for example by triggering animations from the presenter's spatial location instead of requiring remembered gestures."],"supporting_citations":[{"why":"It supplies the technology-probes methodology used to deploy InfoVids as open-ended probes into an unknown design space.","marker":"[33]"},{"why":"It provides the motivating example of a presenter and animated charts co-existing in one video frame during a data talk.","marker":"[58]"},{"why":"It demonstrates interactive body-driven graphics for augmented video performance, the basis for letting the presenter control visualizations.","marker":"[60]"},{"why":"It shows a real-time AR presentation system in which speech and gestures drive 2D visuals, grounding the interaction conditions.","marker":"[41]"},{"why":"It contributes hand-gesture languages and layout experiments for presenting data remotely, informing the design of presenter-visualization interactions.","marker":"[28]"},{"why":"It establishes the precedent of attaching visualizations to the body with a wearable display, supporting the body-vis attachment condition.","marker":"[52]"},{"why":"It overlays data visualizations directly on the presenter's body in a magic-mirror setup, another precedent for body-as-canvas designs.","marker":"[67]"},{"why":"It supplies the autobiographical design method that justifies the paper's long-term designer-perspective lessons.","marker":"[50]"},{"why":"It provides the user-spectator performance framework that the paper extends to account for self-facing cameras.","marker":"[19]"}],"fun_headline_variants":["Co-located presenter and chart reduces split attention","Presenter-chart co-location shifts focus, but not always","When chart and presenter share space, attention moves to presenter","InfoVids: same-frame charts and presenters cut attention splitting","Co-location works only when the chart adds narrative meaning"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The baseline slide videos are a fair comparison condition, even though using the same performance forced the actor into scripted 'common body language'—large open hand gestures and facing the camera—that the paper acknowledges weakened the InfoVid presentation, so some of the measured differences may come from baseline design choices rather than from co-location itself.","fun_headline_variants_meta":{"raw":{"variants":["Co-located presenter and chart reduces split attention","Presenter-chart co-location shifts focus, but not always","When chart and presenter share space, attention moves to presenter","InfoVids: same-frame charts and presenters cut attention splitting","Co-location works only when the chart adds narrative meaning"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000373,"raw_usage":{"total_tokens":2004,"prompt_tokens":966,"completion_tokens":1038,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":582,"completion_tokens_details":{"reasoning_tokens":956}},"tokens_in":582,"tokens_out":1038,"duration_ms":11086,"temperature":1.0,"reasoning_tokens":956,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T23:57:22.526226+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run an eye-tracking study comparing an InfoVid with a baseline where the presenter is allowed to look at and point to the slide content naturally, instead of using the scripted common body language; if switching counts and gaze dwell times are equal, the claimed attention-splitting reduction is an artifact of the baseline style, not of co-location.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"It provides the motivating example of a presenter and animated charts co-existing in one video frame during a data talk."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"It demonstrates interactive body-driven graphics for augmented video performance, the basis for letting the presenter control visualizations."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"It establishes the precedent of attaching visualizations to the body with a wearable display, supporting the body-vis attachment condition."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"It overlays data visualizations directly on the presenter's body in a magic-mirror setup, another precedent for body-as-canvas designs."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"It supplies the autobiographical design method that justifies the paper's long-term designer-perspective lessons."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"It provides the user-spectator performance framework that the paper extends to account for self-facing cameras."}],"review_version":1}