{"id":"9b545651-d67e-4993-a1ef-fac5cfc99905","arxiv_id":"2502.00908","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"People prefer having control over placement of a shared AR whiteboard, but the statistical evidence is weak and automatic placement was often acceptable.","lead":"A user study with 36 participants compared three ways to start a shared whiteboard in simulated AR: drawing it by hand, choosing one of three suggestions, or letting the system pick. The result is a design caution: people value control, but automation is acceptable when it works well.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"AUTOMATIC condition's hand-picked whiteboard placements, not an implemented algorithm, confound level of control with placement quality; the central preference claim therefore rests on an unrepresentative automatic condition.","rationale":"The reader's weakest_assumption correctly identifies the most load-bearing threat: the AUTOMATIC condition is not actually automatic. This threatens the central claim because the comparison between control and no control is entangled with the idiosyncratic quality of one researcher's hand-picked whiteboards. If a real algorithm would place boards differently, the observed preference for user control could be an artifact of the specific placements rather than a general property of level of control. This is more fundamental than the non-significant three-way ranking chi-square (Section 5.1), because the central claim can be read as a binary control-versus-no-control preference; grouping MANUAL and DISCRETE CHOICE against AUTOMATIC gives 26/36 first-choice support, so a reanalysis might support the claim even with the existing data. The automatic-condition concern, by contrast, undermines the construct validity of the independent variable itself. The paper is otherwise careful: the discussion is appropriately cautious, limitations are acknowledged, and the qualitative data are informative. Since the reader already rated this CONDITIONAL, no verdict change is needed; the concern confirms the need for the stated condition that future work implement and validate a real automatic algorithm.","tokens_in":13490,"tokens_out":6549,"duration_ms":69789,"concrete_test":"Ask several independent raters to apply the C-SAW heuristics to the same environment pairs and datasets used in the study, or implement a simple algorithmic version of C-SAW, and compare the resulting whiteboard rectangles to those used in the AUTOMATIC condition (e.g., intersection-over-union, center distance, dimension ratios). If inter-rater variance is high or an algorithmic placement differs materially for any trial, the AUTOMATIC condition is not a stable representative of an automatic system, and the control-preference result would need to be re-examined with a real automated placement pipeline.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The study's central claim is that users prefer having control over whiteboard size, shape, and location during initialization. That claim depends on the AUTOMATIC condition being a fair instantiation of 'no user control.' However, C-SAW was never implemented; a single unaffiliated researcher manually chose whiteboard sizes and locations according to the heuristics (Sections 3.2 and 3.3). The resulting placements are thus one person's idiosyncratic judgment, not a reproducible algorithm's output. Because placement quality varied (some boards were partially occluded by furniture, per participant reports), participants' relative preference for MANUAL and DISCRETE CHOICE may reflect dissatisfaction with those particular hand-picked placements rather than with automation per se. The paper's own discussion concedes that users liked AUTOMATIC when the algorithm 'worked well.' Since the automatic condition cannot be regenerated or varied systematically, the independent variable is not cleanly operationalized, and the conclusion generalizing to real system-assisted AR initialization is not securely supported.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper investigates how the level of user control during the initialization of a shared whiteboard in collaborative AR affects the collaboration experience. Through a within-subjects VR study (36 participants in 18 dyads), the authors compare three initialization techniques: MANUAL (users draw and negotiate a shared whiteboard), DISCRETE CHOICE (users select from three system-generated suggestions), and AUTOMATIC (system provides a whiteboard with no user input). Dependent measures include preference rankings, NASA-TLX, UEQ, SUS (for MANUAL and DISCRETE CHOICE only), initialization time, whiteboard characteristics, and post-study interviews. The authors report that while there were no statistically significant differences in preference rankings, qualitative comments suggested that participants valued control over size, shape, and location, and the paper concludes that users preferred having direct control. The AUTOMATIC condition was operationalized by having a human researcher select whiteboard placements according to the proposed C-SAW heuristics, rather than by an implemented algorithm.","tokens_in":13677,"tokens_out":3985,"duration_ms":38948,"significance":"The paper addresses a genuinely under-explored problem—the initialization of shared surfaces in remote AR collaboration—and contributes a novel comparison of three control levels. The study design is generally careful: counterbalanced within-subjects assignment, varied virtual environments, realistic dyadic collaboration tasks, and both quantitative and qualitative measures. The authors are also transparent about many limitations, including the VR simulation, the absence of digital whiteboard affordances, and specific technical issues in the MANUAL condition. If the central claim about user preference for control were properly supported, the work would offer concrete design guidance for future AR collaboration systems. However, as reported, the quantitative data do not support the abstract's and conclusion's claims of a majority preference for direct control, and the AUTOMATIC condition's reliance on human-picked placements weakens the internal validity of the level-of-control manipulation.","major_comments":[{"comment":"The abstract and Section 8 claim that 'the majority of participants preferred to have direct control' and that participants 'overall preferred having control over the size, shape, and location of whiteboards.' Section 5.1 reports that 18 of 36 participants ranked MANUAL first, 8 ranked DISCRETE CHOICE first, and 10 ranked AUTOMATIC first, with a chi-square test showing no significant difference (p = 0.12). Eighteen out of 36 is exactly half, not a majority, and the non-significant test does not support a preference ordering among the conditions. This is a load-bearing overstatement of the study's central finding.","section":"Abstract, Section 5.1, Section 8"},{"comment":"The AUTOMATIC condition was not produced by an implemented algorithm; rather, a human researcher manually selected whiteboard sizes and locations following the C-SAW heuristics. This confounds 'level of control' with 'placement quality.' Because the researcher's placements were idiosyncratic and, as Section 6.1 documents, sometimes partially occluded by furniture, participants' relative dissatisfaction with AUTOMATIC may reflect these specific hand-picked placements rather than automation per se. The paper's own discussion concedes that users liked AUTOMATIC when the algorithm 'worked well' (Section 6.1). Without an implemented or systematically varied automatic placement generator, the independent variable is not cleanly operationalized, and the conclusion that users prefer manual control over automatic initialization is not securely supported.","section":"Sections 3.2 and 3.3"},{"comment":"The paper states that 'the participant ranking data, as well as comments made to us during the post study interview allow us to conclude that collaborators do prefer to have explicit control.' This inference is problematic because the ranking data show no statistically significant difference among the three techniques, and the interview comments are retrospective and may be influenced by the technical issues reported in the MANUAL condition (e.g., network failures and wall-encroachment errors, also described in Section 6.1). The qualitative evidence is suggestive and useful for generating hypotheses, but it does not by itself establish a preference for control that is independent of the specific implementation issues encountered.","section":"Section 6.1"},{"comment":"The Limitations section does not acknowledge that the AUTOMATIC condition used human-selected placements rather than a reproducible algorithm. This is a significant threat to the internal validity of the comparison between automated and user-controlled initialization and should be explicitly discussed as a limitation, along with its implications for the generality of the results to real-world AR systems that would run an actual C-SAW implementation.","section":"Section 7"}],"minor_comments":[{"comment":"The test statistic is reported as 'F(4) = 7.33' for a chi-square test; it should be reported as a chi-square statistic (e.g., χ²(4) = 7.33) to avoid confusion with an F-test.","section":"Section 5.1"},{"comment":"The text references 'as described in Section 5.3' when discussing whiteboard sizes for different datasets, but the relevant results are in Section 5.4.","section":"Section 6.1"},{"comment":"The word 'majority' is used to describe the preference result, but 18 out of 36 is exactly half (a plurality, not a majority); consider rephrasing to 'half' or 'a plurality' to match the reported numbers.","section":"Abstract and Section 8"},{"comment":"The C-SAW heuristics are stated as a list, but Figure 3, which illustrates C-SAW applied to two environments, is not referenced in the text of Section 3.2; adding an explicit callout would help readers connect the heuristics to the example.","section":"Section 3.2"},{"comment":"The SUS questionnaire was not administered for the AUTOMATIC condition; the paper states this but does not provide a rationale. A brief justification (e.g., no user interaction to rate) would be useful for readers evaluating the completeness of the usability measures.","section":"Section 4.2"}],"recommendation":"major_revision","confidential_remarks":"The paper is within the scope of IEEE VR and addresses a valuable, under-explored topic. The study itself appears to have been conducted with reasonable care, and the qualitative analysis is thoughtful. However, the central claim in the abstract and conclusion is not supported by the paper's own statistics, and the AUTOMATIC condition's reliance on human-selected placements is a substantive confound that the authors should be required to address. With revised claims and an explicit treatment of this limitation, the paper could make a solid contribution. I also note that the reported chi-square statistic appears mislabeled as an F statistic, which should be corrected. I would encourage the editor to request a major revision rather than reject, since the underlying study has merit and the issues are fixable within the manuscript's scope."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nPlain take: this is a legitimate first exploration of a real problem—how to initialize a shared whiteboard in remote collaborative AR—and the authors do some things well. The three conditions are thoughtfully designed, the VR simulation is reasonable, and the qualitative analysis of the post-study interviews is careful and honestly reported. The C-SAW heuristic framework, even as a paper design tool, gives the work a handle for future algorithmic work. That said, the central claim in the abstract and conclusion—that participants overall preferred direct control—is not actually supported by the paper's own statistics. Only 18 of 36 ranked MANUAL first, the chi-square test showed no significant difference in rankings, and DISCRETE CHOICE actually had the fewest first-place votes. The authors do acknowledge this in the discussion, but the abstract overstates it.\n\nThe bigger soft spot is the AUTOMATIC condition. C-SAW was never implemented; a single researcher manually picked whiteboard placements following the heuristics. That introduces a confound: participants' reactions to 'AUTOMATIC' are reactions to those specific placements, not to an algorithmic placement system. Some placements were partially occluded by furniture, which no doubt colored preferences. The paper's own discussion admits users were happy with AUTOMATIC when it worked well. You can't generalize the level-of-control comparison when the automatic condition is a human's one-off judgment.\n\nThere are also minor issues: participants were mostly friends or acquaintances, which likely eased the negotiation overhead in MANUAL; and the task was limited to affinity diagramming. The authors flag most of these limitations themselves.\n\nThe paper is worth a serious referee. It addresses a genuine gap, the study is competently executed, and the qualitative findings are useful design guidance. But it needs revision: either soften the preference claim or gather stronger evidence, and clearly scope the conclusion to 'these hand-generated suggestions' rather than to automatic systems in general.\n\nFor what it's worth, I'd bring this to a reading group as an example of a well-intentioned HCI study where the experimental proxy for 'automatic' undermines the clean conclusion.\n\nRecommendation: send it to peer review, with the expectation of a revision that fixes the abstract and narrows the claims.","headline":"A worthwhile first study of shared-whiteboard initialization in collaborative AR, but the headline claim overreaches its own statistics and the 'automatic' condition is hand-picked rather than algorithmic.","tokens_in":14166,"tokens_out":2171,"would_cite":true,"duration_ms":20697,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Users prefer manual control over shared whiteboard placement in collaborative AR, even when it adds setup work.","keywords":["shared whiteboard","collaborative augmented reality","initialization","level of control","C-SAW","affinity diagramming","user preference","virtual reality simulation"],"falsifier":"Implement a real automatic whiteboard placement algorithm that scans both environments and applies the C-SAW heuristics computationally, run the same paired comparison, and check whether AUTOMATIC's satisfaction and preference rankings remain close to MANUAL's; if AUTOMATIC scores drop when placements are generated by the algorithm rather than hand-selected, the paper's conclusion about automatic acceptability would not generalize.","tokens_in":13303,"feed_emoji":"📋","tokens_out":5265,"duration_ms":49050,"temperature":0.7,"pith_summary":"This paper asks whether letting users control how a shared whiteboard is placed in a collaborative augmented reality session matters for the experience of collaboration. The authors compare three initialization techniques: users draw and negotiate the whiteboard themselves (MANUAL), choose among three system-suggested options (DISCRETE CHOICE), or accept a system-generated whiteboard with no input (AUTOMATIC), in a simulated-AR study inside virtual reality. They find that the majority of participants preferred having direct control over the whiteboard's size, shape, and location, even though MANUAL initialization took longer and induced higher workload. The results also suggest that an automatic system could be acceptable if its placement choices were intelligent enough to match user expectations.","feed_headline":"Users want control over shared AR whiteboard placement","feed_subtitle":"Half of 36 participants ranked manual setup first; automatic was fastest but offered no choices.","key_machinery":"The study's central object is the three-level control design: MANUAL (parallel drawing and overlap negotiation), DISCRETE CHOICE (three C-SAW suggestions), and AUTOMATIC (a single C-SAW placement with no interaction). C-SAW (Collaborative Surface Algorithm for Whiteboarding) is a set of heuristics — reachability, avoidance of important objects, placement in open wall space, minimum usable size, and clear viewing space — for choosing candidate whiteboard placements. The mechanism that carries the argument is a within-subjects user study in which 18 pairs of participants completed an affinity diagramming task after each initialization method, with preference rankings, initialization time, NASA-TLX workload, UEQ, SUS, and interview responses as evidence.","core_discovery":"The paper's central claim is that during the initialization of a shared whiteboarding session in collaborative AR, users overall prefer to have explicit control over where and how the shared surface is placed, and this preference survives the added workload of manual setup. In the study, half of the 36 participants ranked the MANUAL technique as their favorite, and interview comments emphasized the value of tailoring whiteboard dimensions to the task, such as making a wide, short board for a chronological timeline. At the same time, the statistical test found no significant difference in technique rankings, and a large number of participants found the MANUAL workflow unintuitive, with an average of ten whiteboard drafts per pair. The paper concludes that collaborators want control, but the initialization design needs to be more efficient and understandable, and that AUTOMATIC placement is acceptable when the underlying placement logic is good.","pith_inferences":["The preference for control may generalize beyond whiteboards to other auto-placed shared AR anchors such as virtual screens, volumes, or furniture, since the underlying issue is users wanting final say over virtual content in their physical space.","The task-dependence observed here suggests a testable extension: an automatic system that infers whiteboard dimensions from task structure or data volume might close much of the gap between AUTOMATIC and MANUAL preference.","Because the study simulated AR in VR, a real-AR replication with physical occlusions, furniture, and lighting could shift results, as the authors themselves flag.","A direct replication with a larger sample could test whether the observed preference plurality becomes statistically significant, since the chi-square test in this study did not reach significance."],"forward_implications":["Collaborative AR systems should offer users a way to manually adjust or override automatically placed shared surfaces when a session starts.","Automatic initial placement is acceptable when the placement logic is intelligent enough to avoid occlusions and match user expectations, so effort should go into improving automatic placement rather than removing it.","Because whiteboard dimensions were tailored to task content, initialization systems should consider the task or dataset when proposing whiteboard shapes.","A hybrid initialization technique that offers an automatic or suggested starting board plus manual override would combine the speed of AUTOMATIC with the control users preferred.","Keeping collaborators aware of each other's whiteboard dimensions during manual setup could reduce negotiation effort and the number of redraws."],"supporting_citations":[{"why":"Supplies the SUS usability questionnaire administered after the MANUAL and DISCRETE CHOICE conditions.","marker":"[2]"},{"why":"Supplies the NASA-TLX workload measure used to compare workload across the three initialization techniques.","marker":"[8]"},{"why":"Supplies the UEQ user experience questionnaire used to rate each condition across six experience scales.","marker":"[15]"},{"why":"Defines affinity diagramming, the collaboration task used after each initialization in the study.","marker":"[17]"}],"fun_headline_variants":["AR whiteboard users favor manual setup despite extra work","Half of users prefer manual AR whiteboard placement","Manual control wins in AR whiteboard setup but not by much","Users want to steer shared AR whiteboard initialization","AR collaboration: control beats convenience in whiteboard setup"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The AUTOMATIC condition was not produced by a real algorithm: a human researcher picked the whiteboard placements by hand using the C-SAW heuristics, so the results assume those hand-picked placements represent what an actual automatic system would do.","fun_headline_variants_meta":{"raw":{"variants":["AR whiteboard users favor manual setup despite extra work","Half of users prefer manual AR whiteboard placement","Manual control wins in AR whiteboard setup but not by much","Users want to steer shared AR whiteboard initialization","AR collaboration: control beats convenience in whiteboard setup"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000543,"raw_usage":{"total_tokens":2585,"prompt_tokens":917,"completion_tokens":1668,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":533,"completion_tokens_details":{"reasoning_tokens":1592}},"tokens_in":533,"tokens_out":1668,"duration_ms":12479,"temperature":1.0,"reasoning_tokens":1592,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-09T17:16:32.864954+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Implement a real automatic whiteboard placement algorithm that scans both environments and applies the C-SAW heuristics computationally, run the same paired comparison, and check whether AUTOMATIC's satisfaction and preference rankings remain close to MANUAL's; if AUTOMATIC scores drop when placements are generated by the algorithm rather than hand-selected, the paper's conclusion about automatic acceptability would not generalize.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the SUS usability questionnaire administered after the MANUAL and DISCRETE CHOICE conditions."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the NASA-TLX workload measure used to compare workload across the three initialization techniques."},{"cited_title":"Laugwitz, T","cited_arxiv_id":null,"evidence_quote":"Supplies the UEQ user experience questionnaire used to rate each condition across six experience scales."}],"review_version":1}