{"id":"3379bb94-0496-46b7-869a-9a0770e50179","arxiv_id":"2608.00403","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"In a 24-person VR study, proposal-based window placement reduced layouting time but users preferred direct manual control, and the way proposals were visualized changed both performance and preference.","lead":"Three ways of showing users a suggested position for a new window in virtual or mixed reality were tested against freehand placement. The suggestions saved layout time, but users still preferred placing windows by hand, a result that matters for how adaptive interfaces should be built.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Fixed pilot-defined proposal positions (Sec. 3.5) are the load-bearing assumption behind the Manual-preference claim; existing manual-placement logs can test whether users' chosen positions actually diverged from those proposals.","rationale":"The reader's weakest_assumption correctly identifies the fixed proposal positions as the load-bearing premise. I see no additional concern that would force a stronger verdict. The between-proposal comparisons are protected by holding positions constant; the Manual-vs-proposal contrast is not. The paper's limitations section is unusually candid, but candor does not remove the need for a test. The proposal-position concern is testable from data the study should already have, which makes it the right focus. The VR/MR mismatch and the instruction-induced relayouting are real but secondary: they limit generalization rather than threatening the core controlled comparison. Statistical reporting has a labeling inconsistency in Appendix A (the 'Test stat z' column does not match the z values in §5), but the Bonferroni-adjusted p-values and effect sizes are consistent, so I do not treat it as load-bearing. My recommendation is to keep the reader's CONDITIONAL verdict; the condition should be an explicit test of the proposal-set representativeness before the preference claim is generalized to adaptive proposal systems.","tokens_in":23340,"tokens_out":7316,"duration_ms":83062,"concrete_test":"Use the logged final window transforms from the Manual condition (per participant, per window) to compute the distance from each final placement to the nearest of the four fixed proposal positions used in the proposal conditions. If the median distance is small (within one window width) and manual final positions cluster around the pilot positions, the fixed set was representative and the preference result is unlikely to be an artifact of bad proposals; if manual placements systematically fall far from the fixed set, the proposal set was not representative and the reported MA-vs-proposal preference cannot be attributed to proposal-based interaction per se. If logs are not retained, run a follow-up within-subjects condition comparing Manual with the same three visualizations using dynamically generated (e.g., Pareto-optimal) proposals; if the Manual preference gap persists, the fixed-set","verdict_should_be":"UNCHANGED","load_bearing_attack":"All three proposal conditions used identical four expert-chosen positions per window (Sec. 3.5, Sec. 4). This makes the between-proposal contrasts (e.g., SWP faster than 3D) internally clean, but the central MA-vs-proposals contrast is not: if the fixed set was merely 'adequate' rather than representative of what an adaptive system would offer, both the large time savings and the strong preference for Manual could be artifacts of those specific positions. The paper's Sec. 7 concedes that the Manual preference 'may therefore partly reflect the constraint of choosing from a fixed set.' The qualitative data do not resolve this: 11/24 called proposals restrictive, 20/24 wanted manual fine-tuning, and 20/24 found positions reachable/visible—so position adequacy is not established enough to separate 'dislike of proposal interaction' from 'dislike of these proposals.' Secondary concerns (VR vs MR, instruction-induced relayouting) affect external generalization but do not threaten the internal comparison; the fixed-proposal issue is the one whose resolution could change the headline conclusion.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports a within-subjects user study (N=24) comparing three proposal-visualization techniques for placing windows in a VR-based MR environment—Situated Icon Preview, Situated Window Preview, and 3D Preview—against a Manual Positioning baseline. Participants completed a seven-window trip-planning task under each condition. The main quantitative findings are that Situated Icon Preview and Situated Window Preview significantly reduced layouting time relative to Manual Positioning, Situated Window Preview was faster than 3D Preview, and Manual Positioning nevertheless received significantly higher preference ratings. A thematic analysis of interviews identifies perceived control, cognitive cost, familiarity, and proposal informativeness as factors shaping preference, and a majority of participants endorsed combining proposals with manual fine-tuning. The paper includes Bonferroni-adjusted pairwise comparisons, a trial-order analysis, a post-hoc sensitivity analysis, and an open-science statement with data, instruments, and codebook.","tokens_in":23478,"tokens_out":7076,"duration_ms":77466,"significance":"If the results are taken at face value, the paper makes a useful empirical contribution to MR window-management design: it shows that the visualization format of layout proposals is a consequential design choice and that efficiency gains do not automatically translate into user preference, consistent with prior findings on controllability versus automation accuracy. The study is methodologically careful in several respects: full Latin-square counterbalancing, Bonferroni correction, a trial-order analysis, a sensitivity/power analysis, and a transparent limitations section. The qualitative analysis is reported with a detailed codebook. However, two issues temper the headline claims: the use of a single fixed set of expert-defined proposal positions confounds the manual-versus-proposal comparison, and one of the headline pairwise contrasts falls below the paper's own stated sensitivity threshold. Both are addressable, but they need substantive attention before the central claims can be accepted as stated.","major_comments":[{"comment":"The central comparison between proposal conditions and Manual Positioning rests on a fixed set of four expert-defined proposal positions per window, held identical across all three proposal conditions. This is clean for between-proposal contrasts, but the Manual-versus-proposal comparisons are only interpretable if those positions are representative of what users would consider good proposals. The paper itself concedes in §7 that 'the preference for Manual Positioning may therefore partly reflect the constraint of choosing from a fixed set, rather than a principled rejection of proposal-based interaction.' The qualitative data do not resolve the ambiguity: 20/24 found the positions reachable and visible, but 11/24 called the proposals restrictive and 20/24 wanted manual fine-tuning. Because the headline conclusion is that users prefer manual control despite time savings, the manuscript n","section":"§3.5, §4, §7"},{"comment":"Appendix C states that the study has approximately 80% power to detect pairwise effect sizes of r ≥ .50 after Bonferroni correction. Under that threshold, the headline contrast that Situated Window Preview is faster than 3D Preview (Table A.1: z = 2.795, p_adj = .031, r = .40) is below the stated detectable effect. The same applies to the UEQ Perspicuity contrast 3D vs. Situated Window Preview (r = .43) and the Dependability/Efficiency contrasts (r = .41–.44). These are nevertheless reported as significant main findings in §5.1.1 and §5.2.2, while the authors explicitly caution that two nonsignificant omnibus effects (Stimulation, Attractiveness) sit below the detection threshold. This is internally inconsistent. Either the Appendix C thresholds should be recomputed or explained (e.g., in terms of achieved power for the observed effects), or the below-threshold significant contrasts shou","section":"§5.1.1, Appendix C, Table A.1"}],"minor_comments":[{"comment":"The title and abstract say 'Mixed Reality' while the study is conducted in VR with a seated, desk-based setup. The limitations section acknowledges this, but the main text would benefit from explicitly using 'VR' in the title or at least in the abstract's first sentence.","section":"Title/Abstract"},{"comment":"Preference was rated once at the end of the session after all four conditions. The trial-order analysis in Appendix B covers layouting time, layout changes, and overall task completion, but not preference ratings. A sentence explaining how the retrospective preference measure interacts with condition order would strengthen the reporting.","section":"§4.5/§5.2.3"},{"comment":"Minor typographical inconsistency: the Stimulation omnibus reports χ²(3) = 8.12, p = .044 without a Kendall's W value, while the Attractiveness omnibus reports W = .13. Providing W for both would make the effect sizes comparable.","section":"§5.2.2"},{"comment":"The pilot study with four experts is described as determining both the number of proposals and the specific positions. It would be helpful to state explicitly how the 'maximally distant non-occluding positions' were extracted and whether the four experts were authors or external users, since this affects the independence of the proposal set.","section":"§3.5"},{"comment":"The correlations between prior experience and preference are exploratory and based on N=24. The paper already labels them as indicative; consider adding a sentence in the main text to prevent readers from over-interpreting the marginal p = .051 result.","section":"Appendix E"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is a solid empirical study with transparent reporting, but the two load-bearing issues—the fixed proposal positions and the sensitivity-threshold inconsistency—need to be addressed before the central claims are fully convincing. Both are fixable within the scope of the paper: the first by analyzing logged manual positions or reframing the claim, and the second by correcting the sensitivity interpretation. I would not reject; this is a worthwhile contribution if the claims are brought in line with the evidence."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper delivers a genuinely new empirical result: the first direct within-subjects comparison of three proposal-visualization techniques (situated icon, situated window, world-in-miniature) against manual positioning for single-window placement in MR. That is a real gap, and the study is competent. The headline effects are large, survive Bonferroni adjustment, and the authors report a sensitivity analysis showing they are powered only for medium-to-large effects, then restrict their claims accordingly. The auxiliary checks—trial-order analysis, full pairwise tables, codebook—are there. Credit where due: this is a well-executed HCI study with an unusually honest limitations section.\n\nThe internal comparison among the three proposal techniques is clean, since they share the same four pilot-derived positions per window. The between-proposal result (Situated Window Preview faster than 3D Preview, lower perspicuity for 3D) is the most defensible part of the paper. The manual-vs-proposals contrast is where the soft spot sits: all proposal conditions used the same fixed expert-chosen positions, so the large time savings and the strong preference for manual placement could partly be artifacts of that specific set. The authors concede this in Section 7, saying the preference may reflect the constraint of choosing from a fixed set rather than a rejection of proposal interaction. The qualitative data do not fully resolve it either—20 of 24 called the positions reachable and visible, but 11 called proposals restrictive and 20 wanted manual fine-tuning, which is consistent with both readings. This is not a fatal flaw, because the paper frames itself as comparing visualization designs under controlled proposals, not as a test of adaptive algorithms. But the implications section reaches toward hybrid systems and dynamic Pareto-optimal proposals, and that generalization is untested.\n\nTwo smaller issues: the layout-change DV uses different event definitions per condition, and the authors mention rotation corrections may have inflated the manual count; that gap is real but disclosed. And the study is run entirely in VR while the title and framing say MR; the seated desk scenario is a limitation the authors acknowledge. Neither threatens the internal logic.\n\nWho is this for? Anyone working on adaptive MR interfaces, window management in spatial computing, or semi-automated interaction. It deserves a serious referee; the right outcome is likely conditional acceptance with a request to temper the manual-preference generalization and maybe add a short discussion of how the fixed-proposal premise could be tested with dynamic proposals. I would cite it if I were working in this area.","headline":"Solid, cleanly reported user study showing proposal-based placement saves time but loses to manual control on preference; the one load-bearing caveat is the fixed proposal set, which the authors themselves disclose.","tokens_in":24062,"tokens_out":1126,"would_cite":true,"duration_ms":13594,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Proposal-based window placement saves time in MR, but users still prefer direct manual control; the visualization itself is a consequential design choice.","keywords":["mixed reality","window arrangement","layout proposals","user study","perceived control","adaptive user interfaces","world-in-miniature","user experience"],"falsifier":"Run the same seven-window trip-planning task in VR with proposals generated per participant by a multi-objective optimizer (e.g., optimizing reachability, visibility, and ergonomics) instead of the fixed pilot positions, while keeping the three visualizations and manual baseline unchanged. If manual positioning is still preferred over all dynamic-proposal conditions, the controllability explanation is confirmed; if a dynamic-proposal condition matches manual positioning's preference rating, the paper's central preference result was at least partly an artifact of the fixed proposal set.","tokens_in":23121,"feed_emoji":"🪟","tokens_out":8745,"duration_ms":81706,"temperature":0.7,"pith_summary":"This paper asks how layout proposals for windows in mixed reality should be shown to users, and whether proposing positions beats letting them place windows by hand. In a within-subjects VR study, 24 participants planned a trip with seven windows under three proposal visualizations — situated icon previews, full-size situated window previews, and a world-in-miniature 3D preview — plus a manual-positioning baseline. The situated proposal techniques roughly halved layouting time relative to manual positioning, and situated window previews were also faster than the 3D mini-map. Yet the manual condition received the highest preference rating, significantly ahead of all three proposal techniques. The authors conclude that proposal visualization is a genuine design decision, that users trade efficiency for a sense of control, and that hybrid systems combining suggestions with free manual refinement are the road forward.","feed_headline":"Proposals cut MR layout time, but users still pick manual","feed_subtitle":"In 24 VR trip-planning sessions, situated previews halved layouting time; preference still went to direct control.","key_machinery":"The evaluative machinery is a three-dimensional design space for proposal visualizations — degree of automation (semi-automated selection vs. full manual placement), level of detail (position-only icons vs. full-size window frames with content), and interaction-space awareness (first-person situated views vs. a world-in-miniature overview). To isolate visualization from algorithm, all three proposal conditions show the same four predefined positions per window, drawn from a pilot with four experienced MR users, so differences in time and preference are attributable to how the proposal is shown, not what is shown. The qualitative coding of post-study interviews supplies the four explanatory f","core_discovery":"The paper's central claim is a preference–efficiency trade-off. With four fixed proposal positions per window (derived from a four-expert pilot and held constant across proposal conditions), Situated Icon Preview and Situated Window Preview reduced layouting time from a manual mean of 120.3 s to 58.8 s and 54.5 s respectively, and the situated window preview also beat the 3D preview (54.5 s vs 79.4 s) on both layouting time and overall task completion. Despite this, Manual Positioning was preferred by participants (mean 4.46/5) over every proposal technique, and the 3D preview scored worst on UEQ Perspicuity. Interviews attribute the preference to four factors: perceived control over placeme","pith_inferences":["My inference: the preference gap may partly be a fixed-set artifact; a system that generates proposals on the fly from the user's current activity, or from their prior adjustments, could close the gap, but this is precisely the condition the paper deliberately left untested.","My inference: a hybrid interaction where proposals appear only on explicit request, and the user can grab and adjust any proposal before accepting it, would test whether the 'restrictive' feeling disappears while the time savings remain.","My inference: since experienced VR/3D users liked position-only icons less, proposal visualizations could be adapted per-user or per-familiarity, e.g., content-aware icons for novices and minimal icons for experts.","My inference: the same framework could generalize to non-window MR UI elements such as notifications, panels, and annotations, where the trade-off between preview informativeness and clutter is likely even starker."],"forward_implications":["Semi-automated proposal selection roughly halves window-layouting time in MR and cuts the number of layout adjustments by about half relative to manual positioning.","Among proposal visualizations, full-size situated previews are the strongest: they beat the world-in-miniature 3D preview on layouting time and overall task completion, while position-only icons are fastest per selection but least informative.","A world-in-miniature view is not inherently better: its overview benefit is offset by multi-step selection and unfamiliarity, producing worse perspicuity and no preference gain.","User acceptance of automated assistance in spatial layout is governed by perceived control and familiarity, not raw efficiency; any deployed system should include a manual fine-tuning path.","Layout assistance is most valuable when many windows accumulate: readjustments concentrated in the last two of seven task stages, suggesting a threshold near five or more windows."],"supporting_citations":[{"why":"Supplies the semi-automated, one-window-at-a-time proposal paradigm and the Pareto-optimal position sets that the study adopts as state of the art.","marker":"[29]"},{"why":"Motivates semi-automated selection over full automation and provides the situated icon-preview representation that the lowest-detail condition builds on.","marker":"[30]"},{"why":"Supplies the context-aware adaptive-layout objectives (visibility, reachability, ergonomics) and the progressive multi-window task structure the trip-planning task adapts.","marker":"[34]"},{"why":"Sources the central theoretical anchor that controllability outweighs automation accuracy in user preference, which the paper extends from 2D to immersive MR.","marker":"[47]"},{"why":"Supplies the spatially situated window-preview idea behind the Situated Window Preview condition.","marker":"[41]"},{"why":"Contextualizes the proposal-selection paradigm by optimizing which positions are offered, against which the paper argues presentation matters equally.","marker":"[51]"},{"why":"Provides the NASA TLX workload instrument used to measure subjective task load.","marker":"[24]"},{"why":"Provides the UEQ instrument used to measure user-experience subscales.","marker":"[31]"}],"fun_headline_variants":["Proposal previews halve layout time in VR, manual still preferred","MR layout proposals save time, yet users stick with manual","Faster window placement via proposals, but control wins","Proposals speed MR layout, but preference stays manual","Time-efficient proposals in MR fail to beat manual preference"],"cache_read_input_tokens":2688,"weakest_assumption_plain":"The load-bearing premise is that the four fixed proposal positions, set once by a four-person pilot and reused for every proposal condition, are representative of what a real adaptive layout algorithm would suggest; if those positions were merely adequate or mismatched to participants' preferences, both the time savings and the preference for manual control could be artifacts of that fixed set rather than of the visualization designs.","fun_headline_variants_meta":{"raw":{"variants":["Proposal previews halve layout time in VR, manual still preferred","MR layout proposals save time, yet users stick with manual","Faster window placement via proposals, but control wins","Proposals speed MR layout, but preference stays manual","Time-efficient proposals in MR fail to beat manual preference"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000223,"raw_usage":{"total_tokens":1314,"prompt_tokens":781,"completion_tokens":533,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":525,"completion_tokens_details":{"reasoning_tokens":452}},"tokens_in":525,"tokens_out":533,"duration_ms":6281,"temperature":1.0,"reasoning_tokens":452,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T04:13:08.608914+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same seven-window trip-planning task in VR with proposals generated per participant by a multi-objective optimizer (e.g., optimizing reachability, visibility, and ergonomics) instead of the fixed pilot positions, while keeping the three visualizations and manual baseline unchanged. If manual positioning is still preferred over all dynamic-proposal conditions, the controllability explanation is confirmed; if a dynamic-proposal condition matches manual positioning's preference rating, the paper's central preference result was at least partly an artifact of the fixed proposal set.","supporting_citations":[],"review_version":1}