{"id":"41c87a7c-6a9d-4269-b753-c3eda7cb10bb","arxiv_id":"2607.23213","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A 24-user VR study identifies nine types of LLM scene-editing confabulations and shows users systematically underestimate their frequency under high error load.","lead":"This paper runs a lab study where 24 people use a VR scene-editing system driven by a large language model, and catalogs the mistakes the system makes. It introduces the 'perception-reality gap' to show that users notice far fewer errors than actually occur, especially when errors are frequent. The value for a generalist reader is the evidence that user oversight is not a reliable safety net in AI-driven immersive tools, so system-side detection is needed.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Perceived-confabulation questionnaire omits FSC/AEF/MSE categories, so the perception-reality gap and saturation finding may be inflated by instrument mismatch rather than user vigilance breakdown.","rationale":"The paper is a carefully conducted exploratory HCI study, and the reader's CONDITIONAL verdict is appropriate. However, the most load-bearing threat to the central claim is not primarily the lack of inter-rater reliability on the manual confabulation coding, though that is also real. The more direct threat is that the perceived-confabulation questionnaire and the actual-confabulation taxonomy are not aligned: three full taxonomy categories (FSC, AEF, MSE) have no questionnaire items. This means perceived counts are censored from above; a participant who noticed an FSC event could not report it through the structured instrument. The paper explicitly flags this as a 'structural blind spot' in footnote 3, yet the saturation and uncorrelation claims in Section 5 are stated without this caveat. This matters because the central conclusion—'user vigilance is structurally unreliable'—depends on the gap being a perceptual failure rather than an instrument failure. The proposed re-analysis with matched categories is a concrete, feasible check: if the gap and uncorrelated rankings survive after excluding unmapped categories, the conclusion is substantially strengthened; if they do not, the paper's strongest claim must be softened. I therefore keep the CONDITIONAL verdict, with this additional condition on the interpretation of RQ3. This is a partial agreement with the reader because the reader identified manual coding as the weakest assumption, while I see the questionnaire mismatch as more directly tied to the central claim, though both are legitimate concerns.","tokens_in":21574,"tokens_out":3714,"duration_ms":39969,"concrete_test":"Re-run the RQ3 analysis using only actual confabulation categories that have a questionnaire counterpart (SPE/RFA/ORI, ESE, CMD; exclude FSC, AEF, MSE). Recompute scene-level medians, Wilcoxon tests, and Spearman correlations between perceived and restricted actual counts. If the office/city correlations remain near zero and the gap remains significant, the saturation claim survives. If correlations become significant or the gap shrinks materially, the central claim must be qualified as partly a measurement artifact rather than a pure user-perception phenomenon.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that 'user vigilance is structurally unreliable as a detection mechanism' rests on the RQ3 comparison between actual and perceived confabulation counts. The actual counts (Section 4.1) include all nine taxonomy categories, but the perceived counts come from a post-experience questionnaire with predefined categories that, by the authors' own admission (footnote 3), do not include FSC, AEF, or MSE. Consequently, any confabulation of those types is invisible in the perceived-count instrument by construction. This is a systematic asymmetry, not random noise: for example, FSC-DONE (n=28) is a false 'already done' claim that is often salient, yet it has no questionnaire category to report. In the city scene, FSC accounts for 8% of confabulation turns; excluding it from perceived counts artificially widens the actual-vs-perceived gap. The saturation interpretation depends on perceived counts being a valid measure of user awareness. If the questionnaire simply could not register some perceived failures, then the non-significant Spearman correlations in office (rho=0.15, p=.47) and city (rho=0.30, p=.16) may partly reflect a measurement ceiling, not a perceptual ceiling. The paper acknowledges this 'structural blind spot' in a footnote but still generalizes to system-side mitigation as the only reliable safeguard. That conclusion may be correct, but the current data do not cleanly establish it because the perceived instrument cannot detect what it does not ask about.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents an exploratory user study (N=24) of LLM-assisted immersive 3D scene editing, using a JSON-in/JSON-out GPT-5.2 pipeline in three VR scenes (office, loft, city) with distinct task types. Based on manual review of 1,612 interaction turns, the authors construct a nine-type taxonomy of confabulations, report their prevalence and perceived disruptiveness, and introduce the construct of a perception–reality gap by comparing log-based actual confabulation counts with post-experience questionnaire reports. The central claim is that under high confabulation load, user awareness saturates and becomes uncorrelated with actual counts, so human oversight is structurally unreliable and mitigation must be system-side. The paper also derives four design implications (DI1–DI4) and acknowledges its own limitations, including a footnote admitting a structural blind spot in the questionnaire.","tokens_in":21819,"tokens_out":4693,"duration_ms":49144,"significance":"If the empirical foundation holds, the paper makes a timely contribution: it is one of the few studies to characterize LLM confabulations in immersive scene editing from real interaction data, and it offers a concrete taxonomy plus a measurable construct (perception–reality gap) that can be reused by other XR/LLM researchers. The design implications are actionable and grounded in observed failure modes. The authors are also transparent about the questionnaire mismatch and about the limited power for type-level analyses. However, the two load-bearing pillars—the manual coding of ground-truth confabulations and the comparability of perceived versus actual counts—are not yet sufficiently validated. The central conclusion about saturation therefore requires additional evidence before it can be accepted at face value.","major_comments":[{"comment":"The ground-truth actual confabulation counts rest entirely on manual coding by the authors: 'Each interaction turn was manually reviewed ... to identify whether confabulation was present,' and labels were assigned via inductive open coding. No inter-rater reliability, second coder, codebook, or shared annotation dataset is reported. Since every prevalence percentage, every actual-count median, and every perception–reality statistic is computed from this coding, a systematically lenient or strict coder would change the headline results. The authors should report at least a reliability check (e.g., Cohen's kappa on a random subset, dual coding of a sample) and make the coding protocol or annotated data available. This is a necessary validity check for the central claim, not a stylistic suggestion.","section":"§4.1"},{"comment":"The perceived-confabulation instrument is systematically incomplete relative to the taxonomy. Footnote 3 states that FSC, AEF, and MSE 'did not correspond to questionnaire items, which resulted in a structural blind spot.' Yet actual counts include those categories: FSC accounts for 8% of confabulation turns in the city scene, FSC+AEF account for 10% in the office scene, and MSE accounts for 25% in the loft scene. Because the perceived-count measure cannot register these types by construction, the observed underestimation and the non-significant correlations in office/city may be inflated by instrument asymmetry rather than by a genuine perceptual ceiling. The paper acknowledges this in a footnote but still uses the result to claim that 'user vigilance is structurally unreliable as a detection mechanism' (Section 5). The authors should provide a sensitivity analysis that excludes or sepa","section":"§4.3, footnote 3"},{"comment":"The statistical evidence for 'saturation' and 'unrelated' counts is weaker than the prose suggests. The office (ρ=0.15, p=.47) and city (ρ=0.30, p=.16) Spearman correlations are non-significant, but N=24 gives low power, and ρ=0.30 is a moderate effect; non-significance does not establish that perceived and actual counts are unrelated. Moreover, no formal test of saturation is presented—the claim rests on comparing medians and correlations across scenes, with no model of a plateau or a nonlinearity. The authors should either fit a saturation model, test for a difference in correlation magnitudes across high/low-load conditions, or soften the 'structurally unreliable' conclusion to 'not reliably correlated in this sample.'","section":"§4.3"},{"comment":"The operationalization of the perceived confabulation count is underspecified. The post-experience questionnaire asked about 'the type, frequency level of disruption, and a short description of the confabulation,' but it is not stated exactly how these answers were converted into a numeric perceived count, how multi-type reports were handled, or whether participants were forced to choose from predefined categories. Because the perception–reality gap is the paper's central construct, the exact instrument wording, the mapping rules from questionnaire categories to taxonomy types, and the aggregation rule should be reported in enough detail for replication.","section":"§3.4 / §4.3"}],"minor_comments":[{"comment":"Typo: 'χ2 =32,8' should presumably be 'χ2 =32.8'.","section":"§4.2"},{"comment":"The disruption question uses the term 'hallucination' ('When a hallucination occurred, it disrupted my task'), while the paper deliberately defines and uses 'confabulation.' Please make the terminology consistent, or explain why the questionnaire used a different term.","section":"§3.4 / §4.2"},{"comment":"The legend for hatched cells combines two different conditions ('category absent in this scene' and 'questionnaire type could not be mapped to a taxonomy type'). These should be separated visually or in the caption, because they carry different implications for interpretability.","section":"Figure 5"},{"comment":"Table 1 notes that counts do not sum to totals because interaction turns can receive multiple labels. It would help to also report the number of turns with only one label versus multiple labels, since multi-label coding affects how the prevalence percentages should be read.","section":"§4.1"},{"comment":"The Limitations section mentions the small sample and single-system design, but does not mention the absence of inter-rater reliability or the questionnaire blind spot. Adding both to the limitations would help calibrate reader expectations.","section":"§5.2"}],"recommendation":"major_revision","confidential_remarks":"I see this as a potentially valuable empirical contribution rather than a rejected one. The core difficulty is that two fixable validity issues sit directly under the main claim: the unvalidated manual coding and the incomplete perceived-count instrument. If the authors provide a reliability check, a sensitivity analysis that handles unmatched categories, and a more cautious statistical interpretation, the paper could become a solid venue contribution. I would not require a new user study as a condition, but the sensitivity analysis must be substantive, not a one-sentence caveat."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"This is genuinely useful empirical work, but the strongest version of the central claim — that user vigilance saturates and is structurally unreliable — is weaker than the paper's own framing. The taxonomy and the perception-reality gap idea are worth keeping; the saturation conclusion should be softened or better supported.\n\nWhat's new: a grounded taxonomy of LLM confabulations in immersive 3D scene editing, built from 1,612 logged interaction turns across 24 users, with prevalence by scene, disruption ratings, trust/load comparisons, and detection/false-alarm rates. That is real work. The nine-type taxonomy is plausible and task-dependent — city scenes dominated by spatial failures, loft scenes by semantic ones — and it gives the field a usable shared vocabulary. Credit is also due for the honest limitations section and for explicitly flagging the structural blind spot in the perceived-confabulation questionnaire.\n\nSoft spots, in proportion. First, the 'actual' confabulation counts are the authors' own manual coding: no inter-rater reliability, no codebook, no released annotations. Every perception-reality statistic is anchored on those counts, so the validity of the reference measure is an open question. That is addressable, but it is load-bearing. Second, the stress-test note is correct. The perceived counts came from a questionnaire whose predefined categories omit FSC, AEF, and MSE, so those failure types cannot be reported by construction. FSC alone is 8% of confabulation turns in the city scene, and some FSC-DONE instances are salient. The actual-vs-perceived gap is therefore inflated by instrument mismatch, and the non-significant Spearman correlations may reflect a measurement ceiling rather than purely a perceptual ceiling. The authors acknowledge this in a footnote, but then still generalize to system-side mitigation as the only reliable safeguard. That conclusion may well be right, but these data do not cleanly establish it. Third, one proprietary model and one implementation limits generalization; the paper admits this. Minor: type-level analyses are underpowered, and cell-level detection rates are only descriptive.\n\nWould the central finding survive? The overall underestimation is probably real — the median gaps are large. But the saturation/uncorrelated claim is not cleanly supported as stated. A revised version should either use a complete questionnaire, or restrict the comparison to categories present in both instruments.\n\nRecommendation: send to peer review, not desk reject. The taxonomy and construct are solid enough to deserve referee time. Major revision should require inter-rater reliability, a released annotation set or at least a full codebook, and a re-analysis that does not compare apples to oranges.","headline":"Useful empirical taxonomy and a real construct, but the headline claim about saturated user vigilance is partly built on a mismatched questionnaire, and the 'actual' ground truth needs reliability evidence.","tokens_in":22369,"tokens_out":1927,"would_cite":true,"duration_ms":22110,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"In LLM-assisted VR scene editing, users systematically underestimate confabulations, and under high error load their awareness saturates so that perceived error counts become uncorrelated with actual ones—meaning human oversight cannot serv","keywords":["confabulation","hallucination","LLM","virtual reality","3D scene editing","perception-reality gap","human-AI collaboration","error detection"],"falsifier":"Re-code the 1,612 interaction turns with two independent coders following a pre-registered protocol, or automatically detect confabulations by mechanical checks (e.g., geometric overlap detection, out-of-region placement, comparison of claimed vs. applied edits). If a second coding pass yields materially different counts—or if the automatic check argues with the manual labels—the saturation and uncorrelated-counts results could vanish. A simpler check: a direct replication with twice the participants plus instrumented ground-truth logging would show whether perceived counts again plateau near","tokens_in":21378,"feed_emoji":"🥽","tokens_out":5708,"duration_ms":48066,"temperature":0.7,"pith_summary":"This paper argues that when people edit 3D scenes in virtual reality with an LLM-based assistant, the system frequently produces confabulations—plausible but incorrect edits—and, crucially, that users cannot be relied on to notice them. In a study of 24 non-expert users across three VR scenes, 27 percent of interaction turns contained at least one confabulation, and users consistently under-reported them: in the highest-error scene, the median perceived count was 4.0 while the median actual count was 10.5, and individual perceived counts were not significantly correlated with actual counts. The paper concludes that confabulation awareness saturates under load, making human vigilance structurally unreliable as a detection mechanism and putting the burden of mitigation on the system. It also contributes a nine-type taxonomy of confabulations and a construct, the perception-reality gap, for measuring this divergence.","feed_headline":"AI editing errors outpace user awareness in VR test","feed_subtitle":"Perceived and actual error counts become uncorrelated; human oversight alone can't catch AI mistakes.","key_machinery":"The load-bearing construct is the perception-reality gap: the discrepancy between the number and types of confabulations users believe they encountered (measured by post-experience questionnaire) and those that actually occurred (measured by manual log coding). The companion mechanism is saturation: as confabulation load rises, perceived counts plateau and lose correlation with actual counts, rendering human detection ineffective. The taxonomy of nine confabulation types (spatial placement errors, reference frame ambiguity, orientation errors, command misinterpretation without disclosure, false scene claims, edit scope errors, material/style errors, missing information without disclosure, an","core_discovery":"The paper's central discovery is an empirical demonstration that user awareness of LLM errors in immersive editing is bounded: in low-error conditions people track confabulations well (Spearman rho=0.71), but in high-error conditions perceived counts stop tracking actual counts (office rho=0.15; city rho=0.30, n.s.), a pattern the authors call saturation. They arrive at this by building a JSON-in/JSON-out LLM scene-editing workflow, logging 1,612 interaction turns from 24 users, manually coding every turn for nine confabulation types and 24 subtypes (437 confabulation turns total), and comparing these actual counts with users' post-experience questionnaire reports. The discovered asymmetry i","pith_inferences":["The saturation finding likely generalizes beyond VR to any human-AI interface where error rates are high, implying that user vigilance can be modeled as a limited-capacity channel; a quantitative model of that capacity could predict when monitoring will fail.","The 'actual' counts could be obtained automatically in future work—using geometric collision checks and out-of-region tests to flag confabulations by rule—which would let the perception-reality gap be studied at scale without manual labeling and would test whether the manual coding explains the results.","The taxonomy, being task-derived rather than model-derived, may transfer to other structured-data LLM agents (scene graphs, CAD, code) where silent misinterpretation and false completion claims are likely to appear in similar proportions.","The 46 percent false-alarm rate for object omission in the city scene, driven by false 'already done' claims, hints that users infer failure from system denial; studying denial-based inference could turn misleading output into a diagnostic signal."],"forward_implications":["System-side verification becomes the primary safeguard: scene state should be re-checked against the LLM's prior claims before each new interaction round, rather than trusting either the model's text or the user's memory.","Confident textual outputs from the LLM should not be treated as evidence that an edit succeeded; independent structural validation is needed for claimed deletions, undos, and placements.","Mitigation can exploit the perceptual asymmetry: spatial errors are self-revealing in VR, so effort should concentrate on semantic interpretation errors and silent failures, which users rarely notice.","Task framing and the presence of a reference state shape error tolerance: stylistic tasks without a clear reference induced higher trust and lower perceived disruption, suggesting task design can be a mitigation lever.","Out-of-vocabulary internal tokens (like the '_st' suffix) should be flagged as unknown rather than silently mapped to a plausible meaning, since such silent misreadings dominated the stylistic task."],"fun_headline_variants":["VR users can't keep track of AI errors in busy scenes","Under AI error load, VR users' awareness saturates","When AI confabulates in VR, users often don't notice","Perception-reality gap: VR users undercount AI errors","High error rates in VR editing blind users to AI mistakes"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The reference counts of 'actual' confabulations come from a single researcher's manual coding of all 1,612 interaction turns, with no second coder, no inter-rater reliability, and no formal protocol beyond inductive open coding, so every saturation and perception-reality result depends on that labeling being correct.","fun_headline_variants_meta":{"raw":{"variants":["VR users can't keep track of AI errors in busy scenes","Under AI error load, VR users' awareness saturates","When AI confabulates in VR, users often don't notice","Perception-reality gap: VR users undercount AI errors","High error rates in VR editing blind users to AI mistakes"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001134,"raw_usage":{"total_tokens":4533,"prompt_tokens":718,"completion_tokens":3815,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":462,"completion_tokens_details":{"reasoning_tokens":3743}},"tokens_in":462,"tokens_out":3815,"duration_ms":23311,"temperature":1.0,"reasoning_tokens":3743,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-01T03:15:34.686756+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-code the 1,612 interaction turns with two independent coders following a pre-registered protocol, or automatically detect confabulations by mechanical checks (e.g., geometric overlap detection, out-of-region placement, comparison of claimed vs. applied edits). If a second coding pass yields materially different counts—or if the automatic check argues with the manual labels—the saturation and uncorrelated-counts results could vanish. A simpler check: a direct replication with twice the participants plus instrumented ground-truth logging would show whether perceived counts again plateau near","supporting_citations":[],"review_version":1}