{"id":"6305e31a-68ff-41f1-8dea-21740fbfabae","arxiv_id":"2607.21769","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"In a 16-participant study, adding visible objects and grasp interactions to an LLM-driven VR communication trainer did not hurt usability or workload, and the most environment-grounded condition was preferred by autistic trainees and job coaches.","lead":"This paper tested three versions of a VR job-communication trainer for autistic adults: chatting with an AI customer, chatting with objects visible, and chatting while also picking up and placing the objects. Across 9 trainees and 7 job coaches, all versions felt equally usable and easy; trainees and coaches liked the object-grasping version most and said it supported engagement and realistic multi-tasking.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Agent behavior is confounded with condition: the prompt schema changes the LLM avatar's responses, so preference for C+O+G could reflect prompt effects rather than embodied grounding.","rationale":"The reader's weakest_assumption identifies exactly this confound: 'if richer conditions produced systematically different agent behavior (e.g., lengthier or less accessible questions), the central comparison would be confounded.' I considered alternative concerns—small single-site sample, single-item preference measure, missing appendices, and descriptive-only coach ratings. These affect generalizability and statistical strength but do not shake the internal validity of the condition comparison as directly as the agent-behavior confound. The manuscript's own description of the prompt schema (§3.1) makes it clear that the LLM receives different inputs across conditions, and the qualitative section (§5.2.2) explicitly notes that the avatar's responses were constrained by the object schema. Because the LLM is stochastic, condition effects on the avatar's dialogue are plausible and unmeasured. The paper does provide useful independent support: within-subjects counterbalancing, recorded session logs, and null SUS/TLX results that are less sensitive to prompt differences. But the preference outcome—the central claim—remains vulnerable. This is a normal confound in LLM-based HCI studies, not a sign of misconduct. The proposed test is feasible because logs are already stored. If the check shows matched behavior, the reading is upheld; if not, the abstract and conclusion overstate the role of environmental grounding. Since the reader's verdict is already CONDITIONAL and flags the same confound, my recommendation is UNCHANGED: conditional acceptance with a request that the authors either report the transcript-level analysis or soften the causal wording in the abstract and conclusions.","tokens_in":21983,"tokens_out":5468,"duration_ms":57304,"concrete_test":"Use the recorded session logs (§4.5) to compute, per condition, automated agent-response features: utterance length in words, number of questions per turn, fraction of multi-part (compound) questions, number of object/grasp references, and response latency. Compare C, C+O, C+O+G with repeated-measures Friedman tests. If C+O+G differs significantly from C on any feature, the preference finding is plausibly confounded by agent behavior; if features are matched across conditions, the grounding explanation is supported. A complementary human-coding check: two coders blind to condition rate a random sample of agent turns for engagement and complexity; compare distributions by condition.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central causal claim—that C+O+G was preferred while maintaining comparable usability/workload—implicitly requires that the three conditions differ only in environmental grounding. The manuscript does not establish this. §3.1 describes an adapted prompt schema in which C receives no object/action context, while C+O and C+O+G inject an Objects component and, in C+O+G, grasp/release cues. The LLM agent's responses are therefore generated from different prompts across conditions. §5.2.2 acknowledges that agent responses in C+O/C+O+G were 'constrained by the virtual objects defined in the schema,' and §5.3.2 reports the agent sometimes asked multi-part questions that overwhelmed trainees—but no transcript-level check is reported. If the avatar's turns in C+O+G were systematically longer, more object-referencing, or differently structured, the observed preference (trainee Friedman p=.050; 6/7 coaches choosing C+O+G) could be driven by prompt-engineering effects rather than by the embodied interaction modality. This is an unmeasured confound in the independent variable, and it is load-bearing because the headline advantage of C+O+G rests on exactly the kind of effect that differences in agent behavior could produce.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents an exploratory within-subjects study of three VR-based communication training modalities for autistic individuals: conversation-only (C), conversation with environmental objects (C+O), and conversation with objects plus grasp interactions (C+O+G). Nine autistic trainees and seven job coaches used the system; the authors measured System Usability Scale (SUS), raw NASA-TLX, preference ratings, and collected qualitative observations and interviews. The central claim is that usability and workload were comparable across the three modalities, while both trainees and job coaches preferred the most environment-grounded condition (C+O+G), which was perceived as more engaging and as better integrating communication practice into task performance. The paper also offers design considerations for future environment-grounded VR communication training.","tokens_in":22228,"tokens_out":4311,"duration_ms":48474,"significance":"If the finding holds, the paper makes a useful contribution to accessible VR communication training by showing that adding environmental objects and grasp interactions does not necessarily increase workload or reduce usability for autistic trainees, and that such features may be preferred. The study is unusual in including both autistic trainees and job coaches, and the qualitative data are rich. However, the causal interpretation is undermined by a confound: the three conditions differ not only in the user's interaction modality but also in the prompt schema given to the LLM-driven agent, so the agent's dialogue is generated from different contextual inputs across conditions. This unaddressed confound weakens the central preference claim and needs to be resolved before the paper can be accepted.","major_comments":[{"comment":"The independent variable is not limited to the trainee's environmental grounding; it also changes the agent's prompt context. Section 4.1.1 states that the C condition uses a 'basic system context prompt' while C+O and C+O+G receive the full prompting schema, with C+O+G additionally including grasp/release cues. Consequently, the LLM-generated avatar responses differ across conditions by construction. Section 5.2.2 acknowledges that the avatar's responses in C+O and C+O+G were 'constrained by the virtual objects defined in the schema,' and Section 5.3.2 reports that the agent sometimes asked multi-part questions that overwhelmed trainees. No transcript-level analysis of agent behavior (e.g., turn length, number of questions per turn, object reference frequency, response delays) is reported. The preference for C+O+G (Section 5.3.4, Figure 4c) could therefore be driven by prompt-driven dif","section":"§3.1, §4.1.1, §5.2.2, §5.3.2"},{"comment":"The quantitative support for the preference claim is marginal. The Friedman test for overall preference yields χ²(2)=6.00, p=.050, Kendall's W=.33, and the authors appropriately avoid pairwise comparisons. Yet the abstract and conclusion state unqualifiedly that 'both trainees and coaches preferred' C+O+G. This rests on descriptive means (4.89 vs. 4.67 vs. 4.44) and on qualitative reports (6/7 coaches, all trainees). With n=9 and a borderline p-value, the claim should be more carefully qualified, or supported by a pre-planned comparison or effect-size measures, so that readers can judge the strength of the evidence.","section":"§5.1.1, Abstract, §7"}],"minor_comments":[{"comment":"Please clarify how ChatGPT was used to 'assist with organizing the observation notes' and how the first author's coding was verified against the original notes. This transparency is valuable in qualitative work, especially given the importance of establishing that the AI tool did not influence theme development.","section":"§5.2"},{"comment":"Because J01 and J06 each supported two trainees, coach ratings are nested (9 sessions but only 7 coaches). The descriptive analysis is acceptable, but the non-independence should be acknowledged when presenting coach ratings to avoid implying fully independent observations.","section":"§4.3, §5.1.2"},{"comment":"Minor notation inconsistencies: 'Raw NASA-TLX' versus 'raw NASA-TLX'; 'χ²(2)' is used but the degree-of-freedom formatting varies; 'Kendall's W' is sometimes typeset as W=.03, .07, .33 with inconsistent punctuation. A consistent notation style would improve readability.","section":"Throughout"},{"comment":"The scenarios differ in object placement (all front, front/back split, all back). Although scenario-to-condition assignment was randomized, with only 9 participants the randomization may not balance scenario difficulty across conditions. Please report the actual scenario-condition mapping or discuss this as a potential carryover/confound.","section":"§4.1.2"}],"recommendation":"major_revision","confidential_remarks":"The paper is a reasonable exploratory study of an important topic, but the central preference claim is confounded by the simultaneous change in the LLM prompt schema across conditions. This is not merely a stylistic issue; it directly affects the interpretation of the headline result. I would like the authors to add an analysis of the agent's generated behavior by condition or otherwise demonstrate that the observed preference is not an artifact of prompt-induced differences in the avatar's responses. The small sample and the marginal p-value further suggest that the language in the abstract and conclusion should be softened until the evidence is stronger."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick read: this is a modest, carefully reported exploratory study, and it stays on the right side of its evidence until the abstract. The genuinely new piece is the within-subject comparison of three levels of environment grounding (C, C+O, C+O+G) using an LLM-driven VR role-play for autistic trainees and their job coaches. No prior work I know of has that three-condition comparison. The study is honestly run: counterbalanced order, randomized scenario-to-condition pairing, appropriate Friedman tests, and the authors decline to decompose the marginal preference effect. Credit also for reporting the null SUS/TLX result without spin, and for including coach perspectives, which are usually missing from this literature.\n\nThe main soft spot is exactly what the stress-test flagged: the independent variable is not pure modality. The prompt schema differs across conditions, so the LLM avatar's behavior changes with the condition. Section 5.2.2 even says responses in C+O/C+O+G were constrained by the objects defined in the schema. That means the C+O+G preference could be driven by richer agent dialogue rather than by the act of grasping. The paper acknowledges this in places but never treats it as a confound. A transcript-level check of agent turn length, question complexity, or object references across conditions would have helped. Without that, the preference finding stays at hypothesis-generating, not demonstrated.\n\nOther soft spots are less severe: 9 trainees and 7 coaches from one site; thematic analysis with a single coder plus AI assistance, no second coder; and the referenced appendices (full prompts, coach rating items) are missing from the arXiv version, so you cannot fully audit the manipulation. The abstract also says trainees and coaches 'preferred C+O+G' when trainee preference reaches only p=.050 and is carried mostly by interview quotes.\n\nNone of this is fatal. The paper's scaled claims — no usability penalty, comparable workload, qualitative preference themes — hold up as exploratory. The null workload finding is arguably the most useful result, since it suggests adding grasp interaction does not tax this population in a noticeable way.\n\nWho is this for? Anyone designing LLM-driven VR training, especially for autistic users or vocational contexts. It would make for a good reading group discussion. I would send it to peer review, not desk reject. The authors should be asked to share transcripts and prompts, and to address the prompt-behavior confound in revision. My own verdict: conditionally accept with revisions.","headline":"Honest exploratory study with a real confound between prompt schema and modality; worth refereeing but the preference claim should be softened.","tokens_in":22711,"tokens_out":1973,"would_cite":true,"duration_ms":24088,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"An exploratory study with 9 autistic trainees and 7 job coaches finds that adding visible, graspable objects to an LLM-driven VR role-play scenario yields the preferred training mode, with usability and workload unchanged from conversation-","keywords":["communication training","autism spectrum disorder","virtual reality","large language models","workplace simulation","environment grounding","grasp interaction"],"falsifier":"Take the exact system prompts from the three conditions and run the same scripted dialogue with an empty object list in C. Measure the avatar's average number of questions per turn and per-turn length; if the object list appears anywhere in C's system context, or if C+O+G's turns are systematically different in structure or pace from C's, the central comparison is confounded and the preference could be explained by dialogue style rather than grounding.","tokens_in":21840,"feed_emoji":"🥽","tokens_out":9916,"duration_ms":94631,"temperature":0.7,"pith_summary":"This paper tries to establish that job-communication training for autistic individuals can be improved by embedding conversation in the trainee's physical task environment rather than keeping it as isolated verbal role-play. The authors built an LLM-driven virtual customer whose dialogue is conditioned on the VR scene's objects and on the trainee's grasp actions, and compared three levels of grounding — chat-only, chat with visible objects, and chat with objects plus grasping — across café, butcher-shop, and fast-food scenarios. In a study with 9 autistic trainees and 7 job coaches, usability and workload were statistically indistinguishable across the three conditions, yet both groups preferred the most grounded condition (C+O+G), citing sustained engagement, more realistic task-embedded conversation, and the chance to practise multi-tasking. The authors present this as exploratory evidence for flexible, adjustable grounding in VR communication training rather than a one-size-fits-all modality.","feed_headline":"Hands-on VR wins as favorite training mode for autistic users","feed_subtitle":"Hands-on VR practice kept workload steady and boosted engagement for both groups","key_machinery":"The carrying mechanism is a prompting structure that feeds the LLM-driven avatar three layers of context: the scenario and role definition, the catalog of environmental objects and containers, and the user's action cues describing which hand grabs or releases which object from which parent surface. This turns a physical grasp into a text event, letting the avatar produce dialogue that references the shared spatial situation and the work in progress. The three conditions are incremental unlocks of this schema: the conversation-only baseline omits the object list and action cues, C+O adds the object catalog but disables grasping, and C+O+G adds the grasp cues, enabling the avatar to respond to","core_discovery":"The paper's central claim is that environment-grounded VR communication training can raise engagement and perceived workplace relevance for autistic trainees without measurable usability or workload costs. Concretely, when the virtual customer's prompt includes the list of objects in the room and the trainee's grasp/release actions (rendered as descriptive cues such as 'left hand grabs Coffee from Counter'), both trainees and job coaches preferred that condition over conversation-only practice and over a condition with objects but no interaction. The preference was accompanied by qualitative reports that physical interaction gave trainees conversational anchors, encouraged longer and more ob","pith_inferences":["If this preference generalizes beyond the small sample, VR vocational training for autistic individuals may shift from abstract social-skills rehearsal toward task-embedded conversation; the paper gestures at this (calling the hands-on activity a possible anchor) but does not test it directly.","A next experiment could separate task-relevant grounding from mere motor activity, comparing C+O+G with a condition where the physical task is unrelated to the conversation, to see which component drives engagement.","The study stops at experience and preference; the key practical promise — that skills practiced in the grounded condition transfer to real workplaces — remains untested and is a natural target for a longitudinal follow-up.","The coaches' request for per-session flexibility implies an authoring tool that lets them toggle object visibility and grasp interactivity per trainee; the paper lists this as future work, so it is our inference that such a tool is a direct design consequence."],"forward_implications":["Environment-grounded VR role-play can embed communication practice inside task performance, matching the situated nature of real workplace conversation that job coaches emphasized.","Adding visible, graspable objects did not degrade SUS or raise NASA-TLX scores for this sample, suggesting richer interaction need not trade off against usability or cognitive load.","The preference for the C+O+G condition was tied to reports of sustained engagement and more object-referenced talk, implying that grounding can serve as a conversational scaffold for trainees who find open-ended small talk effortful.","Because participants valued different modalities for different trainees and goals, training systems should offer configurable grounding levels rather than a single fixed mode.","The avatar's occasional multi-question turns overwhelmed some trainees, so prompts should constrain the agent to one request or question per turn even when grounding enriches context."],"fun_headline_variants":["Grasping in VR boosts engagement for autistic talk training","Hands-on VR conversation preferred by autistic users and coaches","VR with grasp interactions keeps workload steady, raises engagement","Grasp-based VR talk training more engaging, same workload for autistic users"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The load-bearing premise is that the three conditions differ only in environmental grounding: the conversation-only baseline receives no object or grasp information, and the richer conditions do not change how the avatar talks (response length, number of questions per turn, or pacing). If the prompt schema leaks object context into the baseline, or if the grounded conditions happen to produce friendlier, slower, or simpler dialogue, the observed preference could come from con","fun_headline_variants_meta":{"raw":{"variants":["Grasping in VR boosts engagement for autistic talk training","Hands-on VR conversation preferred by autistic users and coaches","VR with grasp interactions keeps workload steady, raises engagement","Grasp-based VR talk training more engaging, same workload for autistic users"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000201,"raw_usage":{"total_tokens":1180,"prompt_tokens":676,"completion_tokens":504,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":420,"completion_tokens_details":{"reasoning_tokens":434}},"tokens_in":420,"tokens_out":504,"duration_ms":5834,"temperature":1.0,"reasoning_tokens":434,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-01T06:43:41.545636+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take the exact system prompts from the three conditions and run the same scripted dialogue with an empty object list in C. Measure the avatar's average number of questions per turn and per-turn length; if the object list appears anywhere in C's system context, or if C+O+G's turns are systematically different in structure or pace from C's, the central comparison is confounded and the preference could be explained by dialogue style rather than grounding.","supporting_citations":[],"review_version":1}