{"id":"25f9aa50-edfe-4ae4-b244-8e6d892ed41a","arxiv_id":"2508.17124","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"Untrained VR users adapt their multimodal instruction strategies to the spatial clarity of the task, using explicit speech with concrete anchors and implicit speech with prolonged pointing in ambiguous settings.","lead":"A virtual reality study asked untrained users to instruct a simulated robot arm during LEGO and circuit-board assembly while voice, hand gestures, and gaze were recorded. Results show people use more descriptive speech when locations are visually clear and more vague \"put-that-there\" phrasing with longer pointing when locations are hard to describe, and the authors release a dataset of the sessions.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The central causal claim — that spatial anchor availability drives put-that-there language — is underdetermined because the two tasks vary in object type, piece count, board geometry, and collaborative versus instructive mode, with no condition isolating anchor availability.","rationale":"The reader's weakest assumption matches my own. What would have to be true for the central claim is that the LEGO versus PCB contrast is a clean manipulation of spatial anchor availability. It is not. The paper is transparent about the task designs in Section 3.1, but transparency does not remove the confound; Sections 6.2 and 6.3 use the comparison as evidence for an anchor-based mechanism. The observed effects are large, and the dataset is a contribution even if the causal attribution is not established. The concern is not that the results are fabricated or that the descriptive claims are doubtful; it is that the explanatory claim in the abstract and discussion goes beyond what the design can support. A third condition holding all other factors fixed is the natural check. If that condition confirms the anchor effect, the paper's design implications are substantially strengthened; if not, the paper should be reframed as reporting descriptive task differences rather than anchor-driven strategy selection. I therefore keep the reader's CONDITIONAL verdict unchanged rather than moving to accept or reject, because the paper has real empirical value but its central causal interpretation is not yet secured.","tokens_in":20412,"tokens_out":3512,"duration_ms":38726,"concrete_test":"Run a third condition that presents the PCB task with color-coded fiducial markers or pre-placed reference anchors at the target locations, keeping component type, count, board geometry, instruction-only mode, and reference display identical to the existing PCB condition. Compare explicit/implicit utterance ratios and average point durations against the blank-PCB condition and, ideally, against the LEGO condition. If anchor availability changes these measures in the predicted direction, the causal story gains support; if the blank-PCB and anchored-PCB conditions remain similar, the observed task differences are better attributed to task type or other confounds.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The abstract's strongest claim is that users adopt put-that-there language \"in spatially ambiguous contexts\" and descriptive instructions \"in spatially clear ones.\" The evidence for this is a two-condition comparison: LEGO (collaborative, concrete spatial anchors) versus PCB (instructive, ambiguous locations), described in Sections 3.1 and 3.3. The paper then makes causal statements in Section 6.2 (\"due to concrete spatial anchors... due to a lack of spatial anchors\") and Section 6.3. But the conditions differ along many dimensions simultaneously: the user physically manipulates objects only in Task 1; the object sets, numbers of pieces (25 vs 20), board geometry, and reference layouts differ; the instruction mode differs by design; and task completion time differs (M=382.84 vs 451.40 s). The within-subjects design and counterbalancing control for individual differences, but they do not isolate the availability of spatial anchors. The key dependent measures (explicit/implicit ratio, average point duration, unimodal vs trimodal usage) could plausibly track task nature, object familiarity, workload, or instruction complexity instead. Because the design lacks a condition in which anchor availability is varied while other task properties are held fixed, the central causal attribution is underdetermined, even though the descriptive differences are large and the dataset is a useful resource.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents a Wizard-of-Oz elicitation study (N=34) in which participants instruct a virtual robot arm through a collaborative LEGO assembly task and an instructive PCB assembly task in VR. Voice, hand tracking, eye gaze, and head pose are captured and analyzed. The main empirical claims are that explicit/descriptive utterances dominate the LEGO task while implicit \"put-that-there\" language dominates the PCB task, that pointing durations are longer in the PCB task, and that users produce more trimodal commands in the PCB task and more unimodal commands in the LEGO task. The paper also reports correlations between utterance frequency and task completion time and contributes a publicly available annotated dataset.","tokens_in":20657,"tokens_out":6400,"duration_ms":64178,"significance":"If the descriptive findings are reliable, the dataset is a useful resource for future natural multimodal interface design, and the large effects on utterance ratio (t=8.14), pointing time, and trimodal usage are noteworthy. The study is unusual in capturing untrained users' raw multimodal input with a Wizard-of-Oz robot, and the public dataset is a concrete contribution. However, the paper's headline causal interpretation—that spatial-anchor availability causes the observed differences—is not supported by the operationalization, and the central utterance-ratio measure lacks demonstrated coding reliability. The work would be publishable as a descriptive, dataset-focused study once these issues are addressed.","major_comments":[{"comment":"The study is framed as a single-factor comparison of \"task nature,\" but the two conditions differ on many dimensions at once: the participant manually manipulates objects only in Task 1; the object sets, geometries, and board layouts differ; the piece counts are 25 versus 20 (§3.6); and the observed completion times differ as well (M=382.84 s versus 451.40 s, Table 3). The abstract and §6.2/§6.3 nevertheless attribute the differences to the availability of concrete spatial anchors (\"due to concrete spatial anchors... due to a lack of spatial anchors\"). Because no condition manipulates anchor availability while holding the other task properties fixed, the causal attribution in the abstract is underdetermined. The authors should either soften the claims to descriptive differences between two assembly scenarios or add a follow-up condition that isolates anchor availability; the current Limitations section does not acknowledge this confound.","section":"§3.1, §6.2, abstract"},{"comment":"The headline effect—explicit versus implicit utterance ratios differing across tasks with t33 = 8.14—rests entirely on manual annotation of transcripts into nine labels, but the paper provides no information about the number of annotators, the annotation protocol, or inter-rater agreement. Without coding reliability evidence, the large ratio difference could reflect a single annotator's interpretation of the Bolt-based scheme rather than a stable behavioral difference. The authors should report inter-rater reliability (e.g., Cohen's kappa or equivalent) or provide a robustness check such as re-annotation of a subset.","section":"§4.1, §5.1.1"}],"minor_comments":[{"comment":"Statistical reporting needs alignment: §5.4.3 reports F3,30=5.296, R=0.530, R²=0.281 for the Task 1 multimodal correlation, whereas Table 2 reports F3,30=5.438, R=0.535, R²=0.287; please use consistent values and correct notation, since Wilcoxon results are reported with t statistics in places (e.g., §5.2.1 reports \"t33=37.0\").","section":"Tables 1–4 and §5.4.3"},{"comment":"The paper states that participants \"underwent training\" for task objectives and manipulation mechanics, which appears to conflict with the repeated claim that interactions come from untrained users; please clarify what was trained and why this does not undermine the \"untrained\" framing.","section":"§3.3"},{"comment":"The pointing-detection thresholds (z-rotation differences of 10° and 30°, minimum duration 0.5 s) are stated without justification; please cite prior work or report a sensitivity analysis showing the findings are robust to these choices.","section":"§4.2"},{"comment":"The task-sequence visualization is dense and the legend (\"G – Gaze, P – Pointing Gesture, U – Utterance\") does not explain how sequences are encoded; the figure would benefit from a concrete example sequence annotation.","section":"Figure 10"},{"comment":"Minor typographical issues (e.g., \"V oice ismost effective\" in §6.4.1 and \"cuessuch\" in §6.4.2) should be corrected.","section":"§6.4"}],"recommendation":"major_revision","confidential_remarks":"The paper is a reasonable dataset/observational contribution, but the causal framing requires substantial revision. I would be comfortable with major revision; if the authors are unwilling to remove the causal attribution to spatial anchors, I would recommend rejection. The dataset release is the strongest part of the submission."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a workmanlike elicitation study that delivers a genuinely new public dataset (voice, hand tracking, gaze, head pose from 34 untrained users doing LEGO and PCB assembly with a Wizard-of-Oz robot), and the descriptive effects are large and believable. Participants used explicit, descriptive language in the LEGO task and implicit \"put it here\" language with longer pointing in the PCB task. Those differences are robust enough (t=8.14 on explicit/implicit ratio; p<.001 on point duration) that they're probably real. The sequential pattern analysis and phase-wise trends are a reasonable extension of prior elicitation work.\n\nThe soft spot is the one the authors themselves stumble into: they attribute the behavior to \"concrete spatial anchors\" versus \"lack of spatial anchors\" in Section 6.2, but the two tasks differ on many dimensions at once — collaborative vs instructive mode, object type, piece count, board geometry, whether the user physically manipulates anything, even completion time. There is no condition that varies anchor availability while holding other task properties fixed. So the causal claim in the abstract (\"users tended to use put-that-there language in spatially ambiguous contexts\") is underdetermined. What they can legitimately claim is a descriptive difference between two task scenarios; that's still useful, but they should stop saying \"due to.\" A second issue: the explicit/implicit annotation is the load-bearing dependent variable, but there's no inter-rater reliability reported. With nine labels and hand-derived rules, that matters. Multiple comparisons are also uncorrected across the Wilcoxon and t-tests, though most of the headline effects are large enough that correction wouldn't kill them. The dataset link is a view-only OSF link, which undercuts the reproducibility promise until it's made fully public.\n\nOverall, the paper is honest about its scope, the dataset is the real contribution, and the design guidance in Section 6.4 is sensible and clearly grounded in the observed behavior, even if the causal framing goes beyond the evidence. It deserves peer review — a good referee can push for revised causal language, an IRR analysis, and a fully public dataset. I'd assign it conditional, not reject.","headline":"A useful new multimodal VR elicitation dataset with large behavioral differences between two tasks, but the central causal claim about spatial anchors is underdetermined by a confounded two-condition design.","tokens_in":21201,"tokens_out":2378,"would_cite":true,"duration_ms":23634,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"When an assembly workspace offers no nameable landmarks, untrained users fall back on 'put that there' phrasing, longer pointing, and richer multimodal combinations, and the paper argues natural user interfaces should be built to expect…","keywords":["virtual reality","natural user interfaces","multimodal interaction","Wizard-of-Oz","assembly tasks","put-that-there","spatial language","human-robot interaction"],"falsifier":"Run the same circuit-board assembly with visually anchored locations, such as outlined slots or labelled positions, while keeping the board layout and part count identical; if users still produce mostly implicit utterances and long pointing times, the claim that spatial ambiguity drives the behavior is wrong. Alternatively, within the released dataset, compare average pointing duration on implicit versus explicit utterances: if pointing is not longer for implicit commands, the implied linkage between vague language and prolonged deictic compensation fails.","tokens_in":20221,"feed_emoji":"🤖","tokens_out":8336,"duration_ms":81368,"temperature":0.7,"pith_summary":"This paper asks how people who have received no training on the interface naturally tell a virtual robot arm what to assemble, using voice, hand gestures, gaze, and head position. In a Wizard-of-Oz study with 34 participants, the authors compared a collaborative brick assembly, where recently placed pieces give concrete spatial anchors, with an instructive printed-circuit-board assembly, where placement locations are hard to name. Their analyses indicate that users adapt their instructions to the spatial clarity of the scene: descriptive, explicit commands dominate when anchors exist, while spatially vague 'put-that-there' phrasing paired with prolonged pointing and gaze dominates when locations are ambiguous. The same data show task type shifting the modality mix, with more unimodal commands in the anchored task and more trimodal commands in the ambiguous one, and the authors release an annotated multimodal dataset. If the pattern holds, interface designers could let users choose their own instruction style instead of forcing fixed command patterns.","feed_headline":"Untrained users adapt robot commands to spatial clarity in VR assembly","feed_subtitle":"In a Wizard-of-Oz VR study, instruction style tracked whether the scene had nameable locations.","key_machinery":"The load-bearing apparatus is a single-factor within-subjects elicitation design whose two conditions are meant to differ in one property: whether the workspace supplies concrete, nameable spatial anchors. The analytical core is an annotation scheme built on the put-that-there command paradigm, defined as a voice-plus-gesture pattern in which an utterance names an object and a gesture or gaze supplies the location. Each utterance is classified as explicit (descriptive) or implicit (spatially vague, relying on gestures or gaze), and a multimodal sequence miner assembles each voice command with any stationary gaze or pointing gesture occurring within five seconds. This machinery lets the authors compare not only utterance content but also the duration of points, the gaze direction, the number of modalities per instruction, and how these change over task completion phases.","core_discovery":"The paper's central discovery is that untrained users spontaneously follow a put-that-there instruction pattern, an utterance that names an object plus a deictic indication of where it goes, but the linguistic and gestural balance of that pattern shifts with the availability of spatial anchors. In the brick task, where pieces provide nameable reference points, participants used significantly more explicit descriptive commands, pointed for shorter periods, looked more toward the reference model, and leaned on unimodal speech. In the circuit-board task, where the blank board offers no obvious landmarks, participants used significantly more implicit commands such as 'put a resistor here' and compensated with significantly longer pointing, more gaze use, and more trimodal combinations; instruction frequency also predicted completion time far more strongly in this task. These differences were not static: command style evolved across task phases, with brick-task users gradually adding implicit commands while circuit-board users stayed implicit throughout.","pith_inferences":["A direct test that varies only anchor availability, keeping the same board layout and part count while adding or removing visual placeholders, would separate the spatial-anchor explanation from other task differences such as piece count and board geometry; the current study conflates those factors.","Because instruction frequency predicted completion time sharply in the ambiguous task, an interface that auto-completes repeated component types could cut instruction overhead more in layout-heavy tasks than in anchored ones.","Modality mix could serve as a real-time uncertainty signal: rising implicit utterances combined with longer pointing would flag regions the user finds hard to describe, letting an adaptive system highlight candidate placements."],"forward_implications":["Natural multimodal interfaces for assembly should accept put-that-there style instructions that pair vague voice with gesture and gaze instead of requiring users to follow preset command grammars.","In workspaces with concrete anchors, designers can expect explicit descriptive speech to suffice and can keep gesture interpretation lightweight, while in anchor-free workspaces the interface must tolerate long pointing episodes and combine eye and head gaze with speech.","Task phase matters: instruction style changes as an anchored assembly progresses, so session-level assumptions about modality use should give way to phase-adaptive interpretation.","The released annotated dataset, containing voice transcripts, hand tracking, eye gaze, head pose, and command labels, gives other researchers a basis for training models of natural multimodal assembly instruction."],"supporting_citations":[{"why":"Supplies the put-that-there voice-plus-gesture paradigm used to define the explicit/implicit command labels and to interpret user strategies.","marker":"[7]"},{"why":"Provides the multimodal speech-and-gesture design principles that frame the study's analysis of modality integration.","marker":"[50]"},{"why":"Supports the claim that users elaborate spatial context with other modalities, which the discussion uses to interpret anchored versus ambiguous behavior.","marker":"[47]"},{"why":"Reports how spatial context is elaborated with gestures while voice carries description, used to interpret the task differences observed.","marker":"[52]"},{"why":"Informs the research questions about how modalities are integrated in multimodal interaction.","marker":"[46]"},{"why":"Provides a prior natural speech, gesture, and demonstration dataset for robot learning that the elicitation and dataset design build on.","marker":"[63]"},{"why":"Supplies the speech-to-text model used to transcribe voice recordings into timestamped command text for annotation.","marker":"[54]"}],"fun_headline_variants":["Spatial clarity shapes how users command VR robots","In VR assembly, vague spaces trigger put-that-there commands","How spatial cues steer user commands in VR assembly","Users switch command style based on VR scene clarity","VR study: Clear scenes get explicit commands, fuzzy scenes get deictic"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The study assumes the two assembly tasks differ only in the factor being tested, whether the scene provides concrete spatial anchors, so that all observed differences in user behavior can be attributed to that factor, even though the tasks also differ in piece count, part types, board geometry, and interaction demands.","fun_headline_variants_meta":{"raw":{"variants":["Spatial clarity shapes how users command VR robots","In VR assembly, vague spaces trigger put-that-there commands","How spatial cues steer user commands in VR assembly","Users switch command style based on VR scene clarity","VR study: Clear scenes get explicit commands, fuzzy scenes get deictic"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000197,"raw_usage":{"total_tokens":1309,"prompt_tokens":838,"completion_tokens":471,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":454,"completion_tokens_details":{"reasoning_tokens":389}},"tokens_in":454,"tokens_out":471,"duration_ms":4620,"temperature":1.0,"reasoning_tokens":389,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T17:06:38.498964+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same circuit-board assembly with visually anchored locations, such as outlined slots or labelled positions, while keeping the board layout and part count identical; if users still produce mostly implicit utterances and long pointing times, the claim that spatial ambiguity drives the behavior is wrong. Alternatively, within the released dataset, compare average pointing duration on implicit versus explicit utterances: if pointing is not longer for implicit commands, the implied linkage between vague language and prolonged deictic compensation fails.","supporting_citations":[{"cited_title":"put-that-there","cited_arxiv_id":null,"evidence_quote":"Supplies the put-that-there voice-plus-gesture paradigm used to define the explicit/implicit command labels and to interpret user strategies."},{"cited_title":"Oviatt, P","cited_arxiv_id":null,"evidence_quote":"Provides the multimodal speech-and-gesture design principles that frame the study's analysis of modality integration."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supports the claim that users elaborate spatial context with other modalities, which the discussion uses to interpret anchored versus ambiguous behavior."},{"cited_title":"Oviatt, A","cited_arxiv_id":null,"evidence_quote":"Reports how spatial context is elaborated with gestures while voice carries description, used to interpret the task differences observed."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Informs the research questions about how modalities are integrated in multimodal interaction."},{"cited_title":"Radford, J","cited_arxiv_id":null,"evidence_quote":"Supplies the speech-to-text model used to transcribe voice recordings into timestamped command text for annotation."}],"review_version":1}