{"id":"669e07c4-c7e9-4e0b-8821-b3c7a1ed40ba","arxiv_id":"2507.04791","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"Voice commands are grounded in the scene via a vision-language model and converted into collision-avoidance meshes, and a five-person pilot shows the bimanual teleoperation system avoids collisions while replayed commands without it do not.","lead":"This paper presents a teleoperation system where a person controls a two-armed robot with VR controllers and tells it by voice which objects to avoid. The system turned spoken commands into 3D obstacle models and prevented collisions in a five-person pilot, while replays without the safety feature collided in four of five trials.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Replay counterfactual removes operator from the loop, so 4/5 replay collisions do not demonstrate improved live safety.","rationale":"After reading the paper and the reader's verdict, the replay counterfactual is the most load-bearing threat to the central claim because it is the only quantitative evidence for the safety improvement. Without a valid baseline for unassisted teleoperation, zero live collisions alone could simply reflect careful operators or a benign task. The reader's weakest assumption identifies exactly this issue, and my analysis agrees. I do not find additional internal inconsistencies: the system architecture is coherent, the collision-avoidance formulation is standard, and the pilot results are reported transparently. The efficiency claim is also unsupported, but the safety claim is more fundamental. The proposed simulation experiment would directly settle whether the replay is misleading. Since the reader already conditionally accepted the paper on the strength of addressing these experimental gaps, my concern does not change the verdict; it reinforces the need for that condition. Therefore I recommend no change to the reader's CONDITIONAL verdict.","tokens_in":7153,"tokens_out":6796,"duration_ms":74301,"concrete_test":"Run a within-subjects study in a high-fidelity simulator of the same pick-and-place task with at least 10 participants, each performing the task once with collision avoidance active and once with it disabled (order counterbalanced). Instruct participants identically in both conditions, and record collision counts and task completion times. If the live no-safety condition yields a collision rate similar to the replay rate (e.g., 4/5), the replay is a valid proxy and the safety improvement is supported; if the live no-safety rate is significantly lower, the replay overestimates the safety benefit and the central claim must be weakened or reworded.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that the system 'significantly improves operational safety' rests entirely on the replay comparison in Section III. The paper reports that control commands recorded during live trials were replayed on the robot with collision avoidance disabled, and that this caused collisions in four of five trials. This replay is not a valid proxy for live teleoperation without the safety system. In live teleoperation the operator is in the loop and can observe the robot's approach to an obstacle and correct the commands; the replay removes that feedback loop, so the executed trajectory is open-loop and cannot reflect the operator's adaptive response. Moreover, operators' behavior in the live trials may itself have been influenced by the presence of the safety system (risk compensation), so the recorded commands may be more aggressive than they would be without it. Consequently, the observed difference between zero live collisions and four replay collisions does not isolate the system's contribution to safety; it only shows that the raw commanded motions were unsafe without the collision-avoidance constraints.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a bimanual teleoperation system that combines VR controller input with spoken, language-grounded collision avoidance. A speech-to-text and LLM pipeline parses commands into object references; Grounding DINO and SAM segment the referred objects in RGB-D data; a point-cloud pipeline generates convex-hull obstacle meshes; and a whole-body quadratic-programming controller incorporates these meshes as velocity-damping constraints. The empirical section reports a five-participant pilot study of pick-and-place in a cluttered scene: no collisions were observed in live trials, while four out of five replayed trials (with collision avoidance disabled) ended in collisions. The authors also show qualitative demonstrations of more complex tasks. The central claim is that the system significantly improves operational safety without compromising task efficiency.","tokens_in":7315,"tokens_out":6617,"duration_ms":72660,"significance":"The paper's main strength is a clean system integration: it closes the loop from natural-language commands to real-time collision-avoidance constraints in a whole-body controller, and it demonstrates the full pipeline on a physical robot. The live pilot, despite its small size, is genuine evidence that the system can execute pick-and-place without collisions. The replay-without-avoidance condition is an imaginative counterfactual and provides a useful lower bound on how risky the raw commanded trajectories are. However, the abstract's quantitative claims ('significantly improves,' 'without compromising task efficiency') go beyond what the evidence can support: the replay comparison is confounded by removing the operator from the loop, and no efficiency measure is reported. The contribution is therefore best framed as a promising system demonstration whose safety benefit is suggestive rather than established.","major_comments":[{"comment":"The replay comparison does not isolate the safety module's contribution to live teleoperation. In the replay condition, the operator is removed from the loop, whereas in live teleoperation the operator sees the robot approaching an obstacle and can stop or correct the command. The recorded commands may also have been produced under risk compensation: with the safety layer active, operators may have moved more aggressively than they would have without it. The observed contrast (0/5 live collisions versus 4/5 replay collisions) is therefore not a valid proxy for live operation without collision avoidance. I recommend reporting the replay as an open-loop counterfactual that demonstrates the raw commanded motions were unsafe without the constraints, and softening the causal claim about the system's contribution to live safety accordingly.","section":"III (Evaluation)"},{"comment":"The claim that the system improves safety 'without compromising task efficiency' is not supported because no efficiency metric is reported. The pilot section includes no task completion time, no number of corrective interventions, no success rate, and no comparison of any workload or fluency measure. Without at least one quantitative efficiency measure, the 'without compromising' assertion is unverifiable. The authors should either add such a measure (e.g., completion time or time-to-completion relative to a baseline) or explicitly state that efficiency was not measured and remove the claim.","section":"Abstract and Section III"},{"comment":"With five participants and no inferential statistics, the word 'significantly' is not justified. The descriptive comparison (0/5 live collisions versus 4/5 replay collisions) is based on a single binary outcome per trial and does not support a claim of statistical significance. The authors should apply an appropriate test to a defensible paired outcome (for example, an exact McNemar test on paired collision outcomes) or characterize the result as a pilot demonstration rather than a significant improvement.","section":"III (Evaluation)"}],"minor_comments":[{"comment":"The word 'open-vocabulaty' should be 'open-vocabulary'.","section":"II-A.2"},{"comment":"The word 'usfing' should be 'using'.","section":"II-C"},{"comment":"The phrase 'uniformly scaled 2 around its centroid' is unclear; the footnote says the scale was 1.05, so the text should read 'scaled by a factor of 1.05'.","section":"II-A.3"},{"comment":"The participant age description 'aged 22 ± 1 min 21, max 30' is ambiguous and internally implausible; please report mean and standard deviation separately and specify the range clearly.","section":"III"},{"comment":"The system is described as providing 'immersive VR control,' but the interface is monitor-based with no head-mounted display; 'immersive' is misleading and should be replaced with a more neutral term such as 'VR-controller-based.'","section":"Abstract and II-B"},{"comment":"The Spanish phrase 'Agriega la mesa' should be 'Agrega la mesa'.","section":"Fig. 5 caption"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is better framed as a demo/systems paper than as a controlled study. The live pilot and replay data are suggestive, but the abstract overstates the evidence. If the venue accepts system demonstrations, a revision that adds an efficiency metric, uses appropriate statistics, and appropriately qualifies the replay counterfactual would likely make the paper acceptable. The self-citations are not excessive, but the claims in the abstract should be aligned with the pilot's limited scope."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe genuinely new thing here is the integration: voice-command grounding (Whisper, an LLM for parsing, Grounding DINO, SAM) feeding object meshes into a whole-body QP controller that enforces velocity-damping collision constraints during bimanual teleoperation. I have not seen that combination before, and it is a coherent, modular engineering contribution. The controller math is standard and cited correctly, and the authors are candid about perceptual limitations.\n\nThe soft spots are in the evaluation. The pilot has five participants, no statistics, and no efficiency metric, yet the abstract claims the system 'significantly improves operational safety without compromising task efficiency.' Neither half of that claim is supported. Zero live collisions is nice, but there was no live no-safety condition. The replay of recorded commands with avoidance disabled is not a valid counterfactual: it removes the operator from the loop, so the replayed trajectory is open-loop and cannot reflect the operator's adaptive reactions. It also ignores risk compensation—operators who know the safety net is active may command more aggressively. So the 4/5 replay collisions show that the raw commands were unsafe without the filter, but they do not show that the filter improves live safety. The efficiency claim is simply unmeasured.\n\nThe mechanism itself is sound, and the system would likely prevent collisions whenever the segmentation mesh is accurate and the constraint activates. So this is a real contribution looking for a fair test. I would send it to peer review: the experimental gap is fixable, and the integration deserves referee time. But the authors should temper the abstract and, ideally, run a live condition without the safety module, add at least a completion-time measure, and report participant-level data.","headline":"A genuinely novel integration of voice-command grounding with whole-body collision avoidance, undercut by an overclaimed pilot evaluation and an invalid replay counterfactual.","tokens_in":7862,"tokens_out":2889,"would_cite":false,"duration_ms":34878,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A voice command naming an object turns it into a collision barrier for a bimanual teleoperated robot.","keywords":["bimanual teleoperation","collision avoidance","vision-language grounding","whole-body control","virtual reality teleoperation","speech interface","point cloud segmentation","mobile manipulator"],"falsifier":"Run a fresh set of operators on the same cluttered pick-and-place task with no collision avoidance from the start. If those operators, knowing they have no safety net, are measurably more cautious and collide rarely or not at all, then the four-of-five collision rate in the replayed condition does not isolate the system's contribution; conversely, a similar collision rate under live no-safety operation would support the paper's claim.","tokens_in":6962,"feed_emoji":"🤖","tokens_out":4959,"duration_ms":51558,"temperature":0.7,"pith_summary":"Teleoperating two arms in a cluttered scene is hard because the operator cannot reliably judge distances and may knock objects over. This paper proposes a system in which the operator simply says what to avoid, and the robot converts that object into a 3D collision mesh that its whole-body controller actively keeps away from. The paper reports that in a five-participant pilot, no live collision occurred while using the system, while replaying the same recorded controller commands with collision avoidance switched off produced collisions in four of five trials. A sympathetic reading is that the language-triggered avoidance layer, not operator caution alone, accounts for the observed safety gain.","feed_headline":"Spoken obstacle names prevent collisions in bimanual teleop","feed_subtitle":"Five-person pilot: no live collisions; four of five replayed runs without the safety net hit objects.","key_machinery":"The mechanism that carries the argument is a speech-to-mesh-to-controller pipeline. Audio is transcribed, a large language model parses the transcript into structured object references, an open-vocabulary detector finds the referenced object in the camera image, a segmentation model produces its mask, and the mask is projected onto the registered point cloud and reconstructed as a convex hull with a safety margin. That mesh becomes a collision object in a whole-body controller that solves a quadratic program at each control step; the relevant constraint is a velocity-damping inequality that slows the robot's links as their distance to a named object falls below a threshold. The pipeline's defining property is that the obstacle set is dynamic and operator-specified rather than a pre-labeled static map.","core_discovery":"The paper's central claim is that natural-language commands can be used as a real-time control input for collision avoidance: a spoken description such as 'avoid the yellow tool' is grounded to a segmented image region, lifted to a 3D mesh via the camera point cloud, and injected into the whole-body quadratic-programming controller as an obstacle whose proximity damps the robot's velocity. This makes safety task-dependent and intent-aware: only objects the operator names become avoidance constraints, so the robot does not freeze in response to every unmodeled surface. The pilot study is offered as evidence that the approach improves operational safety without compromising task efficiency, since completing the same pick-and-place motions with the avoidance active produced zero collisions and the recorded commands without avoidance collided in most trials.","pith_inferences":["The replay protocol does not include a live no-safety condition, so part of the safety gain may reflect changes in operator caution when the safety layer is known to be active.","A straightforward extension would be periodic re-segmentation so that objects that move or become occluded after the initial voice command are re-meshed, which would be needed before dynamic scenes.","The same voice-to-mesh pipeline could be turned from an avoidance tool into a manipulation tool by naming the object to grasp, reusing the grounding and segmentation stage.","Adding wrist-mounted cameras would likely reduce the single head-mounted camera's occlusion problem and improve mesh completeness, as the paper itself notes as future work."],"forward_implications":["Operators can add or remove obstacles by voice mid-task without pausing or touching a keyboard.","Only user-named objects become avoidance constraints, so the controller can ignore task-irrelevant clutter that a blanket point-cloud risk model would treat as dangerous.","Because the obstacle set is built online from perception, the same controller stack can be deployed in new scenes without manual labeling.","The safety behavior generalizes across languages and phrasings, since the same object can be referred to in several ways and the language layer resolves the reference.","If the pilot result holds, the system offers a path to safer data collection for imitation learning in cluttered domestic settings."],"supporting_citations":[{"why":"Supplies the speech-to-text transcription that turns spoken commands into text for the language parser.","marker":"[9]"},{"why":"Grounds free-form object descriptions to 2D bounding boxes in the camera image.","marker":"[10]"},{"why":"Produces the high-resolution 2D segmentation masks that are projected into 3D.","marker":"[11]"},{"why":"Provides the Cartesian control layer that turns VR end-effector references into joint commands.","marker":"[18]"},{"why":"Solves the hierarchical whole-body quadratic program with tasks and inequality constraints.","marker":"[19]"},{"why":"Supplies the velocity-damping inequality formulation used to slow links near obstacles.","marker":"[20]"},{"why":"Computes the minimum distances between robot links and object meshes used to trigger the constraints.","marker":"[21]"}],"fun_headline_variants":["Voice commands steer bimanual robots clear of clutter","Language-guided teleop reduces collisions in bimanual tasks","Say 'avoid that' to keep dual-arm robots safe","Spoken warnings guide bimanual teleoperation safely","Natural language keeps bimanual teleop collision-free"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The evaluation assumes that replaying the operator's recorded commands with collision avoidance turned off is a fair measure of what would happen live without the safety system, which requires operators to move the same way whether or not avoidance is active.","fun_headline_variants_meta":{"raw":{"variants":["Voice commands steer bimanual robots clear of clutter","Language-guided teleop reduces collisions in bimanual tasks","Say 'avoid that' to keep dual-arm robots safe","Spoken warnings guide bimanual teleoperation safely","Natural language keeps bimanual teleop collision-free"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00019,"raw_usage":{"total_tokens":1289,"prompt_tokens":843,"completion_tokens":446,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":459,"completion_tokens_details":{"reasoning_tokens":368}},"tokens_in":459,"tokens_out":446,"duration_ms":4480,"temperature":1.0,"reasoning_tokens":368,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T19:39:26.164321+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run a fresh set of operators on the same cluttered pick-and-place task with no collision avoidance from the start. If those operators, knowing they have no safety net, are measurably more cautious and collide rarely or not at all, then the four-of-five collision rate in the replayed condition does not isolate the system's contribution; conversely, a similar collision rate under live no-safety operation would support the paper's claim.","supporting_citations":[{"cited_title":"Robust speech recognition via large-scale weak supervision,","cited_arxiv_id":null,"evidence_quote":"Supplies the speech-to-text transcription that turns spoken commands into text for the language parser."},{"cited_title":"Grounding dino: Marrying dino with grounded pre-training for open-set object detection,","cited_arxiv_id":null,"evidence_quote":"Grounds free-form object descriptions to 2D bounding boxes in the camera image."},{"cited_title":"Segment anything,","cited_arxiv_id":null,"evidence_quote":"Produces the high-resolution 2D segmentation masks that are projected into 3D."},{"cited_title":"Cartesi/o: A ros based real-time capable cartesian control framework,","cited_arxiv_id":null,"evidence_quote":"Provides the Cartesian control layer that turns VR end-effector references into joint commands."},{"cited_title":"The open stack of tasks library: Opensot: A software dedicated to hierarchical whole-body control of robots subject to constraints,","cited_arxiv_id":null,"evidence_quote":"Solves the hierarchical whole-body quadratic program with tasks and inequality constraints."},{"cited_title":"Efficient self-collision avoidance based on focus of interest for humanoid robots,","cited_arxiv_id":null,"evidence_quote":"Supplies the velocity-damping inequality formulation used to slow links near obstacles."},{"cited_title":"Coal: an extension of the flexible collision library,","cited_arxiv_id":null,"evidence_quote":"Computes the minimum distances between robot links and object meshes used to trigger the constraints."}],"review_version":1}