{"id":"e431ee46-58c7-4437-b17b-016e6dae07c6","arxiv_id":"2605.30740","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":4.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"GSAM integrates vision perceiver, VLM refiner with CoT, LLM constraint generator, and kinematic planner to raise success rate by 36% and cut variance by 3.1% on 50 hinge tasks across 5 categories.","lead":"GSAM is a robotic system that estimates object kinematics from vision, refines estimates with a VLM using commonsense reasoning, generates collision-avoidance constraints via LLM, and plans reachable trajectories. A smart generalist might read it to see how current AI models are stitched together for safer robot handling of doors and drawers.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.3","headline":"Experimental gains rest on unablated assumption that VLM refiner meaningfully corrects perceiver kinematics","rationale":"The reader's weakest assumption identifies precisely this link; the experimental claim cannot be accepted at face value without evidence that the refiner actually improves the kinematic inputs that downstream modules rely on.","tokens_in":1767,"tokens_out":307,"duration_ms":18353,"concrete_test":"Re-run the 50-task, 50-configuration evaluation with the VLM refiner disabled (raw perceiver outputs fed directly to constraint generation); if success rate drops by more than 15 percentage points or std reduction falls below 1%, the refiner is not load-bearing for the reported gains.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The headline result (36% success-rate lift, 3.1% std reduction on 50 hinge tasks) is presented as evidence of superior generalization and safety. This attribution requires that the fine-tuned VLM refiner, via CoT commonsense reasoning, produces kinematic estimates sufficiently more accurate than the raw vision perceiver to enable better constraint generation and collision-free planning. The abstract supplies no quantitative perception-error metrics (e.g., joint-angle or pose RMSE with vs. without refiner) and no ablation that isolates the refiner. If the refiner step yields only marginal or inconsistent corrections, the observed gains could be driven by the constraint generator, the kinematic planner, or baseline differences rather than the claimed perception-correction mechanism.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The paper claims to introduce GSAM, a framework for articulated object manipulation consisting of a vision-based perceiver for kinematic parameters, a fine-tuned VLM refiner using chain-of-thought commonsense reasoning to correct raw estimations, an interaction constraint function generator, an LLM to functionalize constraints for trajectory and posture planning, and a kinematic-aware manipulation planner. On 50 hinge tasks across 5 object categories with 50 random end-effector-handle configurations, it reports a 36.0% improvement in manipulation success rate and 3.1% reduction in standard deviation over the best baseline.","tokens_in":1930,"tokens_out":441,"duration_ms":18437,"significance":"If validated, the approach could contribute to more generalizable and safer robotic manipulation of articulated objects by mitigating perception errors and collision risks through hybrid vision-VLM-constraint methods. The emphasis on commonsense reasoning in perception refinement is a notable aspect for real-world service robot applications.","major_comments":[{"comment":"The headline experimental result (36.0% success-rate lift, 3.1% std reduction on 50 tasks) is presented without details on baseline implementations, statistical significance, error bars, data splits, or random seeds, preventing verification of the central performance claim.","section":"Abstract"},{"comment":"The framework's claimed gains rest on the unablated assumption that the VLM refiner meaningfully corrects the perceiver's kinematic estimates; no perception-error metrics (e.g., joint-angle RMSE with vs. without refiner) or ablation isolating the refiner are provided, leaving the attribution of improvements to this component unsupported.","section":"Abstract / Experiments"}],"minor_comments":[{"comment":"Typographical error: 'f ine-tuned' appears with a space; should be 'fine-tuned'.","section":"Abstract"},{"comment":"Missing hyphen/space: 'end-effectorhandle' should read 'end-effector-handle'.","section":"Abstract"}],"recommendation":"major_revision","confidential_remarks":"The experimental section appears to lack standard reproducibility elements; this may warrant requesting additional details in revision."},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for the constructive feedback on experimental clarity and component validation. We address each major comment below and commit to revisions that strengthen the presentation without altering the core claims.","responses":[{"response":"We agree that the abstract would benefit from additional context. The full manuscript details baseline implementations, statistical significance testing (including p-values), error bars from repeated trials, data splits across the 5 object categories, and random seeds in Section 4 (Experiments). In the revision we will expand the abstract with a brief clause referencing these elements and the experimental protocol to improve verifiability while respecting length constraints.","revision_made":"yes","referee_comment":"[Abstract] The headline experimental result (36.0% success-rate lift, 3.1% std reduction on 50 tasks) is presented without details on baseline implementations, statistical significance, error bars, data splits, or random seeds, preventing verification of the central performance claim."},{"response":"We acknowledge the value of isolating the VLM refiner's contribution. The current manuscript reports end-to-end success rates but does not include perception-specific metrics or a dedicated ablation. We will add these to the revised Experiments section, including joint-angle RMSE comparisons with and without the refiner, plus an ablation table attributing performance gains to individual components.","revision_made":"yes","referee_comment":"[Abstract / Experiments] The framework's claimed gains rest on the unablated assumption that the VLM refiner meaningfully corrects the perceiver's kinematic estimates; no perception-error metrics (e.g., joint-angle RMSE with vs. without refiner) or ablation isolating the refiner are provided, leaving the attribution of improvements to this component unsupported."}],"tokens_in":1400,"tokens_out":378,"duration_ms":18581,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"GSAM stitches a standard vision perceiver to a fine-tuned VLM that applies chain-of-thought commonsense to clean up raw kinematic estimates, then feeds the result into a constraint function generator and an LLM-based planner that adds collision avoidance. The main reported outcome is a 36% higher success rate and 3.1% lower standard deviation across 50 hinge tasks in five categories with randomized starts.\n\nThe constraint generator that folds object type, interaction pose, and obstacle knowledge into a single base is a practical engineering move that directly targets the safety problem the abstract flags. Running the whole pipeline on multiple object categories and random end-effector placements gives at least a basic check on generalization within the hinge domain.\n\nThe central weakness is the missing evidence for the refiner step. The abstract supplies no joint-angle or pose error numbers comparing the raw perceiver to the VLM-corrected version, and no ablation that removes the refiner. Without those, the 36% gain could come from the planner, the constraint generator, or differences in how baselines were implemented. The text also gives no detail on baseline code, random seeds, or statistical tests.\n\nThis paper is for robotics groups already working with VLMs and LLMs on manipulation who want a concrete pipeline example. It is worth sending to peer review once the perception metrics and ablations are added, because the safety-constraint idea is worth checking properly even if the current numbers are hard to interpret.","headline":"GSAM assembles a vision perceiver, VLM refiner, constraint generator, and LLM planner for hinge manipulation and reports a 36% success lift, but the lift lacks any ablation or perception-error numbers to show the refiner is responsible.","tokens_in":2419,"tokens_out":390,"would_cite":false,"duration_ms":15586,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"GSAM corrects vision-based kinematic estimates with a VLM refiner and uses LLM-generated constraints to plan safer trajectories for articulated objects.","keywords":["articulated object manipulation","robotic framework","vision-based perceiver","VLM refiner","interaction constraints","kinematic planning","generalization","collision avoidance"],"falsifier":"An ablation experiment that removes the VLM refiner and shows that uncorrected kinematic estimates produce no gain in success rate or reduction in standard deviation compared with the best baseline.","tokens_in":2679,"feed_emoji":"🤖","tokens_out":671,"duration_ms":16706,"temperature":0.7,"pith_summary":"The paper introduces GSAM to handle the diversity of articulated objects and the risks of end-effector collisions that limit existing robotic methods. It generates initial kinematic parameters from vision, refines them through a fine-tuned VLM that applies chain-of-thought commonsense reasoning, and creates interaction constraints that an LLM turns into planning rules. A kinematic-aware planner then verifies reachability while avoiding obstacles. Experiments across 50 hinge tasks in five object categories and varied starting positions report higher success rates and lower variability than baselines. Readers would care because service robots need reliable ways to open doors, drawers, and similar items without causing damage.","feed_headline":"GSAM raises articulated object success rate by 36 percent","feed_subtitle":"VLM-refined kinematics and LLM constraints cut variability and collisions across five object types.","key_machinery":"The interaction constraint function generator that folds articulated object geometry, interaction pose, and obstacle avoidance into constraints later functionalized by an LLM.","core_discovery":"GSAM generates kinematic parameters via a vision-based perceiver, refines raw estimates that deviate from commonsense using a fine-tuned VLM with chain-of-thought reasoning, builds an interaction constraint function generator that encodes object, pose, and obstacle knowledge, lets an LLM functionalize those constraints for trajectory and posture planning, and applies a kinematic-aware manipulation planner; this combination yields a 36.0 percent higher success rate and 3.1 percent lower standard deviation on 50 hinge tasks across five object categories and 50 random end-effector-handle configurations.","pith_inferences":["The correction step could apply to other perception-heavy robotic tasks where commonsense priors improve raw sensor data.","Extending the constraint generator beyond hinges would require only new object geometry inputs if the refiner generalizes.","Real-time versions might reduce planning latency if the LLM functionalizer is distilled into faster models."],"forward_implications":["Manipulation succeeds on a wider range of articulated objects from multiple categories.","Random initial end-effector positions lead to fewer failed attempts.","Destructive collisions are reduced through explicit constraint enforcement during planning.","Performance variability across repeated trials drops measurably.","The same pipeline supports both trajectory and posture planning under reachability checks."],"fun_headline_variants":["GSAM achieves 36% higher success in hinge manipulation","Robotic GSAM lowers std by 3.1% on 50 random tasks","Fine-tuned VLM refines kinematics for safer articulations","GSAM uses interaction constraints for collision-free planning"],"cache_read_input_tokens":2112,"weakest_assumption_plain":"The vision-based perceiver produces raw kinematic estimates that deviate from commonsense and require reliable correction by the fine-tuned VLM refiner using chain-of-thought reasoning.","fun_headline_variants_meta":{"raw":{"variants":["GSAM achieves 36% higher success in hinge manipulation","Robotic GSAM lowers std by 3.1% on 50 random tasks","Fine-tuned VLM refines kinematics for safer articulations","GSAM uses interaction constraints for collision-free planning"]},"model":"grok-4.3","cost_usd":0.007245,"raw_usage":{"total_tokens":3367,"prompt_tokens":723,"num_sources_used":0,"completion_tokens":69,"cost_in_usd_ticks":72449500,"prompt_tokens_details":{"text_tokens":723,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":2575,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":723,"tokens_out":69,"duration_ms":16458,"temperature":1.0,"reasoning_tokens":2575,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-06-28T22:38:49.852376+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"An ablation experiment that removes the VLM refiner and shows that uncorrected kinematic estimates produce no gain in success rate or reduction in standard deviation compared with the best baseline.","supporting_citations":[],"review_version":1}