{"id":"bfe30965-077e-4484-9fce-fc36df75350a","arxiv_id":"2605.18617","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"ManiSoft is a new benchmark featuring a soft-body simulator, four deformable control tasks, and an automated pipeline generating 6300 scenes with expert trajectories for training and evaluating vision-language policies on continuum robots.","lead":"The paper introduces ManiSoft, a benchmark with a custom simulator, four tasks, and 6300 expert trajectories for vision-language manipulation using soft continuum robotic arms. This could help develop more adaptable robots for cluttered or confined spaces where rigid arms struggle.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.3","headline":"Simulator fidelity unverified: elastic force constraint may not capture continuum deformation or contact dynamics accurately enough for a representative benchmark.","rationale":"The reader's weakest assumption correctly isolates the simulator and trajectory generator as the load-bearing elements. Full-text access does not remove the concern because the abstract's description of the modeling choice remains the only concrete technical detail supplied, and no independent verification or real-world grounding is indicated. This keeps the paper in a CONDITIONAL state rather than UNVERDICTED or ACCEPT; the benchmark could still be useful if the dynamics check passes, but the central bridging claim cannot be endorsed without it.","tokens_in":1747,"tokens_out":360,"duration_ms":30403,"concrete_test":"For a simple cantilever bending task with known torque sequence, compare simulated tip position, curvature profile, and contact force time series against an independent Cosserat rod or FEM implementation using identical material parameters and boundary conditions; if mean tip error exceeds 8% or contact force RMSE exceeds 15% of peak force, the elastic constraint modeling is insufficient.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The claim that ManiSoft is a valuable testbed rests on the tailored simulator coupling realistic soft-body dynamics with contact-rich interactions via an elastic force constraint, plus the high-level planner + low-level RL pipeline yielding high-quality expert trajectories. Even with full text, no quantitative validation against Cosserat rod models, finite-element references, or physical soft-arm data is described for tip position, curvature, or force profiles under distributed actuation. The reported performance drop under randomization and attribution to visual proprioception errors could equally indicate that the underlying dynamics do not reproduce the distributed compliance and actuation challenges of real continuum robots, undermining transferability of the benchmark.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The paper introduces ManiSoft, a benchmark for vision-language manipulation with soft continuum robotic arms. It features a tailored simulator coupling soft-body dynamics and contact-rich interactions via an elastic force constraint, defines four tasks highlighting deformable control aspects, generates 6300 diverse scenes with expert trajectories via a high-level planner followed by low-level RL policy for torque commands, and benchmarks three representative policy models that show promising results in clean scenes but substantial performance drops under randomization, with failures attributed primarily to inaccurate visual proprioception and limited exploitation of deformability.","tokens_in":1881,"tokens_out":640,"duration_ms":39685,"significance":"If the simulator's dynamics prove representative of real soft continuum robots, ManiSoft could serve as a useful testbed bridging rigid and soft arm research in vision-language manipulation. The automated pipeline for scene and trajectory generation at scale, along with released code and datasets, supports reproducibility and community use. The empirical focus on challenges like unreliable proprioception and distributed actuation in soft arms addresses a genuine gap, though the absence of detailed quantitative validation metrics limits immediate impact assessment.","major_comments":[{"comment":"Simulator description (likely §3 or equivalent): the claim that the tailored simulator 'couples realistic soft-body dynamics with contact-rich interactions via an elastic force constraint' lacks quantitative validation against Cosserat rod models, finite-element references, or physical soft-arm data for metrics such as tip position error, curvature profiles, or force under distributed actuation; this is load-bearing for the central claim that ManiSoft is a representative testbed.","section":"Simulator section"},{"comment":"Benchmarking and results section: the reported 'substantial performance drop under randomization' and attribution to 'inaccurate visual estimation of proprioceptive state' are presented without supporting data tables, ablation studies, or error analysis (e.g., success rates, proprioception error distributions), undermining assessment of whether the drop reflects soft-robot challenges or simulator limitations.","section":"Benchmarking results"},{"comment":"Trajectory generation pipeline: the high-level planner plus low-level RL policy is asserted to produce 'high-quality expert trajectories' suitable for policy training, but no quantitative metrics on waypoint tracking accuracy, trajectory smoothness, or success rates across the 6300 scenes are provided to substantiate this for downstream evaluation.","section":"Expert trajectory generation"}],"minor_comments":[{"comment":"Abstract: 'Out codes and datasets' should be corrected to 'Our codes and datasets'.","section":"Abstract"},{"comment":"Notation and figures: ensure consistent use of task names and that visualizations of failure cases clearly label proprioception errors versus deformability exploitation.","section":"Figures and notation"}],"recommendation":"major_revision","confidential_remarks":"The manuscript fits well within robotics and embodied AI venues given the focus on soft robotics benchmarks. Citation of prior soft-robot simulators and VL manipulation works appears balanced, but ensure the related work section explicitly contrasts against existing rigid-arm benchmarks like those in RLBench or VIMA to clarify novelty."},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for their thoughtful and constructive comments. We address each major comment point by point below, indicating the revisions we will make to strengthen the manuscript.","responses":[{"response":"We agree that quantitative validation would strengthen the simulator's credibility as a representative testbed. In the revised manuscript, we will add a dedicated subsection with direct comparisons to Cosserat rod models and finite-element simulations, reporting metrics such as tip position error and curvature profiles under distributed actuation. For physical soft-arm data, our work is currently simulation-focused; we will explicitly note this scope limitation and discuss it as an avenue for future validation.","revision_made":"yes","referee_comment":"[Simulator section] Simulator description (likely §3 or equivalent): the claim that the tailored simulator 'couples realistic soft-body dynamics with contact-rich interactions via an elastic force constraint' lacks quantitative validation against Cosserat rod models, finite-element references, or physical soft-arm data for metrics such as tip position error, curvature profiles, or force under distributed actuation; this is load-bearing for the central claim that ManiSoft is a representative testbed."},{"response":"We acknowledge that additional empirical support is needed to clarify the sources of performance drops. We will incorporate new tables reporting success rates in clean versus randomized scenes, proprioception error distributions, and ablation studies isolating the contributions of visual estimation inaccuracies and randomization effects in the revised results section.","revision_made":"yes","referee_comment":"[Benchmarking results] Benchmarking and results section: the reported 'substantial performance drop under randomization' and attribution to 'inaccurate visual estimation of proprioceptive state' are presented without supporting data tables, ablation studies, or error analysis (e.g., success rates, proprioception error distributions), undermining assessment of whether the drop reflects soft-robot challenges or simulator limitations."},{"response":"We will expand the trajectory generation section to include quantitative metrics, specifically average waypoint tracking error, trajectory smoothness measures (such as integrated jerk), and overall success rates of the expert policy across all 6300 scenes. These additions will be placed in the main text or supplementary material to better substantiate trajectory quality.","revision_made":"yes","referee_comment":"[Expert trajectory generation] Trajectory generation pipeline: the high-level planner plus low-level RL policy is asserted to produce 'high-quality expert trajectories' suitable for policy training, but no quantitative metrics on waypoint tracking accuracy, trajectory smoothness, or success rates across the 6300 scenes are provided to substantiate this for downstream evaluation."}],"tokens_in":1532,"tokens_out":550,"duration_ms":40149,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"This paper's core offering is ManiSoft, a benchmark that targets vision-language manipulation for soft continuum robots rather than the usual rigid arms. It supplies a custom simulator, four tasks focused on deformable control, and a dataset of 6300 scenes with expert trajectories to let people train and evaluate policies on the specific problems of distributed actuation and unreliable proprioception.","headline":"ManiSoft adds a benchmark and simulator for vision-language tasks on soft continuum arms, with generated trajectories and some policy tests, but the simulator's match to real dynamics stays unverified.","tokens_in":2349,"tokens_out":153,"would_cite":false,"duration_ms":36517,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":{"model":"grok-4.3","evidence":[],"headline":"Soft-robotics VLM benchmark with Cosserat+elastic simulator; no overlap with RS distinction-to-physics chain","alignment":"orthogonal","rationale":"Paper centers on hybrid simulator (Elastica Cosserat rod + MuJoCo via elastic force constraint), hierarchical waypoint+RL torque control, and four manipulation tasks for deformable arms. RS derives J-cost, φ, 8-tick period, D=3 and constants parameter-free from single distinction (reality_from_one_distinction, AbsoluteFloorClosure, Cost/FunctionalEquation, AlexanderDuality). No shared structures, cost functions, periodicity or ratio symmetry present; domain is applied robotics engineering.","tokens_in":56945,"confidence":"high","tokens_out":154,"duration_ms":14129,"cache_read_input_tokens":38528,"cache_creation_input_tokens":0},"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"ManiSoft introduces a benchmark for vision-language manipulation using soft continuum robotic arms.","keywords":["soft robotics","vision-language manipulation","continuum robots","robotic benchmark","soft body simulation","reinforcement learning","deformable control","trajectory generation"],"falsifier":"If policies trained on ManiSoft trajectories show no performance advantage over rigid-arm methods when tested on physical soft robots in confined spaces, the benchmark would fail to demonstrate a useful bridge between the two domains.","tokens_in":2655,"feed_emoji":"🤖","tokens_out":709,"duration_ms":72852,"temperature":0.7,"pith_summary":"Most vision-language manipulation research targets rigid robotic arms whose fixed shape limits work in cluttered or confined spaces. Soft arms can deform to reach such areas but introduce unreliable proprioception and distributed actuation challenges. The paper presents ManiSoft as a benchmark featuring a tailored simulator that models soft-body dynamics and contact interactions through an elastic force constraint, along with four tasks that stress different aspects of deformable control. An automated pipeline creates 6,300 scenes and expert trajectories by combining a high-level planner that sets waypoints with a low-level reinforcement learning policy that outputs torques. Tests of three policy models succeed in clean scenes yet degrade sharply under randomization, mainly from visual errors in estimating arm state and underuse of bending for obstacle avoidance.","feed_headline":"ManiSoft benchmark tests vision-language control on soft robotic arms","feed_subtitle":"A simulator with 6,300 expert trajectories trains policies for deformable manipulation in cluttered spaces.","key_machinery":"The tailored simulator that couples realistic soft-body dynamics with contact-rich interactions via an elastic force constraint, together with the high-level planner and low-level RL policy for expert trajectory generation.","core_discovery":"ManiSoft is a benchmark for vision-language manipulation for soft continuum robotics. It features a tailored simulator that couples realistic soft-body dynamics with contact-rich interactions via an elastic force constraint. On this basis, ManiSoft defines four tasks, each highlighting distinct aspects of deformable control, from basic end-effector coordination to obstacle avoidance. To support policy training and evaluation, ManiSoft includes an automated pipeline that generates 6,300 diverse scenes and corresponding expert trajectories. To produce high-quality trajectories at scale, a high-level planner decomposes each task into a sequence of waypoints, followed by a low-level RL policy to","pith_inferences":["Better vision models focused on estimating continuous arm shape could narrow the performance gap seen under randomization.","Policies that explicitly plan bending sequences around obstacles might make fuller use of soft-arm deformability.","Transferring ManiSoft-trained policies to hardware soft robots would test whether the elastic-force simulator captures real dynamics.","The benchmark could extend to other soft-robot domains such as navigation or inspection in tight environments."],"forward_implications":["The four tasks allow systematic testing of basic coordination through advanced obstacle avoidance in deformable systems.","The pipeline for generating 6,300 scenes and trajectories supports scalable training and evaluation of vision-language policies.","Benchmark results identify concrete failure modes in visual proprioception and deformability exploitation that new methods must solve.","ManiSoft can function as a shared testbed that transfers techniques from rigid-arm research to soft-arm settings."],"fun_headline_variants":["ManiSoft benchmark for vision-language soft arm manipulation","ManiSoft tests vision-language control on soft continuum arms","ManiSoft simulator aids vision-language tasks for deformable robots","ManiSoft generates 6300 trajectories for soft arm vision-language tasks"],"cache_read_input_tokens":2112,"weakest_assumption_plain":"The tailored simulator accurately couples realistic soft-body dynamics with contact-rich interactions via an elastic force constraint, and the high-level planner plus low-level RL policy produces high-quality expert trajectories suitable for policy training and evaluation.","fun_headline_variants_meta":{"raw":{"variants":["ManiSoft benchmark for vision-language soft arm manipulation","ManiSoft tests vision-language control on soft continuum arms","ManiSoft simulator aids vision-language tasks for deformable robots","ManiSoft generates 6300 trajectories for soft arm vision-language tasks"]},"model":"grok-4.3","cost_usd":0.013311,"raw_usage":{"total_tokens":5808,"prompt_tokens":752,"num_sources_used":0,"completion_tokens":67,"cost_in_usd_ticks":133112000,"prompt_tokens_details":{"text_tokens":752,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":4989,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":752,"tokens_out":67,"duration_ms":48928,"temperature":1.0,"reasoning_tokens":4989,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-05-20T09:50:29.047273+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"If policies trained on ManiSoft trajectories show no performance advantage over rigid-arm methods when tested on physical soft robots in confined spaces, the benchmark would fail to demonstrate a useful bridge between the two domains.","supporting_citations":[],"review_version":1}