{"id":"bf35ee38-2f14-4262-8802-1e1f66b6ea47","arxiv_id":"2607.04234","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"Success-only metrics overstate deformable-manipulation performance; tactile sensing raises Safety Success (e.g. 21.4%→35.6% on Object-Soft) while Goal Success stays comparable.","lead":"SoftVTBench is a simulation benchmark that scores robot policies not only on finishing soft-object pick-and-place, but on doing so without drop or excessive crush. It shows success-only scores hide many unsafe completions, and that touch sensing improves safety more than goal rate.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.5","headline":"VO/VT gripper-action confound is the load-bearing threat to the tactile-safety claim.","rationale":"The reader correctly flags the VO/VT gripper-action confound among several reasons for CONDITIONAL and correctly identifies the κ-calibrated FEM safety definition as a weakest assumption. The more immediate load-bearing threat to the paper’s strongest empirical claim (tactile improves Safety Success) is the action-space mismatch in §4.1: continuous gripper width is the control variable that directly trades off grasp stability against D_peak ≤ τ_o, so giving it only to VT confounds the modality comparison. Metric validity and sim-to-real transfer remain important but secondary; they affect generalizability of “physical safety,” whereas the confound threatens the causal interpretation of the main numbers already reported. A single matched-action re-run settles whether tactile itself drives the Safety Success lift. Verdict stays CONDITIONAL (tighten action controls and statistics); no upgrade or rejection is warranted from this check alone. Agreement with the reader is partial because they list the confound but elevate the offline κ/FEM definition as the primary weakest assumption.","tokens_in":16350,"tokens_out":595,"duration_ms":6852,"concrete_test":"Retrain and re-evaluate both policies under identical continuous gripper-width actions (or both under binary), same LoRA schedule and seeds, on Object-Soft and Spatial-Soft; recompute Table 3 Safety Success and Table 4 P95 deformation. If the VT–VO Safety Success gap shrinks below ~5 absolute points or loses significance, the tactile-safety claim weakens and must be restated as a joint modality+action effect.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim that tactile sensing improves Safety Success (Object-Soft 21.4%→35.6%; Spatial-Soft 32.6%→44.6%) while Goal Success stays comparable rests on comparing π0.5-Vision (binary open/close gripper) to π0.5-Visuo-Tactile (continuous gripper-width action) (§4.1). Continuous closure is exactly the degree of freedom that regulates compression against the FEM-RMS threshold τ_o = κ D_ref_o (κ=0.5, Appendix C.3; Eq. 4). The paper itself notes that continuous width would give the vision-only policy “implicit interaction information” and therefore withholds it—yet that same continuous action is given only to VT. Consequently the Safety Success and deformation gains (Table 3–4) cannot be attributed cleanly to tactile RGB/marker history versus to finer gripper control. The rigid-suite pattern (VT sometimes worse on Goal Success) is consistent with a modality/action confound rather than pure tactile benefit. Without an action-matched ablation the headline causal claim about tactile sensing is under-supported.","agreement_with_reader":"partial"},"referee_report":{"model":"grok-4.5","summary":"SoftVTBench is a simulation benchmark (Isaac Sim / PhysX FEM) for physically constrained deformable-object manipulation with visuo-tactile observations. It defines four matched task suites over object type (soft vs. rigid) and variation axis (object vs. spatial), and reports Goal Success separately from Safety Success, where the latter requires no drop/slip and peak object-size-normalized FEM-RMS deformation below an offline-calibrated object-specific threshold (Eqs. 3–4; §3.4). Using π0.5 LoRA baselines, the paper claims that success-only metrics substantially overstate performance because many goal-completing rollouts violate safety, and that adding tactile RGB and marker-motion history improves Safety Success (e.g., Object-Soft 21.4%→35.6%; Spatial-Soft 32.6%→44.6%) while keeping Goal Success comparable and shifting the deformation distribution downward (Tables 3–4, Fig. 4).","tokens_in":16678,"tokens_out":1347,"duration_ms":18601,"significance":"If the Goal–Safety gap and the tactile-safety benefit hold under fair controls, the work fills a clear gap at the intersection of deformable manipulation, visuo-tactile sensing, and process-level physical safety—areas that existing suites (LIBERO, SoftGym, MoDeSuite, UniVTAC, SoGraB, etc.) address only partially (Table 1). Strengths that should be credited: the matched 2×2 suite design, policy-hidden privileged FEM evaluation, offline per-object calibration of τ_o, fixed-seed closed-loop protocol, deformation distribution analysis (not only threshold rates), and public code/website. These make the benchmark a useful, falsifiable testbed for contact-aware policies even if some baseline claims need tightening.","major_comments":[{"comment":"§4.1 and Tables 3–4: the headline causal claim that “incorporating tactile sensing improves Safety Success … while maintaining comparable Goal Success” is confounded by the gripper action interface. VO uses binary open/close; VT uses continuous gripper-width. Continuous closure is precisely the degree of freedom that regulates compression against τ_o = κ D_ref_o (Eq. 4; Appendix C.3, κ=0.5). The paper withholds continuous width from VO to avoid “implicit interaction information,” yet grants it only to VT. Safety Success and the downward shift in FEM-RMS (Table 4) therefore cannot be attributed cleanly to tactile RGB/marker history versus finer gripper control. An action-matched ablation (binary–binary and continuous–continuous, with and without tactile) is load-bearing for the tactile claim; without it the claim should be restated as a joint modality+action effect.","section":null},{"comment":"§4.1–4.2: all reported policies are LoRA fine-tunes of a single architecture (π0.5) with a fixed training recipe (8×A100, batch 256, 7k steps, action horizon 50, execute 10). The abstract and conclusion phrase results as properties of “policies” and of “tactile sensing” in general. With only one backbone, architecture- or training-specific effects cannot be separated from the observation modality. At minimum, the manuscript should either (i) add one independent baseline family under the same protocol, or (ii) explicitly scope all causal language to π0.5-style VLA policies and treat broader claims as hypotheses for future work.","section":null},{"comment":"§3.4 and Appendix C.3: Safety Success hinges on peak FEM-RMS after rigid-motion removal and on τ_o = κ D_ref_o from a compression sweep (default κ=0.5; sensitivity mentioned for {0.3,0.7} but not shown in the main tables). The Goal–Safety gap is large enough that the qualitative overstatement claim is robust, but ranking of methods and the reported absolute Safety Success rates are threshold-dependent. Main-text results should include the κ sensitivity (or an equivalent continuous deformation metric already partially in Table 4) so that the safety ranking is not tied to a single ad-hoc scale factor.","section":null}],"minor_comments":[{"comment":"Abstract and §1: “nearly 2,000 collected episodes” vs. Table 6 total of 2,000 demos / 200 val episodes—align the wording with the table.","section":null},{"comment":"Table 3: Safety Success for rigid suites is marked “–” in the table but described as reducing to Goal Success ∧ NoDrop in §3.4; either report the NoDrop-conditioned rate or state N/A consistently with Fig. 4.","section":null},{"comment":"§4.1: tactile encoding (8-frame history, 4×4 grid, same visual encoder) is underspecified relative to marker-motion history length and normalization; a short appendix table would aid reproducibility.","section":null},{"comment":"Fig. 3 caption and body: useful qualitative evidence, but no quantitative link (e.g., correlation of marker shear with D(t)) is given; a brief note would strengthen the interpretation.","section":null},{"comment":"Typos / polish: “π0.5” vs. “pi0.5” in the abstract; “V ariation” spacing artifacts in Tables 2 and 6; ensure arXiv author list and affiliation markers match the PDF header.","section":null},{"comment":"Limitations (§5) correctly note sim-to-real and asset diversity; consider also stating that NoDrop includes “workspace escape” and “loss of stable containment,” which may mix kinematic failure modes with contact safety.","section":null}],"recommendation":"major_revision","confidential_remarks":"The Goal–Safety gap finding is the paper’s most solid contribution and is largely independent of the VO/VT confound; I would not reject on that basis. The tactile-safety headline is the part that needs an action-matched control before the abstract can keep its current causal wording. Fit for a robotics journal/benchmark track is good if the revision lands. No integrity concerns; code and protocol look reproducible."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The one thing worth knowing: SoftVTBench is a clean, usable Isaac Sim suite that makes the Goal–Safety gap concrete for volumetric soft objects, and the gap is large. Goal Success on Object-Soft is ~70% while Safety Success is 21–36%; success-only numbers really do hide unsafe completions. That is the contribution that sticks.\n\nWhat is new is the combination, not any single piece. Matched 2×2 suites (object/spatial × soft/rigid), dual-finger RGB + marker tactile, multi-view RGB, language, and policy-hidden FEM peak deformation with offline object-specific thresholds. Privileged NoDrop + D_peak ≤ τ_o is a sensible evaluator design, and they ship code/website. Tables 3–4 and Fig. 4 support the descriptive claim that many goal-complete rollouts still over-compress or drop. Deformation distributions shift down under the VT setting, including the P95 tail. Citation pattern is fair: LIBERO, SoftGym, SoGraB, UniVTAC, ManiFeel, Tabero, DefGraspSim are placed correctly; they are not inventing a literature vacuum.\n\nThe soft spot that matters is the VO/VT comparison. Vision-only gets binary open/close; visuo-tactile gets continuous gripper width. Continuous closure is exactly the DOF that regulates compression against τ_o = κ D_ref_o. The paper withholds continuous width from VO to avoid “implicit interaction information,” then gives it only to VT. So the Safety Success lift (21.4%→35.6%, 32.6%→44.6%) and lower FEM-RMS cannot be cleanly attributed to tactile RGB/marker history versus finer gripper control. Rigid suites are mixed, which fits a confound as easily as a tactile story. Single π0.5 LoRA family, no error bars, κ=0.5 default, and pure sim (PhysX corotational FEM + TacEx) are real limits but secondary; the action mismatch is load-bearing for the causal tactile claim.\n\nWho it is for: people building or evaluating soft-object or visuo-tactile policies who need a reproducible safety metric beyond terminal success. Not a dynamics paper and not a real-world transfer result. I would engage with the suite and the gap numbers; I would not yet cite the tactile-safety causal claim without an action-matched ablation. Worth a serious referee—accept-shaped if they fix the action control and add statistics.","headline":"Useful safety-aware soft-object benchmark with a real Goal–Safety gap, but the tactile claim is confounded by continuous vs binary gripper actions.","tokens_in":17344,"tokens_out":612,"would_cite":true,"duration_ms":6641,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"Success-only scores hide unsafe grasps of soft objects; tactile feedback raises safe completion and lowers deformation.","keywords":["deformable object manipulation","visuo-tactile sensing","safety-aware evaluation","SoftVTBench","FEM deformation","Goal Success","Safety Success","robotic grasping"],"falsifier":"Run the same vision-only and visuo-tactile policies on SoftVTBench deformable suites and check whether Safety Success remains far below Goal Success and whether tactile feedback still raises Safety Success and lowers the FEM deformation distribution under the reported protocol.","tokens_in":17229,"feed_emoji":"🦑","tokens_out":629,"duration_ms":7941,"temperature":0.7,"pith_summary":"SoftVTBench argues that finishing a pick-and-place task on a deformable object is not enough: the robot must also keep a stable grasp without drop or slip and keep peak object deformation under an object-specific limit. Existing robot benchmarks mostly score only whether the object ends up in the right place, so they can treat over-squeezed or barely-held rollouts as full successes. SoftVTBench is a closed-loop Isaac Sim suite with finite-element soft bodies, multi-view RGB, dual-finger tactile RGB and marker motion, proprioception, and language instructions. It reports Goal Success and a stricter Safety Success that uses hidden FEM states to reject drops and over-deformation. Matched rigid and soft suites separate ordinary manipulation skill from safe soft-body contact. Baseline policies show that many goal-complete episodes still fail safety, and that adding tactile sensing lifts Safety Success while Goal Success stays comparable and continuous deformation falls.","feed_headline":"Success metrics hide unsafe soft grasps; touch cuts damage","feed_subtitle":"A new sim benchmark shows tactile sensing raises safe completion while goal scores barely move.","key_machinery":"Safety Success: Goal Success and no drop and peak object-size-normalized FEM-RMS deformation D_peak ≤ τ_o, where τ_o is an offline-calibrated object-specific threshold from a compression sweep (default κ = 0.5 times the max stable reference deformation), measured only from policy-hidden privileged FEM states.","core_discovery":"On deformable grasp-and-place tasks, success-only evaluation substantially overstates policy quality because a large share of goal-completing rollouts still drop, slip, or over-deform the object. Under the same protocol, a visuo-tactile policy raises Safety Success relative to a vision-only policy (Object-Soft 21.4% to 35.6%; Spatial-Soft 32.6% to 44.6%) while Goal Success stays comparable or improves modestly, and the full FEM deformation distribution shifts downward.","pith_inferences":[],"forward_implications":[],"fun_headline_variants":["Success scores mask unsafe soft grasps; touch raises Safety Success","Many goal-completing soft rollouts still drop, slip or over-deform","Tactile input lifts Safety Success on deformables as Goal Success holds","SoftVTBench: success-only eval overstates physically safe soft manip","Vision-only finishes soft tasks yet leaves high FEM deformation risk"],"cache_read_input_tokens":128,"weakest_assumption_plain":"That a one-time offline compression-sweep threshold on simulated peak FEM deformation is a faithful stand-in for real physical safety of soft objects under closed-loop robot contact.","fun_headline_variants_meta":{"raw":{"variants":["Success scores mask unsafe soft grasps; touch raises Safety Success","Many goal-completing soft rollouts still drop, slip or over-deform","Tactile input lifts Safety Success on deformables as Goal Success holds","SoftVTBench: success-only eval overstates physically safe soft manip","Vision-only finishes soft tasks yet leaves high FEM deformation risk"]},"model":"grok-4.5","effort":"low","cost_usd":0.004326,"raw_usage":{"total_tokens":1357,"prompt_tokens":858,"num_sources_used":0,"completion_tokens":97,"cost_in_usd_ticks":43260000,"prompt_tokens_details":{"text_tokens":858,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":402,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":858,"tokens_out":97,"duration_ms":5086,"temperature":1.0,"reasoning_tokens":402,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-11T20:45:09.152050+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Run the same vision-only and visuo-tactile policies on SoftVTBench deformable suites and check whether Safety Success remains far below Goal Success and whether tactile feedback still raises Safety Success and lowers the FEM deformation distribution under the reported protocol.","supporting_citations":[],"review_version":1}