{"id":"36ce481a-a85e-4d7b-8ce1-a62f90dfc152","arxiv_id":"2506.18212","paper_version":3,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"A haptic-augmented ACT controller with a soft gripper reaches 80% success on a pseudo-oocyte pick-and-place task, versus 50% for standard ACT.","lead":"This paper adds touch sensing and automatic retry to an imitation-learning robot controller, and tests it on picking up pomegranate seeds that stand in for delicate cells. The authors report higher success than the standard controller, but the tests are small and no code or data is released.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The headline advantage is a 3-trial difference: 8/10 vs 5/10 gives Fisher exact p≈0.35, so the claimed haptic benefit is not statistically distinguishable from chance at the reported sample size.","rationale":"The reader's weakest assumption focused on force-sensor reliability and redundancy; I agree that the paper does not directly test whether the force signal is informative beyond vision, and the observed false-positive retry loops on larger objects (almonds, frozen blueberries) show the signal is not calibration-robust. However, the single most load-bearing issue is more basic: the primary quantitative evidence for the method is a 3-out-of-10 trial difference. With n=10 per cell, a Fisher exact test gives p≈0.35, so the central claim is statistically indistinguishable from chance even before considering sensor reliability or confounds. The 2x2 design (haptic x recovery) is a genuine strength, the soft-gripper evidence is independently useful, and the failure analysis is honest, so I would not reject the paper. But the reported sample size is too small to establish the headline improvement, which supports keeping the reader's CONDITIONAL verdict unchanged pending a larger, pre-registered replication.","tokens_in":7752,"tokens_out":5378,"duration_ms":69317,"concrete_test":"Re-run the known-environment comparison with a pre-registered protocol: at least 30 trials per condition, the same 40-success/10-recovery demonstration split, randomized seed positions and lighting, and a fixed trial-time limit. Report per-trial outcomes and compute an exact binomial confidence interval for the success-rate difference plus a two-sided Fisher exact test. If the 80%-vs-50% gap does not reach p<0.05 or the confidence interval includes zero, the central haptic-benefit claim is not supported at the reported level.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim depends on Table I, where the known-environment comparison is 80% vs 50% with only 10 trials per cell. A Fisher exact test on the 8/10 vs 5/10 counts gives p≈0.35, and the 95% binomial confidence intervals overlap substantially (80% CI ≈ 44–97%; 50% CI ≈ 19–81%). The no-recovery comparison (40% vs 20%, i.e., 4/10 vs 2/10) is even weaker. The paper reports no confidence intervals, no hypothesis tests, and no per-seed replication, so the 30-percentage-point gap could be produced by one or two outcomes changing. The causal story—force input enables failure detection and retry—therefore rests on an effect that is not yet established as real. The unknown-environment results in Table III are also per-object n=10 and mostly low (coffee 40%, almond 20%), so they cannot rescue the claim. This is not a criticism of the experimental design or the soft gripper; it is a statement that the reported sample size cannot support the headline success-rate improvement.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes Haptic-Informed ACT, an imitation-learning policy that augments Action Chunking with Transformers (ACT) with 3-axis gripper force feedback, recovery demonstrations, and a 3D-printed TPU soft gripper, applied to a pseudo-oocyte (pomegranate-seed) pick-and-place task. The authors report that Haptic-Informed ACT outperforms standard ACT in a known environment (Table I: 80% vs 50% success with recovery samples; 40% vs 20% without) and that the system can pick and deliver several novel pseudo-oocyte materials (Table III), while never crushing a target. The paper's central claim is that the haptic modality enables grasp-failure detection and retry, improving robustness in contact-rich fragile-object manipulation.","tokens_in":7988,"tokens_out":6097,"duration_ms":74768,"significance":"If the claimed effect is real, the contribution is a useful, practical integration of haptic feedback, recovery-informed training, and a soft gripper for fine manipulation, and it would support the value of multimodal imitation learning in biomedical automation. The paper is commendable for building a complete physical system with a clear task design, a soft gripper that prevents crushing, and a concrete comparison to an ACT baseline. However, the experimental evidence in its current form is preliminary: the headline success-rate differences rest on 10 trials per condition with no confidence intervals or significance tests, the compared systems differ in retry policy as well as input modality, and the generalization results are not compared against the baseline. The central claim is therefore defensible but not yet established to the standard expected of a journal publication.","major_comments":[{"comment":"The headline improvement is not statistically supported. With 10 trials per condition, the 8/10 vs 5/10 comparison yields a two-tailed Fisher exact p of approximately 0.35, and the 4/10 vs 2/10 comparison yields p of approximately 0.63; the 95% binomial confidence intervals overlap substantially. A change of one or two outcomes could erase the reported 30-percentage-point gap. The paper reports no confidence intervals, no per-seed replications, and no hypothesis tests. This is load-bearing because the abstract and conclusion claim that Haptic-Informed ACT 'significantly' improves success, and Table I is the only direct evidence for that claim in the known environment.","section":"IV-B (Table I)"},{"comment":"The comparison between Haptic-Informed ACT and ACT is confounded by a difference in the deployed retry policy. The text states that Haptic-Informed ACT 'continued attempting to pick the seed until successful,' while the ACT model 'simply attempted to pick the seed and, regardless of success or failure, proceeded to the test tube.' Thus the higher success rate could be due to the extra attempts rather than to the haptic input itself. The sentence 'This proves that haptic feedback is essential' is therefore an overclaim. The paper should either evaluate both methods under an identical retry protocol (e.g., giving ACT a fixed number of retries) or ablate the haptic signal in the proposed architecture, for example by masking the force sensor input or replacing it with a vision-based grasp detector, to isolate the contribution of haptics.","section":"IV-B (Table I)"},{"comment":"The proposed failure-detection mechanism rests on the 3-axis force sensor reliably distinguishing a successful grasp from a failed one, but this is not validated quantitatively. Figure 6 shows an example of the z-axis force trace, but the paper provides no data on the distribution of force readings across successful and failed grasps, no threshold analysis, and no measure of classification accuracy. The reader's weakest-assumption concern is therefore central: if the force signal is noisy, mis-calibrated, or redundant with visual information, the retry behavior may not be driven by haptic feedback at all. This should be addressed with sensor statistics or a direct ablation.","section":"III and Fig. 6"},{"comment":"The unknown-environment generalization claim is not supported by the reported data. The experiments use 10 trials per object, with no confidence intervals and no significance tests, and the success rates are low (e.g., coffee bean 40%, almond 20%). More importantly, there is no ACT baseline in this setting, so the paper cannot show that Haptic-Informed ACT generalizes better than the baseline to new objects. In addition, the sentence 'In all experiments, the robot successfully picked and delivered the seeds' appears inconsistent with Table III, which reports many failures; this wording should be clarified or corrected.","section":"IV-C (Table III)"}],"minor_comments":[{"comment":"The claims that the method 'significantly' improves success and 'proved to be able to successfully execute the task outperforming ACT' should be softened to reflect the absence of statistical tests and the confounding in the comparison.","section":"Abstract and Conclusion"},{"comment":"The sentence 'This proves that haptic feedback is essential for executing fine manipulation tasks' (Section IV-B) should be rephrased as a suggestion or hypothesis, since the controlled evidence is not sufficient for a proof.","section":"II-B"},{"comment":"Figure 6 would be much more informative with labeled axes, units, a time axis, and an explicit indication of the grasp-success threshold in the force trace.","section":"Fig. 6"},{"comment":"The notation 'T a p' in the text describing the action sequence predicted by the transformer is undefined and appears to be a typo; please define it or remove it.","section":"III"},{"comment":"In the conclusion, 'action chucking with transformers' should read 'action chunking with transformers.'","section":"V"}],"recommendation":"major_revision","confidential_remarks":"The paper is likely below the current threshold for a full journal article because the key empirical claims lack statistical support and the main comparison is confounded. However, the proposed system is plausible and the experimental setup is clearly described, so the weaknesses are addressable with additional trials, a controlled retry protocol, and a direct ablation of the haptic input. I would also encourage the authors to release code and data, as no reproducibility artifacts are mentioned."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a nice systems paper, but the headline 30-point success gap is not statistically distinguishable from chance at their sample size. The stress-test note is correct: 8/10 vs 5/10 gives Fisher exact p≈0.35. That doesn't kill the paper, but it means the main claim should be framed as preliminary.\n\nWhat's new: the combination of ACT with a gripper-mounted 3-axis force sensor and recovery demonstrations for a pseudo-oocyte pick-and-place task is genuinely new, as far as I know. The soft TPU gripper is a practical contribution, and the task is a sensible proxy for oocyte transfer. The most compelling evidence is the behavioral difference: ACT with recovery data didn't actually retry, while Haptic-Informed ACT did. That supports the mechanistic story that haptic feedback is what enables failure detection. The paper is also honest about limitations, e.g., the force sensor misinterpreting larger objects as failures, and the almond shape problem.\n\nThe soft spots are around experimental rigor. Ten trials per cell is thin; no error bars, no significance tests, no per-seed runs. The two compared systems differ in both haptic input and the resulting retry behavior, so you can't cleanly isolate the contribution of haptic sensing. The unknown-environment results are mostly low (coffee 40%, almond 20%), which undercuts the generalization claim. No code or data released, which is a practical hurdle. These are all fixable with more experiments and a better analysis, not fundamental flaws.\n\nWho this is for: people working on haptic imitation learning, soft grippers, or biomedical manipulation. I'd send it to a serious referee (e.g., ICRA/IROS) but with the expectation of major revisions. My recommendation: engage with it, but don't accept the success-rate numbers at face value.","headline":"A plausible, honest systems paper whose headline success-rate advantage is not statistically significant at the reported trial counts; worth a serious look but needs more rigorous evaluation.","tokens_in":8514,"tokens_out":2413,"would_cite":false,"duration_ms":27282,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Adding gripper force feedback to ACT doubles success in delicate pick-and-place.","keywords":["haptic feedback","imitation learning","Action Chunking with Transformers","oocyte manipulation","soft gripper","grasp failure recovery","robot manipulation","multimodal learning"],"falsifier":"Train the same Haptic-Informed ACT architecture on the same recovery dataset but feed it only images and proprioception; if success in the trained environment stays at 80% instead of dropping toward 50%, the force channel is not the cause of the improvement. A cheaper check is to record force traces from all successful and failed grasps and test whether the z-axis force at gripper closure cleanly separates the two classes.","tokens_in":7588,"feed_emoji":"🤖","tokens_out":4706,"duration_ms":47993,"temperature":0.7,"pith_summary":"This paper argues that a standard vision-based imitation policy, ACT, fails at delicate oocyte transfer because it cannot tell whether the gripper actually closed on the target. The authors extend ACT with a 3-axis force sensor mounted at the gripper and train the policy on demonstrations that include failed grasps followed by recovery attempts. In a pseudo-oocyte transfer task using a pomegranate seed, the resulting Haptic-Informed ACT succeeds in 80% of trials compared with 50% for ACT trained with the same recovery data, and 40% versus 20% without recovery data. The paper also shows the policy transfers to objects of different size, shape, and color, and that a soft TPU gripper prevents the seeds from being crushed. If correct, the result is evidence that contact-force feedback, not just vision, is the missing signal for dependable fine manipulation.","feed_headline":"Gripper force feedback doubles pick-and-place success","feed_subtitle":"A haptic-aware ACT policy reaches 80% success on fragile seed transfer, versus 50% with vision alone.","key_machinery":"The load-bearing mechanism is the addition of a WACOH Dyn Pick MCF-3 3-axis force sensor at the end effector, whose readings enter the ACT policy alongside camera images and joint positions. ACT is an action-chunking transformer: a conditional variational autoencoder captures the variability of human demonstrations and predicts a sequence of future joint positions, reducing compounding error. The force channel lets the policy distinguish 'closed on the seed' from 'closed on nothing,' while the recovery demonstrations teach the retry behavior. A 3D-printed TPU soft gripper completes the system by deforming around the target so a fully closed gripper does not crush it, and the soft fingers double as a mechanical safety margin during haptic-based grasp detection.","core_discovery":"The central discovery is that adding the gripper's 3-axis force readings as an input channel to ACT gives the policy a reliable grasp-success signal, letting it detect a failed pick and automatically retry instead of blindly proceeding to the delivery tube. The paper demonstrates this in a pseudo-Xenopus-oocyte transfer task: Haptic-Informed ACT reaches 80% success with recovery demonstrations and 40% without, versus 50% and 20% for the visual-only baseline. It also reports that trained models generalize to seven novel pseudo-oocyte materials, with worst-case success on almonds (20%) caused by the object's elongated shape pushing it out of the gripper. Throughout all trials the soft gripper never crushed a seed, which the authors attribute to the TPU fingers bending outward rather than compressing the target.","pith_inferences":["Beyond the paper's stated results, the same force-plus-recovery recipe should transfer to other contact-rich pick-and-place tasks where visual occlusion hides grasp status, such as surgical needle handling or deformable food packing.","A natural testable extension is to replace the learned retry behavior with an explicit force threshold for grasp success; if the threshold version matches 80% success, the haptic channel's contribution is largely the signal itself rather than the recovery demonstrations.","The observed false-failure loops on large objects imply that force readings should be normalized by gripper aperture or object size, or that the training distribution should include larger targets; the authors leave this as future work.","Because the ablation without recovery data still shows a 20-point gain from haptics, the force channel and recovery data appear to contribute independently, though the paper does not run a full factorial ablation to prove it."],"forward_implications":["Haptic-Informed ACT can detect grasp failure in real time and retry until the object is secured, a behavior the visual-only baseline does not exhibit.","Adding recovery demonstrations raises success for both ACT and Haptic-Informed ACT, so failure data is a reusable training resource rather than a contaminant.","The policy generalizes to objects of different size, color, and shape, with the soft gripper absorbing size variation that would break a rigid gripper.","Force magnitude that lies outside the training distribution can be misread as failure, causing retry loops on larger objects such as almonds and frozen blueberries.","The system requires no explicit vision-based grasp verification; the learned policy uses force implicitly to decide when to move on."],"supporting_citations":[{"why":"Supplies the base ACT architecture and the action-chunking paradigm that Haptic-Informed ACT extends with force input and recovery data.","marker":"[9]"},{"why":"Shows that haptic feedback teleoperation can be combined with ACT-style learning for bimanual fine manipulation, motivating the haptic channel.","marker":"[18]"},{"why":"Demonstrates using force information in an ACT-based bilateral control policy, the direct precedent for feeding forces into the transformer.","marker":"[20]"},{"why":"Provides evidence that learning from failed demonstrations is viable, underpinning the recovery-informed training data.","marker":"[23]"},{"why":"Shows error examples from teleoperation can improve mobile manipulation, supporting the inclusion of recovery demonstrations.","marker":"[24]"},{"why":"Gives the action-chunking-as-policy-compression rationale behind ACT's multi-step prediction.","marker":"[25]"},{"why":"Provides the transformer mechanism that ACT uses to model observed features and action sequences.","marker":"[26]"}],"fun_headline_variants":["Haptic ACT with recovery training lifts seed transfer to 80% success","Force-sensing gripper helps ACT auto-retry, hitting 80% on delicate picks","Soft TPU gripper never crushes seeds; haptic ACT retries failed grasps","Haptic ACT detects failed grasps and recovers, boosting success from 50% to 80%"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The method's advantage rests on the assumption that the 3-axis force sensor at the gripper gives a reliable, non-redundant signal for whether the object was actually grasped; the paper only compares the full system against a visual-only baseline, so a noisy or redundant force channel would make the reported gain shrink or disappear.","fun_headline_variants_meta":{"raw":{"variants":["Haptic ACT with recovery training lifts seed transfer to 80% success","Force-sensing gripper helps ACT auto-retry, hitting 80% on delicate picks","Soft TPU gripper never crushes seeds; haptic ACT retries failed grasps","Haptic ACT detects failed grasps and recovers, boosting success from 50% to 80%"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000999,"raw_usage":{"total_tokens":4172,"prompt_tokens":833,"completion_tokens":3339,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":449,"completion_tokens_details":{"reasoning_tokens":3245}},"tokens_in":449,"tokens_out":3339,"duration_ms":27494,"temperature":1.0,"reasoning_tokens":3245,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T23:22:06.765870+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Train the same Haptic-Informed ACT architecture on the same recovery dataset but feed it only images and proprioception; if success in the trained environment stays at 80% instead of dropping toward 50%, the force channel is not the cause of the improvement. A cheaper check is to record force traces from all successful and failed grasps and test whether the z-axis force at gripper closure cleanly separates the two classes.","supporting_citations":[{"cited_title":"Learning variable compliance control from a few demonstrations for bimanual robot with haptic feedback teleoperation system,","cited_arxiv_id":null,"evidence_quote":"Shows that haptic feedback teleoperation can be combined with ACT-style learning for bimanual fine manipulation, motivating the haptic channel."},{"cited_title":"Bi-ACT: Bilateral Control-Based Imitation Learning via Action Chunking with Transformer","cited_arxiv_id":"2401.17698","evidence_quote":"Demonstrates using force information in an ACT-based bilateral control policy, the direct precedent for feeding forces into the transformer."},{"cited_title":"Donut as i do: Learning from failed demonstrations,","cited_arxiv_id":null,"evidence_quote":"Provides evidence that learning from failed demonstrations is viable, underpinning the recovery-informed training data."},{"cited_title":"Error-aware imitation learning from teleopera- tion data for mobile manipulation,","cited_arxiv_id":null,"evidence_quote":"Shows error examples from teleoperation can improve mobile manipulation, supporting the inclusion of recovery demonstrations."},{"cited_title":"Action chunking as policy compression,","cited_arxiv_id":null,"evidence_quote":"Gives the action-chunking-as-policy-compression rationale behind ACT's multi-step prediction."}],"review_version":1}