{"id":"4f80611d-83c6-4649-9d8b-1fee3b7b301c","arxiv_id":"2607.27261","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"Pairing clean and nuisance observations to measure action drift lets CFNBC select 20–30 counterfactual repair examples that outperform matched random selection for robust imitation.","lead":"A new data-selection method, CFNBC, uses simulated visual changes to measure how much a robot's imitation policy changes its actions, then picks the few example situations that best fix its weak spots. In two simulated manipulation tasks, this targeted repair beats random data at the same budget and sometimes matches hundreds of random examples with only 20–30.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Central data-efficiency claim rests on n=3 point estimates with no error bars and is weaker on held-out nuisances; without CIs or released data, the claimed margin over random may not be robust.","rationale":"The reader's verdict is CONDITIONAL, and my concern is consistent with that: the central empirical claim is plausible and well-specified, but the evidence is not yet strong enough for full acceptance. I did not select Eq. (1)'s task-preserving assumption as the primary concern because, for purely visual re-rendering nuisances, the underlying state and expert action are preserved by construction; the more fragile link is the empirical assertion that 20–30 selected candidates are reliably better than random and close to much larger random budgets. The paper is transparent about n=3 and partial held-out transfer, which is creditworthy, but those admissions also mean the headline claims are not yet quantitatively secure. A 20-seed rerun with CIs would settle whether the reported margins are real. If the margins persist, ACCEPT would be justified; if they shrink, CONDITIONAL remains appropriate until stronger evidence is provided.","tokens_in":12592,"tokens_out":6353,"duration_ms":230939,"concrete_test":"Re-run the full CFNBC pipeline with at least 20 independent seeds per condition and report 95% CIs for the Table II and Table IV success metrics. If the Random K=20/30 CI overlaps the Response-guided point estimate, the matched-budget advantage is not established. Report the held-out mean/worst CIs separately; if the high-budget random margin remains large (e.g., ≥0.2), the abstract's 'approaching' claim should be explicitly scoped to seen nuisance conditions.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The load-bearing claim is the quantitative data-efficiency result: K=20–30 response-guided candidates substantially outperform matched-budget random selection and approach 500-sample random repair. The reported support is three-seed point estimates with no variance, no significance test, and no code/data; the paper itself disclaims significance in Section IV. If run-to-run variance is typical for ACT fine-tuning, the Table II margins (e.g., Cube transfer all-nuisance 0.96 vs 0.56 for random K=20) could be partly seed noise rather than a stable selection effect. More importantly, the 'approaching high-budget random' half of the claim is only true on the seen 22-condition evaluation. On held-out nuisance instantiations (Table IV), response-guided K=20 reaches 0.57 held-out mean for Cube transfer vs 0.90 for random K=500, and Cube stacking held-out worst is 0.02 vs 0.28. The method still beats matched-budget random on held-out, but the headline that small selected sets approach much larger random budgets is not supported for the generalization that matters. This does not invalidate the paper; it narrows the central claim to repaired candidate-covered conditions. Eq. (1) is less suspect because nuisances are constructed by scene re-rendering with unchanged state, though 'local support' should be checked for physical side effects.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes Counterfactual Nuisance Behaviour Cloning (CFNBC), an offline data-selection framework for repairing visuomotor imitation policies that are brittle to task-preserving visual nuisances. CFNBC generates paired clean/nuisance observations under an action-preserving assumption, measures the change in the policy's predicted action ('action drift'), and selects a small, response-diverse repair set by greedily maximizing weighted response coverage. The selected counterfactual samples, labelled with inherited expert actions, are mixed with clean demonstrations to fine-tune the nominal policy. Experiments on MuJoCo bimanual cube transfer and SimplerEnv cube stacking report that drift strongly ranks nuisance-induced failure and that K=20--30 selected candidates outperform matched-budget random selection, sometimes approaching the performance of much larger random repair budgets. The core idea is clearly presented and the offline, rollout-free selection signal is appealing.","tokens_in":12924,"tokens_out":2989,"duration_ms":50161,"significance":"If the central data-efficiency claim holds, CFNBC would be a useful contribution to the growing literature on data-centric robustness for imitation learning: it turns a policy-specific sensitivity measure into a practical data-selection rule that does not require online rollouts or success labels. The paper explicitly builds on established ideas (counterfactual/action-preserving augmentation, active imitation, distribution-shift benchmarks) and contributes a well-specified selection objective. The reported correlations (Spearman rho=0.95/0.88) are strong and the controlled comparisons between selection strategies are well motivated. However, the significance is tempered by the empirical evidence base: the main quantitative claims rest on three-seed point estimates with no error bars, and the headline 'approaching much larger random budgets' is not supported on held-out nuisance instantiations. The authors are candid about these limitations, but the abstract and introduction present the stronger reading.","major_comments":[{"comment":"The central data-efficiency claim—that K=20--30 response-guided candidates substantially outperform matched-budget random selection—is supported by three-seed averages with no variance, no error bars, and no significance test. The paper itself states: 'we use these means to reduce dependence on a single run, but do not claim statistical significance with n=3.' Since the margins in Table II (e.g., Cube transfer All nuis. 0.96 vs 0.56 for random K=20; Cube stacking 0.76 vs 0.60 for random K=30) could plausibly change under typical seed variance for ACT fine-tuning, the quantitative strength of the central claim is not yet established. Please provide per-seed results, confidence intervals, or a larger number of seeds, or explicitly downgrade the claim to a preliminary finding.","section":"§IV (Training seeds and uncertainty) and Table II"},{"comment":"The headline claim that selected repair sets 'approach the performance of much larger random repair budgets' is only true on the seen 22-condition evaluation. On held-out nuisance instantiations, Table IV shows a large gap: for Cube transfer, response-guided K=20 attains held-out mean 0.57 vs 0.90 for random K=500; for Cube stacking, held-out worst is 0.02 vs 0.28. The paper acknowledges this in §V-D ('CFNBC is only as good as the candidate response set' and 'held-out transfer is partial'), but the abstract and Section I state the 'approaching larger random budgets' conclusion without this caveat. The central claim should be narrowed to candidate-covered nuisance conditions, or additional evidence is needed that transfer improves with response-guided selection beyond matched budgets.","section":"§V-D and Appendix D (Table IV)"},{"comment":"The load-bearing premise is that all generated nuisances are task-preserving, i.e., a*(s_c)=a*(s_n) for every paired observation. Appendix B asserts this by construction, but for local support changes (a cloth patch under the manipulated objects) there is a real risk that contact geometry or friction changes the feasible or intended action, especially in the bimanual cube transfer task. The paper does not validate this assumption, e.g., by checking that the expert action succeeds under the nuisance condition or that the task state remains unchanged. If any candidate intervention changes the intended action without detection, action drift conflates nuisance sensitivity with task-relevant change and the repair set is mislabelled. Please add an explicit validation protocol or at least an ablation showing results are robust to filtering out candidates with large physical-side-effect risk.","section":"§III-A, Eq. (1), and Appendix B"},{"comment":"The selection objective in Eq. (4) uses response features and an RBF affinity with free parameters (lambda_cf, kernel bandwidth sigma), and the reported drift-weighted score additionally multiplies by normalized drift. The greedy selection procedure is clear, but the choice of summary statistics (mean/std/mean-abs/max-abs) and the median-distance sigma are presented as implementation details without sensitivity analysis. Since the entire method is an offline selection rule, it is important to know how robust the selection is to these choices. A sensitivity study (e.g., varying sigma and the response-feature specification) would strengthen the claim that the gains come from response-guided coverage rather than from incidental properties of the particular affinity function.","section":"§III-C and Appendix A"}],"minor_comments":[{"comment":"The Spearman correlations are computed over nuisance conditions that are also used to build the candidate response set and to evaluate the main repair results. Please clarify the relationship between the plotted conditions and the candidate pool, and state whether the correlation includes all 22 conditions or a subset. This does not invalidate the signal, but it affects how 'offline' and 'prediction' should be interpreted.","section":"Figure 3 and §V-B"},{"comment":"The phrase 'without requiring rollout success labels or online policy execution' is accurate for the selection stage, but the candidate pool is still generated in a simulator with a task-preserving nuisance generator. Consider adding a sentence clarifying that the method is offline with respect to policy rollouts, not necessarily with respect to simulation access.","section":"Abstract and Section I"},{"comment":"The 'Gain vs nominal' column for Random (high-budget) 500 in Cube transfer reports 0.70, while All nuis. is 1.00 and nominal is 0.30. This is consistent, but the column label could be misread as the gain of the high-budget random method over the low-budget random method. Please rename to 'Gain vs nominal (All nuis.)'.","section":"Table II"},{"comment":"The 'Held-out gap' column is defined as the difference between seen and held-out all-nuisance performance, but the column header does not make clear which direction is positive. Please add a footnote or caption explaining the sign convention.","section":"Appendix D, Table IV"}],"recommendation":"major_revision","confidential_remarks":"The paper is a good fit for the robotics/imitation-learning venue and the core idea is worth publishing if the empirical support is strengthened or the claims are appropriately narrowed. The authors are unusually honest about limitations, and the main missing pieces are statistical error bars and a clearer separation between candidate-covered and held-out claims. I would not reject, but the current version overstates the data-efficiency result relative to what Table IV shows. No code/data release is mentioned; given the reliance on three-seed averages, releasing the trained policies and evaluation scripts would materially increase confidence."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Dear colleague,\n\nThe paper to know about: CFNBC selects a small repair set for a brittle visuomotor policy by measuring action drift on paired clean/nuisance counterfactuals, then covering diverse high-drift response modes. The genuinely new bit is selecting data by the policy's induced response rather than by environmental difficulty or visual diversity. The action-drift signal is a sensible offline proxy, and the paper is admirably clear about its limitations.\n\nWhat it does well: the method is fully specified, the drift signal correlates strongly with failure (ρ=0.95 and 0.88), and the low-budget results on the 22 seen nuisance conditions are striking: K=20 response-guided repair gets 0.96 all-nuisance success in cube transfer versus 0.56 for random, and cube stacking 0.76 versus 0.60. The paper explicitly disclaims significance at n=3 and acknowledges the candidate-set dependence and the task-preserving assumption. That honesty is real.\n\nThe soft spots are in proportion. First, the headline 'approaching much larger random budgets' only holds on the seen nuisance conditions. On held-out nuisances not used in candidate generation, response-guided K=20 reaches 0.57 mean in cube transfer while random K=500 reaches 0.90, and cube stacking held-out worst-case is 0.02 versus 0.28. The matched-budget advantage mostly survives—cube transfer 0.57 vs 0.11, cube stacking 0.48 vs 0.42—so the narrow claim is intact, but the abstract overstates the generalization. Second, there are no error bars or significance tests; three seeds is not enough to be confident in the margins, especially for the cube-stacking comparisons. Third, no code or data is released, which makes the empirical claims hard to verify. These are fixable in revision: add more seeds or confidence intervals, report held-out more prominently, and release artifacts. The task-preserving assumption is reasonably well defended by construction—nuisances are re-renderings with unchanged state—but local support changes could plausibly alter contact dynamics, and the paper's own caveat about delayed failures and multimodal policies is worth taking seriously.\n\nWho this is for: anyone working on data-efficient imitation or robustness repair for visuomotor policies. It is a solid within-subfield contribution, not a paradigm shift. I would send it to review; it deserves referee time, but I would expect the authors to narrow the central claim and provide statistical support.\n\nRecommendation: engage with it, but calibrate your expectations to the seen-condition regime.","headline":"A well-specified, honest paper whose data-efficiency result is solid on candidate-covered conditions but does not survive contact with held-out generalization as cleanly as the abstract implies.","tokens_in":13373,"tokens_out":2509,"would_cite":true,"duration_ms":63607,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that the most useful repair examples for a brittle visuomotor policy are the ones exposing its fragile action responses, and that these can be found offline by measuring action drift under task-preserving counterfactual nui","keywords":["imitation learning","visuomotor policy","counterfactual data selection","action drift","robustness repair","behavior cloning","nuisance variation","data efficiency"],"falsifier":"Construct a set of nuisances that deliberately changes the intended goal, violating the equality a*(s_c)=a*(s_n), while keeping the same visual shift magnitudes; if drift ranks these task-changing shifts as high-priority repairs, the signal conflates task change with nuisance fragility. Also, evaluate a held-out nuisance family absent from the candidate pool: if response-guided repair with a budget of 30 cannot beat random selection at the same budget on that condition despite the pool containing useful candidates, the coverage objective has failed.","tokens_in":12467,"feed_emoji":"🤖","tokens_out":5665,"duration_ms":48673,"temperature":0.7,"pith_summary":"The paper sets out to answer a practical question: when a robot policy works in the lab but breaks under minor visual changes, which additional demonstrations will actually fix it? Its answer is an offline audit: generate pairs of clean and nuisance observations that leave the expert action unchanged, record how much the policy's predicted action moves between the pair (action drift), and use that signal to select a small, response-diverse repair set. In the two simulated manipulation tasks studied, fine-tuning on 20-30 selected candidates lifted all-nuisance success by 0.66 and 0.44, while the same budget of random examples gave 0.26 and 0.28, and hundreds of random examples were needed to approach the selected set's performance. If correct, the result reframes robustness as a coverage problem over a policy's failure modes rather than a data-volume problem.","feed_headline":"20 selected demos beat 500 random demos for robot robustness","feed_subtitle":"Offline action-drift audit finds a policy's fragile visual responses, cutting repair data ~25x in two simulated tasks.","key_machinery":"The central object is the paired counterfactual observation (clean scene paired with nuisance scene) constructed to preserve the expert action, together with the action-drift vector, the normalized difference between the policy's predicted actions on the pair. Drift magnitude ranks candidate nuisances; drift response features, summary statistics of that difference over time and action dimensions, define distinct response modes. A kernel affinity over those features feeds a greedy coverage objective that picks a budget-sized repair set spanning diverse high-drift modes. The selected nuisance observations are paired with the original demonstration actions and used to fine-tune the policy from","core_discovery":"The central claim is that a brittle visuomotor policy can serve as its own probe: comparing its predicted actions on paired clean and nuisance observations, where the expert action is identical by construction, yields an offline fragility signal that ranks visual conditions by how much they break the policy. The paper calls this signal action drift and shows that it correlates with rollout failure across nuisance conditions, with rank correlations of 0.95 and 0.88 on the two tasks. It then argues that selecting repairs by drift score alone is suboptimal because high-drift examples can be redundant, and instead selects candidates that cover diverse response shapes. The reported numbers are th","pith_inferences":["A testable extension is to average drift over multiple action samples or output distributions for stochastic or diffusion policies; single-sample drift may understate fragility when the policy is multimodal.","The same signal could be inverted into a live data-acquisition policy: during a distribution shift, redirect data collection toward the highest-drift scenarios rather than re-collecting random demonstrations.","The coverage principle suggests that the best repair set is not intrinsic to the data but depends on the current policy, so the same candidate pool would yield different selections for different checkpoints, which is a checkable prediction.","Pairing drift with rollout-derived or ensemble-disagreement signals might overcome the candidate-coverage limit that the paper explicitly acknowledges."],"forward_implications":["Robustness repair can be planned offline, before deployment, by auditing the trained policy against task-preserving counterfactuals.","At small repair budgets, response-diverse selection is more effective than the same number of random examples, and can match the effect of much larger random budgets.","Action drift is a policy-specific fragility signal: the same nuisance condition that is benign for one policy can be destructive for another, so repair data should be chosen per policy rather than per task.","Held-out transfer is partial and limited by candidate-pool coverage; large random budgets still transfer best, identifying candidate coverage as the main bottleneck."],"fun_headline_variants":["Action drift finds 25x fewer demos for robot repair","Your robot's own predictions pick its best fix data","Offline sensitivity audit cuts robot repair demos 25x","Use the policy's weak spots to choose 20 demos that work","25x less data for robot robustness via action drift"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The load-bearing premise is that every nuisance intervention really is task-preserving, meaning the demonstrated expert action remains the correct action after the visual change; if a subtle change alters the intended task, action drift misreads a legitimate action change as fragility and mislabels the selected repairs.","fun_headline_variants_meta":{"raw":{"variants":["Action drift finds 25x fewer demos for robot repair","Your robot's own predictions pick its best fix data","Offline sensitivity audit cuts robot repair demos 25x","Use the policy's weak spots to choose 20 demos that work","25x less data for robot robustness via action drift"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000146,"raw_usage":{"total_tokens":1044,"prompt_tokens":793,"completion_tokens":251,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":537,"completion_tokens_details":{"reasoning_tokens":181}},"tokens_in":537,"tokens_out":251,"duration_ms":3464,"temperature":1.0,"reasoning_tokens":181,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-01T14:21:49.080379+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Construct a set of nuisances that deliberately changes the intended goal, violating the equality a*(s_c)=a*(s_n), while keeping the same visual shift magnitudes; if drift ranks these task-changing shifts as high-priority repairs, the signal conflates task change with nuisance fragility. Also, evaluate a held-out nuisance family absent from the candidate pool: if response-guided repair with a budget of 30 cannot beat random selection at the same budget on that condition despite the pool containing useful candidates, the coverage objective has failed.","supporting_citations":[],"review_version":1}