REVIEW 3 major objections 5 minor 33 references
A failure-probability model trained on a policy’s own rollouts can steer data collection so finetuning finally escapes the plateau of random sampling.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-31 13:14 UTC pith:NRO5W4DG
load-bearing objection Solid cross-domain transfer of criticality-guided sampling with real ablations; abstract overclaims a single unbiased-IS mechanism that only fully holds in the RL domains. the 3 major comments →
Self-Evolving Learning for Embodied AI with Criticality Model
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The paper’s central claim is that a state-wise criticality model C_φ ≈ P(failure|state), trained only on a policy’s own execution outcomes, can define an importance-sampling proposal that concentrates collection on failure-prone regions; after importance-weight correction (and optional threshold routing between baseline and finetuned policies), the training objective remains unbiased under the natural state distribution while the pool’s information density rises, producing large, consistent failure-rate reductions that matched random-data finetuning does not achieve.
What carries the argument
The criticality model C_φ: a lightweight network that scores P(failure|state) from rollout labels, builds the ε-mixture proposal q(s) ∝ κ(C_φ(s)), and supplies both collection weights and deployment routing thresholds.
Load-bearing premise
That a compact model on hand-chosen state features can keep finding the policy’s remaining failure modes as the policy improves; if those features miss the true rare causes, guided sampling aims at the wrong slice of rarity.
What would settle it
Run the same multi-round loop with a deliberately incomplete state input to C_φ (drop terrain, force, or object-pose channels) and check whether round-over-round failure-rate gains collapse to the random-data control while validation PR-AUC of C_φ stays high on the incomplete features.
If this is right
- Default random finetuning pipelines for locomotion, manipulation, and VLA policies will plateau even with more steps; criticality-guided collection is required to keep cutting residual failures.
- The same P(failure|s) monitor can double as a deployment risk switch that hands high-criticality states to a specialist policy without retraining the base model.
- Failure-rate gains of 51–67% vs trained baselines and 8–25% vs strong VLA checkpoints become available without changing the original RL or imitation loss, only the data proposal.
- Information density—more distinct failure modes in the pool, not a higher critical-to-nominal batch ratio—is the operative lever, as the paper’s controlled toy classification experiment isolates.
Where Pith is reading between the lines
- If criticality were learned in a vision latent space instead of engineered states, the same loop could transfer across embodiments without per-domain feature design.
- Coupling C_φ’s per-step scores as dense process rewards inside online VLA RL could close the loop the paper leaves open under flow-matching instability.
- Domains outside robotics that also suffer a curse of rarity (rare safety events, rare medical outcomes) may admit the same evaluate-to-train transition once a cheap failure predictor exists.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper argues that embodied-policy finetuning plateaus because random data collection undersamples rare failures, and proposes a self-evolving loop built around a lightweight state-wise criticality model C_φ ≈ P(failure|state) trained on the policy’s own rollout outcomes. C_φ defines an importance-sampling proposal that concentrates collection on failure-prone states; the policy is then finetuned on the curated pool, and the loop iterates. At deployment, C_φ optionally routes high-criticality states to the finetuned specialist. Empirically, across Go2 locomotion, ManiSkill multi-task manipulation, LIBERO and RoboTwin VLA benchmarks, and a real banana-on-plate task, the method reports 51–67% relative failure-rate reductions versus trained baselines and 8–25% versus SOTA VLA checkpoints, while matched random-data finetuning yields negligible gain or degradation. A controlled toy experiment (Supp. §1) is used to argue that the mechanism is increased information density of distinct critical scenarios under an unbiased objective, not merely a higher critical:nominal sample ratio.
Significance. If the claims hold under a carefully stated mechanism, this is a practically useful systems contribution: a single cheap failure predictor that couples data selection and deployment routing, validated across RL and IL/VLA regimes including a real robot, with matched random-data ablations (Table 7), multi-round trends (Fig. 3), and an explicit information-density toy study. The work generalizes prior rare-event IS ideas (e.g., NADE-style criticality) from evaluation into iterative embodied finetuning. Strengths include multi-domain consistency, clear random-collection controls, and transparent reporting that Finetune-only can regress on VLA tasks until routing is applied. The main value is empirical and methodological rather than theoretical novelty of IS itself.
major comments (3)
- [Abstract; Method Stage 2–3, Eqs. (2)–(4)] Abstract and Stage 2–3 (Eqs. 2–4) present a unified mechanism: criticality-guided collection plus importance-weight resampling that “preserves an unbiased learning objective under P(s).” This is accurate for the RL domains (Go2, ManiSkill), where episodes are resampled ∝ W = P/q and the loss is unweighted. For the IL/VLA domains (LIBERO, RoboTwin, real robot)—three of five settings and the only SOTA-VLA comparisons—Stage 3 samples expert demos from the biased q(i) (Eq. 3) and runs standard unweighted BC, with the text stating the bias is “absorbed by q(i).” No likelihood-ratio correction is applied. The abstract’s unbiased-objective sentence and the single information-density story therefore overextend the RL construction to the IL results and should be restated domain-conditionally.
- [Deployment Eq. (5); Tables 4–5; Contribution 1] On LIBERO and RoboTwin, Finetune-only underperforms the baseline (Table 4: 2.13% vs 1.87%; Table 5: 16.30% vs 14.40%), and the reported 8–25% SOTA-relative cuts appear only after validation-swept threshold routing (Eq. 5; per-suite/global τ). Routing is therefore load-bearing for the headline VLA numbers, not merely an optional deployment monitor as framed in Contribution 1 and the Deployment section. The paper should lead with this specialization trade-off for IL/VLA, report routed vs unrouted numbers as primary, and avoid attributing those gains solely to unbiased density enrichment.
- [Stage 1; Limitations; Table 1; Supp. §9] The weakest modeling assumption is that pre-defined state features (terrain grid, force grid, padded initial-state vectors, 8-D geometric coords; Table 1 and Supp. §9) suffice for C_φ to surface residual failure modes of the improved policy. Limitations notes manual state engineering, but there is no stress test where the representation is intentionally incomplete or where post-update failure modes leave the scored support. A minimal experiment or quantitative failure analysis (e.g., modes missed by C_φ after Round k) would make the central sampling claim more falsifiable, especially for the VLA initial-state-only scores.
minor comments (5)
- [Method; Supp. §3, §7] Free parameters (ε, β/κ, τ, RoboTwin α and capped-sigmoid schedule) are scattered across Method and Supp.; a single hyperparameter table with selection protocol (validation sweep vs fixed defaults) would aid reproducibility.
- [Figure 1] Figure 1 is helpful but the Stage 2 panel does not visually distinguish RL trajectory-weighting from IL demo-level q(i), which is exactly where the mechanism splits.
- [Tables 2, 6; Supp. §2] Go2 and real-robot results use IS evaluation (Supp. §2). Main-text tables should mark IS-estimated μ explicitly so readers do not compare raw Monte Carlo rates to reweighted rates without noticing.
- [Abstract; Introduction] Minor prose/spacing artifacts appear throughout (e.g., “systemsroutinelyplateauduring,” “state-wisecriticalitymodel”). A full copy-edit pass is needed.
- [Related Work] Related Work cites dense-learning / NADE training extensions appropriately; briefly clarify what is new relative to Feng et al. 2026 beyond multi-perturbation embodied settings (already hinted, but one crisp paragraph would help).
Circularity Check
No derivation-circularity: empirical IS/criticality loop with measured outcomes, not predictions forced by fitted inputs.
full rationale
This is an empirical systems paper. The criticality model is ordinary supervised BCE on observed (state, success/failure) rollouts (Eq. 1); the proposal q and weights w=P/q are standard importance sampling (Eqs. 2–3); reported failure rates are Monte Carlo or likelihood-ratio estimates on held-out episodes, not quantities algebraically identical to a fitted constant. The information-density claim is supported by a controlled toy classification experiment in the supplement, not by renaming the training objective. Self-citations to Feng et al. (NADE / dense learning) supply motivation and prior AV context; they do not import a uniqueness theorem that forces the embodied-AI results. Validation-swept routing thresholds τ are ordinary hyperparameter selection for deployment, not a self-definitional ‘prediction.’ Mechanism overclaim on IL/VLA (biased BC without weight correction) is a correctness/scope issue, not circularity of a derivation chain. No step reduces a claimed first-principles result to its inputs by construction.
Axiom & Free-Parameter Ledger
free parameters (5)
- ε (exploration mixture in q) =
0.01–0.1 typical
- β temperature / κ sharpening =
β=3.0 (ManiSkill); none (Go2)
- deployment routing threshold τ =
domain-specific (Table S3)
- RoboTwin score mix α and capped-sigmoid schedule =
α=0.5; max weight ratio 5×; temp 0.15
- criticality MLP capacity and labeling windows =
varies; 321–1.08M params
axioms (5)
- standard math Importance sampling with w=P/q yields unbiased estimates/objectives when q(s)>0 wherever P(s)>0 (ε-mixture).
- domain assumption Binary episode outcome y is a sufficient supervision target for a useful state-wise failure probability in the chosen state features.
- domain assumption Task perturbation spaces can be discretized into finite candidate sets (terrain grid, force grid, initial-state pools) that cover the failure modes of interest.
- domain assumption Finetuning under the task’s original paradigm (offline AWAC/PPO-style or BC) on IS-curated buffers improves residual failures without needing on-policy VLA RL.
- ad hoc to paper Information density of distinct critical scenarios, not merely critical sample proportion, drives failure-mode learning under an unbiased objective.
invented entities (1)
-
state-wise criticality model C_φ
no independent evidence
read the original abstract
Despite rapid advances in policy pretraining, embodied AI systems routinely plateau during task-specific finetuning. The root cause lies in how finetuning data are collected: the default pipeline gathers data randomly, treating every sample as informative. Datasets become dominated by nominal scenarios, while rare failure cases--the most valuable for improvement--are missed. We propose a self-evolving method that breaks this plateau. Our core insight is that a state-wise criticality model, learned from the policy's own execution outcomes to predict the probability of future failure, can guide importance sampling toward failure-prone scenarios. After replacing redundant nominal scenarios with diverse failure-prone ones, importance weights are used to resample the data during training. This effectively preserves an unbiased learning objective while fundamentally increasing the information density of the training pool. Across quadrupedal locomotion, multi-task manipulation, vision-language-action benchmarks, and a real-robot task, our method reduces failure rates by 51--67% relative to trained baselines and by 8-25% relative to state-of-the-art vision-language-action models.
Figures
Reference graph
Works this paper leans on
-
[1]
Chen, K.; Liu, Z.; Zhang, T.; Guo, Z.; Xu, S.; Lin, H.; Zang, H.; Zhang, Q.; Yu, Z.; Fan, G.; Huang, T.; Wang, Y.; and Yu, C. 2025 a . _ RL : Online RL Fine-tuning for Flow-based Vision-Language-Action Models. arXiv:2510.25889
arXiv 2025
-
[2]
Chen, Z.; Wang, S.; Xiao, T.; Wang, Y.; Chen, S.; Cai, X.; He, J.; and Wang, J. 2025 b . Revisiting Scaling Laws for Language Models: The Role of Data Quality and Training Strategies. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (ACL), volume 1, 23881--23899
2025
-
[3]
Dawson, C.; Parashar, A.; and Fan, C. 2025. RADIUM : Predicting and Repairing End-to-End Robot Failures Using Gradient-Accelerated Sampling. IEEE Transactions on Robotics, 41: 2268--2284
2025
-
[4]
M.; and Kochenderfer, M
Delecki, H.; Katz, S. M.; and Kochenderfer, M. J. 2025. Failure Probability Estimation for Black-Box Autonomous Systems using State-Dependent Importance Sampling Proposals. In Proceedings of the 11th International Conference on Control, Decision and Information Technologies (CoDIT), 1--7
2025
-
[5]
Feng, S.; Sun, H.; Yan, X.; Zhu, H.; Zou, Z.; Shen, S.; and Liu, H. X. 2023. Dense reinforcement learning for safety validation of autonomous vehicles. Nature, 615(7953): 620--627
2023
-
[6]
Feng, S.; Yan, X.; Sun, H.; Feng, Y.; and Liu, H. X. 2021. Intelligent Driving Intelligence Test for Autonomous Vehicles with Naturalistic and Adversarial Environment. Nature Communications, 12(1): 748
2021
-
[7]
Feng, S.; Zhu, H.; Sun, H.; Yan, X.; He, L.; Yang, J.; Su, G.; Li, B.; Li, S.; Wang, L.; Shen, S.; and Liu, H. X. 2026. Breaking through Safety Performance Stagnation in Autonomous Vehicles with Dense Learning. Nature Communications, 17: 3163
2026
-
[8]
Gu, Q.; Ju, Y.; Sun, S.; Gilitschenski, I.; Nishimura, H.; Itkina, M.; and Shkurti, F. 2026. SAFE : Multitask Failure Detection for Vision-Language-Action Models. In Advances in Neural Information Processing Systems (NeurIPS), volume 38, 40041--40076
2026
-
[9]
Huang, C.; Chang, Y.; Lin, J.; Liang, J.; Zeng, R.; and Li, J. 2025 a . Efficient Language-instructed Skill Acquisition via Reward-Policy Co-Evolution. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 39, 14576--14584
2025
-
[10]
Huang, S.; Liao, Y.; Feng, S.; Jiang, S.; Liu, S.; Li, H.; Yao, M.; and Ren, G. 2025 b . Adversarial Data Collection: Human-Collaborative Perturbations for Efficient and Robust Robotic Imitation Learning. arXiv:2503.11646
Pith/arXiv arXiv 2025
-
[11]
P.; and Welling, M
Kingma, D. P.; and Welling, M. 2014. Auto-Encoding Variational B ayes. In International Conference on Learning Representations (ICLR)
2014
-
[12]
o pr \"u l \
K \"o pr \"u l \"u , C.; Li, P.-h.; Qiu, T.; Zhao, R.; Westenbroek, T.; Fridovich-Keil, D.; Chinchali, S.; and Topcu, U. 2025. Dense Dynamics-Aware Reward Synthesis: Integrating Prior Experience with Demonstrations. In Proceedings of the 7th Annual Learning for Dynamics & Control Conference (L4DC), 894--906
2025
-
[13]
Liu, B.; Zhu, Y.; Gao, C.; Feng, Y.; Liu, Q.; Zhu, Y.; and Stone, P. 2023. LIBERO : Benchmarking Knowledge Transfer for Lifelong Robot Learning. In Advances in Neural Information Processing Systems (NeurIPS), volume 36, 44776--44791
2023
-
[14]
X.; and Feng, S
Liu, H. X.; and Feng, S. 2024. Curse of Rarity for Autonomous Vehicles. Nature Communications, 15: 4808
2024
-
[15]
L.; Berg, J.; Sharma, A.; Schaal, S.; Finn, C.; Gupta, A.; and Levine, S
Luo, J.; Hu, Z.; Xu, C.; Tan, Y. L.; Berg, J.; Sharma, A.; Schaal, S.; Finn, C.; Gupta, A.; and Levine, S. 2024. SERL : A Software Suite for Sample-Efficient Robotic Reinforcement Learning. In Proceedings of the IEEE International Conference on Robotics and Automation (ICRA), 16961--16969
2024
-
[16]
Luo, J.; Xu, C.; Wu, J.; and Levine, S. 2025. Precise and Dexterous Robotic Manipulation via Human-in-the-Loop Reinforcement Learning. Science Robotics, 10(105): eads5033
2025
-
[17]
Luo, Y.; Chen, W.; Liang, T.; Wang, B.; and Li, Z. 2026. SimVLA : A Simple VLA Baseline for Robotic Manipulation. arXiv:2602.18224
arXiv 2026
-
[18]
A.; Veness, J.; Bellemare, M
Mnih, V.; Kavukcuoglu, K.; Silver, D.; Rusu, A. A.; Veness, J.; Bellemare, M. G.; Graves, A.; Riedmiller, M.; Fidjeland, A. K.; Ostrovski, G.; Petersen, S.; Beattie, C.; Sadik, A.; Antonoglou, I.; King, H.; Kumaran, D.; Wierstra, D.; Legg, S.; and Hassabis, D. 2015. Human-level control through deep reinforcement learning. Nature, 518(7540): 529--533
2015
-
[19]
Mu, Y.; Chen, T.; Chen, Z.; Peng, S.; Lan, Z.; Gao, Z.; Liang, Z.; Yu, Q.; Zou, Y.; Xu, M.; Lin, L.; Xie, Z.; Ding, M.; and Luo, P. 2025. RoboTwin : Dual-Arm Robot Benchmark with Generative Digital Twins. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 27649--27660
2025
-
[20]
Nair, A.; Gupta, A.; Dalal, M.; and Levine, S. 2020. AWAC : Accelerating Online Reinforcement Learning with Offline Datasets. arXiv:2006.09359
Pith/arXiv arXiv 2020
-
[21]
Physical Intelligence . 2025. ^ 0.6 : A Vision-Language-Action Model that Learns from Experience. arXiv:2511.14759
Pith/arXiv arXiv 2025
-
[22]
Reddi, A.; T \"o lle, M.; Peters, J.; Chalvatzaki, G.; and D'Eramo, C. 2024. Robust Adversarial Reinforcement Learning via Bounded Rationality Curricula. In International Conference on Learning Representations (ICLR)
2024
-
[23]
Schaul, T.; Quan, J.; Antonoglou, I.; and Silver, D. 2016. Prioritized Experience Replay. In International Conference on Learning Representations (ICLR)
2016
-
[24]
K.; Wahid, A.; Tompson, J.; Sanketi, P
Seyed Ghasemipour, S. K.; Wahid, A.; Tompson, J.; Sanketi, P. R.; and Mordatch, I. 2026. Self-Improving Embodied Foundation Models. In Advances in Neural Information Processing Systems (NeurIPS), volume 38, 112651--112686
2026
-
[25]
Stachowicz, K.; Ignatova, L.; and Levine, S. 2025. Lifelong Autonomous Improvement of Navigation Foundation Models in the Wild. In Proceedings of the Conference on Robot Learning (CoRL), volume 270 of Proceedings of Machine Learning Research, 1035--1047. PMLR
2025
-
[26]
StarVLA Community . 2026. StarVLA : A Lego-like Codebase for Vision-Language-Action Model Developing. arXiv:2604.05014
Pith/arXiv arXiv 2026
-
[27]
N.; Choi, Y
Tao, S.; Xiang, F.; Shukla, A.; Qin, Y.; Hinrichsen, X.; Yuan, X.; Bao, C.; Lin, X.; Liu, Y.; Chan, T.-k.; Gao, Y.; Li, X.; Mu, T.; Xiao, N.; Gurha, A.; Rajesh, V. N.; Choi, Y. W.; Chen, Y.-R.; Huang, Z.; Calandra, R.; Chen, R.; Luo, S.; and Su, H. 2025. ManiSkill3 : GPU Parallelized Robotics Simulation and Rendering for Generalizable Embodied AI . In Pro...
2025
-
[28]
Tian, W.; Zhang, S.; Zhang, K.; Chi, X.; Fan, C.; Lu, J.; Luo, Y.; Zhou, Q.; Zhao, Y.; Liu, N.; Lin, S.; Qin, Z.; Ju, X.; Zhang, S.; and Tang, J. 2026. SEEA-R1 : Tree-Structured Reinforcement Fine-Tuning for Self-Evolving Embodied Agents. In Advances in Neural Information Processing Systems (NeurIPS), volume 38, 78458--78499
2026
-
[29]
Todorov, E.; Erez, T.; and Tassa, Y. 2012. MuJoCo : A Physics Engine for Model-Based Control. In Proceedings of IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), 5026--5033
2012
-
[30]
Tsao, R.; Wagenmaker, A.; and Levine, S. 2026. Learning Process Rewards via Success Visitation Matching for Efficient RL . arXiv:2606.23640
Pith/arXiv arXiv 2026
-
[31]
Wu, M.; and Cao, Y. 2025. On-Policy Reinforcement Learning from Failure via Sparse Reward Densification. In Proceedings of the 24th International Conference on Autonomous Agents and Multiagent Systems (AAMAS), 2792--2794
2025
-
[32]
Xiao, W.; Lin, H.; Peng, A.; Xue, H.; He, T.; Xie, Y.; Hu, F.; Wu, J.; Luo, Z.; Fan, L.; Shi, G.; and Zhu, Y. 2026. Self-Improving Vision-Language-Action Models with Data Generation via Residual RL . In International Conference on Learning Representations (ICLR)
2026
-
[33]
Xu, M.; Huang, P.; Li, F.; Zhu, J.; Qi, X.; Oguchi, K.; Huang, Z.; Lam, H.; and Zhao, D. 2022. Scalable Safety-Critical Policy Evaluation with Accelerated Rare Event Sampling. In Proceedings of IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), 12919--12926
2022
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.