REVIEW 3 major objections 5 minor 1 cited by
Reliable value estimates let robots turn messy demos and rollouts into policies that succeed at fine assembly.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-14 14:57 UTC pith:4KPO2WXP
load-bearing objection Solid real-robot offline-to-online system with usable reliability metrics; the reliability o performance story is real but partly confounded by history length and task-tuned thresholds. the 3 major comments →
Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
Downstream offline and online policy performance is strongly associated with value-function reliability: history-conditioned values that capture global task progress and local action preference produce better action-quality labels, so value-guided offline RL scales more effectively than quality-agnostic behavior cloning and online residual adaptation stays stable by prioritizing high-quality rollouts, yielding 86% success on millimeter chip insertion and 84% on block disassembly.
What carries the argument
Robo-ValueRL: a history-conditioned value estimator scored by global Midpoint Ordering Rate and local fluency/error-discrimination metrics, whose value differences become action-quality labels for quality-conditioned consistency-policy pretraining and for filtering online residual adaptation.
Load-bearing premise
The system assumes that normalized remaining-time progress, with a fixed failure penalty, is a good enough training target so that value differences correctly rank which actions help the task in both offline filtering and online residual learning.
What would settle it
Train the same pipeline on chip insertion and block disassembly with matched data scales, but swap in a deliberately poorer value estimator (no history, or remaining-time targets scrambled) while keeping identical quality thresholds and residual training: if success rates no longer track the reliability metrics and value-guided methods stop beating behavior cloning and DAGGER, the central reliability claim fails.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper presents Robo-ValueRL, a full-stack offline-to-online RL framework for robotic manipulation that centers on a history-conditioned value estimator. Value reliability is quantified via global Midpoint Ordering Rate (MOR) and local fluency/error-discrimination metrics (Bump Ratio/Magnitude, Error Sensitivity/Slope). These estimates supply action-quality labels for quality-conditioned consistency-policy VLA pretraining and for filtering online rollouts used to train a gated residual adapter. On two real-robot tasks (millimeter-level chip insertion and generalizable block disassembly), using 240 h of offline data and >3 000 online trajectories, the authors report that higher reliability metrics align with better offline scaling versus BC and more stable online gains versus DAGGER, reaching 86 % and 84 % success respectively.
Significance. If the reliability–performance association holds under cleaner isolation, the work supplies a practical diagnostic suite and an end-to-end recipe for converting mixed-quality robotic experience into policy improvement—an important gap relative to systems that only report final success rates. Strengths include large-scale real-robot evaluation, one-take continuous videos (30/35 and 58/70), public code/data/models, and explicit ablations of history length, quality thresholds, and residual adaptation. The remaining-time supervision target and the proposed metrics are concrete and falsifiable, giving the community a reusable testbed even if some causal claims require tightening.
major comments (3)
- The central claim that value-estimation reliability drives offline scaling and online stability is only partially isolated. Table 1 shows SHORTHISTORY highest on most metrics and on success, yet LONGHISTORY is competitive on Bump Mag. and Slope while collapsing to 30 %/42 % success—already weakening a pure reliability o performance story. The cleanest comparison (VG-STRICT vs VG-NH, same strict labeling rule) still confounds estimator architecture (5-frame history) with the reliability scores themselves. A controlled experiment that holds history length fixed and varies only metric-selected estimators, or that reports reliability of the exact value model used for each policy, is needed before the causal language in the abstract and §4.1–4.3 is fully warranted.
- Quality-selection hyperparameters are task-tuned (strict Δ=20 / top-30 % for chip insertion; soft Δ=60 / top-50 % for block disassembly; §4.2 and Fig. 5). Consequently the reported scaling gains of “reliable-value guided” methods over BC mix reliability with threshold choice. The paper should either (i) fix a single selection rule across both tasks and re-evaluate, or (ii) treat threshold selection as an explicit hyper-parameter study and show that the reliability metrics still predict the best threshold a priori. Without this, the strongest claim overstates what the ablations establish.
- Appendix C.1 (Eqs. 9–10) defines the value target as normalized remaining time with fixed T_max=5000 and an unspecified failure penalty C_fail. This is a free design choice whose sensitivity is never reported. If remaining-time progress poorly reflects true task value under partial observability or multi-step trade-offs (e.g., non-greedy disassembly in Fig. 8a), both the quality labels and the reliability metrics become misaligned with policy needs. A short ablation on C_fail / T_max, or comparison against an alternative progress signal, is required to support the weakest assumption of the pipeline.
minor comments (5)
- Notation for the residual gate and residual bound (r_max, g_t) appears in §3.3 without a clear statement of how r_max is chosen; a sentence in Appendix C.3 would help reproducibility.
- Figure 1 caption and main text use both “Globel” and “Global”; correct the typo.
- The textual quality prompts (“Quality: Low/Medium/High”) and the 10 % dropout to Medium (§C.2) are reasonable but their effect on the consistency head is never ablated; a brief note would strengthen the method section.
- Success rates in Table 1 and Fig. 5 lack error bars or number of evaluation trials; given the real-robot setting this information is important for assessing variance.
- Related-work discussion of concurrent world-value / GVL metrics (§2.3) could more explicitly contrast the proposed MOR/Bump suite against Value-Order Correlation and VIP smoothness on the same trajectories.
Circularity Check
No circular derivation: value targets, reliability metrics, and downstream success are independently defined; the reliability–performance link is an empirical association, not an identity by construction.
full rationale
Robo-ValueRL is an empirical systems paper, not a first-principles derivation. Value supervision is remaining-time progress with a fixed failure penalty (Appendix C.1, Eqs. 9–10), fixed before any policy training. Reliability metrics (MOR, Bump Ratio/Magnitude, Error Sensitivity/Slope, Sec. 3.4) are computed from value sequences against annotated subgoals and errors, independently of final success rates. Action-quality labels from value differences are then used to train quality-conditioned policies and residual adapters—the intended mechanism of value-guided RL, not a definitional identity between inputs and claimed outputs. Downstream success (chip insertion, block disassembly) is measured on real-world rollouts. No equation forces success rates from reliability scores; Table 1 and Figs. 5–6 report measured associations. Self-citations (authors’ prior robotics work) appear only as related work and are not load-bearing uniqueness theorems or smuggled ansatze. Confounding concerns (history length, task-tuned thresholds) affect causal isolation, not circularity of the derivation chain. The paper is self-contained against its own experimental benchmarks.
Axiom & Free-Parameter Ledger
free parameters (5)
- visual history length (SHORT=5 frames vs LONG=30)
- quality-selection window Δ and percentile (Δ=20 top-30% strict vs Δ=60 top-50% soft)
- value-difference margins δ+, δ− for discrete quality labels {0,1,2}
- T_max=5000 and failure penalty C_fail in remaining-time target
- distributional bins K=256, soft-target σ, residual bound r_max, loss weights λ_cons/λ_keep/λ_gate
axioms (4)
- domain assumption Normalized remaining time (with failure penalty) is a valid scalar progress target for supervising a value function that ranks action quality.
- ad hoc to paper Thresholded value differences over fixed temporal windows are sufficient indicators of action quality for both offline filtering and online residual training.
- domain assumption Freezing the offline VLA and training only a gated residual adapter preserves the pretrained prior while allowing targeted correction.
- domain assumption A consistency policy head with one-step denoising is an adequate action expert for real-time quality-conditioned control.
invented entities (3)
-
Midpoint Ordering Rate (MOR) and local fluency/error-discrimination metrics (Bump Ratio, Bump Magnitude, Error Sensitivity, Error Slope)
no independent evidence
-
Quality-conditioned consistency-policy VLA with textual quality prompts
no independent evidence
-
Gated residual online adapter trained with value-filtered high-quality online segments plus offline keep loss
no independent evidence
read the original abstract
Offline-to-online reinforcement learning is promising for generalizable robotic manipulation, yet its full-stack complexity obscures reproduction and diagnosis. Within such systems, value estimation plays a central role in prioritizing heterogeneous data for policy improvement. Despite its importance, the central question remains underexplored: how value-function reliability shapes policy optimization in offline-to-online reinforcement learning. To answer this question, we propose Robo-ValueRL, a unified framework that enables reliable value estimation and systematically traces its downstream effects on policy pretraining and online improvement. Concretely, Robo-ValueRL learns a history-conditioned value estimator and evaluates its reliability through global-progress and local-preference metrics. These resulting value estimates are propagated into quality-conditioned consistency-policy pretraining and a residual adaptation module on online rollouts, providing a unified testbed for analyzing how value reliability shapes downstream policy performance. Across 240 hours of offline demonstrations and over 3,000 online rollout trajectories, our extensive experiments show that downstream performance is strongly associated with value reliability. Reliable value functions provide better action-quality estimates, allowing value-guided offline RL to scale more effectively than quality-agnostic behavior cloning, and stabilize online improvement by prioritizing high-quality rollout data. Integrating reliable value guidance through offline pretraining with online improvement, our system achieves 86% success on millimeter-level precise chip insertion and 84% on generalizable block disassembly. We hope these findings highlight the importance of value-guided data utilization for effective policy improvement from heterogeneous robotic experience.
Forward citations
Cited by 1 Pith paper
-
$N_0$-VTLA: Scaling Vision-Tactile-Language-Action Model with Latent Tactile Tokens
A VLA with predictive latent tactile tokens pretrained on large-scale visuo-tactile data, plus ALTER offline advantage labeling, leads contact-rich real and sim benchmarks.
Reference graph
Works this paper leans on
-
[1]
Awac: Accelerating online reinforcement learning with offline datasets,
A. Nair, A. Gupta, M. Dalal, and S. Levine, “Awac: Accelerating online reinforcement learning with offline datasets,”arXiv preprint arXiv:2006.09359, 2020
Pith/arXiv arXiv 2006
-
[2]
Conservative q-learning for offline reinforcement learning,
A. Kumar, A. Zhou, G. Tucker, and S. Levine, “Conservative q-learning for offline reinforcement learning,”Advances in neural information processing systems, vol. 33, pp. 1179–1191, 2020
2020
-
[3]
Cal-ql: Calibrated offline rl pre-training for efficient online fine-tuning,
M. Nakamoto, S. Zhai, A. Singh, M. Sobol Mark, Y . Ma, C. Finn, A. Kumar, and S. Levine, “Cal-ql: Calibrated offline rl pre-training for efficient online fine-tuning,”Advances in Neural Information Processing Systems, vol. 36, pp. 62 244–62 269, 2023
2023
-
[4]
RT-1: Robotics Transformer for Real-World Control at Scale,
A. Brohan, N. Brown, J. Carbajal, Y . Chebotar, J. Dabis, C. Finn, K. Gopalakrishnan, K. Hausman, A. Herzog, J. Hsu, J. Ibarz, B. Ichter, A. Irpan, T. Jackson, S. Jesmonth, N. Joshi, R. Julian, D. Kalashnikov, Y . Kuang, I. Leal, K.-H. Lee, S. Levine, Y . Lu, U. Malla, D. Manjunath, I. Mordatch, O. Nachum, C. Parada, J. Peralta, E. Perez, K. Pertsch, J. Q...
2023
-
[5]
Rt-2: Vision-language-action models transfer web knowledge to robotic control,
B. Zitkovich, T. Yu, S. Xu, P. Xu, T. Xiao, F. Xia, J. Wu, P. Wohlhart, S. Welker, A. Wahidet al., “Rt-2: Vision-language-action models transfer web knowledge to robotic control,” inConference on Robot Learning. PMLR, 2023, pp. 2165–2183
2023
-
[6]
Open x-embodiment: Robotic learning datasets and rt-x models: Open x-embodiment collaboration 0,
A. O’Neill, A. Rehman, A. Maddukuri, A. Gupta, A. Padalkar, A. Lee, A. Pooley, A. Gupta, A. Mandlekar, A. Jainet al., “Open x-embodiment: Robotic learning datasets and rt-x models: Open x-embodiment collaboration 0,” in2024 IEEE International Conference on Robotics and Automation. IEEE, 2024, pp. 6892–6903
2024
-
[7]
Openvla: An open-source vision-language- action model,
M. J. Kim, K. Pertsch, S. Karamcheti, T. Xiao, A. Balakrishna, S. Nair, R. Rafailov, E. P. Foster, P. R. Sanketi, Q. Vuong, T. Kollar, B. Burchfiel, R. Tedrake, D. Sadigh, S. Levine, P. Liang, and C. Finn, “Openvla: An open-source vision-language- action model,” inProceedings of The 8th Conference on Robot Learning, ser. Proceedings of Machine Learning Re...
2025
-
[8]
π0: A Vision-Language-Action Flow Model for General Robot Control,
K. Black, N. Brown, D. Driess, A. Esmail, M. R. Equi, C. Finn, N. Fusai, L. Groom, K. Hausman, B. Ichter, S. Jakubczak, T. Jones, L. Ke, S. Levine, A. Li-Bell, M. Mothukuri, S. Nair, K. Pertsch, L. X. Shi, L. Smith, J. Tanner, Q. Vuong, A. Walling, H. Wang, and U. Zhilinsky, “ π0: A Vision-Language-Action Flow Model for General Robot Control,” inProceedin...
2025
-
[9]
Gr00t n1: An open foundation model for generalist humanoid robots,
J. Bjorck, F. Castañeda, N. Cherniadev, X. Da, R. Ding, L. Fan, Y . Fang, D. Fox, F. Hu, S. Huanget al., “Gr00t n1: An open foundation model for generalist humanoid robots,”arXiv preprint arXiv:2503.14734, 2025
Pith/arXiv arXiv 2025
-
[10]
Pre-Training for Robots: Offline RL Enables Learning New Tasks in a Handful of Trials,
A. Kumar, A. Singh, F. D. Ebert, M. Nakamoto, Y . Yang, C. Finn, and S. Levine, “Pre-Training for Robots: Offline RL Enables Learning New Tasks in a Handful of Trials,” inProceedings of Robotics: Science and Systems, July 2023
2023
-
[11]
Y . Zhang, K. Wu, Z. Gao, Z. Zhao, P. Ren, Z. Xu, F. Liao, X. Wang, S. Fan, D. Wuet al., “Robogene: Boosting vla pre-training via diversity-driven agentic framework for real-world task generation,”arXiv preprint arXiv:2602.16444, 2026
arXiv 2026
-
[12]
Precise and dexterous robotic manipulation via human-in-the-loop reinforcement learning,
J. Luo, C. Xu, J. Wu, and S. Levine, “Precise and dexterous robotic manipulation via human-in-the-loop reinforcement learning,” Science Robotics, vol. 10, no. 105, p. eads5033, 2025
2025
-
[13]
Robot fine-tuning made easy: Pre-training rewards and policies for autonomous real-world reinforcement learning,
J. Yang, M. S. Mark, B. Vu, A. Sharma, J. Bohg, and C. Finn, “Robot fine-tuning made easy: Pre-training rewards and policies for autonomous real-world reinforcement learning,” in2024 IEEE International Conference on Robotics and Automation. IEEE, 2024, pp. 4804–4811
2024
-
[14]
SimpleVLA-RL: Scaling VLA training via reinforcement learning,
H. Li, Y . Zuo, J. Yu, Y . Zhang, Y . Zhaohui, K. Zhang, X. Zhu, Y . Zhang, T. Chen, G. Cui, D. Wang, D. Luo, Y . Fan, Y . Sun, J. Zeng, J. Pang, S. Zhang, Y . Wang, Y . Mu, B. Zhou, and N. Ding, “SimpleVLA-RL: Scaling VLA training via reinforcement learning,” inThe Fourteenth International Conference on Learning Representations, 2026
2026
-
[15]
Phoenix: A motion-based self-reflection framework for fine-grained robotic action correction,
W. Xia, R. Feng, D. Wang, and D. Hu, “Phoenix: A motion-based self-reflection framework for fine-grained robotic action correction,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2025, pp. 6981–6990. 12
2025
-
[16]
Human-assisted robotic policy refinement via action preference optimization,
W. Xia, Y . Yang, H. Wu, X. Ma, T. Kong, and D. Hu, “Human-assisted robotic policy refinement via action preference optimization,” inAdvances in Neural Information Processing Systems, D. Belgrave, C. Zhang, H. Lin, R. Pascanu, P. Koniusz, M. Ghassemi, and N. Chen, Eds., vol. 38. Curran Associates, Inc., 2025, pp. 36 746–36 768
2025
-
[17]
Geco-srt: Geometry-aware continual adaptation for cross-task sim-to-real transfer,
W. Yu, W. Xia, W. Zhang, and D. Hu, “Geco-srt: Geometry-aware continual adaptation for cross-task sim-to-real transfer,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, June 2026, pp. 42 408–42 417
2026
-
[18]
Deep reinforcement learning that matters,
P. Henderson, R. Islam, P. Bachman, J. Pineau, D. Precup, and D. Meger, “Deep reinforcement learning that matters,” in Proceedings of the AAAI conference on artificial intelligence, vol. 32, no. 1, 2018
2018
-
[19]
Deep reinforcement learning at the edge of the statistical precipice,
R. Agarwal, M. Schwarzer, P. S. Castro, A. C. Courville, and M. Bellemare, “Deep reinforcement learning at the edge of the statistical precipice,”Advances in neural information processing systems, vol. 34, pp. 29 304–29 320, 2021
2021
-
[20]
Learning while deploying: Fleet-scale reinforcement learning for generalist robot policies,
Y . Wang, X. Li, P. Xie, P. Yang, B. Nie, Y . Cai, Q. Zhang, C. Qu, J. Wu, J. Songet al., “Learning while deploying: Fleet-scale reinforcement learning for generalist robot policies,”arXiv preprint arXiv:2605.00416, 2026
Pith/arXiv arXiv 2026
-
[21]
Suf: Stabilized unconstrained fine-tuning for offline-to-online reinforcement learning,
J. Feng, M. Feng, H. Song, W. Zhou, and H. Li, “Suf: Stabilized unconstrained fine-tuning for offline-to-online reinforcement learning,” inProceedings of the AAAI Conference on Artificial Intelligence, vol. 38, no. 11, 2024, pp. 11 961–11 969
2024
-
[22]
Offline-to-online reinforcement learning via balanced replay and pessimistic q-ensemble,
S. Lee, Y . Seo, K. Lee, P. Abbeel, and J. Shin, “Offline-to-online reinforcement learning via balanced replay and pessimistic q-ensemble,” inConference on Robot Learning. PMLR, 2022, pp. 1702–1712
2022
-
[23]
pi*0.6: a vla that learns from experience,
P. Intelligence, A. Amin, R. Aniceto, A. Balakrishna, K. Black, K. Conley, G. Connors, J. Darpinian, K. Dhabalia, J. DiCarlo et al., “pi*0.6: a vla that learns from experience,”arXiv preprint arXiv:2511.14759, 2025
Pith/arXiv arXiv 2025
-
[24]
What matters in learning from offline human demonstrations for robot manipulation,
A. Mandlekar, D. Xu, J. Wong, S. Nasiriany, C. Wang, R. Kulkarni, L. Fei-Fei, S. Savarese, Y . Zhu, and R. Martín-Martín, “What matters in learning from offline human demonstrations for robot manipulation,” inProceedings of the 5th Conference on Robot Learning, A. Faust, D. Hsu, and G. Neumann, Eds., vol. 164. PMLR, 08–11 Nov 2022, pp. 1678–1690
2022
-
[25]
Scalable deep reinforcement learning for vision-based robotic manipulation,
D. Kalashnikov, A. Irpan, P. Pastor, J. Ibarz, A. Herzog, E. Jang, D. Quillen, E. Holly, M. Kalakrishnan, V . Vanhouckeet al., “Scalable deep reinforcement learning for vision-based robotic manipulation,” inConference on robot learning. PMLR, 2018, pp. 651–673
2018
-
[26]
Gr-rl: Going dexterous and precise for long-horizon robotic manipulation,
Y . Li, X. Ma, J. Xu, Y . Cui, Z. Cui, Z. Han, L. Huang, T. Kong, Y . Liu, H. Niuet al., “Gr-rl: Going dexterous and precise for long-horizon robotic manipulation,”arXiv preprint arXiv:2512.01801, 2025
arXiv 2025
-
[27]
VIP: Towards universal visual reward and representation via value-implicit pre-training,
Y . J. Ma, S. Sodhani, D. Jayaraman, O. Bastani, V . Kumar, and A. Zhang, “VIP: Towards universal visual reward and representation via value-implicit pre-training,” inThe Eleventh International Conference on Learning Representations, 2023
2023
-
[28]
Conrft: A reinforced fine-tuning method for vla models via consistency policy,
Y . Chen, S. Tian, S. Liu, Y . Zhou, H. Li, and D. Zhao, “Conrft: A reinforced fine-tuning method for vla models via consistency policy,” inProceedings of Robotics: Science and Systems, 2025, Los Angeles, CA, USA, Jun 21-25, 2025, 2025
2025
-
[29]
Consistency policy: Accelerated visuomotor policies via consistency distillation,
A. Prasad, K. Lin, J. Wu, L. Zhou, and J. Bohg, “Consistency policy: Accelerated visuomotor policies via consistency distillation,” inRobotics: Science and Systems, 2024
2024
-
[30]
Learning from imperfect demonstrations from agents with varying dynamics,
Z. Cao and D. Sadigh, “Learning from imperfect demonstrations from agents with varying dynamics,”IEEE Robotics and Automation Letters, vol. 6, no. 3, pp. 5231–5238, 2021
2021
-
[31]
Curating Demonstrations using Online Experience,
A. S. Chen, A. M. Lessing, Y . Liu, and C. Finn, “Curating Demonstrations using Online Experience,” inProceedings of Robotics: Science and Systems, LosAngeles, CA, USA, June 2025
2025
-
[32]
Octo: An Open-Source Generalist Robot Policy,
D. Ghosh, H. R. Walke, K. Pertsch, K. Black, O. Mees, S. Dasari, J. Hejna, T. Kreiman, C. Xu, J. Luo, Y . L. Tan, L. Y . Chen, Q. Vuong, T. Xiao, P. R. Sanketi, D. Sadigh, C. Finn, and S. Levine, “Octo: An Open-Source Generalist Robot Policy,” in Proceedings of Robotics: Science and Systems, Delft, Netherlands, July 2024
2024
-
[33]
Xr-1: Towards versatile vision- language-action models via learning unified vision-motion representations,
S. Fan, K. Wu, Z. Che, X. Wang, D. Wu, F. Liao, N. Liu, Y . Zhang, Z. Zhao, Z. Xuet al., “Xr-1: Towards versatile vision- language-action models via learning unified vision-motion representations,” inProceedings of the International Conference on Machine Learning, 2026
2026
-
[34]
A survey on offline reinforcement learning: Taxonomy, review, and open problems,
R. F. Prudencio, M. R. Maximo, and E. L. Colombini, “A survey on offline reinforcement learning: Taxonomy, review, and open problems,”IEEE transactions on neural networks and learning systems, vol. 35, no. 8, pp. 10 237–10 257, 2023
2023
-
[35]
Offline reinforcement learning with implicit q-learning,
I. Kostrikov, A. Nair, and S. Levine, “Offline reinforcement learning with implicit q-learning,” inInternational Conference on Learning Representations, 2022
2022
-
[36]
Policy expansion for bridging offline-to-online reinforcement learning,
H. Zhang, W. Xu, and H. Yu, “Policy expansion for bridging offline-to-online reinforcement learning,” inThe Eleventh International Conference on Learning Representations, 2023. 13
2023
-
[37]
Q-transformer: Scalable offline reinforcement learning via autoregressive q-functions,
Y . Chebotar, Q. Vuong, K. Hausman, F. Xia, Y . Lu, A. Irpan, A. Kumar, T. Yu, A. Herzog, K. Pertschet al., “Q-transformer: Scalable offline reinforcement learning via autoregressive q-functions,” inConference on Robot Learning. PMLR, 2023, pp. 3909–3928
2023
-
[38]
Finetuning from offline reinforcement learning: Challenges, trade-offs and practical solutions,
Y . Luo, J. Kay, E. Grefenstette, and M. P. Deisenroth, “Finetuning from offline reinforcement learning: Challenges, trade-offs and practical solutions,”arXiv preprint arXiv:2303.17396, 2023
Pith/arXiv arXiv 2023
-
[39]
Learning to predict by the methods of temporal differences,
R. S. Sutton, “Learning to predict by the methods of temporal differences,”Machine learning, vol. 3, no. 1, pp. 9–44, 1988
1988
-
[40]
Analysis of temporal-diffference learning with function approximation,
J. Tsitsiklis and B. Van Roy, “Analysis of temporal-diffference learning with function approximation,”Advances in neural information processing systems, vol. 9, 1996
1996
-
[41]
Fast gradient-descent methods for temporal-difference learning with linear function approximation,
R. S. Sutton, H. R. Maei, D. Precup, S. Bhatnagar, D. Silver, C. Szepesvári, and E. Wiewiora, “Fast gradient-descent methods for temporal-difference learning with linear function approximation,” inProceedings of the 26th annual international conference on machine learning, 2009, pp. 993–1000
2009
-
[42]
Human-level control through deep reinforcement learning,
V . Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. Riedmiller, A. K. Fidjeland, G. Ostrovskiet al., “Human-level control through deep reinforcement learning,”nature, vol. 518, no. 7540, pp. 529–533, 2015
2015
-
[43]
Benchmarking deep reinforcement learning for continuous control,
Y . Duan, X. Chen, R. Houthooft, J. Schulman, and P. Abbeel, “Benchmarking deep reinforcement learning for continuous control,” inInternational conference on machine learning. PMLR, 2016, pp. 1329–1338
2016
-
[44]
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor,
T. Haarnoja, A. Zhou, P. Abbeel, and S. Levine, “Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor,” inInternational conference on machine learning. Pmlr, 2018, pp. 1861–1870
2018
-
[45]
A vision-language-action-critic model for robotic real-world reinforcement learning,
S. Zhai, Q. Zhang, T. Zhang, F. Huang, H. Zhang, M. Zhou, S. Zhang, L. Liu, S. Lin, and J. Pang, “A vision-language-action-critic model for robotic real-world reinforcement learning,”arXiv preprint arXiv:2509.15937, 2025
arXiv 2025
-
[46]
RL-VLM-f: Reinforcement learning from vision lan- guage foundation model feedback,
Y . Wang, Z. Sun, J. Zhang, Z. Xian, E. Biyik, D. Held, and Z. Erickson, “RL-VLM-f: Reinforcement learning from vision lan- guage foundation model feedback,” inProceedings of the 41st International Conference on Machine Learning, ser. Proceedings of Machine Learning Research, vol. 235. PMLR, 2024, pp. 51 484–51 501
2024
-
[47]
Rank2reward: Learning shaped reward functions from passive video,
D. Yang, D. Tjia, J. Berg, D. Damen, P. Agrawal, and A. Gupta, “Rank2reward: Learning shaped reward functions from passive video,” in2024 IEEE International Conference on Robotics and Automation. IEEE, 2024, pp. 2806–2813
2024
-
[48]
Vision language models are in-context value learners,
Y . J. Ma, J. Hejna, C. Fu, D. Shah, J. Liang, Z. Xu, S. Kirmani, P. Xu, D. Driess, T. Xiaoet al., “Vision language models are in-context value learners,” inInternational Conference on Learning Representations, vol. 2025, 2025, pp. 33 984–34 009
2025
-
[49]
Universal value function approximators,
T. Schaul, D. Horgan, K. Gregor, and D. Silver, “Universal value function approximators,” inInternational conference on machine learning. PMLR, 2015, pp. 1312–1320
2015
-
[50]
Viva: A video-generative value model for robot reinforcement learning,
J. Lv, H. Li, J. Li, Y . Nie, F. Kong, Y . Wang, X. Wang, Z. Zhu, C. Ni, Q. Denget al., “Viva: A video-generative value model for robot reinforcement learning,”arXiv preprint arXiv:2604.08168, 2026
Pith/arXiv arXiv 2026
-
[51]
World value models for robotic manipulation,
Z. Wang, J. Li, Y . Cui, Y . Gao, X. Zhan, J. Yu, and X. Ma, “World value models for robotic manipulation,”arXiv preprint arXiv:2606.24742, 2026
Pith/arXiv arXiv 2026
-
[52]
Paligemma: A versatile 3b vlm for transfer,
L. Beyer, A. Steiner, A. S. Pinto, A. Kolesnikov, X. Wang, D. Salz, M. Neumann, I. Alabdulmohsin, M. Tschannen, E. Bugliarello et al., “Paligemma: A versatile 3b vlm for transfer,”arXiv preprint arXiv:2407.07726, 2024
Pith/arXiv arXiv 2024
-
[53]
Sigmoid loss for language image pre-training,
X. Zhai, B. Mustafa, A. Kolesnikov, and L. Beyer, “Sigmoid loss for language image pre-training,” inProceedings of the IEEE/CVF international conference on computer vision, 2023, pp. 11 975–11 986
2023
-
[54]
Perceiver: General perception with iterative attention,
A. Jaegle, F. Gimeno, A. Brock, O. Vinyals, A. Zisserman, and J. Carreira, “Perceiver: General perception with iterative attention,” inInternational conference on machine learning. PMLR, 2021, pp. 4651–4664
2021
-
[55]
Stop regressing: Training value functions via classification for scalable deep RL,
J. Farebrother, J. Orbay, Q. Vuong, A. Ali Taiga, Y . Chebotar, T. Xiao, A. Irpan, S. Levine, P. S. Castro, A. Faust, A. Kumar, and R. Agarwal, “Stop regressing: Training value functions via classification for scalable deep RL,” inProceedings of the 41st International Conference on Machine Learning, ser. Proceedings of Machine Learning Research, R. Salakh...
2024
-
[56]
A reduction of imitation learning and structured prediction to no-regret online learning,
S. Ross, G. Gordon, and D. Bagnell, “A reduction of imitation learning and structured prediction to no-regret online learning,” inProceedings of the Fourteenth International Conference on Artificial Intelligence and Statistics, ser. Proceedings of Machine Learning Research, G. Gordon, D. Dunson, and M. Dudík, Eds., vol. 15. Fort Lauderdale, FL, USA: PMLR,...
2011
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.