Pith. sign in

REVIEW 3 major objections 5 minor 1 cited by

Reliable value estimates let robots turn messy demos and rollouts into policies that succeed at fine assembly.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-14 14:57 UTC pith:4KPO2WXP

load-bearing objection Solid real-robot offline-to-online system with usable reliability metrics; the reliability o performance story is real but partly confounded by history length and task-tuned thresholds. the 3 major comments →

arxiv 2607.09866 v1 pith:4KPO2WXP submitted 2026-07-10 cs.RO cs.AI

Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning

classification cs.RO cs.AI
keywords offline-to-online reinforcement learningvalue estimationrobotic manipulationquality-conditioned policyconsistency policyresidual adaptationheterogeneous demonstrationsvision-language-action
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

This paper argues that offline-to-online robot learning fails or succeeds largely because of how well the system estimates the value of experience. When value estimates track both long-horizon task progress and local action quality, they can rank mixed-quality demonstrations and online rollouts so the policy is trained on useful segments and spared bad ones. The authors build Robo-ValueRL around a history-conditioned value estimator, simple reliability metrics (global progress ordering and local preference/error response), quality-conditioned consistency-policy pretraining, and a residual online adapter. On two hard real tasks—millimeter chip insertion and generalizable block disassembly—with 240 hours of offline data and thousands of rollouts, more reliable values scale offline learning past plain behavior cloning and stabilize online improvement, reaching 86% and 84% success. The practical message is that robots improve from heterogeneous experience when value reliability is measured and used as the filter for data, not when more data is simply piled on.

Core claim

Downstream offline and online policy performance is strongly associated with value-function reliability: history-conditioned values that capture global task progress and local action preference produce better action-quality labels, so value-guided offline RL scales more effectively than quality-agnostic behavior cloning and online residual adaptation stays stable by prioritizing high-quality rollouts, yielding 86% success on millimeter chip insertion and 84% on block disassembly.

What carries the argument

Robo-ValueRL: a history-conditioned value estimator scored by global Midpoint Ordering Rate and local fluency/error-discrimination metrics, whose value differences become action-quality labels for quality-conditioned consistency-policy pretraining and for filtering online residual adaptation.

Load-bearing premise

The system assumes that normalized remaining-time progress, with a fixed failure penalty, is a good enough training target so that value differences correctly rank which actions help the task in both offline filtering and online residual learning.

What would settle it

Train the same pipeline on chip insertion and block disassembly with matched data scales, but swap in a deliberately poorer value estimator (no history, or remaining-time targets scrambled) while keeping identical quality thresholds and residual training: if success rates no longer track the reliability metrics and value-guided methods stop beating behavior cloning and DAGGER, the central reliability claim fails.

Watch this falsifier — get emailed when new claim-graph text bears on it.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper presents Robo-ValueRL, a full-stack offline-to-online RL framework for robotic manipulation that centers on a history-conditioned value estimator. Value reliability is quantified via global Midpoint Ordering Rate (MOR) and local fluency/error-discrimination metrics (Bump Ratio/Magnitude, Error Sensitivity/Slope). These estimates supply action-quality labels for quality-conditioned consistency-policy VLA pretraining and for filtering online rollouts used to train a gated residual adapter. On two real-robot tasks (millimeter-level chip insertion and generalizable block disassembly), using 240 h of offline data and >3 000 online trajectories, the authors report that higher reliability metrics align with better offline scaling versus BC and more stable online gains versus DAGGER, reaching 86 % and 84 % success respectively.

Significance. If the reliability–performance association holds under cleaner isolation, the work supplies a practical diagnostic suite and an end-to-end recipe for converting mixed-quality robotic experience into policy improvement—an important gap relative to systems that only report final success rates. Strengths include large-scale real-robot evaluation, one-take continuous videos (30/35 and 58/70), public code/data/models, and explicit ablations of history length, quality thresholds, and residual adaptation. The remaining-time supervision target and the proposed metrics are concrete and falsifiable, giving the community a reusable testbed even if some causal claims require tightening.

major comments (3)
  1. The central claim that value-estimation reliability drives offline scaling and online stability is only partially isolated. Table 1 shows SHORTHISTORY highest on most metrics and on success, yet LONGHISTORY is competitive on Bump Mag. and Slope while collapsing to 30 %/42 % success—already weakening a pure reliability o performance story. The cleanest comparison (VG-STRICT vs VG-NH, same strict labeling rule) still confounds estimator architecture (5-frame history) with the reliability scores themselves. A controlled experiment that holds history length fixed and varies only metric-selected estimators, or that reports reliability of the exact value model used for each policy, is needed before the causal language in the abstract and §4.1–4.3 is fully warranted.
  2. Quality-selection hyperparameters are task-tuned (strict Δ=20 / top-30 % for chip insertion; soft Δ=60 / top-50 % for block disassembly; §4.2 and Fig. 5). Consequently the reported scaling gains of “reliable-value guided” methods over BC mix reliability with threshold choice. The paper should either (i) fix a single selection rule across both tasks and re-evaluate, or (ii) treat threshold selection as an explicit hyper-parameter study and show that the reliability metrics still predict the best threshold a priori. Without this, the strongest claim overstates what the ablations establish.
  3. Appendix C.1 (Eqs. 9–10) defines the value target as normalized remaining time with fixed T_max=5000 and an unspecified failure penalty C_fail. This is a free design choice whose sensitivity is never reported. If remaining-time progress poorly reflects true task value under partial observability or multi-step trade-offs (e.g., non-greedy disassembly in Fig. 8a), both the quality labels and the reliability metrics become misaligned with policy needs. A short ablation on C_fail / T_max, or comparison against an alternative progress signal, is required to support the weakest assumption of the pipeline.
minor comments (5)
  1. Notation for the residual gate and residual bound (r_max, g_t) appears in §3.3 without a clear statement of how r_max is chosen; a sentence in Appendix C.3 would help reproducibility.
  2. Figure 1 caption and main text use both “Globel” and “Global”; correct the typo.
  3. The textual quality prompts (“Quality: Low/Medium/High”) and the 10 % dropout to Medium (§C.2) are reasonable but their effect on the consistency head is never ablated; a brief note would strengthen the method section.
  4. Success rates in Table 1 and Fig. 5 lack error bars or number of evaluation trials; given the real-robot setting this information is important for assessing variance.
  5. Related-work discussion of concurrent world-value / GVL metrics (§2.3) could more explicitly contrast the proposed MOR/Bump suite against Value-Order Correlation and VIP smoothness on the same trajectories.

Circularity Check

0 steps flagged

No circular derivation: value targets, reliability metrics, and downstream success are independently defined; the reliability–performance link is an empirical association, not an identity by construction.

full rationale

Robo-ValueRL is an empirical systems paper, not a first-principles derivation. Value supervision is remaining-time progress with a fixed failure penalty (Appendix C.1, Eqs. 9–10), fixed before any policy training. Reliability metrics (MOR, Bump Ratio/Magnitude, Error Sensitivity/Slope, Sec. 3.4) are computed from value sequences against annotated subgoals and errors, independently of final success rates. Action-quality labels from value differences are then used to train quality-conditioned policies and residual adapters—the intended mechanism of value-guided RL, not a definitional identity between inputs and claimed outputs. Downstream success (chip insertion, block disassembly) is measured on real-world rollouts. No equation forces success rates from reliability scores; Table 1 and Figs. 5–6 report measured associations. Self-citations (authors’ prior robotics work) appear only as related work and are not load-bearing uniqueness theorems or smuggled ansatze. Confounding concerns (history length, task-tuned thresholds) affect causal isolation, not circularity of the derivation chain. The paper is self-contained against its own experimental benchmarks.

Axiom & Free-Parameter Ledger

5 free parameters · 4 axioms · 3 invented entities

The central claim rests on a progress-style value target, discrete quality thresholds, and several architectural choices that are not derived from first principles. Free parameters control history length, quality selection, and residual gating; domain assumptions treat remaining time as value and value differences as action quality; the invented metrics and residual module are the paper’s main analytic and engineering contributions.

free parameters (5)
  • visual history length (SHORT=5 frames vs LONG=30)
    Chosen by ablation on the same tasks; SHORTHISTORY is selected as most reliable and used for all main results.
  • quality-selection window Δ and percentile (Δ=20 top-30% strict vs Δ=60 top-50% soft)
    Task-dependent thresholds that directly determine which offline and online segments train the policy; selected via performance on chip vs block tasks.
  • value-difference margins δ+, δ− for discrete quality labels {0,1,2}
    Hand-set positive/negative margins that convert continuous Δv into low/neutral/high labels used for both pretraining and online filtering.
  • T_max=5000 and failure penalty C_fail in remaining-time target
    Normalize progress into [0,1] and down-weight failed trajectories; fixed constants that shape every value label.
  • distributional bins K=256, soft-target σ, residual bound r_max, loss weights λ_cons/λ_keep/λ_gate
    Architectural and optimization hyperparameters that affect value resolution and residual training stability.
axioms (4)
  • domain assumption Normalized remaining time (with failure penalty) is a valid scalar progress target for supervising a value function that ranks action quality.
    Appendix C.1 defines v*_t from trajectory length and C_fail; all quality labels and reliability metrics inherit this choice.
  • ad hoc to paper Thresholded value differences over fixed temporal windows are sufficient indicators of action quality for both offline filtering and online residual training.
    Section 3.1 Action Quality Indicator; the mapping is rule-based rather than learned or theoretically derived.
  • domain assumption Freezing the offline VLA and training only a gated residual adapter preserves the pretrained prior while allowing targeted correction.
    Section 3.3; standard residual-adaptation premise, not proven for the specific VLA backbone used.
  • domain assumption A consistency policy head with one-step denoising is an adequate action expert for real-time quality-conditioned control.
    Section 3.2; relies on prior consistency-policy results without new theoretical guarantees.
invented entities (3)
  • Midpoint Ordering Rate (MOR) and local fluency/error-discrimination metrics (Bump Ratio, Bump Magnitude, Error Sensitivity, Error Slope) no independent evidence
    purpose: Diagnose whether a value function captures global task-stage ordering and local action preference before expensive policy training.
    Section 3.4; new evaluation protocol relative to prior VIP/GVL-style metrics that mainly assessed optimal trajectories.
  • Quality-conditioned consistency-policy VLA with textual quality prompts no independent evidence
    purpose: Inject discrete action-quality labels into language-conditioned action generation so the policy can prefer high-quality modes at inference.
    Section 3.2; combines consistency policy with value-derived quality text.
  • Gated residual online adapter trained with value-filtered high-quality online segments plus offline keep loss no independent evidence
    purpose: Enable targeted online correction without catastrophic forgetting of the offline prior.
    Section 3.3; specific design of gate + residual + mixed keep/corr/gate losses.

pith-pipeline@v1.1.0-grok45 · 23686 in / 3293 out tokens · 56842 ms · 2026-07-14T14:57:03.182560+00:00 · methodology

0 comments
read the original abstract

Offline-to-online reinforcement learning is promising for generalizable robotic manipulation, yet its full-stack complexity obscures reproduction and diagnosis. Within such systems, value estimation plays a central role in prioritizing heterogeneous data for policy improvement. Despite its importance, the central question remains underexplored: how value-function reliability shapes policy optimization in offline-to-online reinforcement learning. To answer this question, we propose Robo-ValueRL, a unified framework that enables reliable value estimation and systematically traces its downstream effects on policy pretraining and online improvement. Concretely, Robo-ValueRL learns a history-conditioned value estimator and evaluates its reliability through global-progress and local-preference metrics. These resulting value estimates are propagated into quality-conditioned consistency-policy pretraining and a residual adaptation module on online rollouts, providing a unified testbed for analyzing how value reliability shapes downstream policy performance. Across 240 hours of offline demonstrations and over 3,000 online rollout trajectories, our extensive experiments show that downstream performance is strongly associated with value reliability. Reliable value functions provide better action-quality estimates, allowing value-guided offline RL to scale more effectively than quality-agnostic behavior cloning, and stabilize online improvement by prioritizing high-quality rollout data. Integrating reliable value guidance through offline pretraining with online improvement, our system achieves 86% success on millimeter-level precise chip insertion and 84% on generalizable block disassembly. We hope these findings highlight the importance of value-guided data utilization for effective policy improvement from heterogeneous robotic experience.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. $N_0$-VTLA: Scaling Vision-Tactile-Language-Action Model with Latent Tactile Tokens

    cs.RO 2026-07 conditional novelty 6.0

    A VLA with predictive latent tactile tokens pretrained on large-scale visuo-tactile data, plus ALTER offline advantage labeling, leads contact-rich real and sim benchmarks.

Reference graph

Works this paper leans on

56 extracted references · 8 linked inside Pith · cited by 1 Pith paper

  1. [1]

    Awac: Accelerating online reinforcement learning with offline datasets,

    A. Nair, A. Gupta, M. Dalal, and S. Levine, “Awac: Accelerating online reinforcement learning with offline datasets,”arXiv preprint arXiv:2006.09359, 2020

  2. [2]

    Conservative q-learning for offline reinforcement learning,

    A. Kumar, A. Zhou, G. Tucker, and S. Levine, “Conservative q-learning for offline reinforcement learning,”Advances in neural information processing systems, vol. 33, pp. 1179–1191, 2020

  3. [3]

    Cal-ql: Calibrated offline rl pre-training for efficient online fine-tuning,

    M. Nakamoto, S. Zhai, A. Singh, M. Sobol Mark, Y . Ma, C. Finn, A. Kumar, and S. Levine, “Cal-ql: Calibrated offline rl pre-training for efficient online fine-tuning,”Advances in Neural Information Processing Systems, vol. 36, pp. 62 244–62 269, 2023

  4. [4]

    RT-1: Robotics Transformer for Real-World Control at Scale,

    A. Brohan, N. Brown, J. Carbajal, Y . Chebotar, J. Dabis, C. Finn, K. Gopalakrishnan, K. Hausman, A. Herzog, J. Hsu, J. Ibarz, B. Ichter, A. Irpan, T. Jackson, S. Jesmonth, N. Joshi, R. Julian, D. Kalashnikov, Y . Kuang, I. Leal, K.-H. Lee, S. Levine, Y . Lu, U. Malla, D. Manjunath, I. Mordatch, O. Nachum, C. Parada, J. Peralta, E. Perez, K. Pertsch, J. Q...

  5. [5]

    Rt-2: Vision-language-action models transfer web knowledge to robotic control,

    B. Zitkovich, T. Yu, S. Xu, P. Xu, T. Xiao, F. Xia, J. Wu, P. Wohlhart, S. Welker, A. Wahidet al., “Rt-2: Vision-language-action models transfer web knowledge to robotic control,” inConference on Robot Learning. PMLR, 2023, pp. 2165–2183

  6. [6]

    Open x-embodiment: Robotic learning datasets and rt-x models: Open x-embodiment collaboration 0,

    A. O’Neill, A. Rehman, A. Maddukuri, A. Gupta, A. Padalkar, A. Lee, A. Pooley, A. Gupta, A. Mandlekar, A. Jainet al., “Open x-embodiment: Robotic learning datasets and rt-x models: Open x-embodiment collaboration 0,” in2024 IEEE International Conference on Robotics and Automation. IEEE, 2024, pp. 6892–6903

  7. [7]

    Openvla: An open-source vision-language- action model,

    M. J. Kim, K. Pertsch, S. Karamcheti, T. Xiao, A. Balakrishna, S. Nair, R. Rafailov, E. P. Foster, P. R. Sanketi, Q. Vuong, T. Kollar, B. Burchfiel, R. Tedrake, D. Sadigh, S. Levine, P. Liang, and C. Finn, “Openvla: An open-source vision-language- action model,” inProceedings of The 8th Conference on Robot Learning, ser. Proceedings of Machine Learning Re...

  8. [8]

    π0: A Vision-Language-Action Flow Model for General Robot Control,

    K. Black, N. Brown, D. Driess, A. Esmail, M. R. Equi, C. Finn, N. Fusai, L. Groom, K. Hausman, B. Ichter, S. Jakubczak, T. Jones, L. Ke, S. Levine, A. Li-Bell, M. Mothukuri, S. Nair, K. Pertsch, L. X. Shi, L. Smith, J. Tanner, Q. Vuong, A. Walling, H. Wang, and U. Zhilinsky, “ π0: A Vision-Language-Action Flow Model for General Robot Control,” inProceedin...

  9. [9]

    Gr00t n1: An open foundation model for generalist humanoid robots,

    J. Bjorck, F. Castañeda, N. Cherniadev, X. Da, R. Ding, L. Fan, Y . Fang, D. Fox, F. Hu, S. Huanget al., “Gr00t n1: An open foundation model for generalist humanoid robots,”arXiv preprint arXiv:2503.14734, 2025

  10. [10]

    Pre-Training for Robots: Offline RL Enables Learning New Tasks in a Handful of Trials,

    A. Kumar, A. Singh, F. D. Ebert, M. Nakamoto, Y . Yang, C. Finn, and S. Levine, “Pre-Training for Robots: Offline RL Enables Learning New Tasks in a Handful of Trials,” inProceedings of Robotics: Science and Systems, July 2023

  11. [11]

    Robogene: Boosting vla pre-training via diversity-driven agentic framework for real-world task generation,

    Y . Zhang, K. Wu, Z. Gao, Z. Zhao, P. Ren, Z. Xu, F. Liao, X. Wang, S. Fan, D. Wuet al., “Robogene: Boosting vla pre-training via diversity-driven agentic framework for real-world task generation,”arXiv preprint arXiv:2602.16444, 2026

  12. [12]

    Precise and dexterous robotic manipulation via human-in-the-loop reinforcement learning,

    J. Luo, C. Xu, J. Wu, and S. Levine, “Precise and dexterous robotic manipulation via human-in-the-loop reinforcement learning,” Science Robotics, vol. 10, no. 105, p. eads5033, 2025

  13. [13]

    Robot fine-tuning made easy: Pre-training rewards and policies for autonomous real-world reinforcement learning,

    J. Yang, M. S. Mark, B. Vu, A. Sharma, J. Bohg, and C. Finn, “Robot fine-tuning made easy: Pre-training rewards and policies for autonomous real-world reinforcement learning,” in2024 IEEE International Conference on Robotics and Automation. IEEE, 2024, pp. 4804–4811

  14. [14]

    SimpleVLA-RL: Scaling VLA training via reinforcement learning,

    H. Li, Y . Zuo, J. Yu, Y . Zhang, Y . Zhaohui, K. Zhang, X. Zhu, Y . Zhang, T. Chen, G. Cui, D. Wang, D. Luo, Y . Fan, Y . Sun, J. Zeng, J. Pang, S. Zhang, Y . Wang, Y . Mu, B. Zhou, and N. Ding, “SimpleVLA-RL: Scaling VLA training via reinforcement learning,” inThe Fourteenth International Conference on Learning Representations, 2026

  15. [15]

    Phoenix: A motion-based self-reflection framework for fine-grained robotic action correction,

    W. Xia, R. Feng, D. Wang, and D. Hu, “Phoenix: A motion-based self-reflection framework for fine-grained robotic action correction,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2025, pp. 6981–6990. 12

  16. [16]

    Human-assisted robotic policy refinement via action preference optimization,

    W. Xia, Y . Yang, H. Wu, X. Ma, T. Kong, and D. Hu, “Human-assisted robotic policy refinement via action preference optimization,” inAdvances in Neural Information Processing Systems, D. Belgrave, C. Zhang, H. Lin, R. Pascanu, P. Koniusz, M. Ghassemi, and N. Chen, Eds., vol. 38. Curran Associates, Inc., 2025, pp. 36 746–36 768

  17. [17]

    Geco-srt: Geometry-aware continual adaptation for cross-task sim-to-real transfer,

    W. Yu, W. Xia, W. Zhang, and D. Hu, “Geco-srt: Geometry-aware continual adaptation for cross-task sim-to-real transfer,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, June 2026, pp. 42 408–42 417

  18. [18]

    Deep reinforcement learning that matters,

    P. Henderson, R. Islam, P. Bachman, J. Pineau, D. Precup, and D. Meger, “Deep reinforcement learning that matters,” in Proceedings of the AAAI conference on artificial intelligence, vol. 32, no. 1, 2018

  19. [19]

    Deep reinforcement learning at the edge of the statistical precipice,

    R. Agarwal, M. Schwarzer, P. S. Castro, A. C. Courville, and M. Bellemare, “Deep reinforcement learning at the edge of the statistical precipice,”Advances in neural information processing systems, vol. 34, pp. 29 304–29 320, 2021

  20. [20]

    Learning while deploying: Fleet-scale reinforcement learning for generalist robot policies,

    Y . Wang, X. Li, P. Xie, P. Yang, B. Nie, Y . Cai, Q. Zhang, C. Qu, J. Wu, J. Songet al., “Learning while deploying: Fleet-scale reinforcement learning for generalist robot policies,”arXiv preprint arXiv:2605.00416, 2026

  21. [21]

    Suf: Stabilized unconstrained fine-tuning for offline-to-online reinforcement learning,

    J. Feng, M. Feng, H. Song, W. Zhou, and H. Li, “Suf: Stabilized unconstrained fine-tuning for offline-to-online reinforcement learning,” inProceedings of the AAAI Conference on Artificial Intelligence, vol. 38, no. 11, 2024, pp. 11 961–11 969

  22. [22]

    Offline-to-online reinforcement learning via balanced replay and pessimistic q-ensemble,

    S. Lee, Y . Seo, K. Lee, P. Abbeel, and J. Shin, “Offline-to-online reinforcement learning via balanced replay and pessimistic q-ensemble,” inConference on Robot Learning. PMLR, 2022, pp. 1702–1712

  23. [23]

    pi*0.6: a vla that learns from experience,

    P. Intelligence, A. Amin, R. Aniceto, A. Balakrishna, K. Black, K. Conley, G. Connors, J. Darpinian, K. Dhabalia, J. DiCarlo et al., “pi*0.6: a vla that learns from experience,”arXiv preprint arXiv:2511.14759, 2025

  24. [24]

    What matters in learning from offline human demonstrations for robot manipulation,

    A. Mandlekar, D. Xu, J. Wong, S. Nasiriany, C. Wang, R. Kulkarni, L. Fei-Fei, S. Savarese, Y . Zhu, and R. Martín-Martín, “What matters in learning from offline human demonstrations for robot manipulation,” inProceedings of the 5th Conference on Robot Learning, A. Faust, D. Hsu, and G. Neumann, Eds., vol. 164. PMLR, 08–11 Nov 2022, pp. 1678–1690

  25. [25]

    Scalable deep reinforcement learning for vision-based robotic manipulation,

    D. Kalashnikov, A. Irpan, P. Pastor, J. Ibarz, A. Herzog, E. Jang, D. Quillen, E. Holly, M. Kalakrishnan, V . Vanhouckeet al., “Scalable deep reinforcement learning for vision-based robotic manipulation,” inConference on robot learning. PMLR, 2018, pp. 651–673

  26. [26]

    Gr-rl: Going dexterous and precise for long-horizon robotic manipulation,

    Y . Li, X. Ma, J. Xu, Y . Cui, Z. Cui, Z. Han, L. Huang, T. Kong, Y . Liu, H. Niuet al., “Gr-rl: Going dexterous and precise for long-horizon robotic manipulation,”arXiv preprint arXiv:2512.01801, 2025

  27. [27]

    VIP: Towards universal visual reward and representation via value-implicit pre-training,

    Y . J. Ma, S. Sodhani, D. Jayaraman, O. Bastani, V . Kumar, and A. Zhang, “VIP: Towards universal visual reward and representation via value-implicit pre-training,” inThe Eleventh International Conference on Learning Representations, 2023

  28. [28]

    Conrft: A reinforced fine-tuning method for vla models via consistency policy,

    Y . Chen, S. Tian, S. Liu, Y . Zhou, H. Li, and D. Zhao, “Conrft: A reinforced fine-tuning method for vla models via consistency policy,” inProceedings of Robotics: Science and Systems, 2025, Los Angeles, CA, USA, Jun 21-25, 2025, 2025

  29. [29]

    Consistency policy: Accelerated visuomotor policies via consistency distillation,

    A. Prasad, K. Lin, J. Wu, L. Zhou, and J. Bohg, “Consistency policy: Accelerated visuomotor policies via consistency distillation,” inRobotics: Science and Systems, 2024

  30. [30]

    Learning from imperfect demonstrations from agents with varying dynamics,

    Z. Cao and D. Sadigh, “Learning from imperfect demonstrations from agents with varying dynamics,”IEEE Robotics and Automation Letters, vol. 6, no. 3, pp. 5231–5238, 2021

  31. [31]

    Curating Demonstrations using Online Experience,

    A. S. Chen, A. M. Lessing, Y . Liu, and C. Finn, “Curating Demonstrations using Online Experience,” inProceedings of Robotics: Science and Systems, LosAngeles, CA, USA, June 2025

  32. [32]

    Octo: An Open-Source Generalist Robot Policy,

    D. Ghosh, H. R. Walke, K. Pertsch, K. Black, O. Mees, S. Dasari, J. Hejna, T. Kreiman, C. Xu, J. Luo, Y . L. Tan, L. Y . Chen, Q. Vuong, T. Xiao, P. R. Sanketi, D. Sadigh, C. Finn, and S. Levine, “Octo: An Open-Source Generalist Robot Policy,” in Proceedings of Robotics: Science and Systems, Delft, Netherlands, July 2024

  33. [33]

    Xr-1: Towards versatile vision- language-action models via learning unified vision-motion representations,

    S. Fan, K. Wu, Z. Che, X. Wang, D. Wu, F. Liao, N. Liu, Y . Zhang, Z. Zhao, Z. Xuet al., “Xr-1: Towards versatile vision- language-action models via learning unified vision-motion representations,” inProceedings of the International Conference on Machine Learning, 2026

  34. [34]

    A survey on offline reinforcement learning: Taxonomy, review, and open problems,

    R. F. Prudencio, M. R. Maximo, and E. L. Colombini, “A survey on offline reinforcement learning: Taxonomy, review, and open problems,”IEEE transactions on neural networks and learning systems, vol. 35, no. 8, pp. 10 237–10 257, 2023

  35. [35]

    Offline reinforcement learning with implicit q-learning,

    I. Kostrikov, A. Nair, and S. Levine, “Offline reinforcement learning with implicit q-learning,” inInternational Conference on Learning Representations, 2022

  36. [36]

    Policy expansion for bridging offline-to-online reinforcement learning,

    H. Zhang, W. Xu, and H. Yu, “Policy expansion for bridging offline-to-online reinforcement learning,” inThe Eleventh International Conference on Learning Representations, 2023. 13

  37. [37]

    Q-transformer: Scalable offline reinforcement learning via autoregressive q-functions,

    Y . Chebotar, Q. Vuong, K. Hausman, F. Xia, Y . Lu, A. Irpan, A. Kumar, T. Yu, A. Herzog, K. Pertschet al., “Q-transformer: Scalable offline reinforcement learning via autoregressive q-functions,” inConference on Robot Learning. PMLR, 2023, pp. 3909–3928

  38. [38]

    Finetuning from offline reinforcement learning: Challenges, trade-offs and practical solutions,

    Y . Luo, J. Kay, E. Grefenstette, and M. P. Deisenroth, “Finetuning from offline reinforcement learning: Challenges, trade-offs and practical solutions,”arXiv preprint arXiv:2303.17396, 2023

  39. [39]

    Learning to predict by the methods of temporal differences,

    R. S. Sutton, “Learning to predict by the methods of temporal differences,”Machine learning, vol. 3, no. 1, pp. 9–44, 1988

  40. [40]

    Analysis of temporal-diffference learning with function approximation,

    J. Tsitsiklis and B. Van Roy, “Analysis of temporal-diffference learning with function approximation,”Advances in neural information processing systems, vol. 9, 1996

  41. [41]

    Fast gradient-descent methods for temporal-difference learning with linear function approximation,

    R. S. Sutton, H. R. Maei, D. Precup, S. Bhatnagar, D. Silver, C. Szepesvári, and E. Wiewiora, “Fast gradient-descent methods for temporal-difference learning with linear function approximation,” inProceedings of the 26th annual international conference on machine learning, 2009, pp. 993–1000

  42. [42]

    Human-level control through deep reinforcement learning,

    V . Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. Riedmiller, A. K. Fidjeland, G. Ostrovskiet al., “Human-level control through deep reinforcement learning,”nature, vol. 518, no. 7540, pp. 529–533, 2015

  43. [43]

    Benchmarking deep reinforcement learning for continuous control,

    Y . Duan, X. Chen, R. Houthooft, J. Schulman, and P. Abbeel, “Benchmarking deep reinforcement learning for continuous control,” inInternational conference on machine learning. PMLR, 2016, pp. 1329–1338

  44. [44]

    Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor,

    T. Haarnoja, A. Zhou, P. Abbeel, and S. Levine, “Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor,” inInternational conference on machine learning. Pmlr, 2018, pp. 1861–1870

  45. [45]

    A vision-language-action-critic model for robotic real-world reinforcement learning,

    S. Zhai, Q. Zhang, T. Zhang, F. Huang, H. Zhang, M. Zhou, S. Zhang, L. Liu, S. Lin, and J. Pang, “A vision-language-action-critic model for robotic real-world reinforcement learning,”arXiv preprint arXiv:2509.15937, 2025

  46. [46]

    RL-VLM-f: Reinforcement learning from vision lan- guage foundation model feedback,

    Y . Wang, Z. Sun, J. Zhang, Z. Xian, E. Biyik, D. Held, and Z. Erickson, “RL-VLM-f: Reinforcement learning from vision lan- guage foundation model feedback,” inProceedings of the 41st International Conference on Machine Learning, ser. Proceedings of Machine Learning Research, vol. 235. PMLR, 2024, pp. 51 484–51 501

  47. [47]

    Rank2reward: Learning shaped reward functions from passive video,

    D. Yang, D. Tjia, J. Berg, D. Damen, P. Agrawal, and A. Gupta, “Rank2reward: Learning shaped reward functions from passive video,” in2024 IEEE International Conference on Robotics and Automation. IEEE, 2024, pp. 2806–2813

  48. [48]

    Vision language models are in-context value learners,

    Y . J. Ma, J. Hejna, C. Fu, D. Shah, J. Liang, Z. Xu, S. Kirmani, P. Xu, D. Driess, T. Xiaoet al., “Vision language models are in-context value learners,” inInternational Conference on Learning Representations, vol. 2025, 2025, pp. 33 984–34 009

  49. [49]

    Universal value function approximators,

    T. Schaul, D. Horgan, K. Gregor, and D. Silver, “Universal value function approximators,” inInternational conference on machine learning. PMLR, 2015, pp. 1312–1320

  50. [50]

    Viva: A video-generative value model for robot reinforcement learning,

    J. Lv, H. Li, J. Li, Y . Nie, F. Kong, Y . Wang, X. Wang, Z. Zhu, C. Ni, Q. Denget al., “Viva: A video-generative value model for robot reinforcement learning,”arXiv preprint arXiv:2604.08168, 2026

  51. [51]

    World value models for robotic manipulation,

    Z. Wang, J. Li, Y . Cui, Y . Gao, X. Zhan, J. Yu, and X. Ma, “World value models for robotic manipulation,”arXiv preprint arXiv:2606.24742, 2026

  52. [52]

    Paligemma: A versatile 3b vlm for transfer,

    L. Beyer, A. Steiner, A. S. Pinto, A. Kolesnikov, X. Wang, D. Salz, M. Neumann, I. Alabdulmohsin, M. Tschannen, E. Bugliarello et al., “Paligemma: A versatile 3b vlm for transfer,”arXiv preprint arXiv:2407.07726, 2024

  53. [53]

    Sigmoid loss for language image pre-training,

    X. Zhai, B. Mustafa, A. Kolesnikov, and L. Beyer, “Sigmoid loss for language image pre-training,” inProceedings of the IEEE/CVF international conference on computer vision, 2023, pp. 11 975–11 986

  54. [54]

    Perceiver: General perception with iterative attention,

    A. Jaegle, F. Gimeno, A. Brock, O. Vinyals, A. Zisserman, and J. Carreira, “Perceiver: General perception with iterative attention,” inInternational conference on machine learning. PMLR, 2021, pp. 4651–4664

  55. [55]

    Stop regressing: Training value functions via classification for scalable deep RL,

    J. Farebrother, J. Orbay, Q. Vuong, A. Ali Taiga, Y . Chebotar, T. Xiao, A. Irpan, S. Levine, P. S. Castro, A. Faust, A. Kumar, and R. Agarwal, “Stop regressing: Training value functions via classification for scalable deep RL,” inProceedings of the 41st International Conference on Machine Learning, ser. Proceedings of Machine Learning Research, R. Salakh...

  56. [56]

    A reduction of imitation learning and structured prediction to no-regret online learning,

    S. Ross, G. Gordon, and D. Bagnell, “A reduction of imitation learning and structured prediction to no-regret online learning,” inProceedings of the Fourteenth International Conference on Artificial Intelligence and Statistics, ser. Proceedings of Machine Learning Research, G. Gordon, D. Dunson, and M. Dudík, Eds., vol. 15. Fort Lauderdale, FL, USA: PMLR,...