Pith. sign in

REVIEW 2 major objections 5 minor 17 references

Counterfactual audit finds no learned adapter worth deploying on the tested frozen locomotion interfaces.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-01 06:25 UTC pith:4ASIDTAD

load-bearing objection The matched-frequency comparator is a genuinely useful attribution tool, and the paper's central negative result is honestly earned — but the headline claim is scoped to an observation-only selector, not the adapters it names. the 2 major comments →

arxiv 2607.21867 v1 pith:4ASIDTAD submitted 2026-07-23 cs.AI cs.SYeess.SY

When Is a Learned Command Adapter Worth It? Closed-Loop Identification and Counterfactual Auditing of Frozen Locomotion Policies

classification cs.AI cs.SYeess.SY
keywords command adapterfrozen policylegged locomotioncounterfactual evaluationstate-allocation gainclosed-loop identificationGO/NO-GO/ABSTAIN auditpersonalization testing
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Learned command adapters are often bolted onto frozen locomotion policies, but this paper asks a prior question: is the improvement real at the same state and recoverable from deployment-time observations? It formalizes an adapter necessity audit that separates global operating-point gain, same-state counterfactual headroom, state-allocation gain over a frequency-matched randomized mixture, and deployment gain over a cross-fitted fixed reference, then maps them to GO/NO-GO/ABSTAIN. The empirical punchline is that local opportunity exists—an exact-replay oracle finds 5.2% work headroom on a three-scale prefix set—but a learned selector recovers only 0.55% allocation gain once action frequencies are matched. The confirmatory twenty-cluster audit returns no real-domain GO: direct Go2 and direct H1 are NO-GO, VGCC and MPC query distributions ABSTAIN. The paper's central structural claim is that Halloc, the allocation gain over a matched mixture, is the quantity that separates genuine state-dependent personalization from a mere global operating-point shift.

Core claim

The paper claims that whether a learned command adapter is worth deploying on a frozen, command-conditioned locomotion policy is a counterfactual question with three separable parts: is there same-state headroom, can a learned selector recover it from deployment features, and does the recovery beat both a fixed reference and a randomized policy with the same action frequencies. The central empirical claim is that on the tested benchmark the answers are 'yes, no, and not enough': an exact-replay oracle finds 5.2% work headroom on a three-scale prefix set, but a leave-one-source-seed-out treatment-effect selector recovers only 0.55% state-allocation gain Halloc, and the confirmatory audit with

What carries the argument

Halloc (Eq. 5) is the load-bearing quantity: Halloc = 1 − E_x[W_x(π̂(x))] / E_{x,a∼p}[W_x(a)], where the denominator is the expected work of a randomized policy that draws actions independently of context with exactly the selector's marginal action frequencies. Because the mixture shares the selector's action composition, Halloc removes gains from changing how often each command is issued; it is positive only when the selector assigns actions to contexts where those actions do better, which the paper treats as the definition of genuine state-dependent adaptation. The accompanying three-way decision rule (Eq. 8) combines deployment gain over a cross-fitted fixed reference, Halloc, and a reali

Load-bearing premise

The audit's conclusions rest on the assumption that replaying a candidate command from the same recorded start state in simulation gives exactly what the frozen policy would do in deployment; if physics noise, hidden state, or the sim-to-real gap breaks that replay, the NO-GO and ABSTAIN decisions are about the wrong counterfactuals.

What would settle it

Take a subset of the 960 Go2/direct query states to physical hardware or a second physics engine, replay each of the seven interventions from identical start states, and compare the ordering of measured mechanical work and success; if the real ordering differs from the simulated ordering, or if the allocation gain with matched frequencies exceeds the 1% threshold, the audit's counterfactual foundation and its NO-GO decision are falsified.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • If the audit is run before adding an adapter, it would have blocked deployment of a learned selector on direct-command Go2 and deployment-representative H1 at 1% thresholds: both return NO-GO.
  • Measured same-state headroom does not imply recoverable value: 5.2% oracle headroom collapses to 0.55% allocation gain, so a positive local-opportunity screen is not sufficient evidence for an adapter.
  • VGCC's apparent gains over direct control are not certified as personalization: although its mean deployment gain reaches 1.34%, its allocation lower bound is 0.09% and its violation upper bound is 6.25%, so the audit returns ABSTAIN.
  • More data alone cannot convert the unresolved rows into GO: the resolution diagnostic shows that with allocation point estimates below 1%, collecting more clusters can narrow bounds but cannot raise the allocation lower bound above the threshold.
  • Average prediction accuracy of the closed-loop model does not certify ranking value: the identified response model is accurate (displacement R² ≥ 0.92, posture ≈ 0.9, cost 0.75–0.8), but the feature ablation finds the observation alone recovers nearly the same limited allocation gain.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • Beyond the paper: the same audit logic transfers to any frozen command-conditioned interface with redundant commands and fast closed-loop response—end-effector servoing, vehicle tracking, or reference-governor settings—provided one can specify an intervention set, an outcome, and an independent deployment unit; exact replay would need replacement with logged or randomized interventions when resets
  • Beyond the paper: the large gap between oracle headroom and recovered allocation gain suggests a practical engineering rule: before investing in adapters, compare the candidate selector against its own frequency-matched random policy; a selector that cannot beat that mixture is not personalizing, no matter how much total gain it shows.
  • Beyond the paper: the NO-GO/ABSTAIN results are stated in simulation with exact replay; if the sim-to-real gap breaks replay fidelity, the decisions would need to be re-audited on hardware, a transfer the paper itself lists as untested.
  • Beyond the paper: the 1% practical-value thresholds are retained via hidden-signal calibration and are not claimed to be universal; a different deployment domain should re-calibrate them, since the paper explicitly warns that thresholds, cluster counts, and action families should be preregistered per domain.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 5 minor

Summary. The paper proposes an adapter necessity audit for frozen, command-conditioned policies. It formalizes four estimands—global operating-point gain, same-state oracle headroom Havail, deployment gain Hdep over a cross-fitted fixed action, and state-allocation gain Halloc over a frequency-matched randomized mixture (Eqs. 4–7)—and a GO/NO-GO/ABSTAIN rule (Eq. 8) driven by cluster-bootstrap learner refits. A closed-loop command-response model is introduced as an optional feature source. On Isaac Lab rough-terrain benchmarks, a scale-prefix diagnostic reports 5.2% same-state work headroom but only 0.55% recovered allocation gain. The confirmatory twenty-cluster, 200-refit audit across direct, scale, heading, and yaw interventions returns NO-GO for Go2/direct and H1/direct, ABSTAIN for Go2/VGCC and Go2/MPC query distributions, and GO for a learner-level synthetic control; the paper's headline is that no real-domain row returns GO.

Significance. If the audit is valid, this is a useful methodological contribution: it separates state-dependent allocation from action-frequency effects, uses exact replay to access potential outcomes in simulation, and embeds uncertainty in a three-way decision with explicit abstention. The empirical work is unusually careful: cross-fitted fixed reference, cluster bootstrap with full learner refits, prespecified thresholds, an excluded non-representative pilot, descriptive/confirmatory labeling, and positive/negative controls. The paper also reports several null or mixed results honestly (VGCC does not beat the fixed-scaling frontier; identified-model features do not decisively improve recovery). The main weakness is scope: the confirmatory decisions are produced by an observation-only selector, while the adapters named in the abstract (VGCC/MPC) use identified-model features. The archived feature ablation cannot close that gap at confirmatory scale. This representativeness concern is load-bearing for the headline negative conclusion, but it is fixable by additional confirmatory runs or by systematically qualifying the abstract and Table 6.

major comments (2)
  1. [§4.7 Stage 3; Appendix A ('Action-aligned confirmatory protocol'); Table 6] The confirmatory learner is explicitly observation-only, so the rows 'Go2/VGCC' and 'Go2/MPC' in Table 6 are query-distribution labels, not evaluations of VGCC/MPC as adapters. VGCC and MPC are driven by fφ/Fη features, which are not used in the confirmatory audit at cluster-refit scale. The abstract's 'VGCC and MPC queries ABSTAIN' and 'no real-domain row returns GO' are therefore easy to misread as adapter-level verdicts. Because Eq. 8 requires L_A > δalloc for GO, a selector using Fη could in principle cross the 1% allocation threshold at twenty clusters; Table 11 does not exclude this, as it is limited to five seeds, resamples fitted fold summaries, and its combined fφ+Fη point estimate (0.76%) is explicitly post hoc. Please either run the confirmatory protocol with the identified-model features or, at minimum, state in the abstract, in Section 4.7, and in the Table 6 caption that al
  2. [§4.7 'Why both comparators matter'; Table 9] The deployment and allocation values reported for the VGCC/MPC rows belong to a conservative treatment-effect learner, not to VGCC or MPC. Table 9 shows this selector activates on 30–39% of queries and matches the same-state oracle on only 13–16% of them. Consequently, the statement that 'the selector improves by 1.34% over direct control' (VGCC row) describes the audited learner on states visited by VGCC, not VGCC's own candidate scoring and bounded blend. The limitations section eventually scopes the result to 'the present representation and data,' but the title and abstract ask when a learned command adapter is worth it. The paper should either test a selector with the adapter's actual features/decision rule in the confirmatory protocol or explicitly say that the audit adjudicates only an observation-only surrogate selector.
minor comments (5)
  1. [Abstract] 'No real-domain row returns GO' should be qualified: no row in Table 6 under the observation-only selector. The current wording invites overgeneralization.
  2. [Table 6 caption] Add a footnote: 'Decisions are for the observation-only treatment-effect learner; rows are query-state distributions induced by the named controller.' This would prevent the misreading discussed in the major comments.
  3. [§4.7 'Replay is exact'] The validity of all potential-outcome quantities rests on bitwise deterministic replay. Please document the simulator determinism settings (fixed RNG streams per branch, no thread nondeterminism, solver settings) in the reproducibility section, or provide a small sensitivity analysis.
  4. [§4.1 / Reproducibility] The repository is described but no URL or DOI is given; include a link for archival reproducibility.
  5. [§4.7 'Decision-resolution diagnostic'] The N^{-1/2} width-scaling calculation is clearly labeled a design aid; this is good. Consider also reporting the width scaling assumption's sensitivity to non-Gaussian tails, since the bootstrap intervals are asymmetric.

Circularity Check

0 steps flagged

No significant circularity: the audit estimates are measured counterfactual outcomes with cross-fitted comparators, not fitted inputs relabeled as predictions.

full rationale

The claimed derivation chain runs from the formal definitions of H_global, H_avail, H_alloc, and H_dep (Eqs. 4–7) through the decision rule (Eq. 8) to the empirical Table 6 decisions. H_alloc is defined as a ratio of measured potential outcomes to a frequency-matched randomized policy; the covariance identity in Eq. 6 is an algebraic restatement of that definition, not an independent assumption. Potential outcomes are measured by exact replay: 'In simulation we observe these potential outcomes by exact start-state replay' (Sec. 3.3). Learner comparisons are cross-fitted: 'the fixed reference is selected from training clusters only, and 200 source-cluster bootstrap replicates refit the complete learner and evaluate it on out-of-bootstrap clusters' (Sec. 4.7). The closed-loop response model is explicitly optional at audit time: 'The confirmatory learner uses observation-only features, thereby avoiding dependence on an auxiliary predictive model at audit time' (Appendix A). The synthetic positive control is constructed with planted signal and is used only to verify procedural reachability, not to certify real-domain gains. The main negative result, 'No real-domain row returns GO,' is a summary of measured bootstrap intervals, not a value fit to those intervals. The only caveat is a scoping mismatch rather than circularity: Table 6 decisions apply to an observation-only selector, while VGCC/MPC also use identified-model features; the paper itself states the combined f_phi+F_eta estimate is 'not a nested or independently confirmed model-selection result' (Appendix A.2). That is a representativeness limitation, not a derivation that reduces to its own inputs.

Axiom & Free-Parameter Ledger

7 free parameters · 6 axioms · 3 invented entities

No new physical entities (forces, particles, dimensions) are postulated. The free parameters are controller and selector hyperparameters and user thresholds; the key axioms are the simulator-faithfulness and exact-replay assumptions that make the counterfactual measurements valid.

free parameters (7)
  • VGCC correction gain α = 0.5
    Bounded blend toward selected candidate; selected on Go2 development seed and fixed; ablation (B.3) shows substitution α=1 costs 5.0 success points.
  • Anneal distance d_anneal = 0.5 m
    Correction fades inside 0.5 m of target; Appendix A controller constants.
  • Progress floor β = 0.9
    Caps instantaneous slowdown at ~10%; removing it raises proxy saving to 19.6% but cuts success by 3.7 points (§4.6).
  • Cost margin ε = 0.1
    Rejects model noise before intervening (Eq. 11).
  • Posture thresholds (eligibility, rescue, restore, δ) = quads (0.85,0.80,0.90), δ=2cm; H1 (0.93,0.90,0.96), δ=0.8cm
    H1 thresholds tightened after observing posture degradation; morphology-dependent (§4.1).
  • Selector eligibility lower bound = 0.5
    Treatment-effect selector only intervenes if predicted eligibility ≥0.5 and work/time upper bounds pass; hand-set conservative criterion (§4.7 Stage 2).
  • Audit thresholds δdep, δalloc, κ = 1%, 1%, 5%
    User-specified practical-value and violation thresholds, prespecified for confirmatory audit; threshold sensitivity reported in A.1.
axioms (6)
  • domain assumption Exact start-state replay in Isaac Lab yields valid potential outcomes (W, T, S) for counterfactual commands on the frozen policy.
    All audit quantities in §3.3 (Havail, Halloc, Hdep, q) are computed from these replayed outcomes; replay is claimed exact after freezing the terrain curriculum (§4.7, Appendix A).
  • domain assumption The frozen policy's behavior is deterministic (or fully reconstructible) after reset; no hidden state or physics stochasticity invalidates the counterfactual branches.
    Exact replay requires that replaying the recorded direct-control prefix leads to the same observation with zero error (§4.7, Appendix A).
  • domain assumption Simulation (Isaac Lab) is a faithful proxy for real deployment on Go2/ANYmal-D/H1.
    All results are simulated; §5 states hardware power measurement remains necessary, and scope excludes hardware transfer.
  • domain assumption The command interface has redundancy and fast closed-loop response for efficiency improvements.
    Scope conditions in §3.1; if false, the audit is inapplicable.
  • domain assumption Public RSL-RL checkpoints are competent, representative frozen policies.
    Used as the base policies on all three embodiments (§4.1).
  • standard math Covariance identity Eq. 6 (selector–mixture attribution) is standard algebra.
    Halloc decomposition follows from expectation algebra; no additional assumption beyond finite expectations.
invented entities (3)
  • Viability-Gated Command Compensation (VGCC) no independent evidence
    purpose: Concrete learned adapter case study: bounded, gate-certified command correction toward lowest-cost candidate from a structured set (~70 candidates)
    A control algorithm, not a physical entity; its value is exactly what the audit tests, so it provides no independent evidence outside the paper.
  • Closed-loop command-response model fϕ no independent evidence
    purpose: Optional decision features predicting short-horizon motion, cost, and terrain-relative posture for candidate commands
    A supervised regression model; accuracy reported but the audit explicitly tests whether it improves selection beyond observations.
  • Matched-frequency randomized mixture µ̂π no independent evidence
    purpose: Comparator that draws actions independently of context with the selector's marginal frequencies, isolating state-allocation gain
    A statistical estimand, not an entity; used to define Halloc (Eq. 5).

pith-pipeline@v1.3.0-alltime-deepseek · 24384 in / 14311 out tokens · 132259 ms · 2026-08-01T06:25:06.687043+00:00 · methodology

0 comments
read the original abstract

Adding a learned adapter to a frozen, command-conditioned locomotion policy is worthwhile only if the interface exposes improvements that are both real and recoverable from deployment-time observations. We introduce an adapter necessity audit that separates global operating-point gain,same-state counterfactual headroom, deployment gain over a cross-fitted fixed action, and state-allocation gain over a frequency-matched randomized policy. Source-cluster learner refits map these quantities and constraint violations to a GO/NO-GO/ABSTAIN decision. Closed-loop command- response identification provides optional decision features. On Go2, an archived scale-prefix diagnostic finds 5.2% same-state headroom but only 0.55% recovered allocation gain. Our confirmatory audit evaluates direct, scale, heading, and yaw interventions on twenty independent clusters for each of three query distributions induced by direct control, VGCC, and MPC, using 200 full learner refits. At 1% deployment and allocation thresholds and a 5% violation tolerance, direct queries return NO-GO, while VGCC and MPC queries ABSTAIN. VGCC has the largest mean deployment gain (1.34%), but its allocation lower bound is 0.09% and its violation upper bound is 6.25%. A deployment-representative twenty-cluster H1 audit also returns NO-GO, whereas a learner-level synthetic control returns GO. The audit therefore tests whether observable signal justifies state-dependent adaptation rather than presuming that an adapter is valuable.

Figures

Figures reproduced from arXiv: 2607.21867 by ZongTan Li.

Figure 1
Figure 1. Figure 1: Adapter necessity audit. Stage 1 measures opportunity ( [PITH_FULL_IMAGE:figures/full_fig_p005_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: VGCC overview (case study under the audit). Offline excitation rollouts train an observation-conditioned [PITH_FULL_IMAGE:figures/full_fig_p007_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: The three embodiments in their Isaac Lab rough-terrain velocity environments: the Go2 and ANYmal-D [PITH_FULL_IMAGE:figures/full_fig_p008_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: Paired per-seed differences (VGCC − direct) for all evaluation seeds. The effort-proxy difference is negative for 17/17 seeds (two-sided exact sign test p = 1.53 × 10−5 ). Success differences straddle zero. Points are descriptive seed aggregates; the pooled sign test combines heterogeneous families. torque-derived mechanical-power channel used for the physical evaluations is identified with the same protoc… view at source ↗
Figure 5
Figure 5. Figure 5: Cost of transport versus completion time on six prespecified seeds (Go2 harder targets, [PITH_FULL_IMAGE:figures/full_fig_p011_5.png] view at source ↗
Figure 6
Figure 6. Figure 6: Stage 1 of the adapter necessity audit, on paired-successful episodes (identical initial states). [PITH_FULL_IMAGE:figures/full_fig_p012_6.png] view at source ↗
Figure 7
Figure 7. Figure 7: Descriptive outcome-level calibration on ten source clusters. [PITH_FULL_IMAGE:figures/full_fig_p015_7.png] view at source ↗
Figure 8
Figure 8. Figure 8: XY paths on the eight harder targets (one evaluation seed, 10 trials per method). VGCC and direct control [PITH_FULL_IMAGE:figures/full_fig_p023_8.png] view at source ↗
Figure 9
Figure 9. Figure 9: Selected paired episode illustrating rescue behavior. This example shows mechanism, not prevalence. [PITH_FULL_IMAGE:figures/full_fig_p023_9.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

17 extracted references · 17 linked inside Pith

  1. [1]

    Residual policy learning.arXiv preprint arXiv:1812.06298,

    Tom Silver, Kelsey Allen, Josh Tenenbaum, and Leslie Kaelbling. Residual policy learning.arXiv preprint arXiv:1812.06298,

  2. [4]

    arXiv:2109.11978. 16 Auditing Learned Adapters for Frozen PoliciesA PREPRINT Viktor Makoviychuk, Lukasz Wawrzyniak, Yunrong Guo, Michelle Lu, Kier Storey, Miles Macklin, David Hoeller, Nikita Rudin, Arthur Allshire, Ankur Handa, and Gavriel State. Isaac Gym: High performance GPU based physics simulation for robot learning. InNeurIPS Datasets and Benchmarks Track,

  3. [5]

    arXiv:2108.10470. Mayank Mittal, Calvin Yu, Qinxi Yu, Jingzhou Liu, Nikita Rudin, David Hoeller, Jia Lin Yuan, Ritvik Singh, Yunrong Guo, Hammad Mazhar, Ajay Mandlekar, Buck Babich, Gavriel State, Marco Hutter, and Animesh Garg. Orbit: A unified simulation framework for interactive robot learning environments.IEEE Robotics and Automation Letters, 8 (6):3740–3747,

  4. [10]

    Anusha Nagabandi, Gregory Kahn, Ronald S

    arXiv:2203.04955. Anusha Nagabandi, Gregory Kahn, Ronald S. Fearing, and Sergey Levine. Neural network dynamics for model-based deep reinforcement learning with model-free fine-tuning. InIEEE International Conference on Robotics and Automation (ICRA), pages 7559–7566,

  5. [11]

    arXiv:1708.02596. Kim P. Wabersich and Melanie N. Zeilinger. A predictive safety filter for learning-based control of constrained nonlinear dynamical systems.Automatica, 129:109597,

  6. [14]

    Archit Sharma, Shixiang Gu, Sergey Levine, Vikash Kumar, and Karol Hausman

    arXiv:1802.06070. Archit Sharma, Shixiang Gu, Sergey Levine, Vikash Kumar, and Karol Hausman. Dynamics-aware unsupervised discovery of skills. InInternational Conference on Learning Representations (ICLR),

  7. [15]

    Xue Bin Peng, Yunrong Guo, Lina Halper, Sergey Levine, and Sanja Fidler

    arXiv:1907.01657. Xue Bin Peng, Yunrong Guo, Lina Halper, Sergey Levine, and Sanja Fidler. ASE: Large-scale reusable adversarial skill embeddings for physically simulated characters. InACM Transactions on Graphics (SIGGRAPH), volume 41,

  8. [16]

    Jean-Baptiste Mouret and Jeff Clune

    arXiv:2205.01906. Jean-Baptiste Mouret and Jeff Clune. Illuminating search spaces by mapping elites.arXiv preprint arXiv:1504.04909,

  9. [17]

    Clemens Schwarke, Mayank Mittal, Nikita Rudin, David Hoeller, and Marco Hutter

    arXiv:2304.13705. Clemens Schwarke, Mayank Mittal, Nikita Rudin, David Hoeller, and Marco Hutter. RSL-RL: A learning library for robotics research.arXiv preprint arXiv:2509.10771,

  10. [2018]

    Michael Janner, Justin Fu, Marvin Zhang, and Sergey Levine

    arXiv:1805.12114. Michael Janner, Justin Fu, Marvin Zhang, and Sergey Levine. When to trust your model: Model-based policy optimization. InAdvances in Neural Information Processing Systems (NeurIPS), volume 32,

  11. [2019]

    Danijar Hafner, Timothy Lillicrap, Jimmy Ba, and Mohammad Norouzi

    arXiv:1906.08253. Danijar Hafner, Timothy Lillicrap, Jimmy Ba, and Mohammad Norouzi. Dream to control: Learning behaviors by latent imagination. InInternational Conference on Learning Representations (ICLR),

  12. [2020]

    Nicklas Hansen, Xiaolong Wang, and Hao Su

    arXiv:1912.01603. Nicklas Hansen, Xiaolong Wang, and Hao Su. Temporal difference learning for model predictive control. In International Conference on Machine Learning (ICML),

  13. [2021]

    Gabriel B

    arXiv:2107.04034. Gabriel B. Margolis and Pulkit Agrawal. Walk these ways: Tuning robot control for generalization with multiplicity of behavior. InConference on Robot Learning (CoRL),

  14. [2022]

    Nikita Rudin, David Hoeller, Philipp Reist, and Marco Hutter

    arXiv:2212.03238. Nikita Rudin, David Hoeller, Philipp Reist, and Marco Hutter. Learning to walk in minutes using massively parallel deep reinforcement learning. InConference on Robot Learning (CoRL), pages 91–100,

  15. [2023]

    Joonho Lee, Jemin Hwangbo, Lorenz Wellhausen, Vladlen Koltun, and Marco Hutter

    arXiv:2301.04195. Joonho Lee, Jemin Hwangbo, Lorenz Wellhausen, Vladlen Koltun, and Marco Hutter. Learning quadrupedal locomotion over challenging terrain.Science Robotics, 5(47):eabc5986,

  16. [2024]

    Jeonghwan Kim, Yunhai Han, Harish Ravichandar, and Sehoon Ha

    arXiv:2401.17583. Jeonghwan Kim, Yunhai Han, Harish Ravichandar, and Sehoon Ha. Learning Koopman dynamics for safe legged locomotion with reinforcement learning-based controller.arXiv preprint arXiv:2409.14736,

  17. [2026]

    arXiv:2607.08951

    doi: 10.1126/science.aeb9506. arXiv:2607.08951. Richard S. Sutton, Doina Precup, and Satinder Singh. Between MDPs and semi-MDPs: A framework for temporal abstraction in reinforcement learning.Artificial Intelligence, 112(1-2):181–211,