Pith. sign in

REVIEW 8 cited by

Zero-Shot Whole-Body Humanoid Control via Behavioral Foundation Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2504.11054 v1 pith:YJ7JREO5 submitted 2025-04-15 cs.LG

Zero-Shot Whole-Body Humanoid Control via Behavioral Foundation Models

classification cs.LG
keywords policiestasksunsuperviseddatasetsdownstreamhumanoidtheyunlabeled
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

Unsupervised reinforcement learning (RL) aims at pre-training agents that can solve a wide range of downstream tasks in complex environments. Despite recent advancements, existing approaches suffer from several limitations: they may require running an RL process on each downstream task to achieve a satisfactory performance, they may need access to datasets with good coverage or well-curated task-specific samples, or they may pre-train policies with unsupervised losses that are poorly correlated with the downstream tasks of interest. In this paper, we introduce a novel algorithm regularizing unsupervised RL towards imitating trajectories from unlabeled behavior datasets. The key technical novelty of our method, called Forward-Backward Representations with Conditional-Policy Regularization, is to train forward-backward representations to embed the unlabeled trajectories to the same latent space used to represent states, rewards, and policies, and use a latent-conditional discriminator to encourage policies to ``cover'' the states in the unlabeled behavior dataset. As a result, we can learn policies that are well aligned with the behaviors in the dataset, while retaining zero-shot generalization capabilities for reward-based and imitation tasks. We demonstrate the effectiveness of this new approach in a challenging humanoid control problem: leveraging observation-only motion capture datasets, we train Meta Motivo, the first humanoid behavioral foundation model that can be prompted to solve a variety of whole-body tasks, including motion tracking, goal reaching, and reward optimization. The resulting model is capable of expressing human-like behaviors and it achieves competitive performance with task-specific methods while outperforming state-of-the-art unsupervised RL and model-based baselines.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Scaling Behavior Foundation Model for Humanoid Robots

    cs.RO 2026-07 conditional novelty 6.0

    A scaling recipe for humanoid behavior foundation models—global-frame motion tracking, on-policy data quantity plus reference diversity, and a transformer with hyperspherical latents—cuts global tracking error by roug...

  2. Exploration and Online Transfer with Behavioral Foundation Models

    cs.AI 2026-06 unverdicted novelty 6.0

    Proposes framing online zero-shot RL transfer as a bandit problem solved by BFMs, deriving eigenvalue minimization of an uncertainty matrix for exploration under linear reward approximation, validated on a simple environment.

  3. Exploration and Online Transfer with Behavioral Foundation Models

    cs.AI 2026-06 unverdicted novelty 6.0

    Frames online zero-shot transfer with BFMs as a bandit problem and derives an eigenvalue-minimization exploration strategy under linear reward approximation.

  4. Reward-free Pretraining for Reinforcement Learning via Occupancy Coverage Maximization

    cs.LG 2026-06 unverdicted novelty 6.0

    ROVER pretrains transferable exploration policies by maximizing occupancy coverage with a learned resolvent world model and virtual sink state, outperforming baselines on sparse navigation tasks.

  5. MotionPyramid: Hierarchical Motion Representation and Residual Interfaces

    cs.CV 2026-06 unverdicted novelty 6.0

    MotionPyramid learns a stack of latent decoders from motion tracking data to create multi-resolution action interfaces for RL policies in humanoid control, with residual interfaces allowing coarse programs and fine co...

  6. Goal-Conditioned Agents that Learn Everything All at Once

    cs.LG 2026-05 unverdicted novelty 6.0

    LEO enables efficient all-goals learning in goal-conditioned RL by jointly predicting for all goals in one network pass, yielding >250x speedup over relabelling and better performance on Craftax.

  7. OMG: Omni-Modal Motion Generation for Generalist Humanoid Control

    cs.RO 2026-06 unverdicted novelty 5.0

    OMG is a diffusion model for omni-modal whole-body humanoid motion generation that uses language, audio, and reference motions after large-scale data curation to achieve state-of-the-art performance and adaptation.

  8. Zero-shot adaptation to order book dynamics

    cs.CE 2026-05 unverdicted novelty 5.0

    Introduces a successor-measure adaptation that separates market dynamics from trading objectives inside the Avellaneda-Stoikov HJB framework to enable zero-shot quote adjustment.