Pith. sign in

REVIEW 8 cited by

Continuously Improving Mobile Manipulation with Autonomous Real-World RL

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2409.20568 v1 pith:FLBQJ2I3 submitted 2024-09-30 cs.RO cs.AIcs.CVcs.LGcs.SYeess.SY

Continuously Improving Mobile Manipulation with Autonomous Real-World RL

classification cs.RO cs.AIcs.CVcs.LGcs.SYeess.SY
keywords manipulationmobileautonomousreal-worldtasksacrossallowsapproach
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

We present a fully autonomous real-world RL framework for mobile manipulation that can learn policies without extensive instrumentation or human supervision. This is enabled by 1) task-relevant autonomy, which guides exploration towards object interactions and prevents stagnation near goal states, 2) efficient policy learning by leveraging basic task knowledge in behavior priors, and 3) formulating generic rewards that combine human-interpretable semantic information with low-level, fine-grained observations. We demonstrate that our approach allows Spot robots to continually improve their performance on a set of four challenging mobile manipulation tasks, obtaining an average success rate of 80% across tasks, a 3-4 improvement over existing approaches. Videos can be found at https://continual-mobile-manip.github.io/

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. SOLE-R1: Video-Language Reasoning as the Sole Reward for On-Robot Reinforcement Learning

    cs.RO 2026-03 conditional novelty 6.5

    A video-language model with per-timestep spatiotemporal CoT and dense progress prediction can serve as the sole reward for zero-shot online robot RL on 24 unseen manipulation tasks.

  2. FT-WBC: Learning Fault-Tolerant Whole-Body Control for Legged Loco-Manipulation

    cs.RO 2026-06 unverdicted novelty 6.0

    FT-WBC introduces a decoupled policy architecture with a Fault Estimator and Posture Adaptation Module that converts unstable arm-driven posture requests into safe base commands under actuator failures in legged manipulators.

  3. Learning Process Rewards via Success Visitation Matching for Efficient RL

    cs.LG 2026-06 unverdicted novelty 6.0

    Success Visitation Matching uses a discriminator to turn sparse outcome rewards into dense process rewards by matching visitations of successful episodes, provably preserving the optimal policy and speeding up robotic...

  4. From Prior to Pro: Efficient Skill Mastery via Distribution Contractive RL Finetuning

    cs.RO 2026-03 accept novelty 6.0

    Residual off-policy RL with selective BC regularization and value-guided sampling contracts a pretrained generative robot policy around successful actions, reaching high success on hard long-horizon tasks from pixels ...

  5. TwinRL: Digital Twin-Driven Reinforcement Learning for Real-World Robotic Manipulation

    cs.RO 2026-02 unverdicted novelty 6.0

    TwinRL expands RL exploration via digital twin reconstruction and twin RL warm-up to guide real-world learning, reaching near-100% success with 20 minutes of on-robot time across four tasks.

  6. UMI-on-Air: Embodiment-Aware Guidance for Embodiment-Agnostic Visuomotor Policies

    cs.RO 2025-10 conditional novelty 6.0

    Embodiment-Aware Diffusion Policy steers a UMI-trained diffusion policy with controller tracking-cost gradients at inference time, improving aerial manipulation success in simulation and real flights.

  7. FT-WBC: Learning Fault-Tolerant Whole-Body Control for Legged Loco-Manipulation

    cs.RO 2026-06 unverdicted novelty 5.0

    FT-WBC is a decoupled-policy framework that uses fault estimation and posture adaptation to synthesize compensatory gaits and preserve arm workspace in legged manipulators under actuator failures.

  8. MoTo: A Zero-shot Plug-in Interaction-aware Navigation for General Mobile Manipulation

    cs.RO 2025-09 conditional novelty 4.0

    MoTo turns existing fixed-base manipulation models into mobile manipulators by using VLM-picked contact keypoints and trajectory optimization to find docking points, with no training of MoTo itself.