Pith. sign in

REVIEW 15 cited by

BAKU: An Efficient Transformer for Multi-Task Policy Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.07539 v2 pith:RHSOEBKK submitted 2024-06-11 cs.RO

classification cs.RO
keywords bakulearningtasksactiondatademonstrationsefficientimprovement
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Training generalist agents capable of solving diverse tasks is challenging, often requiring large datasets of expert demonstrations. This is particularly problematic in robotics, where each data point requires physical execution of actions in the real world. Thus, there is a pressing need for architectures that can effectively leverage the available training data. In this work, we present BAKU, a simple transformer architecture that enables efficient learning of multi-task robot policies. BAKU builds upon recent advancements in offline imitation learning and meticulously combines observation trunks, action chunking, multi-sensory observations, and action heads to substantially improve upon prior work. Our experiments on 129 simulated tasks across LIBERO, Meta-World suite, and the Deepmind Control suite exhibit an overall 18% absolute improvement over RT-1 and MT-ACT, with a 36% improvement on the harder LIBERO benchmark. On 30 real-world manipulation tasks, given an average of just 17 demonstrations per task, BAKU achieves a 91% success rate. Videos of the robot are best viewed at https://baku-robot.github.io/.

Discussion (0). Sign in to comment.

Forward citations

Cited by 15 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Action-Prior Denoising for Smooth Real-Time Chunking

    cs.RO 2026-05 unverdicted novelty 7.0 of 10

    Soft RTC uses partially denoised states for overlap tokens and token-wise blending to reduce action delta and jerk by ~9% versus hard RTC while matching solve rates on Kinetix levels.

  2. Patch Policy: Efficient Embodied Control via Dense Visual Representations

    cs.RO 2026-07 conditional novelty 6.0 of 10

    Patch Policy shows that frozen dense ViT patch tokens, consumed through a block-causal attention mask, let lightweight robot policies beat pooled-feature policies and even a fine-tuned 7B vision-language-action model.

  3. Hierarchical Policy Learning via Spectral Decomposition

    cs.RO 2026-06 unverdicted novelty 6.0 of 10

    Causal Spectral Policy decomposes actions spectrally into coarse motion from obs/language and conditional fine corrections, outperforming baselines on precision manipulation tasks.

  4. TempoVLA: Learning Speed-Controllable Vision-Language-Action Policies

    cs.RO 2026-06 unverdicted novelty 6.0 of 10

    TempoVLA learns a single VLA policy with controllable execution speed via variable-speed trajectory augmentation and explicit speed conditioning.

  5. Continuous Reasoning for Vision-Language-Action

    cs.RO 2026-05 unverdicted novelty 6.0 of 10

    Continuous Reasoning for VLA introduces a shared Gaussian latent for continuous thoughts, trained with self-verification to improve action prediction on LIBERO-PRO and real robots.

  6. DexHoldem: Playing Texas Hold'em with Dexterous Embodied System

    cs.RO 2026-05 unverdicted novelty 6.0 of 10

    DexHoldem is a new benchmark providing 1,470 teleoperated demonstrations across 14 manipulation primitives, plus standardized tests for dexterous policy execution and agentic perception in a physical Texas Hold'em setting.

  7. Nautilus: From One Prompt to Plug-and-Play Robot Learning

    cs.RO 2026-05 conditional novelty 6.0 of 10

    A typed-contract harness with containerized 'chambers' and robotics-specific agent skills lets a coding LLM turn a single natural-language prompt into working reproduction, evaluation, and deployment workflows for rob...

  8. Spotlighting Task-Relevant Features: Object-Centric Representations for Better Generalization in Robotic Manipulation

    cs.RO 2026-01 conditional novelty 6.0 of 10

    Slot-based object-centric visual representations, especially with robot-video pretraining, improve out-of-distribution generalization of robotic manipulation policies compared to global and dense pre-trained features.

  9. AFFORD2ACT: Affordance-Guided Automatic Keypoint Selection for Generalizable and Lightweight Robotic Manipulation

    cs.RO 2025-10 unverdicted novelty 6.0 of 10

    AFFORD2ACT distills a minimal set of affordance-guided 2D keypoints from text and a single image to train a 38-dimensional gated transformer policy that achieves 82% success on unseen objects and scenes.

  10. FLOWER: Democratizing Generalist Robot Policies with Efficient Vision-Language-Action Flow Policies

    cs.RO 2025-09 conditional novelty 6.0 of 10

    A compact 950-million-parameter robot policy trained in about 200 GPU-hours matches or beats multi-billion-parameter baselines on most manipulation benchmarks, including a new best score on CALVIN ABC.

  11. DINO-WM: World Models on Pre-trained Visual Features enable Zero-shot Planning

    cs.RO 2024-11 unverdicted novelty 6.0 of 10

    DINO-WM builds world models on pre-trained DINOv2 features to enable zero-shot planning from offline data without rewards or demonstrations.

  12. TempoVLA: Learning Speed-Controllable Vision-Language-Action Policies

    cs.RO 2026-06 conditional novelty 5.0 of 10

    A single vision-language-action policy can execute robot manipulation at commanded speeds from 0.5x to 2x by training on merged/split demonstration actions conditioned on a speed scalar.

  13. Nautilus: From One Prompt to Plug-and-Play Robot Learning

    cs.RO 2026-05 unverdicted novelty 5.0 of 10

    NAUTILUS is a prompt-driven harness that automates plug-and-play adapters, typed contracts, and validation for policies, benchmarks, and robots in learning research.

  14. ReFineVLA: Multimodal Reasoning-Aware Generalist Robotic Policies via Teacher-Guided Fine-Tuning

    cs.RO 2026-04 unverdicted novelty 5.0 of 10

    ReFineVLA adds teacher-generated reasoning steps to VLA training and reports state-of-the-art success rates on SimplerEnv WidowX and Google Robot benchmarks.

  15. General Covariant Action Modeling: Constructing Generalized Manifolds via Spatio-Temporal Decoupling

    cs.CV 2026-05 unverdicted novelty 4.0 of 10

    GAM framework uses arc-length parameterization for temporal invariance and schema-affine factorization for geometric invariance to build a covariant action manifold integrated into VLA models for improved generalizati...

Pith tools