Pith. sign in

REVIEW 15 cited by

Sparse Diffusion Policy: A Sparse, Reusable, and Flexible Policy for Robot Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.01531 v2 pith:TVMUNXE3 submitted 2024-07-01 cs.RO cs.LG

classification cs.ROcs.LG
keywords policytaskslearningsparsediffusionefficientexpertsactive
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

The increasing complexity of tasks in robotics demands efficient strategies for multitask and continual learning. Traditional models typically rely on a universal policy for all tasks, facing challenges such as high computational costs and catastrophic forgetting when learning new tasks. To address these issues, we introduce a sparse, reusable, and flexible policy, Sparse Diffusion Policy (SDP). By adopting Mixture of Experts (MoE) within a transformer-based diffusion policy, SDP selectively activates experts and skills, enabling efficient and task-specific learning without retraining the entire model. SDP not only reduces the burden of active parameters but also facilitates the seamless integration and reuse of experts across various tasks. Extensive experiments on diverse tasks in both simulations and real world show that SDP 1) excels in multitask scenarios with negligible increases in active parameters, 2) prevents forgetting in continual learning of new tasks, and 3) enables efficient task transfer, offering a promising solution for advanced robotic applications. Demos and codes can be found in https://forrest-110.github.io/sparse_diffusion_policy/.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 15 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. High-Fidelity One-Step Generative Visuomotor Policy via Recursive Correction, Frequency Consistency, and Contrastive Flow Matching

    cs.RO 2026-07 conditional novelty 6.0 of 10

    One-step flow-matching visuomotor policy with recursive correction, dual-timestep spectral consistency, and contrastive mode separation matches or exceeds 10-step baselines at 1 NFE.

  2. Learning Semantic Atomic Skills for Multi-Task Robotic Manipulation

    cs.RO 2025-12 conditional novelty 6.0 of 10

    An imitation-learning system that segments demonstrations into VLM-labeled atomic skills, aligns them with contrastive learning, and uses keypose prediction to chain skills, outperforming prior baselines in multi-task...

  3. RCM-ACT: Imitation Learning with Dynamic RCM Calibration for Autonomous Intraocular Foreign Body Removal

    cs.RO 2025-08 conditional novelty 6.0 of 10

    An imitation-learning robot with dynamic coordinate calibration can grasp and place a 1.22 mm ring in an eye model with 0.686 mm average error, but full-task success was only 5/10 in the reported table.

  4. ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge

    cs.RO 2025-05 conditional novelty 6.0 of 10

    ChatVLA-2 uses dynamic mixture-of-experts and a two-stage training recipe to let a vision-language-action model retain pretrained reasoning while following robot instructions.

  5. Latent Theory of Mind: A Decentralized Diffusion Architecture for Cooperative Manipulation

    cs.RO 2025-05 conditional novelty 6.0 of 10

    A decentralized diffusion policy for two robot arms that aligns a learned consensus embedding across agents and uses theory-of-mind prediction to keep that embedding informative.

  6. Efficient Diffusion Transformer Policies with Mixture of Expert Denoisers for Multitask Learning

    cs.LG 2024-12 conditional novelty 6.0 of 10

    MoDE, a mixture-of-experts diffusion transformer with noise-conditioned routing, reports state-of-the-art results on CALVIN and LIBERO with lower inference FLOPs than dense baselines.

  7. CARP: Visuomotor Policy Learning via Coarse-to-Fine Autoregressive Prediction

    cs.RO 2024-12 conditional novelty 6.0 of 10

    A coarse-to-fine autoregressive policy with multi-scale action tokenization matches or beats diffusion policies on robot manipulation benchmarks at roughly 10x lower inference cost.

  8. Diffusion-VLA: Generalizable and Interpretable Robot Foundation Model via Self-Generated Reasoning

    cs.RO 2024-12 conditional novelty 6.0 of 10

    A robot policy generates its own language reasoning before acting and injects it into a diffusion action decoder, outperforming several VLA baselines on real-robot manipulation.

  9. SkillMemo: Expert-guided Skill Memory Framework for Compositional Embodied Manipulation

    cs.RO 2026-08 conditional novelty 5.0 of 10

    SkillMemo couples MoE-based skill discovery with episodic memory retrieval and reports consistent success-rate gains on diffusion and VLA policies for simulated and real manipulation tasks.

  10. Few-Shot Vision-Language Action-Incremental Policy Learning

    cs.RO 2025-04 conditional novelty 5.0 of 10

    TOPIC adds task-specific prompts and task-similarity-based weight interpolation to transformer policies, improving few-shot action-incremental learning in simulation and on a real robot.

  11. Continual Adaptation for Autonomous Driving with the Mixture of Progressive Experts Network

    cs.RO 2025-02 conditional novelty 5.0 of 10

    A growing mixture-of-experts network trained on reinforcement-learning-generated driving data improves continual adaptation in simulated urban driving over standard continual-learning baselines.

  12. A Novel Scene Coupling Semantic Mask Network for Remote Sensing Image Segmentation

    eess.IV 2025-01 conditional novelty 5.0 of 10

    SCSM, a scene coupling and semantic mask attention decoder, reports higher accuracy than prior methods on four remote sensing segmentation benchmarks with lower computational cost.

  13. Grounding Actions in Camera Space: Observation-Centric Vision-Language-Action Policy

    cs.RO 2025-08 unverdicted novelty 4.0 of 10

    OC-VLA re-labels robot action targets from the robot base frame to the camera frame using the camera's extrinsic calibration, improving cross-view generalization of VLA policies.

  14. Credit Risk Identification in Supply Chains Using Generative Adversarial Networks

    cs.LG 2025-01 reject novelty 4.0 of 10

    A GAN-based model is reported to beat baseline classifiers for supply chain credit risk, but the evaluation uses synthetic test data and no artifacts are provided.

  15. First-place Solution for Streetscape Shop Sign Recognition Competition

    cs.CV 2025-01 reject novelty 2.0 of 10

    A team reports winning a street-view shop sign recognition competition with a multi-stage OCR pipeline built from known components, but provides no code, data, or rigorous ablations.

Pith tools