Pith. sign in

REVIEW 22 cited by

Imitating Human Behaviour with Diffusion Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2301.10677 v2 pith:BRBRMJRL submitted 2023-01-25 cs.AI cs.LGstat.ML

classification cs.AIcs.LGstat.ML
keywords modelsbehaviourdiffusionhumanimitatingactionchoicesenvironments
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Diffusion models have emerged as powerful generative models in the text-to-image domain. This paper studies their application as observation-to-action models for imitating human behaviour in sequential environments. Human behaviour is stochastic and multimodal, with structured correlations between action dimensions. Meanwhile, standard modelling choices in behaviour cloning are limited in their expressiveness and may introduce bias into the cloned policy. We begin by pointing out the limitations of these choices. We then propose that diffusion models are an excellent fit for imitating human behaviour, since they learn an expressive distribution over the joint action space. We introduce several innovations to make diffusion models suitable for sequential environments; designing suitable architectures, investigating the role of guidance, and developing reliable sampling strategies. Experimentally, diffusion models closely match human demonstrations in a simulated robotic control task and a modern 3D gaming environment.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 22 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Action Map Policy: Learning 3D Closed-loop Manipulation via Pixel Classification

    cs.RO 2026-07 conditional novelty 6.0 of 10

    Projecting 3D gripper keypoints onto camera pixels and classifying those pixels yields millimeter-precise, multi-modal closed-loop manipulation faster than diffusion policies.

  2. InDex: Empowering VLA Models with Intent-Conditioned Arm-Hand Coordination for Dexterous Manipulation

    cs.RO 2026-06 unverdicted novelty 6.0 of 10

    InDex adapts VLA models to high-DoF dexterous manipulation via intent-conditioned fine-tuning and a decoupled diffusion head, outperforming monolithic baselines in simulation tasks with minimal data.

  3. FM-IRL: Flow-Matching for Reward Modeling and Policy Regularization in Reinforcement Learning

    cs.LG 2025-10 conditional novelty 6.0 of 10

    An online imitation-learning method uses a flow-matching teacher's class-conditional loss as a reward and a regularizer to train a simple MLP policy, beating cloning and adversarial-imitation baselines on five of six tasks.

  4. CodeDiffuser: Attention-Enhanced Diffusion Policy via VLM-Generated Code for Instruction Ambiguity

    cs.RO 2025-06 conditional novelty 6.0 of 10

    CodeDiffuser uses vision-language-model-generated code to build 3D attention maps that condition a diffusion policy, improving success on ambiguous language manipulation tasks compared with end-to-end baselines.

  5. CDP: Towards Robust Autoregressive Visuomotor Policy Learning via Causal Diffusion

    cs.CV 2025-06 conditional novelty 6.0 of 10

    Causal Diffusion Policy adds historical action conditioning and attention cache sharing to diffusion-based robot policies, improving success rates on most tested manipulation tasks under degraded observations.

  6. ClutterDexGrasp: A Sim-to-Real System for General Dexterous Grasping in Cluttered Scenes

    cs.RO 2025-06 conditional novelty 6.0 of 10

    A simulation-trained teacher-student policy achieves zero-shot sim-to-real closed-loop target-oriented dexterous grasping in cluttered scenes, with 83.9 percent real-world success.

  7. Learning Generalizable Robot Policy with Human Demonstration Video as a Prompt

    cs.RO 2025-05 conditional novelty 6.0 of 10

    A two-stage pipeline trains a robot policy that accepts a human demonstration video as a prompt and generalizes beyond its robot training tasks, with success rates of up to 79 percent on known task variations and unde...

  8. Robust Planning for Autonomous Driving via Mixed Adversarial Diffusion Predictions

    cs.RO 2025-05 conditional novelty 6.0 of 10

    The authors mix normal and adversarially biased diffusion predictions under expected cost, and report a closed-loop score of 86.6 versus 83.5 for the best baseline in three adversarial driving scenarios.

  9. Improving Trajectory Stitching with Flow Models

    cs.RO 2025-05 conditional novelty 6.0 of 10

    Flow Planner combines a UNet with local inpainting conditioning, action-noise data augmentation, and train/inference trajectory splitting to enable flow models to stitch novel trajectories for robotic manipulation.

  10. Analytic Energy-Guided Policy Optimization for Offline Reinforcement Learning

    cs.LG 2025-05 conditional novelty 6.0 of 10

    AEPO approximates the log-expectation in energy-guided diffusion policy sampling using Taylor expansion and the Gaussian moment-generating function, and reports state-of-the-art average scores on D4RL offline RL benchmarks.

  11. SoccerDiffusion: Toward Learning End-to-End Humanoid Robot Soccer from Gameplay Recordings

    cs.RO 2025-04 conditional novelty 6.0 of 10

    An end-to-end transformer diffusion policy, distilled to one inference step, reproduces low-level humanoid soccer behaviors from real RoboCup game recordings but lacks high-level tactical behavior.

  12. Latent Diffusion Planning for Imitation Learning

    cs.RO 2025-04 conditional novelty 6.0 of 10

    Latent Diffusion Planning separates latent-state forecasting from action inference, letting imitation learning use action-free and suboptimal data, and outperforms Diffusion Policy in low-demonstration manipulation tasks.

  13. CordViP: Correspondence-based Visuomotor Policy for Dexterous Manipulation in Real-World

    cs.RO 2025-02 conditional novelty 6.0 of 10

    CordViP achieves strong real-world dexterous manipulation by feeding a diffusion policy with pose-tracked 3D object models and hand point clouds, pretrained on contact maps and arm-hand coordination.

  14. VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks

    cs.RO 2024-12 conditional novelty 6.0 of 10

    VLABench introduces 100 manipulation task categories with long-horizon reasoning, and shows that current VLAs and VLMs fail most of them.

  15. Planning-Guided Diffusion Policy Learning for Generalizable Contact-Rich Bimanual Manipulation

    cs.RO 2024-12 conditional novelty 6.0 of 10

    GLIDE trains one point-cloud diffusion policy on planner-generated simulation trajectories and successfully reorients diverse objects, including out-of-distribution shapes, with two robot arms in the real world.

  16. Enhancing Exploration with Diffusion Policies in Hybrid Off-Policy RL: Application to Non-Prehensile Manipulation

    cs.RO 2024-11 conditional novelty 6.0 of 10

    HyDo combines diffusion-model policies with maximum entropy RL in a hybrid discrete/continuous action space, improving success rates on non-prehensile manipulation tasks.

  17. TROFI: Trajectory-Ranked Offline Inverse Reinforcement Learning

    cs.LG 2025-06 conditional novelty 5.0 of 10

    TROFI learns a reward model from ranked trajectories, labels an offline dataset with it, and trains a TD3+BC policy, matching ground-truth-reward performance on many D4RL tasks without a hand-coded reward or expert de...

  18. Bootstrapping Imitation Learning for Long-horizon Manipulation via Hierarchical Data Collection Space

    cs.RO 2025-05 conditional novelty 5.0 of 10

    Breaking long manipulation tasks into atomic subtasks and collecting demonstrations from varied starting poses improves imitation learning success using fewer demonstration frames.

  19. FlowPolicy: Enabling Fast and Robust 3D Flow-based Policy via Consistency Flow Matching for Robot Manipulation

    cs.RO 2024-12 conditional novelty 5.0 of 10

    A consistency flow matching policy conditioned on 3D point clouds generates robot actions in a single inference step, running 7x faster than DP3 with comparable success rates.

  20. Large Model Empowered Embodied AI: A Survey on Decision-Making and Embodied Learning

    cs.RO 2025-08 reject novelty 4.0 of 10

    A review that categorizes large-model-empowered embodied AI into hierarchical and end-to-end decision-making, imitation and reinforcement learning, and world models.

  21. RoboGrasp: A Universal Grasping Policy for Robust Robotic Control

    cs.RO 2025-02 reject novelty 4.0 of 10

    Conditioning a Diffusion Policy on YOLO-detected grasp boxes improved reported success rates in three grasping tasks, but the evaluation gives RoboGrasp a goal prompt the baseline lacks.

  22. Diffusion-Based Imitation Learning for Social Pose Generation

    cs.LG 2025-01 conditional novelty 3.0 of 10

    Diffusion behavior cloning can generate facilitator poses, and conditioning on plotted pose keypoints lowers MPJPE but increases processing time versus raw images.

Pith tools