Score and Distribution Matching Policy: Advanced Accelerated Visuomotor Policies via Matched Distillation

Bofang Jia; Can Cui; Donglin Wang; Mingyang Sun; Pengfang Qian; Pengxiang Ding; Siteng Huang; Zhaoxin Fan

arxiv: 2412.09265 · v4 · pith:GQANRIBUnew · submitted 2024-12-12 · 💻 cs.RO · cs.LG· stat.ML

Score and Distribution Matching Policy: Advanced Accelerated Visuomotor Policies via Matched Distillation

Bofang Jia , Pengxiang Ding , Can Cui , Mingyang Sun , Pengfang Qian , Siteng Huang , Zhaoxin Fan , Donglin Wang This is my paper

classification 💻 cs.RO cs.LGstat.ML

keywords policymatchingactiondistributioninferencepoliciesscoreadvanced

0 comments

read the original abstract

Visual-motor policy learning has advanced with architectures like diffusion-based policies, known for modeling complex robotic trajectories. However, their prolonged inference times hinder high-frequency control tasks requiring real-time feedback. While consistency distillation (CD) accelerates inference, it introduces errors that compromise action quality. To address these limitations, we propose the Score and Distribution Matching Policy (SDM Policy), which transforms diffusion-based policies into single-step generators through a two-stage optimization process: score matching ensures alignment with true action distributions, and distribution matching minimizes KL divergence for consistency. A dual-teacher mechanism integrates a frozen teacher for stability and an unfrozen teacher for adversarial training, enhancing robustness and alignment with target distributions. Evaluated on a 57-task simulation benchmark, SDM Policy achieves a 6x inference speedup while having state-of-the-art action quality, providing an efficient and reliable framework for high-frequency robotic tasks.

This paper has not been read by Pith yet.

discussion (0)

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

FASTER: Rethinking Real-Time Flow VLAs
cs.RO 2026-03 conditional novelty 6.0

FASTER uses a horizon-aware flow sampling schedule to compress immediate-action denoising to one step, slashing effective reaction latency in real-robot VLA deployments.
FASTER: Rethinking Real-Time Flow VLAs
cs.RO 2026-03 unverdicted novelty 6.0

FASTER adds a Horizon-Aware Schedule to flow VLAs that compresses immediate-action denoising to one step while keeping long-horizon trajectory quality, lowering real-robot reaction latency.
FocalPolicy: Frequency-Optimized Chunking and Locally Anchored Flow Matching for Coherent Visuomotor Policy
cs.RO 2026-05 unverdicted novelty 5.0

FocalPolicy introduces frequency-optimized chunking and locally anchored flow matching with a foresight composite objective to improve inter-chunk coherence in visuomotor policies for manipulation tasks.
FocalPolicy: Frequency-Optimized Chunking and Locally Anchored Flow Matching for Coherent Visuomotor Policy
cs.RO 2026-05 unverdicted novelty 5.0

FocalPolicy introduces frequency-optimized chunking and locally anchored flow matching with a foresight composite objective to reduce inter-chunk discontinuities in visuomotor policies.
Sparse ActionGen: Accelerating Diffusion Policy with Real-time Pruning
cs.RO 2026-01 unverdicted novelty 5.0

Sparse ActionGen accelerates diffusion policies up to 4x for robot control via rollout-adaptive pruning and zig-zag activation reuse without performance loss.