Pith. sign in

REVIEW 3 cited by

SkillMimic-V2: Learning Robust and Generalizable Interaction Skills from Sparse and Noisy Demonstrations

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2505.02094 v1 pith:KUUNMAL6 submitted 2025-05-04 cs.LG cs.CV

classification cs.LGcs.CV
keywords demonstrationdemonstrationsinteractionskilldatalearningnoisyskills
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We address a fundamental challenge in Reinforcement Learning from Interaction Demonstration (RLID): demonstration noise and coverage limitations. While existing data collection approaches provide valuable interaction demonstrations, they often yield sparse, disconnected, and noisy trajectories that fail to capture the full spectrum of possible skill variations and transitions. Our key insight is that despite noisy and sparse demonstrations, there exist infinite physically feasible trajectories that naturally bridge between demonstrated skills or emerge from their neighboring states, forming a continuous space of possible skill variations and transitions. Building upon this insight, we present two data augmentation techniques: a Stitched Trajectory Graph (STG) that discovers potential transitions between demonstration skills, and a State Transition Field (STF) that establishes unique connections for arbitrary states within the demonstration neighborhood. To enable effective RLID with augmented data, we develop an Adaptive Trajectory Sampling (ATS) strategy for dynamic curriculum generation and a historical encoding mechanism for memory-dependent skill learning. Our approach enables robust skill acquisition that significantly generalizes beyond the reference demonstrations. Extensive experiments across diverse interaction tasks demonstrate substantial improvements over state-of-the-art methods in terms of convergence stability, generalization capability, and recovery robustness.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Go to Zero: Towards Zero-shot Motion Generation with Million-scale Data

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A 7B text-to-motion model trained on the new 2M-clip MotionMillion dataset is reported to generalize zero-shot to complex, out-of-domain prompts.

  2. CoDA: Coordinated Diffusion Noise Optimization for Whole-Body Manipulation of Articulated Objects

    cs.GR 2025-05 conditional novelty 6.0 of 10

    CoDA generates coordinated whole-body articulated-object manipulation by optimizing the noise of three decoupled diffusion models, guided by BPS-based end-effector and object trajectories.

  3. Tired Actor: Fatigue-Informed Character Control

    cs.RO 2026-08 conditional novelty 5.0 of 10

    Injecting a muscle-fatigue model into a general physics-based character controller preserves motion imitation accuracy while producing tired, more human-like behaviors such as shorter steps, corner cutting, and fall c...

Pith tools