Pith. sign in

REVIEW 17 cited by

Robot Data Curation with Mutual Information Estimators

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2502.08623 v3 pith:JHUGV5MK submitted 2025-02-12 cs.RO

Robot Data Curation with Mutual Information Estimators

classification cs.RO
keywords datainformationmutualqualityactiondemonstrationsroboticsacross
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

The performance of imitation learning policies often hinges on the datasets with which they are trained. Consequently, investment in data collection for robotics has grown across both industrial and academic labs. However, despite the marked increase in the quantity of demonstrations collected, little work has sought to assess the quality of said data despite mounting evidence of its importance in other areas such as vision and language. In this work, we take a critical step towards addressing the data quality in robotics. Given a dataset of demonstrations, we aim to estimate the relative quality of individual demonstrations in terms of both action diversity and predictability. To do so, we estimate the average contribution of a trajectory towards the mutual information between states and actions in the entire dataset, which captures both the entropy of the marginal action distribution and the state-conditioned action entropy. Though commonly used mutual information estimators require vast amounts of data often beyond the scale available in robotics, we introduce a novel technique based on k-nearest neighbor estimates of mutual information on top of simple VAE embeddings of states and actions. Empirically, we demonstrate that our approach is able to partition demonstration datasets by quality according to human expert scores across a diverse set of benchmarks spanning simulation and real world environments. Moreover, training policies based on data filtered by our method leads to a 5-10% improvement in RoboMimic and better performance on real ALOHA and Franka setups.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 17 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Ambient Diffusion Policy: Imitation Learning from Suboptimal Data in Robotics

    cs.RO 2026-06 unverdicted novelty 7.0

    Ambient Diffusion Policy enables better imitation learning from suboptimal robot data by leveraging spectral properties to restrict data usage to specific diffusion times.

  2. Good in Bad (GiB): Sifting Through End-user Demonstrations for Learning a Better Policy

    cs.RO 2026-05 unverdicted novelty 7.0

    GiB filters erroneous subtasks from mixed-quality human demonstrations using self-supervised latent features and Mahalanobis distance to train more robust imitation learning policies.

  3. ${\pi}_{0.7}$: a Steerable Generalist Robotic Foundation Model with Emergent Capabilities

    cs.LG 2026-04 unverdicted novelty 7.0

    π₀.₇ is a steerable generalist robotic model that uses rich multimodal prompts including language, subgoal images, and performance metadata to achieve out-of-the-box generalization across tasks and robot bodies.

  4. World Action Verifier: Self-Improving World Models via Forward-Inverse Asymmetry

    cs.LG 2026-04 accept novelty 7.0

    WAV self-improves action-conditioned world models by cycle-consistent verification of state plausibility and sparse action reachability, doubling sample efficiency and lifting policy reward by over 22% on nine tasks.

  5. AutoSpeed: Annotation-Free Stage-Adaptive Motion Speed Learning for Robot Manipulation

    cs.RO 2026-07 conditional novelty 6.5

    AutoSpeed learns annotation-free, stage-adaptive robot motion speeds by optimizing policies toward the minimum-cost DCT-retimed multi-speed demonstration target.

  6. SIEVE: Structure-Aware Data Selection for Imitation Learning with VLA Models

    cs.RO 2026-07 conditional novelty 6.0

    Selecting 50% of robot demonstrations by maximizing exposure to reusable primitive-transition patterns outperforms full-data training while halving training steps.

  7. Learning Task-Sufficient World Models by Synergizing Agentic Exploration and Structured Modeling

    cs.LG 2026-07 conditional novelty 6.0

    Closed-loop agentic probing plus minimality/sufficiency masking recovers compact task-sufficient world-model latents that improve sample-efficient policy learning and cross-task generalization.

  8. AutoSpeed: Annotation-Free Stage-Adaptive Motion Speed Learning for Robot Manipulation

    cs.RO 2026-07 unverdicted novelty 6.0

    AutoSpeed optimizes visuomotor policies over candidate trajectories at varying speeds using a composite cost of prediction error versus horizon length, with DCT-based modulation, yielding shorter execution times and h...

  9. Tri-Info: Generalizable, Interpretable Failure Prediction for VLA Models via Information Theory

    cs.RO 2026-06 unverdicted novelty 6.0

    Tri-Info uses three information theory signals on action diversity, temporal consistency, and state coupling to predict VLA model failures with cross-domain generalization to 83% real-world accuracy.

  10. FrameSkip: Learning from Fewer but More Informative Frames in VLA Training

    cs.RO 2026-05 unverdicted novelty 6.0

    FrameSkip improves VLA policy training success from 66.50% to 76.15% by selecting high-importance frames and retaining only 20% of unique frames across three benchmarks.

  11. An Efficient Metric for Data Quality Measurement in Imitation Learning

    cs.RO 2026-05 unverdicted novelty 6.0

    Power spectral density of trajectories ranks demonstration quality for imitation learning, enabling rollout-free curation that improves fine-tuned policy success.

  12. Good in Bad (GiB): Sifting Through End-user Demonstrations for Learning a Better Policy

    cs.RO 2026-05 unverdicted novelty 6.0

    GiB uses self-supervised latent features and Mahalanobis distance to filter erroneous subtasks from mixed-quality human demonstrations, improving robot policy learning in simulation and real-world tasks.

  13. Learning from the Best: Smoothness-Driven Metrics for Data Quality in Imitation Learning

    cs.RO 2026-04 unverdicted novelty 6.0

    RINSE scores robot demonstration trajectories for smoothness via SAL and TED metrics to curate higher-quality data for behavioral cloning, improving success rates with less data on benchmarks and real robots.

  14. TSD: A Physics-Inspired Trajectory Saliency Detector for Efficient Imitation Learning

    cs.RO 2026-06 unverdicted novelty 5.0

    TSD applies two physics metrics to identify salient trajectory segments for dataset compression and expansion in robotic imitation learning, yielding comparable performance with 25% less data on average.

  15. GeoSem-WAM: Geometry- and Semantic-Aware World Action Models

    cs.RO 2026-06 unverdicted novelty 5.0

    GeoSem-WAM adds geometric and semantic auxiliary prediction tasks to World Action Models during training to improve latent representations and action prediction accuracy while keeping inference efficient by avoiding e...

  16. AttenA+: Rectifying Action Inequality in Robotic Foundation Models

    cs.RO 2026-05 unverdicted novelty 5.0

    AttenA+ applies velocity-driven action attention to reweight training objectives toward kinematically critical low-velocity segments, yielding small benchmark gains on Libero and RoboTwin without added parameters.

  17. AttenA+: Rectifying Action Inequality in Robotic Foundation Models

    cs.RO 2026-05 unverdicted novelty 4.0

    AttenA+ reweights action training objectives in VLA and WAM models via inverse velocity attention to prioritize kinematically critical segments, yielding small benchmark gains.