Pith. sign in

REVIEW 13 cited by

Robot Data Curation with Mutual Information Estimators

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2502.08623 v3 pith:JHUGV5MK submitted 2025-02-12 cs.RO

classification cs.RO
keywords datainformationmutualqualityactiondemonstrationsroboticsacross
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The performance of imitation learning policies often hinges on the datasets with which they are trained. Consequently, investment in data collection for robotics has grown across both industrial and academic labs. However, despite the marked increase in the quantity of demonstrations collected, little work has sought to assess the quality of said data despite mounting evidence of its importance in other areas such as vision and language. In this work, we take a critical step towards addressing the data quality in robotics. Given a dataset of demonstrations, we aim to estimate the relative quality of individual demonstrations in terms of both action diversity and predictability. To do so, we estimate the average contribution of a trajectory towards the mutual information between states and actions in the entire dataset, which captures both the entropy of the marginal action distribution and the state-conditioned action entropy. Though commonly used mutual information estimators require vast amounts of data often beyond the scale available in robotics, we introduce a novel technique based on k-nearest neighbor estimates of mutual information on top of simple VAE embeddings of states and actions. Empirically, we demonstrate that our approach is able to partition demonstration datasets by quality according to human expert scores across a diverse set of benchmarks spanning simulation and real world environments. Moreover, training policies based on data filtered by our method leads to a 5-10% improvement in RoboMimic and better performance on real ALOHA and Franka setups.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 13 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. World Action Verifier: Self-Improving World Models via Forward-Inverse Asymmetry

    cs.LG 2026-04 accept novelty 7.0 of 10

    WAV self-improves action-conditioned world models by cycle-consistent verification of state plausibility and sparse action reachability, doubling sample efficiency and lifting policy reward by over 22% on nine tasks.

  2. AutoSpeed: Annotation-Free Stage-Adaptive Motion Speed Learning for Robot Manipulation

    cs.RO 2026-07 unverdicted novelty 6.5 of 10

    AutoSpeed optimizes visuomotor policies over candidate trajectories at varying speeds using a composite cost of prediction error versus horizon length, with DCT-based modulation, yielding shorter execution times and h...

  3. Auditing Instruction-Trajectory Mismatches in Multimodal Robot Demonstrations

    cs.RO 2026-08 conditional novelty 6.0 of 10

    MMPF detects and corrects instruction-trajectory mismatches in robot demonstration datasets using local neighborhood voting, global prototype similarity, and entropy-weighted multimodal fusion.

  4. SIEVE: Structure-Aware Data Selection for Imitation Learning with VLA Models

    cs.RO 2026-07 conditional novelty 6.0 of 10

    Selecting 50% of robot demonstrations by maximizing exposure to reusable primitive-transition patterns outperforms full-data training while halving training steps.

  5. Learning Task-Sufficient World Models by Synergizing Agentic Exploration and Structured Modeling

    cs.LG 2026-07 conditional novelty 6.0 of 10

    Closed-loop agentic probing plus minimality/sufficiency masking recovers compact task-sufficient world-model latents that improve sample-efficient policy learning and cross-task generalization.

  6. Perfect Demo Makes Poor Teacher: Learning Robust Alignment from Critical Motion Segments

    cs.RO 2026-06 conditional novelty 6.0 of 10

    Fluent expert demonstrations under-supervise the short alignment phase that decides success, and a compact spatio-temporal dynamic feature (STAIR) recovers most of the deliberate-demonstration gain from fluent data alone.

  7. Data Retrieval with Importance Weights for Few-Shot Imitation Learning

    cs.RO 2025-09 conditional novelty 6.0 of 10

    Importance Weighted Retrieval scores prior robot data by the ratio of Gaussian kernel density estimates of the target and prior distributions, improving few-shot imitation learning.

  8. Shortcut Learning in Generalist Robot Policies: The Role of Dataset Diversity and Fragmentation

    cs.RO 2025-08 conditional novelty 6.0 of 10

    Low within-subdataset diversity and large between-subdataset differences cause shortcut learning in generalist robot policies, and targeted augmentation can mitigate it.

  9. SCIZOR: A Self-Supervised Approach to Data Curation for Large-Scale Imitation Learning

    cs.RO 2025-05 conditional novelty 6.0 of 10

    SCIZOR filters suboptimal and redundant state-action pairs from robot demonstrations without human labels, improving imitation-learning policy success rates by about 15% on average.

  10. Active Real-World Factor-Based Evaluation for Generalist Robot Policies

    cs.LG 2026-07 conditional novelty 5.0 of 10

    An active evaluation framework selects the most informative task configurations for real-robot tests, matching random testing's accuracy in 20-40% fewer trials.

  11. SPARSE Data, Rich Results: Few-Shot Semi-Supervised Learning via Class-Conditioned Image Translation

    cs.CV 2025-08 unverdicted novelty 5.0 of 10

    A GAN framework that translates unlabeled medical images between classes and fuses ensemble, time-averaged pseudo-labels outperforms six prior GAN semi-supervised methods on MedMNIST at 5-50 labels per class.

  12. Detecting Mislabeled and Corrupted Data via Pointwise Mutual Information

    cs.LG 2025-08 unverdicted novelty 4.0 of 10

    Samples with low pointwise mutual information between image and label are mostly mislabeled or corrupted, and dropping them before training improves MNIST accuracy by up to 15%.

  13. Retrieve-Augmented Generation for Speeding up Diffusion Policy without Additional Training

    cs.LG 2025-07 conditional novelty 4.0 of 10

    RAGDP accelerates pretrained diffusion policies by initializing denoising from the nearest retrieved expert demonstration action, improving accuracy-versus-speed trade-offs without extra training.

Pith tools