REALM jointly learns model parameters and per-annotator expertise scalars during fine-tuning by modeling observed labels as mixtures of model predictions and uniform noise, improving accuracy under simulated annotation noise.
Trail: Near-optimal imitation learning with suboptimal data
2 Pith papers cite this work. Polarity classification is still indexing.
2
Pith papers citing it
citation-role summary
background 1
citation-polarity summary
roles
background 1polarities
unclear 1representative citing papers
MimicGen creates over 50K robot demonstrations from roughly 200 human ones, allowing imitation learning to achieve strong performance on complex long-horizon tasks like assembly and coffee preparation.
citing papers explorer
-
REALM: Reliable Expertise-Aware Language Model Fine-Tuning from Noisy Annotations
REALM jointly learns model parameters and per-annotator expertise scalars during fine-tuning by modeling observed labels as mixtures of model predictions and uniform noise, improving accuracy under simulated annotation noise.
-
MimicGen: A Data Generation System for Scalable Robot Learning using Human Demonstrations
MimicGen creates over 50K robot demonstrations from roughly 200 human ones, allowing imitation learning to achieve strong performance on complex long-horizon tasks like assembly and coffee preparation.