PAIR-VLA adds invariance and sensitivity objectives over paired visual variants during PPO fine-tuning of VLA models, yielding 9-16% average gains on ManiSkill3 under distractors, textures, poses, viewpoints, and lighting shifts.
Invariance co-training for robot visual generalization
3 Pith papers cite this work. Polarity classification is still indexing.
fields
cs.RO 3years
2026 3representative citing papers
A hybrid data collection strategy with a mobile camera arm in dual-arm robots reduces shortcut learning in VLA models and improves spatial generalization to unseen poses and configurations across ACT, Diffusion, and VLA architectures.
HyperSim reports 80% and 95% sim-to-real success on two manipulation policies across 400 real executions by combining synthetic environment synthesis, adversarial trajectories, and co-training.
citing papers explorer
-
What to Ignore, What to React: Visually Robust RL Fine-Tuning of VLA Models
PAIR-VLA adds invariance and sensitivity objectives over paired visual variants during PPO fine-tuning of VLA models, yielding 9-16% average gains on ManiSkill3 under distractors, textures, poses, viewpoints, and lighting shifts.
-
The Moving Eye: Enhancing VLA Spatial Generalization via Hybrid Dynamic Data Collection
A hybrid data collection strategy with a mobile camera arm in dual-arm robots reduces shortcut learning in VLA models and improves spatial generalization to unseen poses and configurations across ACT, Diffusion, and VLA architectures.
-
HyperSim: A Holistic Sim-To-Real Framework For Robust Robotic Manipulation
HyperSim reports 80% and 95% sim-to-real success on two manipulation policies across 400 real executions by combining synthetic environment synthesis, adversarial trajectories, and co-training.