REVIEW 33 cited by
Reconciling Reality through Simulation: A Real-to-Sim-to-Real Approach for Robust Manipulation
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Imitation learning methods need significant human supervision to learn policies robust to changes in object poses, physical disturbances, and visual distractors. Reinforcement learning, on the other hand, can explore the environment autonomously to learn robust behaviors but may require impractical amounts of unsafe real-world data collection. To learn performant, robust policies without the burden of unsafe real-world data collection or extensive human supervision, we propose RialTo, a system for robustifying real-world imitation learning policies via reinforcement learning in "digital twin" simulation environments constructed on the fly from small amounts of real-world data. To enable this real-to-sim-to-real pipeline, RialTo proposes an easy-to-use interface for quickly scanning and constructing digital twins of real-world environments. We also introduce a novel "inverse distillation" procedure for bringing real-world demonstrations into simulated environments for efficient fine-tuning, with minimal human intervention and engineering required. We evaluate RialTo across a variety of robotic manipulation problems in the real world, such as robustly stacking dishes on a rack, placing books on a shelf, and six other tasks. RialTo increases (over 67%) in policy robustness without requiring extensive human data collection. Project website and videos at https://real-to-sim-to-real.github.io/RialTo/
Forward citations
Cited by 33 Pith papers
-
GAUGE: A Measurement-Grounded Benchmark for Physical Fidelity in Simulation Engines and Video World Models
A new real-world-grounded benchmark shows that physics engines and video world models each fail differently, with video models often fitting the shape of a physical law while recovering wrong parameters.
-
Fail2Progress: Learning from Real-World Robot Failures with Stein Variational Inference
Fail2Progress generates failure-targeted simulation data via Stein variational inference and fine-tunes skill effect models, improving long-horizon manipulation success rates and generalizing to unseen object counts a...
-
Prompting with the Future: Open-World Model Predictive Control with Interactive Digital Twins
A VLM-driven model predictive controller that evaluates simulated future outcomes rendered from a physics-based digital twin.
-
V-Simba: Unleashing the Architectural Potential of RL in Visual Continuous Control
V-Simba, a visual RL architecture combining layer normalization, weight decay, and a distributional critic, matches or outperforms complex baselines on 29 continuous control tasks while using less compute.
-
SimFoundry: Modular and Automated Scene Generation for Policy Learning and Evaluation
SimFoundry automates zero-shot real-to-sim scene generation from video, producing digital twins and cousins that enable policy training with 0.911 mean Pearson correlation to real-world results and 17-40% success gain...
-
Preference-Calibrated Human-in-the-Loop Reinforcement Learning for Robotic Manipulation
PACT uses demo-trained progress localization plus intervention preference pairs to correct inflated Bellman targets and align the actor, raising average real-robot success by 24.5% over HIL-SERL.
-
Beyond Imitation: Reinforcement Learning-Based Sim-Real Co-Training for VLA Models
Adding a real-world supervised loss to simulation reinforcement learning improves real-robot success and data efficiency for VLA co-training.
-
TabletopGen: Tabletop Scene Generation and Interactive Simulation for Robotic Manipulation
A training-free pipeline generates instance-level, physically interactive 3D tabletop scenes from text or one image, with a differentiable rotation optimizer and top-view spatial alignment for collision-free layouts.
-
Learning in ImaginationLand: Omnidirectional Policies through 3D Generative Models (OP-Gen)
A robot policy trained on one real demonstration plus AI-generated 3D views succeeds from novel initial poses, including opposite-side starts, across six real manipulation tasks.
-
DEXOP: A Device for Robotic Transfer of Dexterous Human Manipulation
DEXOP implements perioperation with a passive exoskeleton linked to a sensorized robot hand, and DEXOP-collected demonstrations train robot policies more efficiently per unit time than teleoperation.
-
LodeStar: Long-horizon Dexterity via Synthetic Data Augmentation from Human Demonstrations
LodeStar combines automatic skill segmentation with simulation-based reinforcement learning augmentation and a learned routing transformer to let a robotic hand complete long-horizon dexterous tasks from a few human demos.
-
ControlVLA: Few-shot Object-centric Adaptation for Pre-trained Vision-Language-Action Models
ControlVLA adapts a DROID-pretrained diffusion VLA policy to new manipulation tasks with 10 to 20 demos by injecting object-centric features through zero-initialized cross-attention layers, achieving 76.7% success acr...
-
AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making
AntiGrounding lifts candidate robot trajectories into the VLM's visual space via multi-view rendering and structured VQA, and reports 57.5% average success across eight manipulation tasks, beating three intermediate-r...
-
Hearing Hands: Generating Sounds from Physical Interactions in 3D Scenes
A rectified flow model conditioned on 3D hand trajectories and rendered scene video generates realistic hand-scene interaction sounds, with a human study finding near-chance discrimination (47% misclassified).
-
Real2Render2Real: Scaling Robot Data Without Dynamics Simulation or Robot Hardware
A single human demonstration and object scan can be expanded into 1,000 synthetic robot demonstrations that train policies to match teleoperation-trained policies on five tabletop manipulation tasks.
-
HuB: Learning Extreme Humanoid Balance
HuB combines reference motion refinement, balance shaping rewards, and robustness training to enable a G1 humanoid to hold extreme single-leg poses that prior tracking methods fail to maintain.
-
Web2Grasp: Learning Functional Grasps from Web Images of Hand-Object Interactions
A pipeline that reconstructs human hand-object grasps from web images, retargets them to robot hands, filters them with geometry and simulation, and trains an interaction-centric functional grasping model that outperf...
-
MetaScenes: Towards Automated Replica Creation for Real-world 3D Scans
MetaScenes converts 706 ScanNet scenes into simulatable 3D replicas with 15,366 objects and ranked candidate assets, and introduces Scan2Sim for automated asset replacement.
-
PRISM: Projection-based Reward Integration for Scene-Aware Real-to-Sim-to-Real Transfer with Few Demonstrations
Using five demonstrations, PRISM builds a simulator and trains a policy with a vision-language-model reward, reaching 82% success under randomized conditions on six tabletop tasks.
-
MOSAIC: Skill-Centric Manipulation Planning with Physics Simulation
MOSAIC is a multi-directional skill-centric planner that seeds feasible local trajectories with generator skills, links them with connector skills, and uses a statistical oracle and physics simulation to guide the search.
-
Novel Demonstration Generation with Gaussian Splatting Enables Robust One-Shot Manipulation
RoboSplat edits 3D Gaussian scene reconstructions to synthesize diverse robot demonstrations from one expert trajectory, and behavior-cloned policies trained on this data generalize robustly across six disturbance typ...
-
A Real-to-Sim-to-Real Approach to Robotic Manipulation with VLM-Generated Iterative Keypoint Rewards
IKER uses VLM-generated keypoint rewards to train manipulation policies in simulation that transfer to a real robot, enabling multi-step tasks and replanning.
-
Rapidly Adapting Policies to the Real World via Simulation-Guided Fine-Tuning
SGFT uses a simulation-trained value function to guide real-world exploration via potential-based reward shaping and short-horizon objectives, substantially improving fine-tuning sample efficiency.
-
VR-Robo: A Real-to-Sim-to-Real Framework for Visual Robot Navigation and Locomotion
A framework that reconstructs real scenes as interactive 3D Gaussian simulations and trains RGB-only navigation policies for legged robots that transfer to the real world without retraining.
-
Dream to Manipulate: Compositional World Models Empowering Robot Imitation Learning with Imagination
DREMA creates an object-centric Gaussian Splatting plus PyBullet world model and generates equivariant-transformed demonstrations, improving imitation learning from a handful of real demonstrations.
-
Shared Voxel-Map-Based Cooperative Indoor UAV Guidance with a Multi-Agent Soft Actor-Critic Controller
A multi-agent SAC controller using a shared voxel-map BEV representation achieves 90.3% simulated corridor success and 100% success across 50 real two-drone indoor trials after A*-based imitation fine-tuning.
-
Active Real-World Factor-Based Evaluation for Generalist Robot Policies
An active evaluation framework selects the most informative task configurations for real-robot tests, matching random testing's accuracy in 20-40% fewer trials.
-
ObjSplat: Geometry-Aware Gaussian Surfels for Active Object Reconstruction
Coupling Gaussian-surfel reconstruction with back-face-aware uncertainty and next-best-path lookahead yields object scans that are more complete and photorealistic while reducing path length about 4–5× versus greedy planners.
-
SimLauncher: Launching Sample-Efficient Real-world Robotic Reinforcement Learning via Simulation Pre-training
Simulation-pretrained policies, with digital-twin demos for critic bootstrapping and action proposals, cut real-world RL training time while reaching near-perfect success on three manipulation tasks.
-
Scan, Materialize, Simulate: A Generalizable Framework for Physically Grounded Robot Planning
SMS combines 3D Gaussian Splatting, SAM 2 segmentation, GPT-4o material inference, and rigid-body simulation to plan physically dynamic robot actions in billiards and quadrotor landing tasks.
-
Distilling Realizable Students from Unrealizable Teachers
A teacher-student distillation framework that queries the teacher only at critical states or resets RL from teacher recovery states improves performance on partially observable robot tasks.
-
Gaussian Object Carver: Object-Compositional Gaussian Splatting with surfaces completion
A Gaussian splatting pipeline reconstructs indoor scenes as separable objects and uses a trained completion model to fill in occluded surfaces zero-shot.
-
RoomCraft: Controllable and Complete 3D Indoor Scene Generation
RoomCraft generates 3D indoor scenes from text, sketches, or images by extracting structured furniture relations with a VLM and resolving placement conflicts with a weighted positioning heuristic.
Discussion (0). Continue with ORCID to comment.