EgoEngine transforms egocentric human videos into high-fidelity robot data enabling zero-shot visuomotor dexterous policy learning without real-robot demonstrations.
Foundationstereo: Zero-shot stereo matching
10 Pith papers cite this work, alongside 1 external citations. Polarity classification is still indexing.
citation-role summary
citation-polarity summary
representative citing papers
A new benchmark with real lunar stereo ground truth and analog data shows that sim-to-real fine-tuned monocular depth models achieve large in-domain gains but minimal generalization to actual lunar images.
SonarSweep adapts plane sweeping into an end-to-end neural network for sonar-vision fusion to produce dense accurate depth maps that outperform prior methods in high-turbidity underwater conditions.
LiteMatch uses CVCE and HFE encoders plus CVC-Loss to enable lightweight zero-shot stereo matching with competitive EPE and D1 scores on Scene Flow, KITTI, Middlebury, ETH3D, and DrivingStereo using 3.36M-9.58M parameters.
A zero-shot visual world model trained on one child's experience achieves broad competence on physical understanding benchmarks while matching developmental behavioral patterns.
LAS2 is a series of efficient stereo matching models that reach state-of-the-art zero-shot performance among fast methods while running 1.8-2.7x faster than prior iterative approaches on H200 and Orin hardware.
An automated annotation pipeline combining Grounded DINO and SAM produces usable bounding boxes and masks for weakly supervised defect detection in shearography.
TwinOR creates dynamic photorealistic digital twins of operating rooms that generate realistic RGB and depth data enabling embodied AI perception and localization tasks to match real-world performance levels.
A geometry-aware 4D video generation model trained with cross-view pointmap alignment to produce spatio-temporally consistent future videos from novel viewpoints for robot manipulation.
StereoPolicy fuses left-right image features via cross-attention to deliver consistent gains over RGB, RGB-D, point cloud, and multi-view baselines in simulation and real-robot manipulation tasks.
citing papers explorer
-
EgoEngine: From Egocentric Human Videos to High-Fidelity Dexterous Robot Demonstrations
EgoEngine transforms egocentric human videos into high-fidelity robot data enabling zero-shot visuomotor dexterous policy learning without real-robot demonstrations.
-
LuMon: A Comprehensive Benchmark and Development Suite with Novel Datasets for Lunar Monocular Depth Estimation
A new benchmark with real lunar stereo ground truth and analog data shows that sim-to-real fine-tuned monocular depth models achieve large in-domain gains but minimal generalization to actual lunar images.
-
SonarSweep: Fusing Sonar and Vision for Robust 3D Reconstruction via Plane Sweeping
SonarSweep adapts plane sweeping into an end-to-end neural network for sonar-vision fusion to produce dense accurate depth maps that outperform prior methods in high-turbidity underwater conditions.
-
LiteMatch: Lightweight Zero-Shot Stereo Matching via Cost Volume Stabilization
LiteMatch uses CVCE and HFE encoders plus CVC-Loss to enable lightweight zero-shot stereo matching with competitive EPE and D1 scores on Scene Flow, KITTI, Middlebury, ETH3D, and DrivingStereo using 3.36M-9.58M parameters.
-
Zero-shot World Models Are Developmentally Efficient Learners
A zero-shot visual world model trained on one child's experience achieves broad competence on physical understanding benchmarks while matching developmental behavioral patterns.
-
Lite Any Stereo V2: Faster and Stronger Efficient Zero-Shot Stereo Matching
LAS2 is a series of efficient stereo matching models that reach state-of-the-art zero-shot performance among fast methods while running 1.8-2.7x faster than prior iterative approaches on H200 and Orin hardware.
-
Automated Annotation of Shearographic Measurements Enabling Weakly Supervised Defect Detection
An automated annotation pipeline combining Grounded DINO and SAM produces usable bounding boxes and masks for weakly supervised defect detection in shearography.
-
TwinOR: Photorealistic Digital Twins of Dynamic Operating Rooms for Embodied AI Research
TwinOR creates dynamic photorealistic digital twins of operating rooms that generate realistic RGB and depth data enabling embodied AI perception and localization tasks to match real-world performance levels.
-
Geometry-aware 4D Video Generation for Robot Manipulation
A geometry-aware 4D video generation model trained with cross-view pointmap alignment to produce spatio-temporally consistent future videos from novel viewpoints for robot manipulation.
-
StereoPolicy: Improving Robotic Manipulation Policies via Stereo Perception
StereoPolicy fuses left-right image features via cross-attention to deliver consistent gains over RGB, RGB-D, point cloud, and multi-view baselines in simulation and real-robot manipulation tasks.