REVIEW 7 cited by
WiLoR: End-to-end 3D Hand Localization and Reconstruction in-the-wild
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
In recent years, 3D hand pose estimation methods have garnered significant attention due to their extensive applications in human-computer interaction, virtual reality, and robotics. In contrast, there has been a notable gap in hand detection pipelines, posing significant challenges in constructing effective real-world multi-hand reconstruction systems. In this work, we present a data-driven pipeline for efficient multi-hand reconstruction in the wild. The proposed pipeline is composed of two components: a real-time fully convolutional hand localization and a high-fidelity transformer-based 3D hand reconstruction model. To tackle the limitations of previous methods and build a robust and stable detection network, we introduce a large-scale dataset with over than 2M in-the-wild hand images with diverse lighting, illumination, and occlusion conditions. Our approach outperforms previous methods in both efficiency and accuracy on popular 2D and 3D benchmarks. Finally, we showcase the effectiveness of our pipeline to achieve smooth 3D hand tracking from monocular videos, without utilizing any temporal components. Code, models, and dataset are available https://rolpotamias.github.io/WiLoR.
Forward citations
Cited by 7 Pith papers
-
Precise Action-to-Video Generation Through Visual Action Prompts
Skeleton-based visual action prompts give precise, cross-domain action control for video generation of human and robot interactions.
-
Perceiving and Acting in First-Person: A Dataset and Benchmark for Egocentric Human-Object-Human Interactions
Claimed first large-scale egocentric and multi-view dataset of human-object-human assistance (11.4 hours, 1.2M frames) with three benchmarks; only the abstract was assessable because the submitted body text is a diffe...
-
Learning Generalizable Robot Policy with Human Demonstration Video as a Prompt
A two-stage pipeline trains a robot policy that accepts a human demonstration video as a prompt and generalizes beyond its robot training tasks, with success rates of up to 79 percent on known task variations and unde...
-
You Only Teach Once: Learn One-Shot Bimanual Robotic Manipulation from Video Demonstrations
From one human hand demonstration, YOTO generates hundreds of robot demonstrations and trains a bimanual diffusion policy that outperforms ACT, DP, DP3, and EquiBot on five real-world tasks.
-
Self-Supervised Monocular 4D Scene Reconstruction for Egocentric Videos
EgoMono4D estimates depth, camera intrinsics and poses from unlabeled egocentric videos in a single feed-forward pass, reconstructing dense per-frame point clouds better than baseline methods on in-domain and zero-sho...
-
HaWoR: World-Space Hand Motion Reconstruction from Egocentric Videos
HaWoR estimates metric world-space hand trajectories from egocentric video by masking hands from SLAM bundle adjustment, aligning SLAM scale with Metric3D depth, and infilling missing hand frames with a transformer.
-
Hand-Object Contact Detection using Grasp Quality Metrics
A simulation-based pipeline using GraspIt grasp quality metrics detects hand-object contact with 89.3% accuracy on DexYCB frames.
Discussion (0). Continue with ORCID to comment.