Pith. sign in

REVIEW 8 cited by

HO-Cap: A Capture System and Dataset for 3D Reconstruction and Pose Tracking of Hand-Object Interaction

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.06843 v4 pith:YPIDJXI2 submitted 2024-06-10 cs.CV

classification cs.CV
keywords objectshandssystemcapturedatadatasetposetracking
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We introduce a data capture system and a new dataset, HO-Cap, for 3D reconstruction and pose tracking of hands and objects in videos. The system leverages multiple RGBD cameras and a HoloLens headset for data collection, avoiding the use of expensive 3D scanners or mocap systems. We propose a semi-automatic method for annotating the shape and pose of hands and objects in the collected videos, significantly reducing the annotation time compared to manual labeling. With this system, we captured a video dataset of humans interacting with objects to perform various tasks, including simple pick-and-place actions, handovers between hands, and using objects according to their affordance, which can serve as human demonstrations for research in embodied AI and robot manipulation. Our data capture setup and annotation framework will be available for the community to use in reconstructing 3D shapes of objects and human hands and tracking their poses in videos.

Discussion (0). Sign in to comment.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. EgoTL: Egocentric Think-Aloud Chains for Long-Horizon Tasks

    cs.CV 2026-04 unverdicted novelty 7.0 of 10

    EgoTL provides a new egocentric dataset with think-aloud chains and metric labels that benchmarks VLMs on long-horizon tasks and improves their planning, reasoning, and spatial grounding after finetuning.

  2. Event6D: Event-based Novel Object 6D Pose Tracking

    cs.CV 2026-03 conditional novelty 7.0 of 10

    EventTrack6D tracks 6D poses of unseen objects from event cameras by reconstructing dense intensity and depth cues between frames, generalizing from synthetic training to real data at high speed.

  3. AnyHand: A Large-Scale Synthetic Dataset for RGB(-D) Hand Pose Estimation

    cs.CV 2026-03 accept novelty 6.5 of 10

    Co-training HaMeR and WiLoR on AnyHand (2.5M single-hand + 4.1M hand-object RGB-D images) improves FreiHAND/HO-3D metrics and a lightweight depth-fusion model beats prior RGB-D methods.

  4. DemoBridge: A Simulation-in-the-Loop Toolkit for Single-View Human Demonstration Retargeting

    cs.RO 2026-07 conditional novelty 6.0 of 10

    DemoBridge retargets single-view human hand demonstrations into physics-validated, collision-aware robot trajectories via whole-trajectory optimization and simulation-in-the-loop re-planning.

  5. Glove2Hand: Synthesizing Natural Hand-Object Interaction from Multi-Modal Sensing Gloves

    cs.CV 2026-03 conditional novelty 6.0 of 10

    A 3D-Gaussian-plus-diffusion pipeline translates multi-modal glove HOI videos into photorealistic bare-hand videos, yielding the HandSense dataset that improves contact estimation and occluded tracking.

  6. Co-Evolving Latent Action World Models

    cs.LG 2025-10 unverdicted novelty 6.0 of 10

    CoLA-World jointly trains latent action models and world models with a warm-up phase to achieve co-evolution, matching or exceeding prior two-stage methods in video simulation quality and visual planning performance.

  7. villa-X: Enhancing Latent Action Modeling in Vision-Language-Action Models

    cs.RO 2025-07 unverdicted novelty 6.0 of 10

    villa-X enhances latent action modeling in VLA models to support zero-shot action planning for unseen robot embodiments and open-vocabulary instructions, yielding better manipulation results in simulation and real-wor...

  8. OpenEgo: A Large-Scale Multimodal Egocentric Dataset for Dexterous Manipulation

    cs.CV 2025-09 conditional novelty 5.0 of 10

    OpenEgo is a 1,107-hour unified egocentric manipulation dataset with standardized 21-joint hand poses and timestamped action language, plus a small validation showing a language-conditioned policy learns short-horizon...

Pith tools