A dual-contrastive disentanglement method factorizes videos into independent task and embodiment latents, then uses a parameter-efficient adapter on a frozen video diffusion model to synthesize robot executions from single human demonstrations without paired data.
Gigahands: A massive annotated dataset of bimanual hand activities.arXiv preprint arXiv:2412.04244
2 Pith papers cite this work. Polarity classification is still indexing.
years
2026 2verdicts
UNVERDICTED 2representative citing papers
End-to-end neural pipeline extracts hand geometry from unmasked limited-view images and registers it to a personalized tetrahedral model via volumetric offsets, achieving SOTA on over 12,000 sequences.
citing papers explorer
-
Bridging the Embodiment Gap: Disentangled Cross-Embodiment Video Editing
A dual-contrastive disentanglement method factorizes videos into independent task and embodiment latents, then uses a parameter-efficient adapter on a frozen video diffusion model to synthesize robot executions from single human demonstrations without paired data.
-
VEPHand: View-Efficient Photometric Hand Performance Capture at Scale
End-to-end neural pipeline extracts hand geometry from unmasked limited-view images and registers it to a personalized tetrahedral model via volumetric offsets, achieving SOTA on over 12,000 sequences.