Pith. sign in

REVIEW 15 cited by

MediaPipe Hands: On-device Real-time Hand Tracking

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2006.10214 v1 pith:J4RU3MJH submitted 2020-06-18 cs.CV

classification cs.CV
keywords handmediapipepipelinereal-timehandsmodelon-devicetracking
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We present a real-time on-device hand tracking pipeline that predicts hand skeleton from single RGB camera for AR/VR applications. The pipeline consists of two models: 1) a palm detector, 2) a hand landmark model. It's implemented via MediaPipe, a framework for building cross-platform ML solutions. The proposed model and pipeline architecture demonstrates real-time inference speed on mobile GPUs and high prediction quality. MediaPipe Hands is open sourced at https://mediapipe.dev.

Discussion (0). Sign in to comment.

Forward citations

Cited by 15 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 553 citations worldwide. Full citation record

  1. PianoVAM: A Multimodal Piano Performance Dataset

    cs.SD 2025-09 conditional novelty 7.0 of 10

    A multimodal dataset of 21 hours of amateur piano practice with synchronized video, audio, MIDI, hand landmarks, and fingering pseudo-labels, plus benchmarks for audio-only and audio-visual transcription.

  2. EgoTouch: On-Body Touch Input Using AR/VR Headset Cameras

    cs.HC 2025-09 conditional novelty 7.0 of 10

    A headset-mounted RGB camera and a lightweight vision model detect finger-to-skin touches with about 95% frame accuracy and estimate press force, enabling uninstrumented, calibration-free on-body touch input.

  3. DexDirect: Direct Kinesthetic Arm Guidance for Efficient Dexterous Demonstration Collection

    cs.RO 2026-07 conditional novelty 6.0 of 10

    A hybrid kinesthetic-arm-plus-webcam-hand teleoperation interface achieved 17x/3x higher demonstration throughput than vision baselines and trained a 90%-success pick-and-place policy in a ten-person study.

  4. Robotic Contextual Awareness for Human-Robot Collaboration and Environmental Understanding

    cs.RO 2026-07 conditional novelty 6.0 of 10

    Novel person re-identification with continual adaptation plus submap LiDAR SLAM, ground-aware filtering, Gaussian Scan Context, and multi-modal semantic mapping improve robotic contextual awareness for HRC and navigation.

  5. DemoBridge: A Simulation-in-the-Loop Toolkit for Single-View Human Demonstration Retargeting

    cs.RO 2026-07 conditional novelty 6.0 of 10

    DemoBridge retargets single-view human hand demonstrations into physics-validated, collision-aware robot trajectories via whole-trajectory optimization and simulation-in-the-loop re-planning.

  6. From Frame-Level Recognition to Event-Level Confirmation: Repair Traces and Runtime Failure Analysis of Public-Space Gesture Interaction

    cs.AI 2026-05 conditional novelty 6.0 of 10

    Public-space gesture failures form six working classes—model-output degeneration, temporal mismatch, scale instability, coordinate mismatch, lifecycle failure, and feedback mismatch—revealing a recognition-to-interaction gap.

  7. EclipseTouch: Touch Segmentation on Ad Hoc Surfaces using Worn Infrared Shadow Casting

    cs.HC 2025-09 conditional novelty 6.0 of 10

    A headset-integrated system detects touch contact and hover distance on everyday surfaces by analyzing shadows cast by worn infrared emitters, without instrumenting the surface or hands.

  8. 3D Hand Mesh-Guided AI-Generated Malformed Hand Refinement with Hand Pose Transformation via Diffusion Model

    cs.CV 2025-06 conditional novelty 6.0 of 10

    Using 3D hand meshes instead of depth maps to guide diffusion inpainting improves malformed hand refinement and enables training-free hand pose transformation.

  9. Synthetic and Derived Training Images for Campus Waste Detection: A Multi-Seed Evaluation with YOLOv8n

    cs.CV 2026-07 accept novelty 5.0 of 10

    Adding 695 synthetic or derived training images did not improve a YOLOv8n campus-waste detector over the real-only baseline; image source mattered more than image count.

  10. Three ways to share a QPU: Scheduling strategies for hybrid Quantum-HPC applications

    quant-ph 2026-04 unverdicted novelty 5.0 of 10

    Three scheduling strategies for hybrid quantum-HPC systems cut classical resource use by up to 64% or boost QPU utilization depending on workload balance, validated on real hardware.

  11. How Pragmatics Shape Articulation: A Computational Case Study in STEM ASL Discourse

    cs.CL 2025-10 conditional novelty 5.0 of 10

    In a motion-capture study of one ASL instructor-student pair, STEM signs in dialogue were 24.6-44.6% shorter than isolated vocabulary signs, with repeated-mention compression absent in a monologue - evidence of intera...

  12. Analysis of the Dick Effect for AI-based Dynamic Gravimeter

    physics.atom-ph 2025-08 unverdicted novelty 5.0 of 10

    A 0.12 s accelerometer dead time in an atom-interferometer dynamic gravimeter introduces roughly 8 mGal of measurement noise, which the paper attributes to high-frequency aliasing and analyzes with a derived frequency...

  13. A Multimodal Framework for Understanding Collaborative Design Processes

    cs.HC 2025-08 conditional novelty 5.0 of 10

    A multimodal framework and visualization system, reCAPit, automatically extracts activity, attention, and topic segments from workshop recordings to help analysts trace how design decisions emerged.

  14. TrackStudio: An Integrated Toolkit for Markerless Tracking

    cs.CV 2025-11 conditional novelty 4.0 of 10

    TrackStudio packages MediaPipe and Anipose into a GUI with new validation on 76 participants; it shows stable self-consistency, but its error claims lack ground-truth comparison.

  15. Pointing-Guided Target Estimation via Transformer-Based Attention

    cs.RO 2025-09 conditional novelty 4.0 of 10

    A transformer with an engineered finger-to-object angle feature matches a geometric baseline (90% accuracy) for pointing target prediction on a small tabletop dataset.

Pith tools