REVIEW 15 cited by
MediaPipe Hands: On-device Real-time Hand Tracking
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
We present a real-time on-device hand tracking pipeline that predicts hand skeleton from single RGB camera for AR/VR applications. The pipeline consists of two models: 1) a palm detector, 2) a hand landmark model. It's implemented via MediaPipe, a framework for building cross-platform ML solutions. The proposed model and pipeline architecture demonstrates real-time inference speed on mobile GPUs and high prediction quality. MediaPipe Hands is open sourced at https://mediapipe.dev.
Forward citations
Cited by 15 Pith papers
-
PianoVAM: A Multimodal Piano Performance Dataset
A multimodal dataset of 21 hours of amateur piano practice with synchronized video, audio, MIDI, hand landmarks, and fingering pseudo-labels, plus benchmarks for audio-only and audio-visual transcription.
-
EgoTouch: On-Body Touch Input Using AR/VR Headset Cameras
A headset-mounted RGB camera and a lightweight vision model detect finger-to-skin touches with about 95% frame accuracy and estimate press force, enabling uninstrumented, calibration-free on-body touch input.
-
DexDirect: Direct Kinesthetic Arm Guidance for Efficient Dexterous Demonstration Collection
A hybrid kinesthetic-arm-plus-webcam-hand teleoperation interface achieved 17x/3x higher demonstration throughput than vision baselines and trained a 90%-success pick-and-place policy in a ten-person study.
-
Robotic Contextual Awareness for Human-Robot Collaboration and Environmental Understanding
Novel person re-identification with continual adaptation plus submap LiDAR SLAM, ground-aware filtering, Gaussian Scan Context, and multi-modal semantic mapping improve robotic contextual awareness for HRC and navigation.
-
DemoBridge: A Simulation-in-the-Loop Toolkit for Single-View Human Demonstration Retargeting
DemoBridge retargets single-view human hand demonstrations into physics-validated, collision-aware robot trajectories via whole-trajectory optimization and simulation-in-the-loop re-planning.
-
From Frame-Level Recognition to Event-Level Confirmation: Repair Traces and Runtime Failure Analysis of Public-Space Gesture Interaction
Public-space gesture failures form six working classes—model-output degeneration, temporal mismatch, scale instability, coordinate mismatch, lifecycle failure, and feedback mismatch—revealing a recognition-to-interaction gap.
-
EclipseTouch: Touch Segmentation on Ad Hoc Surfaces using Worn Infrared Shadow Casting
A headset-integrated system detects touch contact and hover distance on everyday surfaces by analyzing shadows cast by worn infrared emitters, without instrumenting the surface or hands.
-
3D Hand Mesh-Guided AI-Generated Malformed Hand Refinement with Hand Pose Transformation via Diffusion Model
Using 3D hand meshes instead of depth maps to guide diffusion inpainting improves malformed hand refinement and enables training-free hand pose transformation.
-
Synthetic and Derived Training Images for Campus Waste Detection: A Multi-Seed Evaluation with YOLOv8n
Adding 695 synthetic or derived training images did not improve a YOLOv8n campus-waste detector over the real-only baseline; image source mattered more than image count.
-
Three ways to share a QPU: Scheduling strategies for hybrid Quantum-HPC applications
Three scheduling strategies for hybrid quantum-HPC systems cut classical resource use by up to 64% or boost QPU utilization depending on workload balance, validated on real hardware.
-
How Pragmatics Shape Articulation: A Computational Case Study in STEM ASL Discourse
In a motion-capture study of one ASL instructor-student pair, STEM signs in dialogue were 24.6-44.6% shorter than isolated vocabulary signs, with repeated-mention compression absent in a monologue - evidence of intera...
-
Analysis of the Dick Effect for AI-based Dynamic Gravimeter
A 0.12 s accelerometer dead time in an atom-interferometer dynamic gravimeter introduces roughly 8 mGal of measurement noise, which the paper attributes to high-frequency aliasing and analyzes with a derived frequency...
-
A Multimodal Framework for Understanding Collaborative Design Processes
A multimodal framework and visualization system, reCAPit, automatically extracts activity, attention, and topic segments from workshop recordings to help analysts trace how design decisions emerged.
-
TrackStudio: An Integrated Toolkit for Markerless Tracking
TrackStudio packages MediaPipe and Anipose into a GUI with new validation on 76 participants; it shows stable self-consistency, but its error claims lack ground-truth comparison.
-
Pointing-Guided Target Estimation via Transformer-Based Attention
A transformer with an engineered finger-to-object angle feature matches a geometric baseline (90% accuracy) for pointing target prediction on a small tabletop dataset.
Discussion (0). Sign in to comment.