Pith. sign in

REVIEW 2 major objections 4 minor 1 cited by

emg2pose: A Large and Diverse Benchmark for Surface Electromyographic Hand Pose Estimation

T0 review · 2 major / 4 minor · reviewed 2026-08-11 · deepseek-v4-flash

Pith's one-line read emg2pose releases a 193-user, 370-hour paired wrist-sEMG and hand-pose dataset, with held-out benchmarks showing generalization improves as user and behavior diversity grow.

desk verdict A genuinely large sEMG pose dataset with a novel held-out-stage benchmark; the label-validation gap and double-counted headline numbers are real but fixable. read the letter →

arxiv 2412.02725 v1 pith:IWV4ZJVB submitted 2024-12-02 cs.CV cs.HCcs.LG

classification cs.CVcs.HCcs.LG
keywords emg2posesurfaceelectromyographyhandposeestimationbenchmarkdatasetdomaingeneralizationwrist-wornsEMGmotioncapturelabelsregression
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper's aim is to remove the data bottleneck that keeps sEMG hand-pose models from working for people and gestures they were not trained on. It releases emg2pose, claimed to be the largest public dataset pairing wrist sEMG with hand-pose labels: 193 users, 370 hours, 751 sessions, 29 movement stages, and 80 million labelled frames, a scale it argues is comparable to major computer-vision hand datasets. Alongside the data it defines three held-out generalization tasks (unseen users, unseen stages, and unseen user-stage combinations) and reports baselines, including a new velocity-predicting model whose errors shrink as training users and stages increase. A sympathetic reader would care because this turns sEMG-to-pose from a small-data personalization problem into a community-scale benchmark, bringing the field to the scale standards that have driven vision-based hand tracking.

What carries the argument

The load-bearing mechanism is the paired recording pipeline: a 16-channel, 2 kHz wrist sEMG band worn simultaneously with 19 reflective markers per hand tracked by 26 cameras, whose 3D positions are converted by an inverse-kinematics solver with a personalized hand model into 20 joint-angle degrees of freedom, then filtered and resampled to 2 kHz. On the modelling side, vemg2pose carries the baseline results: a causal strided convolutional featurizer built from time-depth separable convolutions turns sEMG into features at 50 Hz, and an autoregressive LSTM predicts joint angular velocities that are integrated into angles; for tracking the ground-truth initial pose seeds the integrator, while for regression the first 250 ms of angles are also predicted. The velocity representation is what lets one model handle both tasks and keeps predictions smooth.

What would settle it

Take a random sample of several hundred frames from occlusion-heavy and fist-clench stages, have annotators independently label joint angles using a different capture method, and compare them with the dataset's inverse-kinematics angles; if the median per-joint difference approaches the 7 to 15 degree errors reported for the models, the benchmark's headline numbers mostly measure label noise.

Watch

Extended reading notes

Core claim

The central claim is that a dataset large and diverse enough to span user anatomy, sensor placement, and hand kinematics makes it possible to learn continuous hand pose from wrist muscle signals for people and movement types never seen in training. The paper reports 193 users, 370 hours, 29 stages, and 80 million labelled frames, with 16-channel 2 kHz sEMG synchronized to 26-camera motion-capture labels within 10 ms; it calls this the largest open sEMG pose dataset and comparable in scale to large vision hand datasets. On held-out users, stages, and user-stage combinations, the velocity-based vemg2pose baseline achieves mean joint-angle errors of about 12.2, 15.2, and 15.8 degrees in regression and 7.7, 11.2, and 11.0 degrees in tracking, beating reimplementations of prior sEMG pose networks. Scale experiments show held-out error decreasing as training users or stages are added, which the authors present as evidence that breadth across these axes, not just dataset size, drives generalization.

Load-bearing premise

The motion-capture pipeline that turns marker positions into joint angles is accurate enough to serve as ground truth for all 193 users, even though the solver failed on 12.7 percent of frames and its angle estimates were never checked against an independent gold standard.

Editorial extensions

If this is right

  • On emg2pose's held-out splits, sEMG alone supports continuous hand-pose estimation for unseen users at roughly 12.2 degrees mean joint-angle error in regression and 7.7 degrees in tracking, with the velocity model outperforming both reimplemented prior architectures.
  • The three test sets let researchers measure generalization to new anatomy, new kinematics, and both at once; the user-plus-stage condition is the paper's proposed proxy for real-world deployment.
  • Dataset scale is shown to be causally linked to generalization: subsampling training users or stages degrades held-out performance, so further scaling along these axes should keep reducing error.
  • Stages designed to confound vision systems, including occlusion and hand-hand or hand-object interaction, do not degrade sEMG tracking, indicating the modality covers cases where cameras fail.
  • The open benchmark and baselines give the community a shared platform for exploring sequence models, probabilistic decoding, and personalization for biosignal interfaces.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the inverse-kinematics labels contain noise of the same order as the reported errors, and the solver failed on 12.7 percent of frames without an independent gold-standard check, then part of the measured error is label error; an independent label audit on a few hundred frames would separate the two.
  • The velocity-integration design hints that other derivative-sensing wearables, such as inertial, ultrasound, or impedance sensors, could borrow the same architecture whenever the measurement responds to movement rather than to static pose.
  • The user-plus-stage held-out split could be adopted as a general time-series domain-generalization benchmark beyond sEMG, because it cleanly separates shift in the signal source from shift in the output behavior.
  • The paper's own limitation notes imply that adding real-world signal aggressors such as sweat, electrode contact changes, and muscle fatigue, as well as wrist tracking, will be needed before the benchmark reflects in-the-wild performance; those are testable extensions rather than demonstrated results.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

2 major / 4 minor

Summary. The paper introduces emg2pose, a large benchmark dataset of wrist surface electromyography (sEMG) and hand pose labels obtained from a 26-camera motion capture rig. The dataset is reported to span 193 users, 370 hours, 29 kinematic stages, and 80M labeled frames, with pose labels given as joint angles. The authors provide three baselines (NeuroPose, SensingDynamics, and a new velocity-based vemg2pose model), define regression and tracking tasks, and evaluate generalization to held-out users, held-out stages, and held-out user-stage combinations. The paper also includes a datasheet, discussion of limitations, and links to code and data.

Significance. If the dataset and labels are as described, this is a valuable contribution to sEMG-based hand pose estimation and to benchmark research more broadly. The strongest assets are the scale relative to prior open sEMG datasets, the explicit three-axis generalization evaluation (users, stages, and user-stage combinations), the release of code and baseline models, and the unusually detailed documentation of collection, consent, preprocessing, and limitations. The paper is transparent about several weaknesses, including occlusion-induced label degradation, the default hand model used for landmark metrics, and the absence of seed variation in reported results. However, two issues need attention: the accuracy of the IK-derived pose labels is not validated against an independent gold standard, and the headline scale figures double-count the two hands. These issues affect the central claims and should be addressed before publication.

major comments (2)
  1. [Section 3.2, Appendix A, Appendix B.4] The central claim that emg2pose provides "high-quality hand pose labels" is not yet supported because the motion-capture inverse kinematics labels are never validated against an independent gold standard. Appendix A reports a 0.32 degree difference between filtered and unfiltered signals, which quantifies smoothing, not IK accuracy. Section 3.2 states that the IK solver failed on 12.7% of frames, typically due to simultaneously occluded markers, and Section 3.5 says those time-points are skipped during training and evaluation. Since Appendix B.4 concedes that occlusion "hinders label quality for gestures such as fist clenching," the skipped frames are plausibly concentrated in the hardest kinematic conditions; if so, the reported held-out errors are optimistic and the effective kinematic diversity is reduced. I request an independent validation study on a subset (e.g., manual marker annotation, a second sensing modality, or synthetic marker-dropout analysis), a per-stage report of failure rates, and an evaluation of how results change when failed frames are handled differently.
  2. [Section 1, Table 2, Section 3.2] The headline scale figures double-count the two hands. Section 3.2 notes that hours count left- and right-hand data separately, and Table 2 lists both "per hand" and "across hands" frame counts, but the abstract and introduction use the per-hand number (370 hours, 80M frames) without qualification. Unique hours are roughly 185 and unique frames roughly 40M. The largest-sEMG-dataset claim probably survives (Atzori et al. [2014] reports 37 hours), but the comparison with CV datasets such as Sener et al. [2022] (111M frames) is misleading if 80M is used. Please state unique-count figures in the abstract and introduction, or explicitly define 370 hours and 80M frames as per-hand totals.
minor comments (4)
  1. [Section 3.5] There are typos: "sEMG meaures" and "between between predicted and ground truth fingertip locations."
  2. [Appendix A, Section 3.2, Appendix B.2.1] The stated stage duration is inconsistent: Section 3.2 says 45–120 s, Appendix A says 30–120 s, and Appendix B.2.1 says 45–60 s (freeform 60–120 s). Please reconcile.
  3. [Section 3.5 and Table 7] The name "emg2pose" is used both for the dataset and for the positional (non-velocity) baseline model in Table 7, which is confusing next to "vemg2pose."
  4. [Table 4 and Checklist 3(c)] The paper states that seed variance is negligible but does not report the underlying numbers; adding a sentence with the observed spread across seeds would strengthen reproducibility.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the benchmark's empirical claims are self-contained and its held-out splits are independent of baseline fitting.

full rationale

The paper's central contribution is an empirical dataset and benchmark rather than a derivation chain, so there is no predicted quantity that reduces to its own input by construction. The strongest claims are scale statistics (193 users, 370 hours, 80M frames) and held-out generalization errors, all of which are computed from the collected data and fixed train/val/test splits. The self-citations to CTRL-labs at Reality Labs et al. [2024] supply hardware context for the sEMG-RD wristband, and the citation to Han et al. [2018] supplies the marker-based inverse kinematics labeling method; neither citation determines the reported test errors or the dataset scale, and the labeling method is an external, independently developed pipeline rather than a uniqueness theorem or ansatz imposed by this paper. The baselines are trained on the training split and evaluated on fixed held-out users, stages, and user-stage combinations, with no parameter fitted to the test error and renamed as a prediction. The velocity-integration design of vemg2pose is an architectural choice, not a self-defined metric. The acknowledged 12.7% IK failure frames being skipped during training and evaluation, and the default-hand-model bias in landmark distance, are validation and correctness concerns rather than circular reductions. No equation in the paper defines a reported result in terms of itself or of a fitted parameter masquerading as a prediction, so the appropriate finding is no significant circularity.

Assumptions & free parameters 5 free parameters · 4 assumptions · 0 invented entities

The dataset statistics depend on the acquisition hardware and labeling assumptions, not on fitted parameters. The listed free parameters are baseline hyperparameters, which affect the reported error numbers but not the dataset's scale claims. The key axioms are the sufficiency of sEMG for pose inference, the accuracy of the motion-capture IK labels, and the anatomical validity of the default hand model used for landmark metrics.

free parameters (5)
  • Fingertip loss weight = 0.01
    Joint angle L1 loss weight is 1 and fingertip loss weight is 0.01; chosen during training (Section 3.5), affects all baseline results.
  • vemg2pose output scale = 0.01
    LSTM output scaled by 0.01, stated in Appendix C.1 as improving training; a hand-tuned baseline hyperparameter.
  • Regression initial-state window P = 250 ms
    For the regression task the decoder predicts angles for the first P time steps before velocity integration; selected via hyperparameter sweep (Appendix C.1).
  • Joint angle low-pass filter cutoff = 15 Hz
    Applied to motion capture joint angles to remove jitter, following Ingram et al. [2008]; affects label sharpness and all resulting metrics.
  • Evaluation trajectory length = 5 s
    Metrics are evaluated on 5-second trajectories; a benchmark design choice that influences reported errors.
assumptions (4)
  • domain assumption sEMG from the sEMG-RD wristband contains sufficient information about muscle activity to infer hand pose given enough data.
    Motivates the entire benchmark (Sections 1 and 3.1); if false, pose regression from these recordings is impossible regardless of dataset size.
  • domain assumption The motion-capture inverse kinematics pipeline produces accurate ground-truth joint angles; failed frames (12.7%) are skipped without biasing the data.
    Section 3.2 and Appendix A; label quality is not validated against an independent gold standard.
  • domain assumption Software timestamp alignment keeps sEMG and motion capture streams within 10 ms relative latency, approximately the Nyquist limit of the 60 Hz mocap.
    Appendix B.2; the alignment supports the joint angle to sEMG correspondence used for all training and evaluation.
  • domain assumption The default hand model used to convert joint angles to landmark positions is an adequate approximation for all users.
    Section 3.3 and Section 5; the authors note this introduces bias in landmark distance metrics.

how reviews work

0 comments
Cite this review

Pith. "Pith review of emg2pose: A Large and Diverse Benchmark for Surface Electromyographic Hand Pose Estimation." pith.science (2026). https://pith.science/paper/IWV4ZJVB

@misc{pith2026241202725,
  author       = {Pith},
  title        = {Pith review of: emg2pose: A Large and Diverse Benchmark for Surface Electromyographic Hand Pose Estimation},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/IWV4ZJVB}},
  note         = {Machine review of arXiv:2412.02725}
}
read the original abstract

Hands are the primary means through which humans interact with the world. Reliable and always-available hand pose inference could yield new and intuitive control schemes for human-computer interactions, particularly in virtual and augmented reality. Computer vision is effective but requires one or multiple cameras and can struggle with occlusions, limited field of view, and poor lighting. Wearable wrist-based surface electromyography (sEMG) presents a promising alternative as an always-available modality sensing muscle activities that drive hand motion. However, sEMG signals are strongly dependent on user anatomy and sensor placement, and existing sEMG models have required hundreds of users and device placements to effectively generalize. To facilitate progress on sEMG pose inference, we introduce the emg2pose benchmark, the largest publicly available dataset of high-quality hand pose labels and wrist sEMG recordings. emg2pose contains 2kHz, 16 channel sEMG and pose labels from a 26-camera motion capture rig for 193 users, 370 hours, and 29 stages with diverse gestures - a scale comparable to vision-based hand pose datasets. We provide competitive baselines and challenging tasks evaluating real-world generalization scenarios: held-out users, sensor placements, and stages. emg2pose provides the machine learning community a platform for exploring complex generalization problems, holding potential to significantly enhance the development of sEMG-based human-computer interactions.

Figures

Figures reproduced from arXiv: 2412.02725 by the authors.

Figure 1
Figure 1. We introduce the emg2pose dataset and benchmark to facilitate the development of pose estimation models from sEMG. Our vemg2pose model is capable of estimating in real-time hand pose (lower) from held-out users wearing an sEMG wristband (top). See text for further details. Abstract Hands are the primary means through which humans interact with the world. Reliable and always-available hand pose inference could yield … view at source ↗
Figure 2
Figure 2. Dataset composition: a) sEMG-RD wrist-band and motion capture marker (white dots) setup. b) Dataset breakdown. i) Users are prompted to perform a sequence of movement types (gestures), such as counting up and down. sEMG and poses are recorded simultaneously. ii) Groups of specific gesture types comprise a stage, such as counting. Stages are partitioned into train/val/test splits (see Section 3.4). Our dataset consis… view at source ↗
Figure 3
Figure 3. vemg2pose tracking performance break down by stage and generalization condition. Distributions are over users. Note the variability in performance across stages. Each box shows the median and interquartile range (IQR), and whiskers show the minimum and maximum values that are within 1.5 times the IQR of the lower and upper quartiles. 4.2 Analysis on Challenging Stages for Vision-Based Systems [PITH_FULL_IMAGE:figur… view at source ↗
Figures from the paper (12 more)
Figure 4
Figure 4. Figure 4: vemg2pose tracking results with/without occlusion (left) and physical interactions (right). Distributions are over users. See Appendix D.1 for more details. Some stages were specifically designed to test behaviors that are known to be challenging for vision￾based hand …
Figure 5
Figure 5. Figure 5: Median percentile held-out user and stage (Counting2). Top: motion capture; bottom: vemg2pose, tracking predictions. Clips unroll evenly left-to-right over a 2 second segment. We plot vemg2pose, tracking real-time online and offline kinematic predictions for held-out u…
Figure 6
Figure 6. Figure 6: Generalization vs. number of training users (left two) or stages (right three) for vemg2pose tracking. We subsampled the training users/stages but evaluated on the same held-out users/stages. As seen, performance improves with the number of training users/stages, demon…
Figure 7
Figure 7. Figure 7: Excluding stages (left) or users (right) from the training set [PITH_FULL_IMAGE:figures/full_fig_p010_7.png]
Figure 8
Figure 8. Figure 8: Participant demographic information. Train and val users are shown in the top row, and test users on the bottom. Notice that test users are representative of the population of train/val users. B.1 sEMG Sensing sEMG data were collected using the sEMG-RD [CTRL-labs at Re…
Figure 9
Figure 9. Figure 9: vemg2pose vs. emg2pose for tracking and regression tasks. Distributions are over users. Box plots take the same format as [PITH_FULL_IMAGE:figures/full_fig_p023_9.png]
Figure 10
Figure 10. Figure 10: Performance decomposition per finger: for tracking task, vemg2pose. Error per finger is measured by averaging the errors of the joints associated with each finger. Distributions are over users. Box plots take the same format as [PITH_FULL_IMAGE:figures/full_fig_p024_…
Figure 11
Figure 11. Figure 11: Performance decomposition across joint groups: for tracking task, vemg2pose. Perfor￾mance broken down by joint according to their proximal-distal location. Proximal is CMC for the thumb and MCP for other fingers; Mid is MCP for thumb and PIP for other fingers; and Dis…
Figure 12
Figure 12. Figure 12: Held-Out User, Stage tracking, top 15% stage (Gesture2), median user. Top: motion capture; bottom: vemg2pose, tracking predictions. Clips unroll evenly left-to-right over a 2 seconds [PITH_FULL_IMAGE:figures/full_fig_p026_12.png]
Figure 13
Figure 13. Figure 13: Held-Out User, Stage tracking, bottom 15% percentile stage (Counting1), median user. Top: motion capture; bottom: vemg2pose, tracking predictions. Clips unroll evenly left-to-right over a 2 seconds [PITH_FULL_IMAGE:figures/full_fig_p026_13.png]
Figure 14
Figure 14. Figure 14: Held-Out User, Stage tracking, median stage (Counting2), top 15% percentile user. Top: motion capture; bottom: vemg2pose, tracking predictions. Clips unroll evenly left-to-right over a 2 seconds [PITH_FULL_IMAGE:figures/full_fig_p026_14.png]
Figure 15
Figure 15. Figure 15: Held-Out User, Stage tracking, median stage (Wiggling2), bottom 15% percentile user. Top: motion capture; bottom: vemg2pose, tracking predictions. Clips unroll evenly left-to-right over a 2 seconds. 26 [PITH_FULL_IMAGE:figures/full_fig_p026_15.png]

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 3 citations worldwide. Full citation record

  1. Sensor-Placement-Agnostic Sonomyography: Toward Continuous High-Dimensional Control by Users with Tetraplegia

    cs.HC 2026-07 conditional novelty 5.0 of 10

    A three-pose-calibrated optical-flow sonomyography algorithm gives continuous 1-DOF cursor control from arbitrary sensor locations, plus a first pilot of continuous 2-DOF control.

Reference graph

Works this paper leans on

78 extracted references · 68 canonical work pages · cited by 1 Pith paper

  1. [1]

    u ller, and S. G \

    P. Achenbach, S. Laux, D. Purdack, P. N. M \"u ller, and S. G \"o bel. Give me a sign: Using data gloves for static hand-shape recognition. Sensors, 23 0 (24): 0 9847, 2023

  2. [2]

    Achiam, S

    J. Achiam, S. Adler, S. Agarwal, L. Ahmad, I. Akkaya, F. L. Aleman, D. Almeida, J. Altenschmidt, S. Altman, S. Anadkat, et al. Gpt-4 technical report. arXiv preprint arXiv:2303.08774, 2023

  3. [3]

    C. Amma, T. Krings, J. B \"o er, and T. Schultz. Advancing muscle-computer interfaces with high-density electromyography. In Proceedings of the 33rd Annual ACM Conference on Human Factors in Computing Systems, pages 929--938, 2015

  4. [4]

    Atzori, A

    M. Atzori, A. Gijsberts, C. Castellini, B. Caputo, A.-G. M. Hager, S. Elsig, G. Giatsidis, F. Bassetto, and H. M \"u ller. Electromyography data for non-invasive naturally-controlled robotic hand prostheses. Scientific data, 1 0 (1): 0 1--13, 2014

  5. [5]

    Biswas, S

    K. Biswas, S. Kumar, S. Banerjee, and A. K. Pandey. Smu: smooth activation function for deep networks using smoothing maximum technique. arXiv preprint arXiv:2111.04682, 2021

  6. [6]

    Boukhayma, R

    A. Boukhayma, R. d. Bem, and P. H. Torr. 3d hand shape and pose from images in the wild. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 10843--10852, 2019

  7. [7]

    Brahmbhatt, C

    S. Brahmbhatt, C. Tang, C. D. Twigg, C. C. Kemp, and J. Hays. Contactpose: A dataset of grasps with object contact and hand pose. In Computer Vision--ECCV 2020: 16th European Conference, Glasgow, UK, August 23--28, 2020, Proceedings, Part XIII 16, pages 361--378. Springer, 2020

  8. [8]

    Brunetti, D

    A. Brunetti, D. Buongiorno, G. F. Trotta, and V. Bevilacqua. Computer vision and deep learning techniques for pedestrian detection and tracking: A survey. Neurocomputing, 300: 0 17--33, 2018

Show all 78 references
  1. [9]

    Caggiano, H

    V. Caggiano, H. Wang, G. Durandau, M. Sartori, and V. Kumar. Myosuite--a contact-rich simulation suite for musculoskeletal motor control. arXiv preprint arXiv:2205.13600, 2022

  2. [10]

    Y. Cai, L. Ge, J. Cai, and J. Yuan. Weakly-supervised 3d hand pose estimation from monocular rgb images. In Proceedings of the European conference on computer vision (ECCV), pages 666--682, 2018

  3. [11]

    Sussillo, P

    CTRL - labs at Reality Labs , D. Sussillo, P. Kaifosh, and T. Reardon. A generic noninvasive neuromotor interface for human-computer interaction. bioRxiv, 2024. doi:10.1101/2024.02.23.581779. URL https://www.biorxiv.org/content/early/2024/02/28/2024.02.23.581779

  4. [12]

    Danelljan, L

    M. Danelljan, L. V. Gool, and R. Timofte. Probabilistic regression for visual tracking. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 7183--7192, 2020

  5. [13]

    Darvish, L

    K. Darvish, L. Penco, J. Ramos, R. Cisneros, J. Pratt, E. Yoshida, S. Ivaldi, and D. Pucci. Teleoperation of humanoid robots: A survey. IEEE Transactions on Robotics, 39 0 (3): 0 1706--1727, 2023

  6. [14]

    Y. Du, W. Jin, W. Wei, Y. Hu, and W. Geng. Surface emg-based inter-session gesture recognition enhanced by deep domain adaptation. Sensors, 17 0 (3): 0 458, 2017

  7. [15]

    Dunion, T

    M. Dunion, T. McInroe, K. S. Luck, J. Hanna, and S. Albrecht. Conditional mutual information for disentangled representations in reinforcement learning. Advances in Neural Information Processing Systems, 36, 2024

  8. [16]

    Z. Fan, O. Taheri, D. Tzionas, M. Kocabas, M. Kaufmann, M. J. Black, and O. Hilliges. Arctic: A dataset for dexterous bimanual hand-object manipulation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 12943--12954, 2023

  9. [17]

    I. T. Gatt, T. Allen, and J. Wheat. Accuracy and repeatability of wrist joint angles in boxing using an electromagnetic tracking system. Sports Engineering, 23: 0 1--10, 2020

  10. [18]

    L. Ge, H. Liang, J. Yuan, and D. Thalmann. Robust 3d hand pose estimation in single depth images: from single-view cnn to multi-view cnns. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 3593--3601, 2016

  11. [19]

    Gebru, J

    T. Gebru, J. Morgenstern, B. Vecchione, J. W. Vaughan, H. Wallach, H. D. Iii, and K. Crawford. Datasheets for datasets. Communications of the ACM, 64 0 (12): 0 86--92, 2021

  12. [20]

    A. Gu, K. Goel, and C. R \'e . Efficiently modeling long sequences with structured state spaces. arXiv preprint arXiv:2111.00396, 2021

  13. [21]

    Hampali, M

    S. Hampali, M. Rad, M. Oberweger, and V. Lepetit. Honnotate: A method for 3d annotation of hand and object poses. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 3196--3206, 2020

  14. [22]

    S. Han, B. Liu, R. Wang, Y. Ye, C. D. Twigg, and K. Kin. Online optical marker-based hand tracking with deep labels. Acm transactions on graphics (tog), 37 0 (4): 0 1--10, 2018

  15. [23]

    S. Han, B. Liu, R. Cabezas, C. D. Twigg, P. Zhang, J. Petkau, T.-H. Yu, C.-J. Tai, M. Akbay, Z. Wang, et al. Megatrack: monochrome egocentric articulated hand-tracking for virtual reality. ACM Transactions on Graphics (ToG), 39 0 (4): 0 87--1, 2020

  16. [24]

    S. Han, P. Wu, Y. Zhang, B. Liu, L. Zhang, Z. Wang, W. Si, P. Zhang, Y. Cai, T. Hodan, R. Cabezas, L. Tran, M. Akbay, T. Yu, C. Keskin, and R. Wang. Umetrack: Unified multi-view end-to-end hand tracking for VR . In SIGGRAPH Asia 2022 Conference Papers, SA 2022, Daegu, Republic...

  17. [25]

    Hannun, A

    A. Hannun, A. Lee, Q. Xu, and R. Collobert. Sequence-to-sequence speech recognition with time-depth separable convolutions. arXiv preprint arXiv:1904.02619, 2019

  18. [26]

    J. N. Ingram, K. P. K \"o rding, I. S. Howard, and D. M. Wolpert. The statistics of natural hand movements. Experimental brain research, 188: 0 223--236, 2008

  19. [27]

    E. Jang, A. Irpan, M. Khansari, D. Kappler, F. Ebert, C. Lynch, S. Levine, and C. Finn. Bc-z: Zero-shot task generalization with robotic imitation learning. In Conference on Robot Learning, pages 991--1002. PMLR, 2022

  20. [28]

    Jiang, X

    X. Jiang, X. Liu, J. Fan, X. Ye, C. Dai, E. A. Clancy, M. Akay, and W. Chen. Open access dataset, toolbox and benchmark processing results of high-density surface electromyogram recordings. IEEE Transactions on Neural Systems and Rehabilitation Engineering, 29: 0 1035--1046, 2021

  21. [29]

    J. D. M.-W. C. Kenton and L. K. Toutanova. Bert: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of naacL-HLT, volume 1, page 2. Minneapolis, Minnesota, 2019

  22. [30]

    Kirillov, E

    A. Kirillov, E. Mintun, N. Ravi, H. Mao, C. Rolland, L. Gustafson, T. Xiao, S. Whitehead, A. C. Berg, W.-Y. Lo, et al. Segment anything. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 4015--4026, 2023

  23. [31]

    Kojima, S

    T. Kojima, S. S. Gu, M. Reid, Y. Matsuo, and Y. Iwasawa. Large language models are zero-shot reasoners. Advances in neural information processing systems, 35: 0 22199--22213, 2022

  24. [32]

    Krasoulis, I

    A. Krasoulis, I. Kyranou, M. S. Erden, K. Nazarpour, and S. Vijayakumar. Improved prosthetic hand control with concurrent use of myoelectric and inertial measurements. Journal of neuroengineering and rehabilitation, 14: 0 1--14, 2017

  25. [33]

    Laput and C

    G. Laput and C. Harrison. Sensing fine-grained hand activity with smartwatches. In Proceedings of the 2019 CHI Conference on Human Factors in Computing Systems, pages 1--13, 2019

  26. [34]

    Laput, R

    G. Laput, R. Xiao, and C. Harrison. Viband: High-fidelity bio-acoustic sensing using commodity smartwatch accelerometers. In Proceedings of the 29th Annual Symposium on User Interface Software and Technology, UIST '16, page 321–333, New York, NY, USA, 2016. Association for Com...

  27. [35]

    Lauri, D

    M. Lauri, D. Hsu, and J. Pajarinen. Partially observable markov decision processes in robotics: A survey. IEEE Transactions on Robotics, 39 0 (1): 0 21--40, 2022

  28. [36]

    Y. Liu, S. Zhang, and M. Gowda. Neuropose: 3d hand pose tracking using emg wearables. In Proceedings of the Web Conference 2021, pages 1471--1482, 2021

  29. [37]

    Y. Liu, Y. Liu, C. Jiang, K. Lyu, W. Wan, H. Shen, B. Liang, Z. Fu, H. Wang, and L. Yi. Hoi4d: A 4d egocentric dataset for category-level human-object interaction. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 21013--21022, 2022

  30. [38]

    Y. Luo, Y. Li, P. Sharma, W. Shou, K. Wu, M. Foshey, B. Li, T. Palacios, A. Torralba, and W. Matusik. Learning human--environment interactions using conformal tactile textiles. Nature Electronics, 4 0 (3): 0 193--201, 2021

  31. [39]

    McIntosh, A

    J. McIntosh, A. Marzo, M. Fraser, and C. Phillips. Echoflex: Hand gesture recognition using ultrasound imaging. In Proceedings of the 2017 CHI Conference on Human Factors in Computing Systems, pages 1923--1934, 2017

  32. [40]

    Merletti and D

    R. Merletti and D. Farina. Surface electromyography: physiology, engineering, and applications. John Wiley & Sons, 2016

  33. [41]

    Moon, S.-I

    G. Moon, S.-I. Yu, H. Wen, T. Shiratori, and K. M. Lee. Interhand2. 6m: A dataset and baseline for 3d interacting hand pose estimation from a single rgb image. In Computer Vision--ECCV 2020: 16th European Conference, Glasgow, UK, August 23--28, 2020, Proceedings, Part XX 16, p...

  34. [42]

    G. Moon, S. Saito, W. Xu, R. Joshi, J. Buffalini, H. Bellan, N. Rosen, J. Richardson, M. Mize, P. De Bree, et al. A dataset of relighted 3d interacting hands. Advances in Neural Information Processing Systems, 36, 2024

  35. [43]

    Mueller, D

    F. Mueller, D. Mehta, O. Sotnychenko, S. Sridhar, D. Casas, and C. Theobalt. Real-time hand tracking under occlusion from an egocentric rgb-d sensor. In Proceedings of the IEEE International Conference on Computer Vision, pages 1154--1163, 2017

  36. [44]

    Mueller, F

    F. Mueller, F. Bernard, O. Sotnychenko, D. Mehta, S. Sridhar, D. Casas, and C. Theobalt. Ganerated hands for real-time 3d hand tracking from monocular rgb. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 49--59, 2018

  37. [45]

    Palermo, M

    F. Palermo, M. Cognolato, A. Gijsberts, H. M \"u ller, B. Caputo, and M. Atzori. Repeatability of grasp recognition for robotic hand prosthesis control based on semg data. In 2017 International Conference on Rehabilitation Robotics (ICORR), pages 1154--1159. IEEE, 2017

  38. [46]

    X. Pan, P. Luo, J. Shi, and X. Tang. Two at once: Enhancing learning and generalization capacities via ibn-net. In Proceedings of the european conference on computer vision (ECCV), pages 464--479, 2018

  39. [47]

    F. S. Parizi, E. Whitmire, and S. Patel. Auraring: Precise electromagnetic finger tracking. Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies, 3 0 (4): 0 1--28, 2019

  40. [48]

    Park, T.-K

    G. Park, T.-K. Kim, and W. Woo. 3d hand pose estimation with a single infrared camera via domain transfer learning. In 2020 IEEE International Symposium on Mixed and Augmented Reality (ISMAR), pages 588--599. IEEE, 2020

  41. [49]

    Collaborative data science, 2015

    Plotly Technologies Inc. Collaborative data science, 2015. URL https://plot.ly

  42. [50]

    Quivira, T

    F. Quivira, T. Koike-Akino, Y. Wang, and D. Erdogmus. Translating semg signals to continuous hand poses using recurrent neural networks. In 2018 IEEE EMBS International Conference on Biomedical & Health Informatics (BHI), pages 166--169. IEEE, 2018

  43. [51]

    Rawat, S

    S. Rawat, S. Vats, and P. Kumar. Evaluating and exploring the myo armband. In 2016 International Conference System Modeling & Advancement in Research Trends (SMART), pages 115--120. IEEE, 2016

  44. [52]

    Roda-Sales, J

    A. Roda-Sales, J. L. Sancho-Bru, M. Vergara, V. Gracia-Ib \'a \ n ez, and N. J. Jarque-Bou. Effect on manual skills of wearing instrumented gloves during manipulation. Journal of biomechanics, 98: 0 109512, 2020

  45. [53]

    Samarth, T

    B. Samarth, T. Chengcheng, D. T. Christopher, C. K. Charles, and H. James. Contactpose: A dataset of grasps with object contact and hand pose. In European Conference on Computer Vision (ECCV), 2020

  46. [54]

    Santos Carreras

    L. Santos Carreras. Increasing haptic fidelity and ergonomics in teleoperated surgery. Technical report, EPFL, 2012

  47. [55]

    Scheggi, L

    S. Scheggi, L. Meli, C. Pacchierotti, and D. Prattichizzo. Touch the virtual reality: using the leap motion controller for hand tracking and wearable tactile devices for immersive haptic rendering. In ACM SIGGRAPH 2015 Posters, pages 1--1. 2015

  48. [56]

    Sener, D

    F. Sener, D. Chatterjee, D. Shelepov, K. He, D. Singhania, R. Wang, and A. Yao. Assembly101: A large-scale multi-view video dataset for understanding procedural activities. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 21096--21106, 2022

  49. [57]

    Z. Shen, J. Yi, X. Li, M. H. P. Lo, M. Z. Chen, Y. Hu, and Z. Wang. A soft stretchable bending sensor and data glove applications. Robotics and biomimetics, 3 0 (1): 0 22, 2016

  50. [58]

    Shrestha, C

    S. Shrestha, C. Ferm \"u ller, T. Huang, P. T. Win, A. Zukerman, C. M. Parameshwara, and Y. Aloimonos. Aimusicguru: Music assisted human pose correction. arXiv preprint arXiv:2203.12829, 2022

  51. [59]

    R. C. S \^ mpetru, A. Arkudas, D. I. Braun, M. Osswald, D. S. de Oliveira, B. Eskofier, T. M. Kinfe, and A. Del Vecchio. Sensing the full dynamics of the human hand with a neural interface and deep learning. bioRxiv, pages 2022--07, 2022 a

  52. [60]

    R. C. S \^ mpetru, M. Osswald, D. I. Braun, D. S. Oliveira, A. L. Cakici, and A. Del Vecchio. Accurate continuous prediction of 14 degrees of freedom of the hand from myoelectrical signals through convolutive deep learning. In 2022 44th Annual International Conference of the I...

  53. [61]

    Sosin, D

    I. Sosin, D. Kudenko, and A. Shpilman. Continuous gesture recognition from semg sensor data with recurrent neural networks and adversarial domain adaptation. In 2018 15Th international conference on control, automation, robotics and vision (ICARCV), pages 1436--1441. IEEE, 2018

  54. [62]

    M. T. Spaan. Partially observable markov decision processes. In Reinforcement learning: State-of-the-art, pages 387--414. Springer, 2012

  55. [63]

    Spurr, J

    A. Spurr, J. Song, S. Park, and O. Hilliges. Cross-modal deep variational hand pose estimation. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 89--98, 2018

  56. [64]

    Spurr, U

    A. Spurr, U. Iqbal, P. Molchanov, O. Hilliges, and J. Kautz. Weakly supervised 3d hand pose estimation via biomechanical constraints. In European conference on computer vision, pages 211--228. Springer, 2020

  57. [65]

    Spurr, A

    A. Spurr, A. Dahiya, X. Wang, X. Zhang, and O. Hilliges. Self-supervised 3d hand pose estimation from monocular rgb via contrastive learning. In Proceedings of the IEEE/CVF international conference on computer vision, pages 11230--11239, 2021

  58. [66]

    D. Stashuk. Emg signal decomposition: how can it be accomplished and used? Journal of Electromyography and Kinesiology, 11 0 (3): 0 151--173, 2001

  59. [67]

    J. S. Supan c i c , G. Rogez, Y. Yang, J. Shotton, and D. Ramanan. Depth-based hand pose estimation: methods, data, and challenges. International Journal of Computer Vision, 126: 0 1180--1198, 2018

  60. [68]

    Tashakori, Z

    A. Tashakori, Z. Jiang, A. Servati, S. Soltanian, H. Narayana, K. Le, C. Nakayama, C.-l. Yang, Z. J. Wang, J. J. Eng, et al. Capturing complex hand movements and object interactions using machine learning-powered stretchable smart textile gloves. Nature Machine Intelligence, 6...

  61. [69]

    Truong, S

    H. Truong, S. Zhang, U. Muncuk, P. Nguyen, N. Bui, A. Nguyen, Q. Lv, K. Chowdhury, T. Dinh, and T. Vu. Capband: Battery-free successive capacitance sensing wristband for hand gesture recognition. In Proceedings of the 16th ACM Conference on Embedded Networked Sensor Systems, p...

  62. [70]

    C. Wan, T. Probst, L. V. Gool, and A. Yao. Self-supervised 3d hand pose estimation through training by fitting. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 10853--10862, 2019

  63. [71]

    J. Wang, F. Mueller, F. Bernard, S. Sorli, O. Sotnychenko, N. Qian, M. A. Otaduy, D. Casas, and C. Theobalt. Rgb2hands: real-time tracking of 3d hand interactions from monocular rgb video. ACM Transactions on Graphics (ToG), 39 0 (6): 0 1--16, 2020

  64. [72]

    Z. Yang, S. Yan, B.-J. F. van Beijnum, B. Li, and P. H. Veltink. Hand-finger pose estimation using inertial sensors, magnetic sensors and a magnet. IEEE sensors journal, 21 0 (16): 0 18115--18122, 2021

  65. [73]

    Z. Yu, J. S. Yoon, I. K. Lee, P. Venkatesh, J. Park, J. Yu, and H. S. Park. Humbi: A large multiview dataset of human body expressions. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 2990--3000, 2020

  66. [74]

    S. Yuan, Q. Ye, B. Stenger, S. Jain, and T.-K. Kim. Bighand2. 2m benchmark: Hand pose dataset and state of the art analysis. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 4866--4874, 2017

  67. [75]

    Y. Yuan, J. Song, U. Iqbal, A. Vahdat, and J. Kautz. Physdiff: Physics-guided human motion diffusion model. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 16010--16021, 2023

  68. [76]

    Zhang and C

    Y. Zhang and C. Harrison. Tomo: Wearable, low-cost electrical impedance tomography for hand gesture recognition. In Proceedings of the 28th Annual ACM Symposium on User Interface Software & Technology, pages 167--173, 2015

  69. [77]

    Zimmermann and T

    C. Zimmermann and T. Brox. Learning to estimate 3d hand pose from single rgb images. In Proceedings of the IEEE international conference on computer vision, pages 4903--4911, 2017

  70. [78]

    Zimmermann, D

    C. Zimmermann, D. Ceylan, J. Yang, B. Russell, M. Argus, and T. Brox. Freihand: A dataset for markerless capture of hand pose and shape from single rgb images. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 813--822, 2019

Pith tools

Reviewed August 11, 2026 · model on record in the stance chip above.