Pith. sign in

REVIEW 2 major objections 5 minor 51 references

A single-stage biomechanical optimization recovers hand kinematics from multi-view video more robustly than the usual two-stage pipeline, especially with objects.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-12 06:59 UTC pith:BCJNWOHI

load-bearing objection Solid relative comparison showing end-to-end biomechanical fitting beats two-stage IK for unconstrained multi-joint hands with objects; useful methods paper whose main soft spot is the missing 3-D ground truth the authors already flag. the 2 major comments →

arxiv 2607.02796 v1 pith:BCJNWOHI submitted 2026-07-02 cs.CV

Biomechanics-aware Multi-view Markerless Motion Capture of Dexterous Hand Movements

classification cs.CV
keywords markerless motion capturehand kinematicsbiomechanical modelmulti-view pose estimationend-to-end optimizationobject interactioninverse kinematicsdexterous movement
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Markerless motion capture of the hand has lagged behind other body parts because small finger segments amplify keypoint noise and objects create occlusion. This paper shows that feeding multi-view 2-D keypoints into a gradient-based optimizer that keeps a full arm-and-hand biomechanical model inside the loop produces usable joint-angle trajectories for every recording of unconstrained, dexterous tasks. The conventional alternative—first triangulate to 3-D then solve constrained inverse kinematics—fails to converge on 15 % of the same videos and yields less plausible distal-joint angles and poorer keypoint agreement when objects are present. Because the method works without physical markers and across skin tones and hand sizes, it opens a practical route to biomechanically meaningful hand tracking for clinic, rehab, and motor-control studies.

Core claim

An end-to-end, gradient-based optimization that jointly scales a 28-DoF arm-and-hand model and fits its joint trajectories to multi-view 2-D keypoints recovers biomechanically plausible kinematics for all 121 recordings of posture and object-manipulation tasks, whereas the standard two-stage triangulation-plus-inverse-kinematics pipeline fails on 15 % of the same data and produces systematically larger distal flexion errors and lower percentage-of-correct-keypoints scores under occlusion.

What carries the argument

The differentiable biomechanical model inside a multi-layer-perceptron optimizer: joint angles and isotropic scaling factors are updated by back-propagating the weighted 2-D reprojection error of virtual markers across all cameras, so biomechanical constraints act as soft regularizers rather than a separate post-processing stage.

Load-bearing premise

Agreement with high-confidence 2-D keypoints and visual biomechanical plausibility are treated as sufficient evidence of accuracy even though no independent 3-D ground truth (markers or sensors) is available.

What would settle it

Collect simultaneous optical-marker or electromagnetic ground-truth joint angles on the same multi-view video set; if the end-to-end method’s distal-joint or occlusion errors remain larger than or equal to the two-stage method’s, the central superiority claim fails.

Watch this falsifier — get emailed when new claim-graph text bears on it.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 5 minor

Summary. The manuscript evaluates a single-stage, gradient-based end-to-end biomechanical reconstruction pipeline (MLP + differentiable MuJoCo model of a 28-DoF upper-limb/hand system) against a conventional two-stage pipeline (robust triangulation of the same 2-D keypoints followed by constrained OpenSim inverse kinematics) for multi-view markerless tracking of dexterous hand motion. Using an 8-camera setup, 121 recordings from 6 participants performing ASL postures and object-manipulation tasks are analyzed. The end-to-end method converges on every recording while the two-stage method fails on 15 %; on the remaining data it yields smaller distal-joint deviations from full extension during ASL letter B, lower absolute joint-angle differences that grow toward the fingertips, and higher percentage-of-correct-keypoints (PCK) scores that remain robust under object occlusion (significant method imes task interaction). The authors conclude that embedding biomechanical constraints inside the optimization loop better exploits the available 2-D keypoint information.

Significance. If the relative ranking holds, the work supplies a practical, fully automatic route to biomechanically constrained hand kinematics from multi-view video without manual post-processing or marker placement. That capability is directly relevant to clinical hand assessment, rehabilitation monitoring, and motor-control studies where object interaction and proximal-limb motion are essential. Strengths include the clean head-to-head design (identical video, keypoints, calibration, and biomechanical model), the use of linear mixed-effects models that properly handle repeated measures, and the explicit demonstration of non-convergence and occlusion robustness. The absence of independent 3-D ground truth limits absolute accuracy claims but does not erase the comparative evidence that is actually asserted.

major comments (2)
  1. Discussion (and Methods §II.D.3): the sole quantitative accuracy metric is PCK at a 10-pixel threshold against the same high-likelihood 2-D keypoints that the end-to-end optimizer is trained to reproject. While the two-stage pipeline uses identical keypoints, higher PCK is still partly by construction for the end-to-end method. The ASL-B extension analysis and visual overlays supply independent qualitative support, yet the manuscript would be stronger if it either (a) reported an external 3-D reference on a subset of trials or (b) more explicitly framed PCK as a consistency metric rather than an accuracy proxy. This does not reverse the relative ranking, but it is load-bearing for any claim that the kinematics are “more accurate.”
  2. Methods §II.D.2: locking the thorax pose for OpenSim IK is presented as a fair mitigation of kinematic redundancy. The paper should quantify how often and by how much the unlocked two-stage solutions diverge (or fail) relative to the locked ones, so readers can judge whether the 15 % non-convergence rate and the joint-angle differences are inflated by this design choice.
minor comments (5)
  1. Figure 3 caption and color scale: the dual-ring encoding of mean and mean+1 s.d. is dense; a simple table of per-joint means and s.d.s would improve readability.
  2. Equation (1) and surrounding text: the reprojection loss is described as “average estimation error … weighted by the confidence-level,” yet the displayed formula shows only the unweighted Euclidean distance. Clarify the weighting formula.
  3. Table 1: skin-tone labels are given without the reference scale (Fitzpatrick or other); a brief note would aid reproducibility.
  4. Throughout: “end-to-end” and “two-stage” are used consistently, but occasional switches to “single-stage” or “OpenCap-style” could be standardized.
  5. References [16]–[19] and the authors’ prior work [18] are appropriately cited; a short sentence distinguishing the present multi-task, multi-object evaluation from those earlier posture-only results would help readers place the contribution.

Circularity Check

1 steps flagged

Mild circularity confined to PCK: end-to-end directly optimizes the reprojection error that PCK thresholds, so higher PCK is partly by construction; convergence rate and ASL-B joint-angle plausibility remain independent.

specific steps
  1. fitted input called prediction [Methods D.1 (reprojection loss) + D.3 (PCK definition) + Results B / Fig. 6]
    "Reprojection loss was defined as the average estimation error in pixels over all timesteps, cameras, and keypoints weighted by the confidence-level of the keypoints from the pose estimator. … δt,k,c=∥Πc xt,k-yt,k,c∥ (1) … PCK=Ntrue/Total imes100 (2) … the two-stage method’s PCK score was significantly lower for tasks involving objects, and the statistical interaction effect (method imes task paradigm) was significant (p<0.001)"

    The end-to-end optimizer is trained for 40 000 iterations to minimize precisely the pixel-wise Euclidean reprojection error that PCK later thresholds at 10 px. Consequently the higher PCK reported for end-to-end is statistically forced by the training objective itself; the metric does not supply independent confirmation that the method ‘utilizes the available 2-D digital keypoint information’ more effectively.

full rationale

The paper is an empirical head-to-head comparison of two reconstruction pipelines that share identical 2-D keypoints, cameras, calibration, and biomechanical model. Its strongest claims (100 % vs 85 % convergence; more plausible distal kinematics on ASL letter B; significant method imes object interaction on PCK) rest on three distinct metrics. Only the PCK metric is partially circular: the end-to-end MLP is trained by gradient descent on a weighted Euclidean reprojection loss (Eq. 1) whose thresholded version is exactly PCK (Eq. 2). Reporting higher PCK for the method that explicitly minimizes that loss is therefore expected by construction and does not constitute independent evidence of superior utilization of the 2-D information. The two-stage baseline, however, uses the same keypoints yet still under-performs, and the non-convergence rate plus the ASL-B extension analysis (joint angles near 0°) are external to the loss and therefore non-circular. Self-citations to the authors’ prior differentiable-biomechanics papers merely supply the method being evaluated; they do not underwrite a uniqueness claim or force the ranking. No self-definitional loop, uniqueness import, or ansatz smuggling appears. Overall circularity is therefore limited and non-load-bearing for the central relative-performance conclusion.

Axiom & Free-Parameter Ledger

4 free parameters · 3 axioms · 0 invented entities

The central claim rests on standard computer-vision and biomechanics tooling plus a handful of hand-chosen thresholds and architectural choices. No new physical entities are postulated; free parameters are the usual optimization hyper-parameters and detection cut-offs. Domain assumptions about the OpenSim kinematic model and the reliability of commercial pose estimators are inherited from prior literature.

free parameters (4)
  • keypoint likelihood exclusion threshold = 0.25 / 0.5
    Any keypoint with RTMPose confidence < 0.25 is discarded for both methods; PCK further restricts to > 0.5. These cut-offs are chosen by the authors and affect which data enter the comparison.
  • PCK pixel threshold = 10 pixels
    A reprojected marker is counted correct if it lies within 10 pixels of the 2-D detection. The numeric value is arbitrary and directly controls the reported PCK scores.
  • MLP architecture and training schedule = see Methods D.1.a
    Seven-layer MLP sizes (128…4096), 40 000 iterations, AdamW betas, learning-rate decay from 1e-3 to 1e-6, and virtual-marker offset lambda = 1e3 are free design choices that determine the final joint trajectories.
  • model scaling factors and virtual-marker offsets = participant-specific
    Six isotropic segment scales plus 28 marker offsets are optimized per participant via the bilevel AddBiomechanics procedure; they are fitted quantities that absorb residual calibration and anthropometric error.
axioms (3)
  • domain assumption The 21-DoF hand + 7-DoF upper-extremity OpenSim kinematic model (including coupled ring/little CMC flexion and scapulohumeral rhythm) correctly encodes the anatomical constraints of the healthy adult hand and arm.
    Both pipelines use the identical model; any systematic bias in joint limits or segment lengths is shared and therefore invisible to the relative comparison.
  • domain assumption RTMPose and MeTRAbs-ACAE produce sufficiently accurate 2-D keypoints that residual localization error is dominated by triangulation and IK rather than by the detectors themselves.
    All quantitative metrics ultimately rest on these detections; the paper does not independently validate detector accuracy on the collected data.
  • ad hoc to paper Locking the thorax pose for the two-stage OpenSim IK is a fair mitigation of kinematic redundancy and does not systematically disadvantage that method.
    The authors introduced this lock after observing convergence failures; it is not part of the standard OpenCap pipeline they otherwise follow.

pith-pipeline@v1.1.0-grok45 · 18186 in / 3190 out tokens · 35779 ms · 2026-07-12T06:59:13.956257+00:00 · methodology

0 comments
read the original abstract

Markerless motion capture (MMC) techniques have been widely beneficial in biomechanical analysis of human movement; however, application to complex motions of the hand lags other musculoskeletal systems. The primary goal of this study was to evaluate the performance of a biomechanical reconstruction method that implements a gradient-based optimization approach with a biomechanical model in the loop for tracking dexterous, unconstrained hand movements using MMC. Using a custom, 8-camera setup, we acquired 121 video recordings from 6 participants performing 11 different tasks that spanned 6 hand postures, 5 object manipulation tasks, and involved motion of the proximal upper limb joints. Performance of the proposed MMC pipeline was directly compared to a more commonly adopted two-stage reconstruction method that first triangulates 2D keypoints from computer vision pose estimation algorithms to 3D and then enforces biomechanical constraints by solving a constrained inverse kinematics problem. Relative performance was assessed qualitatively by visual inspection and quantitatively using a computer vision metric. Our method generated solutions for all 121 video recordings; the two-stage method did not converge for 15% of the recordings. Across the remaining videos, our method produced more biomechanically plausible hand kinematics than the two-stage method and was more robust to occlusion effects during tasks that involved objects. The relative robustness of the end-to-end method suggests that it is more effective in utilizing the available 2D digital keypoint information. Automatic and biomechanically meaningful tracking of hand kinematics during dexterous movements has the potential to support clinical evaluation, rehabilitation monitoring, and studies of human motor control.

Figures

Figures reproduced from arXiv: 2607.02796 by Anton Sobinov, J.D. Peiffer, Kunal Shah, Lee E. Miller, Pouyan Firouzabadi, R. James Cotton, Wendy M. Murray.

Figure 1
Figure 1. Figure 1: Experimental setup with 8 cameras surrounding the participant replicating the hand pose (the ASL letter B) shown on the video screen [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

51 extracted references · 21 canonical work pages · 1 internal anchor

  1. [1]

    Applications and limitations of current markerless motion capture methods for clinical gait biomechanics,

    L. Wade, L. Needham, P. McGuigan, and J. Bilzon, “Applications and limitations of current markerless motion capture methods for clinical gait biomechanics,” PeerJ, vol. 10, p. e12995, Feb. 2022, doi: 10.7717/peerj.12995

  2. [2]

    The Devil is in the Details: Delving into Unbiased Data Processing for Human Pose Estimation,

    J. Huang, Z. Zhu, F. Guo, and G. Huang, “ The Devil is in the Details: Delving into Unbiased Data Processing for Human Pose Estimation,” Proceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition , pp. 5699 –5708, Nov. 2019, doi: 10.1109/CVPR42600.2020.00574

  3. [3]

    RTMPose: Real -Time Multi-Person Pose Estimation based on MMPose,

    T. Jiang et al., “RTMPose: Real -Time Multi-Person Pose Estimation based on MMPose,” Jul. 02, 2023, arXiv: arXiv:2303.07399. doi: 10.48550/arXiv.2303.07399

  4. [4]

    Deep High -Resolution Representation Learning for Visual Recognition,

    J. Wang et al. , “Deep High -Resolution Representation Learning for Visual Recognition,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 43, no. 10, pp. 3349 –3364, Aug. 2019, doi: 10.1109/TPAMI.2020.2983686

  5. [5]

    Simple Baselines for Human Pose Estimation and Tracking,

    B. Xiao, H. Wu, and Y. Wei, “Simple Baselines for Human Pose Estimation and Tracking,” in Computer Vision – ECCV 2018 , vol. 11210, V. Ferrari, M. Hebert, C. Sminchisescu, and Y. Weiss, Eds., in Lecture Notes in Computer Science, vol. 11210. , Cham: Springer International Publishing, 2018, pp. 472 –487. doi: 10.1007/978-3-030- 01231-1_29

  6. [6]

    Distribution -Aware Coordinate Representation for Human Pose Estimation,

    F. Zhang, X. Zhu, H. Dai, M. Ye, and C. Zhu, “Distribution -Aware Coordinate Representation for Human Pose Estimation,” Proceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition , pp. 7091 –7100, Oct. 2019, doi: 10.1109/CVPR42600.2020.00712

  7. [7]

    Deep Learning-Based Human Pose Estimation: A Survey,

    C. E. Zheng et al., “Deep Learning-Based Human Pose Estimation: A Survey,” Tsinghua Science and Technology , vol. 24, no. 6, pp. 663 – 676, Dec. 2020, doi: 10.26599/TST.2018.9010100

  8. [8]

    Evaluation of drop vertical jump kinematics and kinetics using 3D markerless motion capture in a large cohort,

    T. Templin et al., “Evaluation of drop vertical jump kinematics and kinetics using 3D markerless motion capture in a large cohort,” Front. Bioeng. Biotechnol. , vol. 12, Oct. 2024, doi: 10.3389/fbioe.2024.1426677

  9. [9]

    Variability of in-game markerless and laboratory marker -based baseball pitching biomechanics,

    B. G. Lerch, G. S. Fleisig, J. S. Slowik, and G. D. Oliver, “Variability of in-game markerless and laboratory marker -based baseball pitching biomechanics,” Journal of Biomechanics , vol. 188, p. 112775, Jul. 2025, doi: 10.1016/j.jbiomech.2025.112775. 8

  10. [10]

    Multiple View Geometry in Computer Vision (Cited by: 11343),

    R. Hartley and A. Zisserman, “Multiple View Geometry in Computer Vision (Cited by: 11343),” Cambridge University Press, vol. 2, no. 2, p. 672, 2004

  11. [11]

    OpenCap: Human movement dynamics from smartphone videos,

    S. D. Uhlrich et al. , “OpenCap: Human movement dynamics from smartphone videos,” PLOS Computational Biology, vol. 19, no. 10, p. e1011462, Oct. 2023, doi: 10.1371/journal.pcbi.1011462

  12. [12]

    Optimisation and Comparison of Markerless and Marker -Based Motion Capture Methods for Hand and Finger Movement Analysis,

    V. Maggioni, C. Azevedo -Coste, S. Durand, and F. Bailly, “Optimisation and Comparison of Markerless and Marker -Based Motion Capture Methods for Hand and Finger Movement Analysis,” Sensors 2025, Vol. 25, Page 1079 , vol. 25, no. 4, p. 1079, Feb. 2025, doi: 10.3390/S25041079

  13. [13]

    Hand tracking for clinical applications: validation of the Google MediaPipe Hand (GMH) and the depth -enhanced GMH -D frameworks,

    G. Amprimo, G. Masi, G. Pettiti, G. Olmo, L. Priano, and C. Ferraris, “Hand tracking for clinical applications: validation of the Google MediaPipe Hand (GMH) and the depth -enhanced GMH -D frameworks,” Aug. 2023, Accessed: May 30, 2024. [Online]. Available: https://arxiv.org/abs/2308.01088v1

  14. [14]

    Validation of two -dimensional video -based inference of finger kinematics with pose estimation,

    L. Gionfrida, W. M. R. Rusli, A. A. Bharath, and A. E. Kedgley, “Validation of two -dimensional video -based inference of finger kinematics with pose estimation,” PLOS ONE , vol. 17, no. 11, p. e0276799, Nov. 2022, doi: 10.1371/journal.pone.0276799

  15. [15]

    Multi -view 3D Markerless Hand Motion Capture System with Keypoint Triangulation,

    G. M. Lim, P. Jatesiktat, and W. T. Ang, “Multi -view 3D Markerless Hand Motion Capture System with Keypoint Triangulation,” in 2024 17th International Convention on Rehabilitation Engineering and Assistive Technology (i-CREATe), Aug. 2024, pp. 1–4. doi: 10.1109/i- CREATe62067.2024.10776171

  16. [16]

    Differentiable Biomechanics for Markerless Motion Capture in Upper Limb Stroke Rehabilitation: A Comparison With Optical Motion Capture,

    T. Unger et al., “Differentiable Biomechanics for Markerless Motion Capture in Upper Limb Stroke Rehabilitation: A Comparison With Optical Motion Capture,” IEEE Trans. Med. Robot. Bionics , pp. 1 –1, 2025, doi: 10.1109/TMRB.2025.3605962

  17. [17]

    Differentiable Biomechanics Unlocks Opportunities for Markerless Motion Capture,

    R. J. Cotton, “Differentiable Biomechanics Unlocks Opportunities for Markerless Motion Capture,” in 2025 International Conference On Rehabilitation Robotics (ICORR) , Chicago, IL, USA: IEEE, May 2025, pp. 44–51. doi: 10.1109/ICORR66766.2025.11063174

  18. [18]

    Biomechanical Arm and Hand Tracking with Multiview Markerless Motion Capture,

    P. Firouzabadi et al., “Biomechanical Arm and Hand Tracking with Multiview Markerless Motion Capture,” Proceedings of the IEEE RAS and EMBS International Conference on Biomedical Robotics and Biomechatronics, pp. 1641 –1648, 2024, doi: 10.1109/BIOROB60516.2024.10719940

  19. [19]

    ATHENA: Automatically Tracking Hands Expertly with No Annotations,

    D. M. Mulla, M. Costantino, E. Freud, and J. A. Michaels, “ATHENA: Automatically Tracking Hands Expertly with No Annotations,” Aug. 15, 2025, bioRxiv. doi: 10.1101/2025.08.12.669753

  20. [20]

    Articulated Human Detection with Flexible Mixtures of Parts,

    Y. Yang and D. Ramanan, “Articulated Human Detection with Flexible Mixtures of Parts,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 35, no. 12, pp. 2878 –2890, Dec. 2013, doi: 10.1109/TPAMI.2012.261

  21. [21]

    Vision-based hand pose estimation: A review,

    A. Erol, G. Bebis, M. Nicolescu, R. D. Boyle, and X. Twombly, “Vision-based hand pose estimation: A review,” Computer Vision and Image Understanding, vol. 108, no. 1 –2, pp. 52 –73, Oct. 2007, doi: 10.1016/j.cviu.2006.10.012

  22. [22]

    Research Techniques Made Simple: Cutaneous Colorimetry: A Reliable Technique for Objective Skin Color Measurement,

    B. C. K. Ly, E. B. Dyer, J. L. Feig, A. L. Chien, and S. Del Bino, “Research Techniques Made Simple: Cutaneous Colorimetry: A Reliable Technique for Objective Skin Color Measurement,” Journal of Investigative Dermatology, vol. 140, no. 1, pp. 3 -12.e1, Jan. 2020, doi: 10.1016/j.jid.2019.11.003

  23. [23]

    Markerless Motion Capture and Biomechanical Analysis Pipeline,

    R. J. Cotton et al. , “Markerless Motion Capture and Biomechanical Analysis Pipeline,” Mar. 2023, Accessed: Jan. 31, 2024. [Online]. Available: http://arxiv.org/abs/2303.10654

  24. [24]

    Charuco Board-Based Omnidirectional Camera Calibration Method,

    G. An, S. Lee, M. -W. Seo, K. Yun, W. -S. Cheong, and S. -J. Kang, “Charuco Board-Based Omnidirectional Camera Calibration Method,” Electronics, vol. 7, no. 12, p. 421, Dec. 2018, doi: 10.3390/electronics7120421

  25. [25]

    Anipose: A toolkit for robust markerless 3D pose estimation,

    P. Karashchuk et al. , “Anipose: A toolkit for robust markerless 3D pose estimation,” Cell Reports, vol. 36, no. 13, p. 109730, Sep. 2021, doi: 10.1016/J.CELREP.2021.109730

  26. [26]

    PosePipe: Open -Source Human Pose Estimation Pipeline for Clinical Research,

    R. J. Cotton, “PosePipe: Open -Source Human Pose Estimation Pipeline for Clinical Research,” arXiv.org. Accessed: Feb. 04, 2024. [Online]. Available: https://arxiv.org/abs/2203.08792v1

  27. [27]

    MoVi: A large multi -purpose human motion and video dataset,

    S. Ghorbani et al., “MoVi: A large multi -purpose human motion and video dataset,” PLOS ONE , vol. 16, no. 6, p. e0253157, Jun. 2021, doi: 10.1371/journal.pone.0253157

  28. [28]

    Learning 3D Human Pose Estimation from Dozens of Datasets using a Geometry -Aware Autoencoder to Bridge Between Skeleton Formats,

    I. Sarandi, A. Hermans, and B. Leibe, “Learning 3D Human Pose Estimation from Dozens of Datasets using a Geometry -Aware Autoencoder to Bridge Between Skeleton Formats,” Proceedings - 2023 IEEE Winter Conference on Applications of Computer Vision, WACV 2023 , pp. 2955 –2965, Dec. 2022, doi: 10.1109/WACV56688.2023.00297

  29. [29]

    OpenMMLab Pose Estimation Toolbox and Benchmark

    MMPose Contributers, “OpenMMLab Pose Estimation Toolbox and Benchmark.” Accessed: Apr. 13, 2025. [Online]. Available: https://github.com/open-mmlab/mmpose

  30. [30]

    A model of the upper extremity for simulating musculoskeletal surgery and analyzing neuromuscular control,

    K. R. S. Holzbaur, W. M. Murray, and S. L. Delp, “ A model of the upper extremity for simulating musculoskeletal surgery and analyzing neuromuscular control,” Annals of Biomedical Engineering , vol. 33, no. 6, pp. 829–840, Jun. 2005, doi: 10.1007/s10439-005-3320-7

  31. [31]

    A Musculoskeletal Model of the Hand and Wrist Capable of Simulating Functional Tasks,

    D. C. McFarland, B. I. Binder -Markey, J. A. Nichols, S. J. Wohlman, M. de Bruin, and W. M. Murray, “A Musculoskeletal Model of the Hand and Wrist Capable of Simulating Functional Tasks,” bioRxiv, p. 2021.12.28.474357, Dec. 2021, doi: 10.1101/2021.12.28.474357

  32. [32]

    OpenSim: Simulating musculoskeletal dynamics and neuromuscular control to study human and animal movement,

    A. Seth et al., “OpenSim: Simulating musculoskeletal dynamics and neuromuscular control to study human and animal movement,” PLOS Computational Biology, vol. 14, no. 7, p. e1006223, Jul. 2018, doi: 10.1371/JOURNAL.PCBI.1006223

  33. [33]

    AddBiomechanics: Automating model scaling, inverse kinematics, and inverse dynamics from human motion data through sequential optimization,

    K. Werling et al. , “AddBiomechanics: Automating model scaling, inverse kinematics, and inverse dynamics from human motion data through sequential optimization,” bioRxiv, p. 2023.06.15.545116, Sep. 2023, doi: 10.1101/2023.06.15.545116

  34. [34]

    Compiling machine learning programs via high-level tracing

    R. Frostig, G. Brain, M. J. Johnson, and C. Leary Google, “Compiling machine learning programs via high-level tracing”, Accessed: Jun. 15,

  35. [35]

    Available: https://github.com/jrevels/Cassette.jl

    [Online]. Available: https://github.com/jrevels/Cassette.jl

  36. [36]

    Equinox: neural networks in JAX via callable PyTrees and filtered transformations,

    P. Kidger and C. Garcia, “Equinox: neural networks in JAX via callable PyTrees and filtered transformations,” Oct. 30, 2021, arXiv: arXiv:2111.00254. doi: 10.48550/arXiv.2111.00254

  37. [37]

    Brax -- A Differentiable Physics Engine for Large Scale Rigid Body Simulation,

    C. D. Freeman, E. Frey, A. Raichuk, S. Girgin, I. Mordatch, and O. Bachem, “Brax -- A Differentiable Physics Engine for Large Scale Rigid Body Simulation,” Jun. 2021, Accessed: Feb. 02, 2024. [Online]. Available: https://arxiv.org/abs/2106.13281v1

  38. [38]

    MuJoCo: A physics engine for model-based control,

    E. Todorov, T. Erez, and Y. Tassa, “MuJoCo: A physics engine for model-based control,” IEEE International Conference on Intelligent Robots and Systems , pp. 5026 –5033, 2012, doi: 10.1109/IROS.2012.6386109

  39. [39]

    Converting Biomechanical Models from OpenSim to MuJoCo,

    A. Ikkala and P. Hämäläinen, “Converting Biomechanical Models from OpenSim to MuJoCo,” Biosystems and Biorobotics, vol. 28, pp. 277–281, Jun. 2020, doi: 10.1007/978-3-030-70316-5_45

  40. [40]

    Decoupled Weight Decay Regularization,

    I. Loshchilov and F. Hutter, “Decoupled Weight Decay Regularization,” 7th International Conference on Learning Representations, ICLR 2019 , Nov. 2017, Accessed: Feb. 04, 2024. [Online]. Available: https://arxiv.org/abs/1711.05101v3

  41. [41]

    On Triangulation as a Form of Self -Supervision for 3D Human Pose Estimation,

    S. K. Roy, L. Citraro, S. Honari, and P. Fua, “On Triangulation as a Form of Self -Supervision for 3D Human Pose Estimation,” in 2022 International Conference on 3D Vision (3DV) , Prague, Czech Republic: IEEE, Sep. 2022, pp. 1 –10. doi: 10.1109/3DV57658.2022.00068

  42. [42]

    FreiHAND: A Dataset for Markerless Capture of Hand Pose and Shape from Single RGB Images,

    C. Zimmermann, D. Ceylan, J. Yang, B. Russell, M. J. Argus, and T. Brox, “FreiHAND: A Dataset for Markerless Capture of Hand Pose and Shape from Single RGB Images,” Proceedings of the IEEE International Conference on Computer Vision , vol. 2019 -Octob, pp. 813–822, Sep. 2019, doi: 10.1109/ICCV.2019.00090

  43. [43]

    Whole-Body Human Pose Estimation in the Wild,

    S. Jin et al., “Whole-Body Human Pose Estimation in the Wild,” in Computer Vision – ECCV 2020, A. Vedaldi, H. Bischof, T. Brox, and J.-M. Frahm, Eds., in Lecture Notes in Computer Science. Cham: Springer International Publishing, 2020, pp. 196 –214. doi: 10.1007/978-3-030-58545-7_12

  44. [44]

    Mask -Pose Cascaded CNN for 2D Hand Pose Estimation From Single Color Image,

    Y. Wang, C. Peng, and Y. Liu, “Mask -Pose Cascaded CNN for 2D Hand Pose Estimation From Single Color Image,” IEEE Transactions on Circuits and Systems for Video Technology , vol. 29, no. 11, pp. 3258–3268, Nov. 2019, doi: 10.1109/TCSVT.2018.2879980

  45. [45]

    Learning to Estimate 3D Hand Pose from Single RGB Images,

    C. Zimmermann and T. Brox, “Learning to Estimate 3D Hand Pose from Single RGB Images,” Proceedings of the IEEE International Conference on Computer Vision , vol. 2017 -Octob, pp. 4913 –4921, May 2017, doi: 10.1109/ICCV.2017.525

  46. [46]

    AlphaPose: Whole -Body Regional Multi -Person Pose Estimation and Tracking in Real -Time,

    H. S. Fang et al. , “AlphaPose: Whole -Body Regional Multi -Person Pose Estimation and Tracking in Real -Time,” IEEE Transactions on Pattern Analysis and Machine Intelligence , vol. 45, no. 6, pp. 7157 – 7173, Nov. 2022, doi: 10.1109/TPAMI.2022.3222784

  47. [47]

    Conclusion or Illusion: Quantifying Uncertainty in Inverse Analyses From Marker -Based Motion Capture due to Errors in Marker Registration and Model Scaling,

    T. K. Uchida and A. Seth, “Conclusion or Illusion: Quantifying Uncertainty in Inverse Analyses From Marker -Based Motion Capture due to Errors in Marker Registration and Model Scaling,” Front. Bioeng. Biotechnol. , vol. 10, p. 874725, May 2022, doi: 10.3389/fbioe.2022.874725. 9

  48. [48]

    Effects of hip joint centre mislocation on gait kinematics of children with cerebral palsy calculated using patient -specific direct and inverse kinematic models,

    H. Kainz, C. P. Carty, S. Maine, H. P. J. Walsh, D. G. Lloyd, and L. Modenese, “Effects of hip joint centre mislocation on gait kinematics of children with cerebral palsy calculated using patient -specific direct and inverse kinematic models,” 2017, doi: 10.1016/j.gaitpost.2017.06.002

  49. [49]

    The development and evaluation of a fully automated markerless motion capture workflow,

    L. Needham et al. , “The development and evaluation of a fully automated markerless motion capture workflow,” Journal of Biomechanics, vol. 144, p. 111338, Nov. 2022, doi: 10.1016/J.JBIOMECH.2022.111338

  50. [50]

    SAM 3D Body: Robust Full -Body Human Mesh Recovery,

    X. Yang et al. , “SAM 3D Body: Robust Full -Body Human Mesh Recovery,” Feb. 17, 2026, arXiv: arXiv:2602.15989. doi: 10.48550/arXiv.2602.15989

  51. [51]

    Monocular Biomechanical Tracking of Fingers with Inverse Kinematics to Foundation Models

    R. J. Cotton, P. Firouzabadi, and W. Murray, “Monocular Biomechanical Tracking of Fingers with Inverse Kinematics to Foundation Models,” May 10, 2026, arXiv: arXiv:2605.09258. doi: 10.48550/arXiv.2605.09258