Pith. sign in

REVIEW 41 references

Markerless cameras in routine ARAT sessions can reconstruct upper-limb movement accurately enough to yield valid kinematic metrics that go beyond the ordinal score.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-30 17:42 UTC pith:YVFFOPRA

load-bearing objection Solid routine-clinic MMC validation for ARAT with a clean two-tier known-groups design; H1–H2 hold, H3 is honestly exploratory n=2 and should stay labeled that way.

arxiv 2607.23608 v1 pith:YVFFOPRA submitted 2026-07-26 cs.CV

Markerless Motion Capture in Routine Clinical Upper Limb Assessments: Validity and Insights Beyond Ordinal Scoring

classification cs.CV
keywords markerless motion captureAction Research Arm Testupper limb kinematicsneurorehabilitationconstruct validitystrokeordinal scoringclinical routine
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The Action Research Arm Test scores upper-limb function on a coarse four-point scale that is subjective, insensitive at the extremes, and silent about why a patient got a particular score. This paper asks whether ordinary RGB webcams and AI-based markerless motion capture, dropped into unmodified clinical ARAT sessions, can reconstruct joint motion well enough to produce continuous kinematic metrics that are both trustworthy and clinically informative. In 1,174 tasks from 20 neurological patients, reconstruction error stayed inside an acceptable clinical band and did not worsen with impairment, and the derived metrics discriminated completed from failed tasks strongly while discriminating good from normal completions more weakly—the pattern expected of a measure that tracks real function rather than the noisy scale. In two longitudinal cases, domain scores (range of motion, velocity, compensation) exposed opposite recovery profiles under equal ARAT gains and kept detecting change after the clinical score had hit its ceiling. If this holds, clinics could add objective, multi-dimensional upper-limb measurement without changing how assessments are run.

Core claim

AI-based markerless motion capture embedded in routine ARAT assessments yields biomechanical reconstructions that are accurate and stable across impairment levels, and kinematic metrics derived from them show the known-groups discrimination pattern of a construct-valid continuous measure of upper-limb function, while domain decomposition and post-ceiling tracking supply the specificity and sensitivity the ordinal ARAT lacks.

What carries the argument

Hierarchical compound scores built from ten trial-level kinematic metrics (peak and range values of shoulder, elbow, hand, and trunk signals), normalised to a Score-3 Q75 anchor and averaged into range-of-motion, velocity, compensation, and overall scores; construct validity is read from a two-tier known-groups design (strong Tier-1 completion discrimination, attenuated Tier-2 quality discrimination) rather than simple correlation with the ordinal scale.

Load-bearing premise

The hand-crafted metrics and their fixed domain labels fairly represent movement quality for each ARAT task group, so agreement with clinical scores can be read as true construct validity rather than lucky metric–task match.

What would settle it

A larger multi-session cohort in which overall kinematic compound scores systematically fail to track sub-ceiling ARAT change within one MCID on well-matched task groups, or in which reconstruction error rises sharply in the most impaired score groups, would overturn the validity claim.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Routine ARAT sessions can double as continuous kinematic monitoring without markers, extra setup time, or changes to clinical workflow.
  • Equal ARAT point gains can be decomposed into distinct recovery profiles (e.g. ROM-and-compensation versus velocity), guiding more specific therapy choices.
  • Improvement can still be quantified after patients hit the ARAT ceiling, extending the useful measurement window.
  • The same passive capture stream can feed future reference-based deviation indices and movement foundation models once able-bodied norms are collected at scale.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • Because validity collapsed exactly where domain assignment mismatched the task (pouring, grip without fingers, overhead positioning), task-specific or learned metric sets are likely required before clinic-wide deployment.
  • Replacing the Score-3 Q75 proxy with true able-bodied references would remove a circularity in the normal-performance anchor and tighten MCID estimation.
  • If single-camera reconstruction reaches comparable clinical accuracy, the barrier to in-room continuous monitoring falls from three webcams to ubiquitous devices.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Circularity Check

2 steps flagged

Mild scale-anchoring circularity: compound ‘normal’ is defined from ARAT Score-3 sessions, then longitudinal validity is scored as agreement with ARAT; core discrimination and post-ceiling claims remain independent.

specific steps
  1. self definitional [Methods, Compound Metrics; Figure 1; H3a criterion (Table 1)]
    "Lacking a separate able-bodied cohort, we approximated this reference with Score 3 (clinically normal) sessions. Because Score 3 carries the scoring subjectivity this study documents, we anchored the normal threshold at the 75th percentile (Q75) of Score 3 rather than the median... 100 % corresponds to the per-metric Q75 of Score 3 sessions and 0 % to the per-task cohort minimum... H3a ... Soverall agrees with clinical change (MAD) ... MAD(Soverall)< MCID"

    Soverall’s 100% pole is defined from the same ARAT Score-3 labels the paper criticizes as subjective. Longitudinal ‘construct validity’ is then operationalized as MAD between change in this Score-3-anchored compound and change in rescaled ARAT. Agreement of the overall compound with ARAT is therefore partly built into the shared normal endpoint of the two scales, not an external continuous gold standard. (Domain-level divergence and post-ceiling ΔS are not forced by this anchoring and remain independent content.)

  2. self citation load bearing [H1 threshold statement; Methods Data Quality; citation [23]]
    "cohort-level mean reprojection error within the 10–20 pixel (px) range previously established as acceptable for this reconstruction algorithm in clinical use, modestly above the 5–10 px laboratory standard [23]"

    The binary pass/fail band for ‘adequate’ reconstruction is taken from prior work by overlapping authors on the same algorithm, not from an external standard or independent error model in this study. The measured 12.5 px figure is new data, but declaring H1 supported depends on that self-set acceptability range. This is minor and not load-bearing for H2/H3 clinical claims.

full rationale

The paper is largely self-contained empirical validation, not a first-principles derivation. H1 (reprojection error) and H2 (known-groups AUCs on raw metrics) do not reduce to their inputs by construction: error is an internal geometric residual, and Tier-1/Tier-2 discrimination is tested on un-normalized kinematics with a pre-specified differential pattern rather than forced correlation. The only material circular load is representational: lacking an able-bodied cohort, compound scores set 100% to the Q75 (or Q25) of clinically labeled Score-3 sessions, then H3a treats MAD between those compounds and rescaled ARAT change as longitudinal construct validity. That anchors the ‘normal’ pole of the kinematic scale in the same ordinal system under critique, so overall concordance with ARAT is partly encouraged by scale construction—though not statistically fitted, and domain divergence plus post-ceiling change (H3b/c) are not forced. Self-citations ([17],[23]) supply the reconstruction method and the 10–20 px acceptability band; they are methodological benchmarks, not uniqueness theorems that compel the clinical conclusions. Proportionate score is therefore low-moderate (3), not a collapse of the central claims.

Axiom & Free-Parameter Ledger

7 free parameters · 6 axioms · 2 invented entities

The central clinical-validity claim rests on accepting prior reconstruction quality bands, a literature-derived metric taxonomy, Score-3 percentile anchors as “normal,” an adopted 15 pp MCID, and the premise that differential ARAT-group discrimination equals construct validity. No new physical entities; free choices are thresholds, anchors, and aggregation rules.

free parameters (7)
  • MCID threshold on compound scale = 15 pp
    15 percentage points adopted from consensus recommendations [4], not estimated from test–retest or patient-perceived change in this cohort; gates H3a/H3c pass/fail.
  • Normal-performance anchor percentiles = Q75/Q25 of Score 3; cohort min = 0%
    Q75 of Score-3 sessions (Q25 for inverted compensation metrics) defines 100% on the compound scale; cohort minimum defines 0%. Chosen conservatively because Score-3 is subjective.
  • Reprojection adequacy band = 10–20 px acceptable
    H1 adequacy uses prior clinical-use band 10–20 px (lab 5–10 px) from earlier work rather than a new error-to-metric sensitivity study in this dataset.
  • Outlier exclusion multiplier = IQR × 3.0 (130/1304 trials removed)
    Trials excluded if any metric outside IQR×3.0 within ARAT score group; removes 10% of trials and can shape group distributions.
  • Low-pass filter cutoff = 5 Hz
    Zero-phase 4th-order Butterworth at 5 Hz applied to all kinematic time series before metrics.
  • AUC excellence/chance thresholds for H2 = Tier1 ≥0.85 majority; Tier2 >0.5 and <Tier1
    Tier 1 pass uses conventional AUC≥0.85 in a majority of groups; Tier 2 requires >0.5 and <Tier 1, citing Hosmer et al. discrimination labels.
  • Unweighted domain and overall aggregation = equal (unweighted) means
    Variable scores averaged equally within ROM/Vel/Comp; domains averaged equally into Soverall—hand-chosen equal weights, not fit to outcomes.
axioms (6)
  • domain assumption Differentiable biomechanics MMC reconstruction quality previously validated against optical motion capture is adequate substrate for clinical ARAT metrics when mean reprojection error is ~10–20 px.
    H1 imports acceptability bands and algorithm trust from Cotton/Unger prior work [17,23] rather than re-validating joint angles against markers in this clinic cohort.
  • domain assumption A continuous measure of upper-limb function should discriminate strongly at robust clinical boundaries (0/1 vs 2/3) and more weakly at subjective boundaries (2 vs 3).
    Core construct-validity logic for H2; turns circularity with ARAT into a predicted attenuation pattern (Introduction; Methods H2).
  • ad hoc to paper Trial-level peaks/ROMs of end-effector velocity, elbow/shoulder angles and rates, trunk displacement, and shoulder abduction sufficiently represent ARAT movement quality without phase segmentation or smoothness metrics.
    No ARAT-specific evidence-based metric set exists; authors assemble a general-purpose set from drinking-task and UL literature and assign domains once for the whole battery (Methods: Metric Selection).
  • ad hoc to paper Score-3 sessions can proxy able-bodied normal motor performance for normalization when a separate control cohort is unavailable.
    Explicitly stated limitation used to build all compound scores (Methods: Compound Metrics).
  • domain assumption Pooling tasks that share gross movement patterns into seven analysis groups and averaging yields clinically meaningful continuous scores on [0,3].
    Standard ARAT subgroup practice plus authors’ averaging step (Table 2); underpins all per-group AUCs and longitudinal MAD.
  • standard math Kruskal–Wallis η²_H < 0.06 implies reconstruction quality is robust across impairment for metric extraction purposes.
    Conventional small-effect cutoff applied as H1 robustness criterion (Table 1).
invented entities (2)
  • Q75-anchored hierarchical compound scores (SROM, SVel, SComp, Soverall) no independent evidence
    purpose: Collapse ten raw metrics into clinically interpretable domain and overall percentages for validity and longitudinal claims.
    Aggregation hierarchy and anchors are paper-defined; not a standard ARAT instrument. Independent evidence is only internal consistency with ARAT groups/cases, not external able-bodied or multi-site norms.
  • Two-tier known-groups construct-validity test against ARAT ordinal boundaries independent evidence
    purpose: Claim construct validity without a continuous clinical gold standard by expecting Tier1≫Tier2 discrimination.
    Known-groups validity is standard psychometrics; the specific tiering and pass rules are operationalized here for MMC-ARAT.

pith-pipeline@v1.2.0-grok45-kimik3 · 28329 in / 4323 out tokens · 83231 ms · 2026-07-30T17:42:20.461195+00:00 · methodology

0 comments
read the original abstract

The Action Research Arm Test (ARAT) is a widely-used upper limb outcome measure in neurorehabilitation, but its ordinal scoring is subjective and suffers from limited sensitivity and specificity. We evaluated whether artificial-intelligence (AI)-based markerless motion capture (MMC), embedded into ARAT assessments during clinical routine, accurately reconstructs upper limb movement and yields valid, objective kinematic metrics carrying clinically meaningful information beyond the ordinal score. Across 47 sessions from 20 mixed-neurological patients (1,174 ARAT tasks), biomechanical reconstruction was accurate and robust across impairment levels, and kinematic metrics showed the discrimination pattern expected of a construct-valid measure. In longitudinal case studies, the metrics added the specificity and sensitivity the ordinal score lacks: a domain decomposition exposed patient-specific recovery profiles underlying equal ARAT gains (specificity), and kinematic improvement continued to be detected after the ARAT had saturated (sensitivity). MMC in clinical routine can thus provide valid, objective, sensitive, and specific kinematic measurement complementing ordinal scoring.

Figures

Figures reproduced from arXiv: 2607.23608 by Andreas R. Luft, Chris Easthope Awai, Olivier Lambercy, R. James Cotton, Roger Gassert, Tim Unger.

Figure 1
Figure 1. Figure 1: Hierarchical construction of compound scores from kinematic metrics. Reading right to left: seven kinematic time series (right) are summarised by ten raw metrics (max or range of motion, ROM). Each raw metric is normalised to a Q75-anchored percentage scale on which 100% corresponds to the per-metric 75th percentile (Q75) of Score 3 (clinically normal) cohort sessions and 0% to the per-task cohort minimum,… view at source ↗
Figure 2
Figure 2. Figure 2: Mean per-trial reprojection error (pixels) by clinical ARAT score (0–3) and pooled across all 1304 trials (right). Largely overlapping distributions confirm that reconstruction quality does not systematically depend on patient impairment level (η 2 H=0.0053<0.06; [PITH_FULL_IMAGE:figures/full_fig_p008_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Distribution of kinematic metrics across clinical ARAT scores (0–3) for all trials in the Grasp analysis group. Individual trials shown as jittered points. Outliers removed (IQR × 3.0). 0.96) and Hand to Mouth (0.97) compensation nearly matched ROM, and for Hand on Top of Head it was the strongest single domain (0.87 vs. ROM 0.84, velocity 0.86), indicating that compensatory trunk and shoulder recruitment … view at source ↗
Figure 4
Figure 4. Figure 4: Longitudinal kinematic monitoring of P10’s affected side (ischaemic stroke) across four ARAT sessions, Grasp analysis group. Panel grid: rows are the three kinematic domains (ROM, Compensation, Velocity), each with its raw metrics, a domain compound score (SROM, SComp, SVel), and at far right the overall compound (Soverall). Compound panels are shown on the normalised 0–100 % scale (100 % = per-metric Q75 … view at source ↗
Figure 5
Figure 5. Figure 5: Longitudinal kinematic monitoring of P20’s affected side (stroke, right hemisphere) across nine ARAT sessions over approximately four months, Grasp analysis group. Panel grid, scale, reference bands, and the overlaid mean clinical ARAT score as in [PITH_FULL_IMAGE:figures/full_fig_p012_5.png] view at source ↗
Figure 6
Figure 6. Figure 6: Ad-hoc camera setup around the ARAT assessment table [PITH_FULL_IMAGE:figures/full_fig_p019_6.png] view at source ↗
Figure 7
Figure 7. Figure 7: Sessions per participant 19/23 [PITH_FULL_IMAGE:figures/full_fig_p019_7.png] view at source ↗
Figure 8
Figure 8. Figure 8: Distribution of kinematic metrics across clinical ARAT scores (0–3) for pick-and-place tasks per analysis group. Individual trials shown as jittered points. Outliers removed (IQR × 3.0). 22/23 [PITH_FULL_IMAGE:figures/full_fig_p022_8.png] view at source ↗
Figure 9
Figure 9. Figure 9: Distribution of kinematic metrics across clinical ARAT scores (0–3) for gross movement tasks per analysis group. Individual trials shown as jittered points. Outliers removed (IQR × 3.0). 23/23 [PITH_FULL_IMAGE:figures/full_fig_p023_9.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

41 extracted references · 6 canonical work pages

  1. [1]

    Global, regional, and national burden of disorders affecting the nervous system, 1990-2021: a systematic analysis for the Global Burden of Disease Study 2021,

    GBD 2021 Nervous System Disorders Collaborators, “Global, regional, and national burden of disorders affecting the nervous system, 1990-2021: a systematic analysis for the Global Burden of Disease Study 2021,”The Lancet. Neurology, vol. 23, no. 4, pp. 344–381, Apr. 2024

  2. [2]

    World Stroke Organization: Global Stroke Fact Sheet 2025,

    V . L. Feigin, M. Brainin, B. Norrving, S. O. Martins, J. Pandian, P. Lindsay, M. F Grupper, and I. Rautalin, “World Stroke Organization: Global Stroke Fact Sheet 2025,”International Journal of Stroke, vol. 20, no. 2, pp. 132–144, Feb. 2025. [Online]. Available: https://doi.org/10.1177/17474930241308142

  3. [3]

    Recovery of upper extremity function in stroke patients: The Copenhagen stroke study,

    H. Nakayama, H. S. Jørgensen, H. O. Raaschou, and T. S. Olsen, “Recovery of upper extremity function in stroke patients: The Copenhagen stroke study,”Archives of Physical Medicine and Rehabilitation, vol. 75, no. 4, 1994

  4. [4]

    Standardized measurement of quality of upper limb movement after stroke: Consensus-based core recommendations from the Second Stroke Recovery and Rehabilitation Roundtable,

    G. Kwakkel, E. Van Wegen, J. Burridge, C. Winstein, L. van Dokkum, M. Alt Murphy, M. Levin, and J. Krakauer, “Standardized measurement of quality of upper limb movement after stroke: Consensus-based core recommendations from the Second Stroke Recovery and Rehabilitation Roundtable,”International Journal of Stroke, vol. 14, no. 8, pp. 783–791, Oct. 2019. [...

  5. [5]

    Standardized measurement of sensorimotor recovery in stroke trials: Consensus-based core recommendations from the Stroke Recovery and Rehabilitation Roundtable,

    G. Kwakkel, N. A. Lannin, K. Borschmann, C. English, M. Ali, L. Churilov, G. Saposnik, C. Winstein, E. E. van Wegen, S. L. Wolf, J. W. Krakauer, and J. Bernhardt, “Standardized measurement of sensorimotor recovery in stroke trials: Consensus-based core recommendations from the Stroke Recovery and Rehabilitation Roundtable,”International Journal of Stroke,...

  6. [6]

    Action Research Arm Test (ARAT),

    J. H. van der Lee, “Action Research Arm Test (ARAT),” inOccupational Therapy Assessments for Older Adults, ser. Occupational Therapy Assessments for Older Adults: 100 Instruments for Measuring Occupational Performance. Taylor and Francis, Jan. 2024, pp. 213–214. [Online]. Available: https://www.scopus.com/pages/publications/85199034662

  7. [7]

    Do Activity Level Outcome Measures Commonly Used in Neurological Practice Assess Upper-Limb Movement Quality?

    M. Demers and M. F. Levin, “Do Activity Level Outcome Measures Commonly Used in Neurological Practice Assess Upper-Limb Movement Quality?”Neurorehabilitation and Neural Repair, vol. 31, no. 7, pp. 623–637, Jul. 2017

  8. [8]

    A systematic review of the psychometric properties of the Action Research Arm Test in neurorehabilitation,

    S. Pike, N. A. Lannin, K. Wales, and A. Cusick, “A systematic review of the psychometric properties of the Action Research Arm Test in neurorehabilitation,”Australian Occupational Therapy Journal, vol. 65, no. 5, pp. 449–471, 2018, _eprint: https://onlinelibrary.wiley.com/doi/pdf/10.1111/1440-1630.12527. [Online]. Available: https://onlinelibrary.wiley.co...

  9. [9]

    Intra-rater and inter-rater reliability at the item level of the Action Research Arm Test for patients with stroke,

    A. Nordin, M. Alt Murphy, and A. Danielsson, “Intra-rater and inter-rater reliability at the item level of the Action Research Arm Test for patients with stroke,”Journal of Rehabilitation Medicine, vol. 46, no. 8, pp. 738–745, Sep. 2014

  10. [10]

    G. B. Prange-Lasonder, M. A. Murphy, I. Lamers, A. M. Hughes, J. H. Buurke, P. Feys, T. Keller, V . Klamroth-Marganska, I. M. Tarkka, A. Timmermans, and J. H. Burridge, “European evidence-based recommendations for clinical assessment of upper limb in neurorehabilitation (CAULIN): data synthesis from systematic reviews, clinical practice guidelines and exp...

  11. [11]

    A Standardized Approach to Performing the Action Research Arm Test,

    N. Yozbatiran, L. Der-Yeghiaian, and S. C. Cramer, “A Standardized Approach to Performing the Action Research Arm Test,”Neurorehabilitation and Neural Repair, vol. 22, no. 1, pp. 78–90, Jan. 2008. [Online]. Available: https://doi.org/10.1177/1545968307305353 15/23

  12. [12]

    A Systematic Review of International Clinical Guidelines for Rehabilitation of People With Neurological Conditions: What Recommendations Are Made for Upper Limb Assessment?

    J. Burridge, M. Alt Murphy, J. Buurke, P. Feys, T. Keller, V . Klamroth-Marganska, I. Lamers, L. McNicholas, G. Prange, I. Tarkka, A. Timmermans, and A.-M. Hughes, “A Systematic Review of International Clinical Guidelines for Rehabilitation of People With Neurological Conditions: What Recommendations Are Made for Upper Limb Assessment?”Frontiers in Neurol...

  13. [13]

    Standardized International Manual of the Fugl-Meyer Assessment of Motor Function After Stroke,

    J. Hervé-Colas, S. P. Newton, S. T. Engelter, K. S. Hayward, J. P. O. Held, N. Intering, G. Kwakkel, J. Pohl, D. S. Reisman, A. Schwarz, K. S. Sunnerhagen, J. M. Veerbeek, K. Wiesner, S. B. Zandvliet, and M. Alt Murphy, “Standardized International Manual of the Fugl-Meyer Assessment of Motor Function After Stroke,”Neurorehabilitation and Neural Repair, vo...

  14. [14]

    Analysing the Action Research Arm Test (ARAT): a cautionary tale from the RATULS trial,

    N. Wilson, D. Howel, H. Bosomworth, L. Shaw, and H. Rodgers, “Analysing the Action Research Arm Test (ARAT): a cautionary tale from the RATULS trial,”International Journal of Rehabilitation Research. Internationale Zeitschrift Fur Rehabilitationsforschung. Revue Internationale De Recherches De Readaptation, vol. 44, no. 2, pp. 166–169, Jun. 2021. [Online]...

  15. [15]

    Kinematic Analysis Using 3D Motion Capture of Drinking Task in People With and Without Upper-extremity Impairments,

    M. Alt Murphy, S. Murphy, H. C. Persson, U.-B. Bergström, and K. S. Sunnerhagen, “Kinematic Analysis Using 3D Motion Capture of Drinking Task in People With and Without Upper-extremity Impairments,”Journal of Visualized Experiments : JoVE, no. 133, p. 57228, Mar. 2018. [Online]. Available: https://pmc.ncbi.nlm.nih.gov/articles/PMC5933268/

  16. [16]

    Responsiveness of upper extremity kinematic measures and clinical improvement during the first three months after stroke,

    M. A. Murphy, C. Willén, and K. S. Sunnerhagen, “Responsiveness of upper extremity kinematic measures and clinical improvement during the first three months after stroke,”Neurorehabilitation and Neural Repair, vol. 27, no. 9, pp. 844–853, Nov. 2013

  17. [17]

    Differentiable Biomechanics Unlocks Opportunities for Markerless Motion Capture,

    R. J. Cotton, “Differentiable Biomechanics Unlocks Opportunities for Markerless Motion Capture,” in2025 International Conference On Rehabilitation Robotics (ICORR), May 2025, pp. 44–51, iSSN: 1945-7901. [Online]. Available: https://ieeexplore.ieee.org/document/11063174

  18. [18]

    Pose2Sim: An End-to-End Workflow for 3D Markerless Sports Kinemat- ics—Part 2: Accuracy,

    D. Pagnon, M. Domalain, and L. Reveret, “Pose2Sim: An End-to-End Workflow for 3D Markerless Sports Kinemat- ics—Part 2: Accuracy,”Sensors, vol. 22, no. 7, Apr. 2022

  19. [19]

    A comparison of lower body gait kinematics and kinetics between Theia3D markerless and marker-based models in healthy subjects and clinical patients,

    S. D’Souza, T. Siebert, and V . Fohanno, “A comparison of lower body gait kinematics and kinetics between Theia3D markerless and marker-based models in healthy subjects and clinical patients,”Scientific Reports, vol. 14, no. 1, p. 29154, Nov. 2024. [Online]. Available: https://www.nature.com/articles/s41598-024-80499-8 20.J. Matthis, “FreeMoCap - Free Mot...

  20. [21]

    OpenCap: Human movement dynamics from smartphone videos,

    S. D. Uhlrich, A. Falisse, L. Kidzi ´nski, J. Muccini, M. Ko, A. S. Chaudhari, J. L. Hicks, and S. L. Delp, “OpenCap: Human movement dynamics from smartphone videos,”PLOS Computational Biology, vol. 19, no. 10, p. e1011462, Oct

  21. [22]

    MAMMA: Markerless & Automatic Multi-Person Motion Action Capture,

    H. Cuevas-Velasquez, A. Yiannakidis, S. Shin, G. Becherini, M. Höschle, J. Tesch, T. Obersat, T. Alexiadis, E. Halilaj, and M. J. Black, “MAMMA: Markerless & Automatic Multi-Person Motion Action Capture,” Apr. 2026, arXiv:2506.13040 [cs.CV]. [Online]. Available: http://arxiv.org/abs/2506.13040

  22. [23]

    Differentiable Biomechanics for Markerless Motion Capture in Upper Limb Stroke Rehabilitation: A Comparison With Optical Motion Capture,

    T. Unger, A. S. Moslehian, J. Peiffer, J. Ullrich, R. Gassert, O. Lambercy, R. J. Cotton, and C. A. Easthope, “Differentiable Biomechanics for Markerless Motion Capture in Upper Limb Stroke Rehabilitation: A Comparison With Optical Motion Capture,”IEEE Transactions on Medical Robotics and Bionics, pp. 1–1, 2025. [Online]. Available: https://ieeexplore.iee...

  23. [24]

    Towards routine biomechanical data collection in stroke rehabilitation: a usability comparison of IMU and markerless motion capture systems for functional upper-limb assessments,

    T. Unger, X. Yao, A. Schmitt, L. Cebulla, and C. Easthope Awai, “Towards routine biomechanical data collection in stroke rehabilitation: a usability comparison of IMU and markerless motion capture systems for functional upper-limb assessments,”Journal of NeuroEngineering and Rehabilitation, Jul. 2026. [Online]. Available: https://doi.org/10.1186/s12984-02...

  24. [25]

    D. W. Hosmer, S. Lemeshow, and R. X. Sturdivant,Applied logistic regression, 3rd ed., ser. Wiley series in probability and statistics. Hoboken, New Jersey: Wiley, 2013

  25. [26]

    Systematic Review on Kinematic Assessments of Upper Limb Movements After Stroke,

    A. Schwarz, C. M. Kanzler, O. Lambercy, A. R. Luft, and J. M. Veerbeek, “Systematic Review on Kinematic Assessments of Upper Limb Movements After Stroke,”Stroke, vol. 50, no. 3, pp. 718–727, Mar. 2019

  26. [27]

    How many trials are needed in kinematic analysis of reach-to-grasp?—A study of the drinking task in persons with stroke and non-disabled controls,

    G. E. Frykberg, H. Grip, and M. A. Murphy, “How many trials are needed in kinematic analysis of reach-to-grasp?—A study of the drinking task in persons with stroke and non-disabled controls,”Journal of NeuroEngineering and Rehabilitation, vol. 18, no. 1, pp. number–101, Dec. 2021. 16/23

  27. [28]

    H. C. W. de Vet, C. B. Terwee, L. B. Mokkink, and D. L. Knol,Measurement in Medicine: A Practical Guide, ser. Practical Guides to Biostatistics and Epidemiology. Cambridge: Cambridge University Press, 2011. [Online]. Available: https://www.cambridge.org/core/books/measurement-in-medicine/8BD913A1DA0ECCBA951AC4C1F719BCC5

  28. [29]

    Using deep learning to detect upper limb compensation in individuals post-stroke using consumer-grade webcams—A feasibility study,

    T. Unger, B. Kühnis, L. Sauerzopf, M. R. R. Spiess, A. de Spindler, A. R. R. Luft, C. Easthope Awai, J. G. G. Schönhammer, and E. Gavagnin, “Using deep learning to detect upper limb compensation in individuals post-stroke using consumer-grade webcams—A feasibility study,”Frontiers in Medicine, vol. 12, Nov. 2025. [Online]. Available: https://www.frontiers...

  29. [30]

    Evaluating inter- and intra-rater reliability in assessing upper limb compensatory movements post-stroke: creating a ground truth through video analysis?

    L. Sauerzopf, C. G. C. Panduro, A. R. Luft, B. Kühnis, E. Gavagnin, T. Unger, C. E. Awai, J. G. Schönhammer, J. Degenfellner, and M. R. Spiess, “Evaluating inter- and intra-rater reliability in assessing upper limb compensatory movements post-stroke: creating a ground truth through video analysis?”Journal of NeuroEngineering and Rehabilitation, vol. 21, n...

  30. [31]

    From Metrics to Meaning in Neurological Rehabilitation: Clinicians’ Perspectives on Digital Metrics of Upper Limb Functioning—A Focus Group Study,

    J. Pohl, L. Mayrhuber, O. Lambercy, and C. E. Awai, “From Metrics to Meaning in Neurological Rehabilitation: Clinicians’ Perspectives on Digital Metrics of Upper Limb Functioning—A Focus Group Study,”JMIR Rehabilitation and Assistive Technologies, vol. 13, no. 1, p. e87339, Jun. 2026. [Online]. Available: https://rehab.jmir.org/2026/1/e87339

  31. [32]

    Enabling precision rehabilitation interventions using wearable sensors and machine learning to track motor recovery,

    C. Adans-Dester, N. Hankov, A. O’Brien, G. Vergara-Diaz, R. Black-Schaffer, R. Zafonte, J. Dy, S. I. Lee, and P. Bonato, “Enabling precision rehabilitation interventions using wearable sensors and machine learning to track motor recovery,”npj Digital Medicine, vol. 3, no. 1, p. 121, Sep. 2020. [Online]. Available: https://www.nature.com/articles/s41746-02...

  32. [33]

    Construct validity and responsiveness of clinical upper limb measures and sensor-based arm use within the first year after stroke: a longitudinal cohort study,

    J. Pohl, G. Verheyden, J. P. O. Held, A. R. Luft, C. Easthope Awai, and J. M. Veerbeek, “Construct validity and responsiveness of clinical upper limb measures and sensor-based arm use within the first year after stroke: a longitudinal cohort study,”Journal of NeuroEngineering and Rehabilitation, vol. 22, no. 1, p. 14, Jan. 2025. [Online]. Available: https...

  33. [34]

    The Gait Deviation Index: a new comprehensive index of gait pathology,

    M. H. Schwartz and A. Rozumalski, “The Gait Deviation Index: a new comprehensive index of gait pathology,”Gait & Posture, vol. 28, no. 3, pp. 351–357, Oct. 2008

  34. [35]

    Upper extremity kinematics: development of a quantitative measure of impairment severity and dissimilarity after stroke,

    K. F. Zaidi and M. Harris-Love, “Upper extremity kinematics: development of a quantitative measure of impairment severity and dissimilarity after stroke,”PeerJ, vol. 11, p. e16374, Dec. 2023. [Online]. Available: https://peerj.com/articles/16374

  35. [36]

    BiomechGPT: Towards a Biomechanically Fluent Multimodal Foundation Model for Clinically Relevant Motion Tasks,

    R. Yang, A. Kennedy, and R. J. Cotton, “BiomechGPT: Towards a Biomechanically Fluent Multimodal Foundation Model for Clinically Relevant Motion Tasks,” May 2025, arXiv:2505.18465 [cs]. [Online]. Available: http://arxiv.org/abs/2505.18465

  36. [37]

    Optimisation and Comparison of Markerless and Marker-Based Motion Capture Methods for Hand and Finger Movement Analysis,

    V . Maggioni, C. Azevedo-Coste, S. Durand, and F. Bailly, “Optimisation and Comparison of Markerless and Marker-Based Motion Capture Methods for Hand and Finger Movement Analysis,”Sensors (Basel, Switzerland), vol. 25, no. 4, p. 1079, Feb. 2025. [Online]. Available: https://pmc.ncbi.nlm.nih.gov/articles/PMC11858933/

  37. [38]

    Biomechanical Arm and Hand Tracking with Multiview Markerless Motion Capture,

    P. Firouzabadi, W. Murray, A. R. Sobinov, J. Peiffer, K. Shah, L. E. Miller, and R. J. Cotton, “Biomechanical Arm and Hand Tracking with Multiview Markerless Motion Capture,” in2024 10th IEEE RAS/EMBS International Conference for Biomedical Robotics and Biomechatronics (BioRob). Heidelberg, Germany: IEEE, Sep. 2024, pp. 1641–1648. [Online]. Available: htt...

  38. [39]

    Monocular Markerless Motion Capture Enables Quantitative Assessment of Upper Extremity Reachable Workspace,

    S. Donahue, J. D. Peiffer, R. T. Richardson, Y . Zhong, S. Q. Y . Tan, B. L. Marteau, S. A. Russo, M. D. Wang, R. J. Cotton, and R. Chafetz, “Monocular Markerless Motion Capture Enables Quantitative Assessment of Upper Extremity Reachable Workspace,”Sensors, vol. 26, no. 11, p. 3421, Jan. 2026. [Online]. Available: https://www.mdpi.com/1424-8220/26/11/3421

  39. [40]

    SAM 3D Body: Robust Full-Body Human Mesh Recovery,

    X. Yang, D. Kukreja, D. Pinkus, A. Sagar, T. Fan, J. Park, S. Shin, J. Cao, J. Liu, N. Ugrinovic, M. Feiszli, J. Malik, P. Dollar, and K. Kitani, “SAM 3D Body: Robust Full-Body Human Mesh Recovery,” Feb. 2026, arXiv:2602.15989 [cs.CV]. [Online]. Available: http://arxiv.org/abs/2602.15989

  40. [41]

    Monocular Biomechanical Tracking of Fingers with Inverse Kinematics to Foundation Models,

    R. J. Cotton, P. Firouzabadi, and W. Murray, “Monocular Biomechanical Tracking of Fingers with Inverse Kinematics to Foundation Models,” May 2026. [Online]. Available: http://arxiv.org/abs/2605.09258 17/23 Acknowledgements We thank all participants of the study. This study was enabled through generous funding from the P & K Foundation. Author contribution...

  41. [2023]

    Available: https://journals.plos.org/ploscompbiol/article?id=10.1371/journal.pcbi.1011462

    [Online]. Available: https://journals.plos.org/ploscompbiol/article?id=10.1371/journal.pcbi.1011462