Pith. sign in

REVIEW 4 major objections 4 minor 63 references

Joint angle based learning to refine kinematic human pose estimation

T0 review · 4 major / 4 minor · reviewed 2026-08-06 · deepseek-v4-flash

Pith's one-line read Refining human poses in joint-angle space corrects roughly twice as many keypoint outliers as the leading baseline, and yields smoother sports trajectories.

desk verdict The joint-angle idea is good, but the quantitative claims against SmoothNet are not yet justified by the paper's evidence. read the letter →

arxiv 2507.11075 v2 pith:6HNFVA3O submitted 2025-07-15 cs.CV cs.AI

classification cs.CVcs.AI
keywords humanposeestimationkinematicrefinementjointanglerepresentationFourierseriesmotionmodelbidirectionalGRUattentionmechanismoutliercorrectionsportsanalysis
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper argues that the best place to clean up noisy human pose estimates is joint-angle space rather than pixel-coordinate space. It introduces JAR, a post-processing pipeline that converts keypoint coordinates into 12 joint angles, smooths those angle sequences with a bidirectional recurrent network, and reconstructs keypoint positions under the constraint that limb lengths stay roughly constant. Because manually annotated pose datasets contain errors, the paper avoids using them for refinement training: it synthesizes ground-truth angle sequences with an eighth-order Fourier series fit to open biomechanics data, then adds jitter and outliers. On sprint and standing triple jump videos, JAR corrects 97.77% of erroneous frames, nearly double the rate of the comparison network, and it visibly stabilizes velocities in figure skating and breaking. If correct, the method gives sports analysts a way to turn single-image pose models into biomechanically plausible motion trajectories and even repair inconsistent annotations in existing video datasets.

What carries the argument

The central object is the joint-angle representation of a pose: 12 angles formed by adjacent triples of the 13 keypoints, computed with arctangent. The load-bearing equations are the 8th-order Fourier-series model of angle variation, which supplies synthetic ground-truth training sequences, and the limb-length-constrained reconstruction that maps smoothed angles back to keypoint positions while holding limb lengths and their ratios nearly fixed across frames. The BiGRU-Attention network carries temporal denoising, and sliding-window averaging with distance-based weights combines outputs from overlapping windows.

What would settle it

Take a video of a motion with strong 3D limb rotation, such as a figure skater's spin where the leg rotates toward and away from the camera, alongside synchronized motion-capture ground truth; run JAR on the 2D video and compare each reconstructed limb length and keypoint position to the projected motion-capture positions, and check whether errors grow with the rotation angle.

Watch

Extended reading notes

Core claim

The paper establishes that recognized keypoint errors and jitters are better modeled and removed as errors in joint angles than as errors in pixel coordinates. The pipeline defines 12 joint angles from 13 keypoints using arctangent, treats the nose as a base point, smooths the nose trajectory with a Savitzky-Golay filter, and reconstructs all other keypoints from smoothed angles plus optimized limb lengths, where limb lengths are assumed to vary little across frames and to preserve consistent ratios. Training data are generated by representing joint-angle variation as an 8th-order Fourier series with coefficients fit to open biomechanics datasets; parameters are perturbed to synthesize 512,000 training and 128,000 test segments of 100 frames, with Gaussian jitter and up to 5% outlier frames. A two-layer bidirectional gated recurrent network with attention, applied through sliding windows, denoises each angle sequence. On the reported athletic cases, JAR corrects 95.61% of erroneous frames in standing triple jump and 100% in sprint, versus 53.51% and 45.45% for SmoothNet, and produces smoother velocity curves in figure skating and breaking. The same pipeline is applied to correct jittered annotations in the PoseTrack dataset.

Load-bearing premise

The load-bearing assumption is that the apparent length of each limb in the images stays approximately constant from frame to frame, so that position reconstruction from smoothed angles and fixed limb lengths remains valid; for limbs rotating sharply in 3D, perspective shortening breaks this assumption and the refined pose can be distorted even when the smoothed angles are right.

Editorial extensions

If this is right

  • JAR can be appended to any single-image pose estimator, converting per-frame outputs into temporally consistent motion curves suitable for velocity and acceleration analysis.
  • The method reduces sensitivity to smoothing-window size compared with the SmoothNet baseline, making it more reliable for high-amplitude sports motions without extensive tuning.
  • The Fourier-series data generation scheme means refinement training does not depend on manual annotations for the refinement task, addressing a key bottleneck in pose-refinement learning.
  • The same pipeline can be used to re-annotate video pose datasets such as PoseTrack, reducing inter-frame annotation inconsistencies and potentially improving downstream training data quality.
  • Smoother trajectories yield physiologically coherent velocity profiles, which matters for coaching, biomechanical analysis, and referee-assisted scoring in sports.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Because the method operates on 2D projected angles, its main risk is motions with strong 3D limb rotation; a natural test is to compare JAR reconstruction against motion-capture ground truth for twisting jumps and spins.
  • The Fourier template assumes roughly periodic motion, so non-cyclic or transitional movements such as preparation-to-takeoff may lie outside the synthetic training distribution; mixed cyclic and non-cyclic augmentation could be tested.
  • The dataset-rectification use suggests a closed-loop validation: retrain a pose estimator on JAR-cleaned annotations and check whether downstream pose accuracy improves, which would test whether smoothing actually aids learning and not just visualization.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 4 minor

Summary. The paper proposes JAR, a post-processing refinement pipeline for single-image human pose estimation (HPE) that operates in joint-angle space. Raw 2D keypoints from an HPE model (HRNet or ViTPose++) are converted to joint angles, smoothed by a BiGRU-Attention sequence-to-sequence network, and then mapped back to keypoint coordinates using a limb-length optimization with biomechanical constraints. Training data are synthesized by fitting 8th-order Fourier series to open joint-angle datasets and adding Gaussian jitters and outliers. The authors compare JAR against SmoothNet on real sports videos (standing triple jump, sprint, figure skating, breaking) and report substantially higher outlier correction rates (95.61% vs. 53.51% and 100% vs. 45.45%). They also demonstrate the method's use for correcting annotations in the PoseTrack dataset.

Significance. If the central claims held, a joint-angle-based refinement module that outperforms SmoothNet on challenging, high-amplitude sports poses would be practically valuable for sports biomechanics and video-dataset rectification. The paper proposes a principled representation (joint angles) with a plausible training-data generation strategy, and the architecture (BiGRU-Attention) is simple and reproducible. However, the current evidence for the headline claims is substantially weakened by circular quantitative evaluation, manual error counting, and a reconstruction assumption that may be violated in exactly the scenarios used for demonstration. The idea is worth pursuing, but the paper as written does not yet establish superiority over SmoothNet on real pose data with objective metrics.

major comments (4)
  1. [§5.2, Fig. 8] The quantitative MSE comparison in Fig. 8 is circular: the test set is generated by the same 8th-order Fourier-series pipeline described in Section 3.2/3.3 that produced the training data. This measures how well each model inverts the synthetic generator, not how well the models generalize to the error statistics of real HPE outputs. The MSE numbers therefore cannot support the abstract's claim of 'outstanding performance' or the selection of BiGRU-Attention as the best architecture. Please re-evaluate on independently obtained test data, e.g., real pose sequences with ground-truth 2D/3D keypoints.
  2. [§3.1, Eq. (6)] The reconstruction stage relies on the assumption that projected limb lengths are approximately constant across frames (the second constraint in Eq. (6)). In monocular video, the apparent length of a limb segment is L·cos(α), where α is the angle between the segment and the image plane; α changes substantially during the figure-skating flying spin, breaking handstand, and standing triple jump sequences shown in Figs. 5–7. When out-of-plane rotation occurs, the fidelity term pulls toward the raw projected lengths while the ratio/regularization terms pull toward a constant-length skeleton, so the reconstructed positions are systematically biased even if the smoothed joint angles are perfect. Since no ground-truth 2D/3D positions are provided for these real sequences, the reported correction rates in Table 1 cannot separate this artifact from genuine accuracy. The authors should either validate on sequences with known 3D motion or demonstrate that the assumption holds quantitatively for the test sports.
  3. [Table 1] The outlier correction rates in Table 1 are computed from manually identified 'Erroneous result [frame]' and 'Corrected image [frame]' counts. No criteria are given for what constitutes an erroneous frame (e.g., left-right confusion, joint displacement threshold) or for what counts as 'corrected'. Without a precise, repeatable definition and ideally multiple annotators, the headline numbers (95.61% vs. 53.51%) are not reproducible and should not be used as primary evidence of superiority.
  4. [§5.1, Figs. 5–7] The comparison with SmoothNet on real sports data is qualitative only: the authors show selected trajectories and velocity curves and describe differences as 'smoother' or 'more stable'. No standard quantitative metrics are reported, such as percentage of correct keypoints, mean per-joint position error on a public video benchmark, or a smoothness measure computable without ground truth (e.g., acceleration jerk). Because the central claim is that JAR outperforms SmoothNet, the paper needs at least one quantitative, objective evaluation on real data before that claim can be considered supported.
minor comments (4)
  1. [§3.1, Eq. (2)] The angle definition in Eq. (2) appears to be the angle between vectors OA and OB, but the formula as written gives only the arctangent of a ratio, not the full angle; the text should clarify how the sign and range of the angle are determined and how the two vectors are defined for each keypoint.
  2. [§3.3, Step 3] The description of outlier generation says 'A small number of secondary abnormal frames are placed before and after each outlier. The overall number follows Gaussian distribution with the average of zero and 3σ < 6.' This is unclear: does the Gaussian distribution apply to the number of secondary frames or to their angular values? Please rephrase.
  3. [§5.2, Fig. 8] The MSE values in Fig. 8(c) are presented without error bars or statistical significance tests. Given the small visual differences between BiGRU-Attention and BiGRU, the claim that attention 'contributes considerably' would be stronger with repeated-seed variance or confidence intervals.
  4. [Throughout] The manuscript uses phrases like 'outstanding performance' and 'precisely capture the biomechanical coherency' without operational definitions. Please replace these with quantitative measures or specifically defined qualitative criteria.

Circularity Check

1 steps flagged · score 6.0 of 10

The quantitative benchmark is self-referential: the Fourier-series 'ground truth' used to train JAR is the same generative model used to construct the test set, so the Fig. 8 MSE measures inversion of the paper's own generator rather than accuracy on measured human motion.

  1. fitted input called prediction [Section 3.2 (Eq. 8); Section 3.3 (dataset generation); Section 5.2 (Fig. 8, Eq. 11)]
    "The study ultimately generates 512,000 segments serving as the 'ground truth' for training set, and 128,000 segments for testing set. ... The 'ground truth' sequences are constructed using an 8th-order Fourier series with randomly superimposed jitters and outliers."

    The training target and the quantitative test target are both produced by the same 8th-order Fourier-series model introduced in Eq. (8). The network is trained to output the clean Fourier curve after noise and outliers are added, and Fig. 8's MSE (Eq. 11) is computed against that same constructed curve. Thus the evaluation measures how well BiGRU-Attention inverts the paper's own generator, not how accurately the smoothed angles match independently measured human kinematics. The 'reliable ground truth' is generated by the very approximation the method claims as its key technique, so the predicted smooth joint-angle sequence is, by construction, a regression to the generating model rather than an externally validated quantity.

full rationale

JAR's real-data demonstrations (Figs. 4-7, Table 1) are not circular in the narrow sense: they compare JAR's reconstructed coordinates against raw ViTPose++ outputs and against an external baseline, SmoothNet, on actual sports videos. However, the paper's only quantitative performance measure on the refinement network, Fig. 8, is self-referential: both training and test 'ground truth' sequences are synthesized with the same 8th-order Fourier-series model, so the MSE improvement largely confirms the network can denoise signals drawn from its own assumed generative family. The real-data outlier-correction rates in Table 1 are based on manually identified erroneous frames and visual/manual counting without an independent 2D or 3D ground truth, which weakens the evidence but is not itself a circular derivation. The limb-length constraints in Eq. (6) also make constant-length, smoothed-angle outputs a built-in property of JAR's reconstruction, so some of the apparent advantage over coordinate-space SmoothNet reflects the model's strong priors rather than measured accuracy. No load-bearing self-citation chain is present. Overall, the central quantitative evaluation reduces partly to the paper's own generator, giving partial circularity rather than full equivalence.

Assumptions & free parameters 5 free parameters · 6 assumptions · 0 invented entities

The central claim rests on a chain of assumptions not tested against external ground truth: periodicity of joint angles, negligible foreshortening of limb lengths, stable nose keypoints, and a hand-set noise model. The Fourier coefficients are fit to biomechanics datasets and then varied to synthesize targets, so the quantitative benchmark is largely internal to the pipeline.

free parameters (5)
  • Fourier series coefficients a0, ak, bk and period T per joint = Not tabulated; estimated by least squares from open gait datasets and then manually varied to synthesize samples…
    These coefficients define every synthetic ground-truth angle curve used for both training and testing, so the benchmark outcome depends on them.
  • Noise model amplitudes = Gaussian jitter with 3sigma below 45 degrees; outliers with 3sigma up to 135 degrees on 5% of frames
    Hand-set to approximate HPE output deviations, but not measured from actual HPE error statistics, so the network is tuned to an assumed noise distribution.
  • Savitzky-Golay base-point window w = 50 frames
    Hand-set to about half the frame rate; it controls how strongly the nose trajectory is smoothed and every reconstructed keypoint inherits that smoothing.
  • Inference sliding-window parameters S and lambda = lambda = 0.001; S is not stated
    Used in Eqs. (9)-(10) for weighted averaging of overlapping windows; S being unspecified leaves a free choice that affects smoothness.
  • Per-video limb lengths and ratio matrix from trust-region optimization = Optimized per sequence using Eq. (6)
    Imposes constant limb length and body proportion constraints; this per-video fitting shapes the final reconstructed trajectories.
assumptions (6)
  • domain assumption Joint angle variations in human activities are periodic and representable by an 8th-order Fourier series.
    Invoked in Sec. 3.2 (Eq. 8). This is the basis for the synthetic ground truth, but many evaluated activities such as breaking and handstands are not strictly periodic.
  • domain assumption Projected limb lengths are approximately constant across frames because foreshortening is negligible.
    Stated in Sec. 3.1 constraint (ii) and used in Eq. (6). Strong 3D rotation of limbs can violate this and produce distorted reconstructions.
  • domain assumption Limb length proportions between body segments are approximately consistent across frames.
    Sec. 3.1 constraint (i), used in the objective function Eq. (6). It is a biomechanical regularization, not a measured property of each video.
  • domain assumption The nose keypoint is the most accurate and stable keypoint, so only its trajectory is independently smoothed.
    Sec. 3.1 selects nose as the base point; if nose is wrong, all reconstructed keypoints inherit the error.
  • ad hoc to paper Synthetic Gaussian jitter and outlier distributions match real single-image HPE error distributions.
    Sec. 3.3 Step 3 sets jitter and outlier amplitudes by hand. No calibration against real HPE failure modes is provided.
  • domain assumption A network trained only on Fourier-synthesized angle sequences generalizes to real video poses.
    Sections 4 and 5 assume transfer from synthetic to real; the paper does not demonstrate this with external ground-truth comparisons.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Joint angle based learning to refine kinematic human pose estimation." pith.science (2026). https://pith.science/paper/6HNFVA3O

@misc{pith2026250711075,
  author       = {Pith},
  title        = {Pith review of: Joint angle based learning to refine kinematic human pose estimation},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/6HNFVA3O}},
  note         = {Machine review of arXiv:2507.11075}
}
read the original abstract

Marker-free human pose estimation (HPE) has found increasing applications in various fields. Current HPE suffers from occasional errors in keypoint recognition and random fluctuation in keypoint trajectories when analyzing kinematic human poses. The performance of existing deep learning-based models for HPE refinement is considerably limited by inaccurate training datasets in which the keypoints are manually annotated. This paper proposed a novel method to overcome the difficulty, in which the key techniques include: (i) A robust joint angle-based description of kinematic human poses; (ii) Approximating temporal variation of joint angles using high order Fourier series to get reliable "ground truth"; (iii) A bidirectional recurrent network is designed as a post-processing module to refine the estimation of single image-based HPE models. Trained with the high-quality dataset constructed using our method, the network demonstrates outstanding performance to correct wrongly recognized joints and smooth their spatiotemporal trajectories. Tests show that joint angle-based refinement (JAR) outperforms the state-of-the-art HPE refinement network in challenging cases like figure skating and breaking. JAR also demonstrates great potential to rectify existing datasets.

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

63 extracted references · 27 canonical work pages

  1. [1]

    Zheng, W

    C. Zheng, W. Wu, C. Chen, T. Yang, S. Zhu, J. Shen, N. Kehtarnavaz, M. Shah, Deep Learning -based Human Pose Estimation: A Survey, ACM Comput. Surv. 56 (2024) 1 –37. https://doi.org/10.1145/3603618

  2. [2]

    Srinivasan, J.J

    H.K. Srinivasan, J.J. Mathunny, A. Devaraj, V . Karthik, Validation of an Automated Step Length Measurement Method in Sprinting Athletes Using Computer Vision and Pose Estimation, in: 2023 Int. Conf. Recent Adv. Electr. Electron. Ubiquitous Commun. Comput. Intell. RAEEUCCI, IEEE, Chennai, India, 2023: pp. 1–5. https://doi.org/10.1109/RAEEUCCI57140.2023.10134177

  3. [4]

    J. Wang, K. Qiu, H. Peng, J. Fu, J. Zhu, AI Coach: Deep Human Pose Estimation and Analysis for Personalized Athletic Training Assistance, in: Proc. 27th ACM Int. Conf. Multimed., ACM, Nice France, 2019: pp. 374–382. https://doi.org/10.1145/3343031.3350910. 19

  4. [5]

    Jafarzadeh, P

    P. Jafarzadeh, P. Virjonen, P. Nevalainen, F. Farahnakian, J. Heikkonen, Pose Estimation of Hurdles Athletes using OpenPose, in: 2021 Int. Conf. Electr. Comput. Commun. Mechatron. Eng. ICECCME, IEEE, Mauritius, Mauritius, 2021: pp. 1–6. https://doi.org/10.1109/ICECCME52200.2021.9591066

  5. [6]

    Baumgartner, S

    T. Baumgartner, S. Klatt, Monocular 3D Human Pose Estimation for Sports Broadcasts using Partial Sports Field Registration, in: 2023 IEEECVF Conf. Comput. Vis. Pattern Recognit. Workshop CVPRW, IEEE, Vancouver, BC, Canada, 2023: pp. 5109 –5118. https://doi.org/10.1109/CVPRW59228.2023.00539

  6. [7]

    Stepec, D

    D. Stepec, D. Skocaj, Video-Based Ski Jump Style Scoring from Pose Trajectory, in: 2022 IEEECVF Winter Conf. Appl. Comput. Vis. Workshop WACVW, IEEE, Waikoloa, HI, USA, 2022: pp. 682–690. https://doi.org/10.1109/W ACVW54805.2022.00075

  7. [8]

    J. Wang, K. Sun, T. Cheng, B. Jiang, C. Deng, Y . Zhao, D. Liu, Y . Mu, M. Tan, X. Wang, W. Liu, B. Xiao, Deep High -Resolution Representation Learning for Visual Recognition, IEEE Trans. Pattern Anal. Mach. Intell. 43 (2021) 3349–3364. https://doi.org/10.1109/TPAMI.2020.2983686

  8. [9]

    S. Jin, L. Xu, J. Xu, C. Wang, W. Liu, C. Qian, W. Ouyang, P. Luo, Whole -Body Human Pose Estimation in the Wild, in: A. Vedaldi, H. Bischof, T. Brox, J.-M. Frahm (Eds.), Comput. Vis. – ECCV 2020, Springer International Publishing, Cham, 2020: pp. 196 –214. https://doi.org/10.1007/978 -3- 030-58545-7_12

Show all 63 references
  1. [10]

    Andriluka, L

    M. Andriluka, L. Pishchulin, P. Gehler, B. Schiele, 2D Human Pose Estimation: New Benchmark and State of the Art Analysis, in: 2014 IEEE Conf. Comput. Vis. Pattern Recognit., IEEE, Columbus, OH, USA, 2014: pp. 3686–3693. https://doi.org/10.1109/CVPR.2014.471

  2. [11]

    Eichner, M

    M. Eichner, M. Marin-Jimenez, A. Zisserman, V . Ferrari, 2D Articulated Human Pose Estimation and Retrieval in (Almost) Unconstrained Still Images, Int. J. Comput. Vis. 99 (2012) 190 –214. https://doi.org/10.1007/s11263-012-0524-9

  3. [12]

    Y . Yang, D. Ramanan, Articulated Human Detection with Flexible Mixtures of Parts, IEEE Trans. Pattern Anal. Mach. Intell. 35 (2013) 2878–2890. https://doi.org/10.1109/TPAMI.2012.261

  4. [13]

    Andriluka, U

    M. Andriluka, U. Iqbal, E. Insafutdinov, L. Pishchulin, A. Milan, J. Gall, B. Schiele, PoseTrack: A Benchmark for Human Pose Estimation and Tracking, in: 2018 IEEECVF Conf. Comput. Vis. Pattern Recognit., IEEE, Salt Lake City, UT, USA, 2018: pp. 5167 –5176. https://doi.org/10....

  5. [14]

    W. Lin, H. Liu, S. Liu, Y . Li, H. Xiong, G. Qi, N. Sebe, HiEve: A Large-Scale Benchmark for Human- Centric Video Analysis in Complex Events, Int. J. Comput. Vis. 131 (2023) 2994 –3018. https://doi.org/10.1007/s11263-023-01842-6. 20

  6. [15]

    Iqbal, A

    U. Iqbal, A. Milan, J. Gall, PoseTrack: Joint Multi -Person Pose Estimation and Tracking, (2016). https://doi.org/10.48550/ARXIV .1611.07727

  7. [18]

    D. Xu, R. Zhang, L. Guo, C. Feng, S. Gao, LDNet: Lightweight dynamic convolution network for human pose estimation, Adv. Eng. Inform. 54 (2022) 101785. https://doi.org/10.1016/j.aei.2022.101785

  8. [19]

    Sheng, L

    B. Sheng, L. Chen, J. Cheng, Y . Zhang, Z. Hua, J. Tao, A markless 3D human motion data acquisition method based on the binocular stereo vision and lightweight open pose algorithm, Measurement 225 (2024) 113908. https://doi.org/10.1016/j.measurement.2023.113908

  9. [20]

    Z. Zhao, A. Song, S. Zheng, Q. Xiong, J. Guo, DSC -HRNet: a lightweight teaching pose estimation model with depthwise separable convolution and deep high -resolution representation learning in computer-aided education, Int. J. Inf. Technol. 15 (2023) 2373 –2385. https://doi.or...

  10. [21]

    Nokihara, R

    Y . Nokihara, R. Hachiuma, R. Hori, H. Saito, Future Prediction of Shuttlecock Trajectory in Badminton Using Player’s Information, J. Imaging 9 (2023) 99. https://doi.org/10.3390/jimaging9050099

  11. [22]

    Latreche, R

    A. Latreche, R. Kelaiaia, A. Chemori, A. Kerboua, Reliability and validity analysis of MediaPipe - based measurement system for some human rehabilitation motions, Measurement 214 (2023) 112826. https://doi.org/10.1016/j.measurement.2023.112826

  12. [23]

    Toshev, C

    A. Toshev, C. Szegedy, DeepPose: Human Pose Estimation via Deep Neural Networks, in: 2014 IEEE Conf. Comput. Vis. Pattern Recognit., IEEE, Columbus, OH, USA, 2014: pp. 1653 –1660. https://doi.org/10.1109/CVPR.2014.214

  13. [24]

    Tompson, A

    J. Tompson, A. Jain, Y . LeCun, C. Bregler, Joint Training of a Convolutional Network and a Graphical Model for Human Pose Estimation, (2014). https://doi.org/10.48550/ARXIV .1406.2984

  14. [25]

    S.-E. Wei, V . Ramakrishna, T. Kanade, Y . Sheikh, Convolutional Pose Machines, (2016). https://doi.org/10.48550/ARXIV .1602.00134. 21

  15. [26]

    Newell, K

    A. Newell, K. Yang, J. Deng, Stacked Hourglass Networks for Human Pose Estimation, (2016). https://doi.org/10.48550/ARXIV .1603.06937

  16. [27]

    W. Yang, S. Li, W. Ouyang, H. Li, X. Wang, Learning Feature Pyramids for Human Pose Estimation, in: 2017 IEEE Int. Conf. Comput. Vis. ICCV , IEEE, Venice, 2017: pp. 1290 –1299. https://doi.org/10.1109/ICCV .2017.144

  17. [28]

    L. Ke, M. -C. Chang, H. Qi, S. Lyu, Multi -Scale Structure -Aware Network for Human Pose Estimation, (2018). https://doi.org/10.48550/ARXIV .1803.09894

  18. [29]

    K. Sun, B. Xiao, D. Liu, J. Wang, Deep High -Resolution Representation Learning for Human Pose Estimation, in: 2019 IEEECVF Conf. Comput. Vis. Pattern Recognit. CVPR, IEEE, Long Beach, CA, USA, 2019: pp. 5686–5696. https://doi.org/10.1109/CVPR.2019.00584

  19. [30]

    Bulat, J

    A. Bulat, J. Kossaifi, G. Tzimiropoulos, M. Pantic, Toward fast and accurate human pose estimation via soft -gated skip connections, in: 2020 15th IEEE Int. Conf. Autom. Face Gesture Recognit. FG 2020, IEEE, Buenos Aires, Argentina, 2020: pp. 8–15. https://doi.org/10.1109/FG47...

  20. [31]

    Y . Xu, J. Zhang, Q. Zhang, D. Tao, ViTPose: Simple Vision Transformer Baselines for Human Pose Estimation, (2022). https://doi.org/10.48550/ARXIV .2204.12484

  21. [32]

    Y . Xu, J. Zhang, Q. Zhang, D. Tao, ViTPose++: Vision Transformer for Generic Body Pose Estimation, IEEE Trans. Pattern Anal. Mach. Intell. 46 (2024) 1212 –1230. https://doi.org/10.1109/TPAMI.2023.3330016

  22. [33]

    Y . Luo, J. Ren, Z. Wang, W. Sun, J. Pan, J. Liu, J. Pang, L. Lin, LSTM Pose Machines, (2017). https://doi.org/10.48550/ARXIV .1712.06316

  23. [34]

    W. Li, X. Xu, Y .-J. Zhang, Temporal Feature Correlation for Human Pose Estimation in Videos, in: 2019 IEEE Int. Conf. Image Process. ICIP, IEEE, Taipei, Taiwan, 2019: pp. 599 –603. https://doi.org/10.1109/ICIP.2019.8803797

  24. [35]

    Artacho, A

    B. Artacho, A. Savakis, UniPose: Unified Human Pose Estimation in Single Images and Videos, in: 2020 IEEECVF Conf. Comput. Vis. Pattern Recognit. CVPR, IEEE, Seattle, WA, USA, 2020: pp. 7033–7042. https://doi.org/10.1109/CVPR42600.2020.00706

  25. [36]

    D. Gai, R. Feng, W. Min, X. Yang, P. Su, Q. Wang, Q. Han, Spatiotemporal Learning Transformer for Video-Based Human Pose Estimation, IEEE Trans. Circuits Syst. Video Technol. 33 (2023) 4564 –

  26. [37]

    W. Li, H. Liu, R. Ding, M. Liu, P. Wang, W. Yang, Exploiting Temporal Contexts with Strided Transformer for 3D Human Pose Estimation, (2021). https://doi.org/10.48550/ARXIV .2103.14304. 22

  27. [38]

    Jiang, N.C

    T. Jiang, N.C. Camgoz, R. Bowden, Skeletor: Skeletal Transformers for Robust Body -Pose Estimation, in: 2021 IEEECVF Conf. Comput. Vis. Pattern Recognit. Workshop CVPRW, IEEE, Nashville, TN, USA, 2021: pp. 3389–3397. https://doi.org/10.1109/CVPRW53098.2021.00378

  28. [39]

    J. Li, C. Wang, H. Zhu, Y . Mao, H. -S. Fang, C. Lu, CrowdPose: Efficient Crowded Scenes Pose Estimation and a New Benchmark, in: 2019 IEEECVF Conf. Comput. Vis. Pattern Recognit. CVPR, IEEE, Long Beach, CA, USA, 2019: pp. 10855–10864. https://doi.org/10.1109/CVPR.2019.01112

  29. [40]

    T.-Y . Lin, M. Maire, S. Belongie, L. Bourdev, R. Girshick, J. Hays, P. Perona, D. Ramanan, C.L. Zitnick, P. Dollár, Microsoft COCO: Common Objects in Context, (2014). https://doi.org/10.48550/ARXIV .1405.0312

  30. [41]

    J. Wu, H. Zheng, B. Zhao, Y . Li, B. Yan, R. Liang, W. Wang, S. Zhou, G. Lin, Y . Fu, Y . Wang, Y . Wang, Large-Scale Datasets for Going Deeper in Image Understanding, in: 2019 IEEE Int. Conf. Multimed. Expo ICME, IEEE, Shanghai, China, 2019: pp. 1480 –1485. https://doi.org/10...

  31. [42]

    Fieraru, A

    M. Fieraru, A. Khoreva, L. Pishchulin, B. Schiele, Learning to Refine Human Pose Estimation, in: 2018 IEEECVF Conf. Comput. Vis. Pattern Recognit. Workshop CVPRW, IEEE, Salt Lake City, UT, 2018: pp. 318–31809. https://doi.org/10.1109/CVPRW.2018.00058

  32. [43]

    Moon, J.Y

    G. Moon, J.Y . Chang, K.M. Lee, PoseFix: Model -Agnostic General Human Pose Refinement Network, in: 2019 IEEECVF Conf. Comput. Vis. Pattern Recognit. CVPR, IEEE, Long Beach, CA, USA, 2019: pp. 7765–7773. https://doi.org/10.1109/CVPR.2019.00796

  33. [44]

    Z. Liu, H. Chen, R. Feng, S. Wu, S. Ji, B. Yang, X. Wang, Deep Dual Consecutive Network for Human Pose Estimation, in: 2021 IEEECVF Conf. Comput. Vis. Pattern Recognit. CVPR, IEEE, Nashville, TN, USA, 2021: pp. 525–534. https://doi.org/10.1109/CVPR46437.2021.00059

  34. [45]

    T. Wang, L. Jin, Z. Wang, J. Li, L. Li, F. Zhao, Y . Cheng, L. Yuan, L. Zhou, J. Xing, J. Zhao, SynSP: Synergy of Smoothness and Precision in Pose Sequences Refinement, in: 2024 IEEECVF Conf. Comput. Vis. Pattern Recognit. CVPR, IEEE, Seattle, WA, USA, 2024: pp. 1824 –1833. ht...

  35. [46]

    Véges, A

    M. Véges, A. Lőrincz, Temporal Smoothing for 3D Human Pose Estimation and Localization for Occluded People, in: H. Yang, K. Pasupa, A.C.-S. Leung, J.T. Kwok, J.H. Chan, I. King (Eds.), Neural Inf. Process., Springer International Publishing, Cham, 2020: pp. 557 –568. https://d...

  36. [47]

    A. Zeng, L. Yang, X. Ju, J. Li, J. Wang, Q. Xu, SmoothNet: A Plug -and-Play Network for Refining Human Poses in Videos, in: S. Avidan, G. Brostow, M. Cissé, G.M. Farinella, T. Hassner (Eds.), 23 Comput. Vis. – ECCV 2022, Springer Nature Switzerland, Cham, 2022: pp. 625 –642. h...

  37. [48]

    Scherpereel, D

    K. Scherpereel, D. Molinaro, O. Inan, M. Shepherd, A. Young, A human lower-limb biomechanics and wearable sensors dataset during cyclic and non -cyclic activities, Sci. Data 10 (2023) 924. https://doi.org/10.1038/s41597-023-02840-6

  38. [49]

    Reznick, K.R

    E. Reznick, K.R. Embry, R. Neuman, E. Bolívar -Nieto, N.P. Fey, R.D. Gregg, Lower-limb kinematics and kinetics during continuously varying human locomotion, Sci. Data 8 (2021) 282. https://doi.org/10.1038/s41597-021-01057-9

  39. [50]

    Helwig, K.A

    N.E. Helwig, K.A. Shorter, P. Ma, E.T. Hsiao -Wecksler, Smoothing spline analysis of variance models: A new tool for the analysis of cyclic biomechanical data, J. Biomech. 49 (2016) 3216 –3222. https://doi.org/10.1016/j.jbiomech.2016.07.035

  40. [51]

    Q. Mei, J. Fernandez, L. Xiang, Z. Gao, P. Yu, J.S. Baker, Y . Gu, Dataset of lower extremity joint angles, moments and forces in distance running, Heliyon 8 (2022) e11517. https://doi.org/10.1016/j.heliyon.2022.e11517

  41. [52]

    Mundt, W

    M. Mundt, W. Thomsen, T. Witter, A. Koeppe, S. David, F. Bamer, W. Potthast, B. Markert, Prediction of lower limb joint angles and moments during gait using artificial neural networks, Med. Biol. Eng. Comput. 58 (2020) 211–225. https://doi.org/10.1007/s11517-019-02061-3

  42. [53]

    Sivakumar, A.A

    S. Sivakumar, A.A. Gopalai, K.H. Lim, D. Gouwanda, S. Chauhan, Joint angle estimation with wavelet neural networks, Sci. Rep. 11 (2021) 10306. https://doi.org/10.1038/s41598-021-89580-y

  43. [54]

    J. Shi, J. Zhong, Y . Zhang, B. Xiao, L. Xiao, Y . Zheng, A dual attention LSTM lightweight model based on exponential smoothing for remaining useful life prediction, Reliab. Eng. Syst. Saf. 243 (2024) 109821. https://doi.org/10.1016/j.ress.2023.109821

  44. [55]

    D. Lim, D. Kim, J. Park, Momentum Observer -Based Collision Detection Using LSTM for Model Uncertainty Learning, in: 2021 IEEE Int. Conf. Robot. Autom. ICRA, IEEE, Xi’an, China, 2021: pp. 4516–4522. https://doi.org/10.1109/ICRA48506.2021.9561667

  45. [56]

    T. Xu, H. Tuo, Q. Fang, D. Shan, H. Jin, J. Fan, Y . Zhu, J. Zhao, A novel collision detection method based on current residuals for robots without joint torque sensors: A case study on UR10 robot, Robot. Comput.-Integr. Manuf. 89 (2024) 102777. https://doi.org/10.1016/j.rcim....

  46. [57]

    X. Wang, H. Zhang, Z. Du, Multiscale Noise Reduction Attention Network for Aeroengine Bearing Fault Diagnosis, IEEE Trans. Instrum. Meas. 72 (2023) 1 –10. https://doi.org/10.1109/TIM.2023.3268459. 24

  47. [58]

    S. Mo, H. Wang, B. Li, S. Fan, Y . Wu, X. Liu, TimeSQL: Improving multivariate time series forecasting with multi -scale patching and smooth quadratic loss, Inf. Sci. 671 (2024) 120652. https://doi.org/10.1016/j.ins.2024.120652

  48. [59]

    Lea, M.D

    C. Lea, M.D. Flynn, R. Vidal, A. Reiter, G.D. Hager, Temporal Convolutional Networks for Action Segmentation and Detection, in: 2017 IEEE Conf. Comput. Vis. Pattern Recognit. CVPR, IEEE, Honolulu, HI, 2017: pp. 1003–1012. https://doi.org/10.1109/CVPR.2017.113

  49. [60]

    Hochreiter, J

    S. Hochreiter, J. Schmidhuber, Long Short -Term Memory, Neural Comput. 9 (1997) 1735 –1780. https://doi.org/10.1162/neco.1997.9.8.1735

  50. [61]

    Schuster, K.K

    M. Schuster, K.K. Paliwal, Bidirectional recurrent neural networks, IEEE Trans. Signal Process. 45 (1997) 2673–2681. https://doi.org/10.1109/78.650093

  51. [62]

    Graves, J

    A. Graves, J. Schmidhuber, Framewise phoneme classification with bidirectional LSTM and other neural network architectures, Neural Netw. 18 (2005) 602 –610. https://doi.org/10.1016/j.neunet.2005.06.042

  52. [63]

    Chung, C

    J. Chung, C. Gulcehre, K. Cho, Y . Bengio, Empirical Evaluation of Gated Recurrent Neural Networks on Sequence Modeling, (2014). https://doi.org/10.48550/ARXIV .1412.3555

  53. [64]

    A. Gu, T. Dao, Mamba: Linear -Time Sequence Modeling with Selective State Spaces, (2023). https://doi.org/10.48550/ARXIV .2312.00752

  54. [65]

    T. Dao, A. Gu, Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality, (2024). https://doi.org/10.48550/ARXIV .2405.21060

  55. [4576]

    https://doi.org/10.1109/TCSVT.2023.3269666

Pith tools

Reviewed August 6, 2026 · model on record in the stance chip above.