Pith. sign in

REVIEW 3 cited by

HDFormer: High-order Directed Transformer for 3D Human Pose Estimation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2302.01825 v2 pith:5KL2ICZL submitted 2023-02-03 cs.CV cs.AI

classification cs.CVcs.AI
keywords hdformerhigh-orderjointestimationposeleftrightarrowattentionbone
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

Human pose estimation is a challenging task due to its structured data sequence nature. Existing methods primarily focus on pair-wise interaction of body joints, which is insufficient for scenarios involving overlapping joints and rapidly changing poses. To overcome these issues, we introduce a novel approach, the High-order Directed Transformer (HDFormer), which leverages high-order bone and joint relationships for improved pose estimation. Specifically, HDFormer incorporates both self-attention and high-order attention to formulate a multi-order attention module. This module facilitates first-order "joint$\leftrightarrow$joint", second-order "bone$\leftrightarrow$joint", and high-order "hyperbone$\leftrightarrow$joint" interactions, effectively addressing issues in complex and occlusion-heavy situations. In addition, modern CNN techniques are integrated into the transformer-based architecture, balancing the trade-off between performance and efficiency. HDFormer significantly outperforms state-of-the-art (SOTA) models on Human3.6M and MPI-INF-3DHP datasets, requiring only 1/10 of the parameters and significantly lower computational costs. Moreover, HDFormer demonstrates broad real-world applicability, enabling real-time, accurate 3D pose estimation. The source code is in https://github.com/hyer/HDFormer

Discussion (0). Sign in to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. KASportsFormer: Kinematic Anatomy Enhanced Transformer for 3D Human Pose Estimation on Short Sports Scene Video

    cs.CV 2025-07 conditional novelty 5.0 of 10

    KASportsFormer combines bone and limb tokens with cross-attention in a spatio-temporal transformer, reporting state-of-the-art MPJPE on SportsPose and WorldPose short-video benchmarks.

  2. A Structure-aware and Motion-adaptive Framework for 3D Human Pose Estimation with Mamba

    cs.CV 2025-07 conditional novelty 5.0 of 10

    SAMA adds a structure-aware state integrator and a motion-adaptive timescale modulator to Mamba-based pose lifting, reaching 36.5 mm MPJPE on Human3.6M with lower cost than prior Mamba methods.

  3. A Spatio-temporal Continuous Network for Stochastic 3D Human Motion Prediction

    cs.CV 2025-08 unverdicted novelty 4.0 of 10

    STCN predicts stochastic 3D human futures by combining a spatio-temporal continuous network with anchor-based Gaussian mixture sampling.

Pith tools