Pith. sign in

REVIEW 1 cited by

AggPose: Deep Aggregation Vision Transformer for Infant Pose Estimation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2205.05277 v2 pith:J6RBLXEJ submitted 2022-05-11 cs.CV cs.AI

classification cs.CVcs.AI
keywords poseestimationaggposeinfantdatasettransformeraggregationdeep
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Movement and pose assessment of newborns lets experienced pediatricians predict neurodevelopmental disorders, allowing early intervention for related diseases. However, most of the newest AI approaches for human pose estimation methods focus on adults, lacking publicly benchmark for infant pose estimation. In this paper, we fill this gap by proposing infant pose dataset and Deep Aggregation Vision Transformer for human pose estimation, which introduces a fast trained full transformer framework without using convolution operations to extract features in the early stages. It generalizes Transformer + MLP to high-resolution deep layer aggregation within feature maps, thus enabling information fusion between different vision levels. We pre-train AggPose on COCO pose dataset and apply it on our newly released large-scale infant pose estimation dataset. The results show that AggPose could effectively learn the multi-scale features among different resolutions and significantly improve the performance of infant pose estimation. We show that AggPose outperforms hybrid model HRFormer and TokenPose in the infant pose estimation dataset. Moreover, our AggPose outperforms HRFormer by 0.8 AP on COCO val pose estimation on average. Our code is available at github.com/SZAR-LAB/AggPose.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Decoding Children's Gait Behavior

    cs.CV 2026-08 conditional novelty 6.0 of 10

    A new 1,185-video pediatric gait dataset with EVGS labels, plus a VideoMAE-based model, reaches ~84% average accuracy on 34 clinical gait items — far above MLLMs and prior gait models.

Pith tools