Pith. sign in

REVIEW 7 cited by

AP-10K: A Benchmark for Animal Pose Estimation in the Wild

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2108.12617 v2 pith:WVF5TQEP submitted 2021-08-28 cs.CV

classification cs.CV
keywords animalestimationposeap-10kanimalsbenchmarkgeneralizationlearning
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Accurate animal pose estimation is an essential step towards understanding animal behavior, and can potentially benefit many downstream applications, such as wildlife conservation. Previous works only focus on specific animals while ignoring the diversity of animal species, limiting the generalization ability. In this paper, we propose AP-10K, the first large-scale benchmark for mammal animal pose estimation, to facilitate the research in animal pose estimation. AP-10K consists of 10,015 images collected and filtered from 23 animal families and 54 species following the taxonomic rank and high-quality keypoint annotations labeled and checked manually. Based on AP-10K, we benchmark representative pose estimation models on the following three tracks: (1) supervised learning for animal pose estimation, (2) cross-domain transfer learning from human pose estimation to animal pose estimation, and (3) intra- and inter-family domain generalization for unseen animals. The experimental results provide sound empirical evidence on the superiority of learning from diverse animals species in terms of both accuracy and generalization ability. It opens new directions for facilitating future research in animal pose estimation. AP-10k is publicly available at https://github.com/AlexTheBad/AP10K.

Discussion (0). Sign in to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. MOSAIC: Multi-Subject Personalized Generation via Correspondence-Aware Alignment and Disentanglement

    cs.CV 2025-09 conditional novelty 7.0 of 10

    MOSAIC improves multi-subject personalized image generation by supervising attention maps with semantic point correspondences and a disentanglement loss, and introduces the SemAlign-MS dataset for training.

  2. Promptable Animal Pose Tracking Across Species

    cs.CV 2026-08 conditional novelty 6.0 of 10

    A promptable framework using frozen foundation-model features achieves competitive keypoint tracking accuracy and cross-species generalization on animal video benchmarks, with supervised and unsupervised variants.

  3. Fine-Grained Zero-Shot Object Detection

    cs.CV 2025-07 conditional novelty 6.0 of 10

    The authors define fine-grained zero-shot object detection, build a 1,432-species bird benchmark (FGZSD-Birds), and show their hierarchical MSHC detector outperforms prior ZSD models on that benchmark.

  4. Awakening Diffusion Transformers: Eliciting Stronger Generation and Understanding via Massive Activation Modulation

    cs.CV 2026-07 conditional novelty 5.0 of 10

    Massive activations in DiTs are timestep-driven detail channels; suppressing them guides finer sampling and AdaLN-modulating them yields more discriminative dense features.

  5. Parse Graph-Based Visual-Language Interaction for Human Pose Estimation

    cs.CV 2025-09 conditional novelty 4.0 of 10

    A parse-graph-based visual-language interaction module with a Guided Module improves human and animal pose estimation, especially under occlusion.

  6. Performance is not All You Need: Sustainability Considerations for Algorithms

    cs.CV 2025-08 reject novelty 4.0 of 10

    The paper introduces FMS and ASC, composite sustainability scores that fuse accuracy and energy consumption, and evaluates them on multiple vision tasks.

  7. KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model

    cs.CV 2025-07 conditional novelty 4.0 of 10

    KptLLM++ unifies keypoint semantic understanding, visual-prompt detection, and text-prompt detection in a single multimodal LLM, reporting SOTA accuracy on COCO, AP-10K, Human-Art, and other benchmarks.

Pith tools