Pith. sign in

REVIEW 12 cited by

AI Challenger : A Large-scale Dataset for Going Deeper in Image Understanding

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1711.06475 v1 pith:HCYG5ESB submitted 2017-11-17 cs.CV

classification cs.CV
keywords datasetlarge-scaleimageattributechallengercomputerdatasetskeypoint
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Significant progress has been achieved in Computer Vision by leveraging large-scale image datasets. However, large-scale datasets for complex Computer Vision tasks beyond classification are still limited. This paper proposed a large-scale dataset named AIC (AI Challenger) with three sub-datasets, human keypoint detection (HKD), large-scale attribute dataset (LAD) and image Chinese captioning (ICC). In this dataset, we annotate class labels (LAD), keypoint coordinate (HKD), bounding box (HKD and LAD), attribute (LAD) and caption (ICC). These rich annotations bridge the semantic gap between low-level images and high-level concepts. The proposed dataset is an effective benchmark to evaluate and improve different computational methods. In addition, for related tasks, others can also use our dataset as a new resource to pre-train their models.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 12 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. MetricHMSR:Metric Human Mesh and Scene Recovery from Monocular Images

    cs.CV 2025-06 conditional novelty 6.0 of 10

    From one image, MetricHMSR jointly estimates a metric human mesh, its global 3D position, and a corrected metric scene depth map.

  2. PoseBH: Prototypical Multi-Dataset Training Beyond Human Pose Estimation

    cs.CV 2025-05 conditional novelty 6.0 of 10

    PoseBH unifies pose estimation across human, whole-body, and animal skeletons using nonparametric keypoint prototypes and cross-type self-supervision, improving animal-dataset accuracy while preserving human benchmark...

  3. Adept: Annotation-Denoising Auxiliary Tasks with Discrete Cosine Transform Map and Keypoint for Human-Centric Pretraining

    cs.CV 2025-04 conditional novelty 6.0 of 10

    Adept pretrains visual backbones on RGB images using DCT maps and keypoints as denoising auxiliary targets, improving five downstream human-centric tasks without depth data.

  4. BlanketGen2-Fit3D: Synthetic Blanket Augmentation Towards Improving Real-World In-Bed Blanket Occluded Human Pose Estimation

    cs.CV 2025-01 conditional novelty 6.0 of 10

    Synthetic blanket augmentation of Fit3D improves ViTPose-B pose estimation on real blanket-occluded in-bed images by an absolute 2.3% PCK.

  5. GenHMR: Generative Human Mesh Recovery

    cs.CV 2024-12 conditional novelty 6.0 of 10

    GenHMR applies masked generative token prediction and 2D-pose-guided latent refinement to monocular human mesh recovery, reporting state-of-the-art MPJPE on Human3.6M, 3DPW, and EMDB.

  6. Detection, Pose Estimation and Segmentation for Multiple Bodies: Closing the Virtuous Circle

    cs.CV 2024-12 conditional novelty 6.0 of 10

    An iterative detector-mask-pose loop with a mask-conditioned pose model, MaskPose, sets new state-of-the-art results on OCHuman while matching top-down COCO pose accuracy.

  7. CameraHMR: Aligning People with Perspective

    cs.CV 2024-11 conditional novelty 6.0 of 10

    CameraHMR predicts camera field of view from images of people and uses it to improve pseudo ground-truth training data, achieving state-of-the-art 3D human pose and shape accuracy.

  8. Image Captioning with Very Scarce Supervised Data: Adversarial Semi-Supervised Learning Approach

    cs.CV 2019-09 conditional novelty 6.0 of 10

    An adversarial semi-supervised framework with GAN-based pseudo-label retrieval and confidence weighting substantially improves image captioning when only 1% of paired data is available.

  9. Deep High-Resolution Representation Learning for Visual Recognition

    cs.CV 2019-08 conditional novelty 6.0 of 10

    HRNet maintains high-resolution feature maps in parallel with low-resolution streams and repeatedly fuses them, improving accuracy on pose, segmentation, detection, and face alignment benchmarks.

  10. Human Mesh Recovery from Monocular Images via a Skeleton-disentangled Representation

    cs.CV 2019-08 conditional novelty 6.0 of 10

    A skeleton-disentangled representation with a self-attention temporal network achieves state-of-the-art 3D human mesh recovery on Human3.6M and 3DPW.

  11. Unpaired Cross-lingual Image Caption Generation with Self-Supervised Rewards

    cs.CV 2019-08 conditional novelty 6.0 of 10

    A self-supervised rewarding framework using fluency and multi-level visual relevance rewards improves unpaired cross-lingual image captioning.

  12. Empowering Multimodal LLMs with External Tools: A Comprehensive Survey

    cs.CV 2025-08 unverdicted novelty 2.0 of 10

    A survey paper maps how external tools are used to augment multimodal large language models across data, tasks, evaluation, and future directions.

Pith tools