REVIEW 12 cited by
AI Challenger : A Large-scale Dataset for Going Deeper in Image Understanding
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Significant progress has been achieved in Computer Vision by leveraging large-scale image datasets. However, large-scale datasets for complex Computer Vision tasks beyond classification are still limited. This paper proposed a large-scale dataset named AIC (AI Challenger) with three sub-datasets, human keypoint detection (HKD), large-scale attribute dataset (LAD) and image Chinese captioning (ICC). In this dataset, we annotate class labels (LAD), keypoint coordinate (HKD), bounding box (HKD and LAD), attribute (LAD) and caption (ICC). These rich annotations bridge the semantic gap between low-level images and high-level concepts. The proposed dataset is an effective benchmark to evaluate and improve different computational methods. In addition, for related tasks, others can also use our dataset as a new resource to pre-train their models.
Forward citations
Cited by 12 Pith papers
-
MetricHMSR:Metric Human Mesh and Scene Recovery from Monocular Images
From one image, MetricHMSR jointly estimates a metric human mesh, its global 3D position, and a corrected metric scene depth map.
-
PoseBH: Prototypical Multi-Dataset Training Beyond Human Pose Estimation
PoseBH unifies pose estimation across human, whole-body, and animal skeletons using nonparametric keypoint prototypes and cross-type self-supervision, improving animal-dataset accuracy while preserving human benchmark...
-
Adept: Annotation-Denoising Auxiliary Tasks with Discrete Cosine Transform Map and Keypoint for Human-Centric Pretraining
Adept pretrains visual backbones on RGB images using DCT maps and keypoints as denoising auxiliary targets, improving five downstream human-centric tasks without depth data.
-
BlanketGen2-Fit3D: Synthetic Blanket Augmentation Towards Improving Real-World In-Bed Blanket Occluded Human Pose Estimation
Synthetic blanket augmentation of Fit3D improves ViTPose-B pose estimation on real blanket-occluded in-bed images by an absolute 2.3% PCK.
-
GenHMR: Generative Human Mesh Recovery
GenHMR applies masked generative token prediction and 2D-pose-guided latent refinement to monocular human mesh recovery, reporting state-of-the-art MPJPE on Human3.6M, 3DPW, and EMDB.
-
Detection, Pose Estimation and Segmentation for Multiple Bodies: Closing the Virtuous Circle
An iterative detector-mask-pose loop with a mask-conditioned pose model, MaskPose, sets new state-of-the-art results on OCHuman while matching top-down COCO pose accuracy.
-
CameraHMR: Aligning People with Perspective
CameraHMR predicts camera field of view from images of people and uses it to improve pseudo ground-truth training data, achieving state-of-the-art 3D human pose and shape accuracy.
-
Image Captioning with Very Scarce Supervised Data: Adversarial Semi-Supervised Learning Approach
An adversarial semi-supervised framework with GAN-based pseudo-label retrieval and confidence weighting substantially improves image captioning when only 1% of paired data is available.
-
Deep High-Resolution Representation Learning for Visual Recognition
HRNet maintains high-resolution feature maps in parallel with low-resolution streams and repeatedly fuses them, improving accuracy on pose, segmentation, detection, and face alignment benchmarks.
-
Human Mesh Recovery from Monocular Images via a Skeleton-disentangled Representation
A skeleton-disentangled representation with a self-attention temporal network achieves state-of-the-art 3D human mesh recovery on Human3.6M and 3DPW.
-
Unpaired Cross-lingual Image Caption Generation with Self-Supervised Rewards
A self-supervised rewarding framework using fluency and multi-level visual relevance rewards improves unpaired cross-lingual image captioning.
-
Empowering Multimodal LLMs with External Tools: A Comprehensive Survey
A survey paper maps how external tools are used to augment multimodal large language models across data, tasks, evaluation, and future directions.
Discussion (0). Continue with ORCID to comment.