Pith. sign in

REVIEW 32 cited by

ImageNet-21K Pretraining for the Masses

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2104.10972 v4 pith:522CHFUV submitted 2021-04-22 cs.CV cs.LG

classification cs.CVcs.LG
keywords pretrainingimagenet-21kmodelsavailabledatasetefficienttaskstraining
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

ImageNet-1K serves as the primary dataset for pretraining deep learning models for computer vision tasks. ImageNet-21K dataset, which is bigger and more diverse, is used less frequently for pretraining, mainly due to its complexity, low accessibility, and underestimation of its added value. This paper aims to close this gap, and make high-quality efficient pretraining on ImageNet-21K available for everyone. Via a dedicated preprocessing stage, utilization of WordNet hierarchical structure, and a novel training scheme called semantic softmax, we show that various models significantly benefit from ImageNet-21K pretraining on numerous datasets and tasks, including small mobile-oriented models. We also show that we outperform previous ImageNet-21K pretraining schemes for prominent new models like ViT and Mixer. Our proposed pretraining pipeline is efficient, accessible, and leads to SoTA reproducible results, from a publicly available dataset. The training code and pretrained models are available at: https://github.com/Alibaba-MIIL/ImageNet21K

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 32 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. GVCCS: A Dataset for Contrail Identification and Tracking on Visible Whole Sky Camera Sequences

    cs.CV 2025-07 conditional novelty 7.0 of 10

    GVCCS is the first open dataset of ground-based visible all-sky camera video with instance-level contrail masks, temporal tracking, and flight IDs, plus Mask2Former baselines.

  2. When Model Knowledge meets Diffusion Model: Diffusion-assisted Data-free Image Synthesis with Alignment of Domain and Class

    cs.CV 2025-06 conditional novelty 7.0 of 10

    DDIS generates training-like images from a frozen classifier by steering Stable Diffusion with batch-normalization statistics and an optimized per-class token, improving data-free distillation and pruning.

  3. The Edge-on Galaxies in the DESI survey (EGIDE): sample building and photometry

    astro-ph.GA 2026-06 unverdicted novelty 6.0 of 10

    The EGIDE project releases a tenfold larger catalogue of edge-on galaxies with griz photometry, stellar masses, redshifts and star formation rates, finding that red-sequence galaxies are thicker than blue-cloud ones a...

  4. Towards Continuous Home Cage Monitoring: An Evaluation of Tracking and Identification Strategies for Laboratory Mice

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A real-time mouse tracking and ear-tag identity pipeline reports 95.28% identification accuracy and fewer ID switches than SLEAP and DeepLabCut on a 100-minute home-cage dataset.

  5. Opto-ViT: Architecting a Near-Sensor Region of Interest-Aware Vision Transformer Accelerator with Silicon Photonics

    cs.AR 2025-07 conditional novelty 6.0 of 10

    A near-sensor vision transformer accelerator combines VCSEL-microring photonic matrix multiplication with region-of-interest patch pruning, reporting 100.4 KFPS/W and up to 84% energy savings.

  6. Revisiting Audio-Visual Segmentation with Vision-Centric Transformer

    cs.CV 2025-06 conditional novelty 6.0 of 10

    Vision-derived object queries with prototype prompting achieve new state-of-the-art results on AVSBench audio-visual segmentation.

  7. Learning Along the Arrow of Time: Hyperbolic Geometry for Backward-Compatible Representation Learning

    cs.LG 2025-06 conditional novelty 6.0 of 10

    HBCT lifts embeddings into Lorentz hyperbolic space, uses entailment cones to keep new embeddings inside old ones' cones, and weights contrastive alignment by an uncertainty estimate, improving backward-compatible ret...

  8. SNAP: A Benchmark for Testing the Effects of Capture Conditions on Fundamental Vision Tasks

    cs.CV 2025-05 conditional novelty 6.0 of 10

    A new benchmark shows that camera capture settings and lighting systematically change the performance of image classifiers, object detectors, and VQA models, and that common vision datasets are biased toward narrow ex...

  9. UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation

    cs.CV 2025-05 conditional novelty 6.0 of 10

    UniGen shows a 1.5B model trained on open data can beat larger systems on image understanding and generation once it verifies its own outputs with chain-of-thought and Best-of-N selection.

  10. WILD: a new in-the-Wild Image Linkage Dataset for synthetic image attribution

    cs.MM 2025-04 conditional novelty 6.0 of 10

    WILD is a new 20,000-image benchmark pairing 10 known and 10 unknown generators, prompt-controlled closed set, post-processing chains, and baseline attribution results.

  11. POET: Prompt Offset Tuning for Continual Human Action Adaptation

    cs.CV 2025-04 conditional novelty 6.0 of 10

    POET, a prompt-offset tuning method for frozen graph neural networks, enables few-shot privacy-aware continual action recognition and outperforms adapted baselines on NTU RGB+D and SHREC-2017.

  12. LoRA-Based Continual Learning with Constraints on Critical Parameter Changes

    cs.CV 2025-04 conditional novelty 6.0 of 10

    Freezing the most important ViT parameter matrices before each new task, on top of orthogonal LoRA composition, reduces forgetting and improves average accuracy in class-incremental learning.

  13. Avoiding spurious sharpness minimization broadens applicability of SAM

    cs.LG 2025-02 conditional novelty 6.0 of 10

    SAM's failure in language modeling is traced to a dominant 'logit path' that minimizes sharpness spuriously, and the proposed Functional-SAM, which removes that path, improves validation loss over AdamW and SAM.

  14. A Room to Roam: Reset Prediction Based on Physical Object Placement for Redirected Walking

    cs.HC 2024-12 conditional novelty 6.0 of 10

    A Vision Transformer predicts the number of redirected-walking reset events from a top-down occupancy image of a room, achieving R-squared 0.91 in simulation, and powers a real-time furniture-placement interface.

  15. What makes a good metric? Evaluating automatic metrics for text-to-image consistency

    cs.CL 2024-12 conditional novelty 6.0 of 10

    None of the four tested text-to-image consistency metrics satisfies all proposed validity criteria, and the VQA-based metrics appear to rely largely on text priors such as yes-bias.

  16. EFTViT: Efficient Federated Training of Vision Transformers with Masked Images on Resource-Constrained Clients

    cs.CV 2024-11 conditional novelty 6.0 of 10

    Masking 75% of image patches during client-side training and moving deep layers to the server yields faster, cheaper federated ViT training with modest accuracy gains.

  17. Multimodal Autoregressive Pre-training of Large Vision Encoders

    cs.CV 2024-11 conditional novelty 6.0 of 10

    AIMV2 pre-trains vision encoders by autoregressively predicting both image patches and text tokens, beating CLIP and SigLIP on many recognition and multimodal benchmarks.

  18. Harnessing Vision Foundation Models for High-Performance, Training-Free Open Vocabulary Segmentation

    cs.CV 2024-11 conditional novelty 6.0 of 10

    Trident, a training-free framework combining CLIP, DINO, and SAM, raises state-of-the-art open-vocabulary segmentation mIoU from 44.4 to 48.6 by splicing sub-image features and aggregating them with a SAM affinity matrix.

  19. H3Former: Hypergraph-based Semantic-Aware Aggregation via Hyperbolic Hierarchical Contrastive Loss for Fine-Grained Visual Classification

    cs.CV 2025-11 conditional novelty 5.0 of 10

    A hypergraph-based token-to-region aggregation plus a hyperbolic hierarchical contrastive loss yields reported state-of-the-art fine-grained classification accuracy on four benchmarks.

  20. Revisiting Deepfake Detection: Chronological Continual Learning and the Limits of Generalization

    cs.LG 2025-08 conditional novelty 5.0 of 10

    A chronological continual learning study finds deepfake detectors retain past knowledge but generalize to future generators at near-random AUC around 0.5.

  21. Smelly, dense, and spreaded: The Object Detection for Olfactory References (ODOR) dataset

    cs.CV 2025-07 conditional novelty 5.0 of 10

    The ODOR dataset contributes 38,116 fine-grained object annotations over 4,712 artworks, benchmarked with five detector families, to stress-test object detection on dense, occluded, and off-centre objects in historica...

  22. Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations

    cs.CV 2025-06 conditional novelty 5.0 of 10

    Images are encoded into discrete tokens projected from LLM embeddings, so a single autoregressive model does visual understanding and generation with matched or improved benchmark scores.

  23. RollingQ: Reviving the Cooperation Dynamics in Multimodal Transformer

    cs.LG 2025-06 conditional novelty 5.0 of 10

    RollingQ rotates the classification query in a multimodal Transformer toward a rebalanced direction so attention stops over-favoring a single modality, restoring dynamic fusion and improving accuracy.

  24. PAID: Pairwise Angular-Invariant Decomposition for Continual Test-Time Adaptation

    cs.CV 2025-06 reject novelty 5.0 of 10

    PAID proposes Householder-based orthogonal weight updates for continual test-time adaptation, claiming that preserving pairwise angular structure of pretrained weights is a useful prior, but the math and validation fo...

  25. Eve: Efficient Multimodal Vision Language Models with Elastic Visual Experts

    cs.CV 2025-01 conditional novelty 5.0 of 10

    A 1.8B-parameter vision-language model, Eve, uses elastic visual experts and type-aware token routing to reach a 68.87% average on six VLM benchmarks while preserving language performance.

  26. MoRe: Class Patch Attention Needs Regularization for Weakly Supervised Semantic Segmentation

    cs.CV 2024-12 conditional novelty 5.0 of 10

    MoRe regularizes class-patch attention with a directed graph module and a CAM-informed contrastive loss, improving weakly supervised semantic segmentation.

  27. CascadeFormer: A Family of Two-stage Cascading Transformers for Skeleton-based Human Action Recognition

    cs.CV 2025-08 conditional novelty 4.0 of 10

    A masked-pretrained skeleton transformer with a second fine-tuning transformer and cross-attention fusion reaches 94.66% on Penn Action, 91.16% on N-UCLA, and 81.01%/88.17% on NTU RGB+D 60 cross-subject/cross-view.

  28. SimMIL: A Universal Weakly Supervised Pre-Training Framework for Multi-Instance Learning in Whole Slide Pathology Images

    cs.CV 2025-05 conditional novelty 4.0 of 10

    A weakly supervised pre-training scheme that propagates bag labels to patches improves downstream MIL classification and survival prediction on WSI datasets, but the comparison baselines are not trained on the same ta...

  29. Generalized Single-Image-Based Morphing Attack Detection Using Deep Representations from Vision Transformer

    cs.CV 2025-01 reject novelty 4.0 of 10

    A frozen ImageNet-pretrained Vision Transformer plus a linear classifier improves cross-algorithm morphing detection on digital face images, but not on print-scan images.

  30. Textile Analysis for Recycling Automation using Transfer Learning and Zero-Shot Foundation Models

    cs.CV 2025-06 reject novelty 3.0 of 10

    An RGB-based computer vision pipeline for textile recycling achieves 81.25% accuracy on four fabric classes and 0.90 mIoU when segmenting buttons and zippers with zero-shot foundation models.

  31. Do Language Models Understand Time?

    cs.CV 2024-12 conditional novelty 3.0 of 10

    A survey arguing that video-LLMs rely on pretrained encoders and short-biased datasets, leaving them weak at long-term temporal reasoning such as causality and event progression.

  32. ViT-ProtoNet for Few-Shot Image Classification: A Multi-Benchmark Evaluation

    cs.CV 2025-07 reject novelty 2.0 of 10

    ViT-ProtoNet, a Prototypical Network with a ViT-Small encoder, is reported to reach 95-97% 5-shot accuracy on three benchmarks and 81.88% on FC100, but the evaluation lacks critical baselines.

Pith tools