Pith. sign in

REVIEW 50 cited by

A ConvNet for the 2020s

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2201.03545 v2 pith:KOEQNG3K submitted 2022-01-10 cs.CV

classification cs.CV
keywords transformersconvnetvisionstandardaccuracyconvnetsdesigndetection
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

The "Roaring 20s" of visual recognition began with the introduction of Vision Transformers (ViTs), which quickly superseded ConvNets as the state-of-the-art image classification model. A vanilla ViT, on the other hand, faces difficulties when applied to general computer vision tasks such as object detection and semantic segmentation. It is the hierarchical Transformers (e.g., Swin Transformers) that reintroduced several ConvNet priors, making Transformers practically viable as a generic vision backbone and demonstrating remarkable performance on a wide variety of vision tasks. However, the effectiveness of such hybrid approaches is still largely credited to the intrinsic superiority of Transformers, rather than the inherent inductive biases of convolutions. In this work, we reexamine the design spaces and test the limits of what a pure ConvNet can achieve. We gradually "modernize" a standard ResNet toward the design of a vision Transformer, and discover several key components that contribute to the performance difference along the way. The outcome of this exploration is a family of pure ConvNet models dubbed ConvNeXt. Constructed entirely from standard ConvNet modules, ConvNeXts compete favorably with Transformers in terms of accuracy and scalability, achieving 87.8% ImageNet top-1 accuracy and outperforming Swin Transformers on COCO detection and ADE20K segmentation, while maintaining the simplicity and efficiency of standard ConvNets.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 50 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 178 citations worldwide. Full citation record

  1. AIMIP Phase 1: systematic evaluations of AI weather and climate models

    physics.ao-ph 2026-05 unverdicted novelty 7.0 of 10

    Under one protocol, most AI climate models reproduce historical climatology and ENSO response as well as a CMIP6 model, but some underestimate warming trends and all diverge on +2/+4K SST experiments.

  2. FourCastNet 3: A geometric approach to probabilistic machine-learning weather forecasting at scale

    cs.LG 2025-07 conditional novelty 7.0 of 10

    A purely convolutional, spherical-geometry weather model trained with a combined spatial and spectral CRPS loss delivers GenCast-level skill, IFS-beating accuracy, and stable spectra out to 60 days.

  3. ROMS-IMLE: A Minimalist Approach to Competitive Single-Step Generative Modelling

    cs.LG 2026-07 conditional novelty 6.0 of 10

    A single-step IMLE generator with per-stage supervision and a robust loss reports FID 2.56 on ImageNet-256 by filtering ~5% of samples at test time.

  4. C3DIR: A Deep Learning 3-Dimensional Cloud Property Retrieval Scheme for Passive Satellite Imagers

    physics.ao-ph 2026-07 conditional novelty 6.0 of 10

    C3DIR is a single multi-sensor deep-learning model that retrieves 3-D ice, liquid, and rain water content from passive imagers using voxel-to-voxel collocation with active-sensor profiles.

  5. Continuously Evolving Deepfake Detection: An Architecture and Public-Benchmark Evaluation of a Dynamic Detection System

    cs.CV 2026-07 conditional novelty 6.0 of 10

    A continuously refreshed, incentive-driven deepfake detector beats static detectors on in-the-wild benchmarks and improves on post-export AI-generated media.

  6. Geographically-Weighted Weakly Supervised Bayesian High-Resolution Transformer for 200m Resolution Pan-Arctic Sea Ice Concentration Mapping and Uncertainty Estimation using Sentinel-1, RCM, and AMSR2 Data

    cs.CV 2026-03 conditional novelty 6.0 of 10

    A 200m pan-Arctic sea ice concentration model using a Bayesian Transformer with a geographically-weighted weakly supervised loss; reported 0.70 feature detection accuracy and R²=0.90 vs ASI.

  7. Model soups need only one ingredient

    cs.LG 2026-02 conditional novelty 6.0 of 10

    A single checkpoint, edited by splitting each layer's update with SVD and reweighting the high- and low-energy parts, reaches soup-level OOD robustness without multi-model training.

  8. Khana: A Comprehensive Indian Cuisine Dataset

    cs.CV 2025-09 conditional novelty 6.0 of 10

    Khana is an Indian cuisine image dataset with 131K images across 80 dish categories and baseline classification accuracies up to 86.7% top-1 with ConvNeXT.

  9. RRTO: A High-Performance Transparent Offloading System for Model Inference in Mobile Edge Computing

    cs.NI 2025-07 conditional novelty 6.0 of 10

    RRTO identifies static inference operator sequences from CUDA call logs alone and replays them on an edge GPU, cutting transparent-offloading communication to 11 RPCs per inference instead of thousands, with performan...

  10. Beginning with You: Perceptual-Initialization Improves Vision-Language Representation and Alignment

    cs.CV 2025-05 conditional novelty 6.0 of 10

    Perceptually initializing a CLIP vision encoder with NIGHTS triplet judgments before YFCC15M contrastive training improves zero-shot accuracy and retrieval over an identical random-start baseline.

  11. DiffEx: Explaining a Classifier with Diffusion Models to Identify Microscopic Cellular Variations

    cs.CV 2025-02 conditional novelty 6.0 of 10

    DiffEx builds a classifier-aware latent space with a diffusion model, finds contrastive directions in it, and ranks them to produce visual explanations that reveal cellular phenotype changes.

  12. Approximate Message Passing for Bayesian Neural Networks

    cs.LG 2025-01 conditional novelty 6.0 of 10

    A factor-graph message-passing method for Bayesian neural networks that handles CNNs, avoids double-counting, and shows competitive accuracy with improved calibration on CIFAR-10.

  13. A CNN-Transformer for Classification of Longitudinal 3D MRI Images -- A Case Study on Hepatocellular Carcinoma Prediction

    cs.CV 2025-01 reject novelty 6.0 of 10

    A CNN-Transformer trained on longitudinal 3D MRIs claims high accuracy for predicting next-scan hepatocellular carcinoma, but its time-aware positional encoding reveals the future diagnosis date to the model.

  14. Mantis Shrimp: Exploring Photometric Band Utilization in Computer Vision Networks for Photometric Redshift Estimation

    astro-ph.IM 2025-01 conditional novelty 6.0 of 10

    A multi-survey CNN estimates photometric redshifts from GALEX, PanSTARRS, and UnWISE cutouts, with early and late image fusion performing comparably.

  15. Err on the Side of Texture: Texture Bias on Real Data

    cs.CV 2024-12 conditional novelty 6.0 of 10

    A new metric and texture-identification method show that ImageNet classifiers rely heavily on specific textures, but the claim that texture bias explains natural adversarial examples is largely a consequence of how te...

  16. Ranking-aware adapter for text-driven image ordering with CLIP

    cs.CV 2024-12 conditional novelty 6.0 of 10

    A frozen CLIP plus a small ranking-aware adapter can order images by arbitrary text-specified attributes, using pairwise visual differences as supervision.

  17. Image Generation Diversity Issues and How to Tame Them

    cs.CV 2024-11 conditional novelty 6.0 of 10

    The paper proposes a retrieval-based diversity metric (IRS), finds that state-of-the-art diffusion models retrieve at most 77% of training images, and introduces feature-conditioned DiADM to improve unconditional diversity.

  18. Reliable Evaluation of Attribution Maps in CNNs: A Perturbation-Based Approach

    cs.CV 2024-11 conditional novelty 6.0 of 10

    A perturbation-based metric using FGSM flips of ±1/255 instead of zero-masking gives more consistent and monotonic evaluation of attribution maps across 15 CNN-dataset pairs, with SmoothGrad ranked first.

  19. Introducing Milabench: Benchmarking Accelerators for AI

    cs.LG 2024-11 conditional novelty 6.0 of 10

    Mila releases Milabench, a 42-benchmark open-source suite, and shows that for real AI workloads the NVIDIA H100 generally outperforms AMD MI300X and Intel Gaudi2 despite MI300X's high synthetic FLOP counts.

  20. Small, Bias-Free, Blind and Convolutional Denoiser: A compact ConvNeXt U-Net for blind Gaussian color-image denoising

    eess.IV 2026-07 conditional novelty 5.0 of 10

    BF-ConvUNeXt, a 0.82M-parameter bias-free ConvNeXt U-Net, is degree-1 homogeneous and matches DnCNN/FFDNet in blind color denoising, extrapolating smoothly beyond its training noise range.

  21. REVELIO -- Universal Multimodal Task Load Estimation for Cross-Domain Generalization

    cs.LG 2025-09 conditional novelty 5.0 of 10

    A new cognitive-load dataset with driving, gaming, and n-back tasks shows multimodal models beat unimodal ones, but cross-domain transfer remains poor.

  22. StixelNExT++: Lightweight Monocular Scene Segmentation and Representation for Collective Perception

    cs.CV 2025-07 conditional novelty 5.0 of 10

    A monocular neural network predicts 3D Stixels directly from RGB images in about 10 ms, with a self-defined Waymo evaluation showing competitive performance within 30 m.

  23. Advancing Image-Based Grapevine Variety Classification with a New Benchmark and Evaluation of Masked Autoencoders

    cs.CV 2025-06 conditional novelty 5.0 of 10

    MAE self-supervised pre-training on 54,571 unlabeled grapevine images improves downstream 43-class variety classification (F1 0.7956) over ImageNet-initialized baselines, and the paper releases two multi-season labele...

  24. MIM: Multi-modal Content Interest Modeling Paradigm for User Behavior Modeling

    cs.IR 2025-02 conditional novelty 5.0 of 10

    MIM aligns multi-modal item embeddings with purchase-based user interest signals and combines them with ID-based collaborative filtering, reporting small offline AUC gains and large online CTR and RPM gains at Taobao.

  25. Geometry Matters: Benchmarking Scientific ML Approaches for Flow Prediction around Complex Geometries

    cs.LG 2024-12 conditional novelty 5.0 of 10

    On the FlowBench lid-driven cavity benchmark, vision-transformer foundation models outperform neural operators in data-limited regimes, but all models generalize poorly to out-of-range Reynolds numbers and geometry ge...

  26. ZoRI: Towards Discriminative Zero-Shot Remote Sensing Instance Segmentation

    cs.CV 2024-12 reject novelty 5.0 of 10

    ZoRI combines CLIP text-channel selection, partial fine-tuning, and a pseudo-label cache bank to segment unseen aerial classes, but the cache bank is seeded with the model's own test-set predictions.

  27. From Audio Deepfake Detection to AI-Generated Music Detection -- A Pathway and Overview

    cs.SD 2024-11 conditional novelty 5.0 of 10

    A review of AI-generated music detection that proposes intrinsic music features and multimodal fusion as the basis for adapting audio deepfake detection methods.

  28. Residual Vision Transformer (ResViT) Based Self-Supervised Learning Model for Brain Tumor Classification

    eess.IV 2024-11 conditional novelty 5.0 of 10

    A ResViT-based generative self-supervised pipeline, pretrained on MRI synthesis and fine-tuned for classification, achieves 90.56% on BraTS, 98.53% on Figshare, and 98.47% on Kaggle brain tumor datasets.

  29. WoodYOLO: A Novel Object Detector for Wood Species Detection in Microscopic Images

    cs.CV 2024-11 conditional novelty 5.0 of 10

    A customized YOLO detector, WoodYOLO, reports F2 0.848 at IoU 0.3 for vessel-element detection in wood microscopy, beating YOLOv10 and YOLOv7 on a private dataset.

  30. A Controlled Visual-Backbone Benchmark for Multimodal Short-Term Solar Irradiance Forecasting

    eess.IV 2026-07 conditional novelty 4.0 of 10

    Under one fixed multimodal 10-min irradiance forecaster, VMamba-S and Swin-B nearly tie best Folsom RMSE (~65.4 W/m²) over smart persistence, while NREL’s tiny matched split favors persistence.

  31. Multi-Task Learning for Heterogeneous Prediction from Video Game State with Transfer Learning

    cs.LG 2026-07 conditional novelty 4.0 of 10

    On a large World of Tanks dataset, a shared multi-task model with equal weighting or PCGrad outperforms single-task models on average, and task/map pre-training helps most in low-data regimes.

  32. Revisiting Simple Baselines for In-The-Wild Deepfake Detection

    cs.CV 2025-09 conditional novelty 4.0 of 10

    Finetuned CLIP-pretrained ConvNeXt-base and ViT-b32 classifiers reach 81% accuracy on Deepfake-Eval-2024, within noise of the leading commercial detector's 82%.

  33. Classifying Mitotic Figures in the MIDOG25 Challenge with Deep Ensemble Learning and Rule Based Refinement

    cs.CV 2025-08 conditional novelty 4.0 of 10

    A ConvNeXt ensemble achieves 84% balanced accuracy on MIDOG25 atypical mitotic figure classification, while a rule-based refinement module trades sensitivity for specificity.

  34. Beyond Linear Bottlenecks: Spline-Based Knowledge Distillation for Culturally Diverse Art Style Classification

    cs.CV 2025-07 reject novelty 4.0 of 10

    Replacing MLP projection heads with Kolmogorov-Arnold Network heads in a dual-teacher self-supervised art-style classifier yields Top-1 accuracy gains of around 0.2 to 1.0 percentage points on WikiArt and Pandora18k, ...

  35. Faithful, Interpretable Chest X-ray Diagnosis with Anti-Aliased B-cos Networks

    cs.CV 2025-07 conditional novelty 4.0 of 10

    Combining B-cos networks with anti-aliasing pooling (FLC or BlurPool) reduces grid artifacts in chest X-ray explanation maps while keeping diagnostic accuracy close to baseline networks.

  36. Applying multimodal learning to Classify transient Detections Early (AppleCiDEr) I: Data set, methods, and infrastructure

    astro-ph.IM 2025-07 conditional novelty 4.0 of 10

    AppleCiDEr combines photometry, images, metadata, and spectra in one deep learning pipeline to classify ZTF transients and variable stars, with high accuracy on common classes but poor performance on tidal disruption events.

  37. First-of-its-kind AI model for bioacoustic detection using a lightweight associative memory Hopfield neural network

    cs.LG 2025-07 reject novelty 4.0 of 10

    A Hopfield neural network trained on two bat calls classifies 10,384 recordings in 5.4 seconds with claimed accuracy up to 80%, though the headline numbers hinge on removing ambiguous calls.

  38. Bridging the gap in FER: addressing age bias in deep learning

    cs.CV 2025-07 conditional novelty 4.0 of 10

    Age-aware loss reweighting and age-conditioned training reduce elderly recognition errors in facial expression models, even when age labels are automatically estimated, but the evidence rests on a small elderly benchmark.

  39. SIM-Net: A Multimodal Fusion Network Using Inferred 3D Object Shape Point Clouds from RGB Images for 2D Classification

    cs.CV 2025-06 conditional novelty 4.0 of 10

    A dual-branch network that fuses ResNet image features with PointNet features over point clouds built from RGB values and pixel coordinates improves herbarium trait classification in most, but not all, of the tested settings.

  40. Comparison of ConvNeXt and Vision-Language Models for Breast Density Assessment in Screening Mammography

    eess.IV 2025-06 conditional novelty 4.0 of 10

    Fine-tuned ConvNeXt achieves 0.73 accuracy and 0.78 F1, beating BioMedCLIP linear probe (0.64/0.63) and zero-shot (0.47/0.31) for BI-RADS breast density classification.

  41. Predicting Genetic Mutations from Single-Cell Bone Marrow Images in Acute Myeloid Leukemia Using Noise-Robust Deep Learning Models

    eess.IV 2025-06 reject novelty 4.0 of 10

    A deep learning pipeline claims 85% accuracy for classifying four AML-associated mutations from single-cell blood and bone marrow images, but the evaluation has unresolved validity issues.

  42. CASE: Contrastive Activation for Saliency Estimation

    cs.CV 2025-06 conditional novelty 4.0 of 10

    CASE removes gradient components shared with confused classes to produce more class-distinct saliency maps, validated on a top-k overlap diagnostic where many existing methods show class-insensitive behavior.

  43. DAGNet: A Dual-View Attention-Guided Network for Efficient X-ray Security Inspection

    cs.CV 2025-02 conditional novelty 4.0 of 10

    DAGNet, a three-module architecture, improves multi-label contraband classification mAP on the DvXray dual-view X-ray dataset by about 1.5 to 2.4 points over AHCR and 2.5 to 4.8 points over the dual-view baseline.

  44. Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey

    cs.CL 2024-12 conditional novelty 4.0 of 10

    A comprehensive survey that frames multimodal understanding and generation as next token prediction and proposes a five-part taxonomy.

  45. Automated Fetal Biometry Assessment with Deep Ensembles using Sparse-Sampling of 2D Intrapartum Ultrasound Images

    eess.IV 2025-05 reject novelty 3.0 of 10

    A fetal biometry pipeline using sparse sampling and deep ensembles reports high accuracy, but its published measurement errors conflict with its own results tables.

  46. Thoughts on Objectives of Sparse and Hierarchical Masked Image Model

    eess.IV 2025-05 reject novelty 3.0 of 10

    A checkerboard-like 'mesh mask' for SparK pre-training ties the random mask at F1 87.7 on brain CT tumor classification, with no improvement.

  47. Scalable Whole Slide Image Representation Using K-Mean Clustering and Fisher Vector Aggregation

    cs.CV 2025-01 reject novelty 3.0 of 10

    A WSI classification pipeline that clusters patch embeddings with K-means and encodes each cluster with Fisher vectors, evaluated on four benchmark datasets.

  48. The Good, The Efficient and the Inductive Biases: Exploring Efficiency in Deep Learning Through the Use of Inductive Biases

    cs.LG 2024-11 conditional novelty 3.0 of 10

    A dissertation synthesizing the author's papers on continuous kernel convolutions and symmetry-preserving architectures, claiming these inductive biases improve deep learning efficiency.

  49. A Survey on Training-free Open-Vocabulary Semantic Segmentation

    cs.CV 2025-05 conditional novelty 2.0 of 10

    A structured review of over 30 training-free open-vocabulary semantic segmentation methods, organized by whether they rely on CLIP alone, auxiliary visual foundation models, or generative models.

  50. Multilabel Classification for Lung Disease Detection: Integrating Deep Learning and Natural Language Processing

    cs.CV 2024-12 reject novelty 2.0 of 10

    A transfer-learning benchmark on CheXpert chest X-rays reports AUROC 0.86 with ConvNeXt, but the claimed NLP integration is not demonstrated.

Pith tools