Pith. sign in

REVIEW 33 cited by

A ConvNet for the 2020s

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2201.03545 v2 pith:KOEQNG3K submitted 2022-01-10 cs.CV

A ConvNet for the 2020s

classification cs.CV
keywords transformersconvnetvisionstandardaccuracyconvnetsdesigndetection
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

The "Roaring 20s" of visual recognition began with the introduction of Vision Transformers (ViTs), which quickly superseded ConvNets as the state-of-the-art image classification model. A vanilla ViT, on the other hand, faces difficulties when applied to general computer vision tasks such as object detection and semantic segmentation. It is the hierarchical Transformers (e.g., Swin Transformers) that reintroduced several ConvNet priors, making Transformers practically viable as a generic vision backbone and demonstrating remarkable performance on a wide variety of vision tasks. However, the effectiveness of such hybrid approaches is still largely credited to the intrinsic superiority of Transformers, rather than the inherent inductive biases of convolutions. In this work, we reexamine the design spaces and test the limits of what a pure ConvNet can achieve. We gradually "modernize" a standard ResNet toward the design of a vision Transformer, and discover several key components that contribute to the performance difference along the way. The outcome of this exploration is a family of pure ConvNet models dubbed ConvNeXt. Constructed entirely from standard ConvNet modules, ConvNeXts compete favorably with Transformers in terms of accuracy and scalability, achieving 87.8% ImageNet top-1 accuracy and outperforming Swin Transformers on COCO detection and ADE20K segmentation, while maintaining the simplicity and efficiency of standard ConvNets.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 33 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. AIMIP Phase 1: systematic evaluations of AI weather and climate models

    physics.ao-ph 2026-05 accept novelty 7.0

    AI atmosphere models trained only on ERA5 reproduce historical climate and ENSO as well as GFDL-CM4, but several under-estimate warming trends and produce implausible responses to uniform +2 K/+4 K SST perturbations.

  2. AIMIP Phase 1: systematic evaluations of AI weather and climate models

    physics.ao-ph 2026-05 conditional novelty 7.0

    Under one protocol, most AI climate models reproduce historical climatology and ENSO response as well as a CMIP6 model, but some underestimate warming trends and all diverge on +2/+4K SST experiments.

  3. Generative diffusion models for spatiotemporal influenza forecasting

    cs.LG 2026-04 unverdicted novelty 7.0

    Influpaint uses generative diffusion models on image-encoded influenza data to produce realistic and diverse epidemic trajectories that match leading ensemble methods in accuracy.

  4. Descriptor: Parasitoid Wasps and Associated Hymenoptera Dataset (DAPWH)

    cs.CV 2026-02 unverdicted novelty 7.0

    Releases the DAPWH dataset of 3556 wasp images including 1739 COCO-annotated examples to enable AI models for identifying Ichneumonoidea and associated families.

  5. ROMS-IMLE: A Minimalist Approach to Competitive Single-Step Generative Modelling

    cs.LG 2026-07 conditional novelty 6.0

    A single-step IMLE generator with per-stage supervision and a robust loss reports FID 2.56 on ImageNet-256 by filtering ~5% of samples at test time.

  6. C3DIR: A Deep Learning 3-Dimensional Cloud Property Retrieval Scheme for Passive Satellite Imagers

    physics.ao-ph 2026-07 conditional novelty 6.0

    C3DIR is a single multi-sensor deep-learning model that retrieves 3-D ice, liquid, and rain water content from passive imagers using voxel-to-voxel collocation with active-sensor profiles.

  7. Continuously Evolving Deepfake Detection: An Architecture and Public-Benchmark Evaluation of a Dynamic Detection System

    cs.CV 2026-07 conditional novelty 6.0

    A continuously refreshed, incentive-driven deepfake detector beats static detectors on in-the-wild benchmarks and improves on post-export AI-generated media.

  8. ScaleAware-JEPA: Latent Representation for Discovery in Multiscale Physical Fields

    cs.LG 2026-06 unverdicted novelty 6.0

    ScaleAware-JEPA combines Constrained Diffusion Decomposition with a scale-tied JEPA objective to learn label-free latent coordinates that recover coherent morphology in multiscale fields such as MHD turbulence and int...

  9. Modeling Local, Global, and Cross-Modal Context in Multimodal 3D MRI

    cs.CV 2026-06 unverdicted novelty 6.0

    MICViT outperforms CNN and transformer baselines on brain age prediction from multimodal 3D MRI by combining modality-specific and cross-modal local/global attention across three heterogeneous datasets.

  10. TimeProVe: Propose, then Verify for Efficient Long Video Temporal Reasoning in Activities of Daily Living

    cs.CV 2026-06 unverdicted novelty 6.0

    TimeProVe proposes a propose-then-verify framework using lightweight action-based candidate evidence generation followed by targeted VLM verification for efficient long video temporal reasoning, achieving 7.3% improve...

  11. Spiral arms across cosmic time: JWST measurements of the pitch angles of spiral galaxies at $z<3.5$

    astro-ph.GA 2026-06 unverdicted novelty 6.0

    JWST measurements of pitch angles in 593 spiral galaxies to z=3.5 show no overall redshift evolution but reveal correlations with mass and sSFR only below z=1.25, implying a transition from locally driven to globally ...

  12. Generating synthetic computed tomography for radiotherapy: SynthRAD2025 challenge report

    physics.med-ph 2026-05 accept novelty 6.0

    SynthRAD2025 shows deep learning produces synthetic CTs with MAE 48-65 HU and high dosimetric gamma passing rates for radiotherapy, performing better on CBCT-to-CT than MRI-to-CT tasks.

  13. Birds of a Feather Flock Together: Background-Invariant Representations via Linear Structure in VLMs

    cs.CV 2026-05 unverdicted novelty 6.0

    Exploiting linear structure in VLM embeddings, a synthetic-data pre-training method yields background-invariant representations that exceed 90% worst-group accuracy on Waterbirds even under 100% spurious correlation w...

  14. AIMIP Phase 1: systematic evaluations of AI weather and climate models

    physics.ao-ph 2026-05 unverdicted novelty 6.0

    AIMIP Phase 1 sets up a common experiment and five evaluation criteria for AI atmosphere models forced by historical sea surface temperatures, finding they match conventional models on most metrics but underestimate s...

  15. AIMIP Phase 1: systematic evaluations of AI weather and climate models

    physics.ao-ph 2026-05 unverdicted novelty 6.0

    AIMIP Phase 1 shows AI models simulate historical climate and El Niño responses as well as traditional models, though some underestimate trends and diverge in generalization tests, with a public dataset released for f...

  16. A Novel Graph-Regulated Disentangling Mamba Model with Sparse Tokens for Enhanced Tree Species Classification from MODIS Time Series

    cs.CV 2026-05 unverdicted novelty 6.0

    A graph-regulated disentangling Mamba model with sparse tokens achieves 93.94% accuracy classifying tree species from MODIS time series in Alberta and outperforms twelve prior models.

  17. Only Brains Align with Brains: Cross-Region Alignment Patterns Expose Limits of Normative Models

    q-bio.NC 2026-04 unverdicted novelty 6.0

    Alignment pattern analysis reveals that models aligned to individual brain ROIs do not reproduce the stable cross-region alignment profiles observed across human subjects.

  18. Towards Real-Time ECG and EMG Modeling on $\mu$NPUs

    cs.LG 2026-04 unverdicted novelty 6.0

    PhysioLite delivers Transformer-comparable ECG/EMG performance using learnable wavelet filters and hardware-aware design at ~370KB quantized size on μNPUs.

  19. Geographically-Weighted Weakly Supervised Bayesian High-Resolution Transformer for 200m Resolution Pan-Arctic Sea Ice Concentration Mapping and Uncertainty Estimation using Sentinel-1, RCM, and AMSR2 Data

    cs.CV 2026-03 conditional novelty 6.0

    A 200m pan-Arctic sea ice concentration model using a Bayesian Transformer with a geographically-weighted weakly supervised loss; reported 0.70 feature detection accuracy and R²=0.90 vs ASI.

  20. Model soups need only one ingredient

    cs.LG 2026-02 conditional novelty 6.0

    A single checkpoint, edited by splitting each layer's update with SVD and reweighting the high- and low-energy parts, reaches soup-level OOD robustness without multi-model training.

  21. Small, Bias-Free, Blind and Convolutional Denoiser: A compact ConvNeXt U-Net for blind Gaussian color-image denoising

    eess.IV 2026-07 conditional novelty 5.0

    BF-ConvUNeXt, a 0.82M-parameter bias-free ConvNeXt U-Net, is degree-1 homogeneous and matches DnCNN/FFDNet in blind color denoising, extrapolating smoothly beyond its training noise range.

  22. Leveraging Multimodality for Real-Time Classification of Transients and Variables found by the Zwicky Transient Facility

    astro-ph.IM 2026-06 unverdicted novelty 5.0

    ORACLE-2 multimodal classifiers raise macro F1 from 0.52-0.66 (light-curve only) to 0.73 on ZTF Bright Transient Survey data and reach 0.88 on simulated ELAsTiCC data.

  23. SPADE: Sketch-guided Path Planning Augmented with Diffusion Experts

    cs.RO 2026-06 unverdicted novelty 5.0

    SPADE combines sketch-guided path planning with diffusion-augmented imitation learning to achieve better generalization and lower error with fewer parameters than prior methods.

  24. Hist2Style: Histogram-Guided Stylization with Bilateral Grids

    cs.CV 2026-06 unverdicted novelty 5.0

    Hist2Style introduces a lightweight bilateral-grid network conditioned on histogram embeddings for distilling large-model stylization into real-time, structure-preserving, user-controllable photorealistic edits.

  25. Physics-informed convolutional neural networks for fluid flow through porous media

    cs.LG 2026-05 unverdicted novelty 5.0

    A physics-informed CNN predicts pore-scale velocity fields from geometry and serves as a warm-start to accelerate Lattice-Boltzmann solvers in over 90% of tested cases.

  26. Federated Medical Image Classification under Class and Domain Imbalance exploiting Synthetic Sample Generation

    cs.CV 2026-04 unverdicted novelty 5.0

    FedSSG generates and shares synthetic samples within a federated setup to reduce class imbalance and domain shift problems in medical image classification.

  27. HLGFA: High-Low Resolution Guided Feature Alignment for Unsupervised Anomaly Detection

    cs.CV 2026-02 unverdicted novelty 5.0

    HLGFA detects anomalies by identifying breakdowns in cross-resolution feature consistency between high- and low-resolution views of normal samples, guided by structure and detail priors, and reports 97.9% pixel AUROC ...

  28. InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks

    cs.CV 2023-12 unverdicted novelty 5.0

    InternVL scales a vision model to 6B parameters and aligns it with LLMs using web data to achieve state-of-the-art results on 32 visual-linguistic benchmarks.

  29. A Controlled Visual-Backbone Benchmark for Multimodal Short-Term Solar Irradiance Forecasting

    eess.IV 2026-07 conditional novelty 4.0

    Under one fixed multimodal 10-min irradiance forecaster, VMamba-S and Swin-B nearly tie best Folsom RMSE (~65.4 W/m²) over smart persistence, while NREL’s tiny matched split favors persistence.

  30. Multi-Task Learning for Heterogeneous Prediction from Video Game State with Transfer Learning

    cs.LG 2026-07 conditional novelty 4.0

    On a large World of Tanks dataset, a shared multi-task model with equal weighting or PCGrad outperforms single-task models on average, and task/map pre-training helps most in low-data regimes.

  31. LETT-NeXt: A Lightweight RECIST-Guided Model for 3D CT Lesion Segmentation

    cs.CV 2026-06 unverdicted novelty 4.0

    LETT-NeXt uses RECIST line prompts in a cropped MedNeXt-v2 encoder-decoder to predict 3D lesion masks, reaching DSC 73.9 on hidden test data for a CVPR 2026 segmentation competition.

  32. NTIRE 2026 The Second Challenge on Day and Night Raindrop Removal for Dual-Focused Images: Methods and Results

    cs.CV 2026-04 unverdicted novelty 2.0

    The NTIRE 2026 challenge reports strong performance from 17 teams on raindrop removal for dual-focused day and night images using an adjusted real-world dataset with 14,139 training images.

  33. NTIRE 2026 The Second Challenge on Day and Night Raindrop Removal for Dual-Focused Images: Methods and Results

    cs.CV 2026-04 unverdicted novelty 2.0

    The second NTIRE challenge on day and night raindrop removal for dual-focused images received 17 valid team submissions that demonstrated strong performance on the Raindrop Clarity dataset.