Pith. sign in

REVIEW 25 cited by

Adaptive Fourier Neural Operators: Efficient Token Mixers for Transformers

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2111.13587 v2 pith:TYPLYCY3 submitted 2021-11-24 cs.CV cs.LG

classification cs.CVcs.LG
keywords afnofourierlearningtokenefficientmixingadaptiveconvolution
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Vision transformers have delivered tremendous success in representation learning. This is primarily due to effective token mixing through self attention. However, this scales quadratically with the number of pixels, which becomes infeasible for high-resolution inputs. To cope with this challenge, we propose Adaptive Fourier Neural Operator (AFNO) as an efficient token mixer that learns to mix in the Fourier domain. AFNO is based on a principled foundation of operator learning which allows us to frame token mixing as a continuous global convolution without any dependence on the input resolution. This principle was previously used to design FNO, which solves global convolution efficiently in the Fourier domain and has shown promise in learning challenging PDEs. To handle challenges in visual representation learning such as discontinuities in images and high resolution inputs, we propose principled architectural modifications to FNO which results in memory and computational efficiency. This includes imposing a block-diagonal structure on the channel mixing weights, adaptively sharing weights across tokens, and sparsifying the frequency modes via soft-thresholding and shrinkage. The resulting model is highly parallel with a quasi-linear complexity and has linear memory in the sequence size. AFNO outperforms self-attention mechanisms for few-shot segmentation in terms of both efficiency and accuracy. For Cityscapes segmentation with the Segformer-B3 backbone, AFNO can handle a sequence size of 65k and outperforms other efficient self-attention mechanisms.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 25 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. From Points to Edges: Edge-Conditioned Spectral Operators for Physics-Sensitive PDE Learning

    cs.AI 2026-08 conditional novelty 6.0 of 10

    ESO conditions spectral mode mixing on local pairwise variation statistics, improving neural operator accuracy on PDEs with sharp local structures.

  2. Adaptive Mamba Neural Operators

    cs.LG 2026-07 reject novelty 6.0 of 10

    AMO builds adaptive Takenaka-Malmquist bases inside a Mamba state-space model for PDE operator learning, but the claimed equivalence to adaptive Fourier decomposition is not supported by the implemented recurrence.

  3. Practical Quantum Advantage before Fault Tolerance via Quantum-Informed Machine Learning

    quant-ph 2026-06 unverdicted novelty 6.0 of 10

    Higher-order Q-Priors plus two-copy Bell readout give a claimed copy-complexity separation for invariant-measure correlators of chaotic systems, with reported skill gains on turbulence and ERA5 weather tasks.

  4. Hierarchical Physics-Embedded Learning for Partially Known Spatiotemporal Dynamics

    cs.LG 2025-10 reject novelty 6.0 of 10

    A hierarchical physics-embedded Fourier neural operator reduces long-horizon prediction error on Cahn-Hilliard, Allen-Cahn, dKPZ and CGL benchmarks while enabling symbolic recovery of unknown constitutive terms.

  5. Fourier Neural Operators for Time-Periodic Quantum Systems: Learning Floquet Hamiltonians, Observable Dynamics, and Operator Growth

    quant-ph 2025-09 conditional novelty 6.0 of 10

    FNOs learn three maps for time-periodic spin chains (Floquet Hamiltonian, local observables, operator growth) with high accuracy, zero-shot transfer across time grids and driving frequencies, and extrapolation beyond ...

  6. OpenBreastUS: Benchmarking Neural Operators for Wave Imaging Using Breast Ultrasound Computed Tomography

    cs.CV 2025-07 conditional novelty 6.0 of 10

    OpenBreastUS provides a large-scale, anatomically realistic benchmark of 16 million breast ultrasound simulations and demonstrates neural-operator-based full-waveform inversion on clinical in vivo breast data.

  7. DEF: Diffusion-augmented Ensemble Forecasting

    cs.LG 2025-06 conditional novelty 6.0 of 10

    A conditional diffusion model that generates perturbed initial states can turn any deterministic neural weather forecast model into an ensemble, with measured error reduction on a single ERA5 case study.

  8. Tokenizing Electron Cloud in Protein-Ligand Interaction Learning

    cs.LG 2025-05 conditional novelty 6.0 of 10

    ECBind tokenizes electron cloud densities via quantized embeddings and improves protein-ligand binding affinity prediction, especially per-structure correlations on MISATO.

  9. DGenNO: A Novel Physics-aware Neural Operator for Solving Forward and Inverse PDE Problems based on Deep, Generative Probabilistic Modeling

    cs.LG 2025-02 conditional novelty 6.0 of 10

    A physics-driven neural operator with latent generative encoding solves forward and inverse PDE problems without labeled data, using weak-form residuals to handle discontinuous coefficients.

  10. Implicit factorized transformer approach to fast prediction of turbulent channel flows

    physics.flu-dyn 2024-12 conditional novelty 6.0 of 10

    IFactFormer-m, using parallel factorized attention, achieves stable long-term predictions of turbulent channel flows at Re_tau = 180, 395, and 590, outperforming FNO, IFNO, IFactFormer-o, and traditional LES models.

  11. DRIFT: Direct Reduced Fourier Transforms for Distributed Spectral Neural Operators

    cs.DC 2026-07 conditional novelty 5.0 of 10

    Distributed Fourier Neural Operators can compute their truncated spectra with local partial DFTs and two collectives on the kept modes, giving exact results with communication independent of grid resolution.

  12. Integrating Fourier Neural Operator with Diffusion Model for Autoregressive Predictions of Three-dimensional Turbulence

    physics.flu-dyn 2025-12 conditional novelty 5.0 of 10

    DiAFNO, an implicit adaptive Fourier neural operator used as the denoiser inside an EDM diffusion model, gives more accurate autoregressive predictions of 3D turbulence than EDM or dynamic Smagorinsky LES.

  13. A Geometry-Aware AI Emulator for the Coupled Whole Atmosphere from Earth Surface to the Ionosphere and Thermosphere

    physics.space-ph 2025-06 reject novelty 5.0 of 10

    CAM-NET, an SFNO-based neural emulator, is claimed to reproduce WACCM-X whole-atmosphere variability at over 1000x speedup, but the paper provides no quantitative validation, code, or data to support the claim.

  14. FourierFlow: Frequency-aware Flow Matching for Generative Turbulence Modeling

    cs.LG 2025-06 conditional novelty 5.0 of 10

    FourierFlow improves multi-step generative turbulence modeling by adding a differential-attention branch, a frequency-weighted Fourier mixing branch, and MAE-based feature alignment to a flow matching model.

  15. Neural Interpretable PDEs: Harmonizing Fourier Insights with Attention for Scalable and Interpretable Physics Discovery

    cs.LG 2025-05 conditional novelty 5.0 of 10

    NIPS is a neural operator that uses linear attention and Fourier kernels to simultaneously predict PDE solutions and recover hidden material properties from limited data.

  16. Latent Mamba Operator for Partial Differential Equations

    cs.LG 2025-05 conditional novelty 5.0 of 10

    LaMO replaces attention in latent-token neural operators with bidirectional state-space models and reports consistent accuracy gains on six PDE benchmarks.

  17. Physics-Guided Learning of Meteorological Dynamics for Weather Downscaling and Forecasting

    cs.LG 2025-05 conditional novelty 5.0 of 10

    PhyDL-NWP trains neural surrogates with a fitted PDE regularizer and a latent force term, improving weather downscaling and fine-tuning forecasts over 17 baselines.

  18. Fourier-enhanced Neural Networks For Systems Biology Applications

    cs.LG 2025-02 conditional novelty 5.0 of 10

    SB-FNN, a Fourier-neural-operator-based physics-informed solver with adaptive activations and a variance penalty, reports lower N-MSE than vanilla PINN on six systems biology models.

  19. AirRadar: Inferring Nationwide Air Quality in China with Deep Neural Networks

    cs.LG 2025-01 conditional novelty 5.0 of 10

    A deep network with mask tokens, local and global spatial learners, and learned context weights infers PM2.5 at unmonitored locations across China with reported MAE of 6.41 to 8.11 at 25% to 75% missing stations.

  20. Physics-guided spatiotemporal neural models for fuel density prediction

    cs.LG 2026-07 conditional novelty 4.0 of 10

    Adding physics-guided loss terms to ConvLSTM, AFNONet, and ViViT improves fuel density prediction accuracy and stability over purely data-driven baselines on simulated prescribed-fire data.

  21. GeoTransolver: Learning Physics on Irregular Domains Using Multi-scale Geometry Aware Physics Attention Transformer

    cs.LG 2025-12 conditional novelty 4.0 of 10

    GeoTransolver, a geometry-aware attention transformer, improves surrogate CFD accuracy over existing baselines on three automotive/aerospace datasets, but the paper has major reporting gaps.

  22. Principled Approaches for Extending Neural Architectures to Function Spaces for Operator Learning

    cs.LG 2025-06 conditional novelty 4.0 of 10

    A practical recipe to convert common neural architectures into discretization-agnostic neural operators, validated by Navier-Stokes experiments showing cross-resolution generalization of FNO-style models.

  23. Deep Learning and Foundation Models for Weather Prediction: A Survey

    cs.LG 2025-01 conditional novelty 4.0 of 10

    A survey that organizes deep learning weather prediction models into three training paradigms: deterministic, generative, and pre-train-fine-tune.

  24. Two-flow Feedback Multi-scale Progressive Generative Adversarial Network

    cs.CV 2025-08 reject novelty 3.0 of 10

    A GAN paper that proposes several new modules but reports no actual experimental results, with placeholder dataset names and percentages.

  25. Dynamic Double Space Tower

    cs.CV 2025-06 reject novelty 3.0 of 10

    The paper claims a four-layer Gestalt-based tower can replace attention in VQA and lift a 3B model to state-of-the-art spatial reasoning, but provides no reproducible method or consistent results.

Pith tools