REVIEW 3 major objections 5 minor 57 references
Mondrian: Transformer Operators via Domain Decomposition
T0 review · 3 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read Mondrian computes exact softmax attention over subdomain-restricted functions, decoupling transformer operator learning from grid resolution and enabling PDE models trained at 32×32 to run at 128×128 without retraining.
desk verdict Useful architecture paper, but the cross-subdomain inner product is underspecified and likely vanishes, so the central attention-over-functions claim needs a fix before it can stand. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is subdomain self-attention over functions (Algorithm 1). Instead of patching points as vision transformers do, it uses a domain decomposition $D_s(\Omega)$ with restriction and zero-extension operators, computes queries, keys, and values as functions via neural operators, and replaces the vector dot product $QK^T$ with the $L^2$ inner product $\langle Q_k,K_j\rangle$ to build a finite score matrix. A second component is the mixture operator, whose kernel $\kappa(x,y)=\sum_i M_i C_i(x,y)$ mixes learned matrices with coefficients from a small network; the paper shows this subdomain operator is the choice that stays accurate as resolution grows.
What would settle it
Compute the score matrix from Algorithm 1 for actual trained queries and keys: if $S_{k,j}$ is evaluated as the standard $L^2$ inner product over $\Omega$ (or over the union of subdomains), all off-diagonal entries with $k \neq j$ are exactly zero for functions with disjoint support, so the model would reduce to independent per-subdomain processing. Checking whether off-diagonal scores are nonzero, and over what domain the integration runs, settles whether the claimed global attention is real.
Extended reading notes
Core claim
Mondrian's central claim is that attention can be computed over functions rather than vectors: decompose $\Omega$ into $s$ non-overlapping subdomains, map each restricted input through a neural operator $I_{QKV}$ to query, key, and value functions, form the score matrix $S_{k,j}=\langle Q_k,K_j\rangle$ with the $L^2$ inner product, and recompose softmax-weighted values across subdomains. Because the subdomain partition is fixed in physical space, the sequence length $s$ is independent of how finely the function is sampled, which the paper identifies as the mechanism behind resolution generalization. Experimentally, the mixture subdomain operator—a kernel built as a learned mixture of trainable matrices—gives the most stable resolution scaling, and a Swin-style shifted-window version performs best on turbulent Navier-Stokes rollouts.
Load-bearing premise
The argument depends on the score $S_{k,j}=\langle Q_k,K_j\rangle$ being a meaningful, nonzero inner product between functions defined on different, non-overlapping subdomains; the paper never states the integration domain or function space, and with the usual $L^2$ inner product over the whole domain those cross-subdomain scores vanish, which would collapse attention to a per-subdomain diagonal and remove the cross-subdomain communication the method relies on.
Editorial extensions
If this is right
- Exact softmax attention over subdomains stays quadratic in the number of subdomains rather than grid points, so finer discretizations no longer increase attention sequence length.
- A Mondrian model trained only on $32\times32$ Allen-Cahn data can be evaluated at $64\times64$ and $128\times128$ without retraining, with the best variants keeping MAE below $10^{-2}$ at the finest resolution.
- Swin-style shifted-window attention pays off on turbulent Navier-Stokes: the Swin-NO with the mixture operator reaches the lowest two-step and rollout errors among the compared models.
- The choice of subdomain operator is decisive: mixture and separable-mixture kernels beat spectral convolution, low-rank kernels, and in-subdomain attention, with the separable variant doing so at roughly five times fewer parameters.
Reading between the lines
- The paper does not specify the integration domain for $\langle Q_k,K_j\rangle$; if cross-subdomain terms are implemented as zero-extension inner products over $\Omega$, the model's resolution transfer becomes a claim about the kernel's sampling behavior that the mixture operator's learned coefficients could be tested against directly.
- Because vanishing off-diagonal scores would reduce the method to per-subdomain processing, a cheap check is to permute subdomain indices at inference: truly global attention would change outputs, while an effectively diagonal score matrix would not.
- The localized cross-attention idea for grid-to-grid mapping could be lifted out as a standalone interpolation or coarsening operator; measuring its error against standard interpolation on an irregular grid would quantify how much function-level attention contributes beyond geometry-aware pooling.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes Mondrian, a family of transformer architectures for PDE operator learning. The input function is decomposed into non-overlapping subdomains, and the transformer sequence is a set of subdomain-restricted functions rather than point vectors. Attention between subdomains is defined through L2 inner products of query and key functions, with optional relative position bias; local variants use windowed and neighborhood attention. Subdomain-level operators, including a new Mixture Operator, replace linear layers. The authors evaluate on Allen-Cahn and 2D Navier-Stokes, reporting generalization from 32x32 training to 64x64 and 128x128 test resolutions and comparing with FNO, FFNO, Galerkin Transformer, and FactFormer.
Significance. The architectural idea is attractive: it keeps exact softmax attention while making the sequence length depend on the number of subdomains rather than the discretization, and it provides a modular set of subdomain operators. The experimental protocol is careful, with held-out resolutions, validation-based hyperparameter selection, parameter counts, and extensive ablations, and the paper is transparent about limitations such as spectral bias and zero-shot superresolution. However, the central definition of the score inner product is ambiguous, and under the only extension operator defined in the paper the cross-subdomain attention scores vanish. This gap directly affects the claimed mechanism for global interaction and must be repaired before the architectural claims can be assessed.
major comments (3)
- [Section 3.2, Algorithm 1 (line 3)] The score definition S_{k,j}=<Q_k,K_j> is either undefined or degenerate because no common domain for Q_k and K_j is specified. The only function-space convention introduced in Section 2.2 uses restrictions to subdomains and zero extensions E_{Omega_i} to Omega. Under that convention, Q_k and K_j with k != j have disjoint supports, so <Q_k,K_j>_{L2(Omega)}=0, softmax attention is diagonal, and cross-subdomain content communication vanishes; the off-diagonal scores would then reduce to the learned position bias B described in Section 3.2, which the paper states is used in the experiments. The notation Omega_i ≃ Omega_j suggests that a pullback to a reference subdomain is intended, but no such construction, shared basis, or quadrature is given. This is load-bearing because it is the mechanism for the abstract's claim of attention across subdomains via softmax-based inner products over functions.
- [Algorithm 1, lines 4-5] The recombination step u_{Omega_i}=softmax(tau S)_i V is also not well-defined as written: each V_j is a function on Omega_j or its zero extension, so forming a weighted sum over j and then applying E_{Omega_i} from Omega_i to Omega either requires adding functions on different domains or discards the V_j contributions outside Omega_i. A consistent construction, such as pullback to a common reference subdomain followed by pushforward, must be supplied. This is not a notational issue because the output function of the architecture is defined through this step.
- [Abstract, Section 4.3, Section 4.4, Tables 4 and 5] The claim that Mondrian 'matches state-of-the-art accuracy' is stated more strongly than the evidence. In Table 4, the best Mondrian variant has MAE 3.58e-3 at 32x32 and 8.12e-3 at 128x128, whereas FFNO has 4.35e-3 and 5.11e-3; at the highest resolution FFNO is both more accurate and more stable across resolutions. In Table 5, the best Mondrian model (Swin-NO-MO) has MAE 7.07e-2 compared with FFNO's 5.05e-2. The results show competitive performance and nontrivial resolution transfer, but the stated match to state-of-the-art accuracy should be qualified accordingly.
minor comments (5)
- [Section 2.4] The compressed notation with the domain 'assumed to be clear from the context' is exactly what makes Algorithm 1 ambiguous; a concrete domain should be attached to every function in Algorithm 1.
- [Section 3.2] The expression 'softmax(tau S)_i V' should include the subscript j on V and the summation over j to match standard attention; as written it is unclear what is being summed.
- [Section 4.3 and Appendix C.1] The range for gamma is stated as [1e-4, 5e-3] in Section 4.3 but as [1e-5, 5e-3] in Appendix C.1; these should be reconciled.
- [Appendix E.2, Table 7] Swin-NO-SpecConv reports a two-step MSE of 5.25e-1 and a rollout MSE of 1.26e-1; since rollout error compounds over multiple steps, this inversion is unexpected and warrants verification.
- [Section 5] The comparison with Continuum Attention [6] would benefit from an explicit statement of whether that work's score inner product uses a common reference domain; the current phrasing gives no basis for the claimed difference.
Circularity Check
No significant circularity: the central claims are empirically tested against external baselines and held-out resolutions, not derived from fitted inputs or self-citations.
full rationale
The paper's load-bearing claims are empirical architecture comparisons. Resolution scaling is evaluated on held-out 64x64 and 128x128 test sets after training only at 32x32 (Section 4.3, Table 4), and hyperparameters are selected on a validation split (Appendix C.2). The mixture operator's kernel is a genuine function of coordinates, and the paper reports that some subdomain operators, such as the Interpolating operator, degrade at higher resolutions (Table 6), which shows resolution generalization is not forced by construction. No self-citation is load-bearing: the closest related work, Continuum attention [6], has disjoint authorship and is cited as related background, not as justification for Mondrian's design. The ambiguity about the domain of integration in the L2 inner product in Section 3.2 (Algorithm 1, line 3) is a soundness and rigor concern about whether cross-subdomain attention is content-based; if the inner product vanishes on disjoint subdomains, that undermines the global-interaction claim rather than tautologically establishing it. This is not a circularity: the empirical claims do not reduce to this definition, and the paper's own limitations section concedes zero-shot superresolution and high-frequency issues. Therefore no circular step can be exhibited.
Assumptions & free parameters
free parameters (4)
- number of subdomains s (subdomain grid decomposition) =
64 subdomains (8x8) for Allen-Cahn and ViT-NO Navier-Stokes; 32x32 grid for Swin-NO Navier-Stokes
- Mixture Operator number of matrices n =
tuned from {32,64,128} for MO and {32,64,128,256} for Separable MO (Table 3)
- embedding dimension =
tuned from {32,64,128,256} (Table 3)
- Swin window size and shift =
window of 4x4 subdomains, shift of 2 subdomains (Allen-Cahn); 16x16 subdomain windows with shift 4 (Navier-Stokes)
assumptions (3)
- ad hoc to paper S_kj = <Q_k, K_j> is a well-defined L2 inner product and is nonzero for k != j, with Q_k and K_j functions on disjoint subdomains
- domain assumption The kernel integral operator's discretized quadrature on test resolutions accurately approximates the continuous operator trained at 32x32
- domain assumption The training distribution of 32x32 subsampled initial conditions covers the higher-resolution function space
Cite this review
Pith. "Pith review of Mondrian: Transformer Operators via Domain Decomposition." pith.science (2026). https://pith.science/paper/KMQLKCR4
@misc{pith2026250608226,
author = {Pith},
title = {Pith review of: Mondrian: Transformer Operators via Domain Decomposition},
year = {2026},
howpublished = {\url{https://pith.science/paper/KMQLKCR4}},
note = {Machine review of arXiv:2506.08226}
}
read the original abstract
Operator learning enables data-driven modeling of partial differential equations (PDEs) by learning mappings between function spaces. However, scaling transformer-based operator models to high-resolution, multiscale domains remains a challenge due to the quadratic cost of attention and its coupling to discretization. We introduce \textbf{Mondrian}, transformer operators that decompose a domain into non-overlapping subdomains and apply attention over sequences of subdomain-restricted functions. Leveraging principles from domain decomposition, Mondrian decouples attention from discretization. Within each subdomain, it replaces standard layers with expressive neural operators, and attention across subdomains is computed via softmax-based inner products over functions. The formulation naturally extends to hierarchical windowed and neighborhood attention, supporting both local and global interactions. Mondrian achieves strong performance on Allen-Cahn and Navier-Stokes PDEs, demonstrating resolution scaling without retraining. These results highlight the promise of domain-decomposed attention for scalable and general-purpose neural operators.
Figures
Figures from the paper (5 more)
Reference graph
Works this paper leans on
-
[1]
S. M. Allen and J. W. Cahn. Mechanisms of phase transformations within the miscibility gap of fe-rich fe-al alloys.Acta Metallurgica, 24(5):425–437, 1976
work page 1976
-
[2]
K. Azizzadenesheli, N. Kovachki, Z. Li, M. Liu-Schiaffini, J. Kossaifi, and A. Anandkumar. Neural operators for accelerating scientific simulations and design.Nature Reviews Physics, pages 1–9, 2024
work page 2024
-
[3]
F. Bartolucci, E. de Bézenac, B. Raonic, R. Molinaro, S. Mishra, and R. Alaifari. Representation equivalent neural operators: a framework for alias-free operator learning.Advances in Neural Information Processing Systems, 36, 2024
work page 2024
-
[4]
M. J. Berger and P. Colella. Local adaptive mesh refinement for shock hydrodynamics.Journal of computational Physics, 82(1):64–84, 1989
1989
-
[5]
M. J. Berger and J. Oliger. Adaptive mesh refinement for hyperbolic partial differential equations. Journal of computational Physics, 53(3):484–512, 1984
work page 1984
-
[6]
E. Calvello, N. B. Kovachki, M. E. Levine, and A. M. Stuart. Continuum attention for neural operators, 2024
work page 2024
-
[7]
S. Cao. Choose a transformer: Fourier or Galerkin. InAdvances in Neural Information Processing Systems (NeurIPS 2021), volume 34, 2021
work page 2021
-
[8]
S. Cao, P. Xu, and D. A. Clifton. How to understand masked autoencoders, 2022
work page 2022
Show all 57 references
-
[9]
T. F. Chan and T. P. Mathew. Domain decomposition algorithms.Acta Numerica, 3:61–143, 1994
1994
-
[10]
F. Chollet. Xception: Deep learning with depthwise separable convolutions. InProceedings of the IEEE conference on computer vision and pattern recognition, pages 1251–1258, 2017
2017
-
[11]
K. M. Choromanski, V . Likhosherstov, D. Dohan, X. Song, A. Gane, T. Sarlos, P. Hawkins, J. Q. Davis, A. Mohiuddin, L. Kaiser, D. B. Belanger, L. J. Colwell, and A. Weller. Rethinking attention with performers. InInternational Conference on Learning Representations, 2021
2021
-
[12]
T. Dao. FlashAttention-2: Faster attention with better parallelism and work partitioning. In International Conference on Learning Representations (ICLR), 2024
2024
-
[13]
T. Dao, D. Y . Fu, S. Ermon, A. Rudra, and C. Ré. FlashAttention: Fast and memory-efficient exact attention with IO-awareness. InAdvances in Neural Information Processing Systems (NeurIPS), 2022
2022
-
[14]
Dao and A
T. Dao and A. Gu. Transformers are SSMs: Generalized models and efficient algorithms through structured state space duality. InInternational Conference on Machine Learning (ICML), 2024. 10
2024
-
[15]
Dehghani, J
M. Dehghani, J. Djolonga, B. Mustafa, P. Padlewski, J. Heek, J. Gilmer, A. Steiner, M. Caron, R. Geirhos, I. Alabdulmohsin, R. Jenatton, L. Beyer, M. Tschannen, A. Arnab, X. Wang, C. Riquelme, M. Minderer, J. Puigcerver, U. Evci, M. Kumar, S. van Steenkiste, G. F. Elsayed, A. ...
2023
-
[16]
Dolean, P
V . Dolean, P. Jolivet, and F. Nataf.An Introduction to Domain Decomposition Methods. Society for Industrial and Applied Mathematics, Philadelphia, PA, 2015
2015
-
[17]
Dosovitskiy, L
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, J. Uszkoreit, and N. Houlsby. An image is worth 16x16 words: Transformers for image recognition at scale, 2021
2021
-
[18]
Dresdner, D
G. Dresdner, D. Kochkov, P. Norgaard, L. Zepeda-Núñez, J. A. Smith, M. P. Brenner, and S. Hoyer. Learning to correct spectral methods for simulating turbulent flows.arXiv preprint arXiv:2207.00556, 2022
2022 arXiv
-
[19]
Dubois, M
O. Dubois, M. J. Gander, S. Loisel, A. St-Cyr, and D. B. Szyld. The optimized schwarz method with a coarse grid correction.SIAM Journal on Scientific Computing, 34(1):A421–A458, 2012
2012
-
[20]
Fanaskov and I
V . Fanaskov and I. Oseledets. Spectral neural operators, 2024
2024
-
[21]
M. J. Gander. Optimized schwarz methods.SIAM Journal on Numerical Analysis, 44(2):699– 731, 2006
2006
-
[22]
Gu and T
A. Gu and T. Dao. Mamba: Linear-time sequence modeling with selective state spaces.arXiv preprint arXiv:2312.00752, 2023
2023 arXiv
-
[23]
D. Han, Y . Pu, Z. Xia, Y . Han, X. Pan, X. Li, J. Lu, S. Song, and G. Huang. Bridging the divide: Reconsidering softmax and linear attention. InThe Thirty-eighth Annual Conference on Neural Information Processing Systems, 2024
2024
-
[24]
Z. Hao, C. Su, S. Liu, J. Berner, C. Ying, H. Su, A. Anandkumar, J. Song, and J. Zhu. DPOT: Auto-regressive denoising operator transformer for large-scale PDE pre-training. InForty-first International Conference on Machine Learning, 2024
2024
-
[25]
Z. Hao, Z. Wang, H. Su, C. Ying, Y . Dong, S. Liu, Z. Cheng, J. Song, and J. Zhu. Gnot: A general neural operator transformer for operator learning. InInternational Conference on Machine Learning, pages 12556–12569. PMLR, 2023
2023
-
[26]
K. He, X. Chen, S. Xie, Y . Li, P. Dollár, and R. Girshick. Masked autoencoders are scalable vision learners. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 16000–16009, 2022
2022
-
[27]
Herde, B
M. Herde, B. Raoni ´c, T. Rohner, R. Käppeli, R. Molinaro, E. de Bézenac, and S. Mishra. Poseidon: Efficient foundation models for pdes.arXiv preprint arXiv:2405.19101, 2024
2024 arXiv
-
[28]
Hersbach, B
H. Hersbach, B. Bell, P. Berrisford, S. Hirahara, A. Horányi, J. Muñoz-Sabater, J. Nicolas, C. Peubey, R. Radu, D. Schepers, et al. The era5 global reanalysis.Quarterly Journal of the Royal Meteorological Society, 146(730):1999–2049, 2020
1999
-
[29]
J. Ho, N. Kalchbrenner, D. Weissenborn, and T. Salimans. Axial attention in multidimensional transformers.arXiv preprint arXiv:1912.12180, 2019
1912 arXiv
-
[30]
Karypis and V
G. Karypis and V . Kumar. A fast and high quality multilevel scheme for partitioning irregular graphs.SIAM Journal on scientific Computing, 20(1):359–392, 1998
1998
-
[31]
Kirillov, E
A. Kirillov, E. Mintun, N. Ravi, H. Mao, C. Rolland, L. Gustafson, T. Xiao, S. Whitehead, A. C. Berg, W.-Y . Lo, P. Dollár, and R. Girshick. Segment anything.arXiv:2304.02643, 2023
2023 arXiv
-
[32]
N. B. Kovachki, S. Lanthaler, and A. M. Stuart. Operator learning: Algorithms and analysis. arXiv preprint arXiv:2402.15715, 2024. 11
2024 arXiv
-
[33]
N. B. Kovachki, Z. Li, B. Liu, K. Azizzadenesheli, K. Bhattacharya, A. M. Stuart, and A. Anand- kumar. Neural operator: Learning maps between function spaces.CoRR, abs/2108.08481, 2021
2021 arXiv
-
[34]
Z. Li, N. Kovachki, K. Azizzadenesheli, B. Liu, K. Bhattacharya, A. Stuart, and A. Anandkumar. Fourier neural operator for parametric partial differential equations, 2020
2020
-
[35]
Z. Li, N. Kovachki, K. Azizzadenesheli, B. Liu, K. Bhattacharya, A. Stuart, and A. Anandkumar. Neural operator: Graph kernel network for partial differential equations, 2020
2020
-
[36]
Z. Li, N. B. Kovachki, C. Choy, B. Li, J. Kossaifi, S. P. Otta, M. A. Nabian, M. Stadler, C. Hundt, K. Azizzadenesheli, and A. Anandkumar. Geometry-informed neural operator for large-scale 3d PDEs. InThirty-seventh Conference on Neural Information Processing Systems, 2023
2023
-
[37]
Z. Li, D. Shu, and A. Barati Farimani. Scalable transformer for pde surrogate modeling. Advances in Neural Information Processing Systems, 36, 2024
2024
-
[38]
Lions et al
P.-L. Lions et al. On the schwarz alternating method. i. InFirst international symposium on domain decomposition methods for partial differential equations, volume 1, page 42. Paris, France, 1988
1988
-
[39]
Z. Liu, H. Hu, Y . Lin, Z. Yao, Z. Xie, Y . Wei, J. Ning, Y . Cao, Z. Zhang, L. Dong, F. Wei, and B. Guo. Swin transformer v2: Scaling up capacity and resolution, 2022
2022
-
[40]
Z. Liu, Y . Lin, Y . Cao, H. Hu, Y . Wei, Z. Zhang, S. Lin, and B. Guo. Swin transformer: Hierarchical vision transformer using shifted windows, 2021
2021
-
[41]
Liu-Schiaffini, J
M. Liu-Schiaffini, J. Berner, B. Bonev, T. Kurth, K. Azizzadenesheli, and A. Anandkumar. Neural operators with localized integral and differential kernels. InICLR 2024 Workshop on AI4DifferentialEquations In Science, 2024
2024
-
[42]
Loshchilov, F
I. Loshchilov, F. Hutter, et al. Fixing weight decay regularization in adam.arXiv preprint arXiv:1711.05101, 5, 2017
2017 arXiv
-
[43]
Oquab, T
M. Oquab, T. Darcet, T. Moutakanni, H. V . V o, M. Szafraniec, V . Khalidov, P. Fernandez, D. Haziza, F. Massa, A. El-Nouby, R. Howes, P.-Y . Huang, H. Xu, V . Sharma, S.-W. Li, W. Galuba, M. Rabbat, M. Assran, N. Ballas, G. Synnaeve, I. Misra, H. Jegou, J. Mairal, P. Labatut,...
2023
-
[44]
Pathak, S
J. Pathak, S. Subramanian, P. Harrington, S. Raja, A. Chattopadhyay, M. Mardani, T. Kurth, D. Hall, Z. Li, K. Azizzadenesheli, et al. Fourcastnet: A global data-driven high-resolution weather model using adaptive fourier neural operators.arXiv preprint arXiv:2202.11214, 2022
2022 arXiv
-
[45]
P. Pope, C. Zhu, A. Abdelkader, M. Goldblum, and T. Goldstein. The intrinsic dimension of images and its impact on learning. InInternational Conference on Learning Representations, 2021
2021
-
[46]
M. A. Rahman, R. J. George, M. Elleithy, D. Leibovici, Z. Li, B. Bonev, C. White, J. Berner, R. A. Yeh, J. Kossaifi, K. Azizzadenesheli, and A. Anandkumar. Pretraining codomain attention neural operators for solving multiphysics pdes, 2024
2024
-
[47]
Raonic, R
B. Raonic, R. Molinaro, T. De Ryck, T. Rohner, F. Bartolucci, R. Alaifari, S. Mishra, and E. de Bézenac. Convolutional neural operators for robust and accurate learning of pdes. Advances in Neural Information Processing Systems, 36, 2024
2024
-
[48]
P. Shaw, J. Uszkoreit, and A. Vaswani. Self-attention with relative position representations. arXiv preprint arXiv:1803.02155, 2018
2018 arXiv
-
[49]
Solodskikh, A
K. Solodskikh, A. Kurbanov, R. Aydarkhanov, I. Zhelavskaya, Y . Parfenov, D. Song, and S. Lefkimmiatis. Integral neural networks. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 16113–16122, June 2023
2023
-
[50]
J. Su, M. Ahmed, Y . Lu, S. Pan, W. Bo, and Y . Liu. Roformer: Enhanced transformer with rotary position embedding.Neurocomputing, 568:127063, 2024. 12
2024
-
[51]
Touvron, T
H. Touvron, T. Lavril, G. Izacard, X. Martinet, M.-A. Lachaux, T. Lacroix, B. Rozière, N. Goyal, E. Hambro, F. Azhar, A. Rodriguez, A. Joulin, E. Grave, and G. Lample. Llama: Open and efficient foundation language models, 2023
2023
-
[52]
A. Tran, A. Mathews, L. Xie, and C. S. Ong. Factorized fourier neural operators. InThe Eleventh International Conference on Learning Representations, 2023
2023
-
[53]
Vaswani, N
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin. Attention is all you need, 2023
2023
-
[54]
Villaescusa-Navarro
F. Villaescusa-Navarro. Pylians: Python libraries for the analysis of numerical simulations. Astrophysics Source Code Library, record ascl:1811.008, Nov. 2018
2018
-
[55]
P. Wang, W. Zheng, T. Chen, and Z. Wang. Anti-oversmoothing in deep vision transformers via the fourier domain analysis: From theory to practice. InInternational Conference on Learning Representations, 2022
2022
-
[56]
C. Yun, S. Bhojanapalli, A. S. Rawat, S. Reddi, and S. Kumar. Are transformers universal approximators of sequence-to-sequence functions? InInternational Conference on Learning Representations, 2020
2020
-
[57]
subdomain operators
D. Zwicker. py-pde.https://github.com/zwicker-group/py-pde, 2025. 13 A Notation and Terminology Ω⊂R n A bounded domain Ds(Ω) A decomposition of the domainΩintosnon-overlapping subdomains Ωi Thei-th subdomain of a decomposition Θi Thei-th window used in window attention fΩi The...
2025
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.