Pith. sign in

REVIEW 3 major objections 1 minor 37 references

Physics- and geometry-aware spatio-spectral graph neural operator for time-independent and time-dependent PDEs

T0 review · 3 major / 1 minor · reviewed 2026-08-05 · deepseek-v4-flash

Pith's one-line read The abstract claims a physics-aware graph neural operator for PDEs that outperforms benchmarks; the supplied full text is a different paper and contains none of it.

desk verdict Desk reject: the abstract announces a PDE graph-neural-operator paper, but the body is an unrelated math paper on self-attention dynamics with different authors, so there is nothing to referee. read the letter →

arxiv 2508.09627 v1 pith:PVJXSE63 submitted 2025-08-13 cs.LG

classification cs.LG
keywords graphneuraloperatorPDEsolutionoperatorsphysics-informedlossspatio-spectrallearninggeometry-awaredeepmultiscaleself-attentiondynamicsFrank–Wolfe
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper's abstract asks the reader to believe that a new architecture, the Physics- and Geometry-Aware Spatio-Spectral Graph Neural Operator (πG-Sp²GNO), learns solution operators for time-independent and time-dependent PDEs on complex and varying geometries with limited labeled data, using a physics-informed loss that combines higher-order time marching with a stochastic projection inspired by upscaled theory. If true, this would matter because PDE surrogate models could be trained without simulation data and would transfer across geometry changes at inference time. The attached full text, however, is a different manuscript: a mathematical study of the hardmax limit of self-attention dynamics as Frank–Wolfe steps, with Voronoi cells, convergence rates, and metastability. It contains no description of πG-Sp²GNO, no spatio-spectral graph construction, no hybrid loss, and no PDE benchmarks. The central claim is therefore asserted in the abstract but unsupported by the supplied body.

What carries the argument

According to the abstract, the machinery is the spatio-spectral graph architecture inherited from Sp²GNO, extended with geometry awareness, together with a hybrid physics-informed loss. The spatio-spectral structure is meant to enable multiscale learning; the geometry-awareness strategies are meant to handle irregular and changing domains; the loss combines a higher-order time-marching scheme with an upscaled-theory-inspired stochastic projection to fold the governing PDE into training. In the attached text, by contrast, the operative machinery is the identification of the hardmax self-attention update as a Frank–Wolfe step, with cell structure given by a Voronoi tessellation when the key–qu

What would settle it

Read the supplied body for any occurrence of 'Sp²GNO', 'physics-informed', 'benchmark', or the geometry-aware graph construction; none appears, because the body is a separate paper on self-attention. To test the substantive claim, implement πG-Sp²GNO as described in the abstract and run it on the listed PDE benchmarks, comparing error and training cost against existing physics-informed neural operators; the abstract's claim stands or falls on those numbers, which are not present.

Watch

Extended reading notes

Core claim

The abstract's central claim is that πG-Sp²GNO learns PDE solution operators in a simulation-free setup and beats existing physics-informed neural operators on benchmarks that include regular and complex domains, geometry changes during inference, and time-dependent problems. The mechanisms named for this are a geometry-aware spatio-spectral graph structure for multiscale learning and a hybrid physics-informed loss combining a higher-order time-marching scheme with an upscaled-theory-inspired stochastic projection. The supplied body, however, is a different work on the hardmax limit of self-attention dynamics, where the update rule is shown to be a Frank–Wolfe step, with Voronoi-cell geometr

Load-bearing premise

The load-bearing premise is that the attached full text is the paper described in the abstract; in the submitted document this is false, so the claimed architecture, loss, and experiments have no supporting text.

Editorial extensions

If this is right

  • If πG-Sp²GNO works as claimed, PDE solution operators could be trained without labeled simulation data, using only the governing equations and geometry information, on irregular and changing domains.
  • Time-dependent PDEs would benefit from a hybrid loss where higher-order time marching and a stochastic projection of the upscaled physics are both enforced, improving physical consistency with limited data.
  • Geometry-awareness strategies would allow inference on domains not seen during training, since the graph representation adapts to the new geometry.
  • On the stated benchmarks, the method would outperform existing physics-informed neural operators, providing a practical accuracy gain.
  • None of these implications can be verified from the supplied text, since the body is a different manuscript and contains no experiments or architecture details.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • One testable extension of the abstract's claim would be to stress the geometry-aware operator on topologically changing domains (e.g., holes appearing or merging), not just smooth shape variation; the abstract does not promise this, but it is a natural next stress test.
  • The body's Frank–Wolfe picture suggests that attention-based operators may only need the extreme tokens of the embedding cloud; if that transfers, spatio-spectral graph operators could be made cheaper by pruning non-extreme tokens, but the paper does not make this connection.
  • An editorial inference: the mismatch indicates that structural checks—comparing abstract and body, author lists, and theorem statements—should precede any evaluation of scientific claims; this is a process lesson, not a scientific result of the paper.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 1 minor

Summary. The abstract of the submitted manuscript announces a physics- and geometry-aware spatio-spectral graph neural operator (πG-Sp²GNO) for learning solution operators of time-independent and time-dependent PDEs, together with a hybrid physics-informed loss combining higher-order time marching and an upscaled-theory-inspired stochastic projection, and reports benchmark improvements over state-of-the-art physics-informed neural operators. However, the full text provided is a different paper, 'Attention's forward pass and Frank-Wolfe' (arXiv:2508.09628), by different authors, with its own abstract, contents, and bibliography. The body contains no graph neural operator, no spatio-spectral architecture, no geometry-awareness mechanism, no PDE experiments, and no discussion of the proposed loss. As assembled, the document does not describe or test the claimed method.

Significance. If the claimed πG-Sp²GNO method were fully described and validated, it could be a meaningful contribution to operator learning for PDEs on complex geometries with limited data, particularly the proposed hybrid loss and geometry-awareness strategies. However, the submitted document contains no derivation, algorithm, or experimental evidence supporting these claims. The attached full text is a mathematical study of self-attention dynamics with a different title, author list, and subject matter. Consequently, the significance of the claimed result cannot be assessed from this submission. The only verifiable content is the unrelated attention-dynamics paper, which is not the work described in the abstract.

major comments (3)
  1. [Abstract vs. full text] The central claim of the abstract—that the paper introduces πG-Sp²GNO and demonstrates its superiority on PDE benchmarks—has no supporting material in the document body. The full text is 'Attention's forward pass and Frank-Wolfe' (arXiv:2508.09628), with a different title, authors, abstract, and bibliography. There is no occurrence of the proposed architecture, the hybrid physics-informed loss, the stochastic projection, or any PDE experiment. Thus the load-bearing premise of the submission is false: the document under review does not describe the method it claims to present.
  2. [§1.1, Eqs. (1.1)–(1.3)] The mathematical content begins with self-attention dynamics for token embeddings, not with a graph neural operator for PDEs. Equation (1.1) is the standard softmax self-attention update, and the subsequent analysis concerns Frank-Wolfe steps, Voronoi cells, and metastability. None of these objects is connected to solving PDEs, spatio-spectral graph operators, or geometry-aware learning. The claimed method therefore has no derivations in the text.
  3. [Abstract (hybrid loss and geometry-awareness)] The abstract asserts two novel components—two geometry-awareness strategies and a hybrid loss combining higher-order time marching with an upscaled-theory-inspired stochastic projection—and claims they improve upon state-of-the-art physics-informed neural operators. Because the body contains no definitions, no algorithmic descriptions, and no benchmark tables or figures, these assertions are unverifiable. This is not a minor omission; it removes the entire evidential basis for the abstract's conclusions.
minor comments (1)
  1. [General] The manuscript's title, author list, and abstract should be consistent with the body text. As submitted, the metadata identifies one paper while the content is a different paper. This must be resolved administratively, but it is not a presentational issue that can be corrected locally.

Circularity Check

0 steps flagged · score 0.0 of 10

No circular derivation is identifiable; the abstract's claimed PDE-operator derivation is simply absent from the supplied text, which is a disjoint attention-dynamics paper.

full rationale

The submitted text for arXiv:2508.09627 consists of an abstract announcing piG-Sp2GNO, a graph neural operator for PDEs, followed by the full text of "Attention's forward pass and Frank-Wolfe" (arXiv:2508.09628) by different authors. None of the claimed components -- spatio-spectral graph convolution, geometry-awareness strategies, upscaled-theory stochastic projection, physics-informed loss, or benchmarks -- appears in the body. There is therefore no derivation chain to reduce: no equation in the body defines the proposed operator, and no fitted parameter is renamed as a prediction. The absence of the method is a verifiability/integrity problem, not a circularity, since circularity requires exhibiting a specific reduction of a claimed result to its own inputs. The attention-dynamics text itself contains no circular steps: its Frank-Wolfe/Voronoi theorems are derived from the stated dynamics (eqs. 1.1-1.2, SA-infinity) via proofs, and its citations to [GLPR23, GLPR25] are background context rather than load-bearing uniqueness assumptions. Accordingly, the circularity score is 0, while the central claim's support is entirely missing -- a defect that must be handled as a correctness/verification issue outside this rubric.

Assumptions & free parameters 3 free parameters · 3 assumptions · 2 invented entities

Inventory taken from the abstract only, because the full text is an unrelated manuscript. The central claims rest on unstated architectural and training choices: the spatio-spectral graph construction, the two geometry-awareness encodings, the hybrid loss weights, and the projection parameters. All fitted values are unknown. The clearest ad hoc component is the upscaled-theory-inspired stochastic projection: it is presented as novel and its correctness as an integrator is assumed rather than derived or benchmarked.

free parameters (3)
  • πG-Sp²GNO architecture hyperparameters (depth, width, spectral modes, graph construction) = not specified
    Not given in the abstract; such choices are hand-tuned or fitted and materially affect the claimed benchmark accuracy.
  • Hybrid physics-informed loss combination weights and scheme orders = not specified
    The abstract announces a loss that merges a higher-order time-marching scheme with a stochastic projection scheme; the weights and orders are free choices that determine reported performance.
  • Geometry-awareness encoding parameters (two strategies) = not specified
    Two strategies for geometry awareness are introduced; their encoding scales and network placements are unstated at the abstract level.
assumptions (3)
  • domain assumption The governing PDE can be encoded as a differentiable residual or equivalent physics constraint usable in training.
    The abstract's 'simulation-free setup' depends on this; no equation or discretization is given.
  • domain assumption Spatio-spectral (graph spectral) decomposition preserves multiscale solution structure on complex and varying geometries.
    The claimed multiscale learning rests on this; spectral expressiveness on arbitrary domains is assumed.
  • ad hoc to paper The upscaled-theory-inspired stochastic projection is an accurate integration scheme for the physics-informed loss.
    Introduced as novel in the abstract; its accuracy as an integrator is asserted without demonstration.
invented entities (2)
  • πG-Sp²GNO architecture
    purpose: Learns solution operators of time-independent and time-dependent PDEs on complex geometries with limited labeled data.
    A new neural architecture asserted in the abstract; no equations, experiments, or external falsifiable handle appears in the attached manuscript.
  • Upscaled-theory-inspired stochastic projection scheme
    purpose: Component of the hybrid physics-informed loss intended to integrate physics accurately for time-dependent problems.
    Declared novel in the abstract with no mechanism described; there is no evidence outside this paper, and none inside the provided text.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Physics- and geometry-aware spatio-spectral graph neural operator for time-independent and time-dependent PDEs." pith.science (2026). https://pith.science/paper/PVJXSE63

@misc{pith2026250809627,
  author       = {Pith},
  title        = {Pith review of: Physics- and geometry-aware spatio-spectral graph neural operator for time-independent and time-dependent PDEs},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/PVJXSE63}},
  note         = {Machine review of arXiv:2508.09627}
}
abstract

Solving partial differential equations (PDEs) efficiently and accurately remains a cornerstone challenge in science and engineering, especially for problems involving complex geometries and limited labeled data. We introduce a Physics- and Geometry- Aware Spatio-Spectral Graph Neural Operator ($\pi$G-Sp$^2$GNO) for learning the solution operators of time-independent and time-dependent PDEs. The proposed approach first improves upon the recently developed Sp$^2$GNO by enabling geometry awareness and subsequently exploits the governing physics to learn the underlying solution operator in a simulation-free setup. While the spatio-spectral structure present in the proposed architecture allows multiscale learning, two separate strategies for enabling geometry awareness is introduced in this paper. For time dependent problems, we also introduce a novel hybrid physics informed loss function that combines higher-order time-marching scheme with upscaled theory inspired stochastic projection scheme. This allows accurate integration of the physics-information into the loss function. The performance of the proposed approach is illustrated on number of benchmark examples involving regular and complex domains, variation in geometry during inference, and time-independent and time-dependent problems. The results obtained illustrate the efficacy of the proposed approach as compared to the state-of-the-art physics-informed neural operator algorithms in the literature.

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

37 extracted references · 7 canonical work pages

  1. [1]

    mechanistic interpretability

    Attention’s forward pass and Frank-Wolfe Albert Alcalde FAU Erlangen–Nürnberg Borjan Geshkovski Inria & Sorbonne Université Domènec Ruiz-Balet Université Paris-Dauphine September 16, 2025 Abstract We study thehardmax limit of self-attention dynamics for token embed- dings obtained in the zero-temperature (β→ +∞) regime, and relate it to the finite-β setti...

  2. [4]

    France 2030

    Then, particles remain within a ball of radius𝜀 around this configuration at time𝑇1 until elapsing𝛾𝑡/𝜀∼𝑒β steps (Theo- rem 5.4). Interestingly, unlike for (SA∞), particles at finite-β do not remain stationary—in infinite time, they collapse to a single cluster at the origin (Proposition 5.1). Thus (SA∞) can be seen as a valid approximation of (1.1) up to𝑂...

  3. [8]

    What does bert look at? an analysis of bert’s attention

    [CKLM19] Kevin Clark, Urvashi Khandelwal, Omer Levy, and Christopher D Manning. What does bert look at? an analysis of bert’s attention. arXiv preprint arXiv:1906.04341,

  4. [10]

    Code available at https://github.com/borjanG/2025-transformers-frank-wolfe

    Axes are suppressed for visual clarity; all plots are rendered on the same met- ric scale. Code available at https://github.com/borjanG/2025-transformers-frank-wolfe. Taking a convex combination, ⟨𝐵(𝜆𝑥1 + (1−𝜆)𝑥2),𝑣𝑖⟩ =𝜆⟨𝐵𝑥1,𝑣𝑖⟩ + (1−𝜆)⟨𝐵𝑥2,𝑣𝑖⟩ and similarly for⟨𝐵(𝜆𝑥1 + (1−𝜆)𝑥2),𝑦⟩. Thus, ⟨𝐵(𝜆𝑥1 + (1−𝜆)𝑥2),𝑣𝑖⟩ ⩾⟨𝐵(𝜆𝑥1 + (1−𝜆)𝑥2),𝑦⟩, showing that𝜆𝑥1 + (1−𝜆...

  5. [11]

    Geometric dynamics of signal propagation predict trainability of transformers

    [CNQG24] Aditya Cowsik, Tamra Nebabu, Xiao-Liang Qi, and Surya Ganguli. Geometric dynamics of signal propagation predict trainability of transformers. arXiv preprint arXiv:2403.02579,

  6. [12]

    Synchronization on circles and spheres with non- linear interactions

    [CRMB24] Christopher Criscitiello, Quentin Rebjock, Andrew D McRae, and Nicolas Boumal. Synchronization on circles and spheres with non- linear interactions. arXiv preprint arXiv:2405.18273,

  7. [13]

    Setting the record straight on transformer oversmoothing

    [DBK24] Gbètondji JS Dovonon, Michael M Bronstein, and Matt J Kusner. Setting the record straight on transformer oversmoothing. arXiv preprint arXiv:2401.04301,

  8. [15]

    Hashattention: Semantic spar- sity for faster inference

    [DYC+24] Aditya Desai, Shuo Yang, Alejandro Cuadron, Matei Zaharia, Joseph E Gonzalez, and Ion Stoica. Hashattention: Semantic spar- sity for faster inference. arXiv preprint arXiv:2412.14468,

Show all 37 references
  1. [16]

    Large-time asymptotics in deep learning

    [EGPZ20] Carlos Esteve, Borjan Geshkovski, Dario Pighin, and Enrique Zuazua. Large-time asymptotics in deep learning. arXiv preprint arXiv:2008.02491,

  2. [18]

    Dynamic metastability in the self-attention model

    [GKPR24] Borjan Geshkovski, Hugo Koubbi, Yury Polyanskiy, and Philippe Rigollet. Dynamic metastability in the self-attention model. arXiv preprint arXiv:2410.06833,

  3. [19]

    When attention sink emerges in language models: An empirical view.arXiv preprint arXiv:2410.10781,

    [GPD+24] Xiangming Gu, Tianyu Pang, Chao Du, Qian Liu, Fengzhuo Zhang, Cunxiao Du, Ye Wang, and Min Lin. When attention sink emerges in language models: An empirical view.arXiv preprint arXiv:2410.10781,

  4. [20]

    Measure-to-measure interpolation using Transformers

    [GRRB24] Borjan Geshkovski, Philippe Rigollet, and Domènec Ruiz-Balet. Measure-to-measure interpolation using Transformers. arXiv preprint arXiv:2411.04551,

  5. [21]

    On the num- ber of modes of Gaussian kernel density estimators

    [GRS24] Borjan Geshkovski, Philippe Rigollet, and Yihang Sun. On the num- ber of modes of Gaussian kernel density estimators. arXiv preprint arXiv:2412.09080,

  6. [23]

    Ot-transformer: a continuous-time transformer architecture with optimal transport regularization

    [KLO25] Kelvin Kan, Xingjian Li, and Stanley Osher. Ot-transformer: a continuous-time transformer architecture with optimal transport regularization. arXiv preprint arXiv:2501.18793,

  7. [24]

    Convergence rate of Frank-Wolfe for non- convex objectives

    [LJ16] Simon Lacoste-Julien. Convergence rate of Frank-Wolfe for non- convex objectives. arXiv preprint arXiv:1607.00345,

  8. [26]

    The sparse frontier: Sparse atten- tion trade-offs in transformer llms

    [NLH+25] Piotr Nawrot, Robert Li, Renjie Huang, Sebastian Ruder, Kelly Marchisio, and Edoardo M Ponti. The sparse frontier: Sparse atten- tion trade-offs in transformer llms. arXiv preprint arXiv:2504.17768 ,

  9. [29]

    Revisiting over-smoothing in bert from the perspective of graph

    [SGX+22] Han Shi, Jiahui Gao, Hang Xu, Xiaodan Liang, Zhenguo Li, Lingpeng Kong, Stephen Lee, and James T Kwok. Revisiting over-smoothing in bert from the perspective of graph. arXiv preprint arXiv:2202.08625 ,

  10. [30]

    Exploiting sparsity for long context inference: Million token contexts on commodity gpus

    [SHK+25] Ryan Synk, Monte Hoover, John Kirchenbauer, Neel Jain, Alex Stein, Manli Shu, Josue Melendez Sanchez, Ramani Duraiswami, and Tom Goldstein. Exploiting sparsity for long context inference: Million token contexts on commodity gpus. arXiv preprint arXiv:2502.06766,

  11. [31]

    Prefixing attention sinks can mitigate activation outliers for large language model quantization

    [SPH+24] Seungwoo Son, Wonpyo Park, Woohyun Han, Kyuyeun Kim, and Jaeho Lee. Prefixing attention sinks can mitigate activation outliers for large language model quantization. arXiv preprint arXiv:2406.12016,

  12. [32]

    Solutions of stationary McKean-Vlasov equation on a high-dimensional sphere and other Riemannian manifolds

    [SS24] Anna Shalova and André Schlichting. Solutions of stationary McKean-Vlasov equation on a high-dimensional sphere and other Riemannian manifolds. arXiv preprint arXiv:2412.14813,

  13. [33]

    Residual connections and normalization can provably pre- vent oversmoothing in gnns

    51 [SWJS24] Michael Scholkemper, Xinyi Wu, Ali Jadbabaie, and Michael T Schaub. Residual connections and normalization can provably pre- vent oversmoothing in gnns. arXiv preprint arXiv:2406.02997,

  14. [34]

    Recurrent self-attention dy- namics: An energy-agnostic perspective from Jacobians

    [TK25] Akiyoshi Tomihari and Ryo Karakida. Recurrent self-attention dy- namics: An energy-agnostic perspective from Jacobians. arXiv preprint arXiv:2505.19458,

  15. [36]

    Transformer-based causal language models perform clustering

    [WV24] Xinbo Wu and Lav R Varshney. Transformer-based causal language models perform clustering. arXiv preprint arXiv:2402.12151,

  16. [37]

    Learning Gaussian Mixture Models via Transformer Mea- sure Flows

    52 [ZKPR25] Aleksandr Zimin, Anastasiia Kutakh, Yury Polyanskiy, and Philippe Rigollet. Learning Gaussian Mixture Models via Transformer Mea- sure Flows. In ICML 2025 Workshop on Methods and Opportunities at Small Scale,

  17. [1973]

    The impact of lora on the emergence of clusters in transformers

    [KBH24] Hugo Koubbi, Matthieu Boussard, and Louis Hernandez. The impact of lora on the emergence of clusters in transformers. arXiv preprint arXiv:2402.15415,

  18. [1996]

    Finding clustering algorithms in the transformer ar- chitecture

    [CHI+25] Kenneth L Clarkson, Lior Horesh, Takuya Ito, Charlotte Park, and Parikshit Ram. Finding clustering algorithms in the transformer ar- chitecture. arXiv preprint arXiv:2506.19125,

  19. [2004]

    Two failure modes of deep transformers and how to avoid them: a unified theory of signal prop- agation at initialisation

    [GG25] Alessio Giorlandino and Sebastian Goldt. Two failure modes of deep transformers and how to avoid them: a unified theory of signal prop- agation at initialisation. arXiv preprint arXiv:2505.24333,

  20. [2005]

    Self-attention net- works localize when qk-eigenspectrum concentrates

    [BHK24] Han Bao, Ryuichiro Hataya, and Ryo Karakida. Self-attention net- works localize when qk-eigenspectrum concentrates. arXiv preprint arXiv:2402.02098,

  21. [2007]

    Synchro- nization of mean-field models on the circle

    [PRY25] Yury Polyanskiy, Philippe Rigollet, and Andrew Yao. Synchro- nization of mean-field models on the circle. arXiv preprint arXiv:2507.22857,

  22. [2012]

    Towards understanding how attention mechanism works in deep learning

    [RZ24] Tianyu Ruan and Shihua Zhang. Towards understanding how attention mechanism works in deep learning. arXiv preprint arXiv:2412.18288,

  23. [2015]

    The geometry of tokens in in- ternal representations of large language models

    [VGP+25] Karthik Viswanathan, Yuri Gardinazzi, Giada Panerai, Alberto Caz- zaniga, and Matteo Biagetti. The geometry of tokens in in- ternal representations of large language models. arXiv preprint arXiv:2501.10573,

  24. [2016]

    Understanding and improving transformer from a multi-particle dynamic system point of view

    [LLH+19] Yiping Lu, Zhuohan Li, Di He, Zhiqing Sun, Bin Dong, Tao Qin, Liwei Wang, and Tie-Yan Liu. Understanding and improving transformer from a multi-particle dynamic system point of view. arXiv preprint arXiv:1906.02762,

  25. [2019]

    Continuum attention for neural operators

    [CKLS24] Edoardo Calvello, Nikola B Kovachki, Matthew E Levine, and An- drew M Stuart. Continuum attention for neural operators. arXiv preprint arXiv:2406.06486,

  26. [2021]

    Deepseekmoe: Towards ultimate expert specialization in mixture- of-experts language models

    [DDZ+24] Damai Dai, Chengqi Deng, Chenggang Zhao, RX Xu, Huazuo Gao, Deli Chen, Jiashi Li, Wangding Zeng, Xingkai Yu, Yu Wu, et al. Deepseekmoe: Towards ultimate expert specialization in mixture- of-experts language models. arXiv preprint arXiv:2401.06066,

  27. [2022]

    A unified perspective on the dynamics of deep transformers

    [CACP25] Valérie Castin, Pierre Ablin, José Antonio Carrillo, and Gabriel Peyré. A unified perspective on the dynamics of deep transformers. arXiv preprint arXiv:2501.18322,

  28. [2024]

    Quantitative clustering in mean-field transformer models

    47 [CLPR25] Shi Chen, Zhengjiang Lin, Yury Polyanskiy, and Philippe Rigollet. Quantitative clustering in mean-field transformer models. arXiv preprint arXiv:2504.14697,

  29. [2025]

    Breaking BERT: evaluating and optimizing sparsified attention

    [BZM22] Siddhartha Brahma, Polina Zablotskaia, and David Mimno. Breaking BERT: evaluating and optimizing sparsified attention. arXiv preprint arXiv:2210.03841,

Pith tools

Reviewed August 5, 2026 · model on record in the stance chip above.