Pith. sign in

Paper Citation Record · LEDGER

Standard Transformers Achieve the Minimax Rate in Nonparametric Regression with $C^{s,\lambda}$ Targets

As of 15 August 2026, this Paper Citation Record lists 62 of 62 outbound references and 1 inbound Pith citation observation for arXiv:2602.20555.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2602.20555 v2

Coverage vector

measured 62 of 62 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-02T21:28:35.859983Z

measured 63 of 63 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-27T05:08:00.711576Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T16:38:40.636427Z

Reference resolution

62 of 62 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved62
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2e501e7c-54b3-422f-bb71-a1e2a5cf5896 · outbound

This paper cites Pearson, 2 edition, 1974.

Standard Transformers Achieve the Minimax Rate in Nonparametric Regression with $C^{s,\lambda}$ Targets Pearson, 2 edition, 1974

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-02T21:28:27.605327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:28:27.605327Z digest=sha256:6c506f5fbce14f5b5dc1cdd806fd5d5fb1330202fa0b1322c4d4206e233d908b

Observation 2cb93dd9-59ed-4fef-814b-6c73a60efb61 · outbound

This paper cites Low-rank bottleneck in multi-head attention models.

Standard Transformers Achieve the Minimax Rate in Nonparametric Regression with $C^{s,\lambda}$ Targets Low-rank bottleneck in multi-head attention models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-02T21:28:27.662144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:28:27.662144Z digest=sha256:bf68d4bb9921689dc46ed3f001884fe25e300dc1cfe60001385e15046d0908bb

Observation 9dcec2f4-ee2f-438d-8a71-8b2b4c50cf3f · outbound

This paper cites an unresolved cited work.

Standard Transformers Achieve the Minimax Rate in Nonparametric Regression with $C^{s,\lambda}$ Targets Unresolved cited work

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-02T21:28:27.738468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:28:27.738468Z digest=sha256:39ea34bb3935a298e03e3e064ed541560096d3d3c02d91da40bd58280ef95c4b

Observation 3a154508-52de-4dab-b3d7-8c63331d65bd · outbound

This paper cites A unified framework for establishing the universal approximation of transformer-type architectures.

Standard Transformers Achieve the Minimax Rate in Nonparametric Regression with $C^{s,\lambda}$ Targets A unified framework for establishing the universal approximation of transformer-type architectures

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-02T21:28:27.890674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:28:27.890674Z digest=sha256:6b4fbbb0f9c6534632f4ea5ae17b5b2999f4ee380b6b3bc5e072f5065c60008d

Observation 0631de7d-335e-43e3-b836-7243964006c8 · outbound

This paper cites Efficient and Minimax Optimal In-context Nonparametric Regression with Transformers.

Standard Transformers Achieve the Minimax Rate in Nonparametric Regression with $C^{s,\lambda}$ Targets Efficient and Minimax Optimal In-context Nonparametric Regression with Transformers

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-02T21:28:28.002565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:28:28.002565Z digest=sha256:e2d90bf357ffb24d3e59290a643d90050ed59d374fdc048c4d5fa4f84eee5568

Observation 3015893b-4405-4162-b14b-292db3e5d535 · outbound

This paper cites Bert: Pre- training of deep bidirectional transformers for language understanding.

Standard Transformers Achieve the Minimax Rate in Nonparametric Regression with $C^{s,\lambda}$ Targets Bert: Pre- training of deep bidirectional transformers for language understanding

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-02T21:28:28.198176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:28:28.198176Z digest=sha256:d959490e02f5f14c18d3c201606f2132af89a87531345efa2e23629f35ca7637

Observation 74b568fe-6acc-4a49-a4bd-22c86ff064c3 · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale.

Standard Transformers Achieve the Minimax Rate in Nonparametric Regression with $C^{s,\lambda}$ Targets An image is worth 16x16 words: Transformers for image recognition at scale

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-02T21:28:28.302711Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:28:28.302711Z digest=sha256:e6b4971f3e9c057f3f2db1c65bb228e85fc4dc5067d03e4cb12719ee7b3da32f

Observation 8460b952-662f-4668-992d-b767a01fc7ef · outbound

This paper cites Inductive biases and variable creation in self-attention mechanisms.

Standard Transformers Achieve the Minimax Rate in Nonparametric Regression with $C^{s,\lambda}$ Targets Inductive biases and variable creation in self-attention mechanisms

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-02T21:28:28.421735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:28:28.421735Z digest=sha256:8cb13158d10001762da4c40ad71ba4d19a82fec651c74f41d29115915e6d5bbf

Observation c42b2c03-6729-453e-b86e-560979d2eaeb · outbound

This paper cites How do noise tails impact on deep relu networks?The Annals of Statistics, 52(4):1845–1871, 2024.

Standard Transformers Achieve the Minimax Rate in Nonparametric Regression with $C^{s,\lambda}$ Targets How do noise tails impact on deep relu networks?The Annals of Statistics, 52(4):1845–1871, 2024

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-02T21:28:28.506109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:28:28.506109Z digest=sha256:7c03f35a611b6e94bb34eafaf15fc2c0bec9010b554d4048648e90837a5679f4

Observation 21fca2ad-42cd-474f-8f6e-442a5c6a24a2 · outbound

This paper cites Deep neural networks for esti- mation and inference.Econometrica, 89(1):181–213, 2021.

Standard Transformers Achieve the Minimax Rate in Nonparametric Regression with $C^{s,\lambda}$ Targets Deep neural networks for esti- mation and inference.Econometrica, 89(1):181–213, 2021

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-02T21:28:28.660201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:28:28.660201Z digest=sha256:f6cd5ff805cf435010c5bfff5d26d3d9ea010ae560984d3636e88c29cc34ecc8

Observation 52cf793c-4063-4ea4-a189-cc84f8b183f8 · outbound

This paper cites Cambridge university press, 2021.

Standard Transformers Achieve the Minimax Rate in Nonparametric Regression with $C^{s,\lambda}$ Targets Cambridge university press, 2021

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-02T21:28:28.761626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:28:28.761626Z digest=sha256:eba9fa74a4220f53d45a3a2f712444e1301c1efc3c9df854c0826a90070385b9

Observation 50bfc956-8029-4e2f-9fa0-23f6530e8275 · outbound

This paper cites Approximation rates for neural networks with encodable weights in smoothness spaces.Neural Networks, 134:107–130, 2021.

Standard Transformers Achieve the Minimax Rate in Nonparametric Regression with $C^{s,\lambda}$ Targets Approximation rates for neural networks with encodable weights in smoothness spaces.Neural Networks, 134:107–130, 2021

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-02T21:28:28.870976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:28:28.870976Z digest=sha256:14fd3bd0054ab729a2d3c50f7a080fdff3c957d695e75d9a9f07651307dfe07a

Observation bcfbd33e-b72c-4476-9626-f494cb8692e3 · outbound

This paper cites On the rate of convergence of a classifier based on a transformer encoder.IEEE Transactions on Information Theory, 68(12):8139–8155, 2022.

Standard Transformers Achieve the Minimax Rate in Nonparametric Regression with $C^{s,\lambda}$ Targets On the rate of convergence of a classifier based on a transformer encoder.IEEE Transactions on Information Theory, 68(12):8139–8155, 2022

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-02T21:28:28.932299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:28:28.932299Z digest=sha256:f7aad19af16867d50a68101e40a6e0605e822764979d71e64047056d97e02521

Observation 29a4e6c4-1d11-4e20-8767-9cf148ade33f · outbound

This paper cites Understanding scaling laws with statisti- cal and approximation theory for transformer neural networks on intrinsically low- dimensional data.

Standard Transformers Achieve the Minimax Rate in Nonparametric Regression with $C^{s,\lambda}$ Targets Understanding scaling laws with statisti- cal and approximation theory for transformer neural networks on intrinsically low- dimensional data

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-02T21:28:29.050031Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:28:29.050031Z digest=sha256:8cc3b8557a9c463e5d36910eca17eba71455fce757f9dc40ecd51af686046607

Observation 84ba33a6-a5b1-478a-a558-35c34085ea7d · outbound

This paper cites Minimal width for universal property of deep rnn.Journal of Machine Learning Research, 24(121):1–41, 2023.

Standard Transformers Achieve the Minimax Rate in Nonparametric Regression with $C^{s,\lambda}$ Targets Minimal width for universal property of deep rnn.Journal of Machine Learning Research, 24(121):1–41, 2023

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-02T21:28:29.159551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:28:29.159551Z digest=sha256:e3201c70685581febeb59bbbcc36d171c97b91e2e6033c97ef453d705a9ac7f2

Observation 59978333-edd3-4ecd-a1e5-d7b4ee9e32ca · outbound

This paper cites Universal approximation with softmax attention.arXiv preprint arXiv:2504.15956, 2025.

Standard Transformers Achieve the Minimax Rate in Nonparametric Regression with $C^{s,\lambda}$ Targets Universal approximation with softmax attention.arXiv preprint arXiv:2504.15956, 2025

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-02T21:28:29.291247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:28:29.291247Z digest=sha256:19cf3c46e6e626784a61c425399679d7ba4dbc4117caf2c1ce4bbf464f1fe505

Observation a1ac003f-e4f8-457f-a66e-5c90c914d2ce · outbound

This paper cites Approximation rate of the transformer architecture for sequence modeling.

Standard Transformers Achieve the Minimax Rate in Nonparametric Regression with $C^{s,\lambda}$ Targets Approximation rate of the transformer architecture for sequence modeling

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-02T21:28:29.408007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:28:29.408007Z digest=sha256:a0a6f17284dc0bb4415fb4a6736c828e7fa178f63d78c2f598988a5db1bfb12f

Observation 4ae2c7f0-c677-4dca-b5c1-d7f38b197cb3 · outbound

This paper cites Approximation Bounds for Transformer Networks with Application to Regression.

Standard Transformers Achieve the Minimax Rate in Nonparametric Regression with $C^{s,\lambda}$ Targets Approximation Bounds for Transformer Networks with Application to Regression

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-02T21:28:29.502605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:28:29.502605Z digest=sha256:72bda9853b862f940a2d40101ac5a3a3416203ebbac466fe8539dea1c8f911c4

Observation 53c04e93-8632-4176-9f9f-cd0403ff7dca · outbound

This paper cites Transformers Can Overcome the Curse of Dimensionality: A Theoretical Study from an Approximation Perspective.

Standard Transformers Achieve the Minimax Rate in Nonparametric Regression with $C^{s,\lambda}$ Targets Transformers Can Overcome the Curse of Dimensionality: A Theoretical Study from an Approximation Perspective

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-02T21:28:29.665918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:28:29.665918Z digest=sha256:18b7fac094c219878351267c03cbb61044a5b4594e35ce162e3d604a197014f2

Observation edb1e354-523b-40eb-93c1-43e29e926da5 · outbound

This paper cites Deep nonparametric regression on approximate manifolds: Nonasymptotic error bounds with polynomial prefactors.The Annals of Statistics, 51(2):691–716, 2023.

Standard Transformers Achieve the Minimax Rate in Nonparametric Regression with $C^{s,\lambda}$ Targets Deep nonparametric regression on approximate manifolds: Nonasymptotic error bounds with polynomial prefactors.The Annals of Statistics, 51(2):691–716, 2023

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-02T21:28:29.779421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:28:29.779421Z digest=sha256:156c83b5fa67c64c12de4a968bf7c7e04b0a0504e98c5e72aeeae3bf25056e44

Observation 4a49078c-b132-492e-a8b4-c2e8951e1e92 · outbound

This paper cites Approximation bounds for recurrent neural networks with application to regression.arXiv preprint arXiv:2409.05577, 2024.

Standard Transformers Achieve the Minimax Rate in Nonparametric Regression with $C^{s,\lambda}$ Targets Approximation bounds for recurrent neural networks with application to regression.arXiv preprint arXiv:2409.05577, 2024

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-02T21:28:29.879046Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:28:29.879046Z digest=sha256:9f0cdeaf42364d1be3be57fe51de04a91dff00d1caef8c86cd94cabb8d03b790

Observation 03d88569-eff9-45f2-9fd9-c2b44b11131b · outbound

This paper cites Are transformers with one layer self-attention using low-rank weight matrices universal approximators? InThe Twelfth International Conference on Learning Representations, 2024.

Standard Transformers Achieve the Minimax Rate in Nonparametric Regression with $C^{s,\lambda}$ Targets Are transformers with one layer self-attention using low-rank weight matrices universal approximators? InThe Twelfth International Conference on Learning Representations, 2024

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-02T21:28:30.030422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:28:30.030422Z digest=sha256:eab6772518c51c75a5b3f9b493548b2fa64872203c4e0dca5ff53ff24415b738

Observation 5781e412-70e1-47c4-9cbf-75993ee51e79 · outbound

This paper cites On the optimal memorization capacity of trans- formers.

Standard Transformers Achieve the Minimax Rate in Nonparametric Regression with $C^{s,\lambda}$ Targets On the optimal memorization capacity of trans- formers

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-02T21:28:30.176620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:28:30.176620Z digest=sha256:6db4025306caadddce0cb07a395e32d3e52cd8dc8c22b286ceabc9503ef938b1

Observation b82921ca-47a1-4901-b929-5422a6d5df4a · outbound

This paper cites The lipschitz constant of self-attention.

Standard Transformers Achieve the Minimax Rate in Nonparametric Regression with $C^{s,\lambda}$ Targets The lipschitz constant of self-attention

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-02T21:28:30.303270Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:28:30.303270Z digest=sha256:098ad6f85b13388946f30a3f5f78cb69556499c36d9efd623ffb1d96bde5d52c

Observation 650bce6b-9287-4c58-9aca-20b450104f33 · outbound

This paper cites Provable memorization capac- ity of transformers.

Standard Transformers Achieve the Minimax Rate in Nonparametric Regression with $C^{s,\lambda}$ Targets Provable memorization capac- ity of transformers

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-02T21:28:30.429646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:28:30.429646Z digest=sha256:17905334cb0337c751972f21370c9b5ab1e2aca2e76785d74d99f31637ca1a3e

Observation 3e335400-d249-410d-a1f5-27555456dd6b · outbound

This paper cites Transformers are minimax optimal non- parametric in-context learners.

Standard Transformers Achieve the Minimax Rate in Nonparametric Regression with $C^{s,\lambda}$ Targets Transformers are minimax optimal non- parametric in-context learners

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-02T21:28:30.607066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:28:30.607066Z digest=sha256:028e4717815a879ea6bd4eb085a662d29d3d81c0bfb1993c05200168abe76876

Observation 4f45b3b4-d6e9-4c0e-8bb9-0d60c0de16be · outbound

This paper cites On the rate of convergence of fully connected deep neural network regression estimates.The Annals of Statistics, 49(4):2231–2249, 2021.

Standard Transformers Achieve the Minimax Rate in Nonparametric Regression with $C^{s,\lambda}$ Targets On the rate of convergence of fully connected deep neural network regression estimates.The Annals of Statistics, 49(4):2231–2249, 2021

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-02T21:28:30.744229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:28:30.744229Z digest=sha256:c03f64e896bc9ede0c9e5a39a93a4778ad71f9835bb8f21f7f0b899fef74bdad

Observation b668f606-30f5-4b2e-ae07-52c39ce70c3f · outbound

This paper cites Univer- sal approximation under constraints is possible with transformers.

Standard Transformers Achieve the Minimax Rate in Nonparametric Regression with $C^{s,\lambda}$ Targets Univer- sal approximation under constraints is possible with transformers

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-02T21:28:30.872899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:28:30.872899Z digest=sha256:567834f57e61af8253f8d5c9b29d5204208bff4d30b49faf4b23de687f9ae91f

Observation 260ffc38-d2a9-45d3-ae22-ae304c2de87f · outbound

This paper cites Approximation and optimization theory for linear continuous-time recurrent neural networks.Journal of Machine Learning Research, 23(42):1–85, 2022.

Standard Transformers Achieve the Minimax Rate in Nonparametric Regression with $C^{s,\lambda}$ Targets Approximation and optimization theory for linear continuous-time recurrent neural networks.Journal of Machine Learning Research, 23(42):1–85, 2022

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-02T21:28:31.044785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:28:31.044785Z digest=sha256:82800752b85022b3a195c44af79e5b9f07d8c9d4d143422375f5b9b074892cf4

Observation bda47ccb-5703-41b4-b655-844792366c2d · outbound

This paper cites Generalization analysis of transformers in distribu- tion regression.Neural Computation, 37(2):260–293, 2025.

Standard Transformers Achieve the Minimax Rate in Nonparametric Regression with $C^{s,\lambda}$ Targets Generalization analysis of transformers in distribu- tion regression.Neural Computation, 37(2):260–293, 2025

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-02T21:28:31.160241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:28:31.160241Z digest=sha256:7a0e0fea2590a2db8c4ea3a5dc4d8531084fc91d2d6cd2704e7dd3d0ad365f33

Observation 108ea5fe-2b55-4c77-9472-2eaa21ea1ebd · outbound

This paper cites Deep network approxi- mation for smooth functions.SIAM Journal on Mathematical Analysis, 53(5):5465– 5506, 2021.

Standard Transformers Achieve the Minimax Rate in Nonparametric Regression with $C^{s,\lambda}$ Targets Deep network approxi- mation for smooth functions.SIAM Journal on Mathematical Analysis, 53(5):5465– 5506, 2021

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-02T21:28:31.326706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:28:31.326706Z digest=sha256:7995b7ddf1c55cabd24274f325743c9e4a8bc26fb650b4d2e88260fe9e8d505d

Observation 5839742c-0e87-4728-bc4d-4e84d1d4e510 · outbound

This paper cites Upper and lower memory capacity bounds of transformers for next- token prediction.arXiv preprint arXiv:2405.13718, 2024.

Standard Transformers Achieve the Minimax Rate in Nonparametric Regression with $C^{s,\lambda}$ Targets Upper and lower memory capacity bounds of transformers for next- token prediction.arXiv preprint arXiv:2405.13718, 2024

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-02T21:28:31.429497Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:28:31.429497Z digest=sha256:7f5fdd46ca3fbc7038fbffce3cc76fefc54d4cea41e1727bbdb9482238bfb428

Observation 251460d0-5f51-4561-9b09-e150acfe6811 · outbound

This paper cites Memorization capacity of multi-head attention in transformers.

Standard Transformers Achieve the Minimax Rate in Nonparametric Regression with $C^{s,\lambda}$ Targets Memorization capacity of multi-head attention in transformers

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-02T21:28:31.589919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:28:31.589919Z digest=sha256:2507a1f221d418428227b39a8447611a2b648370fc29e1ceede7ed6af96d8f3e

Observation 5c1cf47c-8ded-4184-bbb6-c2965444c604 · outbound

This paper cites Rates of approximation by relu shallow neural networks.Journal of Complexity, 79:101784, 2023.

Standard Transformers Achieve the Minimax Rate in Nonparametric Regression with $C^{s,\lambda}$ Targets Rates of approximation by relu shallow neural networks.Journal of Complexity, 79:101784, 2023

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-02T21:28:31.739676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:28:31.739676Z digest=sha256:f9fafdd1795aa3b298c53ac322c55fc15540c6ea96c98b69325dfdbbfec6f7fb

Observation cf525ea2-cadb-48a4-95e8-8298ca2ab187 · outbound

This paper cites Adaptive approximation and generalization of deep neural network with intrinsic dimensionality.Journal of Machine Learning Research, 21(174):1–38, 2020.

Standard Transformers Achieve the Minimax Rate in Nonparametric Regression with $C^{s,\lambda}$ Targets Adaptive approximation and generalization of deep neural network with intrinsic dimensionality.Journal of Machine Learning Research, 21(174):1–38, 2020

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-02T21:28:31.921512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:28:31.921512Z digest=sha256:c33a4a0933e39edaf64a541ecb9bf51c19c459aac14071c8a6d0e2af6f21c54e

Observation 43b61371-d162-4bca-97b7-97550f2f9e86 · outbound

This paper cites Provable memorization via deep neural networks using sub-linear parameters.

Standard Transformers Achieve the Minimax Rate in Nonparametric Regression with $C^{s,\lambda}$ Targets Provable memorization via deep neural networks using sub-linear parameters

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-02T21:28:32.126644Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:28:32.126644Z digest=sha256:a89bf5393ea85118d980aac27f6f3dea2f1dee068e641d600af47de5340f39d4

Observation 742af754-03ef-4462-9503-40a2cfa02e69 · outbound

This paper cites Equivalence of approximation by convolu- tional neural networks and fully-connected networks.Proceedings of the American Mathematical Society, 148(4):1567–1581, 2020.

Standard Transformers Achieve the Minimax Rate in Nonparametric Regression with $C^{s,\lambda}$ Targets Equivalence of approximation by convolu- tional neural networks and fully-connected networks.Proceedings of the American Mathematical Society, 148(4):1567–1581, 2020

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-02T21:28:32.361578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:28:32.361578Z digest=sha256:ba1eb1e39311192ee2ec79377c53df43470a3370744f639117dce49442f9c359

Observation d868eeef-7558-442d-b0dc-6b00fee595bd · outbound

This paper cites Representational strengths and limitations of transformers.

Standard Transformers Achieve the Minimax Rate in Nonparametric Regression with $C^{s,\lambda}$ Targets Representational strengths and limitations of transformers

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-02T21:28:32.551226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:28:32.551226Z digest=sha256:c6c1cc0cb4d7c8e73a830a6ee6353e639e884764538466c3fc120cdafdb1e4d2

Observation 7eb5b0ad-7a9a-465e-a2b1-a5decf1adfa0 · outbound

This paper cites Nonparametric regression using deep neural net- works with relu activation function.Annals of statistics, 48(4):1875–1897, 2020.

Standard Transformers Achieve the Minimax Rate in Nonparametric Regression with $C^{s,\lambda}$ Targets Nonparametric regression using deep neural net- works with relu activation function.Annals of statistics, 48(4):1875–1897, 2020

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-02T21:28:32.650600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:28:32.650600Z digest=sha256:5b85e2da9fcd5bc2ffc76d0af24c0ab3d9796c0751c3be451fcb0292fa45642e

Observation dacd7337-14ac-4d67-9a69-ff3c8d659772 · outbound

This paper cites The kolmogorov–arnold representation theorem revisited.

Standard Transformers Achieve the Minimax Rate in Nonparametric Regression with $C^{s,\lambda}$ Targets The kolmogorov–arnold representation theorem revisited

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-02T21:28:32.770414Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:28:32.770414Z digest=sha256:50a51dfb22cfee3615fc577f7d543fcab44e9debca0cbf0514fc95ed7acd6699

Observation de14478b-5b9a-48a7-8353-971145ea79f6 · outbound

This paper cites Understanding in- context learning on structured manifolds: Bridging attention to kernel methods.

Standard Transformers Achieve the Minimax Rate in Nonparametric Regression with $C^{s,\lambda}$ Targets Understanding in- context learning on structured manifolds: Bridging attention to kernel methods

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-02T21:28:32.935975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:28:32.935975Z digest=sha256:8f4aa9c35fd56e796865cfadac60006e07a2aa894be2f2b5827c0bab7795f4f8

Observation 1bf983f1-d5f5-433a-a9ec-1c0b5da649fa · outbound

This paper cites Deep network approximation char- acterized by number of neurons.Communications in Computational Physics, 28(5), 2020.

Standard Transformers Achieve the Minimax Rate in Nonparametric Regression with $C^{s,\lambda}$ Targets Deep network approximation char- acterized by number of neurons.Communications in Computational Physics, 28(5), 2020

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-02T21:28:33.054527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:28:33.054527Z digest=sha256:4faff358c60e4a8212b5865c22627dfa597a1b90ac4465212ef90db94f646ad9

Observation d17c58a7-8caf-4b31-9793-6dd6a74c6613 · outbound

This paper cites Optimal approximation rate of relu networks in terms of width and depth.Journal de Math´ ematiques Pures et Appliqu´ ees, 157:101–135, 2022.

Standard Transformers Achieve the Minimax Rate in Nonparametric Regression with $C^{s,\lambda}$ Targets Optimal approximation rate of relu networks in terms of width and depth.Journal de Math´ ematiques Pures et Appliqu´ ees, 157:101–135, 2022

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-02T21:28:33.200221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:28:33.200221Z digest=sha256:4e0e8f866f9914972ccb0c8c90717f455ab88ab80a9a51c1ef6268b53cb56222

Observation e4e39434-fbd4-46d0-a726-e7376d1f7765 · outbound

This paper cites Optimal approximation rates for deep relu neural networks on sobolev and besov spaces.Journal of Machine Learning Research, 24(357):1–52, 2023.

Standard Transformers Achieve the Minimax Rate in Nonparametric Regression with $C^{s,\lambda}$ Targets Optimal approximation rates for deep relu neural networks on sobolev and besov spaces.Journal of Machine Learning Research, 24(357):1–52, 2023

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-02T21:28:33.378677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:28:33.378677Z digest=sha256:cba6412102646ed86a0da438f25d39bcea3c131edb93dfc4c679ca6cd637761f

Observation 73080394-9a7f-4dec-bd37-0cb148359b01 · outbound

This paper cites Optimal rates of convergence for nonparametric estimators.The annals of Statistics, pages 1348–1360, 1980.

Standard Transformers Achieve the Minimax Rate in Nonparametric Regression with $C^{s,\lambda}$ Targets Optimal rates of convergence for nonparametric estimators.The annals of Statistics, pages 1348–1360, 1980

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-02T21:28:33.503352Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:28:33.503352Z digest=sha256:98465daac693b449cb3a7a37604cda5a1375c1dae392e96211e04d2aaca52b0f

Observation 0dcf055d-49fd-40e5-87b4-b8361a039fd6 · outbound

This paper cites Adaptivity of deep reLU network for learning in besov and mixed smooth besov spaces: optimal rate and curse of dimensionality.

Standard Transformers Achieve the Minimax Rate in Nonparametric Regression with $C^{s,\lambda}$ Targets Adaptivity of deep reLU network for learning in besov and mixed smooth besov spaces: optimal rate and curse of dimensionality

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-02T21:28:33.641883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:28:33.641883Z digest=sha256:5f32d42c588b84a70da884a02ac5120bceb23c978c54e713b9ac7fd5a771d79c

Observation d7ad0ad1-320c-4f5a-96d9-0591b72e1610 · outbound

This paper cites Approximation and estimation ability of trans- formers for sequence-to-sequence functions with infinite dimensional input.

Standard Transformers Achieve the Minimax Rate in Nonparametric Regression with $C^{s,\lambda}$ Targets Approximation and estimation ability of trans- formers for sequence-to-sequence functions with infinite dimensional input

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-02T21:28:33.794727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:28:33.794727Z digest=sha256:14083de114a87478ca1af609a411d2c88f995f9a26bfca97ef231a78e835c415

Observation ec702a1a-f943-401f-8d3b-7efa7b87a55f · outbound

This paper cites Approximation of Permutation Invariant Polynomials by Transformers: Efficient Construction in Column-Size.

Standard Transformers Achieve the Minimax Rate in Nonparametric Regression with $C^{s,\lambda}$ Targets Approximation of Permutation Invariant Polynomials by Transformers: Efficient Construction in Column-Size

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-02T21:28:33.955158Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:28:33.955158Z digest=sha256:367a3319e09b2991d01c2a29deece3ea1d708a222b30e6c60b228a4171c74f45

Observation 83aa814b-aab6-44a8-9754-a959086b241a · outbound

This paper cites Weak convergence.

Standard Transformers Achieve the Minimax Rate in Nonparametric Regression with $C^{s,\lambda}$ Targets Weak convergence

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-02T21:28:34.094329Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:28:34.094329Z digest=sha256:66ed35a071924eacac962a8906d69bad5c1db0332df4ffa18252cf4dc4572189

Observation 99685651-1abc-433d-86a7-2760e342b9b3 · outbound

This paper cites On the optimal memorization power of reLU neural networks.

Standard Transformers Achieve the Minimax Rate in Nonparametric Regression with $C^{s,\lambda}$ Targets On the optimal memorization power of reLU neural networks

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-02T21:28:34.244520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:28:34.244520Z digest=sha256:1ba545a476461217af1f0c61df121c157f200a63d77583e8d8ba36f3cd8b9f5d

Observation 89b4f471-10f4-4ce3-9e09-92b1fd3deeec · outbound

This paper cites Attention is all you need.Advances in neural information processing systems, 30, 2017.

Standard Transformers Achieve the Minimax Rate in Nonparametric Regression with $C^{s,\lambda}$ Targets Attention is all you need.Advances in neural information processing systems, 30, 2017

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-02T21:28:34.372655Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:28:34.372655Z digest=sha256:7ffdecd249f0565056b4870b5d036fdfb0ef7a320304faa5a3a01cc52b948e3e

Observation 56d032e3-a846-4c70-a37c-2f027e3cdcb7 · outbound

This paper cites Prompt tuning transformers for data memorization.

Standard Transformers Achieve the Minimax Rate in Nonparametric Regression with $C^{s,\lambda}$ Targets Prompt tuning transformers for data memorization

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-02T21:28:34.546484Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:28:34.546484Z digest=sha256:2105318efe599eb59c6ef8ab8cd617175bc8dafb67de7268406a6cce2e955c21

Observation c22ce1a9-de37-476e-9d5d-3a8213da738f · outbound

This paper cites an unresolved cited work.

Standard Transformers Achieve the Minimax Rate in Nonparametric Regression with $C^{s,\lambda}$ Targets Unresolved cited work

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-02T21:28:34.685239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:28:34.685239Z digest=sha256:865029d8e35dc562522d961714426b96499ae257dae4c2bd7108660f5428d746

Observation f2100755-2050-4fab-b5ba-e3c77e50f39b · outbound

This paper cites Statistically meaningful approximation: a case study on approximating turing machines with transformers.

Standard Transformers Achieve the Minimax Rate in Nonparametric Regression with $C^{s,\lambda}$ Targets Statistically meaningful approximation: a case study on approximating turing machines with transformers

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-02T21:28:34.856079Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:28:34.856079Z digest=sha256:24e13b1ca204be967b976890b8ee80747f33f090cb6dc5885440f033f95e818b

Observation 986e253b-7c69-4463-8e21-357e00d1a999 · outbound

This paper cites On the optimal approximation of Sobolev and Besov functions using deep ReLU neural networks.

Standard Transformers Achieve the Minimax Rate in Nonparametric Regression with $C^{s,\lambda}$ Targets On the optimal approximation of Sobolev and Besov functions using deep ReLU neural networks

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-02T21:28:34.968470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:28:34.968470Z digest=sha256:560a46a5dfa6c2696e48e8004366a2c243bded301286c45167af94c832be177e

Observation 14739f9b-4530-47ab-b1cb-451be7a90070 · outbound

This paper cites Nonparametric regression using over- parameterized shallow relu neural networks.Journal of Machine Learning Research, 25(165):1–35, 2024.

Standard Transformers Achieve the Minimax Rate in Nonparametric Regression with $C^{s,\lambda}$ Targets Nonparametric regression using over- parameterized shallow relu neural networks.Journal of Machine Learning Research, 25(165):1–35, 2024

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-02T21:28:35.048382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:28:35.048382Z digest=sha256:9e65fbfb99b4d69cd6b00855dce6b6686988c8d15b1f1b5eca900d3352546804

Observation f319b704-d617-4bd8-85bb-28ec3cf8b13d · outbound

This paper cites Error bounds for approximations with deep relu networks.Neural networks, 94:103–114, 2017.

Standard Transformers Achieve the Minimax Rate in Nonparametric Regression with $C^{s,\lambda}$ Targets Error bounds for approximations with deep relu networks.Neural networks, 94:103–114, 2017

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-02T21:28:35.137363Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:28:35.137363Z digest=sha256:3d65065c2af29248cf5b12113afb1920fd763fdafe806fc556d11b6bf8105db4

Observation f64b5c00-4260-436c-9af7-3131b2b60151 · outbound

This paper cites Optimal approximation of continuous functions by very deep relu networks.

Standard Transformers Achieve the Minimax Rate in Nonparametric Regression with $C^{s,\lambda}$ Targets Optimal approximation of continuous functions by very deep relu networks

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-02T21:28:35.306182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:28:35.306182Z digest=sha256:0d7588d66728fafc8622c4c177cf47cb24346354ceb1ca51c2f7f4e578d6a390

Observation e94b9401-45af-4c6b-8d36-f462d85f6de9 · outbound

This paper cites Are transformers universal approximators of sequence-to-sequence func- tions? InInternational Conference on Learning Representations, 2020.

Standard Transformers Achieve the Minimax Rate in Nonparametric Regression with $C^{s,\lambda}$ Targets Are transformers universal approximators of sequence-to-sequence func- tions? InInternational Conference on Learning Representations, 2020

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-02T21:28:35.487880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:28:35.487880Z digest=sha256:fcb4a554c0a2ea7974e63da8736967c984425f5457a0c2114c41015d0553b236

Observation ce2b942a-2dfd-40b5-8197-b5f2af50374d · outbound

This paper cites O (n) connections are expressive enough: Universal approximability of sparse transformers.Advances in Neural Information Processing Systems, 33:13783–13794, 2020.

Standard Transformers Achieve the Minimax Rate in Nonparametric Regression with $C^{s,\lambda}$ Targets O (n) connections are expressive enough: Universal approximability of sparse transformers.Advances in Neural Information Processing Systems, 33:13783–13794, 2020

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-02T21:28:35.601184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:28:35.601184Z digest=sha256:0ab3a7b814a3492223a46657e29fbde68ec3d93d08135192b00fabce9f512d73

Observation 4ac02717-490a-42a6-bc5b-d1643ab170dc · outbound

This paper cites Theory of deep convolutional neural networks: Downsampling.

Standard Transformers Achieve the Minimax Rate in Nonparametric Regression with $C^{s,\lambda}$ Targets Theory of deep convolutional neural networks: Downsampling

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-02T21:28:35.721401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:28:35.721401Z digest=sha256:5555b7b20bcce3ef42ee0469a4ee24cd5d75da2335c736c9ff047609ba180f0e

Observation 89d356a7-abf0-415b-bbb9-5f6fefffdbca · outbound

This paper cites Universality of deep convolutional neural networks.Applied and computational harmonic analysis, 48(2):787–794, 2020.

Standard Transformers Achieve the Minimax Rate in Nonparametric Regression with $C^{s,\lambda}$ Targets Universality of deep convolutional neural networks.Applied and computational harmonic analysis, 48(2):787–794, 2020

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-02T21:28:35.859983Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:28:35.859983Z digest=sha256:ba6d7901262cd5bd3f9045eb83ebd4939e76cc7a384fab586c3c704791ddc29e

Pith citing papers

Observation 68c41094-657e-4173-97e4-2c0684c89151 · inbound

Generalization Bounds for Transformer-Based Next-Token Prediction in a Language Model cites this paper.

Generalization Bounds for Transformer-Based Next-Token Prediction in a Language Model Standard Transformers Achieve the Minimax Rate in Nonparametric Regression with $C^{s,\lambda}$ Targets

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-07-29T02:24:15.834827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-27T05:08:00.711576Z digest=sha256:6bf81f92af1445e42233be73c9a4fbb91dca423d766991a0f6b19d0f3dd10590