Pith. sign in

Paper Citation Record · LEDGER

Are Transformers universal approximators of sequence-to-sequence functions?

As of 20 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 38 inbound Pith citation observations for arXiv:1912.10077.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
1912.10077 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 38 of 38 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 38 of 38 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T05:49:39.647398Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T10:29:44.309933Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation cb614542-6b76-4fc1-a6f8-a55fb35bd9f5 · inbound

Learning Elementary Cellular Automata with Transformers cites this paper.

Learning Elementary Cellular Automata with Transformers Are Transformers universal approximators of sequence-to-sequence functions?

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T04:27:33.578985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:27:33.578985Z digest=sha256:c65f8c7d5a336a2f40a12cf1fe3bab1467d032182ade40d7d54d37ec571e43fa

Observation f2936d46-d892-46e3-80a6-da59772f8ad4 · inbound

Rethinking Addressing in Language Models via Contexualized Equivariant Positional Encoding cites this paper.

Rethinking Addressing in Language Models via Contexualized Equivariant Positional Encoding Are Transformers universal approximators of sequence-to-sequence functions?

Reference 117

Resolution
unresolved
no resolver link, observed 2026-08-10T22:49:45.195898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:49:45.195898Z digest=sha256:779b5c703a167667dc6b2be836fe77b547573c641af4d3310b7c6b8bda281b29

Observation 49785f9d-c8a3-4613-90e8-314de27e1ca1 · inbound

Learning Spectral Methods by Transformers cites this paper.

Learning Spectral Methods by Transformers Are Transformers universal approximators of sequence-to-sequence functions?

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-10T22:36:51.338488Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:36:51.338488Z digest=sha256:ef2ae591ce8cd84badd7bd2251cfe6d5cab29b4297adebb37eb6004a4919c35f

Observation 88569954-70d3-4fc7-8e03-801cf35823aa · inbound

Video Latent Flow Matching: Optimal Polynomial Projections for Video Interpolation and Extrapolation cites this paper.

Video Latent Flow Matching: Optimal Polynomial Projections for Video Interpolation and Extrapolation Are Transformers universal approximators of sequence-to-sequence functions?

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-09T18:51:12.699988Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T18:51:12.699988Z digest=sha256:e2d44fefefb7ffb8aa3a84743a0bad38870dbb765b64a44989c80c3bc9acc980

Observation 59f2d4e1-2304-4a1e-888d-d92c10fbcea7 · inbound

Emergent Stack Representations in Modeling Counter Languages Using Transformers cites this paper.

Emergent Stack Representations in Modeling Counter Languages Using Transformers Are Transformers universal approximators of sequence-to-sequence functions?

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-09T15:26:53.482240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T15:26:53.482240Z digest=sha256:b16354cdc6013611976b9aa2133119ab75dc39c3466f1c1bc447ab90fc7d9c5f

Observation 70a371b3-fb9d-458f-818c-9b2bd6025f81 · inbound

Transformers versus the EM Algorithm in Multi-class Clustering cites this paper.

Transformers versus the EM Algorithm in Multi-class Clustering Are Transformers universal approximators of sequence-to-sequence functions?

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-08T17:13:25.754027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T17:13:25.754027Z digest=sha256:4c73fa0f8d1fd40d0426d47a65991b3c04d7eaf2b680647ec6dfa9c929604464

Observation db5b7be2-258f-440b-a1a8-592ca5d26ecb · inbound

Solving Empirical Bayes via Transformers cites this paper.

Solving Empirical Bayes via Transformers Are Transformers universal approximators of sequence-to-sequence functions?

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-07T20:21:58.396306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:21:58.396306Z digest=sha256:3e01c8c39ea8fcd638785fb6f8197d956d61d72f7f631b6565887921d8470fa3

Observation 3bfb064f-9405-46f3-af0f-de6310a4f201 · inbound

Attention Mechanism, Max-Affine Partition, and Universal Approximation cites this paper.

Attention Mechanism, Max-Affine Partition, and Universal Approximation Are Transformers universal approximators of sequence-to-sequence functions?

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-16T05:49:39.647398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:49:39.647398Z digest=sha256:dfbb74b0390fc3a5931b5277fa01fe0656c5e0de99e0f2d61418db9b8732d9f5

Observation deadf153-e3be-404f-ad12-0701ed62cc0b · inbound

Rethinking Causal Mask Attention for Vision-Language Inference cites this paper.

Rethinking Causal Mask Attention for Vision-Language Inference Are Transformers universal approximators of sequence-to-sequence functions?

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T14:33:33.425489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:33:33.425489Z digest=sha256:2ee8edc823ae85e289a06434cb150bd696a32be8cb1bfb57da08342d7f077382

Observation 908a9664-6277-4c40-80e4-df1aa557d70f · inbound

GSA-TTS : Toward Zero-Shot Speech Synthesis based on Gradual Style Adaptor cites this paper.

GSA-TTS : Toward Zero-Shot Speech Synthesis based on Gradual Style Adaptor Are Transformers universal approximators of sequence-to-sequence functions?

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:07.766917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:07.766917Z digest=sha256:72a1de6ce2883472b32cf58a908e21f503cedc4145b756c6ef44a13c79d1f7b6

Observation f8a408ed-f361-4636-a630-e614ce3f8825 · inbound

Transformers Are Universally Consistent cites this paper.

Transformers Are Universally Consistent Are Transformers universal approximators of sequence-to-sequence functions?

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:06.521514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:35:06.521514Z digest=sha256:884ac46e365524b91ace99d2fe7163be793b3f6f334560b3ca9926d162ef2493

Observation a2084ee9-b7bd-45a0-8515-4d9860c7f7c7 · inbound

Existing Large Language Model Unlearning Evaluations Are Inconclusive cites this paper.

Existing Large Language Model Unlearning Evaluations Are Inconclusive Are Transformers universal approximators of sequence-to-sequence functions?

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T12:05:56.351628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:05:56.351628Z digest=sha256:419962e0b3be18c69df96ad45edecc223a40e10181a28e36d531d4e047665d04

Observation e11c36a9-ba77-42ee-894a-2d3458285132 · inbound

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization cites this paper.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Are Transformers universal approximators of sequence-to-sequence functions?

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T06:10:21.358940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:10:21.358940Z digest=sha256:2217962f683aa22a292e691cb8765700954f4ebf51b7b301acea056363a2495b

Observation 84a07a37-1edc-468f-b252-8735743043ca · inbound

Theoretical Analysis of Positional Encodings in Transformer Models: Impact on Expressiveness and Generalization cites this paper.

Theoretical Analysis of Positional Encodings in Transformer Models: Impact on Expressiveness and Generalization Are Transformers universal approximators of sequence-to-sequence functions?

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:31.995796Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:19:31.995796Z digest=sha256:32e8ea3d2f67161c10766f57d1a75f7b08e2aa8668d02d2444dd2b78d9e56b54

Observation 036772fa-2944-4984-b8d5-206c295f6300 · inbound

Eliciting Fine-Tuned Transformer Capabilities via Inference-Time Techniques cites this paper.

Eliciting Fine-Tuned Transformer Capabilities via Inference-Time Techniques Are Transformers universal approximators of sequence-to-sequence functions?

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T05:38:35.295805Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:38:35.295805Z digest=sha256:981246fe0ae1b7c8ddbd1d5c3b154ffe1cdea2aa8825b7ea9109f86d173275b1

Observation defe9de6-530e-4ce4-ae4c-6a86dcb06e0b · inbound

Intrinsic and Extrinsic Organized Attention: Softmax Invariance and Network Sparsity cites this paper.

Intrinsic and Extrinsic Organized Attention: Softmax Invariance and Network Sparsity Are Transformers universal approximators of sequence-to-sequence functions?

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T19:38:27.124508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:38:27.124508Z digest=sha256:a8915359170ecca4418090a9a897597c4394b0b5be00a56572f831c780ecc619

Observation e24742c2-24f2-4c94-9fb9-1d5775b6f79a · inbound

Time Resolution Independent Operator Learning cites this paper.

Time Resolution Independent Operator Learning Are Transformers universal approximators of sequence-to-sequence functions?

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T20:34:09.205665Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:34:09.205665Z digest=sha256:25271e925a1d46435d5847b68a3bc0d167c546e0878b165275035d0f8828223e

Observation 36941140-0c54-4d2b-a716-20ac10fdcd4c · inbound

On the Mathematical Impossibility of Safe Universal Approximators cites this paper.

On the Mathematical Impossibility of Safe Universal Approximators Are Transformers universal approximators of sequence-to-sequence functions?

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T20:41:55.587713Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:41:55.587713Z digest=sha256:0c4747cc197f329340f5a46473941a4b0d46c0d75730184e6a8c1147f6becd92

Observation 53fa0876-b32a-4d23-8dda-130b4925efb4 · inbound

Decoding Consumer Preferences Using Attention-Based Language Models cites this paper.

Decoding Consumer Preferences Using Attention-Based Language Models Are Transformers universal approximators of sequence-to-sequence functions?

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-06T14:51:53.078612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:51:53.078612Z digest=sha256:1544c23766efe89e260f85617c957c84628764a381981fe986a2ed59c7485ea5

Observation 23bd166a-b334-430f-83da-9aa628c163ad · inbound

Graded Transformers cites this paper.

Graded Transformers Are Transformers universal approximators of sequence-to-sequence functions?

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T17:54:10.144066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:54:10.144066Z digest=sha256:9e55ef0b727592855b95d8bfda33b598bf056c0d8f175e21773a798e298e4a3a

Observation 50e79fa4-7a8d-4670-8b8a-57c90e4b4820 · inbound

Context-aware Rotary Position Embedding cites this paper.

Context-aware Rotary Position Embedding Are Transformers universal approximators of sequence-to-sequence functions?

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T11:09:08.432049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T11:09:08.432049Z digest=sha256:10bb13b390274564e83c2bf83c9656ddf04b3d9a5cfda52edd850441902192ba

Observation 5f78fd59-44ff-486a-a621-39ee8527af08 · inbound

Beyond Memorization: Extending Reasoning Depth with Recurrence, Memory and Test-Time Compute Scaling cites this paper.

Beyond Memorization: Extending Reasoning Depth with Recurrence, Memory and Test-Time Compute Scaling Are Transformers universal approximators of sequence-to-sequence functions?

Reference 76

Resolution
verified exact
arxiv_id, observed 2026-05-18T20:51:50.868831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-18T20:49:26.966293Z digest=sha256:3e856c83376ce4cc59b01d29d33b345943c6e3a47257a6e12838058faafe7c30

Observation fcd3909e-358c-4cc6-a20f-584c92128715 · inbound

Unraveling Syntax: Language Modeling and the Substructure of Grammars cites this paper.

Unraveling Syntax: Language Modeling and the Substructure of Grammars Are Transformers universal approximators of sequence-to-sequence functions?

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-04T12:47:21.283300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:47:21.283300Z digest=sha256:14866d1f8208f740b33ed959e0625a9fc8552dc2258ad9cc5781bcff18badc68

Observation 0f10c16c-47ba-4cd8-a5e9-c11b220203b2 · inbound

Key and Value Weights Are Probably All You Need: On the Necessity of the Query, Key, Value weight Triplet in Self-Attention Transformers cites this paper.

Key and Value Weights Are Probably All You Need: On the Necessity of the Query, Key, Value weight Triplet in Self-Attention Transformers Are Transformers universal approximators of sequence-to-sequence functions?

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-18T03:40:50.496470Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-18T03:38:36.932424Z digest=sha256:5342188ebe3c00985546225fc8879e4ddfaedb2045fcb494182e57a6f073cbe2

Observation d6286ac3-109e-4068-89e0-8b0de8876a56 · inbound

Universal Approximation Theorems for Dynamical Systems with Infinite-Time Horizon Guarantees cites this paper.

Universal Approximation Theorems for Dynamical Systems with Infinite-Time Horizon Guarantees Are Transformers universal approximators of sequence-to-sequence functions?

Reference 174

Resolution
unresolved
no resolver link, observed 2026-08-03T03:21:49.824957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:21:49.824957Z digest=sha256:d0ddcc60755762422e2b2ffeeca5ab7e60decd97e2665aa6fecf380d6dc2fe4b

Observation f01a6dbb-b2cd-48d4-9662-d8b2c6245f2a · inbound

Gating Enables Curvature: A Geometric Expressivity Gap in Attention cites this paper.

Gating Enables Curvature: A Geometric Expressivity Gap in Attention Are Transformers universal approximators of sequence-to-sequence functions?

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-10T11:20:11.376181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-10T11:10:47.526833Z digest=sha256:94eeff7e4e00d24555dbb04ce2ba6f5e8666899566e357cecb0ea2d4ac3e60b8

Observation 2ecaae8f-ef3e-4e72-b0e9-bd0d77e79bfc · inbound

Continuous transformations of probability measures and their transport representations cites this paper.

Continuous transformations of probability measures and their transport representations Are Transformers universal approximators of sequence-to-sequence functions?

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-10T06:56:47.232493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-10T06:56:24.472027Z digest=sha256:a7ec823298631bb136470aae79e7eadec320800d881a3eee30f6d9b20a5d0e46

Observation 79ee8e96-593f-4e46-a052-e5a04f30d8ae · inbound

Progressive Approximation in Deep Residual Networks: Theory and Validation cites this paper.

Progressive Approximation in Deep Residual Networks: Theory and Validation Are Transformers universal approximators of sequence-to-sequence functions?

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T21:41:16.148117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-08T04:35:14.636981Z digest=sha256:9d6d67d6a9ae638c8fc59fdff07fa366cbc700832c66f465c8638541a2a2a0c7

Observation ec91d9a6-9208-4372-b079-1b6ddd27492c · inbound

How Does Attention Help? Insights from Random Matrices on Signal Recovery from Sequence Models cites this paper.

How Does Attention Help? Insights from Random Matrices on Signal Recovery from Sequence Models Are Transformers universal approximators of sequence-to-sequence functions?

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:05:58.700918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-11T00:52:08.399401Z digest=sha256:db0099047771def9c899683a96129eb1f71d4d3a3fea941cca3c184c5e1494a9

Observation c2348a66-465a-4e22-8a00-5423acb01eb5 · inbound

A generative pre-trained transformer with Kerr-soliton attention cites this paper.

A generative pre-trained transformer with Kerr-soliton attention Are Transformers universal approximators of sequence-to-sequence functions?

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-06-30T15:04:46.486774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-30T14:57:08.127650Z digest=sha256:3d13700458eb7ceae7cd7a8b9bfb60438c5102dc229e569e0ad3f4d538d07562

Observation 31e7d9a7-ca2e-478e-a3d4-09830604fd58 · inbound

The Variance Brain Foundation Models Forgot: Third-Order Statistics Predict Cognition Where Billion-Parameter Models Fail cites this paper.

The Variance Brain Foundation Models Forgot: Third-Order Statistics Predict Cognition Where Billion-Parameter Models Fail Are Transformers universal approximators of sequence-to-sequence functions?

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-06-28T19:12:34.913911Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-06-28T19:07:34.240009Z digest=sha256:a891daff00f22c148b32729d12c4181f6fe9ac22d530fdbc2cb7e3e8857a26d6

Observation d4e3d0f4-2e32-4f57-a427-31644db06327 · inbound

Imbuing Large Language Models with Bidirectional Logic for Robust Chain Repair cites this paper.

Imbuing Large Language Models with Bidirectional Logic for Robust Chain Repair Are Transformers universal approximators of sequence-to-sequence functions?

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-07-02T08:46:48.999772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-28T05:46:17.481698Z digest=sha256:18da1dda736595a90ac4ad04ff094fc625ef54bebf2760a3498581e5e5838809

Observation f16d0fac-d2c0-41da-9f12-fc68533b825a · inbound

Convergence of Gradient Descent for General Neural Network Architectures Beyond the NTK Regime cites this paper.

Convergence of Gradient Descent for General Neural Network Architectures Beyond the NTK Regime Are Transformers universal approximators of sequence-to-sequence functions?

Reference 110

Resolution
verified exact
arxiv_id, observed 2026-07-04T10:29:44.311684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-06-26T08:53:46.285233Z digest=sha256:ac849796b71a8abe9ae3498208fbc7bca8b25c5b67d9cdf4b0bedca77fc54220

Observation 21bb62ce-7cc0-4e08-8803-e89fb4644a7a · inbound

Pre-Strings Lectures on Artificial Intelligence cites this paper.

Pre-Strings Lectures on Artificial Intelligence Are Transformers universal approximators of sequence-to-sequence functions?

Reference 29

Resolution
unresolved
no resolver link, observed 2026-07-12T06:14:03.658427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T06:14:03.658427Z digest=sha256:dbfa11efd3fd5af3beeb5920394092e3a37fb30a480ae87556676f65ebb7699d

Observation f95da0d6-cff4-4aa6-8be6-00989bc4ecdb · inbound

On Transformer Dynamics cites this paper.

On Transformer Dynamics Are Transformers universal approximators of sequence-to-sequence functions?

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-02T05:49:28.691589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:49:28.691589Z digest=sha256:2c4ab1f64f174a45417ce4de5b33e0e5ad4b6bb1f5cf2b144511e407a7359811

Observation 847e8761-c387-4425-be05-f1e3f3272ba1 · inbound

Training with (Swap) Regret Loss in a Single-Layer Self-Attention Model: A Case Study on the Probability Simplex cites this paper.

Training with (Swap) Regret Loss in a Single-Layer Self-Attention Model: A Case Study on the Probability Simplex Are Transformers universal approximators of sequence-to-sequence functions?

Reference 213

Resolution
unresolved
no resolver link, observed 2026-07-31T23:52:08.233130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T23:52:08.233130Z digest=sha256:9684b5ca09e996f022b412173af6ccf452e86d14f1b2235c52c486010e746d84

Observation 4257aa79-2604-4872-89a4-6e4b6ef5cb9a · inbound

Identifiability-Aware Source Apportionment in City-Scale Advection-Diffusion Systems cites this paper.

Identifiability-Aware Source Apportionment in City-Scale Advection-Diffusion Systems Are Transformers universal approximators of sequence-to-sequence functions?

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-04T01:37:40.900908Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T01:37:40.900908Z digest=sha256:9fe4c6127ac440d653b34561440ce5cdf9a19cbe9bf11ee8bc3e2e4e0f7b47f4

Observation b029877b-2e9f-4876-aa2c-e7bad65ca26b · inbound

Faster Query-Key Learning Sharpens Attention in Self-Attention Models cites this paper.

Faster Query-Key Learning Sharpens Attention in Self-Attention Models Are Transformers universal approximators of sequence-to-sequence functions?

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-15T14:39:18.186873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:39:18.186873Z digest=sha256:435aba9fae4edbad526997f8ea1590529ebc4f5bfc20a70349630c5f6cc2ed87