Pith. sign in

Paper Citation Record · LEDGER

Are Transformers universal approximators of sequence-to-sequence functions?

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 28 inbound Pith citation observations for arXiv:1912.10077.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
1912.10077 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 28 of 28 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 28 of 28 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T20:21:58.396306Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T10:29:44.309933Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation db5b7be2-258f-440b-a1a8-592ca5d26ecb · inbound

Solving Empirical Bayes via Transformers cites this paper.

Solving Empirical Bayes via Transformers Are Transformers universal approximators of sequence-to-sequence functions?

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-07T20:21:58.396306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:21:58.396306Z digest=sha256:8ecf3a9f68848cdac0ed7d451d69fd70ebf05a82d21d936f5ceab162245fe917

Observation deadf153-e3be-404f-ad12-0701ed62cc0b · inbound

Rethinking Causal Mask Attention for Vision-Language Inference cites this paper.

Rethinking Causal Mask Attention for Vision-Language Inference Are Transformers universal approximators of sequence-to-sequence functions?

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T14:33:33.425489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:33:33.425489Z digest=sha256:b6d7f68373c0a76e60344d91ff5677290b42569f1943d87115b288f53cc6f2e7

Observation 908a9664-6277-4c40-80e4-df1aa557d70f · inbound

GSA-TTS : Toward Zero-Shot Speech Synthesis based on Gradual Style Adaptor cites this paper.

GSA-TTS : Toward Zero-Shot Speech Synthesis based on Gradual Style Adaptor Are Transformers universal approximators of sequence-to-sequence functions?

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:07.766917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:07.766917Z digest=sha256:8241cd813a5909fcb8bbfba0a86f90f5c8aa35ab5fa2ff4de25148fd46d66506

Observation f8a408ed-f361-4636-a630-e614ce3f8825 · inbound

Transformers Are Universally Consistent cites this paper.

Transformers Are Universally Consistent Are Transformers universal approximators of sequence-to-sequence functions?

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:06.521514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:35:06.521514Z digest=sha256:b357745cc68878022e4c285203804413b4ce3461171e108dc96a1b8f90833490

Observation a2084ee9-b7bd-45a0-8515-4d9860c7f7c7 · inbound

Existing Large Language Model Unlearning Evaluations Are Inconclusive cites this paper.

Existing Large Language Model Unlearning Evaluations Are Inconclusive Are Transformers universal approximators of sequence-to-sequence functions?

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T12:05:56.351628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:05:56.351628Z digest=sha256:a50e011baf0239606415c8350b3609bcc1b2f180b0247afc5da2e3db83729125

Observation e11c36a9-ba77-42ee-894a-2d3458285132 · inbound

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization cites this paper.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Are Transformers universal approximators of sequence-to-sequence functions?

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T06:10:21.358940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:10:21.358940Z digest=sha256:ed44eda78bbea3de002aad41cf498044e21e105a73fe10b8084daa57f3e482e7

Observation 84a07a37-1edc-468f-b252-8735743043ca · inbound

Theoretical Analysis of Positional Encodings in Transformer Models: Impact on Expressiveness and Generalization cites this paper.

Theoretical Analysis of Positional Encodings in Transformer Models: Impact on Expressiveness and Generalization Are Transformers universal approximators of sequence-to-sequence functions?

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:31.995796Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:19:31.995796Z digest=sha256:13cc70bcc38db7527800d28c6340350418cbc81879cae30f00ef6dbaaa3db425

Observation 036772fa-2944-4984-b8d5-206c295f6300 · inbound

Eliciting Fine-Tuned Transformer Capabilities via Inference-Time Techniques cites this paper.

Eliciting Fine-Tuned Transformer Capabilities via Inference-Time Techniques Are Transformers universal approximators of sequence-to-sequence functions?

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T05:38:35.295805Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:38:35.295805Z digest=sha256:e9e1e7589685a4b797a43ec2557b61a64b18b552bf25ca04c36356d77d35e7b5

Observation e24742c2-24f2-4c94-9fb9-1d5775b6f79a · inbound

Time Resolution Independent Operator Learning cites this paper.

Time Resolution Independent Operator Learning Are Transformers universal approximators of sequence-to-sequence functions?

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T20:34:09.205665Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:34:09.205665Z digest=sha256:077e0483e9440d85e8cea5664519c5cb9f26db59ed78ea03366bf0b6e6e4c5d6

Observation 36941140-0c54-4d2b-a716-20ac10fdcd4c · inbound

On the Mathematical Impossibility of Safe Universal Approximators cites this paper.

On the Mathematical Impossibility of Safe Universal Approximators Are Transformers universal approximators of sequence-to-sequence functions?

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T20:41:55.587713Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:41:55.587713Z digest=sha256:1a2b35c706c13fdddab9e40f9d02ef51655d58031513cd4527eb81db14b3efb9

Observation 53fa0876-b32a-4d23-8dda-130b4925efb4 · inbound

Decoding Consumer Preferences Using Attention-Based Language Models cites this paper.

Decoding Consumer Preferences Using Attention-Based Language Models Are Transformers universal approximators of sequence-to-sequence functions?

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-06T14:51:53.078612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:51:53.078612Z digest=sha256:5b2d3cb3ef7b9d07e546c63dbf582677db75abc4d63d1329bbe8dfcd035a0ed5

Observation 50e79fa4-7a8d-4670-8b8a-57c90e4b4820 · inbound

Context-aware Rotary Position Embedding cites this paper.

Context-aware Rotary Position Embedding Are Transformers universal approximators of sequence-to-sequence functions?

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T11:09:08.432049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T11:09:08.432049Z digest=sha256:c7040c4f750281e8e049040ae2c17501f786e660dabed6fd0a1a8f209d4a9d7c

Observation 5f78fd59-44ff-486a-a621-39ee8527af08 · inbound

Beyond Memorization: Extending Reasoning Depth with Recurrence, Memory and Test-Time Compute Scaling cites this paper.

Beyond Memorization: Extending Reasoning Depth with Recurrence, Memory and Test-Time Compute Scaling Are Transformers universal approximators of sequence-to-sequence functions?

Reference 76

Resolution
verified exact
arxiv_id, observed 2026-05-18T20:51:50.868831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-18T20:49:26.966293Z digest=sha256:be0cb5a5343caa2436431b26cc4a8e1fcc738ed4822cda0e8faf806662ce50bc

Observation fcd3909e-358c-4cc6-a20f-584c92128715 · inbound

Unraveling Syntax: Language Modeling and the Substructure of Grammars cites this paper.

Unraveling Syntax: Language Modeling and the Substructure of Grammars Are Transformers universal approximators of sequence-to-sequence functions?

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-04T12:47:21.283300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:47:21.283300Z digest=sha256:f924a90cf51b75adfc21fca0f743d9d7fc1b52aa99dcf138a12862298d9a2c41

Observation 0f10c16c-47ba-4cd8-a5e9-c11b220203b2 · inbound

Key and Value Weights Are Probably All You Need: On the Necessity of the Query, Key, Value weight Triplet in Self-Attention Transformers cites this paper.

Key and Value Weights Are Probably All You Need: On the Necessity of the Query, Key, Value weight Triplet in Self-Attention Transformers Are Transformers universal approximators of sequence-to-sequence functions?

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-18T03:40:50.496470Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-18T03:38:36.932424Z digest=sha256:8730893ba98da9ac09292ea2d8c60ca5414ad2dce743532b496570cad2737675

Observation d6286ac3-109e-4068-89e0-8b0de8876a56 · inbound

Universal Approximation Theorems for Dynamical Systems with Infinite-Time Horizon Guarantees cites this paper.

Universal Approximation Theorems for Dynamical Systems with Infinite-Time Horizon Guarantees Are Transformers universal approximators of sequence-to-sequence functions?

Reference 174

Resolution
unresolved
no resolver link, observed 2026-08-03T03:21:49.824957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:21:49.824957Z digest=sha256:78ef83ffb1af6e5f6f16af8f52b6efbb10d84e83767d6ec9e9fc5bc9506ed4e9

Observation f01a6dbb-b2cd-48d4-9662-d8b2c6245f2a · inbound

Gating Enables Curvature: A Geometric Expressivity Gap in Attention cites this paper.

Gating Enables Curvature: A Geometric Expressivity Gap in Attention Are Transformers universal approximators of sequence-to-sequence functions?

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-10T11:20:11.376181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T11:10:47.526833Z digest=sha256:d885c70d4efa4ce6ac6c32b5ef02fb780b1d8d7ee403e376528fd5eee3087c74

Observation 2ecaae8f-ef3e-4e72-b0e9-bd0d77e79bfc · inbound

Continuous transformations of probability measures and their transport representations cites this paper.

Continuous transformations of probability measures and their transport representations Are Transformers universal approximators of sequence-to-sequence functions?

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-10T06:56:47.232493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-10T06:56:24.472027Z digest=sha256:0933d64e21e86aa07362a5ff38ebd9eb0f8089e5618d30f63676b5a56fd1045c

Observation 79ee8e96-593f-4e46-a052-e5a04f30d8ae · inbound

Progressive Approximation in Deep Residual Networks: Theory and Validation cites this paper.

Progressive Approximation in Deep Residual Networks: Theory and Validation Are Transformers universal approximators of sequence-to-sequence functions?

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T21:41:16.148117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-08T04:35:14.636981Z digest=sha256:346e0f1abebb7511d5aabbd614079e0665cf6de0be6ceda47f6221f9bc97880b

Observation ec91d9a6-9208-4372-b079-1b6ddd27492c · inbound

How Does Attention Help? Insights from Random Matrices on Signal Recovery from Sequence Models cites this paper.

How Does Attention Help? Insights from Random Matrices on Signal Recovery from Sequence Models Are Transformers universal approximators of sequence-to-sequence functions?

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:05:58.700918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-11T00:52:08.399401Z digest=sha256:d81087c8a60c316755208385d4b8683d2d5e49226e475b4b43200cbbbd045fcc

Observation c2348a66-465a-4e22-8a00-5423acb01eb5 · inbound

A generative pre-trained transformer with Kerr-soliton attention cites this paper.

A generative pre-trained transformer with Kerr-soliton attention Are Transformers universal approximators of sequence-to-sequence functions?

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-06-30T15:04:46.486774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-30T14:57:08.127650Z digest=sha256:9d3ecc63a38c948d96789419fab9c0f2bd2c742f6eee46cf1e81d4271044aca9

Observation 31e7d9a7-ca2e-478e-a3d4-09830604fd58 · inbound

The Variance Brain Foundation Models Forgot: Third-Order Statistics Predict Cognition Where Billion-Parameter Models Fail cites this paper.

The Variance Brain Foundation Models Forgot: Third-Order Statistics Predict Cognition Where Billion-Parameter Models Fail Are Transformers universal approximators of sequence-to-sequence functions?

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-06-28T19:12:34.913911Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-28T19:07:34.240009Z digest=sha256:71819913df8070ef5ca3842c269147b8f1f8f1a30b746d43bcec0b0eb90361b8

Observation d4e3d0f4-2e32-4f57-a427-31644db06327 · inbound

Imbuing Large Language Models with Bidirectional Logic for Robust Chain Repair cites this paper.

Imbuing Large Language Models with Bidirectional Logic for Robust Chain Repair Are Transformers universal approximators of sequence-to-sequence functions?

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-07-02T08:46:48.999772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T05:46:17.481698Z digest=sha256:0d01bce610e09aec17a31bd31762f174a5d0fd3b4d5293e291418744fd3701bd

Observation f16d0fac-d2c0-41da-9f12-fc68533b825a · inbound

Convergence of Gradient Descent for General Neural Network Architectures Beyond the NTK Regime cites this paper.

Convergence of Gradient Descent for General Neural Network Architectures Beyond the NTK Regime Are Transformers universal approximators of sequence-to-sequence functions?

Reference 110

Resolution
verified exact
arxiv_id, observed 2026-07-04T10:29:44.311684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-26T08:53:46.285233Z digest=sha256:1d4190c16eb0b6600a5d843ff7bf453ea54bb000df6ae006c669eebf2a5eba35

Observation 21bb62ce-7cc0-4e08-8803-e89fb4644a7a · inbound

Pre-Strings Lectures on Artificial Intelligence cites this paper.

Pre-Strings Lectures on Artificial Intelligence Are Transformers universal approximators of sequence-to-sequence functions?

Reference 29

Resolution
unresolved
no resolver link, observed 2026-07-12T06:14:03.658427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T06:14:03.658427Z digest=sha256:559faf67638b514b69d091f681b622bbe1b64fdd7e44d326945ea9529cbea255

Observation f95da0d6-cff4-4aa6-8be6-00989bc4ecdb · inbound

On Transformer Dynamics cites this paper.

On Transformer Dynamics Are Transformers universal approximators of sequence-to-sequence functions?

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-02T05:49:28.691589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:49:28.691589Z digest=sha256:1dd147af462d0bf1a0f5ab6c291afea8ed125b05bfe807092586e4510a3d7e9f

Observation 847e8761-c387-4425-be05-f1e3f3272ba1 · inbound

Training with (Swap) Regret Loss in a Single-Layer Self-Attention Model: A Case Study on the Probability Simplex cites this paper.

Training with (Swap) Regret Loss in a Single-Layer Self-Attention Model: A Case Study on the Probability Simplex Are Transformers universal approximators of sequence-to-sequence functions?

Reference 213

Resolution
unresolved
no resolver link, observed 2026-07-31T23:52:08.233130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T23:52:08.233130Z digest=sha256:4f76ec36538c6e8aa3fb2c7727c30a6d1d55ae0ec5ee04f603ca4ef0edf74eeb

Observation 4257aa79-2604-4872-89a4-6e4b6ef5cb9a · inbound

Identifiability-Aware Source Apportionment in City-Scale Advection-Diffusion Systems cites this paper.

Identifiability-Aware Source Apportionment in City-Scale Advection-Diffusion Systems Are Transformers universal approximators of sequence-to-sequence functions?

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-04T01:37:40.900908Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T01:37:40.900908Z digest=sha256:d7d6e562b5a2faccbed752037742dede3aa772dc190c8fd876b955315da97ca5