Pith. sign in

Paper Citation Record · LEDGER

Theoretical limitations of multi-layer Transformer

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 17 inbound Pith citation observations for arXiv:2412.02975.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.02975 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 17 of 17 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 17 of 17 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T00:18:21.263783Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T22:06:16.211261Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation b0de0330-ed76-4636-9e78-5e132706a0e8 · inbound

Pause Tokens Strictly Increase the Expressivity of Constant-Depth Transformers cites this paper.

Pause Tokens Strictly Increase the Expressivity of Constant-Depth Transformers Theoretical limitations of multi-layer Transformer

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T13:52:57.267988Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:52:57.267988Z digest=sha256:001a2f327cbc9d1e2ab6246da89ffa76b0642d6618131ebf6247f3ab10a01c44

Observation 23b29014-a63e-4589-8fc7-e3563d33c822 · inbound

Learning Compositional Functions with Transformers from Easy-to-Hard Data cites this paper.

Learning Compositional Functions with Transformers from Easy-to-Hard Data Theoretical limitations of multi-layer Transformer

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T12:46:33.978754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:46:33.978754Z digest=sha256:d143305f0b0f44f392ba1e2eeec0465b06219e918463f886c5328141dcd5b6fe

Observation 0892a673-8ad9-4bb9-8ba8-6a3aa63cfb42 · inbound

Transformers Meet In-Context Learning: A Universal Approximation Theory cites this paper.

Transformers Meet In-Context Learning: A Universal Approximation Theory Theoretical limitations of multi-layer Transformer

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T10:33:35.409540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:33:35.409540Z digest=sha256:fa9cf19ac8e7976f0cc4e377c2037815f3cdb74e6c2100fa0e576e31b374a0cc

Observation c7ad6197-2c7d-4432-9d5b-2249cf7b5a9e · inbound

Understanding Input Selectivity in Mamba: Impact on Approximation Power, Memorization, and Associative Recall Capacity cites this paper.

Understanding Input Selectivity in Mamba: Impact on Approximation Power, Memorization, and Associative Recall Capacity Theoretical limitations of multi-layer Transformer

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T01:18:51.437766Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T01:18:51.437766Z digest=sha256:8bb480a4147c2dbeb13478b42f35bc9302d1f98f12bfbcc9dedd9fc8b8784044

Observation 38a38488-424e-417c-9efa-9dbe9551f1bc · inbound

The Serial Scaling Hypothesis cites this paper.

The Serial Scaling Hypothesis Theoretical limitations of multi-layer Transformer

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-19T04:12:02.486141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-19T04:08:11.344622Z digest=sha256:615b4d3ac102f4089ffb86bf538f8684b53cb3cd2c8121666c3df4392ae605b3

Observation fea037c4-e0f2-44f6-83db-7d76371b6564 · inbound

Falcon-H1: A Family of Hybrid-Head Language Models Redefining Efficiency and Performance cites this paper.

Falcon-H1: A Family of Hybrid-Head Language Models Redefining Efficiency and Performance Theoretical limitations of multi-layer Transformer

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T11:44:04.879218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T11:44:04.879218Z digest=sha256:c4738da3d38d919c4da76fa6f6f0d5556dde543b601fcfc9d9e11df73dc66e72

Observation 87de76e6-4d0e-4296-a207-0af19beb2fa0 · inbound

Deep sequence models tend to memorize geometrically; it is unclear why cites this paper.

Deep sequence models tend to memorize geometrically; it is unclear why Theoretical limitations of multi-layer Transformer

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-21T20:40:36.194169Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-21T20:38:18.005002Z digest=sha256:ca419e215e6c152c7c1b89ed31fcc148588f153477dc15b804a0cdbc5c12d92c

Observation 06813133-1349-45ba-96d5-f3f7f8f2f588 · inbound

Transformers with RL or SFT Provably Learn Sparse Boolean Functions, But Differently cites this paper.

Transformers with RL or SFT Provably Learn Sparse Boolean Functions, But Differently Theoretical limitations of multi-layer Transformer

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-03T20:57:09.914249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T20:57:09.914249Z digest=sha256:b2191408073ad67bd945c24ab541d70133d9afab77c00fdcb4317c083ec7c189

Observation ef74e4a7-f24d-426e-a99e-21cdeb40f1fb · inbound

When Do Hallucinations Arise? A Graph Perspective on the Evolution of Path Reuse and Path Compression cites this paper.

When Do Hallucinations Arise? A Graph Perspective on the Evolution of Path Reuse and Path Compression Theoretical limitations of multi-layer Transformer

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-13T13:07:00.928770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T13:07:00.928770Z digest=sha256:b29b68f84fe6215386a16eb436c70179329b8f171556f6aae778fb6757441af7

Observation 58cf521f-8294-41cc-9335-9906260b1043 · inbound

The Power of Power Law: Asymmetry Enables Compositional Reasoning cites this paper.

The Power of Power Law: Asymmetry Enables Compositional Reasoning Theoretical limitations of multi-layer Transformer

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:31:08.240882Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-08T11:49:49.787123Z digest=sha256:8e77c86ce3e0dbdf7fc7c22a5ac451d6dc51369ebb48a017fd5f540ad754501b

Observation 346494fb-6aff-40b1-af21-3c6719f9c40d · inbound

The Power of Power Law: Asymmetry Enables Compositional Reasoning cites this paper.

The Power of Power Law: Asymmetry Enables Compositional Reasoning Theoretical limitations of multi-layer Transformer

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-12T18:26:05.728364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-12T18:26:05.728364Z digest=sha256:55e84475b3e0cb107c61ddb6e491b0b271fed0488b865184cb5e3cc960747cd7

Observation 0f6a387a-5972-444b-ad9e-d0e0bd7ea9f1 · inbound

Continuous Latent Contexts Enable Efficient Online Learning in Transformers cites this paper.

Continuous Latent Contexts Enable Efficient Online Learning in Transformers Theoretical limitations of multi-layer Transformer

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T05:41:23.574147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T05:06:54.598357Z digest=sha256:d042521b1b53423926376f3f334bb33f643c63266dcad9fb82d03b8d5941d337

Observation cb428005-a4ad-4090-8f1a-d89ede65f0a7 · inbound

Agentic Transformers Provably Learn to Search via Reinforcement Learning cites this paper.

Agentic Transformers Provably Learn to Search via Reinforcement Learning Theoretical limitations of multi-layer Transformer

Reference 24

Resolution
metadata mismatch
arxiv_id, observed 2026-06-28T23:42:49.947749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-28T23:26:28.158991Z digest=sha256:1e34c4d76ff75424675fe91dc7ef4ff452ab60b3d0a3953feb832f24d48c90d9

Observation 4b62c3fd-4b9e-4246-80fe-07705c18a1f7 · inbound

Rethinking the Role of Positional Encoding: Sliding-Window Transformers without PE Remain Turing Complete cites this paper.

Rethinking the Role of Positional Encoding: Sliding-Window Transformers without PE Remain Turing Complete Theoretical limitations of multi-layer Transformer

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T22:06:16.212640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-28T15:48:48.046003Z digest=sha256:a681b49fe199f40c53e99e16e31d2a0c2db2fca4fe7026e44d1894a158a222db

Observation 8bfa3f70-6a48-42be-9ed7-5e34a0907aa8 · inbound

Frontier Language Models Struggle to Copy: Text Can Be Better Viewed in 2D cites this paper.

Frontier Language Models Struggle to Copy: Text Can Be Better Viewed in 2D Theoretical limitations of multi-layer Transformer

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-01T21:31:00.903239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T21:31:00.903239Z digest=sha256:05c72b2c7fadf6c6dc75ea3b1448b72749442af66aed42475760af5f1fbb7823

Observation d52774c1-9cf4-4ffe-bffc-6140e67f8c26 · inbound

Hierarchical Domain Generalization cites this paper.

Hierarchical Domain Generalization Theoretical limitations of multi-layer Transformer

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-01T20:54:05.553053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T20:54:05.553053Z digest=sha256:893238db0b896f2d94f4f762a2d30067fe709abdec29c3abe014a00efd6e8475

Observation 2b3d7df0-27b4-4bd4-bce4-d824ce19a4e0 · inbound

Attention-based representations for multi-task computation cites this paper.

Attention-based representations for multi-task computation Theoretical limitations of multi-layer Transformer

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-08T00:18:21.263783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T00:18:21.263783Z digest=sha256:96cc7369473397532b2c9038383b858ec76f54b0b96f9a7b8f8c9de70b46b4d0