Pith. sign in

Paper Citation Record · LEDGER

Theoretical limitations of multi-layer Transformer

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 17 inbound Pith citation observations for arXiv:2412.02975.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.02975 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 17 of 17 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 17 of 17 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T00:18:21.263783Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T22:06:16.211261Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation b0de0330-ed76-4636-9e78-5e132706a0e8 · inbound

Pause Tokens Strictly Increase the Expressivity of Constant-Depth Transformers cites this paper.

Pause Tokens Strictly Increase the Expressivity of Constant-Depth Transformers Theoretical limitations of multi-layer Transformer

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T13:52:57.267988Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:52:57.267988Z digest=sha256:f16b116a8f370469bc913e9fd8abb5a573e994dc35edf7f24be60f6b6fbaa342

Observation 23b29014-a63e-4589-8fc7-e3563d33c822 · inbound

Learning Compositional Functions with Transformers from Easy-to-Hard Data cites this paper.

Learning Compositional Functions with Transformers from Easy-to-Hard Data Theoretical limitations of multi-layer Transformer

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T12:46:33.978754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:46:33.978754Z digest=sha256:d143305f0b0f44f392ba1e2eeec0465b06219e918463f886c5328141dcd5b6fe

Observation 0892a673-8ad9-4bb9-8ba8-6a3aa63cfb42 · inbound

Transformers Meet In-Context Learning: A Universal Approximation Theory cites this paper.

Transformers Meet In-Context Learning: A Universal Approximation Theory Theoretical limitations of multi-layer Transformer

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T10:33:35.409540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:33:35.409540Z digest=sha256:854a8d3f2d7683de898df4fc68632281b6a3f7b6fa8ec1db56245698d697b193

Observation c7ad6197-2c7d-4432-9d5b-2249cf7b5a9e · inbound

Understanding Input Selectivity in Mamba: Impact on Approximation Power, Memorization, and Associative Recall Capacity cites this paper.

Understanding Input Selectivity in Mamba: Impact on Approximation Power, Memorization, and Associative Recall Capacity Theoretical limitations of multi-layer Transformer

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T01:18:51.437766Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T01:18:51.437766Z digest=sha256:8bb480a4147c2dbeb13478b42f35bc9302d1f98f12bfbcc9dedd9fc8b8784044

Observation 38a38488-424e-417c-9efa-9dbe9551f1bc · inbound

The Serial Scaling Hypothesis cites this paper.

The Serial Scaling Hypothesis Theoretical limitations of multi-layer Transformer

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-19T04:12:02.486141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-19T04:08:11.344622Z digest=sha256:90731461863f752a99adcee5f2df3f02ef904f7112b3ecbeb2341228b72133a5

Observation fea037c4-e0f2-44f6-83db-7d76371b6564 · inbound

Falcon-H1: A Family of Hybrid-Head Language Models Redefining Efficiency and Performance cites this paper.

Falcon-H1: A Family of Hybrid-Head Language Models Redefining Efficiency and Performance Theoretical limitations of multi-layer Transformer

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T11:44:04.879218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T11:44:04.879218Z digest=sha256:c4738da3d38d919c4da76fa6f6f0d5556dde543b601fcfc9d9e11df73dc66e72

Observation 87de76e6-4d0e-4296-a207-0af19beb2fa0 · inbound

Deep sequence models tend to memorize geometrically; it is unclear why cites this paper.

Deep sequence models tend to memorize geometrically; it is unclear why Theoretical limitations of multi-layer Transformer

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-21T20:40:36.194169Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-21T20:38:18.005002Z digest=sha256:79fc32af43a9d7f36b4b4590f16dd8f2f669d9233819a64c7c6f4cb336db9d49

Observation 06813133-1349-45ba-96d5-f3f7f8f2f588 · inbound

Transformers with RL or SFT Provably Learn Sparse Boolean Functions, But Differently cites this paper.

Transformers with RL or SFT Provably Learn Sparse Boolean Functions, But Differently Theoretical limitations of multi-layer Transformer

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-03T20:57:09.914249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T20:57:09.914249Z digest=sha256:11ae6f372a81eb97504d8b69f08f54e408ee0d5e742364d1aa2cebf4336ab54c

Observation ef74e4a7-f24d-426e-a99e-21cdeb40f1fb · inbound

When Do Hallucinations Arise? A Graph Perspective on the Evolution of Path Reuse and Path Compression cites this paper.

When Do Hallucinations Arise? A Graph Perspective on the Evolution of Path Reuse and Path Compression Theoretical limitations of multi-layer Transformer

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-13T13:07:00.928770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T13:07:00.928770Z digest=sha256:b29b68f84fe6215386a16eb436c70179329b8f171556f6aae778fb6757441af7

Observation 58cf521f-8294-41cc-9335-9906260b1043 · inbound

The Power of Power Law: Asymmetry Enables Compositional Reasoning cites this paper.

The Power of Power Law: Asymmetry Enables Compositional Reasoning Theoretical limitations of multi-layer Transformer

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:31:08.240882Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-08T11:49:49.787123Z digest=sha256:1348939d6c07cab6311b324dff1e551918f3b99fac88c34249c6a4012f4e33c5

Observation 346494fb-6aff-40b1-af21-3c6719f9c40d · inbound

The Power of Power Law: Asymmetry Enables Compositional Reasoning cites this paper.

The Power of Power Law: Asymmetry Enables Compositional Reasoning Theoretical limitations of multi-layer Transformer

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-12T18:26:05.728364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-12T18:26:05.728364Z digest=sha256:55e84475b3e0cb107c61ddb6e491b0b271fed0488b865184cb5e3cc960747cd7

Observation 0f6a387a-5972-444b-ad9e-d0e0bd7ea9f1 · inbound

Continuous Latent Contexts Enable Efficient Online Learning in Transformers cites this paper.

Continuous Latent Contexts Enable Efficient Online Learning in Transformers Theoretical limitations of multi-layer Transformer

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T05:41:23.574147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-12T05:06:54.598357Z digest=sha256:dead402227ec5074fbc6e3c9e0a0dd28d60c80aa4d3fa882aac675c5489fc74f

Observation cb428005-a4ad-4090-8f1a-d89ede65f0a7 · inbound

Agentic Transformers Provably Learn to Search via Reinforcement Learning cites this paper.

Agentic Transformers Provably Learn to Search via Reinforcement Learning Theoretical limitations of multi-layer Transformer

Reference 24

Resolution
metadata mismatch
arxiv_id, observed 2026-06-28T23:42:49.947749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-28T23:26:28.158991Z digest=sha256:309c8e75319c933a7f6d535107504e9c8672880d90c9cc44f858d270990e3616

Observation 4b62c3fd-4b9e-4246-80fe-07705c18a1f7 · inbound

Rethinking the Role of Positional Encoding: Sliding-Window Transformers without PE Remain Turing Complete cites this paper.

Rethinking the Role of Positional Encoding: Sliding-Window Transformers without PE Remain Turing Complete Theoretical limitations of multi-layer Transformer

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T22:06:16.212640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-28T15:48:48.046003Z digest=sha256:7103b2516c7dc1e843fd9d03a78a4aa3b16e0c9c6d100d8800e3419649db44bb

Observation 8bfa3f70-6a48-42be-9ed7-5e34a0907aa8 · inbound

Frontier Language Models Struggle to Copy: Text Can Be Better Viewed in 2D cites this paper.

Frontier Language Models Struggle to Copy: Text Can Be Better Viewed in 2D Theoretical limitations of multi-layer Transformer

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-01T21:31:00.903239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T21:31:00.903239Z digest=sha256:05c72b2c7fadf6c6dc75ea3b1448b72749442af66aed42475760af5f1fbb7823

Observation d52774c1-9cf4-4ffe-bffc-6140e67f8c26 · inbound

Hierarchical Domain Generalization cites this paper.

Hierarchical Domain Generalization Theoretical limitations of multi-layer Transformer

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-01T20:54:05.553053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T20:54:05.553053Z digest=sha256:893238db0b896f2d94f4f762a2d30067fe709abdec29c3abe014a00efd6e8475

Observation 2b3d7df0-27b4-4bd4-bce4-d824ce19a4e0 · inbound

Attention-based representations for multi-task computation cites this paper.

Attention-based representations for multi-task computation Theoretical limitations of multi-layer Transformer

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-08T00:18:21.263783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T00:18:21.263783Z digest=sha256:3b28e43e838b70afdf15ff345cd15a733d5cc5e17dba7ac0bb667b6b3136b578