Pith. sign in

Paper Citation Record · LEDGER

Repeat After Me: Transformers are Better than State Space Models at Copying

As of 6 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 30 inbound Pith citation observations for arXiv:2402.01032.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2402.01032 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 30 of 30 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 30 of 30 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T16:23:14.385860Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T07:49:39.775848Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 4785e0e2-ffce-4e15-87fe-1fa1c43050d1 · inbound

Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models cites this paper.

Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models Repeat After Me: Transformers are Better than State Space Models at Copying

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:58:17.519265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T06:58:17.370396Z digest=sha256:a68ae2cb5e4309fdc5fd1b969bc130962abf807afbfa7c5d383c494f286b2c0d

Observation f6b2ad84-3d20-4dd8-bc39-137bce7e2a5b · inbound

An Empirical Study of Mamba-based Language Models cites this paper.

An Empirical Study of Mamba-based Language Models Repeat After Me: Transformers are Better than State Space Models at Copying

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-18T10:31:03.861393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T10:31:03.777169Z digest=sha256:617e2d9605bcf5ab75af7acdd94dfab3c7145fc9d396e3510d7386b6fc26bda4

Observation 5e71ee86-8927-4fc7-b8d0-87a84a262889 · inbound

TTT3R: 3D Reconstruction as Test-Time Training cites this paper.

TTT3R: 3D Reconstruction as Test-Time Training Repeat After Me: Transformers are Better than State Space Models at Copying

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-17T06:41:16.436769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T06:41:16.306593Z digest=sha256:e13a8233c16439c046094786ebc74880e8352ca94a64033abd98acc6d58f4120

Observation 4ec361f8-de23-415e-8b48-2abadecf5c13 · inbound

To model human linguistic prediction, make LLMs less superhuman cites this paper.

To model human linguistic prediction, make LLMs less superhuman Repeat After Me: Transformers are Better than State Space Models at Copying

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-04T12:55:17.799547Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:55:17.799547Z digest=sha256:dd05eb88fe9f7c5964849a7b26a67c1db6471fa4925537ffe5b3c74fee0aedd9

Observation f3f7ac2c-eccf-44b4-984b-62f8888dbd4a · inbound

Kimi Linear: An Expressive, Efficient Attention Architecture cites this paper.

Kimi Linear: An Expressive, Efficient Attention Architecture Repeat After Me: Transformers are Better than State Space Models at Copying

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-13T23:49:11.163140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T23:49:10.555255Z digest=sha256:58186892a57fb2c6c9bb390c770a750792674322cb399c81e97fec0d76368880

Observation 8103b078-7df7-4329-8404-fb9a9755987a · inbound

Controllably Efficient Language Models cites this paper.

Controllably Efficient Language Models Repeat After Me: Transformers are Better than State Space Models at Copying

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-03T23:33:56.548151Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:33:56.548151Z digest=sha256:7cc31580050948910d93d7a31ace3f11576235c3bd3b72f815d999df5d58cab2

Observation 7cd71372-04b3-411a-86a3-38862c2577d6 · inbound

Gated KalmaNet: A Fading Memory Layer Through Test-Time Ridge Regression cites this paper.

Gated KalmaNet: A Fading Memory Layer Through Test-Time Ridge Regression Repeat After Me: Transformers are Better than State Space Models at Copying

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-21T18:00:27.090936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T17:59:23.826110Z digest=sha256:7198a29f9e786127d53af61433d61950b59b9a25b9e7ec81871b619ab5d62eaf

Observation 31bba847-f7ed-4acd-9271-7a3a672b259d · inbound

Physics of Language Models: Part 4.1, Architecture Design and the Magic of Canon Layers cites this paper.

Physics of Language Models: Part 4.1, Architecture Design and the Magic of Canon Layers Repeat After Me: Transformers are Better than State Space Models at Copying

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-03T15:22:50.161253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T15:22:50.161253Z digest=sha256:5c00c7c6a2120c2c078ca50ff88d8f1a16baf28b7b9c07731bc3216bdfd945d7

Observation 31684462-8ffd-4905-adcd-93b9b894d592 · inbound

The Bayesian Geometry of Transformer Attention cites this paper.

The Bayesian Geometry of Transformer Attention Repeat After Me: Transformers are Better than State Space Models at Copying

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T16:20:20.659781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T16:15:57.408685Z digest=sha256:5857593505ba9f19c96819e9474c0966e5055e9cad35bbd8857e059bbc53724f

Observation 52bab556-5e0a-451b-974f-9fde3c04def9 · inbound

Towards Understanding What State Space Models Learn About Code cites this paper.

Towards Understanding What State Space Models Learn About Code Repeat After Me: Transformers are Better than State Space Models at Copying

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-03T03:51:43.479561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T03:51:43.479561Z digest=sha256:87f18dd8c49d7d1b00915d09298280ff419e26ae522bb88f8f98ffc7e639afaa

Observation effe2026-80e9-4099-9641-bacc9a687709 · inbound

The UNDO Flip-Flop: A Controlled Probe for Reversible Semantic State Management in State Space Model cites this paper.

The UNDO Flip-Flop: A Controlled Probe for Reversible Semantic State Management in State Space Model Repeat After Me: Transformers are Better than State Space Models at Copying

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:00:52.709421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T18:45:35.405345Z digest=sha256:5809d1cca77dff0b817c525e1bbd8b91475c87424fa03bc75adc4243d5a9c91d

Observation 5215fe50-c4af-46e9-bf7e-663a5d4856ef · inbound

The Recurrent Transformer: Greater Effective Depth and Efficient Decoding cites this paper.

The Recurrent Transformer: Greater Effective Depth and Efficient Decoding Repeat After Me: Transformers are Better than State Space Models at Copying

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-05-11T14:21:04.627239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-09T22:03:35.260373Z digest=sha256:5e8ccd14ab98071520430d3a6f964b50eb944ac87fce377861c178ee672b409b

Observation a966edb1-f8a6-4fb8-b3dd-aac28ccb0c79 · inbound

Toeplitz MLP Mixers are Low Complexity, Information-Rich Sequence Models cites this paper.

Toeplitz MLP Mixers are Low Complexity, Information-Rich Sequence Models Repeat After Me: Transformers are Better than State Space Models at Copying

Reference 53

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T04:45:56.203904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-11T01:07:33.756985Z digest=sha256:c8b726d03b6467f9e903668ee258224af1c0a8a55b8ab96a2ecee47d043708a8

Observation ad50ff5c-3d7c-4fe5-bd79-3d8a57f0aca5 · inbound

Echo: KV-Cache-Free Associative Recall with Spectral Koopman Operators cites this paper.

Echo: KV-Cache-Free Associative Recall with Spectral Koopman Operators Repeat After Me: Transformers are Better than State Space Models at Copying

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:25:56.325867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-11T01:26:48.026538Z digest=sha256:26d95a1741df5b99c2f0dfb97353fdc1ff7edf9824c4917d3733b8ada03cf052

Observation bc1e3a80-bddc-4fb0-b861-9bad26bb08db · inbound

Positional LSH: Binary Block Matrix Approximation for Attention with Linear Biases cites this paper.

Positional LSH: Binary Block Matrix Approximation for Attention with Linear Biases Repeat After Me: Transformers are Better than State Space Models at Copying

Reference 38

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T06:11:27.007244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-12T04:29:37.989500Z digest=sha256:c51e5444feb03b356a729c23071dd97209d3a3e979ac08cf441064909fc1d122

Observation e827a593-b678-41e4-b8a5-15afeea5ad62 · inbound

OSDN: Improving Delta Rule with Provable Online Preconditioning in Linear Attention cites this paper.

OSDN: Improving Delta Rule with Provable Online Preconditioning in Linear Attention Repeat After Me: Transformers are Better than State Space Models at Copying

Reference 25

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T19:09:23.228532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-14T19:08:18.768344Z digest=sha256:07a0e83092da03ed482016f19e1f47c5c02b2f74ac10041f65cf1711fef8e383

Observation 7606210a-7535-4e13-a8d3-435571d897c7 · inbound

Do Language Models Need Sleep? Offline Recurrence for Improved Online Inference cites this paper.

Do Language Models Need Sleep? Offline Recurrence for Improved Online Inference Repeat After Me: Transformers are Better than State Space Models at Copying

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-06-29T21:43:59.528847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-29T21:37:51.638904Z digest=sha256:14a985c0c44192045f4a3b9e5b2925a812d8907b2473ed9286c2e60e31376030

Observation f65382ec-1ece-4ce6-9e6f-5f4802358b7d · inbound

Zamba2-VL Technical Report cites this paper.

Zamba2-VL Technical Report Repeat After Me: Transformers are Better than State Space Models at Copying

Reference 65

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T19:26:00.673759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T22:34:20.970856Z digest=sha256:6f22181e0f53c1052543299b7e0e9c267a4991961fe36a3fbc27e0bd14afce57

Observation 3c7b0e8c-ea93-4311-8fa3-f457973d9ff0 · inbound

Attention Amnesia in Hybrid LLMs: When CoT Fine-Tuning Breaks Long-Range Recall, and How to Fix It cites this paper.

Attention Amnesia in Hybrid LLMs: When CoT Fine-Tuning Breaks Long-Range Recall, and How to Fix It Repeat After Me: Transformers are Better than State Space Models at Copying

Reference 112

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T05:37:40.365982Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-06-27T13:08:57.218711Z digest=sha256:e1eaca7a90906054eb749b7a60e3abfa558d5ec30d93982a62eb49c78558e99e

Observation dff055e0-ab10-4f46-938d-37dd97af1876 · inbound

A Verifiable Search Is Not a Learnable Chain-of-Thought cites this paper.

A Verifiable Search Is Not a Learnable Chain-of-Thought Repeat After Me: Transformers are Better than State Space Models at Copying

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T07:49:39.777090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-26T12:35:20.698118Z digest=sha256:8251eb53f8ddd8e9c24908d587784fa21ba239e5a1c497a1217042093b8610a5

Observation ed936055-c6b7-4404-84e4-ce80ef412ce5 · inbound

A Hippocampus for Linear Attention: An Exact Memory for What the Recurrent State Forgets cites this paper.

A Hippocampus for Linear Attention: An Exact Memory for What the Recurrent State Forgets Repeat After Me: Transformers are Better than State Space Models at Copying

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-07-03T13:48:20.316429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-03T13:46:18.862925Z digest=sha256:cb08b890ed9c2e2b9a4e9922e6b3229188d4f92e5ec6aeecd7b72f3fb7e4cb11

Observation c975545d-840b-4086-ad3b-027c90be10fd · inbound

Remembering Distinct Items, Not Tokens: A Learnable Dirichlet-Process Cache Between State-Space Models and Attention cites this paper.

Remembering Distinct Items, Not Tokens: A Learnable Dirichlet-Process Cache Between State-Space Models and Attention Repeat After Me: Transformers are Better than State Space Models at Copying

Reference 13

Resolution
unresolved
no resolver link, observed 2026-07-14T14:50:03.831572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T14:50:03.831572Z digest=sha256:673f0a636d0840c68690490db0fe4595b31b7b47ccc36813c98282959aa51a57

Observation 0aea3c71-0acc-4a0a-831f-6f5ac09bd726 · inbound

DSSMs: State Space Models with Explicit Memory via Delay Differential Equations cites this paper.

DSSMs: State Space Models with Explicit Memory via Delay Differential Equations Repeat After Me: Transformers are Better than State Space Models at Copying

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-14T13:15:38.524121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T13:15:38.524121Z digest=sha256:802a4ccf8c790e255648a1a7144b2fcc3b452412bf1f4989f1cf3615e93a2dbb

Observation 1ce7cf71-0aa1-429a-acce-aae0ba99215a · inbound

Context by Distinct Information: An Auditable Dirichlet-Process Working Memory for Long, Redundant Context Streams cites this paper.

Context by Distinct Information: An Auditable Dirichlet-Process Working Memory for Long, Redundant Context Streams Repeat After Me: Transformers are Better than State Space Models at Copying

Reference 5

Resolution
unresolved
no resolver link, observed 2026-07-14T11:45:09.424370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T11:45:09.424370Z digest=sha256:b9cd33a641443b2e096439c684e52a371d6e1f08e3144c2f93f386c31445f6b6

Observation e90e7497-c8d5-4d09-bdfa-ee34ab6c017d · inbound

The Capability Convergence Hypothesis: Capability from Access Structure, Not Scale cites this paper.

The Capability Convergence Hypothesis: Capability from Access Structure, Not Scale Repeat After Me: Transformers are Better than State Space Models at Copying

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-02T06:39:22.140957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:39:22.140957Z digest=sha256:ed752d8f0a8fbc9ab0d846e4c4beff474ed150b7313ecf340394d5952c031175

Observation a3aa00d8-fc6d-4ebf-b856-35e685c3f74f · inbound

The Capability Convergence Hypothesis: Capability from Access Structure, Not Scale cites this paper.

The Capability Convergence Hypothesis: Capability from Access Structure, Not Scale Repeat After Me: Transformers are Better than State Space Models at Copying

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-03T02:03:24.068474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:03:24.068474Z digest=sha256:567dc1255c521ac044e5c15fe84a83c1557fa44917b6b98243ba6c3df82d6a64

Observation 86d98c18-52eb-4932-92da-eaadab7a7c71 · inbound

Kernelized Linear Attention: Breaking the Capacity Wall with Symmetric Cones cites this paper.

Kernelized Linear Attention: Breaking the Capacity Wall with Symmetric Cones Repeat After Me: Transformers are Better than State Space Models at Copying

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-01T18:06:48.827702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T18:06:48.827702Z digest=sha256:d1b9626eb3d1f0db3ca0d5117aaa81d725e54f66fe91ffab8d813fda0b2ba65a

Observation d26c5665-9749-4773-b242-650aa8297dc6 · inbound

Can Transformers Really Do It All? On the Compatibility of Inductive Biases Across Tasks cites this paper.

Can Transformers Really Do It All? On the Compatibility of Inductive Biases Across Tasks Repeat After Me: Transformers are Better than State Space Models at Copying

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-01T17:31:47.771481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:31:47.771481Z digest=sha256:f62aa1dc84f5d6b841cf60d3d4f26946c2931442fa7d5e59e9b0f9578d7d7f94

Observation 653d6092-2031-4eaf-9b84-61e86a5c15f9 · inbound

Raven: High-Recall Sequence Modeling with Sparse Memory Routing cites this paper.

Raven: High-Recall Sequence Modeling with Sparse Memory Routing Repeat After Me: Transformers are Better than State Space Models at Copying

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-01T02:44:01.258380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:44:01.258380Z digest=sha256:57c8ce5f610e2d69ed60cfb07e2ceecf1a23339de365e3b72697672dce49b484

Observation 74de6bad-a1d0-4c22-af69-25568eff1569 · inbound

DART: Decoded Attention over Recurrent States for Efficient Long-Context Sequence Modeling cites this paper.

DART: Decoded Attention over Recurrent States for Efficient Long-Context Sequence Modeling Repeat After Me: Transformers are Better than State Space Models at Copying

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T16:23:14.385860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:23:14.385860Z digest=sha256:75cc132423e1ba59181ba3d4c70076a22bc5d4864b34533281ea4138ac797ec2