Pith. sign in

Paper Citation Record · LEDGER

Latent Recurrent Transformer: Architecture Exploration, Training Strategies, and Scaling Behavior

As of 4 August 2026, this Paper Citation Record lists 27 of 27 outbound references and 0 inbound Pith citation observations for arXiv:2605.26797.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.26797 v1

Coverage vector

measured 27 of 27 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-29T19:36:13.559393Z

measured 27 of 27 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

27 of 27 outbound references displayed

  • verified exact18
  • verified fuzzy0
  • unresolved8
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 6bdb8679-bbd1-4df7-b8e6-e25b92d6532a · outbound

This paper cites Language models are few-shot learners.

Latent Recurrent Transformer: Architecture Exploration, Training Strategies, and Scaling Behavior Language models are few-shot learners

Reference 1

Resolution
unresolved
no resolver link, observed 2026-06-29T19:36:13.559393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T19:36:13.559393Z digest=sha256:306eb1625f1f69ebc6e78428669a0bd4e3328cfcf405b50070bef69b9ce14bdc

Observation 2efbb355-3685-403c-8cb5-003ce1c758cb · outbound

This paper cites Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models.

Latent Recurrent Transformer: Architecture Exploration, Training Strategies, and Scaling Behavior Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-06-29T19:43:54.937865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-29T19:36:13.559393Z digest=sha256:eedcc920bf6a7293e94084ae09cd400718327cf44303f9962fa5f4a27147dd93

Observation 99c4b0bc-06bf-42fb-8651-5f97cf7b9336 · outbound

This paper cites Addressing Some Limitations of Transformers with Feedback Memory.

Latent Recurrent Transformer: Architecture Exploration, Training Strategies, and Scaling Behavior Addressing Some Limitations of Transformers with Feedback Memory

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-06-29T19:43:54.933744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-29T19:36:13.559393Z digest=sha256:5e75b740abb057e7af55ace6960943517bfad2e51060b1cd11b07b810fbd8e16

Observation dab136e2-21ca-4c32-a2b2-f48ea4cd2812 · outbound

This paper cites Hungry Hungry Hippos: Towards Language Modeling with State Space Models.

Latent Recurrent Transformer: Architecture Exploration, Training Strategies, and Scaling Behavior Hungry Hungry Hippos: Towards Language Modeling with State Space Models

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-06-29T19:43:54.934978Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-29T19:36:13.559393Z digest=sha256:29c9b23169ae050fc8ff0356f3265da3260fb69086ede5d2c1a6e740e8b09d16

Observation 8a7b088d-d4cd-41b9-a7d4-4cfea3c4b9c8 · outbound

This paper cites Think before you speak: Training language models with pause tokens.

Latent Recurrent Transformer: Architecture Exploration, Training Strategies, and Scaling Behavior Think before you speak: Training language models with pause tokens

Reference 5

Resolution
unresolved
no resolver link, observed 2026-06-29T19:36:13.559393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T19:36:13.559393Z digest=sha256:9d59236da91b0af4228e841d4cb150d5d3435f11ab5f861ae6f6c0259d0b30c2

Observation f6801829-3569-41c0-9877-f7a91b56e039 · outbound

This paper cites Mamba: Linear-Time Sequence Modeling with Selective State Spaces.

Latent Recurrent Transformer: Architecture Exploration, Training Strategies, and Scaling Behavior Mamba: Linear-Time Sequence Modeling with Selective State Spaces

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-06-29T19:43:54.980929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-29T19:36:13.559393Z digest=sha256:da8f68523bf8fb5eeb34f1f1dd0b4cbd42a4f280131a434bc0d45e423d510c09

Observation ff661fe8-d849-4d21-ba1d-b9ee4cbe9c2f · outbound

This paper cites Training Large Language Models to Reason in a Continuous Latent Space.

Latent Recurrent Transformer: Architecture Exploration, Training Strategies, and Scaling Behavior Training Large Language Models to Reason in a Continuous Latent Space

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-06-29T19:43:54.969758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-29T19:36:13.559393Z digest=sha256:c8bbbecda53c3a9efade6792fd1d90e9dd17a9b071a47c198f9045b53585795f

Observation cc9c82eb-31f1-4a78-b661-ab6fa781d8fa · outbound

This paper cites Thinking Tokens for Language Modeling.

Latent Recurrent Transformer: Architecture Exploration, Training Strategies, and Scaling Behavior Thinking Tokens for Language Modeling

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T19:43:54.967076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-29T19:36:13.559393Z digest=sha256:bdbe39981ca8961b88caacdb6d5307b1ff1b757a1dc4f6b9a3be3b80401ff00b

Observation 8f438508-705f-47b5-9e66-0fe83f8cde05 · outbound

This paper cites TransformerFAM: Feedback attention is working memory.

Latent Recurrent Transformer: Architecture Exploration, Training Strategies, and Scaling Behavior TransformerFAM: Feedback attention is working memory

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-06-29T19:43:54.961412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-29T19:36:13.559393Z digest=sha256:90980844f2b118ebfc4a331e18055cde11fa3fdb7d3d6e3c7d28a8c9f66c3ce3

Observation 788d6530-a7dd-4091-af6d-179219879769 · outbound

This paper cites Muon: An optimizer for hidden layers in neural networks, 2024.URL https://kellerjordan.

Latent Recurrent Transformer: Architecture Exploration, Training Strategies, and Scaling Behavior Muon: An optimizer for hidden layers in neural networks, 2024.URL https://kellerjordan

Reference 10

Resolution
unresolved
no resolver link, observed 2026-06-29T19:36:13.559393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T19:36:13.559393Z digest=sha256:9fa024b02fb250f45861c21114f048cb44d8d2f1c80ea6f76828e11052fd208f

Observation 7a1b5163-bc86-4017-b0ac-ae3dd66d0c2e · outbound

This paper cites Jamba: A Hybrid Transformer-Mamba Language Model.

Latent Recurrent Transformer: Architecture Exploration, Training Strategies, and Scaling Behavior Jamba: A Hybrid Transformer-Mamba Language Model

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-06-29T19:43:54.983792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-29T19:36:13.559393Z digest=sha256:89279e421fb7a9b65b2eed409d1902ddcee4d68366c5d02dea00e2988abae477

Observation 1d46141e-98e7-48cb-a689-2edde5adce72 · outbound

This paper cites Fineweb-edu: the finest collection of educational content, 2024.URL https://huggingface.

Latent Recurrent Transformer: Architecture Exploration, Training Strategies, and Scaling Behavior Fineweb-edu: the finest collection of educational content, 2024.URL https://huggingface

Reference 12

Resolution
unresolved
no resolver link, observed 2026-06-29T19:36:13.559393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T19:36:13.559393Z digest=sha256:2dba3e5d32e1690c81653d99720c471771644b5b330d15ee37565b782d794e0d

Observation b602e78d-e8f5-434d-aa23-e7433373862a · outbound

This paper cites Landmark Attention: Random-Access Infinite Context Length for Transformers.

Latent Recurrent Transformer: Architecture Exploration, Training Strategies, and Scaling Behavior Landmark Attention: Random-Access Infinite Context Length for Transformers

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-06-29T19:43:54.977792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-29T19:36:13.559393Z digest=sha256:7f8f5aed05a04cd99481097ff83c9a4064423c70b71c557f71f2069de38cd98d

Observation 16341c4e-7eed-455c-a832-a1231670ebf2 · outbound

This paper cites Leave No Context Behind: Efficient Infinite Context Transformers with Infini-attention.

Latent Recurrent Transformer: Architecture Exploration, Training Strategies, and Scaling Behavior Leave No Context Behind: Efficient Infinite Context Transformers with Infini-attention

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-06-29T19:43:54.969393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-29T19:36:13.559393Z digest=sha256:cddc4af9ca5ecf602449ad83b333597adb7dede9057f02a764f025417d36a2e6

Observation 5d043d28-2feb-4069-a115-80f40c889473 · outbound

This paper cites RWKV: Reinventing RNNs for the Transformer Era.

Latent Recurrent Transformer: Architecture Exploration, Training Strategies, and Scaling Behavior RWKV: Reinventing RNNs for the Transformer Era

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-06-29T19:43:54.971790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-29T19:36:13.559393Z digest=sha256:4edb4cb262093907b9dfa0e54c4e66c50fb479b4563c263ff6bf196348e119cc

Observation efb1b40d-02f1-405a-85e2-3ba3052b2561 · outbound

This paper cites Let's Think Dot by Dot: Hidden Computation in Transformer Language Models.

Latent Recurrent Transformer: Architecture Exploration, Training Strategies, and Scaling Behavior Let's Think Dot by Dot: Hidden Computation in Transformer Language Models

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-06-29T19:43:54.963801Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-29T19:36:13.559393Z digest=sha256:df733053cfebd19d1a62cc33a8b2becaf002ba70472501621f753102366fa754

Observation eb9e9bd5-3a79-4c8d-9263-11c7c44f33b2 · outbound

This paper cites Compressive Transformers for Long-Range Sequence Modelling.

Latent Recurrent Transformer: Architecture Exploration, Training Strategies, and Scaling Behavior Compressive Transformers for Long-Range Sequence Modelling

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-06-29T19:43:54.972100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-29T19:36:13.559393Z digest=sha256:2996e2b49dd415b0d4c5bbfda4c1f821b30241e49d5e67881c3ce956e1b6d848

Observation e95820fa-5fb5-4599-9941-b201aaeb1536 · outbound

This paper cites Simplified State Space Layers for Sequence Modeling.

Latent Recurrent Transformer: Architecture Exploration, Training Strategies, and Scaling Behavior Simplified State Space Layers for Sequence Modeling

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-06-29T19:43:54.974514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-29T19:36:13.559393Z digest=sha256:7d16acd50cd23a21dcadc03c7f7430bff4d17fdb57cd49133f25038da6f1f586

Observation f1089c86-5ba0-421b-8a27-89a4c9df6810 · outbound

This paper cites Learning to (Learn at Test Time): RNNs with Expressive Hidden States.

Latent Recurrent Transformer: Architecture Exploration, Training Strategies, and Scaling Behavior Learning to (Learn at Test Time): RNNs with Expressive Hidden States

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-06-29T19:43:54.986638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-29T19:36:13.559393Z digest=sha256:82553e00841eae229f2facc78e72a3afc8e0e588b0c0e360b7bc6e0156a6b37a

Observation 0b0f7268-8c3f-4441-8b75-20e48d214f03 · outbound

This paper cites Retentive Network: A Successor to Transformer for Large Language Models.

Latent Recurrent Transformer: Architecture Exploration, Training Strategies, and Scaling Behavior Retentive Network: A Successor to Transformer for Large Language Models

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-06-29T19:43:54.948230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-29T19:36:13.559393Z digest=sha256:f1ecde6a8a53c4edf38b52b75b5aa6380ca67594bdb73902cf9aa0ff825566ef

Observation 4dd94e11-ba68-4b2f-a850-a88bbdcf7c16 · outbound

This paper cites Gemma 2: Improving Open Language Models at a Practical Size.

Latent Recurrent Transformer: Architecture Exploration, Training Strategies, and Scaling Behavior Gemma 2: Improving Open Language Models at a Practical Size

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-06-29T19:43:54.951166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-29T19:36:13.559393Z digest=sha256:6dec38b6ef3e5a9918f350b2fb8078e89e44d4f83fc3854a70d356c4d60a5fbe

Observation 3f0b871a-b078-46b8-92de-780ac6053677 · outbound

This paper cites Memorizing Transformers.

Latent Recurrent Transformer: Architecture Exploration, Training Strategies, and Scaling Behavior Memorizing Transformers

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-06-29T19:43:54.966815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-29T19:36:13.559393Z digest=sha256:af679ebe3511d87aa5f21d68b59b006047bbc470615c3cb9273a1de11753b84e

Observation f52560f1-2bb2-4c67-a37e-0bd395dd25f4 · outbound

This paper cites Quiet-STaR: Language Models Can Teach Themselves to Think Before Speaking.

Latent Recurrent Transformer: Architecture Exploration, Training Strategies, and Scaling Behavior Quiet-STaR: Language Models Can Teach Themselves to Think Before Speaking

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-06-29T19:43:54.977708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-29T19:36:13.559393Z digest=sha256:65ef00db70e966c84d064de71d7735fd0714a8fd7ce19c7c038d3147fc94e8b0

Observation 9b03380b-245f-4a4c-86cf-9f83cd2b7cf3 · outbound

This paper cites an unresolved cited work.

Latent Recurrent Transformer: Architecture Exploration, Training Strategies, and Scaling Behavior Unresolved cited work

Reference 24

Resolution
unresolved
no resolver link, observed 2026-06-29T19:36:13.559393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T19:36:13.559393Z digest=sha256:de337c3f908bb6a75f453dd3bb6a5034cfaa8d9ffb76ff3cffb1b6b6092b5f71

Observation 6060a32c-504e-448a-9b92-adc50a8054ea · outbound

This paper cites The final shard is held out as validation and the remaining shards are used for training.

Latent Recurrent Transformer: Architecture Exploration, Training Strategies, and Scaling Behavior The final shard is held out as validation and the remaining shards are used for training

Reference 25

Resolution
unresolved
no resolver link, observed 2026-06-29T19:36:13.559393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T19:36:13.559393Z digest=sha256:1a3b6bf2d6836fee1609dcf9800487eaf96f6e8f553f6c8504d1634c62b5a0da

Observation 2519ddef-6310-46c1-8f96-81a4c414cbed · outbound

This paper cites WithC= 256, the recurrent signal becomes too sparse and the result is nearly identical to the baseline.

Latent Recurrent Transformer: Architecture Exploration, Training Strategies, and Scaling Behavior WithC= 256, the recurrent signal becomes too sparse and the result is nearly identical to the baseline

Reference 26

Resolution
unresolved
no resolver link, observed 2026-06-29T19:36:13.559393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T19:36:13.559393Z digest=sha256:e956f5c46199e6b4f7422d8127b8705135d097f50adb9d436d61808de7c50549

Observation 89b9ffb8-b897-42dc-b563-6ec85fa23f8c · outbound

This paper cites IncreasingScreates a longer sparse refinement chain, but does not improve BPB in this setting.

Latent Recurrent Transformer: Architecture Exploration, Training Strategies, and Scaling Behavior IncreasingScreates a longer sparse refinement chain, but does not improve BPB in this setting

Reference 27

Resolution
unresolved
no resolver link, observed 2026-06-29T19:36:13.559393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T19:36:13.559393Z digest=sha256:3c93b28f85870ebf4df072245cb30c209b26a70535da99b65328db4080349bba

Pith citing papers

No inbound Pith citation observations are available.