Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-06-29T19:36:13.559393Z
Paper Citation Record · LEDGER
As of 4 August 2026, this Paper Citation Record lists 27 of 27 outbound references and 0 inbound Pith citation observations for arXiv:2605.26797.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-06-29T19:36:13.559393Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
27 of 27 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 6bdb8679-bbd1-4df7-b8e6-e25b92d6532a · outbound
Latent Recurrent Transformer: Architecture Exploration, Training Strategies, and Scaling Behavior Language models are few-shot learners
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2efbb355-3685-403c-8cb5-003ce1c758cb · outbound
Latent Recurrent Transformer: Architecture Exploration, Training Strategies, and Scaling Behavior Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 99c4b0bc-06bf-42fb-8651-5f97cf7b9336 · outbound
Latent Recurrent Transformer: Architecture Exploration, Training Strategies, and Scaling Behavior Addressing Some Limitations of Transformers with Feedback Memory
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation dab136e2-21ca-4c32-a2b2-f48ea4cd2812 · outbound
Latent Recurrent Transformer: Architecture Exploration, Training Strategies, and Scaling Behavior Hungry Hungry Hippos: Towards Language Modeling with State Space Models
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 8a7b088d-d4cd-41b9-a7d4-4cfea3c4b9c8 · outbound
Latent Recurrent Transformer: Architecture Exploration, Training Strategies, and Scaling Behavior Think before you speak: Training language models with pause tokens
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f6801829-3569-41c0-9877-f7a91b56e039 · outbound
Latent Recurrent Transformer: Architecture Exploration, Training Strategies, and Scaling Behavior Mamba: Linear-Time Sequence Modeling with Selective State Spaces
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation ff661fe8-d849-4d21-ba1d-b9ee4cbe9c2f · outbound
Latent Recurrent Transformer: Architecture Exploration, Training Strategies, and Scaling Behavior Training Large Language Models to Reason in a Continuous Latent Space
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation cc9c82eb-31f1-4a78-b661-ab6fa781d8fa · outbound
Latent Recurrent Transformer: Architecture Exploration, Training Strategies, and Scaling Behavior Thinking Tokens for Language Modeling
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 8f438508-705f-47b5-9e66-0fe83f8cde05 · outbound
Latent Recurrent Transformer: Architecture Exploration, Training Strategies, and Scaling Behavior TransformerFAM: Feedback attention is working memory
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 788d6530-a7dd-4091-af6d-179219879769 · outbound
Latent Recurrent Transformer: Architecture Exploration, Training Strategies, and Scaling Behavior Muon: An optimizer for hidden layers in neural networks, 2024.URL https://kellerjordan
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7a1b5163-bc86-4017-b0ac-ae3dd66d0c2e · outbound
Latent Recurrent Transformer: Architecture Exploration, Training Strategies, and Scaling Behavior Jamba: A Hybrid Transformer-Mamba Language Model
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 1d46141e-98e7-48cb-a689-2edde5adce72 · outbound
Latent Recurrent Transformer: Architecture Exploration, Training Strategies, and Scaling Behavior Fineweb-edu: the finest collection of educational content, 2024.URL https://huggingface
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b602e78d-e8f5-434d-aa23-e7433373862a · outbound
Latent Recurrent Transformer: Architecture Exploration, Training Strategies, and Scaling Behavior Landmark Attention: Random-Access Infinite Context Length for Transformers
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 16341c4e-7eed-455c-a832-a1231670ebf2 · outbound
Latent Recurrent Transformer: Architecture Exploration, Training Strategies, and Scaling Behavior Leave No Context Behind: Efficient Infinite Context Transformers with Infini-attention
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 5d043d28-2feb-4069-a115-80f40c889473 · outbound
Latent Recurrent Transformer: Architecture Exploration, Training Strategies, and Scaling Behavior RWKV: Reinventing RNNs for the Transformer Era
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation efb1b40d-02f1-405a-85e2-3ba3052b2561 · outbound
Latent Recurrent Transformer: Architecture Exploration, Training Strategies, and Scaling Behavior Let's Think Dot by Dot: Hidden Computation in Transformer Language Models
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation eb9e9bd5-3a79-4c8d-9263-11c7c44f33b2 · outbound
Latent Recurrent Transformer: Architecture Exploration, Training Strategies, and Scaling Behavior Compressive Transformers for Long-Range Sequence Modelling
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation e95820fa-5fb5-4599-9941-b201aaeb1536 · outbound
Latent Recurrent Transformer: Architecture Exploration, Training Strategies, and Scaling Behavior Simplified State Space Layers for Sequence Modeling
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation f1089c86-5ba0-421b-8a27-89a4c9df6810 · outbound
Latent Recurrent Transformer: Architecture Exploration, Training Strategies, and Scaling Behavior Learning to (Learn at Test Time): RNNs with Expressive Hidden States
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 0b0f7268-8c3f-4441-8b75-20e48d214f03 · outbound
Latent Recurrent Transformer: Architecture Exploration, Training Strategies, and Scaling Behavior Retentive Network: A Successor to Transformer for Large Language Models
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 4dd94e11-ba68-4b2f-a850-a88bbdcf7c16 · outbound
Latent Recurrent Transformer: Architecture Exploration, Training Strategies, and Scaling Behavior Gemma 2: Improving Open Language Models at a Practical Size
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 3f0b871a-b078-46b8-92de-780ac6053677 · outbound
Latent Recurrent Transformer: Architecture Exploration, Training Strategies, and Scaling Behavior Memorizing Transformers
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation f52560f1-2bb2-4c67-a37e-0bd395dd25f4 · outbound
Latent Recurrent Transformer: Architecture Exploration, Training Strategies, and Scaling Behavior Quiet-STaR: Language Models Can Teach Themselves to Think Before Speaking
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 9b03380b-245f-4a4c-86cf-9f83cd2b7cf3 · outbound
Latent Recurrent Transformer: Architecture Exploration, Training Strategies, and Scaling Behavior Unresolved cited work
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6060a32c-504e-448a-9b92-adc50a8054ea · outbound
Latent Recurrent Transformer: Architecture Exploration, Training Strategies, and Scaling Behavior The final shard is held out as validation and the remaining shards are used for training
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2519ddef-6310-46c1-8f96-81a4c414cbed · outbound
Latent Recurrent Transformer: Architecture Exploration, Training Strategies, and Scaling Behavior WithC= 256, the recurrent signal becomes too sparse and the result is nearly identical to the baseline
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 89b9ffb8-b897-42dc-b563-6ec85fa23f8c · outbound
Latent Recurrent Transformer: Architecture Exploration, Training Strategies, and Scaling Behavior IncreasingScreates a longer sparse refinement chain, but does not improve BPB in this setting
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.