Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-04T16:23:16.094743Z
Paper Citation Record · LEDGER
As of 23 August 2026, this Paper Citation Record lists 23 of 23 outbound references and 0 inbound Pith citation observations for arXiv:2608.02032.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-04T16:23:16.094743Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
23 of 23 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 8539423f-5bdd-4887-909c-806ed7a90f7f · outbound
DART: Decoded Attention over Recurrent States for Efficient Long-Context Sequence Modeling Attamba: Attending To Multi-Token States
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 29191822-520d-40a5-8d09-d9625091e65f · outbound
DART: Decoded Attention over Recurrent States for Efficient Long-Context Sequence Modeling Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d42f4839-a8b2-46af-a012-08198c9eaa30 · outbound
DART: Decoded Attention over Recurrent States for Efficient Long-Context Sequence Modeling Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0dc0081a-e71a-46bd-9e91-054b53168889 · outbound
DART: Decoded Attention over Recurrent States for Efficient Long-Context Sequence Modeling Zamba: A Compact 7B SSM Hybrid Model
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 41529fa1-3355-4e8e-8964-13217bb32efb · outbound
DART: Decoded Attention over Recurrent States for Efficient Long-Context Sequence Modeling Mamba: Linear-Time Sequence Modeling with Selective State Spaces
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 74de6bad-a1d0-4c22-af69-25568eff1569 · outbound
DART: Decoded Attention over Recurrent States for Efficient Long-Context Sequence Modeling Repeat After Me: Transformers are Better than State Space Models at Copying
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4a83bfa8-a121-4b4c-8a1a-bab5ebd7fcad · outbound
DART: Decoded Attention over Recurrent States for Efficient Long-Context Sequence Modeling Jamba: A Hybrid Transformer-Mamba Language Model
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6cc8213f-f69f-4e40-9a7e-30d14b9e1961 · outbound
DART: Decoded Attention over Recurrent States for Efficient Long-Context Sequence Modeling DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 68e105c5-4834-49ea-ba4e-94abf07cf497 · outbound
DART: Decoded Attention over Recurrent States for Efficient Long-Context Sequence Modeling Samba: Simple hybrid state space models for efficient unlimited context language modeling
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bc972461-5a1d-46a1-9019-c3053e32186c · outbound
DART: Decoded Attention over Recurrent States for Efficient Long-Context Sequence Modeling Retentive Network: A Successor to Transformer for Large Language Models
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e1cfaf4f-82ef-4323-bf9f-613962ef5469 · outbound
DART: Decoded Attention over Recurrent States for Efficient Long-Context Sequence Modeling LLaMA: Open and Efficient Foundation Language Models
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fe103f6b-2403-4d91-bb98-c93bdb2355fd · outbound
DART: Decoded Attention over Recurrent States for Efficient Long-Context Sequence Modeling The Pile: An 800GB Dataset of Diverse Text for Language Modeling
Reference 1990
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 93d5b400-ebab-43e1-8af2-716264a7eea2 · outbound
DART: Decoded Attention over Recurrent States for Efficient Long-Context Sequence Modeling RULER: What's the Real Context Size of Your Long-Context Language Models?
Reference 1997
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0644f03e-5047-4866-89af-e7bf7fa4b2cd · outbound
DART: Decoded Attention over Recurrent States for Efficient Long-Context Sequence Modeling Gated DeltaNet-2: Decoupling Erase and Write in Linear Attention
Reference 2011
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f96f38e5-48f7-4b0a-9bf8-08ff6006fe25 · outbound
DART: Decoded Attention over Recurrent States for Efficient Long-Context Sequence Modeling RWKV: Reinventing RNNs for the transformer era
Reference 2016
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2fd0833c-eee5-47a5-be9e-e8a9435eed6d · outbound
DART: Decoded Attention over Recurrent States for Efficient Long-Context Sequence Modeling An Empirical Study of Mamba-based Language Models
Reference 2017
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7a1bbf1a-3a5d-4124-9580-439ac0650f1e · outbound
DART: Decoded Attention over Recurrent States for Efficient Long-Context Sequence Modeling FlashAttention-2: Faster attention with better parallelism and work partitioning
Reference 2018
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a92629ae-e9ac-434b-98a6-be7c136bffa2 · outbound
DART: Decoded Attention over Recurrent States for Efficient Long-Context Sequence Modeling Li, Berlin Chen, Caitlin Wang, Aviv Bick, J
Reference 2019
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 52be26f5-8845-4aa7-8f6a-614cf4133422 · outbound
DART: Decoded Attention over Recurrent States for Efficient Long-Context Sequence Modeling Language models are few-shot learners.Advances in neural information processing systems, 33:1877–1901,
Reference 2020
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 462c56b4-1714-4ddc-abfd-d3c088b1e945 · outbound
DART: Decoded Attention over Recurrent States for Efficient Long-Context Sequence Modeling Simplified State Space Layers for Sequence Modeling
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a5ac6537-4fa8-4327-a38a-e30bb26ea72a · outbound
DART: Decoded Attention over Recurrent States for Efficient Long-Context Sequence Modeling DROP: A reading comprehension benchmark requiring discrete reasoning over paragraphs
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8c6b10b9-dc50-4311-b5f6-2b2c1d1e50a1 · outbound
DART: Decoded Attention over Recurrent States for Efficient Long-Context Sequence Modeling Efficiently Modeling Long Sequences with Structured State Spaces
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 965468b8-0054-4ffc-9657-b2b217569011 · outbound
DART: Decoded Attention over Recurrent States for Efficient Long-Context Sequence Modeling Simple linear attention language models balance the recall-throughput tradeoff
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.