Pith. sign in

Paper Citation Record · LEDGER

DART: Decoded Attention over Recurrent States for Efficient Long-Context Sequence Modeling

As of 23 August 2026, this Paper Citation Record lists 23 of 23 outbound references and 0 inbound Pith citation observations for arXiv:2608.02032.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.02032 v1

Coverage vector

measured 23 of 23 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T16:23:16.094743Z

measured 23 of 23 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

23 of 23 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved23
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8539423f-5bdd-4887-909c-806ed7a90f7f · outbound

This paper cites Attamba: Attending To Multi-Token States.

DART: Decoded Attention over Recurrent States for Efficient Long-Context Sequence Modeling Attamba: Attending To Multi-Token States

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-04T16:23:12.468968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:23:12.468968Z digest=sha256:421da64b68b01c8a11df46973739e880259b9170f9369d12abf0dc9c115f77a9

Observation 29191822-520d-40a5-8d09-d9625091e65f · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

DART: Decoded Attention over Recurrent States for Efficient Long-Context Sequence Modeling Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-04T16:23:12.884752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:23:12.884752Z digest=sha256:22590c76177bdd81b48ff143d117abc347911f18fa9ad6ebece7dd83c57ee070

Observation d42f4839-a8b2-46af-a012-08198c9eaa30 · outbound

This paper cites Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality.

DART: Decoded Attention over Recurrent States for Efficient Long-Context Sequence Modeling Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T16:23:13.174744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:23:13.174744Z digest=sha256:ece1661f989e5d0138b6a9afab7c2665f56d273208730eb9b8665072e495ab2b

Observation 0dc0081a-e71a-46bd-9e91-054b53168889 · outbound

This paper cites Zamba: A Compact 7B SSM Hybrid Model.

DART: Decoded Attention over Recurrent States for Efficient Long-Context Sequence Modeling Zamba: A Compact 7B SSM Hybrid Model

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T16:23:13.534743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:23:13.534743Z digest=sha256:336f0c2e01440a1b570199b0bb430009f1ed08a093dc860077b6b5679970300c

Observation 41529fa1-3355-4e8e-8964-13217bb32efb · outbound

This paper cites Mamba: Linear-Time Sequence Modeling with Selective State Spaces.

DART: Decoded Attention over Recurrent States for Efficient Long-Context Sequence Modeling Mamba: Linear-Time Sequence Modeling with Selective State Spaces

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T16:23:13.654907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:23:13.654907Z digest=sha256:c421ad986104e76b22b4d44524e7c75898a9d2f22c2cc946930f110d931d2bba

Observation 74de6bad-a1d0-4c22-af69-25568eff1569 · outbound

This paper cites Repeat After Me: Transformers are Better than State Space Models at Copying.

DART: Decoded Attention over Recurrent States for Efficient Long-Context Sequence Modeling Repeat After Me: Transformers are Better than State Space Models at Copying

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T16:23:14.385860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:23:14.385860Z digest=sha256:3cb830c85d43c94162107d43489839a8b1c84702876a8803456627b7868fab21

Observation 4a83bfa8-a121-4b4c-8a1a-bab5ebd7fcad · outbound

This paper cites Jamba: A Hybrid Transformer-Mamba Language Model.

DART: Decoded Attention over Recurrent States for Efficient Long-Context Sequence Modeling Jamba: A Hybrid Transformer-Mamba Language Model

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T16:23:14.806287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:23:14.806287Z digest=sha256:2889bf340f05c1ea171de42b185aa201ca2abf27d1f95edd121f43f69cef7b5d

Observation 6cc8213f-f69f-4e40-9a7e-30d14b9e1961 · outbound

This paper cites DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model.

DART: Decoded Attention over Recurrent States for Efficient Long-Context Sequence Modeling DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T16:23:15.004747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:23:15.004747Z digest=sha256:8403a4ce6c37b0b8df67013c88389c59560fff4f6161bcab91adcf465282830e

Observation 68e105c5-4834-49ea-ba4e-94abf07cf497 · outbound

This paper cites Samba: Simple hybrid state space models for efficient unlimited context language modeling.

DART: Decoded Attention over Recurrent States for Efficient Long-Context Sequence Modeling Samba: Simple hybrid state space models for efficient unlimited context language modeling

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-04T16:23:15.545025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:23:15.545025Z digest=sha256:9820a2c8b3d96e4e2900357af63ee61b0f5b7c4cbc5bfc3493a38271896431c9

Observation bc972461-5a1d-46a1-9019-c3053e32186c · outbound

This paper cites Retentive Network: A Successor to Transformer for Large Language Models.

DART: Decoded Attention over Recurrent States for Efficient Long-Context Sequence Modeling Retentive Network: A Successor to Transformer for Large Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-04T16:23:15.815056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:23:15.815056Z digest=sha256:4d9dc2823fac1229c6f3e20afe22b9130936cfa55044fb5a54675cd1f3f7935e

Observation e1cfaf4f-82ef-4323-bf9f-613962ef5469 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

DART: Decoded Attention over Recurrent States for Efficient Long-Context Sequence Modeling LLaMA: Open and Efficient Foundation Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-04T16:23:15.954841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:23:15.954841Z digest=sha256:f11ee5a65c0e19d4f5f8f08aca79f54a1dc17ddafa871c8b039689ca41ee4357

Observation fe103f6b-2403-4d91-bb98-c93bdb2355fd · outbound

This paper cites The Pile: An 800GB Dataset of Diverse Text for Language Modeling.

DART: Decoded Attention over Recurrent States for Efficient Long-Context Sequence Modeling The Pile: An 800GB Dataset of Diverse Text for Language Modeling

Reference 1990

Resolution
unresolved
no resolver link, observed 2026-08-04T16:23:13.365450Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:23:13.365450Z digest=sha256:48f457daaacefcad34d6ee75c9090d147647e320feb4541b2a14615b0bd2de2f

Observation 93d5b400-ebab-43e1-8af2-716264a7eea2 · outbound

This paper cites RULER: What's the Real Context Size of Your Long-Context Language Models?.

DART: Decoded Attention over Recurrent States for Efficient Long-Context Sequence Modeling RULER: What's the Real Context Size of Your Long-Context Language Models?

Reference 1997

Resolution
unresolved
no resolver link, observed 2026-08-04T16:23:14.121784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:23:14.121784Z digest=sha256:bbd1d721805e99f8442bc65cb3d6ac275b4902c129acfff88270f626ff255e33

Observation 0644f03e-5047-4866-89af-e7bf7fa4b2cd · outbound

This paper cites Gated DeltaNet-2: Decoupling Erase and Write in Linear Attention.

DART: Decoded Attention over Recurrent States for Efficient Long-Context Sequence Modeling Gated DeltaNet-2: Decoupling Erase and Write in Linear Attention

Reference 2011

Resolution
unresolved
no resolver link, observed 2026-08-04T16:23:13.984846Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:23:13.984846Z digest=sha256:9ebed565adaf8ed78e07cf7c29f6b073c4c21d434c2e05020b212a937b8a00c6

Observation f96f38e5-48f7-4b0a-9bf8-08ff6006fe25 · outbound

This paper cites RWKV: Reinventing RNNs for the transformer era.

DART: Decoded Attention over Recurrent States for Efficient Long-Context Sequence Modeling RWKV: Reinventing RNNs for the transformer era

Reference 2016

Resolution
unresolved
no resolver link, observed 2026-08-04T16:23:15.234741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:23:15.234741Z digest=sha256:7ed1770d505b9f01ff1ce55c0c84f2ed3884d7596ed67b6393ec472a0860b905

Observation 2fd0833c-eee5-47a5-be9e-e8a9435eed6d · outbound

This paper cites An Empirical Study of Mamba-based Language Models.

DART: Decoded Attention over Recurrent States for Efficient Long-Context Sequence Modeling An Empirical Study of Mamba-based Language Models

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-04T16:23:16.094743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:23:16.094743Z digest=sha256:bfb703eccce8af254ea569b55174b28e0c65e37e700076f06eb86358de6ce415

Observation 7a1bbf1a-3a5d-4124-9580-439ac0650f1e · outbound

This paper cites FlashAttention-2: Faster attention with better parallelism and work partitioning.

DART: Decoded Attention over Recurrent States for Efficient Long-Context Sequence Modeling FlashAttention-2: Faster attention with better parallelism and work partitioning

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-04T16:23:12.984838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:23:12.984838Z digest=sha256:49ec3703d3dd605953d203f649c8ae0aa4b4836761282f0414bbdc6fe599a123

Observation a92629ae-e9ac-434b-98a6-be7c136bffa2 · outbound

This paper cites Li, Berlin Chen, Caitlin Wang, Aviv Bick, J.

DART: Decoded Attention over Recurrent States for Efficient Long-Context Sequence Modeling Li, Berlin Chen, Caitlin Wang, Aviv Bick, J

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-04T16:23:14.614811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:23:14.614811Z digest=sha256:3ef1365687ae58aa50514ca9d1ce91f8077352e09a242ef707c5ce0d98635fde

Observation 52be26f5-8845-4aa7-8f6a-614cf4133422 · outbound

This paper cites Language models are few-shot learners.Advances in neural information processing systems, 33:1877–1901,.

DART: Decoded Attention over Recurrent States for Efficient Long-Context Sequence Modeling Language models are few-shot learners.Advances in neural information processing systems, 33:1877–1901,

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-04T16:23:12.744337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:23:12.744337Z digest=sha256:d4c9e35de9a191d0ae8000ca8fe5afa19aeb5cddc44b62c7e2775228ff67e615

Observation 462c56b4-1714-4ddc-abfd-d3c088b1e945 · outbound

This paper cites Simplified State Space Layers for Sequence Modeling.

DART: Decoded Attention over Recurrent States for Efficient Long-Context Sequence Modeling Simplified State Space Layers for Sequence Modeling

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-04T16:23:15.687548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:23:15.687548Z digest=sha256:8ff8c6628f4f82e34d74824a1c6e5d144d194d2ce3961df97184aa7f27a6fda5

Observation a5ac6537-4fa8-4327-a38a-e30bb26ea72a · outbound

This paper cites DROP: A reading comprehension benchmark requiring discrete reasoning over paragraphs.

DART: Decoded Attention over Recurrent States for Efficient Long-Context Sequence Modeling DROP: A reading comprehension benchmark requiring discrete reasoning over paragraphs

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-04T16:23:13.234740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:23:13.234740Z digest=sha256:457e7cd74ad6e3913d9e19403ab234fe2cf9cf73dd52c3864d4c65dc299d5104

Observation 8c6b10b9-dc50-4311-b5f6-2b2c1d1e50a1 · outbound

This paper cites Efficiently Modeling Long Sequences with Structured State Spaces.

DART: Decoded Attention over Recurrent States for Efficient Long-Context Sequence Modeling Efficiently Modeling Long Sequences with Structured State Spaces

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-04T16:23:13.798657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:23:13.798657Z digest=sha256:7e3f47f742d25b904b4a232d850174f8856a19c9e9cd3c01581480afaf0be0fe

Observation 965468b8-0054-4ffc-9657-b2b217569011 · outbound

This paper cites Simple linear attention language models balance the recall-throughput tradeoff.

DART: Decoded Attention over Recurrent States for Efficient Long-Context Sequence Modeling Simple linear attention language models balance the recall-throughput tradeoff

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-04T16:23:12.574467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:23:12.574467Z digest=sha256:a58b6801bc3978f5e82c60f32210e9b79d2d65749ada9dcaffcaf8ab7bf71fb0

Pith citing papers

No inbound Pith citation observations are available.