Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T22:50:28.835675Z
Paper Citation Record · LEDGER
As of 11 August 2026, this Paper Citation Record lists 67 of 67 outbound references and 0 inbound Pith citation observations for arXiv:2501.00692.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T22:50:28.835675Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
67 of 67 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation e99b3f86-32a1-4820-8131-f8ff9c1e4fd6 · outbound
Adjoint sharding for very long context training of state space models BlackMamba: Mixture of Experts for State-Space Models
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bf8d5a77-5e96-4a43-ab94-2b34547d5a1f · outbound
Adjoint sharding for very long context training of state space models Fast Jacobian-Vector Product for Deep Networks
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ae67fdfe-2068-4f89-b70c-69ecafbfff53 · outbound
Adjoint sharding for very long context training of state space models Automatic differentiation in machine learning: a survey
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f895bba9-0913-44b8-9667-381d1a53c7a4 · outbound
Adjoint sharding for very long context training of state space models xLSTM: Extended Long Short-Term Memory
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 85553efb-415d-49f1-b461-85f0f64bcfff · outbound
Adjoint sharding for very long context training of state space models Longformer: The Long-Document Transformer
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 070ce06c-5a9b-4957-a7d4-58d66004f45b · outbound
Adjoint sharding for very long context training of state space models Internlm2 technical report,
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 0baf7150-8e83-4678-818c-fd34b341cff6 · outbound
Adjoint sharding for very long context training of state space models Adjoint sensitivity analysis for differential-algebraic equations: algorithms and software
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation aecbb979-4b57-4028-99f2-0f5bb7231634 · outbound
Adjoint sharding for very long context training of state space models Neural Ordinary Differential Equations
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation edbbc815-26f6-4d04-a078-fd8f06a5d57f · outbound
Adjoint sharding for very long context training of state space models Extending Context Window of Large Language Models via Positional Interpolation
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 624eaeba-c5b0-4f3a-8862-1ecbeae640df · outbound
Adjoint sharding for very long context training of state space models LongLoRA: Efficient Fine-tuning of Long-Context Large Language Models
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c5668d3e-01c5-47dd-b180-8403a30e76b6 · outbound
Adjoint sharding for very long context training of state space models The Backpropagation algorithm for a math student
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation ef15a127-c9d4-4244-b772-366c06a7db7d · outbound
Adjoint sharding for very long context training of state space models FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1094675c-c2f8-4d8b-a53c-f348e20cbd90 · outbound
Adjoint sharding for very long context training of state space models Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8a49de87-00d2-443a-b115-f71f4beb563f · outbound
Adjoint sharding for very long context training of state space models FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c36509f4-1270-44cc-af80-1555ab80fb33 · outbound
Adjoint sharding for very long context training of state space models Unresolved cited work
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 4277793a-c7d5-4b8b-a87b-e385df3b5f10 · outbound
Adjoint sharding for very long context training of state space models LongRoPE: Extending LLM Context Window Beyond 2 Million Tokens
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4ef388c6-e212-4c3e-961a-2c65891d49d6 · outbound
Adjoint sharding for very long context training of state space models Augmented Neural ODEs
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6b230c12-468d-4b42-a2f2-a1733d9d5ff2 · outbound
Adjoint sharding for very long context training of state space models Fu, Tri Dao, Khaled K
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 42dda1f6-ddd6-4dd3-83db-7ee9bf0e7251 · outbound
Adjoint sharding for very long context training of state space models Mamba: Linear-Time Sequence Modeling with Selective State Spaces
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 283b746e-b320-4c1b-bf46-a6f3ebf6f618 · outbound
Adjoint sharding for very long context training of state space models Combining Recurrent, Convolutional, and Continuous-time Models with Linear State-Space Layers
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dd6a2829-5bc1-4f5a-897c-8eee0cf5a9c5 · outbound
Adjoint sharding for very long context training of state space models How to Train Your HiPPO: State Space Models with Generalized Orthogonal Basis Projections
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 43a4e018-a215-4ca5-a18a-c0572863726e · outbound
Adjoint sharding for very long context training of state space models Attention mechanisms in computer vision: A survey
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 22e3801c-b225-49bc-8bc2-f8d06774a95d · outbound
Adjoint sharding for very long context training of state space models Simplifying and Understanding State Space Models with Diagonal Linear RNNs
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2bfc6337-d651-443b-aba5-7c2030a42c8b · outbound
Adjoint sharding for very long context training of state space models Deep Residual Learning for Image Recognition
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1d7e40de-c191-407a-bcb0-97e1b28d4108 · outbound
Adjoint sharding for very long context training of state space models Deep residual learning for image recognition
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4e9c8b16-e551-43a8-8a21-be6f31b3a379 · outbound
Adjoint sharding for very long context training of state space models Optimal checkpointing for heterogeneous chains: how to train deep neural networks with limited memory, 2019
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6cf268e0-d752-4456-8bdd-b6f75035f474 · outbound
Adjoint sharding for very long context training of state space models A tutorial on training recurrent neural networks , covering bppt , rtrl , ekf and the ” echo state network ” approach - semantic scholar
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 1ab5ffaa-01d8-464f-bbe5-57e8008714ae · outbound
Adjoint sharding for very long context training of state space models Adjoint methods and sensitivity analysis for recurrence, 01 2007
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 5bcfbff1-cf11-448a-81fb-c1315c223a02 · outbound
Adjoint sharding for very long context training of state space models Linear dynamical systems as a core computational primitive
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 55b31091-3be3-4176-ad29-9abff81abd8c · outbound
Adjoint sharding for very long context training of state space models Segment anything
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b45875e8-6c89-4120-814a-ba7e78919f9e · outbound
Adjoint sharding for very long context training of state space models Gonzalez, Ion Stoica, Xuezhe Ma, and Hao Zhang
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 94d34a26-4942-442e-8c0f-5424ce90dda9 · outbound
Adjoint sharding for very long context training of state space models Long-context LLMs Struggle with Long In-context Learning
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e761159d-6db4-4d7b-8acb-01dfb22f5555 · outbound
Adjoint sharding for very long context training of state space models Jamba: A Hybrid Transformer-Mamba Language Model
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d8cefb4f-686a-4888-9fbc-bd8788478359 · outbound
Adjoint sharding for very long context training of state space models Ring attention with blockwise transformers for near-infinite context,
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 3d1ec21e-95f6-47a5-9577-58695bad7677 · outbound
Adjoint sharding for very long context training of state space models World Model on Million-Length Video And Language With Blockwise RingAttention
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 009032f5-bff7-4f07-85d2-7600c62db912 · outbound
Adjoint sharding for very long context training of state space models The Llama 3 Herd of Models
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c83296ba-ed55-46b3-9b3f-d7be246c9d41 · outbound
Adjoint sharding for very long context training of state space models Mixed Precision Training
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0f8e9868-8353-4e7d-9dd9-99b5e5303970 · outbound
Adjoint sharding for very long context training of state space models Fast Finite Width Neural Tangent Kernel
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c8b882df-8c3f-48d7-a072-4ebc5c82c2db · outbound
Adjoint sharding for very long context training of state space models Matrix multiplication background user’s guide, 2024
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation a9f0488d-6a2a-473b-854b-f8504d364620 · outbound
Adjoint sharding for very long context training of state space models GPT-4 Technical Report
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6bb42725-7c29-4c50-85c3-f8604c0edca3 · outbound
Adjoint sharding for very long context training of state space models Resurrecting recurrent neural networks for long sequences, 2023
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 231e8049-bb44-49ad-a252-7d3d0182f369 · outbound
Adjoint sharding for very long context training of state space models On the difficulty of training recurrent neural networks,
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 7d849226-e6e1-4061-81f0-efbe283bd521 · outbound
Adjoint sharding for very long context training of state space models PyTorch: An Imperative Style, High-Performance Deep Learning Library
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b15134ab-d40e-45b2-ae64-92ec8b38bb1d · outbound
Adjoint sharding for very long context training of state space models Scalable diffusion models with transformers
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c1e48a4e-0276-419f-a9fe-05bf2653e47c · outbound
Adjoint sharding for very long context training of state space models RWKV: Reinventing RNNs for the Transformer Era
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cbeac19b-314b-42a3-9dcf-3144c5f2fe51 · outbound
Adjoint sharding for very long context training of state space models YaRN: Efficient Context Window Extension of Large Language Models
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f511de94-55b9-4601-b157-f583ccb982dc · outbound
Adjoint sharding for very long context training of state space models MoE-Mamba: Efficient Selective State Space Models with Mixture of Experts
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 45c9323e-e783-49d2-89b4-90629e0a17f6 · outbound
Adjoint sharding for very long context training of state space models ZeRO: Memory Optimizations Toward Training Trillion Parameter Models
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 87cd75ca-9024-4b70-9036-d4edcb65e7c0 · outbound
Adjoint sharding for very long context training of state space models ZeRO-Offload: Democratizing Billion-Scale Model Training
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a6a2acd4-bcb0-4569-8169-b00b439bd6fd · outbound
Adjoint sharding for very long context training of state space models Flashattention-3: Fast and accurate attention with asynchrony and low-precision, 2024
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation c968568b-da2d-4c22-8a6d-6eb31d3dd92b · outbound
Adjoint sharding for very long context training of state space models Low-Memory Neural Network Training: A Technical Report
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6c7a2c8c-4fcf-473d-84be-56c2712d83d7 · outbound
Adjoint sharding for very long context training of state space models Unbiasing Truncated Backpropagation Through Time
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f2eb26b8-3649-473e-97e5-0ab9ec96c32f · outbound
Adjoint sharding for very long context training of state space models Focused transformer: Contrastive training for context scaling, 2023
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 522d017f-3044-4443-8c6c-126d5db9fe90 · outbound
Adjoint sharding for very long context training of state space models Ntk-aware scaled rope, 2023
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation b8ba6447-40e7-4c8f-8362-e8d8c5f2d1d7 · outbound
Adjoint sharding for very long context training of state space models Attention Is All You Need
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 18190afd-0c24-4dfd-b0e1-9a504fd379e8 · outbound
Adjoint sharding for very long context training of state space models Rellermeyer
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ba32e0c4-e6db-46c1-9f36-5966078f0baf · outbound
Adjoint sharding for very long context training of state space models An Empirical Study of Mamba-based Language Models
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f0b74476-409f-4245-98ff-b089b3a59bd1 · outbound
Adjoint sharding for very long context training of state space models State-space Models with Layer-wise Nonlinearity are Universal Approximators with Exponential Decaying Memory
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cf45a7c3-f255-4093-b336-27e0016e5151 · outbound
Adjoint sharding for very long context training of state space models Unresolved cited work
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bc78160b-148b-433f-b4a6-75ee14b43cba · outbound
Adjoint sharding for very long context training of state space models Efficient Streaming Language Models with Attention Sinks
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b91f716d-7110-4e49-8df6-3f18f13b72f3 · outbound
Adjoint sharding for very long context training of state space models Characteristic Neural Ordinary Differential Equations
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation b399dfae-e754-4ec6-9c80-17ff7042c01a · outbound
Adjoint sharding for very long context training of state space models Focal self- attention for local-global interactions in vision transformers, 2021
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 422cebae-62f4-4436-84fe-5189484914e8 · outbound
Adjoint sharding for very long context training of state space models Long Context Compression with Activation Beacon
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b940b398-90de-4f1e-8780-c9364d5ff600 · outbound
Adjoint sharding for very long context training of state space models PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 51b3fa41-2ff7-4646-b78c-2eb3f411c157 · outbound
Adjoint sharding for very long context training of state space models On the difficulty of training Recurrent Neural Networks
Reference 2013
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6517988b-e2bb-4e52-9ceb-6487f6dc82f1 · outbound
Adjoint sharding for very long context training of state space models Ring Attention with Blockwise Transformers for Near-Infinite Context
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 29b6cc77-4fc1-4b37-974e-fff6fe7095ba · outbound
Adjoint sharding for very long context training of state space models InternLM2 Technical Report
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.