Pith. sign in

Paper Citation Record · LEDGER

Adjoint sharding for very long context training of state space models

As of 11 August 2026, this Paper Citation Record lists 67 of 67 outbound references and 0 inbound Pith citation observations for arXiv:2501.00692.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.00692 v1

Coverage vector

measured 67 of 67 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T22:50:28.835675Z

measured 67 of 67 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

67 of 67 outbound references displayed

  • verified exact3
  • verified fuzzy15
  • unresolved49
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e99b3f86-32a1-4820-8131-f8ff9c1e4fd6 · outbound

This paper cites BlackMamba: Mixture of Experts for State-Space Models.

Adjoint sharding for very long context training of state space models BlackMamba: Mixture of Experts for State-Space Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T22:50:28.484069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:50:28.484069Z digest=sha256:603d9e1685f95bf94fca62c039d4a53198015497d6b1b2c422a5e150b1181d2c

Observation bf8d5a77-5e96-4a43-ab94-2b34547d5a1f · outbound

This paper cites Fast Jacobian-Vector Product for Deep Networks.

Adjoint sharding for very long context training of state space models Fast Jacobian-Vector Product for Deep Networks

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T22:50:28.491015Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:50:28.491015Z digest=sha256:adf022414c933a4c1893323620b83b6d86c95cffbf69b51d1d2f5a84720d418b

Observation ae67fdfe-2068-4f89-b70c-69ecafbfff53 · outbound

This paper cites Automatic differentiation in machine learning: a survey.

Adjoint sharding for very long context training of state space models Automatic differentiation in machine learning: a survey

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T22:50:28.502732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:50:28.502732Z digest=sha256:e9f23d8b80caffd75d8f0f30acb50a9d3c8a20c7bce8376b5cec856856c49fef

Observation f895bba9-0913-44b8-9667-381d1a53c7a4 · outbound

This paper cites xLSTM: Extended Long Short-Term Memory.

Adjoint sharding for very long context training of state space models xLSTM: Extended Long Short-Term Memory

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T22:50:28.508274Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:50:28.508274Z digest=sha256:273a9b5e8236c1907becd20694f7bac7730b4dd705275182e22b8fa9ec46cec3

Observation 85553efb-415d-49f1-b461-85f0f64bcfff · outbound

This paper cites Longformer: The Long-Document Transformer.

Adjoint sharding for very long context training of state space models Longformer: The Long-Document Transformer

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T22:50:28.514062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:50:28.514062Z digest=sha256:f226168e8da645813c384e4caa1ae5251e2dcbe94270be90d3333c6dd160772c

Observation 070ce06c-5a9b-4957-a7d4-58d66004f45b · outbound

This paper cites Internlm2 technical report,.

Adjoint sharding for very long context training of state space models Internlm2 technical report,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:50:29.968339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T22:50:28.519633Z digest=sha256:372217f51b7cb58b1e864d3b9413d21233d0ff5a8815158419821dea47bbb095

Observation 0baf7150-8e83-4678-818c-fd34b341cff6 · outbound

This paper cites Adjoint sensitivity analysis for differential-algebraic equations: algorithms and software.

Adjoint sharding for very long context training of state space models Adjoint sensitivity analysis for differential-algebraic equations: algorithms and software

Reference 8

Resolution
verified exact
doi, observed 2026-08-10T22:50:28.898216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T22:50:28.529339Z digest=sha256:19247e3616c645521ab18445691496b8a2aa5228e2ebb9c097aef8be165638ca

Observation aecbb979-4b57-4028-99f2-0f5bb7231634 · outbound

This paper cites Neural Ordinary Differential Equations.

Adjoint sharding for very long context training of state space models Neural Ordinary Differential Equations

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T22:50:28.535476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:50:28.535476Z digest=sha256:99132bab7364f08636a1aa107a34b841b67a181bdec0f5cbe972366a229d55bc

Observation edbbc815-26f6-4d04-a078-fd8f06a5d57f · outbound

This paper cites Extending Context Window of Large Language Models via Positional Interpolation.

Adjoint sharding for very long context training of state space models Extending Context Window of Large Language Models via Positional Interpolation

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T22:50:28.540632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:50:28.540632Z digest=sha256:4e92c67e3b4b4711cefa6c5fed065780f0c43f70d9d28a49e5e870c9efca58dc

Observation 624eaeba-c5b0-4f3a-8862-1ecbeae640df · outbound

This paper cites LongLoRA: Efficient Fine-tuning of Long-Context Large Language Models.

Adjoint sharding for very long context training of state space models LongLoRA: Efficient Fine-tuning of Long-Context Large Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T22:50:28.546514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:50:28.546514Z digest=sha256:3c3e7af201a176d85f280d0c3b5c8b3cbd652963fdce69e77d773ff8116c498b

Observation c5668d3e-01c5-47dd-b180-8403a30e76b6 · outbound

This paper cites The Backpropagation algorithm for a math student.

Adjoint sharding for very long context training of state space models The Backpropagation algorithm for a math student

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-08-10T22:50:29.547231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T22:50:28.551587Z digest=sha256:c4d3665463d2f3893de4dc2e10bfe60ade5d2b3d9a6613b1017a9011d12876a1

Observation ef15a127-c9d4-4244-b772-366c06a7db7d · outbound

This paper cites FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning.

Adjoint sharding for very long context training of state space models FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T22:50:28.557753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:50:28.557753Z digest=sha256:efdf2ebbb82afe86f7c58dbf58bd4f70740a6951bfc2c0eb473af97c7a7ce990

Observation 1094675c-c2f8-4d8b-a53c-f348e20cbd90 · outbound

This paper cites Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality.

Adjoint sharding for very long context training of state space models Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T22:50:28.568578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:50:28.568578Z digest=sha256:4d333bde8a040ca561cae23b5060aff4690d3db6a4241d7174c8d8d3039cbf5c

Observation 8a49de87-00d2-443a-b115-f71f4beb563f · outbound

This paper cites FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness.

Adjoint sharding for very long context training of state space models FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T22:50:28.573058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:50:28.573058Z digest=sha256:869cbc5341b9981f4fb2a10829e7dfc3cccc44c1aa46f2e26c7c49c1d41bae9c

Observation c36509f4-1270-44cc-af80-1555ab80fb33 · outbound

This paper cites an unresolved cited work.

Adjoint sharding for very long context training of state space models Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:50:29.953519Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T22:50:28.579575Z digest=sha256:5c7c44a982b8797947a514f29f4d66de8192fdfbff1b9b73652edda4dea12df5

Observation 4277793a-c7d5-4b8b-a87b-e385df3b5f10 · outbound

This paper cites LongRoPE: Extending LLM Context Window Beyond 2 Million Tokens.

Adjoint sharding for very long context training of state space models LongRoPE: Extending LLM Context Window Beyond 2 Million Tokens

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T22:50:28.589583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:50:28.589583Z digest=sha256:6273ead86a414245597a94462ccc6a4c3ed39ffa33a7164c2826a86ea59e88a2

Observation 4ef388c6-e212-4c3e-961a-2c65891d49d6 · outbound

This paper cites Augmented Neural ODEs.

Adjoint sharding for very long context training of state space models Augmented Neural ODEs

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T22:50:28.593837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:50:28.593837Z digest=sha256:651f93d846c58ae28059194b05613bed652e1107222abf4d8708a81cd37ab125

Observation 6b230c12-468d-4b42-a2f2-a1733d9d5ff2 · outbound

This paper cites Fu, Tri Dao, Khaled K.

Adjoint sharding for very long context training of state space models Fu, Tri Dao, Khaled K

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:50:29.937435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T22:50:28.598096Z digest=sha256:791b6f7341f451f24590a613227a1b3db7b673e0770430a7d2828fd8be6965b9

Observation 42dda1f6-ddd6-4dd3-83db-7ee9bf0e7251 · outbound

This paper cites Mamba: Linear-Time Sequence Modeling with Selective State Spaces.

Adjoint sharding for very long context training of state space models Mamba: Linear-Time Sequence Modeling with Selective State Spaces

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T22:50:28.606419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:50:28.606419Z digest=sha256:701e5219b68846e6a236babb43bd1c17ab5931944675cd586a5b8419574a0a4c

Observation 283b746e-b320-4c1b-bf46-a6f3ebf6f618 · outbound

This paper cites Combining Recurrent, Convolutional, and Continuous-time Models with Linear State-Space Layers.

Adjoint sharding for very long context training of state space models Combining Recurrent, Convolutional, and Continuous-time Models with Linear State-Space Layers

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T22:50:28.610398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:50:28.610398Z digest=sha256:3d3ffc054f0d546c2a53f396785bdb6142c0e60efcffb44c48a85d82b42dfc0b

Observation dd6a2829-5bc1-4f5a-897c-8eee0cf5a9c5 · outbound

This paper cites How to Train Your HiPPO: State Space Models with Generalized Orthogonal Basis Projections.

Adjoint sharding for very long context training of state space models How to Train Your HiPPO: State Space Models with Generalized Orthogonal Basis Projections

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-10T22:50:28.614831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:50:28.614831Z digest=sha256:1d095fcedc10ae185f67115891e6428506427513250b7d3c7ef71d6ae4d2340b

Observation 43a4e018-a215-4ca5-a18a-c0572863726e · outbound

This paper cites Attention mechanisms in computer vision: A survey.

Adjoint sharding for very long context training of state space models Attention mechanisms in computer vision: A survey

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:50:29.921783Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T22:50:28.619200Z digest=sha256:aea1b00c45d24160d5732f881040dd2d3a37c71beeb232a85d6320f19d434a40

Observation 22e3801c-b225-49bc-8bc2-f8d06774a95d · outbound

This paper cites Simplifying and Understanding State Space Models with Diagonal Linear RNNs.

Adjoint sharding for very long context training of state space models Simplifying and Understanding State Space Models with Diagonal Linear RNNs

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-10T22:50:28.623895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:50:28.623895Z digest=sha256:dd8b124372cb53e32b8c543eb9cfaaaa54cb63896d9dabe3b917a06868219ad4

Observation 2bfc6337-d651-443b-aba5-7c2030a42c8b · outbound

This paper cites Deep Residual Learning for Image Recognition.

Adjoint sharding for very long context training of state space models Deep Residual Learning for Image Recognition

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-10T22:50:28.629714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:50:28.629714Z digest=sha256:2d5d092227bac48c463fe586c427147aec19558828f79477c7878cfa4a4da1ba

Observation 1d7e40de-c191-407a-bcb0-97e1b28d4108 · outbound

This paper cites Deep residual learning for image recognition.

Adjoint sharding for very long context training of state space models Deep residual learning for image recognition

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-10T22:50:28.634649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:50:28.634649Z digest=sha256:40b213b0d5a5e485abd48ec84dd598233d755f894d2a74950d6dccd20dd15764

Observation 4e9c8b16-e551-43a8-8a21-be6f31b3a379 · outbound

This paper cites Optimal checkpointing for heterogeneous chains: how to train deep neural networks with limited memory, 2019.

Adjoint sharding for very long context training of state space models Optimal checkpointing for heterogeneous chains: how to train deep neural networks with limited memory, 2019

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-10T22:50:28.639358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:50:28.639358Z digest=sha256:dd712ee87e6b4627fd78553d6a78037d6cc90b107080270bf83e72631a750d3e

Observation 6cf268e0-d752-4456-8bdd-b6f75035f474 · outbound

This paper cites A tutorial on training recurrent neural networks , covering bppt , rtrl , ekf and the ” echo state network ” approach - semantic scholar.

Adjoint sharding for very long context training of state space models A tutorial on training recurrent neural networks , covering bppt , rtrl , ekf and the ” echo state network ” approach - semantic scholar

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:50:29.896896Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T22:50:28.644010Z digest=sha256:f67faa550fd9b3a054cf38e8507e2bc8be774a8aa4336ce7d61a6f637ddefb43

Observation 1ab5ffaa-01d8-464f-bbe5-57e8008714ae · outbound

This paper cites Adjoint methods and sensitivity analysis for recurrence, 01 2007.

Adjoint sharding for very long context training of state space models Adjoint methods and sensitivity analysis for recurrence, 01 2007

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:50:29.880912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T22:50:28.648655Z digest=sha256:701bc82bb31bf48eacdd1e991d94f349321c7ad3fe831cb08a30e414f6c52903

Observation 5bcfbff1-cf11-448a-81fb-c1315c223a02 · outbound

This paper cites Linear dynamical systems as a core computational primitive.

Adjoint sharding for very long context training of state space models Linear dynamical systems as a core computational primitive

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:50:29.865266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T22:50:28.653859Z digest=sha256:42f359d2ee3a13404dc256dfd017eb7c87fe266a8942251c893c1c1cee38df75

Observation 55b31091-3be3-4176-ad29-9abff81abd8c · outbound

This paper cites Segment anything.

Adjoint sharding for very long context training of state space models Segment anything

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-10T22:50:28.658687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:50:28.658687Z digest=sha256:b4180a2d8406003eec55feb607939b41d37d98cf4c2f052856f4379dc653ff2d

Observation b45875e8-6c89-4120-814a-ba7e78919f9e · outbound

This paper cites Gonzalez, Ion Stoica, Xuezhe Ma, and Hao Zhang.

Adjoint sharding for very long context training of state space models Gonzalez, Ion Stoica, Xuezhe Ma, and Hao Zhang

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:50:29.838923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T22:50:28.663400Z digest=sha256:73b7310e1475121ad35ee515a0ff985a96881708170c0f25b0c702b37aaab6ac

Observation 94d34a26-4942-442e-8c0f-5424ce90dda9 · outbound

This paper cites Long-context LLMs Struggle with Long In-context Learning.

Adjoint sharding for very long context training of state space models Long-context LLMs Struggle with Long In-context Learning

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-10T22:50:28.668601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:50:28.668601Z digest=sha256:7bb81e185f3aad2b06d0cbe1edf0a74654067960e8b422b70be3f0a6b5a94daa

Observation e761159d-6db4-4d7b-8acb-01dfb22f5555 · outbound

This paper cites Jamba: A Hybrid Transformer-Mamba Language Model.

Adjoint sharding for very long context training of state space models Jamba: A Hybrid Transformer-Mamba Language Model

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-10T22:50:28.673258Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:50:28.673258Z digest=sha256:f360f25963cf20fd032a0ca316f4e53d41ca020818f085731a17ebdf80022f82

Observation d8cefb4f-686a-4888-9fbc-bd8788478359 · outbound

This paper cites Ring attention with blockwise transformers for near-infinite context,.

Adjoint sharding for very long context training of state space models Ring attention with blockwise transformers for near-infinite context,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:50:29.822170Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T22:50:28.678219Z digest=sha256:baa957d83260bcad67b4ae11a1a183bb12923d8a4113209344804550397e6b99

Observation 3d1ec21e-95f6-47a5-9577-58695bad7677 · outbound

This paper cites World Model on Million-Length Video And Language With Blockwise RingAttention.

Adjoint sharding for very long context training of state space models World Model on Million-Length Video And Language With Blockwise RingAttention

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-10T22:50:28.688390Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:50:28.688390Z digest=sha256:942f6e8e4eae157a901127ca5c6ef4259b855ef78880810348202a0d4e326142

Observation 009032f5-bff7-4f07-85d2-7600c62db912 · outbound

This paper cites The Llama 3 Herd of Models.

Adjoint sharding for very long context training of state space models The Llama 3 Herd of Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-10T22:50:28.693596Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:50:28.693596Z digest=sha256:7cf08b6dfe351ed7e7c028fd9c1fbf6907086da55538baa184bc135dba5b99c1

Observation c83296ba-ed55-46b3-9b3f-d7be246c9d41 · outbound

This paper cites Mixed Precision Training.

Adjoint sharding for very long context training of state space models Mixed Precision Training

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-10T22:50:28.698249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:50:28.698249Z digest=sha256:d5857501f7d88c64845775f7c206ed71758c28bc5ed6919ca39dbada297230a5

Observation 0f8e9868-8353-4e7d-9dd9-99b5e5303970 · outbound

This paper cites Fast Finite Width Neural Tangent Kernel.

Adjoint sharding for very long context training of state space models Fast Finite Width Neural Tangent Kernel

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-10T22:50:28.703665Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:50:28.703665Z digest=sha256:864eb057533abae2a48069d85e6838478a8718f5390b0efd3a4e35b020fd0a4f

Observation c8b882df-8c3f-48d7-a072-4ebc5c82c2db · outbound

This paper cites Matrix multiplication background user’s guide, 2024.

Adjoint sharding for very long context training of state space models Matrix multiplication background user’s guide, 2024

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:50:29.806652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T22:50:28.708445Z digest=sha256:ac7f4ac74e9922a0b504be7c4f71523594c78387ba31bdf8f8bca140215256bb

Observation a9f0488d-6a2a-473b-854b-f8504d364620 · outbound

This paper cites GPT-4 Technical Report.

Adjoint sharding for very long context training of state space models GPT-4 Technical Report

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-10T22:50:28.713364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:50:28.713364Z digest=sha256:b365360fde31c8f11a6e3f5de8316765b86a177770b415eb118374ee365e4961

Observation 6bb42725-7c29-4c50-85c3-f8604c0edca3 · outbound

This paper cites Resurrecting recurrent neural networks for long sequences, 2023.

Adjoint sharding for very long context training of state space models Resurrecting recurrent neural networks for long sequences, 2023

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:50:29.790576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T22:50:28.717938Z digest=sha256:91f9afa0705039c1be1da70d5bd4ebbd88617093f079e8678faaadf12c612c0c

Observation 231e8049-bb44-49ad-a252-7d3d0182f369 · outbound

This paper cites On the difficulty of training recurrent neural networks,.

Adjoint sharding for very long context training of state space models On the difficulty of training recurrent neural networks,

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:50:29.774729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T22:50:28.722494Z digest=sha256:5678f31a17f2dddcd3c415bd5f41765ac5b63b3370a8d80dc82f0a71792194a1

Observation 7d849226-e6e1-4061-81f0-efbe283bd521 · outbound

This paper cites PyTorch: An Imperative Style, High-Performance Deep Learning Library.

Adjoint sharding for very long context training of state space models PyTorch: An Imperative Style, High-Performance Deep Learning Library

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-10T22:50:28.733647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:50:28.733647Z digest=sha256:068b6f9439472aaead8c10fe7720afc54b3d679b8514ffc070ca21ef4851488c

Observation b15134ab-d40e-45b2-ae64-92ec8b38bb1d · outbound

This paper cites Scalable diffusion models with transformers.

Adjoint sharding for very long context training of state space models Scalable diffusion models with transformers

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-10T22:50:28.738118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:50:28.738118Z digest=sha256:f3ee5833c1f02fbb449547b39dc10bba15ed9087522ecf4adbfa2a5c3e365b6a

Observation c1e48a4e-0276-419f-a9fe-05bf2653e47c · outbound

This paper cites RWKV: Reinventing RNNs for the Transformer Era.

Adjoint sharding for very long context training of state space models RWKV: Reinventing RNNs for the Transformer Era

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-10T22:50:28.742462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:50:28.742462Z digest=sha256:2185b4012ca0d89ac86ffe487747cd217fa45cc07f6d5f493cd655365cd7f23a

Observation cbeac19b-314b-42a3-9dcf-3144c5f2fe51 · outbound

This paper cites YaRN: Efficient Context Window Extension of Large Language Models.

Adjoint sharding for very long context training of state space models YaRN: Efficient Context Window Extension of Large Language Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-10T22:50:28.746548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:50:28.746548Z digest=sha256:e150ca451286b99c54e63e40b901d7c3727b81143bf87438b83542a36ca04fb5

Observation f511de94-55b9-4601-b157-f583ccb982dc · outbound

This paper cites MoE-Mamba: Efficient Selective State Space Models with Mixture of Experts.

Adjoint sharding for very long context training of state space models MoE-Mamba: Efficient Selective State Space Models with Mixture of Experts

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-10T22:50:28.750951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:50:28.750951Z digest=sha256:c0694c2b8e8f84bbe806d82887b805fc6c94f14f2b2f0caa3a23aaa278dae214

Observation 45c9323e-e783-49d2-89b4-90629e0a17f6 · outbound

This paper cites ZeRO: Memory Optimizations Toward Training Trillion Parameter Models.

Adjoint sharding for very long context training of state space models ZeRO: Memory Optimizations Toward Training Trillion Parameter Models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-10T22:50:28.760927Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:50:28.760927Z digest=sha256:4ade69836f2010042b3e96654e3d1409ec348bfd25ef2b411395fd386c999480

Observation 87cd75ca-9024-4b70-9036-d4edcb65e7c0 · outbound

This paper cites ZeRO-Offload: Democratizing Billion-Scale Model Training.

Adjoint sharding for very long context training of state space models ZeRO-Offload: Democratizing Billion-Scale Model Training

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-10T22:50:28.765359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:50:28.765359Z digest=sha256:b27f37927ab2fb5c6912faafcd6be89d9a662e1914db2455f32c4269961563f7

Observation a6a2acd4-bcb0-4569-8169-b00b439bd6fd · outbound

This paper cites Flashattention-3: Fast and accurate attention with asynchrony and low-precision, 2024.

Adjoint sharding for very long context training of state space models Flashattention-3: Fast and accurate attention with asynchrony and low-precision, 2024

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:50:29.749372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T22:50:28.770760Z digest=sha256:5d4c685aff3fc6547df24ae2b729d60d851797bd8fc4e5dc0ae87f47b8e329ff

Observation c968568b-da2d-4c22-8a6d-6eb31d3dd92b · outbound

This paper cites Low-Memory Neural Network Training: A Technical Report.

Adjoint sharding for very long context training of state space models Low-Memory Neural Network Training: A Technical Report

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-10T22:50:28.775364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:50:28.775364Z digest=sha256:e070880e7a6d1fff6443301d144d3fec17d5a4b6c6bf0c4f05aa17a3825e4801

Observation 6c7a2c8c-4fcf-473d-84be-56c2712d83d7 · outbound

This paper cites Unbiasing Truncated Backpropagation Through Time.

Adjoint sharding for very long context training of state space models Unbiasing Truncated Backpropagation Through Time

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-10T22:50:28.780156Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:50:28.780156Z digest=sha256:0e9e197b34f6f070313bdafb396831ab41d5adc48f825e66f2f1591058dc6232

Observation f2eb26b8-3649-473e-97e5-0ab9ec96c32f · outbound

This paper cites Focused transformer: Contrastive training for context scaling, 2023.

Adjoint sharding for very long context training of state space models Focused transformer: Contrastive training for context scaling, 2023

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:50:29.733806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T22:50:28.785242Z digest=sha256:10e02721d32917181e87137146a74ca198d1685d480a87afea632cd7c5d22259

Observation 522d017f-3044-4443-8c6c-126d5db9fe90 · outbound

This paper cites Ntk-aware scaled rope, 2023.

Adjoint sharding for very long context training of state space models Ntk-aware scaled rope, 2023

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:50:29.719374Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T22:50:28.790111Z digest=sha256:823d7071b4346f2a60671c3f99c91dcd457c66c811b43935826e00167117fda8

Observation b8ba6447-40e7-4c8f-8362-e8d8c5f2d1d7 · outbound

This paper cites Attention Is All You Need.

Adjoint sharding for very long context training of state space models Attention Is All You Need

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-10T22:50:28.794937Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:50:28.794937Z digest=sha256:99360a41ea2950ef14ca5ba962afacb365135242064a5501f4adf1f553da438c

Observation 18190afd-0c24-4dfd-b0e1-9a504fd379e8 · outbound

This paper cites Rellermeyer.

Adjoint sharding for very long context training of state space models Rellermeyer

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-10T22:50:28.799594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:50:28.799594Z digest=sha256:d88df6762fb5fb98f863baa87b462f38afea38cf91b4ead566569b9ef677bb1f

Observation ba32e0c4-e6db-46c1-9f36-5966078f0baf · outbound

This paper cites An Empirical Study of Mamba-based Language Models.

Adjoint sharding for very long context training of state space models An Empirical Study of Mamba-based Language Models

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-10T22:50:28.804313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:50:28.804313Z digest=sha256:247be3f4022cd254263148c92d456eec005b7bd0af8f41b97c447fdfacd2684c

Observation f0b74476-409f-4245-98ff-b089b3a59bd1 · outbound

This paper cites State-space Models with Layer-wise Nonlinearity are Universal Approximators with Exponential Decaying Memory.

Adjoint sharding for very long context training of state space models State-space Models with Layer-wise Nonlinearity are Universal Approximators with Exponential Decaying Memory

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-10T22:50:28.809266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:50:28.809266Z digest=sha256:dc5bfc10407c8cd5efb7ce55542e939c7fbdb252f5bc9393a1a347d5f12e47b5

Observation cf45a7c3-f255-4093-b336-27e0016e5151 · outbound

This paper cites an unresolved cited work.

Adjoint sharding for very long context training of state space models Unresolved cited work

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-10T22:50:28.814054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:50:28.814054Z digest=sha256:d750274054e0dc09a38acaeac55f200c9c5bdd7402f9fe28d0337a174df448c6

Observation bc78160b-148b-433f-b4a6-75ee14b43cba · outbound

This paper cites Efficient Streaming Language Models with Attention Sinks.

Adjoint sharding for very long context training of state space models Efficient Streaming Language Models with Attention Sinks

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-10T22:50:28.818518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:50:28.818518Z digest=sha256:9ce09cff05bb9c88c5142c2d8b6d502d658ef00e42d952981abe0660f65e7ad0

Observation b91f716d-7110-4e49-8df6-3f18f13b72f3 · outbound

This paper cites Characteristic Neural Ordinary Differential Equations.

Adjoint sharding for very long context training of state space models Characteristic Neural Ordinary Differential Equations

Reference 66

Resolution
verified exact
local_arxiv, observed 2026-08-10T22:50:28.951016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T22:50:28.822823Z digest=sha256:7709577d84f354906ff6f18be6ee32e9bb87cf986229e4039df31308d5b82b09

Observation b399dfae-e754-4ec6-9c80-17ff7042c01a · outbound

This paper cites Focal self- attention for local-global interactions in vision transformers, 2021.

Adjoint sharding for very long context training of state space models Focal self- attention for local-global interactions in vision transformers, 2021

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:50:29.703945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T22:50:28.827158Z digest=sha256:4faf927cfd0dfa38b3cd3ffac6c4f0afd31a25c7bcdc60ca51eff2dd3e954e40

Observation 422cebae-62f4-4436-84fe-5189484914e8 · outbound

This paper cites Long Context Compression with Activation Beacon.

Adjoint sharding for very long context training of state space models Long Context Compression with Activation Beacon

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-10T22:50:28.831365Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:50:28.831365Z digest=sha256:3770dc5a1f4f7ca4453e8b4beaa5b9292cff36d723c51bd0a06d2ecddca8541f

Observation b940b398-90de-4f1e-8780-c9364d5ff600 · outbound

This paper cites PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel.

Adjoint sharding for very long context training of state space models PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-10T22:50:28.835675Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:50:28.835675Z digest=sha256:004ee701912d05c1edacea7debfbca2f0a6af729a859d0f2040edd64a83cba2a

Observation 51b3fa41-2ff7-4646-b78c-2eb3f411c157 · outbound

This paper cites On the difficulty of training Recurrent Neural Networks.

Adjoint sharding for very long context training of state space models On the difficulty of training Recurrent Neural Networks

Reference 2013

Resolution
unresolved
no resolver link, observed 2026-08-10T22:50:28.727533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:50:28.727533Z digest=sha256:6f75459daf76809c714254160581ff8de20266e4f80c96889c34c95e25ce1939

Observation 6517988b-e2bb-4e52-9ceb-6487f6dc82f1 · outbound

This paper cites Ring Attention with Blockwise Transformers for Near-Infinite Context.

Adjoint sharding for very long context training of state space models Ring Attention with Blockwise Transformers for Near-Infinite Context

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-10T22:50:28.683339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:50:28.683339Z digest=sha256:278e36e93e1c93831d58e197b38de12c00caa2a147a06107ef7b2592894dff98

Observation 29b6cc77-4fc1-4b37-974e-fff6fe7095ba · outbound

This paper cites InternLM2 Technical Report.

Adjoint sharding for very long context training of state space models InternLM2 Technical Report

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-10T22:50:28.524381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:50:28.524381Z digest=sha256:6530bd21cb451b6e06ee714242e87d62b1786fea59bb06303469f3c303bc4442

Pith citing papers

No inbound Pith citation observations are available.