Pith. sign in

Paper Citation Record · LEDGER

Adjoint sharding for very long context training of state space models

As of 22 August 2026, this Paper Citation Record lists 67 of 67 outbound references and 0 inbound Pith citation observations for arXiv:2501.00692.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.00692 v1

Coverage vector

measured 67 of 67 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T22:50:28.835675Z

measured 67 of 67 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

67 of 67 outbound references displayed

  • verified exact3
  • verified fuzzy15
  • unresolved49
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e99b3f86-32a1-4820-8131-f8ff9c1e4fd6 · outbound

This paper cites BlackMamba: Mixture of Experts for State-Space Models.

Adjoint sharding for very long context training of state space models BlackMamba: Mixture of Experts for State-Space Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T22:50:28.484069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:50:28.484069Z digest=sha256:4edcb5e62b62e5004a6ce6416a2ea1a77400d5f5bc2b453354be04495573ae89

Observation bf8d5a77-5e96-4a43-ab94-2b34547d5a1f · outbound

This paper cites Fast Jacobian-Vector Product for Deep Networks.

Adjoint sharding for very long context training of state space models Fast Jacobian-Vector Product for Deep Networks

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T22:50:28.491015Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:50:28.491015Z digest=sha256:a5b74d1940f8c0780ea1133fd782dcbc62bdc72082d486f8c263b1e5401a5229

Observation ae67fdfe-2068-4f89-b70c-69ecafbfff53 · outbound

This paper cites Automatic differentiation in machine learning: a survey.

Adjoint sharding for very long context training of state space models Automatic differentiation in machine learning: a survey

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T22:50:28.502732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:50:28.502732Z digest=sha256:628620535b86deb390c26d422c3dc90fbcec3c65d58e25a2057f669753a8aab1

Observation f895bba9-0913-44b8-9667-381d1a53c7a4 · outbound

This paper cites xLSTM: Extended Long Short-Term Memory.

Adjoint sharding for very long context training of state space models xLSTM: Extended Long Short-Term Memory

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T22:50:28.508274Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:50:28.508274Z digest=sha256:c1ae1906cde449e5255c1b0153047b3559510d2e82b8fa3c316d1e066f5351c1

Observation 85553efb-415d-49f1-b461-85f0f64bcfff · outbound

This paper cites Longformer: The Long-Document Transformer.

Adjoint sharding for very long context training of state space models Longformer: The Long-Document Transformer

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T22:50:28.514062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:50:28.514062Z digest=sha256:966dc5577eea8d995405f83d39f284235b98f2381b9d800aa4e7da2ef549d322

Observation 070ce06c-5a9b-4957-a7d4-58d66004f45b · outbound

This paper cites Internlm2 technical report,.

Adjoint sharding for very long context training of state space models Internlm2 technical report,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:50:29.968339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T22:50:28.519633Z digest=sha256:39fb1c764e0a6f1a5da851c287777212ae54939da7d8f2db57b365649e8f3a03

Observation 0baf7150-8e83-4678-818c-fd34b341cff6 · outbound

This paper cites Adjoint sensitivity analysis for differential-algebraic equations: algorithms and software.

Adjoint sharding for very long context training of state space models Adjoint sensitivity analysis for differential-algebraic equations: algorithms and software

Reference 8

Resolution
verified exact
doi, observed 2026-08-10T22:50:28.898216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T22:50:28.529339Z digest=sha256:6052a78503712419c79280f5e0ab95aece0e948d43c965aa7d6323535db72d7d

Observation aecbb979-4b57-4028-99f2-0f5bb7231634 · outbound

This paper cites Neural Ordinary Differential Equations.

Adjoint sharding for very long context training of state space models Neural Ordinary Differential Equations

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T22:50:28.535476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:50:28.535476Z digest=sha256:85d6805023a0fa30ba7e8b32055367414a9342eab37e6529981d602f363a39c3

Observation edbbc815-26f6-4d04-a078-fd8f06a5d57f · outbound

This paper cites Extending Context Window of Large Language Models via Positional Interpolation.

Adjoint sharding for very long context training of state space models Extending Context Window of Large Language Models via Positional Interpolation

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T22:50:28.540632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:50:28.540632Z digest=sha256:a8162f8cf4085552427771f44252830185b5483b63c3d24af21ea5f5667f52a2

Observation 624eaeba-c5b0-4f3a-8862-1ecbeae640df · outbound

This paper cites LongLoRA: Efficient Fine-tuning of Long-Context Large Language Models.

Adjoint sharding for very long context training of state space models LongLoRA: Efficient Fine-tuning of Long-Context Large Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T22:50:28.546514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:50:28.546514Z digest=sha256:c14e58030150b95c5c222758195f4dd8ada8a2900172d7b37f0cd13a904fc600

Observation c5668d3e-01c5-47dd-b180-8403a30e76b6 · outbound

This paper cites The Backpropagation algorithm for a math student.

Adjoint sharding for very long context training of state space models The Backpropagation algorithm for a math student

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-08-10T22:50:29.547231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T22:50:28.551587Z digest=sha256:808dadfbba554d30f7a6b14896c9296d7ee7c7d22cca05cfd7702b9ca8aa01b1

Observation ef15a127-c9d4-4244-b772-366c06a7db7d · outbound

This paper cites FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning.

Adjoint sharding for very long context training of state space models FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T22:50:28.557753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:50:28.557753Z digest=sha256:77d433b8a527bcfce94e4ae5a06b04e0c7780b845e6afd23351c542283eb4bdd

Observation 1094675c-c2f8-4d8b-a53c-f348e20cbd90 · outbound

This paper cites Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality.

Adjoint sharding for very long context training of state space models Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T22:50:28.568578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:50:28.568578Z digest=sha256:7585699b5a2f3cad96f1a2075f9e33702ea9482f7cc9c56868186006993de2ce

Observation 8a49de87-00d2-443a-b115-f71f4beb563f · outbound

This paper cites FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness.

Adjoint sharding for very long context training of state space models FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T22:50:28.573058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:50:28.573058Z digest=sha256:563fd5d6b14fe55c2337537cddbc8066c322734d3fcb586adca1781b91599e66

Observation c36509f4-1270-44cc-af80-1555ab80fb33 · outbound

This paper cites an unresolved cited work.

Adjoint sharding for very long context training of state space models Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:50:29.953519Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T22:50:28.579575Z digest=sha256:8d028c4a115ede0818124d20caf27c628ae60154636ff257b4b7a58dc6179008

Observation 4277793a-c7d5-4b8b-a87b-e385df3b5f10 · outbound

This paper cites LongRoPE: Extending LLM Context Window Beyond 2 Million Tokens.

Adjoint sharding for very long context training of state space models LongRoPE: Extending LLM Context Window Beyond 2 Million Tokens

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T22:50:28.589583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:50:28.589583Z digest=sha256:72cc53f49972d0f42722363c9f5d33da5c868e32b91c6dbfdd1db998110a9e6e

Observation 4ef388c6-e212-4c3e-961a-2c65891d49d6 · outbound

This paper cites Augmented Neural ODEs.

Adjoint sharding for very long context training of state space models Augmented Neural ODEs

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T22:50:28.593837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:50:28.593837Z digest=sha256:f5c6e3523807ae8c92104f159876205528c99e48e082bb7840812d334a4f8abf

Observation 6b230c12-468d-4b42-a2f2-a1733d9d5ff2 · outbound

This paper cites Fu, Tri Dao, Khaled K.

Adjoint sharding for very long context training of state space models Fu, Tri Dao, Khaled K

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:50:29.937435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T22:50:28.598096Z digest=sha256:c5a912a93cd53dbb96157364287323930a3e03aaf84d715bb2a1a2e0b462ad1e

Observation 42dda1f6-ddd6-4dd3-83db-7ee9bf0e7251 · outbound

This paper cites Mamba: Linear-Time Sequence Modeling with Selective State Spaces.

Adjoint sharding for very long context training of state space models Mamba: Linear-Time Sequence Modeling with Selective State Spaces

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T22:50:28.606419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:50:28.606419Z digest=sha256:2031b93dee907123a32dba0aedd4c271851f288515e9340ea46342b1d9605b95

Observation 283b746e-b320-4c1b-bf46-a6f3ebf6f618 · outbound

This paper cites Combining Recurrent, Convolutional, and Continuous-time Models with Linear State-Space Layers.

Adjoint sharding for very long context training of state space models Combining Recurrent, Convolutional, and Continuous-time Models with Linear State-Space Layers

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T22:50:28.610398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:50:28.610398Z digest=sha256:0a7196d3a144c68eec1b8688c0f9009e67947a210ac71d05c9ae46f6d65e6f9b

Observation dd6a2829-5bc1-4f5a-897c-8eee0cf5a9c5 · outbound

This paper cites How to Train Your HiPPO: State Space Models with Generalized Orthogonal Basis Projections.

Adjoint sharding for very long context training of state space models How to Train Your HiPPO: State Space Models with Generalized Orthogonal Basis Projections

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-10T22:50:28.614831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:50:28.614831Z digest=sha256:e2aa0b073bc0b785c4a6541f2c7034bef0afa14e8a8f43d807ecfccf98092d08

Observation 43a4e018-a215-4ca5-a18a-c0572863726e · outbound

This paper cites Attention mechanisms in computer vision: A survey.

Adjoint sharding for very long context training of state space models Attention mechanisms in computer vision: A survey

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:50:29.921783Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T22:50:28.619200Z digest=sha256:a2f35232d6d96a6f25a86a252225c1285cadc7912d6a5e054d52ccdb76dbdf3c

Observation 22e3801c-b225-49bc-8bc2-f8d06774a95d · outbound

This paper cites Simplifying and Understanding State Space Models with Diagonal Linear RNNs.

Adjoint sharding for very long context training of state space models Simplifying and Understanding State Space Models with Diagonal Linear RNNs

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-10T22:50:28.623895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:50:28.623895Z digest=sha256:17d50c04100bea64f2212732fa7e793c7f2e233a88a4d3ecb270c56dee28b026

Observation 2bfc6337-d651-443b-aba5-7c2030a42c8b · outbound

This paper cites Deep Residual Learning for Image Recognition.

Adjoint sharding for very long context training of state space models Deep Residual Learning for Image Recognition

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-10T22:50:28.629714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:50:28.629714Z digest=sha256:1a95bf9d8c4f99da48b715f2923e510c25e228eaf1ff3a0add951a92fc48d07b

Observation 1d7e40de-c191-407a-bcb0-97e1b28d4108 · outbound

This paper cites Deep residual learning for image recognition.

Adjoint sharding for very long context training of state space models Deep residual learning for image recognition

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-10T22:50:28.634649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:50:28.634649Z digest=sha256:8575ca05e2791871692f877623d64ef4a5cf95a54fbe2a115c538285dc944a48

Observation 4e9c8b16-e551-43a8-8a21-be6f31b3a379 · outbound

This paper cites Optimal checkpointing for heterogeneous chains: how to train deep neural networks with limited memory, 2019.

Adjoint sharding for very long context training of state space models Optimal checkpointing for heterogeneous chains: how to train deep neural networks with limited memory, 2019

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-10T22:50:28.639358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:50:28.639358Z digest=sha256:5fa2a483e338564091acca690e0bf5c35609b3923ba634f6ceea2c393368e952

Observation 6cf268e0-d752-4456-8bdd-b6f75035f474 · outbound

This paper cites A tutorial on training recurrent neural networks , covering bppt , rtrl , ekf and the ” echo state network ” approach - semantic scholar.

Adjoint sharding for very long context training of state space models A tutorial on training recurrent neural networks , covering bppt , rtrl , ekf and the ” echo state network ” approach - semantic scholar

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:50:29.896896Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T22:50:28.644010Z digest=sha256:c6aca3f29c0f8aeed65988bb745c40bbc90a1f46fe90f060161df6b1ba3efc7b

Observation 1ab5ffaa-01d8-464f-bbe5-57e8008714ae · outbound

This paper cites Adjoint methods and sensitivity analysis for recurrence, 01 2007.

Adjoint sharding for very long context training of state space models Adjoint methods and sensitivity analysis for recurrence, 01 2007

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:50:29.880912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T22:50:28.648655Z digest=sha256:d7fe02c20b4d99e939c3198b13b39dd4f2eebcf595e780e4d4cbabc894b3c9a9

Observation 5bcfbff1-cf11-448a-81fb-c1315c223a02 · outbound

This paper cites Linear dynamical systems as a core computational primitive.

Adjoint sharding for very long context training of state space models Linear dynamical systems as a core computational primitive

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:50:29.865266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T22:50:28.653859Z digest=sha256:4ea11708582a9c672e331b11689471ffbbad7fed3765e6048bc70ee9adc6544b

Observation 55b31091-3be3-4176-ad29-9abff81abd8c · outbound

This paper cites Segment anything.

Adjoint sharding for very long context training of state space models Segment anything

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-10T22:50:28.658687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:50:28.658687Z digest=sha256:e7d7b688af311a2fe6f21279d2448bf23c64bbe291daf1ab8c78f7ecfa656e74

Observation b45875e8-6c89-4120-814a-ba7e78919f9e · outbound

This paper cites Gonzalez, Ion Stoica, Xuezhe Ma, and Hao Zhang.

Adjoint sharding for very long context training of state space models Gonzalez, Ion Stoica, Xuezhe Ma, and Hao Zhang

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:50:29.838923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T22:50:28.663400Z digest=sha256:44d747c8ae6f82a97216d6279b22f37fd1b215e1756565fb063df5e47f3e041a

Observation 94d34a26-4942-442e-8c0f-5424ce90dda9 · outbound

This paper cites Long-context LLMs Struggle with Long In-context Learning.

Adjoint sharding for very long context training of state space models Long-context LLMs Struggle with Long In-context Learning

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-10T22:50:28.668601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:50:28.668601Z digest=sha256:dc7fe0502c965e602b8811776696ba4dc4ff80f3f7e3e106a7d666ba083c8246

Observation e761159d-6db4-4d7b-8acb-01dfb22f5555 · outbound

This paper cites Jamba: A Hybrid Transformer-Mamba Language Model.

Adjoint sharding for very long context training of state space models Jamba: A Hybrid Transformer-Mamba Language Model

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-10T22:50:28.673258Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:50:28.673258Z digest=sha256:7802e11cda84fb441179550ced31da76573421292ddb5506d692aaa07525df65

Observation d8cefb4f-686a-4888-9fbc-bd8788478359 · outbound

This paper cites Ring attention with blockwise transformers for near-infinite context,.

Adjoint sharding for very long context training of state space models Ring attention with blockwise transformers for near-infinite context,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:50:29.822170Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T22:50:28.678219Z digest=sha256:6aa1a1ecea92355d400c145e0f6ac111e2f31df64b5b8547b1e442237f6559d0

Observation 3d1ec21e-95f6-47a5-9577-58695bad7677 · outbound

This paper cites World Model on Million-Length Video And Language With Blockwise RingAttention.

Adjoint sharding for very long context training of state space models World Model on Million-Length Video And Language With Blockwise RingAttention

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-10T22:50:28.688390Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:50:28.688390Z digest=sha256:2b598be737d57b985ba50dfd41a17b3e3d9e617e4ec0e034c16b0b6786f3cf4d

Observation 009032f5-bff7-4f07-85d2-7600c62db912 · outbound

This paper cites The Llama 3 Herd of Models.

Adjoint sharding for very long context training of state space models The Llama 3 Herd of Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-10T22:50:28.693596Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:50:28.693596Z digest=sha256:fe339bef0dd835193b64ca53c39db80d3d597ad915880855689c62579515052d

Observation c83296ba-ed55-46b3-9b3f-d7be246c9d41 · outbound

This paper cites Mixed Precision Training.

Adjoint sharding for very long context training of state space models Mixed Precision Training

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-10T22:50:28.698249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:50:28.698249Z digest=sha256:154ae8b748fe8fb645eaa179574a2db4e973d76e0bbc5415e5b05b43f038487d

Observation 0f8e9868-8353-4e7d-9dd9-99b5e5303970 · outbound

This paper cites Fast Finite Width Neural Tangent Kernel.

Adjoint sharding for very long context training of state space models Fast Finite Width Neural Tangent Kernel

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-10T22:50:28.703665Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:50:28.703665Z digest=sha256:e73f5f5efce10e8e068846281a092c3264092ad2099eb66d78c7399e2f4bb24d

Observation c8b882df-8c3f-48d7-a072-4ebc5c82c2db · outbound

This paper cites Matrix multiplication background user’s guide, 2024.

Adjoint sharding for very long context training of state space models Matrix multiplication background user’s guide, 2024

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:50:29.806652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T22:50:28.708445Z digest=sha256:1b8003e0db1105426477aa536a638fd625477fe42a11c7b361e760abf2fb49de

Observation a9f0488d-6a2a-473b-854b-f8504d364620 · outbound

This paper cites GPT-4 Technical Report.

Adjoint sharding for very long context training of state space models GPT-4 Technical Report

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-10T22:50:28.713364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:50:28.713364Z digest=sha256:927191abc7a2651cecf2c3043b8c3e6d91faa9a95d0c69f6449923dabd94e545

Observation 6bb42725-7c29-4c50-85c3-f8604c0edca3 · outbound

This paper cites Resurrecting recurrent neural networks for long sequences, 2023.

Adjoint sharding for very long context training of state space models Resurrecting recurrent neural networks for long sequences, 2023

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:50:29.790576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T22:50:28.717938Z digest=sha256:f024b20b47999742176d6d2bdaad76e2682397aff176d306ff7ef39405dc3b2e

Observation 231e8049-bb44-49ad-a252-7d3d0182f369 · outbound

This paper cites On the difficulty of training recurrent neural networks,.

Adjoint sharding for very long context training of state space models On the difficulty of training recurrent neural networks,

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:50:29.774729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T22:50:28.722494Z digest=sha256:c5476968610ab77aca594b9d73bef1be18a9ca953d55554d9d3cdc4e3e13f9eb

Observation 7d849226-e6e1-4061-81f0-efbe283bd521 · outbound

This paper cites PyTorch: An Imperative Style, High-Performance Deep Learning Library.

Adjoint sharding for very long context training of state space models PyTorch: An Imperative Style, High-Performance Deep Learning Library

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-10T22:50:28.733647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:50:28.733647Z digest=sha256:13747b55f7b95e2b32ceb8a8b5951d17938dfaffac1b1a98f756e71eae37783c

Observation b15134ab-d40e-45b2-ae64-92ec8b38bb1d · outbound

This paper cites Scalable diffusion models with transformers.

Adjoint sharding for very long context training of state space models Scalable diffusion models with transformers

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-10T22:50:28.738118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:50:28.738118Z digest=sha256:17ca76130f557a02cb8a4e8406d85c578f1080b1531a71dc7b5b1f6b70e2ab3f

Observation c1e48a4e-0276-419f-a9fe-05bf2653e47c · outbound

This paper cites RWKV: Reinventing RNNs for the Transformer Era.

Adjoint sharding for very long context training of state space models RWKV: Reinventing RNNs for the Transformer Era

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-10T22:50:28.742462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:50:28.742462Z digest=sha256:159064395369d053246da86eab017344afe240458d543a7e388fe7f6c52b324a

Observation cbeac19b-314b-42a3-9dcf-3144c5f2fe51 · outbound

This paper cites YaRN: Efficient Context Window Extension of Large Language Models.

Adjoint sharding for very long context training of state space models YaRN: Efficient Context Window Extension of Large Language Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-10T22:50:28.746548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:50:28.746548Z digest=sha256:f5696e099e6d1e49f6ba45d2a478bbbdde90b7f8357c5a4d218066b3bf1eae8e

Observation f511de94-55b9-4601-b157-f583ccb982dc · outbound

This paper cites MoE-Mamba: Efficient Selective State Space Models with Mixture of Experts.

Adjoint sharding for very long context training of state space models MoE-Mamba: Efficient Selective State Space Models with Mixture of Experts

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-10T22:50:28.750951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:50:28.750951Z digest=sha256:e30569ea0467e8f59e824d9c50eecf04f7834d68ac7a12088ea5128208d20a4c

Observation 45c9323e-e783-49d2-89b4-90629e0a17f6 · outbound

This paper cites ZeRO: Memory Optimizations Toward Training Trillion Parameter Models.

Adjoint sharding for very long context training of state space models ZeRO: Memory Optimizations Toward Training Trillion Parameter Models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-10T22:50:28.760927Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:50:28.760927Z digest=sha256:2890d288574cdb0c8d34c0b68b903345126a4776b22ad90829fc377bc369f7bd

Observation 87cd75ca-9024-4b70-9036-d4edcb65e7c0 · outbound

This paper cites ZeRO-Offload: Democratizing Billion-Scale Model Training.

Adjoint sharding for very long context training of state space models ZeRO-Offload: Democratizing Billion-Scale Model Training

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-10T22:50:28.765359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:50:28.765359Z digest=sha256:ab2d4151a8e87270aa5db469ed5202508319be9767071cb535773df4fb5af920

Observation a6a2acd4-bcb0-4569-8169-b00b439bd6fd · outbound

This paper cites Flashattention-3: Fast and accurate attention with asynchrony and low-precision, 2024.

Adjoint sharding for very long context training of state space models Flashattention-3: Fast and accurate attention with asynchrony and low-precision, 2024

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:50:29.749372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T22:50:28.770760Z digest=sha256:5cefd6f660c6d02731df7497f14f927f582d8d3a6ac4cfa06f06f29503464755

Observation c968568b-da2d-4c22-8a6d-6eb31d3dd92b · outbound

This paper cites Low-Memory Neural Network Training: A Technical Report.

Adjoint sharding for very long context training of state space models Low-Memory Neural Network Training: A Technical Report

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-10T22:50:28.775364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:50:28.775364Z digest=sha256:6c51e8fa453798007303cb0c2b084753e209cc1990835fcdb7de6a67242c805a

Observation 6c7a2c8c-4fcf-473d-84be-56c2712d83d7 · outbound

This paper cites Unbiasing Truncated Backpropagation Through Time.

Adjoint sharding for very long context training of state space models Unbiasing Truncated Backpropagation Through Time

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-10T22:50:28.780156Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:50:28.780156Z digest=sha256:66e2fde2da3ba2f939702982035761b8e88e66ee8005459c35ab181425a35403

Observation f2eb26b8-3649-473e-97e5-0ab9ec96c32f · outbound

This paper cites Focused transformer: Contrastive training for context scaling, 2023.

Adjoint sharding for very long context training of state space models Focused transformer: Contrastive training for context scaling, 2023

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:50:29.733806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T22:50:28.785242Z digest=sha256:d24b88a4ab524139b2f826227a97ef0c56e2400d45d3089fdc7cb2c7446967a9

Observation 522d017f-3044-4443-8c6c-126d5db9fe90 · outbound

This paper cites Ntk-aware scaled rope, 2023.

Adjoint sharding for very long context training of state space models Ntk-aware scaled rope, 2023

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:50:29.719374Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T22:50:28.790111Z digest=sha256:524860721a3dc8befbc2805f9aef3f192a425afd4fbf3163b9fdcab91a0a311e

Observation b8ba6447-40e7-4c8f-8362-e8d8c5f2d1d7 · outbound

This paper cites Attention Is All You Need.

Adjoint sharding for very long context training of state space models Attention Is All You Need

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-10T22:50:28.794937Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:50:28.794937Z digest=sha256:db49ddc9f97cd99913e4f04b9a3d5bd5164840486f9cc75bf47b6024b6295a62

Observation 18190afd-0c24-4dfd-b0e1-9a504fd379e8 · outbound

This paper cites Rellermeyer.

Adjoint sharding for very long context training of state space models Rellermeyer

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-10T22:50:28.799594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:50:28.799594Z digest=sha256:367f9a5a7af9d90386436f11ed8a7531af7ff5b9305ce40eb75875b08afc6cd6

Observation ba32e0c4-e6db-46c1-9f36-5966078f0baf · outbound

This paper cites An Empirical Study of Mamba-based Language Models.

Adjoint sharding for very long context training of state space models An Empirical Study of Mamba-based Language Models

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-10T22:50:28.804313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:50:28.804313Z digest=sha256:11ae9b68d5808dc198637b1fac6fcdb6e2bff64ff5862f3ddddeeefb9a4ffa7b

Observation f0b74476-409f-4245-98ff-b089b3a59bd1 · outbound

This paper cites State-space Models with Layer-wise Nonlinearity are Universal Approximators with Exponential Decaying Memory.

Adjoint sharding for very long context training of state space models State-space Models with Layer-wise Nonlinearity are Universal Approximators with Exponential Decaying Memory

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-10T22:50:28.809266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:50:28.809266Z digest=sha256:8d302462c434f6a683452bb0d7f673906846944d75f9d2ecd2009e335d97b60b

Observation cf45a7c3-f255-4093-b336-27e0016e5151 · outbound

This paper cites an unresolved cited work.

Adjoint sharding for very long context training of state space models Unresolved cited work

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-10T22:50:28.814054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:50:28.814054Z digest=sha256:7f1529f15543656daf7ad8fef3dd48d839065c07ed820eeb8174bae9587de508

Observation bc78160b-148b-433f-b4a6-75ee14b43cba · outbound

This paper cites Efficient Streaming Language Models with Attention Sinks.

Adjoint sharding for very long context training of state space models Efficient Streaming Language Models with Attention Sinks

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-10T22:50:28.818518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:50:28.818518Z digest=sha256:0356e86c47d7800ba347914563525973b63433337440ffcee7b6a26295b8fad1

Observation b91f716d-7110-4e49-8df6-3f18f13b72f3 · outbound

This paper cites Characteristic Neural Ordinary Differential Equations.

Adjoint sharding for very long context training of state space models Characteristic Neural Ordinary Differential Equations

Reference 66

Resolution
verified exact
local_arxiv, observed 2026-08-10T22:50:28.951016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T22:50:28.822823Z digest=sha256:9cda9bf831f4d223e30e03e75123099d2704ca91143ab5c86a68a63f6d6784ca

Observation b399dfae-e754-4ec6-9c80-17ff7042c01a · outbound

This paper cites Focal self- attention for local-global interactions in vision transformers, 2021.

Adjoint sharding for very long context training of state space models Focal self- attention for local-global interactions in vision transformers, 2021

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:50:29.703945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T22:50:28.827158Z digest=sha256:92e79c2bc7f2d80bb26bee72abff5cbf7f8e1865b1953cafc956b04c0bcaa15f

Observation 422cebae-62f4-4436-84fe-5189484914e8 · outbound

This paper cites Long Context Compression with Activation Beacon.

Adjoint sharding for very long context training of state space models Long Context Compression with Activation Beacon

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-10T22:50:28.831365Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:50:28.831365Z digest=sha256:aac537d3e17290c0d6624bd8475694b29c58aca47217c4aa4c3e2cd6f311534a

Observation b940b398-90de-4f1e-8780-c9364d5ff600 · outbound

This paper cites PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel.

Adjoint sharding for very long context training of state space models PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-10T22:50:28.835675Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:50:28.835675Z digest=sha256:67858aed2c434f55ca1c7bf89fd9bccddeb3dc5233835ddb684350e240834aa3

Observation 51b3fa41-2ff7-4646-b78c-2eb3f411c157 · outbound

This paper cites On the difficulty of training Recurrent Neural Networks.

Adjoint sharding for very long context training of state space models On the difficulty of training Recurrent Neural Networks

Reference 2013

Resolution
unresolved
no resolver link, observed 2026-08-10T22:50:28.727533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:50:28.727533Z digest=sha256:7f0a8f3cfeb91bc4e072a8226f35345e6c32a8f0713f613167d90124b83c568a

Observation 6517988b-e2bb-4e52-9ceb-6487f6dc82f1 · outbound

This paper cites Ring Attention with Blockwise Transformers for Near-Infinite Context.

Adjoint sharding for very long context training of state space models Ring Attention with Blockwise Transformers for Near-Infinite Context

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-10T22:50:28.683339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:50:28.683339Z digest=sha256:b5bb6ba2f955e2d9b3daec4fa7f408d5387aba834776ab6a21eda91110267a19

Observation 29b6cc77-4fc1-4b37-974e-fff6fe7095ba · outbound

This paper cites InternLM2 Technical Report.

Adjoint sharding for very long context training of state space models InternLM2 Technical Report

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-10T22:50:28.524381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:50:28.524381Z digest=sha256:61043543377b3918af0fb2d587320cd34c6249099da6a5e2a6e4f89d81730ea3

Pith citing papers

No inbound Pith citation observations are available.