Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-04T23:14:12.487890Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 57 of 57 outbound references and 0 inbound Pith citation observations for arXiv:2608.01672.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-04T23:14:12.487890Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
57 of 57 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 25ca538f-d972-44a1-beb3-6a536ead5d5b · outbound
Learning What to Remember: Test-Time Training via Context Distillation SmolLM2: When Smol Goes Big -- Data-Centric Training of a Small Language Model
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6034f659-c302-4f72-9d5c-6be2765af8ab · outbound
Learning What to Remember: Test-Time Training via Context Distillation Using fast weights to attend to the recent past.Advances in neural information processing systems, 29, 2016
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4a8c3380-a083-4314-8a25-23d679cf89af · outbound
Learning What to Remember: Test-Time Training via Context Distillation An Empirical Evaluation of Generic Convolutional and Recurrent Networks for Sequence Modeling
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7fb360cc-2a80-4f60-a6ad-6f8c3679c091 · outbound
Learning What to Remember: Test-Time Training via Context Distillation Titans: Learning to Memorize at Test Time
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c15e55c8-6b23-48cf-96b8-62549536d869 · outbound
Learning What to Remember: Test-Time Training via Context Distillation Lessons from the Trenches on Reproducible Evaluation of Language Models
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e0926956-1441-459c-9c26-2b041cf3cdf9 · outbound
Learning What to Remember: Test-Time Training via Context Distillation KV-Distill: Nearly Lossless Learnable Context Compression for LLMs
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 34ad7e52-2280-443b-aab7-5fe5745e20c7 · outbound
Learning What to Remember: Test-Time Training via Context Distillation Learning to compress prompt in natural language formats
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a5cf0b1b-9438-46ae-b6ed-0992a4b44ffb · outbound
Learning What to Remember: Test-Time Training via Context Distillation Empirical Evaluation of Gated Recurrent Neural Networks on Sequence Modeling
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 80705b35-3476-4122-acde-bde493fc613d · outbound
Learning What to Remember: Test-Time Training via Context Distillation Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e27ad568-460b-442e-834f-a93ff175d8ad · outbound
Learning What to Remember: Test-Time Training via Context Distillation Finch: Prompt-guided key-value cache compression for large language models.Transactions of the Association for Computational Linguistics, 12: 1517–1532, 2024
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation b5f6d980-f486-4050-9bff-b202d9aeeab7 · outbound
Learning What to Remember: Test-Time Training via Context Distillation FlashAttention-2: Faster attention with better parallelism and work partitioning
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 397eb15a-02c7-4554-994a-8acdc000f1a3 · outbound
Learning What to Remember: Test-Time Training via Context Distillation Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0d3d0072-0426-4262-8add-27db456a8b19 · outbound
Learning What to Remember: Test-Time Training via Context Distillation Flashattention: Fast and memory-efficient exact attention with io-awareness.Advances in neural information processing systems, 35:16344–16359, 2022
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6893cd2a-08e9-45ba-9d42-0f4d153f44cc · outbound
Learning What to Remember: Test-Time Training via Context Distillation Cartridges: Lightweight and general-purpose long context representations via self-study
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ab97a2d9-9dc7-428e-8d4e-58ab5386247c · outbound
Learning What to Remember: Test-Time Training via Context Distillation In-place test-time training
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 4db8bf9e-d62c-41f2-81bb-8cbed9025657 · outbound
Learning What to Remember: Test-Time Training via Context Distillation Hungry Hungry Hippos: Towards Language Modeling with State Space Models
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8a5ee083-a876-497a-98c1-0df6f402a9a2 · outbound
Learning What to Remember: Test-Time Training via Context Distillation The Pile: An 800GB Dataset of Diverse Text for Language Modeling
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3018c465-52a7-4b15-a788-52b17a2974f7 · outbound
Learning What to Remember: Test-Time Training via Context Distillation How to train long-context language models (effectively)
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e8cb5d61-4e89-454f-8152-0cd605d7d62e · outbound
Learning What to Remember: Test-Time Training via Context Distillation Model Tells You What to Discard: Adaptive KV Cache Compression for LLMs
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6622a355-149e-4d57-98bc-2a62bfb730e7 · outbound
Learning What to Remember: Test-Time Training via Context Distillation Mamba: Linear-Time Sequence Modeling with Selective State Spaces
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bd5b79d2-829d-4207-a2d3-d610b49eed4c · outbound
Learning What to Remember: Test-Time Training via Context Distillation Efficiently Modeling Long Sequences with Structured State Spaces
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f4a080cd-867c-445a-8964-9940e2d65195 · outbound
Learning What to Remember: Test-Time Training via Context Distillation Log-linear attention.arXiv preprint arXiv:2506.04761, 2025
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7ab99549-c75e-45c0-856a-64fdb2f73595 · outbound
Learning What to Remember: Test-Time Training via Context Distillation Measuring Massive Multitask Language Understanding
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 919ae80d-7d08-4bf1-bdd6-7359d9d19f49 · outbound
Learning What to Remember: Test-Time Training via Context Distillation Long short-term memory.Neural computation, 9(8): 1735–1780, 1997
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 226c7399-347e-414d-8446-1d605e707c2f · outbound
Learning What to Remember: Test-Time Training via Context Distillation RULER: What's the Real Context Size of Your Long-Context Language Models?
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 64a8102d-d7e0-4921-9932-e63d85c1ac63 · outbound
Learning What to Remember: Test-Time Training via Context Distillation MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 47553360-1a86-4eaa-980a-4375973611b4 · outbound
Learning What to Remember: Test-Time Training via Context Distillation Characterizing Prompt Compression Methods for Long Context Inference
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0a3b789e-b136-4bbc-8b7d-9dd3a816c0bb · outbound
Learning What to Remember: Test-Time Training via Context Distillation Llmlingua: Compress- ing prompts for accelerated inference of large language models
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cc9b1233-617f-4401-89d8-72104f18e534 · outbound
Learning What to Remember: Test-Time Training via Context Distillation Babilong: Testing the limits of llms with long context reasoning-in-a-haystack
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a79c02d6-a830-47a0-b543-7f2bf52a67c2 · outbound
Learning What to Remember: Test-Time Training via Context Distillation The power of scale for parameter-efficient prompt tuning
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 13d8cc07-5037-4a8c-b89e-312003fc2828 · outbound
Learning What to Remember: Test-Time Training via Context Distillation Prefix-tuning: Optimizing continuous prompts for generation
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation adbd24ab-f684-4a90-bf00-1f1f4e15b14d · outbound
Learning What to Remember: Test-Time Training via Context Distillation Snapkv: Llm knows what you are looking for before generation.Advances in Neural Information Processing Systems, 37:22947–22970, 2024
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 816d684a-addd-49aa-b6bc-8a610d2e6e2c · outbound
Learning What to Remember: Test-Time Training via Context Distillation Leave No Context Behind: Efficient Infinite Context Transformers with Infini-attention
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aae81bf1-aa7d-47b8-80a4-fe9c14319503 · outbound
Learning What to Remember: Test-Time Training via Context Distillation Transformers are multi-state rnns
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 5d86c09b-944d-4fd0-9f10-4e067ac893eb · outbound
Learning What to Remember: Test-Time Training via Context Distillation Rwkv: Reinventing rnns for the transformer era
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 40fd6dab-2f54-4d83-8ed8-b9ffee9cfd62 · outbound
Learning What to Remember: Test-Time Training via Context Distillation Mechanistic Design and Scaling of Hybrid Architectures
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 448891ea-e8db-4747-9d84-708f7e159209 · outbound
Learning What to Remember: Test-Time Training via Context Distillation Improving language understanding by generative pre-training
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cf6f7410-8dd1-49e4-add4-177088d2587b · outbound
Learning What to Remember: Test-Time Training via Context Distillation Linear transformers are secretly fast weight programmers
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a209b04f-5702-4b86-bd83-513447cfdfd9 · outbound
Learning What to Remember: Test-Time Training via Context Distillation FlashAttention-3: Fast and accurate attention with asynchrony and low-precision
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation d702b95c-08f7-4a5f-ae0c-2adbb5ce5d7b · outbound
Learning What to Remember: Test-Time Training via Context Distillation Learning by Distilling Context
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ef0a8900-bbce-4f3e-81af-7cc143090630 · outbound
Learning What to Remember: Test-Time Training via Context Distillation Roformer: Enhanced transformer with rotary position embedding.Neurocomputing, 568:127063, 2024
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 18a50eff-afb5-4a4c-87b0-6a62d78006cc · outbound
Learning What to Remember: Test-Time Training via Context Distillation Test-time training with self-supervision for generalization under distribution shifts
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d84e59b2-544a-4af1-832a-fc036152c07c · outbound
Learning What to Remember: Test-Time Training via Context Distillation Learning to (Learn at Test Time): RNNs with Expressive Hidden States
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d8f1b53f-9c2b-4fee-a59c-f391f420b1b1 · outbound
Learning What to Remember: Test-Time Training via Context Distillation End-to-end test-time training for long context.arXiv preprint arXiv:2512.23675, 2025
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 21e81452-944f-4526-99ff-b77dd6ce18f2 · outbound
Learning What to Remember: Test-Time Training via Context Distillation Kimi Linear: An Expressive, Efficient Attention Architecture
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 57870a3d-587b-49cf-8ef6-e46b09f61007 · outbound
Learning What to Remember: Test-Time Training via Context Distillation Long data collections database, 2024
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4315f292-e6c7-41fe-8d5d-f0ea54d47e33 · outbound
Learning What to Remember: Test-Time Training via Context Distillation Attention is all you need.Advances in neural information processing systems, 30, 2017
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6c056be1-e655-4eef-b1f3-07ef8161b56e · outbound
Learning What to Remember: Test-Time Training via Context Distillation Rattention: Towards the minimal sliding window size in local-global attention models.arXiv preprint arXiv:2506.15545, 2025
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 90465926-40c6-49b7-91bd-65939b77e029 · outbound
Learning What to Remember: Test-Time Training via Context Distillation Fantastic Pretraining Optimizers and Where to Find Them
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d2f92442-bf3a-438a-bc24-b9d17918de0a · outbound
Learning What to Remember: Test-Time Training via Context Distillation Gated Linear Attention Transformers with Hardware-Efficient Training
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 43063023-1f75-43e9-a0ab-6981d13021b3 · outbound
Learning What to Remember: Test-Time Training via Context Distillation Gated Delta Networks: Improving Mamba2 with Delta Rule
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 610f9608-bdd7-4f16-a98b-796fb62e02d6 · outbound
Learning What to Remember: Test-Time Training via Context Distillation Parallelizing linear transformers with the delta rule over sequence length.Advances in neural information processing systems, 37:115491–115522, 2024
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 171207e8-4b27-4a4f-8e1c-4e1c333b3367 · outbound
Learning What to Remember: Test-Time Training via Context Distillation Native sparse attention: Hardware-aligned and natively trainable sparse attention
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9aed4c75-ec93-48a7-af42-4a9a1aff78fc · outbound
Learning What to Remember: Test-Time Training via Context Distillation Hellaswag: Can a machine really finish your sentence? InProceedings of the 57th annual meeting of the association for computational linguistics, pages 4791–4800, 2019
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 941c35c7-1044-4412-8a89-17dec2777047 · outbound
Learning What to Remember: Test-Time Training via Context Distillation Test-Time Training Done Right
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 76baa423-df23-4828-8c47-e691ab3d19a8 · outbound
Learning What to Remember: Test-Time Training via Context Distillation up to 1.3×
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 129dc4f6-bd6b-4602-9d85-ca55dde4b2bd · outbound
Learning What to Remember: Test-Time Training via Context Distillation Unresolved cited work
Reference 2026
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
No inbound Pith citation observations are available.