Pith. sign in

Paper Citation Record · LEDGER

Learning What to Remember: Test-Time Training via Context Distillation

As of 8 August 2026, this Paper Citation Record lists 57 of 57 outbound references and 0 inbound Pith citation observations for arXiv:2608.01672.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.01672 v1

Coverage vector

measured 57 of 57 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T23:14:12.487890Z

measured 57 of 57 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

57 of 57 outbound references displayed

  • verified exact0
  • verified fuzzy6
  • unresolved50
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 25ca538f-d972-44a1-beb3-6a536ead5d5b · outbound

This paper cites SmolLM2: When Smol Goes Big -- Data-Centric Training of a Small Language Model.

Learning What to Remember: Test-Time Training via Context Distillation SmolLM2: When Smol Goes Big -- Data-Centric Training of a Small Language Model

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-04T23:14:04.620717Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:14:04.620717Z digest=sha256:621153402d0aa4a2b32bbbe6d901d631bc6297c35aafb7430567162d75617c00

Observation 6034f659-c302-4f72-9d5c-6be2765af8ab · outbound

This paper cites Using fast weights to attend to the recent past.Advances in neural information processing systems, 29, 2016.

Learning What to Remember: Test-Time Training via Context Distillation Using fast weights to attend to the recent past.Advances in neural information processing systems, 29, 2016

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-04T23:14:04.756074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:14:04.756074Z digest=sha256:083ad6135d558e95afb0ac0b48bc4d2cb27063415093d582921f9b94f0fe7edf

Observation 4a8c3380-a083-4314-8a25-23d679cf89af · outbound

This paper cites An Empirical Evaluation of Generic Convolutional and Recurrent Networks for Sequence Modeling.

Learning What to Remember: Test-Time Training via Context Distillation An Empirical Evaluation of Generic Convolutional and Recurrent Networks for Sequence Modeling

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-04T23:14:04.918437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:14:04.918437Z digest=sha256:7c7a33ef14cb411883fa2769f30bb7271ab1389dee95d9918d81e523b1b78595

Observation 7fb360cc-2a80-4f60-a6ad-6f8c3679c091 · outbound

This paper cites Titans: Learning to Memorize at Test Time.

Learning What to Remember: Test-Time Training via Context Distillation Titans: Learning to Memorize at Test Time

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-04T23:14:05.016701Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:14:05.016701Z digest=sha256:89a2a58a94d9fc841721126c26831434ecd4e8adcb0922449a1c4fdba9e0a3b9

Observation c15e55c8-6b23-48cf-96b8-62549536d869 · outbound

This paper cites Lessons from the Trenches on Reproducible Evaluation of Language Models.

Learning What to Remember: Test-Time Training via Context Distillation Lessons from the Trenches on Reproducible Evaluation of Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T23:14:05.158429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:14:05.158429Z digest=sha256:871dae7ae157e6eda595417e95285a4feaf52d404eabb962e4c35396031156f8

Observation e0926956-1441-459c-9c26-2b041cf3cdf9 · outbound

This paper cites KV-Distill: Nearly Lossless Learnable Context Compression for LLMs.

Learning What to Remember: Test-Time Training via Context Distillation KV-Distill: Nearly Lossless Learnable Context Compression for LLMs

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T23:14:05.348841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:14:05.348841Z digest=sha256:872764be24e019aaa2a59a8218c536958a6122bf31112b9c364d25e9e8df395d

Observation 34ad7e52-2280-443b-aab7-5fe5745e20c7 · outbound

This paper cites Learning to compress prompt in natural language formats.

Learning What to Remember: Test-Time Training via Context Distillation Learning to compress prompt in natural language formats

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:14:15.205397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-04T23:14:05.477552Z digest=sha256:b336ea666c74ae256f7e0ae52198e6089ab1a2d1275708876990a742091c72c4

Observation a5cf0b1b-9438-46ae-b6ed-0992a4b44ffb · outbound

This paper cites Empirical Evaluation of Gated Recurrent Neural Networks on Sequence Modeling.

Learning What to Remember: Test-Time Training via Context Distillation Empirical Evaluation of Gated Recurrent Neural Networks on Sequence Modeling

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T23:14:05.666221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:14:05.666221Z digest=sha256:3efe7a09fdd89429b6c1aa0d10ab3e1c90fece934b00ca6dadf5f0e46f64dc50

Observation 80705b35-3476-4122-acde-bde493fc613d · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

Learning What to Remember: Test-Time Training via Context Distillation Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T23:14:05.790838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:14:05.790838Z digest=sha256:7af0df9c480a2fe1a080e92cc6d5e726ae7bf75b7357b093246e760df8795a7a

Observation e27ad568-460b-442e-834f-a93ff175d8ad · outbound

This paper cites Finch: Prompt-guided key-value cache compression for large language models.Transactions of the Association for Computational Linguistics, 12: 1517–1532, 2024.

Learning What to Remember: Test-Time Training via Context Distillation Finch: Prompt-guided key-value cache compression for large language models.Transactions of the Association for Computational Linguistics, 12: 1517–1532, 2024

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:14:14.905011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-04T23:14:05.921463Z digest=sha256:4c1c61aac9d8c30c40dd84357ea3c422fa9860fbc5fa428414fbdaaf8629ff37

Observation b5f6d980-f486-4050-9bff-b202d9aeeab7 · outbound

This paper cites FlashAttention-2: Faster attention with better parallelism and work partitioning.

Learning What to Remember: Test-Time Training via Context Distillation FlashAttention-2: Faster attention with better parallelism and work partitioning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T23:14:06.021143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:14:06.021143Z digest=sha256:75aaabcc2ba6ce36ad096377b946f5c91e32696b512df8c8a58bb78ea8cb0bf3

Observation 397eb15a-02c7-4554-994a-8acdc000f1a3 · outbound

This paper cites Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality.

Learning What to Remember: Test-Time Training via Context Distillation Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T23:14:06.123574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:14:06.123574Z digest=sha256:fe791538b298687d764582ce5b443c6407af7468eaf403b2d7241bff578c52a0

Observation 0d3d0072-0426-4262-8add-27db456a8b19 · outbound

This paper cites Flashattention: Fast and memory-efficient exact attention with io-awareness.Advances in neural information processing systems, 35:16344–16359, 2022.

Learning What to Remember: Test-Time Training via Context Distillation Flashattention: Fast and memory-efficient exact attention with io-awareness.Advances in neural information processing systems, 35:16344–16359, 2022

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T23:14:06.256065Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:14:06.256065Z digest=sha256:23d1574e213ff17fa357717e5fd232e3a71957197eddc4aee27054d9e4dca156

Observation 6893cd2a-08e9-45ba-9d42-0f4d153f44cc · outbound

This paper cites Cartridges: Lightweight and general-purpose long context representations via self-study.

Learning What to Remember: Test-Time Training via Context Distillation Cartridges: Lightweight and general-purpose long context representations via self-study

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T23:14:06.389567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:14:06.389567Z digest=sha256:cd8e6fd4a9620d55fbf6403079a96931a919b221d8c76eafcc492013f606d935

Observation ab97a2d9-9dc7-428e-8d4e-58ab5386247c · outbound

This paper cites In-place test-time training.

Learning What to Remember: Test-Time Training via Context Distillation In-place test-time training

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:14:14.540371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-04T23:14:06.547378Z digest=sha256:f0e1115770b76d58118f213d6550a77c631a4d0ca39659f572ed920f3e2f8362

Observation 4db8bf9e-d62c-41f2-81bb-8cbed9025657 · outbound

This paper cites Hungry Hungry Hippos: Towards Language Modeling with State Space Models.

Learning What to Remember: Test-Time Training via Context Distillation Hungry Hungry Hippos: Towards Language Modeling with State Space Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T23:14:06.803422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:14:06.803422Z digest=sha256:29b900e477d4f944b54b206c2e19118e6cc3e3db3b9bcf13ffb072e5ad6e5fd9

Observation 8a5ee083-a876-497a-98c1-0df6f402a9a2 · outbound

This paper cites The Pile: An 800GB Dataset of Diverse Text for Language Modeling.

Learning What to Remember: Test-Time Training via Context Distillation The Pile: An 800GB Dataset of Diverse Text for Language Modeling

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T23:14:06.952421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:14:06.952421Z digest=sha256:590240402478e3450c3d21e879c574911cfc1bc028499fad01f4da3a5fcdaa2a

Observation 3018c465-52a7-4b15-a788-52b17a2974f7 · outbound

This paper cites How to train long-context language models (effectively).

Learning What to Remember: Test-Time Training via Context Distillation How to train long-context language models (effectively)

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T23:14:07.065138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:14:07.065138Z digest=sha256:c164637f5ac3860d3d133770fdb13283ab91a0c926870c1ef176d3c11e98e4d0

Observation e8cb5d61-4e89-454f-8152-0cd605d7d62e · outbound

This paper cites Model Tells You What to Discard: Adaptive KV Cache Compression for LLMs.

Learning What to Remember: Test-Time Training via Context Distillation Model Tells You What to Discard: Adaptive KV Cache Compression for LLMs

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-04T23:14:07.185124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:14:07.185124Z digest=sha256:65dcbd4089ac94b528eafe69635a762f04a36a197d665878b78ddbb33e7a5842

Observation 6622a355-149e-4d57-98bc-2a62bfb730e7 · outbound

This paper cites Mamba: Linear-Time Sequence Modeling with Selective State Spaces.

Learning What to Remember: Test-Time Training via Context Distillation Mamba: Linear-Time Sequence Modeling with Selective State Spaces

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T23:14:07.335199Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:14:07.335199Z digest=sha256:7b01558fe5f3d0e944618cc257c83cbef89a8c6484aff5fdbd2eaab176d76a9b

Observation bd5b79d2-829d-4207-a2d3-d610b49eed4c · outbound

This paper cites Efficiently Modeling Long Sequences with Structured State Spaces.

Learning What to Remember: Test-Time Training via Context Distillation Efficiently Modeling Long Sequences with Structured State Spaces

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-04T23:14:07.450289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:14:07.450289Z digest=sha256:d551ff366b762eee075d88d7b1d46f96ab5e996b768ea1904ce76c40404670db

Observation f4a080cd-867c-445a-8964-9940e2d65195 · outbound

This paper cites Log-linear attention.arXiv preprint arXiv:2506.04761, 2025.

Learning What to Remember: Test-Time Training via Context Distillation Log-linear attention.arXiv preprint arXiv:2506.04761, 2025

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-04T23:14:07.584943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:14:07.584943Z digest=sha256:ac7d45168b162c263d3a9306613df54149b0b2f772c2e3d22d0afca3e15ee741

Observation 7ab99549-c75e-45c0-856a-64fdb2f73595 · outbound

This paper cites Measuring Massive Multitask Language Understanding.

Learning What to Remember: Test-Time Training via Context Distillation Measuring Massive Multitask Language Understanding

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-04T23:14:07.739124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:14:07.739124Z digest=sha256:1d4498ec74598b17228fcfa1a8c9946785315af53cb9ec3edd92782d66f5ad4f

Observation 919ae80d-7d08-4bf1-bdd6-7359d9d19f49 · outbound

This paper cites Long short-term memory.Neural computation, 9(8): 1735–1780, 1997.

Learning What to Remember: Test-Time Training via Context Distillation Long short-term memory.Neural computation, 9(8): 1735–1780, 1997

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-04T23:14:07.880464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:14:07.880464Z digest=sha256:9125318fe0a8f0aa4d1952ea197d837fce1361fd2952f67c49ca9ca0092a93ad

Observation 226c7399-347e-414d-8446-1d605e707c2f · outbound

This paper cites RULER: What's the Real Context Size of Your Long-Context Language Models?.

Learning What to Remember: Test-Time Training via Context Distillation RULER: What's the Real Context Size of Your Long-Context Language Models?

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-04T23:14:08.026195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:14:08.026195Z digest=sha256:9d734244ab14b48136e88d7dc46868ddb71b532fab53be8572060bc11a9fdc95

Observation 64a8102d-d7e0-4921-9932-e63d85c1ac63 · outbound

This paper cites MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies.

Learning What to Remember: Test-Time Training via Context Distillation MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-04T23:14:08.177389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:14:08.177389Z digest=sha256:f658e736cdf9455cbab0ef3e37e96f469a460ef9e887aba1b5d0ccfd52ceac64

Observation 47553360-1a86-4eaa-980a-4375973611b4 · outbound

This paper cites Characterizing Prompt Compression Methods for Long Context Inference.

Learning What to Remember: Test-Time Training via Context Distillation Characterizing Prompt Compression Methods for Long Context Inference

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-04T23:14:08.287634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:14:08.287634Z digest=sha256:084e3048a83c095f7a18455b3ab0633df4ce9e49a115ebf3a7faf1d6ca2a94e1

Observation 0a3b789e-b136-4bbc-8b7d-9dd3a816c0bb · outbound

This paper cites Llmlingua: Compress- ing prompts for accelerated inference of large language models.

Learning What to Remember: Test-Time Training via Context Distillation Llmlingua: Compress- ing prompts for accelerated inference of large language models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-04T23:14:08.403619Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:14:08.403619Z digest=sha256:a82de27897181e9c08d33f4a0488f084d0db83645ee88166f5a906ff6b0b67a4

Observation cc9b1233-617f-4401-89d8-72104f18e534 · outbound

This paper cites Babilong: Testing the limits of llms with long context reasoning-in-a-haystack.

Learning What to Remember: Test-Time Training via Context Distillation Babilong: Testing the limits of llms with long context reasoning-in-a-haystack

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-04T23:14:08.528498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:14:08.528498Z digest=sha256:661a1f44a0829062dd5c2fc11d5f7c46238086e659032c3e5bbc293361cf5b30

Observation a79c02d6-a830-47a0-b543-7f2bf52a67c2 · outbound

This paper cites The power of scale for parameter-efficient prompt tuning.

Learning What to Remember: Test-Time Training via Context Distillation The power of scale for parameter-efficient prompt tuning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-04T23:14:08.662989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:14:08.662989Z digest=sha256:114438ab7ae66c6bc83f9c1832bf34f9c33404f4ac9a0d925d7f7228fa852b0e

Observation 13d8cc07-5037-4a8c-b89e-312003fc2828 · outbound

This paper cites Prefix-tuning: Optimizing continuous prompts for generation.

Learning What to Remember: Test-Time Training via Context Distillation Prefix-tuning: Optimizing continuous prompts for generation

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-04T23:14:08.800032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:14:08.800032Z digest=sha256:7b6e8783e3daa8a658bf202c2648b894abc9dfc6950996ec5a290f378bfedb0b

Observation adbd24ab-f684-4a90-bf00-1f1f4e15b14d · outbound

This paper cites Snapkv: Llm knows what you are looking for before generation.Advances in Neural Information Processing Systems, 37:22947–22970, 2024.

Learning What to Remember: Test-Time Training via Context Distillation Snapkv: Llm knows what you are looking for before generation.Advances in Neural Information Processing Systems, 37:22947–22970, 2024

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-04T23:14:08.961052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:14:08.961052Z digest=sha256:3b9313d6e694d33f793701eb5f51732412c0c6dffbf43a7d2949d17e067058b6

Observation 816d684a-addd-49aa-b6bc-8a610d2e6e2c · outbound

This paper cites Leave No Context Behind: Efficient Infinite Context Transformers with Infini-attention.

Learning What to Remember: Test-Time Training via Context Distillation Leave No Context Behind: Efficient Infinite Context Transformers with Infini-attention

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-04T23:14:09.106654Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:14:09.106654Z digest=sha256:94257c01a2c4720f35556a18ed1daeca66e0a70d8bd906dd4a109e1afd349441

Observation aae81bf1-aa7d-47b8-80a4-fe9c14319503 · outbound

This paper cites Transformers are multi-state rnns.

Learning What to Remember: Test-Time Training via Context Distillation Transformers are multi-state rnns

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:14:13.881753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-04T23:14:09.274202Z digest=sha256:fdc8dbb6a4ff59ecd63201798e48e05cf6d1c270eb3161c13a771e0108b9a319

Observation 5d86c09b-944d-4fd0-9f10-4e067ac893eb · outbound

This paper cites Rwkv: Reinventing rnns for the transformer era.

Learning What to Remember: Test-Time Training via Context Distillation Rwkv: Reinventing rnns for the transformer era

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-04T23:14:09.470969Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:14:09.470969Z digest=sha256:01b9662eb6e9ade04f20beb5fb59a6d6663ef13312edbdd0eb6bea71991944e1

Observation 40fd6dab-2f54-4d83-8ed8-b9ffee9cfd62 · outbound

This paper cites Mechanistic Design and Scaling of Hybrid Architectures.

Learning What to Remember: Test-Time Training via Context Distillation Mechanistic Design and Scaling of Hybrid Architectures

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-04T23:14:09.604533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:14:09.604533Z digest=sha256:c682796ccdf71f9736d64c8ec0a2349ca939ea4c5335fa1c2c5d7cfc8d3e7411

Observation 448891ea-e8db-4747-9d84-708f7e159209 · outbound

This paper cites Improving language understanding by generative pre-training.

Learning What to Remember: Test-Time Training via Context Distillation Improving language understanding by generative pre-training

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-04T23:14:09.767096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:14:09.767096Z digest=sha256:bb6e5b19423e74df4840835e48ef9b580ae7d138cb669bb56b95bf747ab7f298

Observation cf6f7410-8dd1-49e4-add4-177088d2587b · outbound

This paper cites Linear transformers are secretly fast weight programmers.

Learning What to Remember: Test-Time Training via Context Distillation Linear transformers are secretly fast weight programmers

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:14:13.672391Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-04T23:14:09.931300Z digest=sha256:72721209138f160fe93020f47fbba2f47856227b650681a7d1c904bc5aec2842

Observation a209b04f-5702-4b86-bd83-513447cfdfd9 · outbound

This paper cites FlashAttention-3: Fast and accurate attention with asynchrony and low-precision.

Learning What to Remember: Test-Time Training via Context Distillation FlashAttention-3: Fast and accurate attention with asynchrony and low-precision

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:14:13.354172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-04T23:14:10.050140Z digest=sha256:0a61e561c3e48d80f93f9b30564572a8c7f8dc0e65d9eff527f65fde9d4a7db6

Observation d702b95c-08f7-4a5f-ae0c-2adbb5ce5d7b · outbound

This paper cites Learning by Distilling Context.

Learning What to Remember: Test-Time Training via Context Distillation Learning by Distilling Context

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-04T23:14:10.159095Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:14:10.159095Z digest=sha256:54abe25f5fbcd0d2d17a3dc21701640160ee241f422db86fb36de3a865d2e9b6

Observation ef0a8900-bbce-4f3e-81af-7cc143090630 · outbound

This paper cites Roformer: Enhanced transformer with rotary position embedding.Neurocomputing, 568:127063, 2024.

Learning What to Remember: Test-Time Training via Context Distillation Roformer: Enhanced transformer with rotary position embedding.Neurocomputing, 568:127063, 2024

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-04T23:14:10.276672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:14:10.276672Z digest=sha256:565838b1fdf54cb5f620a36936467bd85d879fba1a3eae47c5d3515a0cf62b87

Observation 18a50eff-afb5-4a4c-87b0-6a62d78006cc · outbound

This paper cites Test-time training with self-supervision for generalization under distribution shifts.

Learning What to Remember: Test-Time Training via Context Distillation Test-time training with self-supervision for generalization under distribution shifts

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-04T23:14:10.362549Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:14:10.362549Z digest=sha256:a27282386a82b61f8ed90d0447dc07b7f38b6320f84076d1312d59450a4c23e3

Observation d84e59b2-544a-4af1-832a-fc036152c07c · outbound

This paper cites Learning to (Learn at Test Time): RNNs with Expressive Hidden States.

Learning What to Remember: Test-Time Training via Context Distillation Learning to (Learn at Test Time): RNNs with Expressive Hidden States

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-04T23:14:10.523720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:14:10.523720Z digest=sha256:91a9d0624727f4297bd872d7dec34eb4ce0b218ea83d679dd5503ae3fe08f0bc

Observation d8f1b53f-9c2b-4fee-a59c-f391f420b1b1 · outbound

This paper cites End-to-end test-time training for long context.arXiv preprint arXiv:2512.23675, 2025.

Learning What to Remember: Test-Time Training via Context Distillation End-to-end test-time training for long context.arXiv preprint arXiv:2512.23675, 2025

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-04T23:14:10.646589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:14:10.646589Z digest=sha256:191fd5916b4b91ec71cf654f74f08928eb0e8ff7e14c02bdc2eb422b17d02ca2

Observation 21e81452-944f-4526-99ff-b77dd6ce18f2 · outbound

This paper cites Kimi Linear: An Expressive, Efficient Attention Architecture.

Learning What to Remember: Test-Time Training via Context Distillation Kimi Linear: An Expressive, Efficient Attention Architecture

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-04T23:14:10.779384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:14:10.779384Z digest=sha256:4f40d0ea1ebe539ec335c41f1489179af61dc8b7d2de80d062b2899ef759747f

Observation 57870a3d-587b-49cf-8ef6-e46b09f61007 · outbound

This paper cites Long data collections database, 2024.

Learning What to Remember: Test-Time Training via Context Distillation Long data collections database, 2024

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-04T23:14:10.965292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:14:10.965292Z digest=sha256:40742157f0fd5a83987b9572d832715cb26ac4ec90a85faac04aa5c4c8da67a9

Observation 4315f292-e6c7-41fe-8d5d-f0ea54d47e33 · outbound

This paper cites Attention is all you need.Advances in neural information processing systems, 30, 2017.

Learning What to Remember: Test-Time Training via Context Distillation Attention is all you need.Advances in neural information processing systems, 30, 2017

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-04T23:14:11.133096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:14:11.133096Z digest=sha256:ee0ad3b2cd4acde5e4842ab9327f002a60cdeb3ebfc1f86a26734adf57fb1cfd

Observation 6c056be1-e655-4eef-b1f3-07ef8161b56e · outbound

This paper cites Rattention: Towards the minimal sliding window size in local-global attention models.arXiv preprint arXiv:2506.15545, 2025.

Learning What to Remember: Test-Time Training via Context Distillation Rattention: Towards the minimal sliding window size in local-global attention models.arXiv preprint arXiv:2506.15545, 2025

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-04T23:14:11.286715Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:14:11.286715Z digest=sha256:a9491b227e4266142f3e0aaf540f99de249349609886b5236dabc8a55fd01177

Observation 90465926-40c6-49b7-91bd-65939b77e029 · outbound

This paper cites Fantastic Pretraining Optimizers and Where to Find Them.

Learning What to Remember: Test-Time Training via Context Distillation Fantastic Pretraining Optimizers and Where to Find Them

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-04T23:14:11.425305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:14:11.425305Z digest=sha256:1ba2ebe2fdc68d09b8c5a0f391bfd95c544479cbd9f00a9ee7a6cfa707060028

Observation d2f92442-bf3a-438a-bc24-b9d17918de0a · outbound

This paper cites Gated Linear Attention Transformers with Hardware-Efficient Training.

Learning What to Remember: Test-Time Training via Context Distillation Gated Linear Attention Transformers with Hardware-Efficient Training

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-04T23:14:11.560413Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:14:11.560413Z digest=sha256:5b698189f0c355e801c3b04acb033820f1a1ccbcc9944c5902b3c7200852b3bb

Observation 43063023-1f75-43e9-a0ab-6981d13021b3 · outbound

This paper cites Gated Delta Networks: Improving Mamba2 with Delta Rule.

Learning What to Remember: Test-Time Training via Context Distillation Gated Delta Networks: Improving Mamba2 with Delta Rule

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-04T23:14:11.743368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:14:11.743368Z digest=sha256:f64e7446b9af51b20f151a9e26aa080b06cbb58190434fab2dbbee24ff324776

Observation 610f9608-bdd7-4f16-a98b-796fb62e02d6 · outbound

This paper cites Parallelizing linear transformers with the delta rule over sequence length.Advances in neural information processing systems, 37:115491–115522, 2024.

Learning What to Remember: Test-Time Training via Context Distillation Parallelizing linear transformers with the delta rule over sequence length.Advances in neural information processing systems, 37:115491–115522, 2024

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-04T23:14:11.883687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:14:11.883687Z digest=sha256:ac2f0c681c4b992d3b85f05e44154e8da12273b3ebbe179ed8153c0a6acc89b2

Observation 171207e8-4b27-4a4f-8e1c-4e1c333b3367 · outbound

This paper cites Native sparse attention: Hardware-aligned and natively trainable sparse attention.

Learning What to Remember: Test-Time Training via Context Distillation Native sparse attention: Hardware-aligned and natively trainable sparse attention

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-04T23:14:12.075110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:14:12.075110Z digest=sha256:ee29be48f0f2ed837bb4f741fd2a9a6e27629acc040f1620f7413ab02e5dbe35

Observation 9aed4c75-ec93-48a7-af42-4a9a1aff78fc · outbound

This paper cites Hellaswag: Can a machine really finish your sentence? InProceedings of the 57th annual meeting of the association for computational linguistics, pages 4791–4800, 2019.

Learning What to Remember: Test-Time Training via Context Distillation Hellaswag: Can a machine really finish your sentence? InProceedings of the 57th annual meeting of the association for computational linguistics, pages 4791–4800, 2019

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-04T23:14:12.201229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:14:12.201229Z digest=sha256:36f406d2e08a8fb753ecccbff28ad609db13c0ef5ae1d3a2b512091256b8a951

Observation 941c35c7-1044-4412-8a89-17dec2777047 · outbound

This paper cites Test-Time Training Done Right.

Learning What to Remember: Test-Time Training via Context Distillation Test-Time Training Done Right

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-04T23:14:12.314176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:14:12.314176Z digest=sha256:aa245ab0601d153d4fa76fe9f08e8cfdc51bbbb7d50bb6b46ebe1da924d5f631

Observation 76baa423-df23-4828-8c47-e691ab3d19a8 · outbound

This paper cites up to 1.3×.

Learning What to Remember: Test-Time Training via Context Distillation up to 1.3×

Reference 56

Resolution
malformed identifier
arxiv_id, observed 2026-08-04T23:14:12.722467Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-04T23:14:12.487890Z digest=sha256:ba75928362271b4ba74c0cc2bedca3687129ce1c672b897f2a84eea9f05c4dc0

Observation 129dc4f6-bd6b-4602-9d85-ca55dde4b2bd · outbound

This paper cites an unresolved cited work.

Learning What to Remember: Test-Time Training via Context Distillation Unresolved cited work

Reference 2026

Resolution
unresolved
raw_fallback, observed 2026-08-04T23:14:14.291009Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-04T23:14:06.676734Z digest=sha256:7f2455c4032b0a547c76fb07e18aed36b3f415e02757355e7a2b5d935ad6d7b8

Pith citing papers

No inbound Pith citation observations are available.