Pith. sign in

Paper Citation Record · LEDGER

Distilled Pretraining: A modern lens of Data, In-Context Learning and Test-Time Scaling

As of 9 August 2026, this Paper Citation Record lists 68 of 68 outbound references and 3 inbound Pith citation observations for arXiv:2509.01649.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.01649 v1

Coverage vector

measured 68 of 68 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T12:25:28.508547Z

measured 71 of 71 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-01T03:02:07.425724Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-21T05:49:40.779767Z

Reference resolution

68 of 68 outbound references displayed

  • verified exact4
  • verified fuzzy16
  • unresolved47
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation bebfc12a-0e3c-4625-95a5-75834a84312f · outbound

This paper cites write newline.

Distilled Pretraining: A modern lens of Data, In-Context Learning and Test-Time Scaling write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T12:25:22.958457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T12:25:22.958457Z digest=sha256:6d1eaefc36f6a7a425feb7b6d270f20d8ce44e42003a0e41e3e4a74e5d442b3f

Observation 170ad7b5-b0ea-44cd-8d91-ef4ff033c4fa · outbound

This paper cites write newline.

Distilled Pretraining: A modern lens of Data, In-Context Learning and Test-Time Scaling write newline

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-05T12:25:23.019471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T12:25:23.019471Z digest=sha256:86f47a9795bfe3b27affee8055b36bfb63c9ce6897cc31f88fd2d9d2c164c2bf

Observation f6ef6a1f-07ce-4a2a-9ca5-c55a343986d3 · outbound

This paper cites @esa (Ref.

Distilled Pretraining: A modern lens of Data, In-Context Learning and Test-Time Scaling @esa (Ref

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-05T12:25:23.125793Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T12:25:23.125793Z digest=sha256:56b311ab578483419f87b3090b8ac2b7945dd97ae26506fb08ca364d7a289352

Observation 73c6fc5c-8d25-41fa-a7e3-621d5c5e05e2 · outbound

This paper cites an unresolved cited work.

Distilled Pretraining: A modern lens of Data, In-Context Learning and Test-Time Scaling Unresolved cited work

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T12:25:23.209699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T12:25:23.209699Z digest=sha256:c56b175be19206fb7c1b80bdadcaae402a220ab7974ff25f0b50aecb14a0313a

Observation f9d525d0-7efe-4c5a-8840-c8115731ff7a · outbound

This paper cites an unresolved cited work.

Distilled Pretraining: A modern lens of Data, In-Context Learning and Test-Time Scaling Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-05T12:25:29.878023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T12:25:23.286873Z digest=sha256:b99e914c5b80a10f680660ee8c49fbd8b42ad908ffc5f59e483d9fcebc7af652

Observation 13ee03da-a6ba-4052-8e00-1052b8bd690e · outbound

This paper cites On-Policy Distillation of Language Models: Learning from Self-Generated Mistakes.

Distilled Pretraining: A modern lens of Data, In-Context Learning and Test-Time Scaling On-Policy Distillation of Language Models: Learning from Self-Generated Mistakes

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T12:25:23.384000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T12:25:23.384000Z digest=sha256:7bae05ae24c65ebf11ec0c67011bcb8e18c8c88c694d8ed89d616b5b350677e2

Observation 0f4d3823-fe0f-494b-b090-8a614bc9f5fc · outbound

This paper cites Alphaevolve: A gemini-powered coding agent for designing advanced algorithms, 2025.

Distilled Pretraining: A modern lens of Data, In-Context Learning and Test-Time Scaling Alphaevolve: A gemini-powered coding agent for designing advanced algorithms, 2025

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:25:29.864076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T12:25:23.494428Z digest=sha256:7114db7a7660dc3dfc7a4bb0bc2f1cabe392f6809f8560d9e361082e05e62cca

Observation b877d89e-b5d5-443a-b6d2-bdf26a6d18a4 · outbound

This paper cites Program Synthesis with Large Language Models.

Distilled Pretraining: A modern lens of Data, In-Context Learning and Test-Time Scaling Program Synthesis with Large Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-05T12:25:23.571192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T12:25:23.571192Z digest=sha256:f6f9250caf0471d435f9660498c231f0d9e62fb36e294ac46da0c5337911cd06

Observation 7a7d6cf2-48c2-4210-b588-804baacad62c · outbound

This paper cites Jiang, Jia Deng, Stella Biderman, and Sean Welleck.

Distilled Pretraining: A modern lens of Data, In-Context Learning and Test-Time Scaling Jiang, Jia Deng, Stella Biderman, and Sean Welleck

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:25:29.849538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T12:25:23.667642Z digest=sha256:ad5603b3520862888a6ff992e4cacba4c193318cc740abcf239fffa82f7b31c9

Observation d9d4f780-ffd3-4fc5-b148-a943f3ff6a63 · outbound

This paper cites Do deep nets really need to be deep? Advances in neural information processing systems, 27, 2014.

Distilled Pretraining: A modern lens of Data, In-Context Learning and Test-Time Scaling Do deep nets really need to be deep? Advances in neural information processing systems, 27, 2014

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T12:25:23.760335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T12:25:23.760335Z digest=sha256:4071797866ea2f252647453b9b9e8bae508d3aa080440bf6f9fa477ef5e300c0

Observation 37f6152d-fbcd-4502-9ded-cd4546b6deb2 · outbound

This paper cites Scaling test-time compute with open models, 2024.

Distilled Pretraining: A modern lens of Data, In-Context Learning and Test-Time Scaling Scaling test-time compute with open models, 2024

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:25:29.824431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T12:25:23.842264Z digest=sha256:db924150f6865897c830be03c23d545fe8dc9078fa97451d15802c1441731d7c

Observation e31731bc-8330-4580-aa69-80e800748b77 · outbound

This paper cites Knowledge distillation: A good teacher is patient and consistent.

Distilled Pretraining: A modern lens of Data, In-Context Learning and Test-Time Scaling Knowledge distillation: A good teacher is patient and consistent

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T12:25:23.941000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T12:25:23.941000Z digest=sha256:328034a1ef8238149395ba571e6050eddb8a813db3473de0c308ac2c4ec18beb

Observation d4477826-47e3-49cd-bffa-20fb3c8a19ee · outbound

This paper cites Birth of a transformer: A memory viewpoint.

Distilled Pretraining: A modern lens of Data, In-Context Learning and Test-Time Scaling Birth of a transformer: A memory viewpoint

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:25:29.808008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T12:25:24.039161Z digest=sha256:856af66ae1b9ca04868c688a592bd4149411afc6e4feb6cff95b5557c738b93e

Observation 88af22cc-8883-4c65-8514-7f7a59104101 · outbound

This paper cites Model compression.

Distilled Pretraining: A modern lens of Data, In-Context Learning and Test-Time Scaling Model compression

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-05T12:25:24.111826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T12:25:24.111826Z digest=sha256:fa37c6ed69e0bea40093cbebade101f640f413e90d3641b36f6175e313601a79

Observation d9287584-c60c-4e1f-b10f-1b5121db1396 · outbound

This paper cites Distillation Scaling Laws.

Distilled Pretraining: A modern lens of Data, In-Context Learning and Test-Time Scaling Distillation Scaling Laws

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-05T12:25:24.197937Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T12:25:24.197937Z digest=sha256:c7922e58045dafd627df7441adc3d1da5e164340924f3182fc6be17ad3eb919c

Observation 13688017-0268-4d3b-aba3-de9cf28eacea · outbound

This paper cites Why knowledge distillation works in generative models: A minimal working explanation.

Distilled Pretraining: A modern lens of Data, In-Context Learning and Test-Time Scaling Why knowledge distillation works in generative models: A minimal working explanation

Reference 16

Resolution
verified exact
raw_fallback, observed 2026-08-05T12:25:29.535717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T12:25:24.290047Z digest=sha256:531366d438e18bf75a35e9382801f735883af32d674b824e7c03a44c194cc2ef

Observation 8a84ea49-d432-49ff-813d-59b1cf6abef4 · outbound

This paper cites Rethinking fine-tuning when scaling test-time compute: Limiting confidence improves mathematical reasoning, 2025.

Distilled Pretraining: A modern lens of Data, In-Context Learning and Test-Time Scaling Rethinking fine-tuning when scaling test-time compute: Limiting confidence improves mathematical reasoning, 2025

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T12:25:24.382110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T12:25:24.382110Z digest=sha256:a9785a58771a28963c9ad924a3c82693d27fb85784bb64b4e1c3bc6b92a0b67e

Observation 3751ea62-5e02-48cc-bd39-ce387294c4a2 · outbound

This paper cites AlphaMath Almost Zero: Process Supervision without Process.

Distilled Pretraining: A modern lens of Data, In-Context Learning and Test-Time Scaling AlphaMath Almost Zero: Process Supervision without Process

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-05T12:25:24.465741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T12:25:24.465741Z digest=sha256:c388b1c117ffd7a05d5adadf349b7835f39c6991c506958d4c5ff8cc1bd77eeb

Observation bd39482b-8f99-48b4-b476-9d495338e925 · outbound

This paper cites On the Efficacy of Knowledge Distillation.

Distilled Pretraining: A modern lens of Data, In-Context Learning and Test-Time Scaling On the Efficacy of Knowledge Distillation

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-05T12:25:24.562713Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T12:25:24.562713Z digest=sha256:360b2fa2fb34236288c48ad6a12f2d3705f2e3af68f36ca14d5264412a5de365

Observation 868e1558-4f58-4504-a633-676f96595b73 · outbound

This paper cites Inference-aware fine-tuning for best-of-n sampling in large language models, 2024.

Distilled Pretraining: A modern lens of Data, In-Context Learning and Test-Time Scaling Inference-aware fine-tuning for best-of-n sampling in large language models, 2024

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-05T12:25:24.661688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T12:25:24.661688Z digest=sha256:644d65026a32b7ac16daa78534bb535b803f9fd15372210302a97c9faa59fbc2

Observation 8c19ebc7-857c-426c-9bde-ab3eaeeee94c · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Distilled Pretraining: A modern lens of Data, In-Context Learning and Test-Time Scaling Training Verifiers to Solve Math Word Problems

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-05T12:25:24.753080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T12:25:24.753080Z digest=sha256:57231dc134883b71722a8a045d08083d42ec1a33e61cca180c2d640f922fc424

Observation 40bcfe5b-aacf-4f6e-8f25-f3186fb1aa5e · outbound

This paper cites Weight ensembling improves reasoning in language models, 2025.

Distilled Pretraining: A modern lens of Data, In-Context Learning and Test-Time Scaling Weight ensembling improves reasoning in language models, 2025

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-05T12:25:24.878084Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T12:25:24.878084Z digest=sha256:dadc84dd6edbb9fa42b85b7b3566375bf49859bcdb715b1004f5cdef2d57df3e

Observation 2331a9c0-d643-47c0-a4e9-eeb4b0349675 · outbound

This paper cites DROP : A reading comprehension benchmark requiring discrete reasoning over paragraphs.

Distilled Pretraining: A modern lens of Data, In-Context Learning and Test-Time Scaling DROP : A reading comprehension benchmark requiring discrete reasoning over paragraphs

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:25:29.782758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T12:25:25.003825Z digest=sha256:4936941156c52b1cd0a385dd58bc8aef469251898eacabaf16433c3c3c514518

Observation 3162a577-bca9-4c0a-891f-828d349a23c3 · outbound

This paper cites The Evolution of Statistical Induction Heads: In-Context Learning Markov Chains.

Distilled Pretraining: A modern lens of Data, In-Context Learning and Test-Time Scaling The Evolution of Statistical Induction Heads: In-Context Learning Markov Chains

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-05T12:25:25.085058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T12:25:25.085058Z digest=sha256:0bcee7c158226eeadc87c16cb5258c73a1475db4e7dcd915a2c2b83c98f0e8b1

Observation 0b7e0d21-d7d3-4292-833c-b36f0f938a44 · outbound

This paper cites Born again neural networks.

Distilled Pretraining: A modern lens of Data, In-Context Learning and Test-Time Scaling Born again neural networks

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:25:29.768885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T12:25:25.189309Z digest=sha256:e39a37437a580c6844e7549104abdb918027304a6375d1e2ddee5b886ddf0589

Observation 7740d992-e655-4e5b-bcf1-92634caac227 · outbound

This paper cites The Pile: An 800GB Dataset of Diverse Text for Language Modeling.

Distilled Pretraining: A modern lens of Data, In-Context Learning and Test-Time Scaling The Pile: An 800GB Dataset of Diverse Text for Language Modeling

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-05T12:25:25.292590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T12:25:25.292590Z digest=sha256:e26500af15733be7b621a70ac0e9ba7a55c659acdc506094e4c97a618d7b879f

Observation 63fa1951-723a-4c0d-b665-b4b1e524d4a2 · outbound

This paper cites Gemma: Open Models Based on Gemini Research and Technology.

Distilled Pretraining: A modern lens of Data, In-Context Learning and Test-Time Scaling Gemma: Open Models Based on Gemini Research and Technology

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-05T12:25:25.407357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T12:25:25.407357Z digest=sha256:d5edd0f2148863be341203822346e61ada4528615104bdf456b141c88f82ea94

Observation 15800d26-2908-495b-b1a6-dc7e07adcffe · outbound

This paper cites Gemma 3 Technical Report.

Distilled Pretraining: A modern lens of Data, In-Context Learning and Test-Time Scaling Gemma 3 Technical Report

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-05T12:25:25.533802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T12:25:25.533802Z digest=sha256:11d32dfc0e8ba47f5f0f4c47dae4b345f836a1b9566b3e390454ce85a0e06be7

Observation efb07af1-884e-43c0-9ec7-f8dbe9355320 · outbound

This paper cites Multi-Token Prediction Needs Registers.

Distilled Pretraining: A modern lens of Data, In-Context Learning and Test-Time Scaling Multi-Token Prediction Needs Registers

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-08-05T12:25:29.120281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T12:25:25.612205Z digest=sha256:d129370414cb7e815e150c99484058566d5de4b104af40f96851356d6cac0d54

Observation c07f773e-84fc-4187-a908-a2a26fe6a00a · outbound

This paper cites Better & Faster Large Language Models via Multi-token Prediction.

Distilled Pretraining: A modern lens of Data, In-Context Learning and Test-Time Scaling Better & Faster Large Language Models via Multi-token Prediction

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-05T12:25:25.693098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T12:25:25.693098Z digest=sha256:d34f685a00908858853556d30ae6350366da46fb25f6ac9cc9b65802d7923f52

Observation 916976b7-535c-48aa-a155-adb098a62325 · outbound

This paper cites Knowledge distillation: A survey.

Distilled Pretraining: A modern lens of Data, In-Context Learning and Test-Time Scaling Knowledge distillation: A survey

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:25:29.754807Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T12:25:25.826857Z digest=sha256:ada0e0fe53d4189b925a0c320bb28b402e8da17064e82ed6e6036ffc1d471bcc

Observation 950d04fa-88ce-495d-9437-76ff921952f3 · outbound

This paper cites Context-Parametric Inversion: Why Instruction Finetuning Can Worsen Context Reliance.

Distilled Pretraining: A modern lens of Data, In-Context Learning and Test-Time Scaling Context-Parametric Inversion: Why Instruction Finetuning Can Worsen Context Reliance

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-05T12:25:25.983146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T12:25:25.983146Z digest=sha256:03040362882ee52e724b16142c0364f488bd14461780ef878986fa903ec102bb

Observation a16dd5d6-37af-4f07-bb5b-3bcafd78c977 · outbound

This paper cites MiniPLM: Knowledge Distillation for Pre-Training Language Models.

Distilled Pretraining: A modern lens of Data, In-Context Learning and Test-Time Scaling MiniPLM: Knowledge Distillation for Pre-Training Language Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-05T12:25:26.071158Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T12:25:26.071158Z digest=sha256:df5e853f211b8d8f6beda184ecba3ed1d5eb4c9f1e3a9714cc454bc85df73b31

Observation e3603938-50c9-446d-9107-271dc73a8ef6 · outbound

This paper cites OpenThoughts: Data Recipes for Reasoning Models.

Distilled Pretraining: A modern lens of Data, In-Context Learning and Test-Time Scaling OpenThoughts: Data Recipes for Reasoning Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-05T12:25:26.151714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T12:25:26.151714Z digest=sha256:d6ba5f819009834679866896a0b06fad5005bc834b44fd4dc1e02df1cf0429ab

Observation 6a2fbe04-2c5c-4787-8fd7-a197b2a4f0e3 · outbound

This paper cites Distilling the Knowledge in a Neural Network.

Distilled Pretraining: A modern lens of Data, In-Context Learning and Test-Time Scaling Distilling the Knowledge in a Neural Network

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-05T12:25:26.237527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T12:25:26.237527Z digest=sha256:fcc3f91a759949e0a6399d929c7ef6a1934c8c979bddd02f9b2ebcef288689c5

Observation 319d01d5-180e-4cec-b10a-5bd0d5c4394d · outbound

This paper cites Babilong: Testing the limits of llms with long context reasoning-in-a-haystack, 2024.

Distilled Pretraining: A modern lens of Data, In-Context Learning and Test-Time Scaling Babilong: Testing the limits of llms with long context reasoning-in-a-haystack, 2024

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:25:29.740240Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T12:25:26.317897Z digest=sha256:d5b8dab84928925bfd5328cca2e3f787b54e5546fb477db9c11f87f29b92a703

Observation 9b800b41-3613-4b6f-94ae-bf1e4b6442ab · outbound

This paper cites RACE : Large-scale R e A ding comprehension dataset from examinations.

Distilled Pretraining: A modern lens of Data, In-Context Learning and Test-Time Scaling RACE : Large-scale R e A ding comprehension dataset from examinations

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-05T12:25:26.460696Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T12:25:26.460696Z digest=sha256:a52299eef237e1c0b2cb16cddf3e52683277c7c8487d49d502f60e3f38359c39

Observation 8f8129e7-95e5-4f7f-bfa2-116c3889867b · outbound

This paper cites DataComp-LM: In search of the next generation of training sets for language models.

Distilled Pretraining: A modern lens of Data, In-Context Learning and Test-Time Scaling DataComp-LM: In search of the next generation of training sets for language models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-05T12:25:26.534257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T12:25:26.534257Z digest=sha256:a9c2c4830f81fe7614188e7994284522b283973f975255f0e6ab2fc70ebcf958

Observation d1bfaa4f-9ee2-4737-9226-6586500a26b9 · outbound

This paper cites Dynamic Knowledge Distillation for Pre-trained Language Models.

Distilled Pretraining: A modern lens of Data, In-Context Learning and Test-Time Scaling Dynamic Knowledge Distillation for Pre-trained Language Models

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-08-05T12:25:28.999083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T12:25:26.646130Z digest=sha256:839a5b47c5a6ebdf77f35c57df2d916c3b9950bd1e1163a6ea96f855a498d607

Observation 75528aac-f872-4454-b9ca-66c4769d2145 · outbound

This paper cites Multi-Agent Verification: Scaling Test-Time Compute with Multiple Verifiers.

Distilled Pretraining: A modern lens of Data, In-Context Learning and Test-Time Scaling Multi-Agent Verification: Scaling Test-Time Compute with Multiple Verifiers

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-05T12:25:26.743517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T12:25:26.743517Z digest=sha256:754b57b94492ffc71def444bdd8d5774b0c230e5f68efc45e0833d1f9a258a6a

Observation da525c05-8ae1-47c4-afa0-52d6583dfb8e · outbound

This paper cites Let's Verify Step by Step.

Distilled Pretraining: A modern lens of Data, In-Context Learning and Test-Time Scaling Let's Verify Step by Step

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-05T12:25:26.814614Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T12:25:26.814614Z digest=sha256:ae2cdc9e5666c33285b15c4a17de578c30fb5356d1a66d3b2e235fbce783d183

Observation 1dfec237-e73a-4b93-9c3c-da7387c8bdb7 · outbound

This paper cites Program Induction by Rationale Generation : Learning to Solve and Explain Algebraic Word Problems.

Distilled Pretraining: A modern lens of Data, In-Context Learning and Test-Time Scaling Program Induction by Rationale Generation : Learning to Solve and Explain Algebraic Word Problems

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-05T12:25:26.957355Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T12:25:26.957355Z digest=sha256:c2a1d37197e5f0f0b167f3750fb50fff21c6dcb6f25c2a3852bfd04c59fec87d

Observation f66d842b-8d0d-4730-8b0f-226a3dbc9704 · outbound

This paper cites Unifying distillation and privileged information.

Distilled Pretraining: A modern lens of Data, In-Context Learning and Test-Time Scaling Unifying distillation and privileged information

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-05T12:25:27.075470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T12:25:27.075470Z digest=sha256:e20cd37742891b22ee15f9c4b3c6975e2adecd237584fd507b2eebaaac3a279b

Observation 11ea56a0-aa12-4035-9a03-e35ba87fecee · outbound

This paper cites Fineweb-edu: the finest collection of educational content, 2024.

Distilled Pretraining: A modern lens of Data, In-Context Learning and Test-Time Scaling Fineweb-edu: the finest collection of educational content, 2024

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:25:29.725818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T12:25:27.160458Z digest=sha256:c114b77354e99a5addb433d403e948e12fc7576a8c3c787b7e2f4f51bc4fc364

Observation c30b19de-14d5-441a-bb7b-ba5a9fbb23a2 · outbound

This paper cites A statistical perspective on distillation.

Distilled Pretraining: A modern lens of Data, In-Context Learning and Test-Time Scaling A statistical perspective on distillation

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:25:29.711070Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T12:25:27.208168Z digest=sha256:3704b54149a8c33702bb0c8be5715b25d518487c47ff72de93e95fb90b71d465

Observation a0767308-b056-4601-8786-1bf4d746dc22 · outbound

This paper cites The llama 4 herd: The beginning of a new era of natively multimodal ai innovation.

Distilled Pretraining: A modern lens of Data, In-Context Learning and Test-Time Scaling The llama 4 herd: The beginning of a new era of natively multimodal ai innovation

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:25:29.695815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T12:25:27.320379Z digest=sha256:4a5fdbc362d4b901684cdfd744baa39af37cf75663b9b35f9be437307df37999

Observation b85ab490-80ee-40a4-a470-343191edce28 · outbound

This paper cites Llama 3.2: Revolutionizing edge ai and vision with open, customizable models.

Distilled Pretraining: A modern lens of Data, In-Context Learning and Test-Time Scaling Llama 3.2: Revolutionizing edge ai and vision with open, customizable models

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:25:29.678919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T12:25:27.424089Z digest=sha256:4fc032bbdc294ab2bdba9c733937f326685089413cf5dbcdbed7cb49c0790050

Observation 557e08d9-99b3-4932-9822-923375655c2b · outbound

This paper cites Improved Knowledge Distillation via Teacher Assistant.

Distilled Pretraining: A modern lens of Data, In-Context Learning and Test-Time Scaling Improved Knowledge Distillation via Teacher Assistant

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-05T12:25:27.560427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T12:25:27.560427Z digest=sha256:1a053134f4b22cffa9fad73145118ce53720c023285032cde5e8cc36fb474901

Observation 905b4434-242f-48ea-a140-62e4deacf45c · outbound

This paper cites Improved knowledge distillation via teacher assistant.

Distilled Pretraining: A modern lens of Data, In-Context Learning and Test-Time Scaling Improved knowledge distillation via teacher assistant

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:25:29.662757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T12:25:27.725332Z digest=sha256:d5b876c8882139de4aadd4afe16729e997a4032b2148b79f70a9e13f31048806

Observation 8e753b6c-bbe5-4a56-bbd5-639b7da58a5c · outbound

This paper cites Self-Distillation Amplifies Regularization in Hilbert Space.

Distilled Pretraining: A modern lens of Data, In-Context Learning and Test-Time Scaling Self-Distillation Amplifies Regularization in Hilbert Space

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-05T12:25:27.847217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T12:25:27.847217Z digest=sha256:3728676e14e4ca307456fff1867e5e0ab2ada9f9d59d28ecc9378496fb41ece5

Observation ad9e1f80-2dc3-4368-91a8-1417616395fd · outbound

This paper cites s1: Simple test-time scaling.

Distilled Pretraining: A modern lens of Data, In-Context Learning and Test-Time Scaling s1: Simple test-time scaling

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-05T12:25:27.955742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T12:25:27.955742Z digest=sha256:9ebfa9e25f279eb08b332538ee081860e2723a668dfe9fbbc6ef781bfd7b2b8f

Observation 4abc85d8-de09-4335-b395-1adace19a8e8 · outbound

This paper cites On student-teacher deviations in distillation: does it pay to disobey?.

Distilled Pretraining: A modern lens of Data, In-Context Learning and Test-Time Scaling On student-teacher deviations in distillation: does it pay to disobey?

Reference 52

Resolution
verified exact
local_arxiv, observed 2026-08-05T12:25:28.870109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T12:25:28.121806Z digest=sha256:6cc428cf05a08edf0d9002eeda828fb9c4308772e4c21e3f34c5e235fb176fe3

Observation f0d95de9-d0a1-4a0b-930c-063cb50ab965 · outbound

This paper cites Roll the dice & look before you leap: Going beyond the creative limits of next-token prediction.

Distilled Pretraining: A modern lens of Data, In-Context Learning and Test-Time Scaling Roll the dice & look before you leap: Going beyond the creative limits of next-token prediction

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-05T12:25:28.192200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T12:25:28.192200Z digest=sha256:d6864070891cba27debce726be6425571c0c7c5692fd480ec85748434d41e171

Observation 17dc5935-242b-443f-80fa-80f87e8db6ee · outbound

This paper cites Github code dataset, 2022.

Distilled Pretraining: A modern lens of Data, In-Context Learning and Test-Time Scaling Github code dataset, 2022

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:25:29.649243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T12:25:28.300845Z digest=sha256:602d3f43475eae3e219a469ddc037a9677cdc0673e76a7aace9bad3d30dc3ddc

Observation 3b8cad1c-5e9b-4692-a04e-fd7d04d2e1d5 · outbound

This paper cites In-context Learning and Induction Heads.

Distilled Pretraining: A modern lens of Data, In-Context Learning and Test-Time Scaling In-context Learning and Induction Heads

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-05T12:25:28.393471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T12:25:28.393471Z digest=sha256:a404ff9c0e099dd8a908b07702c3b00ee070422e4799051e1ad8c970a872577e

Observation 1024990e-3795-47dc-8873-a393f8cc9b43 · outbound

This paper cites Towards understanding knowledge distillation.

Distilled Pretraining: A modern lens of Data, In-Context Learning and Test-Time Scaling Towards understanding knowledge distillation

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:25:29.635210Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T12:25:28.453398Z digest=sha256:137e56416bd8775b995d28212143554d24064a9674ff54b90346dfab5be2ab5a

Observation e5f48848-06da-4b1d-9e02-33e09e48f46b · outbound

This paper cites Knowledge distillation performs partial variance reduction.

Distilled Pretraining: A modern lens of Data, In-Context Learning and Test-Time Scaling Knowledge distillation performs partial variance reduction

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:25:29.619969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T12:25:28.458085Z digest=sha256:121cb026b47f868d6530998cffd2e5b19e0e62cfc539204a2d91bd8b22cc8742

Observation 52a39e37-f54e-4ff0-8073-df32c9f50524 · outbound

This paper cites Analysing Mathematical Reasoning Abilities of Neural Models.

Distilled Pretraining: A modern lens of Data, In-Context Learning and Test-Time Scaling Analysing Mathematical Reasoning Abilities of Neural Models

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-05T12:25:28.462546Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T12:25:28.462546Z digest=sha256:d4d182e3795b9e3aefd8242a1f59787d49dab2a6eeb552593c9628baa55beac8

Observation f223f575-148c-4056-abc9-c696b265925b · outbound

This paper cites BOND: Aligning LLMs with Best-of-N Distillation.

Distilled Pretraining: A modern lens of Data, In-Context Learning and Test-Time Scaling BOND: Aligning LLMs with Best-of-N Distillation

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-05T12:25:28.468096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T12:25:28.468096Z digest=sha256:015cc250e14fee277a4c5f0ccaa9bd541a9a945a6aab54c638c00865ff248ec8

Observation 794df7d2-8c04-4523-afba-0c37c3b052a7 · outbound

This paper cites Rewarding Progress: Scaling Automated Process Verifiers for LLM Reasoning.

Distilled Pretraining: A modern lens of Data, In-Context Learning and Test-Time Scaling Rewarding Progress: Scaling Automated Process Verifiers for LLM Reasoning

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-05T12:25:28.473347Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T12:25:28.473347Z digest=sha256:f60c07ef992ab0ddf6fc5cef538018f4e9fadddc501822721c959ac22e521527

Observation cadc9b2e-20f5-4ea4-938f-aeb4cdea2e7d · outbound

This paper cites Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters.

Distilled Pretraining: A modern lens of Data, In-Context Learning and Test-Time Scaling Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-05T12:25:28.477571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T12:25:28.477571Z digest=sha256:4db95e7c32012a1d74f37898333121daf53597d84ac213ba114ae979b49c699b

Observation b0e3b286-442b-4737-9c7e-39e4b5a9fd4c · outbound

This paper cites Looking beyond the next token.

Distilled Pretraining: A modern lens of Data, In-Context Learning and Test-Time Scaling Looking beyond the next token

Reference 62

Resolution
metadata mismatch
local_arxiv, observed 2026-08-05T12:25:28.757794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T12:25:28.481689Z digest=sha256:79742ef3f3114f3cfddd3b82b6ab16611d5839c8d37462084d75f6cd10b068db

Observation 8ff7a89c-ae5c-49ee-849d-0e5cdf7798b2 · outbound

This paper cites Qwen3 Technical Report.

Distilled Pretraining: A modern lens of Data, In-Context Learning and Test-Time Scaling Qwen3 Technical Report

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-05T12:25:28.485825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T12:25:28.485825Z digest=sha256:920e54315756d7d1a3e2e84e34d53cdbcefa02c01b7e88c5361259cb72c0ab94

Observation 8db1bde8-d7fc-41c4-8ebd-99f2ca04cbfc · outbound

This paper cites Naturalreasoning: Reasoning in the wild with 2.8m challenging questions, 2025.

Distilled Pretraining: A modern lens of Data, In-Context Learning and Test-Time Scaling Naturalreasoning: Reasoning in the wild with 2.8m challenging questions, 2025

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-05T12:25:28.490690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T12:25:28.490690Z digest=sha256:64f0123f4bf09069837c39764884ca01de3d23bfb8e4ca071588deab2a9d86a9

Observation 2f7f2f36-07a8-4106-bf54-b3fae88e91f3 · outbound

This paper cites Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?.

Distilled Pretraining: A modern lens of Data, In-Context Learning and Test-Time Scaling Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-05T12:25:28.494576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T12:25:28.494576Z digest=sha256:5d9b4b169eec7daf1c74e3c476c7f2347e85fa3f16cae31582a3f85e5e159ddd

Observation 45fc63d8-b021-47dd-ad0c-245369d337a8 · outbound

This paper cites Lifting the Curse of Capacity Gap in Distilling Language Models.

Distilled Pretraining: A modern lens of Data, In-Context Learning and Test-Time Scaling Lifting the Curse of Capacity Gap in Distilling Language Models

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-05T12:25:28.499159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T12:25:28.499159Z digest=sha256:7cf931fd26ddc671ebc3449ccf2113b3ef0d33d329ac41d599e6e0b67d4f58e1

Observation 520022cd-cc73-4caa-9412-06279f1c3ca1 · outbound

This paper cites Towards the Law of Capacity Gap in Distilling Language Models.

Distilled Pretraining: A modern lens of Data, In-Context Learning and Test-Time Scaling Towards the Law of Capacity Gap in Distilling Language Models

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-05T12:25:28.503732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T12:25:28.503732Z digest=sha256:531e4ddf99de8140a50e7e829ff6c85b9ad1fd0f362e1f89754d12f067812ba0

Observation 57f74805-85f5-41e1-a581-74370717f587 · outbound

This paper cites Forcing Diffuse Distributions out of Language Models.

Distilled Pretraining: A modern lens of Data, In-Context Learning and Test-Time Scaling Forcing Diffuse Distributions out of Language Models

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-05T12:25:28.508547Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T12:25:28.508547Z digest=sha256:bca32dc3dd562d6d52bb80bc1184279a3da65d45636eaf6e450b0b8931da60f1

Pith citing papers

Observation 04476116-ecc1-4122-944c-2d0c395bdfa6 · inbound

Ministral 3 cites this paper.

Ministral 3 Distilled Pretraining: A modern lens of Data, In-Context Learning and Test-Time Scaling

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T19:12:24.696963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-14T19:12:24.627033Z digest=sha256:98eb4bfa6eaf0ce4ae1fd326e41e425e6e7421f1a5207ea2a3b2840d2327f0dd

Observation bc6c8d4f-1d84-421f-85d3-a20dac254e15 · inbound

Memory Grafting: Scaling Language Model Pre-training via Offline Conditional Memory cites this paper.

Memory Grafting: Scaling Language Model Pre-training via Offline Conditional Memory Distilled Pretraining: A modern lens of Data, In-Context Learning and Test-Time Scaling

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-21T05:49:40.781393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-21T05:49:14.789955Z digest=sha256:bb17c5da63d77438679adb0d5c562fe3024e250af607d36eafe4fe686004ac43

Observation fb3df0d2-ba17-40a7-98f1-6e64c0bf2d94 · inbound

Bridging Compute- and Data-Optimal Pretraining cites this paper.

Bridging Compute- and Data-Optimal Pretraining Distilled Pretraining: A modern lens of Data, In-Context Learning and Test-Time Scaling

Reference 102

Resolution
unresolved
no resolver link, observed 2026-08-01T03:02:07.425724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T03:02:07.425724Z digest=sha256:0ced7db6fb8059f83ea48d46614e57c74968c1b33efe1d8c118bdac195932672