Pith. sign in

Paper Citation Record · LEDGER

TREK: Distill to Explore, Reinforce to Refine

As of 10 August 2026, this Paper Citation Record lists 29 of 29 outbound references and 0 inbound Pith citation observations for arXiv:2607.05339.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.05339 v1

Coverage vector

measured 29 of 29 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-07T16:22:55.342495Z

measured 29 of 29 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

29 of 29 outbound references displayed

  • verified exact21
  • verified fuzzy1
  • unresolved2
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch5

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 80fd619e-8d26-41d4-bcfd-4500c6ca129d · outbound

This paper cites an unresolved cited work.

TREK: Distill to Explore, Reinforce to Refine Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-07-07T16:24:01.114846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-07T16:22:55.342495Z digest=sha256:09308585e800379cc4e8d95ff6f0ac7ab6d923acb94da0d0a8ff9a81f44c8b0f

Observation 5352be6f-0bc7-4a50-833a-a960d1605d4f · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

TREK: Distill to Explore, Reinforce to Refine DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-07-07T16:24:01.004931Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-07T16:22:55.342495Z digest=sha256:03edf99363c66bcdf3f1b02df1765688f10139d593135f980fffc0e18a34a39d

Observation 6633503d-8a18-4a48-818e-a0717298c2f7 · outbound

This paper cites On-Policy Distillation of Language Models: Learning from Self-Generated Mistakes.

TREK: Distill to Explore, Reinforce to Refine On-Policy Distillation of Language Models: Learning from Self-Generated Mistakes

Reference 3

Resolution
metadata mismatch
local_arxiv, observed 2026-07-07T16:24:01.039917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-07T16:22:55.342495Z digest=sha256:a41f9aed65c937e7f4db3a347a6b8fe35198a07c6b5bd3cb7e7245c7e5c83d1f

Observation ff8e90aa-e930-490e-9926-7509fb356e77 · outbound

This paper cites Not All Disagreement Is Learnable: Token Teachability in On-Policy Distillation.

TREK: Distill to Explore, Reinforce to Refine Not All Disagreement Is Learnable: Token Teachability in On-Policy Distillation

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-07-07T16:24:01.079034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-07T16:22:55.342495Z digest=sha256:30ffeb01007e634e879680ddbb28e3135e684665306a4db523e7effde2a0c7ee

Observation 074eaec9-0a3d-4dea-b54b-31bc0f77c52d · outbound

This paper cites PACED: Distillation and On-Policy Self-Distillation at the Frontier of Student Competence.

TREK: Distill to Explore, Reinforce to Refine PACED: Distillation and On-Policy Self-Distillation at the Frontier of Student Competence

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-07-07T16:24:01.042864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-07T16:22:55.342495Z digest=sha256:72fc7c89feab1f92188870be0bcba32a7478987a41f520a6e25e1c1c6c92df2d

Observation 751d1dbc-e511-49da-ab79-ee3f3978eab6 · outbound

This paper cites ReAct: Synergizing Reasoning and Acting in Language Models.

TREK: Distill to Explore, Reinforce to Refine ReAct: Synergizing Reasoning and Acting in Language Models

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-07-07T16:24:01.097345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-07T16:22:55.342495Z digest=sha256:c40e3174fba61106d41f5501aa515693f8b617a20eeaa7215f93696f2bc2f064

Observation c63a4003-8702-4dab-83f6-e875f8d999fe · outbound

This paper cites Tree of Thoughts: Deliberate Problem Solving with Large Language Models.

TREK: Distill to Explore, Reinforce to Refine Tree of Thoughts: Deliberate Problem Solving with Large Language Models

Reference 7

Resolution
metadata mismatch
local_arxiv, observed 2026-07-07T16:24:00.996759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-07T16:22:55.342495Z digest=sha256:d17ee6c7254700fe3c304bc2f2c11dd5d3abb76418bb99111634ae5647b75a98

Observation ce4c47c0-24b1-47a2-92af-e5d57f20c716 · outbound

This paper cites Qwen3 Technical Report.

TREK: Distill to Explore, Reinforce to Refine Qwen3 Technical Report

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-07-07T16:24:01.047960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-07T16:22:55.342495Z digest=sha256:a98777ac6d217167b2420e4c9037622048be2094ded22294c1f55ee466b27089

Observation 248af4a3-40d6-4158-b03e-eb2f46dacadb · outbound

This paper cites Qwen2.5 Technical Report.

TREK: Distill to Explore, Reinforce to Refine Qwen2.5 Technical Report

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-07-07T16:24:01.085334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-07T16:22:55.342495Z digest=sha256:9a1312755fd205721b7ab36c2bc99cbb984386dff1ba4d4df7f38fe469343c11

Observation a7fba036-a782-4d1e-bef0-4fb3cb71fdeb · outbound

This paper cites ScienceWorld: Is your Agent Smarter than a 5th Grader?.

TREK: Distill to Explore, Reinforce to Refine ScienceWorld: Is your Agent Smarter than a 5th Grader?

Reference 10

Resolution
metadata mismatch
local_arxiv, observed 2026-07-07T16:24:01.073239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-07T16:22:55.342495Z digest=sha256:462557b0f5c85c9ac7254498a3283b837c71457dde601c10bb09853a27cb4dc0

Observation c0781527-49b7-4801-b333-8456bfad9bdf · outbound

This paper cites Learn hard problems during rl with reference guided fine-tuning.

TREK: Distill to Explore, Reinforce to Refine Learn hard problems during rl with reference guided fine-tuning

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-07-07T16:24:01.015130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-07T16:22:55.342495Z digest=sha256:2d9b2fad3d3d1187488c878a2a21143a43433b25f2d54100f7a7677f0430e6dd

Observation f0e7e388-ff58-4425-a764-6195cda5f8ce · outbound

This paper cites Privileged Information Distillation for Language Models.

TREK: Distill to Explore, Reinforce to Refine Privileged Information Distillation for Language Models

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-07-07T16:24:01.067687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-07T16:22:55.342495Z digest=sha256:308bc11047d5539cb5d566a8df1ecc81b031a7376ab4dacac13f5686279c9bff

Observation 7e044e34-5c57-4bce-8caa-255f67e81ed3 · outbound

This paper cites Self-Distillation Zero: Self-Revision Turns Binary Rewards into Dense Supervision.

TREK: Distill to Explore, Reinforce to Refine Self-Distillation Zero: Self-Revision Turns Binary Rewards into Dense Supervision

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-07-07T16:24:01.033844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-07T16:22:55.342495Z digest=sha256:c810f2e0efc590ee9524e00c3d3f14604e955f8a4ac6e0b18f559eb06cfe9bd1

Observation 68754d8a-47c3-4562-9926-71451d911559 · outbound

This paper cites Unifying group-relative and self-distillation policy optimization via sample routing.

TREK: Distill to Explore, Reinforce to Refine Unifying group-relative and self-distillation policy optimization via sample routing

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-07-07T16:24:01.053970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-07T16:22:55.342495Z digest=sha256:bd64b31baafbb1f6299d1bcfee6aaeb088e5a2c1b501ca343fa43bf2023e6c3e

Observation 0e522810-a5e4-4478-a6ae-0e83353208e8 · outbound

This paper cites Uni-OPD: Unifying On-Policy Distillation with a Dual-Perspective Recipe.

TREK: Distill to Explore, Reinforce to Refine Uni-OPD: Unifying On-Policy Distillation with a Dual-Perspective Recipe

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-07-07T16:24:01.103354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-07T16:22:55.342495Z digest=sha256:bf6631ff16cd467cbab6d389f78909b712056d34841c2a5c7944f2fdeb9b2a55

Observation e49b3531-1c2d-4a64-a419-f497a9e0e943 · outbound

This paper cites Rethinking On-Policy Distillation of Large Language Models: Phenomenology, Mechanism, and Recipe.

TREK: Distill to Explore, Reinforce to Refine Rethinking On-Policy Distillation of Large Language Models: Phenomenology, Mechanism, and Recipe

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-07-07T16:24:01.037165Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-07T16:22:55.342495Z digest=sha256:484cd01bc3f414ac95cdfcae5b749ae66c4451022da8c1f2c0bfe183c94f21d6

Observation f84730e4-135e-4d18-ac39-ede5209b6f0a · outbound

This paper cites Black-box on-policy distillation of large language models.arXiv preprint.

TREK: Distill to Explore, Reinforce to Refine Black-box on-policy distillation of large language models.arXiv preprint

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-07-07T16:24:01.023987Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-07T16:22:55.342495Z digest=sha256:f914b1a24ce281a10f5a8744feec8a9aa6acc9f7e0472e62faff766aed12e348

Observation b9e19399-1ae3-4466-ab36-2a952f75ea88 · outbound

This paper cites Self-Refine: Iterative Refinement with Self-Feedback.

TREK: Distill to Explore, Reinforce to Refine Self-Refine: Iterative Refinement with Self-Feedback

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-07-07T16:24:01.030689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-07T16:22:55.342495Z digest=sha256:9076c976a23ac3222450e27ffe7db10a60c0626b6c92c18a063f96dbf1c79efa

Observation b249f26c-14a1-4331-a12d-4ee15eaae7d8 · outbound

This paper cites Let's Verify Step by Step.

TREK: Distill to Explore, Reinforce to Refine Let's Verify Step by Step

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-07-07T16:24:01.000144Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-07T16:22:55.342495Z digest=sha256:fc67b6905b842bf419695dc645dd3f8957420ebe89fcf495a223e9c0435ef0d9

Observation 5dba23c9-7194-410c-8aa0-94905afb19d5 · outbound

This paper cites Large Language Models Can Self-Improve.

TREK: Distill to Explore, Reinforce to Refine Large Language Models Can Self-Improve

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-07-07T16:24:01.011439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-07T16:22:55.342495Z digest=sha256:9e7092767d35c9652ecfe084bb90ebf5852ebd6cd8431836f7f786aece7c2042

Observation 4edc7a00-cbb0-46c6-bd20-dbf8595f699c · outbound

This paper cites Explanations from Large Language Models Make Small Reasoners Better.

TREK: Distill to Explore, Reinforce to Refine Explanations from Large Language Models Make Small Reasoners Better

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-07-07T16:24:01.020296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-07T16:22:55.342495Z digest=sha256:0b3a3988c9cb3f9f9a8a3a4cc6cc0bdfc8443b9a00f1b489024610e999d99414

Observation c88f63a8-01cb-4ef6-a526-ebeaa3148fb1 · outbound

This paper cites Distilling step-by-step! Outperforming larger language models with less training data and smaller model sizes.

TREK: Distill to Explore, Reinforce to Refine Distilling step-by-step! Outperforming larger language models with less training data and smaller model sizes

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T16:24:01.109437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-07T16:22:55.342495Z digest=sha256:2259b3b220fbf538c660cec599a4d12aa5963776fbaa4985f9ea725a0669bff0

Observation d0388e6d-e7ad-4743-8324-16ab7979439e · outbound

This paper cites Knowledge Distillation with Training Wheels.

TREK: Distill to Explore, Reinforce to Refine Knowledge Distillation with Training Wheels

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-07-07T16:24:00.992883Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-07T16:22:55.342495Z digest=sha256:6139abdffbd6608daf4a4d0e686da9b15243e2135db00a6f79eb40146b705806

Observation 95e4f408-5e2b-441d-b11d-8da9c6048c83 · outbound

This paper cites Codes: A context-efficient framework for enhancing small language models via domain-specific adaptation and model ensembling.

TREK: Distill to Explore, Reinforce to Refine Codes: A context-efficient framework for enhancing small language models via domain-specific adaptation and model ensembling

Reference 24

Resolution
metadata mismatch
arxiv_id, observed 2026-07-07T16:24:00.985351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-07T16:22:55.342495Z digest=sha256:809ae45bde1f798173a06ba48099b79cbf30d455af2fd2cb99d62674962e960c

Observation 811c02bb-30bc-4395-9d72-89f511531161 · outbound

This paper cites DriftGuard: Safety-Aware Multi-Monitor Detection and Selective Adaptation for Evolving Toxicity Moderation.

TREK: Distill to Explore, Reinforce to Refine DriftGuard: Safety-Aware Multi-Monitor Detection and Selective Adaptation for Evolving Toxicity Moderation

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-07-07T16:24:01.090823Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-07T16:22:55.342495Z digest=sha256:41b49a8b075f2509933c017d7d9f96134ce6f557c939ee51391d0b9904260fab

Observation 49cc97da-1f47-49f0-b1ac-383d63a66841 · outbound

This paper cites Reinforcement-aware Knowledge Distillation for LLM Reasoning.

TREK: Distill to Explore, Reinforce to Refine Reinforcement-aware Knowledge Distillation for LLM Reasoning

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-07-07T16:24:01.060405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-07T16:22:55.342495Z digest=sha256:43c92e2509ae698e4f403b23ef0946bd6dc61f109e7742540f61f5f84e093b8f

Observation 4e3e0f65-0415-4218-a492-0faf2c693271 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

TREK: Distill to Explore, Reinforce to Refine DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-07-07T16:24:01.026792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-07T16:22:55.342495Z digest=sha256:11e336d05bf585ee0c42372db571b9b01cb6dc9bcbac21451704dbed0217635b

Observation 82644a5a-4930-4ea6-9580-ee56a8ce3e35 · outbound

This paper cites Group Sequence Policy Optimization.

TREK: Distill to Explore, Reinforce to Refine Group Sequence Policy Optimization

Reference 28

Resolution
metadata mismatch
local_arxiv, observed 2026-07-07T16:24:01.008182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-07T16:22:55.342495Z digest=sha256:d73105a4dbcc1b811ae4d9f432a343915b090bf608bc2ca423d579ba6dbee199

Observation cd089bc9-d9f3-4a1e-8fc8-79639bc469e3 · outbound

This paper cites an unresolved cited work.

TREK: Distill to Explore, Reinforce to Refine Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-07-07T16:24:01.119416Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-07T16:22:55.342495Z digest=sha256:7aa2de2eed7e695d03ec2abb2c2421454c496c9eb228a52e8baac72da5250be8

Pith citing papers

No inbound Pith citation observations are available.