Pith. sign in

Paper Citation Record · LEDGER

SOD: Step-wise On-policy Distillation for Small Language Model Agents

As of 1 August 2026, this Paper Citation Record lists 81 of 81 outbound references and 5 inbound Pith citation observations for arXiv:2605.07725.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.07725 v1

Coverage vector

measured 81 of 81 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-11T02:25:59.056181Z

measured 86 of 86 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-01T06:32:01.292127+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-31T18:36:19.656189Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-09T00:35:48.710213Z

Reference resolution

81 of 81 outbound references displayed

  • verified exact60
  • verified fuzzy19
  • unresolved1
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 387058b7-0712-420c-b4f0-57f4cc8c728a · outbound

This paper cites The Rise and Potential of Large Language Model Based Agents: A Survey.

SOD: Step-wise On-policy Distillation for Small Language Model Agents The Rise and Potential of Large Language Model Based Agents: A Survey

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:48:03.736032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-11T02:25:59.056181Z digest=sha256:71f6035fe62233c5b4c488ddd04434130106fa7fe08144f8f4c156762f37dbdf

Observation beb73551-eb2d-4192-8626-c56a05dfec20 · outbound

This paper cites Distilling llm agent into small models with retrieval and code tools.

SOD: Step-wise On-policy Distillation for Small Language Model Agents Distilling llm agent into small models with retrieval and code tools

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:40:54.392401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-11T02:25:59.056181Z digest=sha256:93b9b2d43a2e0a48bff88cf142b723cd9f30cf76d1f4daec52b5ed64665a8e5c

Observation df63bdf5-522a-4aa1-810b-3571257e9a95 · outbound

This paper cites Narasimhan, and Yuan Cao.

SOD: Step-wise On-policy Distillation for Small Language Model Agents Narasimhan, and Yuan Cao

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T12:40:17.356444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-11T02:25:59.056181Z digest=sha256:e2d90c82eb04ad9f6d0f94c6b132e7de2ea038edac35cce97b184039535888a6

Observation 23c89fc3-b040-45c0-b960-65ee38d9a09a · outbound

This paper cites Toolformer: Language models can teach themselves to use tools.

SOD: Step-wise On-policy Distillation for Small Language Model Agents Toolformer: Language models can teach themselves to use tools

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T12:40:17.361447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-11T02:25:59.056181Z digest=sha256:65bf6b659857aa5859f1667a1be84ada496674ac91ad936151d234e65add5f40

Observation 4c94ae1f-2a38-492b-9b90-cf5594b2b0f7 · outbound

This paper cites Agentic Reasoning and Tool Integration for LLMs via Reinforcement Learning.

SOD: Step-wise On-policy Distillation for Small Language Model Agents Agentic Reasoning and Tool Integration for LLMs via Reinforcement Learning

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-18T02:58:56.460093Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-11T02:25:59.056181Z digest=sha256:52d7348a19420fb68d9d2eb420be573ccc144f06e0ba22383ecb3557639ea62d

Observation 855ef61a-7595-49f8-b181-799fd44a9543 · outbound

This paper cites Mixed distillation helps smaller language models reason better.

SOD: Step-wise On-policy Distillation for Small Language Model Agents Mixed distillation helps smaller language models reason better

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T12:40:17.284002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-11T02:25:59.056181Z digest=sha256:d70e0c82b6b194326dd92786b9d9ea894cfa8cffd541d2b03db0f1ea080bd0b9

Observation 29e19b0a-d9ad-4791-bc49-3f8139fcb030 · outbound

This paper cites AgentProcessBench: Diagnosing Step-Level Process Quality in Tool-Using Agents.

SOD: Step-wise On-policy Distillation for Small Language Model Agents AgentProcessBench: Diagnosing Step-Level Process Quality in Tool-Using Agents

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-06-02T03:04:40.253850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-11T02:25:59.056181Z digest=sha256:9ec1a82cd205dc099677fd1e31448da1f27d9de15bcda9a57ed3409b360b51e3

Observation 11cae7de-8f7d-42f5-9db1-2518e889257a · outbound

This paper cites On-Device Language Models: A Comprehensive Review.

SOD: Step-wise On-policy Distillation for Small Language Model Agents On-Device Language Models: A Comprehensive Review

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:40:54.363916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-11T02:25:59.056181Z digest=sha256:fceaeca485cc850723c01ebc21f35b83050df7842d1e303978f43d601baf70f9

Observation 5874f411-31f7-41df-9fe7-93bde841b2df · outbound

This paper cites A Survey on Knowledge Distillation of Large Language Models.

SOD: Step-wise On-policy Distillation for Small Language Model Agents A Survey on Knowledge Distillation of Large Language Models

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-17T23:31:12.027711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-11T02:25:59.056181Z digest=sha256:36a870084d6d91e8efc49062e0b5869dd3eec2e001cf722dd068ab5056205568

Observation 78e615e7-2449-4ac4-aa66-e2da5269144a · outbound

This paper cites AgentDistill: Training-Free Agent Distillation with Generalizable MCP Boxes.

SOD: Step-wise On-policy Distillation for Small Language Model Agents AgentDistill: Training-Free Agent Distillation with Generalizable MCP Boxes

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:40:54.398511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-11T02:25:59.056181Z digest=sha256:d4980d32adf22ff61af3a5200405990852e37682f8a73b0f65d9ffe094d6dc4a

Observation 926ac0a8-bc32-4a68-91f0-e76bcab2e798 · outbound

This paper cites O-researcher: An open ended deep research model via multi-agent distillation and agentic rl.

SOD: Step-wise On-policy Distillation for Small Language Model Agents O-researcher: An open ended deep research model via multi-agent distillation and agentic rl

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:40:54.409941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-11T02:25:59.056181Z digest=sha256:059c0231213c51c6772ee507dd696cbc7d43605f62d5b92c3483d76d20e5a74e

Observation 1e23cc51-8da2-4836-a10b-219e68569a74 · outbound

This paper cites Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL.

SOD: Step-wise On-policy Distillation for Small Language Model Agents Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:40:54.388059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-11T02:25:59.056181Z digest=sha256:9e5733c6b92bc3d30317a298e6c55b623ad2fc2b21565389a172d281d5ae728d

Observation d5a8dc9d-55e9-4b10-9237-374ab7b93e3c · outbound

This paper cites ToolRL: Reward is All Tool Learning Needs.

SOD: Step-wise On-policy Distillation for Small Language Model Agents ToolRL: Reward is All Tool Learning Needs

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-14T00:26:48.594660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-11T02:25:59.056181Z digest=sha256:e4a2689b0ff106027ffd4ccef96627de67d696d5cf3358e8b28139b2025574bc

Observation 2eb7cbae-7bad-446d-a269-775a3c3e9e82 · outbound

This paper cites Replacing thinking with tool usage enables reasoning in small language models.

SOD: Step-wise On-policy Distillation for Small Language Model Agents Replacing thinking with tool usage enables reasoning in small language models

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:40:54.443222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-11T02:25:59.056181Z digest=sha256:c86154b8d5c4cc57171ddf8fbc28b0767b512fe3ade2bb3550805da4c3816e97

Observation 31f4b51b-df37-4077-ba13-c1f579eaefd2 · outbound

This paper cites SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning.

SOD: Step-wise On-policy Distillation for Small Language Model Agents SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:30:58.767379Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-11T02:25:59.056181Z digest=sha256:783bc56ef563c92924c0cf7e177977807347751ae6b0bd7f893d0613c2b0229d

Observation f89fdfb4-7e0c-438f-bbc0-a5aedf41b9c1 · outbound

This paper cites Structured Agent Distillation for Large Language Model.

SOD: Step-wise On-policy Distillation for Small Language Model Agents Structured Agent Distillation for Large Language Model

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-28T03:04:41.842017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-11T02:25:59.056181Z digest=sha256:9b5e216567d85757812ddc97c1bfec09653301784c46d0529076914a4b648d4d

Observation 2001a2f7-cab5-4e06-993e-153caaf1c8bb · outbound

This paper cites ToRL: Scaling Tool-Integrated RL.

SOD: Step-wise On-policy Distillation for Small Language Model Agents ToRL: Scaling Tool-Integrated RL

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:30:58.761107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-11T02:25:59.056181Z digest=sha256:9bd8389c2d257ad9bf843758aa50bdc9dc2e8db5956bbfe3fb167883a66f4cc2

Observation beef77d6-6fe2-43eb-b68a-87c2f95821bd · outbound

This paper cites Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning.

SOD: Step-wise On-policy Distillation for Small Language Model Agents Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:40:54.416063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-11T02:25:59.056181Z digest=sha256:4283054df7842ae2c856322dfeed17401d7f9dda4647967d5eab54763fa831b4

Observation 7686b27a-49c1-479a-a0b6-312c162e2e0e · outbound

This paper cites R1-Searcher: Incentivizing the Search Capability in LLMs via Reinforcement Learning.

SOD: Step-wise On-policy Distillation for Small Language Model Agents R1-Searcher: Incentivizing the Search Capability in LLMs via Reinforcement Learning

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-13T18:37:17.866238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-11T02:25:59.056181Z digest=sha256:00626ff511cb29098678ea9e2cd3f0cb60ea695cb3825b23e543d57de84b361a

Observation 23f75ed2-d476-4844-9b51-79f526810541 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

SOD: Step-wise On-policy Distillation for Small Language Model Agents DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-05-11T03:40:54.422229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-11T02:25:59.056181Z digest=sha256:868be694d44596878e074cfb84cf5a126265014391a54282a909b6d7cc2248db

Observation 3eecbc68-9747-4ded-a5fd-07abd909f807 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

SOD: Step-wise On-policy Distillation for Small Language Model Agents DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-05-11T03:30:58.717508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-11T02:25:59.056181Z digest=sha256:caa516b33bf315174ede542211c060620b6c26a79aed0f977e41538cdf1130d9

Observation 47e026fd-5b20-462c-8edd-9895967bc0b6 · outbound

This paper cites KEPO: Knowledge-Enhanced Preference Optimization for Multimodal Reasoning with Applications to Medical VQA.

SOD: Step-wise On-policy Distillation for Small Language Model Agents KEPO: Knowledge-Enhanced Preference Optimization for Multimodal Reasoning with Applications to Medical VQA

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:45:20.706329Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-11T02:25:59.056181Z digest=sha256:d75e8ccc2c27ec30852192589e7a81db08ab9c1dfb10520323c2242322d30b56

Observation 7b6694bd-55df-4c1c-bba9-4116d3cfb541 · outbound

This paper cites On-policy distillation of language models: Learning from self-generated mistakes.

SOD: Step-wise On-policy Distillation for Small Language Model Agents On-policy distillation of language models: Learning from self-generated mistakes

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T12:40:17.331988Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-11T02:25:59.056181Z digest=sha256:b8ca43101487cc92b9d9a8c55006419437398ee58936c4a57f87dd35cf8e1cb0

Observation 9a63733c-d9bd-4e83-88dc-acf17807d0ee · outbound

This paper cites Minillm: Knowledge distillation of large language models.

SOD: Step-wise On-policy Distillation for Small Language Model Agents Minillm: Knowledge distillation of large language models

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T12:40:17.287809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-11T02:25:59.056181Z digest=sha256:ae54a0017fb080bba999e07dda211bd030859f13afb2480e4be200f1578cc2fe

Observation d5d8abe9-5ccc-44ed-b3de-de977da1bad6 · outbound

This paper cites Entropy-Aware On-Policy Distillation of Language Models.

SOD: Step-wise On-policy Distillation for Small Language Model Agents Entropy-Aware On-Policy Distillation of Language Models

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-25T03:02:00.585154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-11T02:25:59.056181Z digest=sha256:7a4ddf4f870bcb11e7a96f39f5f41bbf62e3a98ca36ed22e106ab996557356d8

Observation f5126f51-8d44-4e58-a339-d567eb72a11c · outbound

This paper cites Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph E.

SOD: Step-wise On-policy Distillation for Small Language Model Agents Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph E

Reference 26

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T03:30:58.700899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-11T02:25:59.056181Z digest=sha256:f01355af48e09920272da072e0118c5722b24527d7596efae5c193d74c034dfa

Observation cb307289-c7a3-4025-b26e-a7473d5e8e3e · outbound

This paper cites Lightning OPD: Efficient Post-Training for Large Reasoning Models with Offline On-Policy Distillation.

SOD: Step-wise On-policy Distillation for Small Language Model Agents Lightning OPD: Efficient Post-Training for Large Reasoning Models with Offline On-Policy Distillation

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-05-11T03:30:58.645663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-11T02:25:59.056181Z digest=sha256:30abeba6f19e85297fe3c7007a04d26bc5aca22c15e32d5ed862fb07660fed3c

Observation edc8b8d2-f5fe-4ff9-8217-b0442b85d68f · outbound

This paper cites Qwen3 Technical Report.

SOD: Step-wise On-policy Distillation for Small Language Model Agents Qwen3 Technical Report

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-05-11T03:30:58.637592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-11T02:25:59.056181Z digest=sha256:f0f37f0aae88be21d1d9ce71dee532993a334093f0bfd991033e0db3db28915e

Observation 23a42cb7-8442-4368-b85d-fee896d4f1d3 · outbound

This paper cites Rethinking On-Policy Distillation of Large Language Models: Phenomenology, Mechanism, and Recipe.

SOD: Step-wise On-policy Distillation for Small Language Model Agents Rethinking On-Policy Distillation of Large Language Models: Phenomenology, Mechanism, and Recipe

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-05-11T03:30:58.651270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-11T02:25:59.056181Z digest=sha256:4d737c0d900fd0d13d3ff0dee9b57fb6afae3fe4a50b7a053689233015d22f78

Observation 72d33e95-19a6-408b-877d-0fd22a7d21d2 · outbound

This paper cites Rlkd: Distilling llms’ reasoning via reinforcement learning.

SOD: Step-wise On-policy Distillation for Small Language Model Agents Rlkd: Distilling llms’ reasoning via reinforcement learning

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T12:45:15.823707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-11T02:25:59.056181Z digest=sha256:1b89f8d6ab7140d548d6a3f1dc12c28134205c4e256d71108d59bf278759a149

Observation 9da3d397-d60f-4dad-b43a-2b0d990d1715 · outbound

This paper cites Unifying group-relative and self-distillation policy optimization via sample routing.

SOD: Step-wise On-policy Distillation for Small Language Model Agents Unifying group-relative and self-distillation policy optimization via sample routing

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:30:58.659483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-11T02:25:59.056181Z digest=sha256:5574092b9b29e48e7cfb461a76b328f1d255ca2a399141302461f9e74e8002a6

Observation 9fd0ddf1-23d1-4275-a3a9-da68959a41de · outbound

This paper cites Self-Distilled RLVR.

SOD: Step-wise On-policy Distillation for Small Language Model Agents Self-Distilled RLVR

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-05-11T03:30:58.671489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-11T02:25:59.056181Z digest=sha256:316e8cc8aa68d135f44764e772f35ff381080dfe499812cbe7ca7c1c6523c989

Observation cd7b84f1-02bd-47d0-bc5d-1a4ac02fd924 · outbound

This paper cites VOLD: Reasoning Transfer from LLMs to Vision-Language Models via On-Policy Distillation.

SOD: Step-wise On-policy Distillation for Small Language Model Agents VOLD: Reasoning Transfer from LLMs to Vision-Language Models via On-Policy Distillation

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-06-05T01:14:30.340920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-11T02:25:59.056181Z digest=sha256:d7d80d2a2027388f84ab983c4dd9bd54ccaa340ecfd29b02d7d67a20de3e1ff5

Observation aa642c6a-7684-46d0-9e09-54e20b7bdac5 · outbound

This paper cites OpenClaw-RL: Train Any Agent Simply by Talking.

SOD: Step-wise On-policy Distillation for Small Language Model Agents OpenClaw-RL: Train Any Agent Simply by Talking

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:30:58.731803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-11T02:25:59.056181Z digest=sha256:bce98fb418631e11c3db0d8b1df49d477fe4ccbb710251ad7f37f161b01893c1

Observation 8c64a382-1919-40d1-9c8d-c6d215e01efd · outbound

This paper cites Revisiting On-Policy Distillation: Empirical Failure Modes and Simple Fixes.

SOD: Step-wise On-policy Distillation for Small Language Model Agents Revisiting On-Policy Distillation: Empirical Failure Modes and Simple Fixes

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-05-11T03:30:58.665532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-11T02:25:59.056181Z digest=sha256:a93f36821294fb26a5acae293b29ba387ab818ba6c41ac068122a169181a9ad4

Observation 4e0097b7-59c4-4418-ba8d-b16f8aec5f5e · outbound

This paper cites A Survey of On-Policy Distillation for Large Language Models.

SOD: Step-wise On-policy Distillation for Small Language Model Agents A Survey of On-Policy Distillation for Large Language Models

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-13T02:47:19.238105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-11T02:25:59.056181Z digest=sha256:b742c59badf11ddf2772f569dbb3ec2132c8083f584ce160eb666b7580079285

Observation 64e62c2a-e126-4abb-9d5e-164b900f48f8 · outbound

This paper cites Gordon, and Drew Bagnell.

SOD: Step-wise On-policy Distillation for Small Language Model Agents Gordon, and Drew Bagnell

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T12:45:15.818675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-11T02:25:59.056181Z digest=sha256:fa490d871b2f87a06ddd04d2e581049712b12c5163e9d04d3d0e44c2193674dd

Observation 0f6b9d9e-c074-43df-a60f-7103d8d1622b · outbound

This paper cites The False Promise of Imitating Proprietary LLMs.

SOD: Step-wise On-policy Distillation for Small Language Model Agents The False Promise of Imitating Proprietary LLMs

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-18T06:54:31.902059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-11T02:25:59.056181Z digest=sha256:4f60e8e0c8b2719733344bcbfac0e9ebc214252d718c402e1714d8a3e921313a

Observation e21d29eb-c765-4c6b-87e3-d433cc682f34 · outbound

This paper cites TCOD: Exploring Temporal Curriculum in On-Policy Distillation for Multi-turn Autonomous Agents.

SOD: Step-wise On-policy Distillation for Small Language Model Agents TCOD: Exploring Temporal Curriculum in On-Policy Distillation for Multi-turn Autonomous Agents

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-05-11T03:30:58.690400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-11T02:25:59.056181Z digest=sha256:d503737eb44c548f9eb5db7ee2a7250e0f1db19b215134d48f29f2c56f94f54f

Observation 399b77e4-ff4b-421f-88c7-caa536af039e · outbound

This paper cites Stable On-Policy Distillation through Adaptive Target Reformulation.

SOD: Step-wise On-policy Distillation for Small Language Model Agents Stable On-Policy Distillation through Adaptive Target Reformulation

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-05-11T03:40:54.452232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-11T02:25:59.056181Z digest=sha256:bce81ac33127c7f6017074895983794b2899f3cb0c971e13cfecf52b0710c02b

Observation cfbe5751-f9fc-473e-91ad-8609a9615974 · outbound

This paper cites TIP: Token Importance in On-Policy Distillation.

SOD: Step-wise On-policy Distillation for Small Language Model Agents TIP: Token Importance in On-Policy Distillation

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-05-11T03:40:54.419097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-11T02:25:59.056181Z digest=sha256:4a88bbcbeb8b6a34a41906333ce2944904c52feaae0a889eb65f54922299ff86

Observation e32f4091-39c5-48bb-9eeb-20851908a470 · outbound

This paper cites Deep reinforcement learning from human preferences.Advances in neural information processing systems, 30.

SOD: Step-wise On-policy Distillation for Small Language Model Agents Deep reinforcement learning from human preferences.Advances in neural information processing systems, 30

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T12:45:15.814230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-11T02:25:59.056181Z digest=sha256:416c09715ee82dffb6a94eb0881961f5a6a3aa591f68c5a8436325da8ef15a6e

Observation de8e4467-84a9-432f-ade2-78b797f71d54 · outbound

This paper cites Fine-Tuning Language Models from Human Preferences.

SOD: Step-wise On-policy Distillation for Small Language Model Agents Fine-Tuning Language Models from Human Preferences

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-05-11T03:40:54.456154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-11T02:25:59.056181Z digest=sha256:f40b9c6c6f993b38ddb58d28d9ff39a195b44d8f0a4d8885446d4c7cb04651c7

Observation 5d9a04a2-6c28-439c-bd9b-52794975e857 · outbound

This paper cites Proximal Policy Optimization Algorithms.

SOD: Step-wise On-policy Distillation for Small Language Model Agents Proximal Policy Optimization Algorithms

Reference 44

Resolution
verified exact
local_arxiv, observed 2026-05-11T03:40:54.481179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-11T02:25:59.056181Z digest=sha256:74733af3b9cda989c3b99dca5b53f04bf13013e1d7455536dd7c5bff98974a41

Observation 414876f4-f394-4d5f-ae80-fb49c6d560d0 · outbound

This paper cites FireAct: Toward Language Agent Fine-tuning.

SOD: Step-wise On-policy Distillation for Small Language Model Agents FireAct: Toward Language Agent Fine-tuning

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:40:54.506607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-11T02:25:59.056181Z digest=sha256:f0e4e76917bb1e5ca0fa9f276228305cc51d939858ba0f043df3e9e3a7da0d74

Observation 9c81871f-9215-4d1e-8c71-5d595e9883ca · outbound

This paper cites Swe-agent: Agent-computer interfaces enable automated software engineering.Advances in Neural Information Processing Systems, 37:50528–50652.

SOD: Step-wise On-policy Distillation for Small Language Model Agents Swe-agent: Agent-computer interfaces enable automated software engineering.Advances in Neural Information Processing Systems, 37:50528–50652

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T12:45:15.809847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-11T02:25:59.056181Z digest=sha256:6764963b34efaa00c2b39a86898f569ca3698d193a12da2cddcd21bb9b73f993

Observation a77abe9c-843d-4664-a59f-8614182031ed · outbound

This paper cites ReTool: Reinforcement Learning for Strategic Tool Use in LLMs.

SOD: Step-wise On-policy Distillation for Small Language Model Agents ReTool: Reinforcement Learning for Strategic Tool Use in LLMs

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-13T18:42:39.160125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-11T02:25:59.056181Z digest=sha256:73c29192f9ebbfe0a9296ba06f33eef0fe8c9aa5e53c65d7730146b999cd64b7

Observation 4efebc59-7037-405d-8524-6744e16f0995 · outbound

This paper cites Digirl: Training in-the-wild device-control agents with autonomous reinforcement learning.

SOD: Step-wise On-policy Distillation for Small Language Model Agents Digirl: Training in-the-wild device-control agents with autonomous reinforcement learning

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T12:40:17.366234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-11T02:25:59.056181Z digest=sha256:799631f9fd629fadc2918d1df0117ff31d59a514b4bbb97790a8482ca6618c18

Observation 11ce60e9-4099-45ee-8ddb-43cd23bb6da5 · outbound

This paper cites WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning.

SOD: Step-wise On-policy Distillation for Small Language Model Agents WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:40:54.542352Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-11T02:25:59.056181Z digest=sha256:a2955ae71a91dbd7ec77385fe141ba4acc96fd20e4a5b7e2ec5d8a9a973c7be0

Observation 830e555e-e74d-40ef-84fb-4290fa4652ad · outbound

This paper cites ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL.

SOD: Step-wise On-policy Distillation for Small Language Model Agents ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:40:54.475555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-11T02:25:59.056181Z digest=sha256:1c53571bc381ee1164bb3e20507b50d0f419c500efbcb1d118fb1e9c8c25fb94

Observation f4304d3e-0096-4a0d-9b56-6c8036d8642a · outbound

This paper cites Reinforcement Learning for Long-Horizon Interactive LLM Agents.

SOD: Step-wise On-policy Distillation for Small Language Model Agents Reinforcement Learning for Long-Horizon Interactive LLM Agents

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:40:54.478495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-11T02:25:59.056181Z digest=sha256:bf15954fdb554c1c2d4798aac5afd2441c9ba24b374b12534bd82fc6800d09f2

Observation 73c0740e-bc66-468b-a498-943e5494ab26 · outbound

This paper cites Demystifying reinforcement learning in agentic reasoning.

SOD: Step-wise On-policy Distillation for Small Language Model Agents Demystifying reinforcement learning in agentic reasoning

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:40:54.496029Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-11T02:25:59.056181Z digest=sha256:c96cb7573ca6651a5a2b8a78c9b5ade326be655cda542cc1143753a5b826398f

Observation 2cc8d805-46bc-4f27-8e89-1c239d8a5be1 · outbound

This paper cites Rlanything: Forge environment, policy, and reward model in completely dynamic rl system.

SOD: Step-wise On-policy Distillation for Small Language Model Agents Rlanything: Forge environment, policy, and reward model in completely dynamic rl system

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:40:54.488729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-11T02:25:59.056181Z digest=sha256:d223dbad58addf617225c01306011f0a5daf3175513b34394acbf394707f6377

Observation 8335d5dd-2469-4c5f-a4ea-591b10668fca · outbound

This paper cites Co-evolving llm coder and unit tester via reinforcement learning.arXiv preprint arXiv:2506.03136.

SOD: Step-wise On-policy Distillation for Small Language Model Agents Co-evolving llm coder and unit tester via reinforcement learning.arXiv preprint arXiv:2506.03136

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:40:54.514453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-11T02:25:59.056181Z digest=sha256:5586a22a79cba59c542aa21d157ffcf3d9aeeb1fc7bd2f179b551abccccecc60

Observation 1df2d2f0-9e89-40ad-8337-899c60069ee7 · outbound

This paper cites On-Policy Context Distillation for Language Models.

SOD: Step-wise On-policy Distillation for Small Language Model Agents On-Policy Context Distillation for Language Models

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-05-13T20:48:41.492604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-11T02:25:59.056181Z digest=sha256:157cb95f82d510ca1db22ab9e9c043083d53af29c26ed149418e4df2f5fc91a8

Observation 5c1881bd-45b1-464c-a3b3-e9b09c88dcc2 · outbound

This paper cites Black-box on-policy distillation of large language models.arXiv preprint.

SOD: Step-wise On-policy Distillation for Small Language Model Agents Black-box on-policy distillation of large language models.arXiv preprint

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:40:54.551709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-11T02:25:59.056181Z digest=sha256:e123c4d26196399a8f79f9f355118d68e03b591cdc73a1a670bc5965075d0885

Observation 81440573-cb7e-4fe5-b359-ad3e7c5a5fca · outbound

This paper cites Hybrid Policy Distillation for LLMs.

SOD: Step-wise On-policy Distillation for Small Language Model Agents Hybrid Policy Distillation for LLMs

Reference 57

Resolution
verified exact
local_arxiv, observed 2026-05-11T03:40:54.531672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-11T02:25:59.056181Z digest=sha256:e5e31ac77292c6df952a0339d9c3282d307772aef49f757ee671578c3a3962c8

Observation 1a1229c6-63d0-4ef8-a2e8-ef3f7b015bc5 · outbound

This paper cites SCOPE: Signal-Calibrated On-Policy Distillation Enhancement with Dual-Path Adaptive Weighting.

SOD: Step-wise On-policy Distillation for Small Language Model Agents SCOPE: Signal-Calibrated On-Policy Distillation Enhancement with Dual-Path Adaptive Weighting

Reference 58

Resolution
verified exact
local_arxiv, observed 2026-05-11T03:40:54.538197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-11T02:25:59.056181Z digest=sha256:ec1a45dedf0d33fd02d015e37b83ca16fa32169f2c48c3330b373351763caa76

Observation 019ee0cb-692a-4c8c-8c7b-cc7f51187955 · outbound

This paper cites SODA: Semi On-Policy Black-Box Distillation for Large Language Models.

SOD: Step-wise On-policy Distillation for Small Language Model Agents SODA: Semi On-Policy Black-Box Distillation for Large Language Models

Reference 59

Resolution
verified exact
local_arxiv, observed 2026-05-11T03:40:54.517741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-11T02:25:59.056181Z digest=sha256:fb43be70a71634916744cf14cb3db6346a8aa1d6f4882dc20e26f670524bddea

Observation 61b0c8b4-a0e2-4045-8dc7-d3430afc24cd · outbound

This paper cites Learning beyond Teacher: Generalized On-Policy Distillation with Reward Extrapolation.

SOD: Step-wise On-policy Distillation for Small Language Model Agents Learning beyond Teacher: Generalized On-Policy Distillation with Reward Extrapolation

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-05-16T05:11:07.754859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-11T02:25:59.056181Z digest=sha256:3c45db9db2b1dbeb5a1f74723e6d4e27c582db5bf35b0693b0291b21e8b1fbdc

Observation 28f76bcb-17ac-48ef-a40e-23b074abbde9 · outbound

This paper cites Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models.

SOD: Step-wise On-policy Distillation for Small Language Model Agents Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models

Reference 61

Resolution
verified exact
arxiv_id, observed 2026-05-12T03:54:31.187112Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-11T02:25:59.056181Z digest=sha256:d491c28b11476b3ce3d194b3da7fe4fde5b5e0851bdcbea90720197c3dcee97e

Observation b1e518f7-5015-4e73-8957-0c918aa3f206 · outbound

This paper cites Privileged Information Distillation for Language Models.

SOD: Step-wise On-policy Distillation for Small Language Model Agents Privileged Information Distillation for Language Models

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-05-22T08:25:38.085126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-11T02:25:59.056181Z digest=sha256:3d9b4be7c99eddc4211a5d936dde1de037c1b73a2bd2265f180e60bd300d4e29

Observation 0dce8468-16a9-4efd-822f-7d08368f4c1f · outbound

This paper cites Reinforcement Learning via Self-Distillation.

SOD: Step-wise On-policy Distillation for Small Language Model Agents Reinforcement Learning via Self-Distillation

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-05-12T04:29:18.795354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-11T02:25:59.056181Z digest=sha256:fe9eca9116209af2e38d4b8cb54e1a1ddeffa8c73adc367e33221b59efc55e0d

Observation 5e3f0643-4f42-40f2-b21f-cade4f97822e · outbound

This paper cites Self-Distillation Enables Continual Learning.

SOD: Step-wise On-policy Distillation for Small Language Model Agents Self-Distillation Enables Continual Learning

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-05-12T05:31:52.723793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-11T02:25:59.056181Z digest=sha256:ff3ca96b00be6bba53d85727eb6a1a27c541efde8a448bdef67239e07acc5536

Observation 0d8b62f4-f970-43a0-8743-3cf917334732 · outbound

This paper cites Self-Distillation Zero: Self-Revision Turns Binary Rewards into Dense Supervision.

SOD: Step-wise On-policy Distillation for Small Language Model Agents Self-Distillation Zero: Self-Revision Turns Binary Rewards into Dense Supervision

Reference 65

Resolution
verified exact
local_arxiv, observed 2026-05-11T03:40:54.563012Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-11T02:25:59.056181Z digest=sha256:057d7c4533be6a89bf75beeec799854293e26ce9638dc4980d9edabccbe55a96

Observation 7e4dbbc3-1559-41b4-b645-4368493b19a8 · outbound

This paper cites Skill-SD: Skill-Conditioned Self-Distillation for Multi-turn LLM Agents.

SOD: Step-wise On-policy Distillation for Small Language Model Agents Skill-SD: Skill-Conditioned Self-Distillation for Multi-turn LLM Agents

Reference 66

Resolution
verified exact
local_arxiv, observed 2026-05-11T03:40:54.463182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-11T02:25:59.056181Z digest=sha256:d1d523eb9a66525b4e0d26256eeac4f3395e92c0048e29768a25eca2a86838cc

Observation 8f63b949-d0fc-410e-acb1-7657a83b2a75 · outbound

This paper cites On-policy distillation.Thinking Machines Lab: Con- nectionism.

SOD: Step-wise On-policy Distillation for Small Language Model Agents On-policy distillation.Thinking Machines Lab: Con- nectionism

Reference 67

Resolution
verified exact
doi, observed 2026-05-11T02:30:53.275162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T07:08:02.609025+00:00.

source=pdf_text observed=2026-05-11T02:25:59.056181Z digest=sha256:49d7779a84033138f7f787bf897571ab1790a7b683cb46965b2d0836c3c787d7

Observation 204e1e9e-7f5a-41fa-a845-8a5c60ff7e9a · outbound

This paper cites s1: Simple test-time scaling.

SOD: Step-wise On-policy Distillation for Small Language Model Agents s1: Simple test-time scaling

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T12:45:15.800279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-11T02:25:59.056181Z digest=sha256:b0c6b08a74b47a16dea37f80752f66f02431bcc184f83a9f0516081bdaaef6d9

Observation a652af40-8885-414d-b31f-fab7ff4f8893 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

SOD: Step-wise On-policy Distillation for Small Language Model Agents DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 69

Resolution
verified exact
local_arxiv, observed 2026-05-11T03:40:54.502449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-11T02:25:59.056181Z digest=sha256:aed388b064b561d1baf9b842143424be6c13a703c6bd91a91127e4b671e840f1

Observation 222de1d9-2cac-4ee5-8a52-ed3b652667bd · outbound

This paper cites Skywork Open Reasoner 1 Technical Report.

SOD: Step-wise On-policy Distillation for Small Language Model Agents Skywork Open Reasoner 1 Technical Report

Reference 70

Resolution
verified exact
arxiv_id, observed 2026-05-17T04:26:47.432522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-11T02:25:59.056181Z digest=sha256:654531305520277f9e5e2349a35f6bf82b95286d780b025e11f36e7ed3eec678

Observation 6683b0ae-5a2a-4ff1-bc95-d89a5a8bb99c · outbound

This paper cites MegaScience: Pushing the Frontiers of Post-Training Datasets for Science Reasoning.

SOD: Step-wise On-policy Distillation for Small Language Model Agents MegaScience: Pushing the Frontiers of Post-Training Datasets for Science Reasoning

Reference 71

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:40:54.499255Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-11T02:25:59.056181Z digest=sha256:877a31fee7f0557dd8eadeda43ad7004c823ff83300b539432f546749015a1b1

Observation 71dbaed3-d187-45ef-9b4f-d5c89cf25ebf · outbound

This paper cites Gpqa: A graduate-level google-proof q&a benchmark.

SOD: Step-wise On-policy Distillation for Small Language Model Agents Gpqa: A graduate-level google-proof q&a benchmark

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T12:40:17.352075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-11T02:25:59.056181Z digest=sha256:68f347203aefef43dcd6d109c09f4a236f0c7f0fc77666365cddd7d7a1f3b8fe

Observation b727a3d1-5a32-47cb-b835-92cd5b622f19 · outbound

This paper cites Livecodebench: Holistic and contamination free evaluation of large language models for code.

SOD: Step-wise On-policy Distillation for Small Language Model Agents Livecodebench: Holistic and contamination free evaluation of large language models for code

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T12:40:17.336221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-11T02:25:59.056181Z digest=sha256:21d9773570845aa46ac7d3ca692ab937461782ae1bf744e4b61fdbf14f396e13

Observation 8f5af7ef-36d8-4eea-83d3-c8e54f24ec26 · outbound

This paper cites Reasonflux-prm: Trajectory-aware prms for long chain-of-thought reasoning in llms.The Thirty-ninth Annual Conference on Neural Information Processing Systems.

SOD: Step-wise On-policy Distillation for Small Language Model Agents Reasonflux-prm: Trajectory-aware prms for long chain-of-thought reasoning in llms.The Thirty-ninth Annual Conference on Neural Information Processing Systems

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T12:40:17.327393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-11T02:25:59.056181Z digest=sha256:f7f8a53c468a435c2e5c95e7dc4855e8648ff889f0802cb27993d35ca5cb5b6a

Observation fe0f54f0-85b2-46d0-b37e-68742d50efe2 · outbound

This paper cites LeetCodeDataset: A Temporal Dataset for Robust Evaluation and Efficient Training of Code LLMs.

SOD: Step-wise On-policy Distillation for Small Language Model Agents LeetCodeDataset: A Temporal Dataset for Robust Evaluation and Efficient Training of Code LLMs

Reference 75

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:40:54.525214Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-11T02:25:59.056181Z digest=sha256:d73c123558d430801ffd8fc34b66ab08dea740148972b412a5cdfac0d12f15f2

Observation 008ff677-4787-48fe-b2bf-68abdbc4b69f · outbound

This paper cites Google-proof.

SOD: Step-wise On-policy Distillation for Small Language Model Agents Google-proof

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T12:40:17.319367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-11T02:25:59.056181Z digest=sha256:3815a8eb67de77e341a0cd85e90db5dd6fb98922f54b75c0d319692c8a486e6e

Observation 0cef3d08-67c9-4bda-8942-b3772d3e0c51 · outbound

This paper cites The maximum prompt length is set to 2,560 tokens and the maximum response length to 20,480 tokens.

SOD: Step-wise On-policy Distillation for Small Language Model Agents The maximum prompt length is set to 2,560 tokens and the maximum response length to 20,480 tokens

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T12:40:17.323281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-11T02:25:59.056181Z digest=sha256:c2eab6541ae203474d2b35f50e923d9ff1e197dfad9f2537b80af38ec9ae6d93

Observation 24d4001e-5d5f-4a8f-ae83-53989b50f71c · outbound

This paper cites This already introduces a divergence jump substantially larger than text-only drift ( Ω(m·η tool) vs.

SOD: Step-wise On-policy Distillation for Small Language Model Agents This already introduces a divergence jump substantially larger than text-only drift ( Ω(m·η tool) vs

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T12:40:17.315121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-11T02:25:59.056181Z digest=sha256:1dd4454e146d5843fdeca22662245941f95e31ed97cc00729a289474595dfdb5

Observation bf207ca7-6bbb-4179-ab9f-cd198129964f · outbound

This paper cites an unresolved cited work.

SOD: Step-wise On-policy Distillation for Small Language Model Agents Unresolved cited work

Reference 79

Resolution
unresolved
raw_fallback, observed 2026-05-14T12:40:17.310420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-11T02:25:59.056181Z digest=sha256:d9a259749fb3c803339623ef26e84df3579d1132c2248ff7f51e0e7d38c19c19

Observation 84286c90-ee92-41c8-a4e4-ce1442914c7d · outbound

This paper cites Updates become dominated by uninformative, high-magnitude contributions from tokens where the teacher provides no meaningful guidance.

SOD: Step-wise On-policy Distillation for Small Language Model Agents Updates become dominated by uninformative, high-magnitude contributions from tokens where the teacher provides no meaningful guidance

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T12:40:17.342690Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-11T02:25:59.056181Z digest=sha256:5f5106e38cead1df3de79a88e64ec4f3ae565b15f6a644093e5305ce74bec992

Observation 8ca1ccc9-8c17-48b1-af18-98abebe27f3b · outbound

This paper cites 36”, changes to “\boxed{66}.

SOD: Step-wise On-policy Distillation for Small Language Model Agents 36”, changes to “\boxed{66}

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T12:40:17.347410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-11T02:25:59.056181Z digest=sha256:a203805514f197d2c2c7747b061f3ff245345fb4ac2b9597bbc47a98509dd460

Pith citing papers

Observation 2d033aab-3719-44d3-b4af-1b427e359efe · inbound

SimCT: Recovering Lost Supervision for Cross-Tokenizer On-Policy Distillation cites this paper.

SimCT: Recovering Lost Supervision for Cross-Tokenizer On-Policy Distillation SOD: Step-wise On-policy Distillation for Small Language Model Agents

Reference 69

Resolution
verified exact
local_arxiv, observed 2026-05-22T10:31:25.020076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-22T10:30:11.309910Z digest=sha256:54b0b631bbfc301a27e241e9074d06d0dec969111df038ca2eb81983962193e4

Observation 6af974ae-4e7a-4354-926b-45edfd26ca2b · inbound

ATOD: Annealed Turn-Aware On-Policy Distillation for Multi-Turn Agentic Tasks cites this paper.

ATOD: Annealed Turn-Aware On-Policy Distillation for Multi-Turn Agentic Tasks SOD: Step-wise On-policy Distillation for Small Language Model Agents

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-06-29T20:03:57.356711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-06-29T04:30:43.424100Z digest=sha256:9f84fff54c1b0e9f1bb811c429291d8fc4f7cb96b2e17bc3e06b95e90ce90a50

Observation 1a3bf83e-6221-4a24-9c75-3282cdfdba3c · inbound

TurnOPD: Making On-Policy Distillation Turn-Aware for Efficient Long-Horizon Agent Training cites this paper.

TurnOPD: Making On-Policy Distillation Turn-Aware for Efficient Long-Horizon Agent Training SOD: Step-wise On-policy Distillation for Small Language Model Agents

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-07-09T00:35:48.712044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-07-09T00:29:35.583300Z digest=sha256:c296544cfc946b9a535130b12607b16f56f25bf3a3bf9b40c92958376de4357c

Observation 1c9dc3f0-37b5-4dc9-9a73-420cf8b81308 · inbound

Rethinking Classifier-Free Guidance in On-Policy Diffusion Distillation cites this paper.

Rethinking Classifier-Free Guidance in On-Policy Diffusion Distillation SOD: Step-wise On-policy Distillation for Small Language Model Agents

Reference 12

Resolution
unresolved
no resolver link, observed 2026-07-31T06:35:52.597753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T06:35:52.597753Z digest=sha256:a4545dfac4778fa44ef9b057fa1b792c14c02a4efc99f598844dd034ac55c7a7

Observation a825be51-2b0e-434c-abc2-4677470ea796 · inbound

Group-Reflective Self-Distillation for Agentic Reinforcement Learning cites this paper.

Group-Reflective Self-Distillation for Agentic Reinforcement Learning SOD: Step-wise On-policy Distillation for Small Language Model Agents

Reference 40

Resolution
unresolved
no resolver link, observed 2026-07-31T18:36:19.656189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T18:36:19.656189Z digest=sha256:daccdcc65dd5f05dbe5124af5ae0210a49b34c0e02c8e7762208f2ab168f36df