Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-11T02:25:59.056181Z
Paper Citation Record · LEDGER
As of 1 August 2026, this Paper Citation Record lists 81 of 81 outbound references and 5 inbound Pith citation observations for arXiv:2605.07725.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-11T02:25:59.056181Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-01T06:32:01.292127+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-07-31T18:36:19.656189Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-09T00:35:48.710213Z
81 of 81 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 387058b7-0712-420c-b4f0-57f4cc8c728a · outbound
SOD: Step-wise On-policy Distillation for Small Language Model Agents The Rise and Potential of Large Language Model Based Agents: A Survey
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation beb73551-eb2d-4192-8626-c56a05dfec20 · outbound
SOD: Step-wise On-policy Distillation for Small Language Model Agents Distilling llm agent into small models with retrieval and code tools
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation df63bdf5-522a-4aa1-810b-3571257e9a95 · outbound
SOD: Step-wise On-policy Distillation for Small Language Model Agents Narasimhan, and Yuan Cao
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 23c89fc3-b040-45c0-b960-65ee38d9a09a · outbound
SOD: Step-wise On-policy Distillation for Small Language Model Agents Toolformer: Language models can teach themselves to use tools
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 4c94ae1f-2a38-492b-9b90-cf5594b2b0f7 · outbound
SOD: Step-wise On-policy Distillation for Small Language Model Agents Agentic Reasoning and Tool Integration for LLMs via Reinforcement Learning
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 855ef61a-7595-49f8-b181-799fd44a9543 · outbound
SOD: Step-wise On-policy Distillation for Small Language Model Agents Mixed distillation helps smaller language models reason better
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 29e19b0a-d9ad-4791-bc49-3f8139fcb030 · outbound
SOD: Step-wise On-policy Distillation for Small Language Model Agents AgentProcessBench: Diagnosing Step-Level Process Quality in Tool-Using Agents
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 11cae7de-8f7d-42f5-9db1-2518e889257a · outbound
SOD: Step-wise On-policy Distillation for Small Language Model Agents On-Device Language Models: A Comprehensive Review
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 5874f411-31f7-41df-9fe7-93bde841b2df · outbound
SOD: Step-wise On-policy Distillation for Small Language Model Agents A Survey on Knowledge Distillation of Large Language Models
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 78e615e7-2449-4ac4-aa66-e2da5269144a · outbound
SOD: Step-wise On-policy Distillation for Small Language Model Agents AgentDistill: Training-Free Agent Distillation with Generalizable MCP Boxes
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 926ac0a8-bc32-4a68-91f0-e76bcab2e798 · outbound
SOD: Step-wise On-policy Distillation for Small Language Model Agents O-researcher: An open ended deep research model via multi-agent distillation and agentic rl
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 1e23cc51-8da2-4836-a10b-219e68569a74 · outbound
SOD: Step-wise On-policy Distillation for Small Language Model Agents Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation d5a8dc9d-55e9-4b10-9237-374ab7b93e3c · outbound
SOD: Step-wise On-policy Distillation for Small Language Model Agents ToolRL: Reward is All Tool Learning Needs
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 2eb7cbae-7bad-446d-a269-775a3c3e9e82 · outbound
SOD: Step-wise On-policy Distillation for Small Language Model Agents Replacing thinking with tool usage enables reasoning in small language models
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 31f4b51b-df37-4077-ba13-c1f579eaefd2 · outbound
SOD: Step-wise On-policy Distillation for Small Language Model Agents SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation f89fdfb4-7e0c-438f-bbc0-a5aedf41b9c1 · outbound
SOD: Step-wise On-policy Distillation for Small Language Model Agents Structured Agent Distillation for Large Language Model
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 2001a2f7-cab5-4e06-993e-153caaf1c8bb · outbound
SOD: Step-wise On-policy Distillation for Small Language Model Agents ToRL: Scaling Tool-Integrated RL
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation beef77d6-6fe2-43eb-b68a-87c2f95821bd · outbound
SOD: Step-wise On-policy Distillation for Small Language Model Agents Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 7686b27a-49c1-479a-a0b6-312c162e2e0e · outbound
SOD: Step-wise On-policy Distillation for Small Language Model Agents R1-Searcher: Incentivizing the Search Capability in LLMs via Reinforcement Learning
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 23f75ed2-d476-4844-9b51-79f526810541 · outbound
SOD: Step-wise On-policy Distillation for Small Language Model Agents DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 3eecbc68-9747-4ded-a5fd-07abd909f807 · outbound
SOD: Step-wise On-policy Distillation for Small Language Model Agents DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 47e026fd-5b20-462c-8edd-9895967bc0b6 · outbound
SOD: Step-wise On-policy Distillation for Small Language Model Agents KEPO: Knowledge-Enhanced Preference Optimization for Multimodal Reasoning with Applications to Medical VQA
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 7b6694bd-55df-4c1c-bba9-4116d3cfb541 · outbound
SOD: Step-wise On-policy Distillation for Small Language Model Agents On-policy distillation of language models: Learning from self-generated mistakes
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 9a63733c-d9bd-4e83-88dc-acf17807d0ee · outbound
SOD: Step-wise On-policy Distillation for Small Language Model Agents Minillm: Knowledge distillation of large language models
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation d5d8abe9-5ccc-44ed-b3de-de977da1bad6 · outbound
SOD: Step-wise On-policy Distillation for Small Language Model Agents Entropy-Aware On-Policy Distillation of Language Models
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation f5126f51-8d44-4e58-a339-d567eb72a11c · outbound
SOD: Step-wise On-policy Distillation for Small Language Model Agents Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph E
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation cb307289-c7a3-4025-b26e-a7473d5e8e3e · outbound
SOD: Step-wise On-policy Distillation for Small Language Model Agents Lightning OPD: Efficient Post-Training for Large Reasoning Models with Offline On-Policy Distillation
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation edc8b8d2-f5fe-4ff9-8217-b0442b85d68f · outbound
SOD: Step-wise On-policy Distillation for Small Language Model Agents Qwen3 Technical Report
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 23a42cb7-8442-4368-b85d-fee896d4f1d3 · outbound
SOD: Step-wise On-policy Distillation for Small Language Model Agents Rethinking On-Policy Distillation of Large Language Models: Phenomenology, Mechanism, and Recipe
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 72d33e95-19a6-408b-877d-0fd22a7d21d2 · outbound
SOD: Step-wise On-policy Distillation for Small Language Model Agents Rlkd: Distilling llms’ reasoning via reinforcement learning
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 9da3d397-d60f-4dad-b43a-2b0d990d1715 · outbound
SOD: Step-wise On-policy Distillation for Small Language Model Agents Unifying group-relative and self-distillation policy optimization via sample routing
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 9fd0ddf1-23d1-4275-a3a9-da68959a41de · outbound
SOD: Step-wise On-policy Distillation for Small Language Model Agents Self-Distilled RLVR
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation cd7b84f1-02bd-47d0-bc5d-1a4ac02fd924 · outbound
SOD: Step-wise On-policy Distillation for Small Language Model Agents VOLD: Reasoning Transfer from LLMs to Vision-Language Models via On-Policy Distillation
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation aa642c6a-7684-46d0-9e09-54e20b7bdac5 · outbound
SOD: Step-wise On-policy Distillation for Small Language Model Agents OpenClaw-RL: Train Any Agent Simply by Talking
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 8c64a382-1919-40d1-9c8d-c6d215e01efd · outbound
SOD: Step-wise On-policy Distillation for Small Language Model Agents Revisiting On-Policy Distillation: Empirical Failure Modes and Simple Fixes
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 4e0097b7-59c4-4418-ba8d-b16f8aec5f5e · outbound
SOD: Step-wise On-policy Distillation for Small Language Model Agents A Survey of On-Policy Distillation for Large Language Models
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 64e62c2a-e126-4abb-9d5e-164b900f48f8 · outbound
SOD: Step-wise On-policy Distillation for Small Language Model Agents Gordon, and Drew Bagnell
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 0f6b9d9e-c074-43df-a60f-7103d8d1622b · outbound
SOD: Step-wise On-policy Distillation for Small Language Model Agents The False Promise of Imitating Proprietary LLMs
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation e21d29eb-c765-4c6b-87e3-d433cc682f34 · outbound
SOD: Step-wise On-policy Distillation for Small Language Model Agents TCOD: Exploring Temporal Curriculum in On-Policy Distillation for Multi-turn Autonomous Agents
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 399b77e4-ff4b-421f-88c7-caa536af039e · outbound
SOD: Step-wise On-policy Distillation for Small Language Model Agents Stable On-Policy Distillation through Adaptive Target Reformulation
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation cfbe5751-f9fc-473e-91ad-8609a9615974 · outbound
SOD: Step-wise On-policy Distillation for Small Language Model Agents TIP: Token Importance in On-Policy Distillation
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation e32f4091-39c5-48bb-9eeb-20851908a470 · outbound
SOD: Step-wise On-policy Distillation for Small Language Model Agents Deep reinforcement learning from human preferences.Advances in neural information processing systems, 30
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation de8e4467-84a9-432f-ade2-78b797f71d54 · outbound
SOD: Step-wise On-policy Distillation for Small Language Model Agents Fine-Tuning Language Models from Human Preferences
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 5d9a04a2-6c28-439c-bd9b-52794975e857 · outbound
SOD: Step-wise On-policy Distillation for Small Language Model Agents Proximal Policy Optimization Algorithms
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 414876f4-f394-4d5f-ae80-fb49c6d560d0 · outbound
SOD: Step-wise On-policy Distillation for Small Language Model Agents FireAct: Toward Language Agent Fine-tuning
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 9c81871f-9215-4d1e-8c71-5d595e9883ca · outbound
SOD: Step-wise On-policy Distillation for Small Language Model Agents Swe-agent: Agent-computer interfaces enable automated software engineering.Advances in Neural Information Processing Systems, 37:50528–50652
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation a77abe9c-843d-4664-a59f-8614182031ed · outbound
SOD: Step-wise On-policy Distillation for Small Language Model Agents ReTool: Reinforcement Learning for Strategic Tool Use in LLMs
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 4efebc59-7037-405d-8524-6744e16f0995 · outbound
SOD: Step-wise On-policy Distillation for Small Language Model Agents Digirl: Training in-the-wild device-control agents with autonomous reinforcement learning
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 11ce60e9-4099-45ee-8ddb-43cd23bb6da5 · outbound
SOD: Step-wise On-policy Distillation for Small Language Model Agents WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 830e555e-e74d-40ef-84fb-4290fa4652ad · outbound
SOD: Step-wise On-policy Distillation for Small Language Model Agents ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation f4304d3e-0096-4a0d-9b56-6c8036d8642a · outbound
SOD: Step-wise On-policy Distillation for Small Language Model Agents Reinforcement Learning for Long-Horizon Interactive LLM Agents
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 73c0740e-bc66-468b-a498-943e5494ab26 · outbound
SOD: Step-wise On-policy Distillation for Small Language Model Agents Demystifying reinforcement learning in agentic reasoning
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 2cc8d805-46bc-4f27-8e89-1c239d8a5be1 · outbound
SOD: Step-wise On-policy Distillation for Small Language Model Agents Rlanything: Forge environment, policy, and reward model in completely dynamic rl system
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 8335d5dd-2469-4c5f-a4ea-591b10668fca · outbound
SOD: Step-wise On-policy Distillation for Small Language Model Agents Co-evolving llm coder and unit tester via reinforcement learning.arXiv preprint arXiv:2506.03136
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 1df2d2f0-9e89-40ad-8337-899c60069ee7 · outbound
SOD: Step-wise On-policy Distillation for Small Language Model Agents On-Policy Context Distillation for Language Models
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 5c1881bd-45b1-464c-a3b3-e9b09c88dcc2 · outbound
SOD: Step-wise On-policy Distillation for Small Language Model Agents Black-box on-policy distillation of large language models.arXiv preprint
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 81440573-cb7e-4fe5-b359-ad3e7c5a5fca · outbound
SOD: Step-wise On-policy Distillation for Small Language Model Agents Hybrid Policy Distillation for LLMs
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 1a1229c6-63d0-4ef8-a2e8-ef3f7b015bc5 · outbound
SOD: Step-wise On-policy Distillation for Small Language Model Agents SCOPE: Signal-Calibrated On-Policy Distillation Enhancement with Dual-Path Adaptive Weighting
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 019ee0cb-692a-4c8c-8c7b-cc7f51187955 · outbound
SOD: Step-wise On-policy Distillation for Small Language Model Agents SODA: Semi On-Policy Black-Box Distillation for Large Language Models
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 61b0c8b4-a0e2-4045-8dc7-d3430afc24cd · outbound
SOD: Step-wise On-policy Distillation for Small Language Model Agents Learning beyond Teacher: Generalized On-Policy Distillation with Reward Extrapolation
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 28f76bcb-17ac-48ef-a40e-23b074abbde9 · outbound
SOD: Step-wise On-policy Distillation for Small Language Model Agents Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation b1e518f7-5015-4e73-8957-0c918aa3f206 · outbound
SOD: Step-wise On-policy Distillation for Small Language Model Agents Privileged Information Distillation for Language Models
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 0dce8468-16a9-4efd-822f-7d08368f4c1f · outbound
SOD: Step-wise On-policy Distillation for Small Language Model Agents Reinforcement Learning via Self-Distillation
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 5e3f0643-4f42-40f2-b21f-cade4f97822e · outbound
SOD: Step-wise On-policy Distillation for Small Language Model Agents Self-Distillation Enables Continual Learning
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 0d8b62f4-f970-43a0-8743-3cf917334732 · outbound
SOD: Step-wise On-policy Distillation for Small Language Model Agents Self-Distillation Zero: Self-Revision Turns Binary Rewards into Dense Supervision
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 7e4dbbc3-1559-41b4-b645-4368493b19a8 · outbound
SOD: Step-wise On-policy Distillation for Small Language Model Agents Skill-SD: Skill-Conditioned Self-Distillation for Multi-turn LLM Agents
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 8f63b949-d0fc-410e-acb1-7657a83b2a75 · outbound
SOD: Step-wise On-policy Distillation for Small Language Model Agents On-policy distillation.Thinking Machines Lab: Con- nectionism
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T07:08:02.609025+00:00.
Observation 204e1e9e-7f5a-41fa-a845-8a5c60ff7e9a · outbound
SOD: Step-wise On-policy Distillation for Small Language Model Agents s1: Simple test-time scaling
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation a652af40-8885-414d-b31f-fab7ff4f8893 · outbound
SOD: Step-wise On-policy Distillation for Small Language Model Agents DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 222de1d9-2cac-4ee5-8a52-ed3b652667bd · outbound
SOD: Step-wise On-policy Distillation for Small Language Model Agents Skywork Open Reasoner 1 Technical Report
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 6683b0ae-5a2a-4ff1-bc95-d89a5a8bb99c · outbound
SOD: Step-wise On-policy Distillation for Small Language Model Agents MegaScience: Pushing the Frontiers of Post-Training Datasets for Science Reasoning
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 71dbaed3-d187-45ef-9b4f-d5c89cf25ebf · outbound
SOD: Step-wise On-policy Distillation for Small Language Model Agents Gpqa: A graduate-level google-proof q&a benchmark
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation b727a3d1-5a32-47cb-b835-92cd5b622f19 · outbound
SOD: Step-wise On-policy Distillation for Small Language Model Agents Livecodebench: Holistic and contamination free evaluation of large language models for code
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 8f5af7ef-36d8-4eea-83d3-c8e54f24ec26 · outbound
SOD: Step-wise On-policy Distillation for Small Language Model Agents Reasonflux-prm: Trajectory-aware prms for long chain-of-thought reasoning in llms.The Thirty-ninth Annual Conference on Neural Information Processing Systems
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation fe0f54f0-85b2-46d0-b37e-68742d50efe2 · outbound
SOD: Step-wise On-policy Distillation for Small Language Model Agents LeetCodeDataset: A Temporal Dataset for Robust Evaluation and Efficient Training of Code LLMs
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 008ff677-4787-48fe-b2bf-68abdbc4b69f · outbound
SOD: Step-wise On-policy Distillation for Small Language Model Agents Google-proof
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 0cef3d08-67c9-4bda-8942-b3772d3e0c51 · outbound
SOD: Step-wise On-policy Distillation for Small Language Model Agents The maximum prompt length is set to 2,560 tokens and the maximum response length to 20,480 tokens
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 24d4001e-5d5f-4a8f-ae83-53989b50f71c · outbound
SOD: Step-wise On-policy Distillation for Small Language Model Agents This already introduces a divergence jump substantially larger than text-only drift ( Ω(m·η tool) vs
Reference 78
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation bf207ca7-6bbb-4179-ab9f-cd198129964f · outbound
SOD: Step-wise On-policy Distillation for Small Language Model Agents Unresolved cited work
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 84286c90-ee92-41c8-a4e4-ce1442914c7d · outbound
SOD: Step-wise On-policy Distillation for Small Language Model Agents Updates become dominated by uninformative, high-magnitude contributions from tokens where the teacher provides no meaningful guidance
Reference 80
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 8ca1ccc9-8c17-48b1-af18-98abebe27f3b · outbound
SOD: Step-wise On-policy Distillation for Small Language Model Agents 36”, changes to “\boxed{66}
Reference 81
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 2d033aab-3719-44d3-b4af-1b427e359efe · inbound
SimCT: Recovering Lost Supervision for Cross-Tokenizer On-Policy Distillation SOD: Step-wise On-policy Distillation for Small Language Model Agents
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 6af974ae-4e7a-4354-926b-45edfd26ca2b · inbound
ATOD: Annealed Turn-Aware On-Policy Distillation for Multi-Turn Agentic Tasks SOD: Step-wise On-policy Distillation for Small Language Model Agents
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 1a3bf83e-6221-4a24-9c75-3282cdfdba3c · inbound
TurnOPD: Making On-Policy Distillation Turn-Aware for Efficient Long-Horizon Agent Training SOD: Step-wise On-policy Distillation for Small Language Model Agents
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 1c9dc3f0-37b5-4dc9-9a73-420cf8b81308 · inbound
Rethinking Classifier-Free Guidance in On-Policy Diffusion Distillation SOD: Step-wise On-policy Distillation for Small Language Model Agents
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a825be51-2b0e-434c-abc2-4677470ea796 · inbound
Group-Reflective Self-Distillation for Agentic Reinforcement Learning SOD: Step-wise On-policy Distillation for Small Language Model Agents
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.