Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T23:51:43.368701Z
Paper Citation Record · LEDGER
As of 22 August 2026, this Paper Citation Record lists 74 of 74 outbound references and 0 inbound Pith citation observations for arXiv:2608.12764.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T23:51:43.368701Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
74 of 74 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 548ae55b-55e2-41eb-954d-b76e5e591933 · outbound
Beyond Outcome Rewards: Step-Level Self-Distilled Policy Optimization for Deep Search Agents Glm-4.6.https://docs.z.ai/guides/llm/glm-4.6, 2025
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation c1e85063-e048-44c7-b5a2-a1c149ddbf0a · outbound
Beyond Outcome Rewards: Step-Level Self-Distilled Policy Optimization for Deep Search Agents Glm-5.1.https://docs.z.ai/guides/llm/glm-5.1, 2026
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 9ed6c33b-fb1f-40be-93b9-0ba7e7fc2e4f · outbound
Beyond Outcome Rewards: Step-Level Self-Distilled Policy Optimization for Deep Search Agents Introducing Claude 4 — anthropic.com
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation dd43c5cd-cd67-4d67-8ec2-b4b3ee26c669 · outbound
Beyond Outcome Rewards: Step-Level Self-Distilled Policy Optimization for Deep Search Agents Introducing claude opus 4.7
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 61ac298d-1de2-46b8-80eb-5bac227bbd23 · outbound
Beyond Outcome Rewards: Step-Level Self-Distilled Policy Optimization for Deep Search Agents Is gpt-oss good? a comprehensive evaluation of openai’s latest open source models, 2025
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation c25e1169-fde7-4057-85bb-fbf7c94c7f7d · outbound
Beyond Outcome Rewards: Step-Level Self-Distilled Policy Optimization for Deep Search Agents SFT memorizes, RL generalizes
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 418d6d98-af6d-414e-bdaa-a607bb5b60bd · outbound
Beyond Outcome Rewards: Step-Level Self-Distilled Policy Optimization for Deep Search Agents Gemini 3.1 pro: Best for complex tasks and bringing creative concepts to life.https://deepmind.google/models/gemini/pro/, 2026
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation ffdab16d-2070-416f-ba54-76c838ebfb19 · outbound
Beyond Outcome Rewards: Step-Level Self-Distilled Policy Optimization for Deep Search Agents Deepseek-v4: Towards highly efficient million-token context intelligence, 2026
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6345e824-41d7-403e-8924-25c31cd738c2 · outbound
Beyond Outcome Rewards: Step-Level Self-Distilled Policy Optimization for Deep Search Agents Deepseek-v3.2: Pushing the frontier of open large language models, 2025
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 83d43f61-deb0-4545-aa73-af2292a24da3 · outbound
Beyond Outcome Rewards: Step-Level Self-Distilled Policy Optimization for Deep Search Agents Openthoughts: Data recipes for reasoning models, 2025
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation ff42beca-82c0-4741-b8d7-724ef4aea5cd · outbound
Beyond Outcome Rewards: Step-Level Self-Distilled Policy Optimization for Deep Search Agents Unresolved cited work
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 2904fb17-06de-41a6-a14e-4fa7bd5933af · outbound
Beyond Outcome Rewards: Step-Level Self-Distilled Policy Optimization for Deep Search Agents Constructing a multi-hop QA dataset for comprehensive evaluation of reasoning steps
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cf34c1f6-2774-42bd-9aa3-68e0632b0f3c · outbound
Beyond Outcome Rewards: Step-Level Self-Distilled Policy Optimization for Deep Search Agents Reinforcement Learning via Self-Distillation
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bbcfe5a8-7c4d-4729-bd5c-0faf5bb04904 · outbound
Beyond Outcome Rewards: Step-Level Self-Distilled Policy Optimization for Deep Search Agents An Empirical Study on Reinforcement Learning for Reasoning-Search Interleaved LLM Agents
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 97fbe675-2bfe-49eb-9b8a-d2e9faf3e251 · outbound
Beyond Outcome Rewards: Step-Level Self-Distilled Policy Optimization for Deep Search Agents Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 92f2cbd0-5407-4c46-816c-332d28b43ff2 · outbound
Beyond Outcome Rewards: Step-Level Self-Distilled Policy Optimization for Deep Search Agents Fact, fetch, and reason: A unified evaluation of retrieval-augmented generation, 2024
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation f05af0a4-9eb9-41e7-bfb8-5d88a1596e67 · outbound
Beyond Outcome Rewards: Step-Level Self-Distilled Policy Optimization for Deep Search Agents Unifying group-relative and self-distillation policy optimization via sample routing, 2026
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e4e0c641-23b8-482d-9d52-34c178b8b2c5 · outbound
Beyond Outcome Rewards: Step-Level Self-Distilled Policy Optimization for Deep Search Agents The tool decathlon: Benchmarking language agents for diverse, realistic, and long-horizon task execution, 2026
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation a140c354-1a2d-4ab6-b8f4-a0801f3f26f2 · outbound
Beyond Outcome Rewards: Step-Level Self-Distilled Policy Optimization for Deep Search Agents Websailor-v2: Bridging the chasm to proprietary agents via synthetic data and scalable reinforcement learning, 2025
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 1b758483-8e39-456d-af58-47c56d82d287 · outbound
Beyond Outcome Rewards: Step-Level Self-Distilled Policy Optimization for Deep Search Agents Websailor: Navigating super-human reasoning for web agent, 2025
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 8a65a022-0282-4aa2-a37a-067dcfa28e0f · outbound
Beyond Outcome Rewards: Step-Level Self-Distilled Policy Optimization for Deep Search Agents Chain-of-agents: End-to-end agent foundation models via multi-agent distillation and agentic rl, 2025
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation d7383d2c-d8df-4543-ad74-e1e76aa73a4a · outbound
Beyond Outcome Rewards: Step-Level Self-Distilled Policy Optimization for Deep Search Agents Webthinker: Empowering large reasoning models with deep research capability, 2025
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation ee6b7ef5-bd7e-4eaa-b120-b617099f344a · outbound
Beyond Outcome Rewards: Step-Level Self-Distilled Policy Optimization for Deep Search Agents Rethinking On-Policy Distillation of Large Language Models: Phenomenology, Mechanism, and Recipe
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d4bf8ac0-7843-4c0e-87d0-d91b7edf8f46 · outbound
Beyond Outcome Rewards: Step-Level Self-Distilled Policy Optimization for Deep Search Agents Code-r1: Reproducing r1 for code with reliable rewards
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d2e772a2-db15-4381-a0e3-384c9ba65cdd · outbound
Beyond Outcome Rewards: Step-Level Self-Distilled Policy Optimization for Deep Search Agents Synlogic: Synthesizing verifiable reasoning data at scale for learning logical reasoning and beyond, 2025
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ddeb2804-c677-44c9-a553-71645a4bf174 · outbound
Beyond Outcome Rewards: Step-Level Self-Distilled Policy Optimization for Deep Search Agents Webexplorer: Explore and evolve for training long-horizon web agents, 2025
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation efd09cfe-6597-4652-a0fc-febcdad64339 · outbound
Beyond Outcome Rewards: Step-Level Self-Distilled Policy Optimization for Deep Search Agents On-policy distillation.Thinking Machines Lab: Connec- tionism, 2025
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 67d044b7-5e1d-4cfe-80b8-677198465c45 · outbound
Beyond Outcome Rewards: Step-Level Self-Distilled Policy Optimization for Deep Search Agents Deepdive: Advancing deep search agents with knowledge graphs and multi-turn rl, 2025
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 14420623-eb4a-42d0-b418-dee0e3931b24 · outbound
Beyond Outcome Rewards: Step-Level Self-Distilled Policy Optimization for Deep Search Agents Gaia: a benchmark for general ai assistants, 2023
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0e369d84-6b97-4c31-b204-86d92bc8f893 · outbound
Beyond Outcome Rewards: Step-Level Self-Distilled Policy Optimization for Deep Search Agents Minimax m2.7: Early echoes of self-evolution, 2026
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation b4d8224f-499a-4c67-a16d-4c557279332b · outbound
Beyond Outcome Rewards: Step-Level Self-Distilled Policy Optimization for Deep Search Agents Gpt-5-nano
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 1b80dfb7-199c-4a0d-8cb6-d9cb74fe7f51 · outbound
Beyond Outcome Rewards: Step-Level Self-Distilled Policy Optimization for Deep Search Agents Introducing gpt-oss
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation c7a3c941-fbae-4ded-854e-614078dc060d · outbound
Beyond Outcome Rewards: Step-Level Self-Distilled Policy Optimization for Deep Search Agents Introducing gpt -5.5
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 57bd4ec2-cd31-4f68-b75c-2c7bf272849f · outbound
Beyond Outcome Rewards: Step-Level Self-Distilled Policy Optimization for Deep Search Agents Qwen3.6-Plus: Towards real world agents, April 2026
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6a18700d-d4b5-4446-b0bf-2a88035e295b · outbound
Beyond Outcome Rewards: Step-Level Self-Distilled Policy Optimization for Deep Search Agents Crisp: Compressed reasoning via iterative self-policy distillation, 2026
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 4c5b4d03-3257-41e3-9e70-c77e4a31c06a · outbound
Beyond Outcome Rewards: Step-Level Self-Distilled Policy Optimization for Deep Search Agents Generalization in generation: A closer look at exposure bias
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 4cf2727c-0a05-4416-914f-0597dda75701 · outbound
Beyond Outcome Rewards: Step-Level Self-Distilled Policy Optimization for Deep Search Agents Unresolved cited work
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f54d5f89-7ecc-4ad0-bb98-e67927e0f891 · outbound
Beyond Outcome Rewards: Step-Level Self-Distilled Policy Optimization for Deep Search Agents Self-distillation enables continual learning, 2026
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b310cec0-6933-4c75-9c8a-faa11b68b926 · outbound
Beyond Outcome Rewards: Step-Level Self-Distilled Policy Optimization for Deep Search Agents A survey of on-policy distillation for large language models, 2026
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 83067a0b-77ad-4284-bc08-cb3c031d2576 · outbound
Beyond Outcome Rewards: Step-Level Self-Distilled Policy Optimization for Deep Search Agents Webshaper: Agentically data synthesizing via information-seeking formalization, 2025
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 9558d3d7-cf9a-4498-a7cc-94240f65c2e3 · outbound
Beyond Outcome Rewards: Step-Level Self-Distilled Policy Optimization for Deep Search Agents Kimi k2.6: Advancing open-source coding
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 3c517d4d-a00d-4422-8f88-bc637c8986c4 · outbound
Beyond Outcome Rewards: Step-Level Self-Distilled Policy Optimization for Deep Search Agents MiroThinker: Pushing the Performance Boundaries of Open-Source Research Agents via Model, Context, and Interactive Scaling
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8a772cdf-7424-4a82-accf-872f59ca70a2 · outbound
Beyond Outcome Rewards: Step-Level Self-Distilled Policy Optimization for Deep Search Agents Browsecomp: A simple yet challenging benchmark for browsing agents, 2025
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4788aa76-868c-4c01-bf0f-2843fee84807 · outbound
Beyond Outcome Rewards: Step-Level Self-Distilled Policy Optimization for Deep Search Agents Chain-of-thought prompting elicits reasoning in large language models, 2023
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 29ff97a0-2086-4e2b-9d1e-c04652128481 · outbound
Beyond Outcome Rewards: Step-Level Self-Distilled Policy Optimization for Deep Search Agents Smartsearch: Process reward-guided query refinement for search agents, 2026
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 191dff8d-fe0f-433c-9749-f9147ab6ad48 · outbound
Beyond Outcome Rewards: Step-Level Self-Distilled Policy Optimization for Deep Search Agents Mirage or method? how model-task alignment induces divergent rl conclusions, 2025
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 8005269b-060e-43ba-82d3-6734b9123600 · outbound
Beyond Outcome Rewards: Step-Level Self-Distilled Policy Optimization for Deep Search Agents Recode: Updating code api knowledge with reinforcement learning, 2025
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation c6938df4-51c4-4d03-94bd-24fb7aaad3fc · outbound
Beyond Outcome Rewards: Step-Level Self-Distilled Policy Optimization for Deep Search Agents On the generalization of sft: A reinforcement learning perspective with reward rectification, 2026
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 4f3c1a6e-fe1b-4e02-9fb5-1422126c4526 · outbound
Beyond Outcome Rewards: Step-Level Self-Distilled Policy Optimization for Deep Search Agents Grok 4.1 fast and agent tools api.https://x.ai/news/grok-4-1-fast, 2025
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 8f6fd0bb-f01d-4ffc-8a13-8bae0310a77d · outbound
Beyond Outcome Rewards: Step-Level Self-Distilled Policy Optimization for Deep Search Agents Logic-rl: Unleashing llm reasoning with rule-based reinforcement learning, 2025
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f3271436-d47e-4138-85c1-e3730eee63c4 · outbound
Beyond Outcome Rewards: Step-Level Self-Distilled Policy Optimization for Deep Search Agents Principle process reward for search agents, 2026
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 12da9215-0b9a-420e-a5d5-b714f08c0a3f · outbound
Beyond Outcome Rewards: Step-Level Self-Distilled Policy Optimization for Deep Search Agents Tip: Token importance in on-policy distillation, 2026
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 42f2809f-408f-4a3b-9531-c6ea05356dee · outbound
Beyond Outcome Rewards: Step-Level Self-Distilled Policy Optimization for Deep Search Agents Qwen3 technical report, 2025
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 312d3551-dfb5-4e9a-9e0d-4382c12354ac · outbound
Beyond Outcome Rewards: Step-Level Self-Distilled Policy Optimization for Deep Search Agents Self-distilled rlvr, 2026
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 35d6e048-4992-48ca-a0d3-0e0bc682f0b5 · outbound
Beyond Outcome Rewards: Step-Level Self-Distilled Policy Optimization for Deep Search Agents Cohen, Ruslan Salakhut- dinov, and Christopher D
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d742c681-1544-463c-a09f-331a18f4816b · outbound
Beyond Outcome Rewards: Step-Level Self-Distilled Policy Optimization for Deep Search Agents React: Synergizing reasoning and acting in language models, 2023
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 840f85e9-87c7-4670-86a4-15b0d5d55b88 · outbound
Beyond Outcome Rewards: Step-Level Self-Distilled Policy Optimization for Deep Search Agents On-policy context distillation for language models, 2026
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e179a6bd-ed91-47fb-8320-b9a32f1fb84b · outbound
Beyond Outcome Rewards: Step-Level Self-Distilled Policy Optimization for Deep Search Agents Dapo: An open-source llm reinforcement learning system at scale, 2025
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eba60e96-0f9f-486a-90bd-0fb221ce7f8a · outbound
Beyond Outcome Rewards: Step-Level Self-Distilled Policy Optimization for Deep Search Agents Simplerl-zoo: Investigating and taming zero reinforcement learning for open base models in the wild
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 1847a17f-81e8-46b7-a374-781a1ebd4558 · outbound
Beyond Outcome Rewards: Step-Level Self-Distilled Policy Optimization for Deep Search Agents The landscape of agentic reinforcement learning for LLMs: A survey.Transactions on Machine Learning Research,
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 625fdb2f-a8e4-489a-8087-d9f3ab320c85 · outbound
Beyond Outcome Rewards: Step-Level Self-Distilled Policy Optimization for Deep Search Agents Tool-r1: Sample-efficient reinforcement learning for agentic tool use, 2025
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 8eb1cad7-5bfc-4ba1-bb2d-71eb96edfe97 · outbound
Beyond Outcome Rewards: Step-Level Self-Distilled Policy Optimization for Deep Search Agents Criticsearch: Fine-grained credit assignment for search agents via a retrospective critic, 2025
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 5569005a-0110-47dc-accd-32fd96bfa726 · outbound
Beyond Outcome Rewards: Step-Level Self-Distilled Policy Optimization for Deep Search Agents Self-distilled reasoner: On-policy self-distillation for large language models, 2026
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 375f757c-d9e1-4c0b-b105-7a844f775214 · outbound
Beyond Outcome Rewards: Step-Level Self-Distilled Policy Optimization for Deep Search Agents aha moments
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 895914a5-9a82-4a58-901a-763b5af889fa · outbound
Beyond Outcome Rewards: Step-Level Self-Distilled Policy Optimization for Deep Search Agents Unresolved cited work
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 1a19588e-4b42-4fa6-9b8d-2366147fa7e7 · outbound
Beyond Outcome Rewards: Step-Level Self-Distilled Policy Optimization for Deep Search Agents Unresolved cited work
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation b0742e58-0293-4282-beb3-d51d6bcd15dd · outbound
Beyond Outcome Rewards: Step-Level Self-Distilled Policy Optimization for Deep Search Agents Unresolved cited work
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 45cf3b16-9885-4a08-98ce-3cbc662b9405 · outbound
Beyond Outcome Rewards: Step-Level Self-Distilled Policy Optimization for Deep Search Agents Unresolved cited work
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 9bbcc85a-0281-456a-9068-466bffddca48 · outbound
Beyond Outcome Rewards: Step-Level Self-Distilled Policy Optimization for Deep Search Agents Unresolved cited work
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 6636605a-6564-4d6c-86f4-90b27da15767 · outbound
Beyond Outcome Rewards: Step-Level Self-Distilled Policy Optimization for Deep Search Agents Unresolved cited work
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 21f1c732-01b6-48bf-a80d-1f7f50ab5d17 · outbound
Beyond Outcome Rewards: Step-Level Self-Distilled Policy Optimization for Deep Search Agents Unresolved cited work
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 11882a9e-120b-4e7f-865e-3da02dbbca80 · outbound
Beyond Outcome Rewards: Step-Level Self-Distilled Policy Optimization for Deep Search Agents who is the COOP leader in Amarillo
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 00f96cd1-8200-47cd-83c8-1ec6ff9fe91b · outbound
Beyond Outcome Rewards: Step-Level Self-Distilled Policy Optimization for Deep Search Agents Unresolved cited work
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 3668f22e-e1fc-404e-806f-09de3c9ffead · outbound
Beyond Outcome Rewards: Step-Level Self-Distilled Policy Optimization for Deep Search Agents Unresolved cited work
Reference 2026
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.