Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-12T14:37:28.660097Z
Paper Citation Record · LEDGER
As of 11 August 2026, this Paper Citation Record lists 40 of 40 outbound references and 2 inbound Pith citation observations for arXiv:2606.09348.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-12T14:37:28.660097Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-06T16:39:06.584525Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-06T16:39:08.145196Z
40 of 40 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 271df1c1-b80c-4a90-8d7a-cb2c275ae08e · outbound
PBSD: Privileged Bayesian Self-Distillation for Long-Horizon Credit Assignment DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b7838f1c-c8e0-472d-a45d-f95616640fa4 · outbound
PBSD: Privileged Bayesian Self-Distillation for Long-Horizon Credit Assignment DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d41e2f70-1e2a-4bb4-982c-1fc5d55cb135 · outbound
PBSD: Privileged Bayesian Self-Distillation for Long-Horizon Credit Assignment Tongyi DeepResearch Technical Report
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8eaafdb3-a240-47d2-b17b-b3e3fc1c1559 · outbound
PBSD: Privileged Bayesian Self-Distillation for Long-Horizon Credit Assignment DR-Venus: Towards Frontier Edge-Scale Deep Research Agents with Only 10K Open Data
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 99a86f73-10da-42c2-8c65-79136130593a · outbound
PBSD: Privileged Bayesian Self-Distillation for Long-Horizon Credit Assignment OpenSearch-VL: An Open Recipe for Frontier Multimodal Search Agents
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a3dbaaa8-63a2-4dda-a5d8-d83ef8d110e2 · outbound
PBSD: Privileged Bayesian Self-Distillation for Long-Horizon Credit Assignment Kimi K2.5: Visual Agentic Intelligence
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 58275e48-cd67-4043-9df3-9981930d8b7a · outbound
PBSD: Privileged Bayesian Self-Distillation for Long-Horizon Credit Assignment Mind DeepResearch Technical Report
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3bc5bcf0-576b-4e6e-9b8d-6c589443d14a · outbound
PBSD: Privileged Bayesian Self-Distillation for Long-Horizon Credit Assignment Tree search for llm agent reinforcement learning.arXiv preprint arXiv:2509.21240, 2025
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bcba510f-3d50-43de-90b2-ecabedac86b2 · outbound
PBSD: Privileged Bayesian Self-Distillation for Long-Horizon Credit Assignment Treerpo: Tree relative policy optimization
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9783179d-9fc5-4e62-a5b6-1fb73bd651b9 · outbound
PBSD: Privileged Bayesian Self-Distillation for Long-Horizon Credit Assignment On-policy distillation
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 56b852a0-4995-46d8-96a7-ab4d6300a29d · outbound
PBSD: Privileged Bayesian Self-Distillation for Long-Horizon Credit Assignment Rethinking On-Policy Distillation of Large Language Models: Phenomenology, Mechanism, and Recipe
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3f5ce916-b75e-4304-b226-4d294cf7e800 · outbound
PBSD: Privileged Bayesian Self-Distillation for Long-Horizon Credit Assignment Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f2dfa813-82f0-481b-a68e-f9b11a85bd46 · outbound
PBSD: Privileged Bayesian Self-Distillation for Long-Horizon Credit Assignment Self-Distilled RLVR
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 18c5bcec-c127-4fb6-85a2-f1f253c40672 · outbound
PBSD: Privileged Bayesian Self-Distillation for Long-Horizon Credit Assignment Criticsearch: Fine-grained credit assignment for search agents via a retrospective critic.arXiv preprint arXiv:2511.12159, 2025
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bb4595ac-9b8f-46c8-83b3-5f2d32aaf597 · outbound
PBSD: Privileged Bayesian Self-Distillation for Long-Horizon Credit Assignment RubricEM: Meta-RL with Rubric-guided Policy Decomposition beyond Verifiable Rewards
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 76d81e20-033a-4f34-8979-905514cf5d4a · outbound
PBSD: Privileged Bayesian Self-Distillation for Long-Horizon Credit Assignment Reward Hacking in Rubric-Based Reinforcement Learning
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d52464a9-a930-4b0c-8ca9-24415466a351 · outbound
PBSD: Privileged Bayesian Self-Distillation for Long-Horizon Credit Assignment StepSearch: Igniting LLMs Search Ability via Step-Wise Proximal Policy Optimization
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 41f54242-efd5-4ed6-86d4-1e1e6898a2bf · outbound
PBSD: Privileged Bayesian Self-Distillation for Long-Horizon Credit Assignment Reinforcing multi-turn reasoning in llm agents via turn- level reward design.arXiv preprint arXiv:2505.11821, 2025
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9d89e011-167c-4caa-8076-607574450878 · outbound
PBSD: Privileged Bayesian Self-Distillation for Long-Horizon Credit Assignment On-policy distillation of language models: Learning from self-generated mistakes
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 50974593-1df5-4bcb-9e25-714769de15db · outbound
PBSD: Privileged Bayesian Self-Distillation for Long-Horizon Credit Assignment Revisiting On-Policy Distillation: Empirical Failure Modes and Simple Fixes
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 80f0789b-3f89-47f5-beef-54f3a5c6a7bf · outbound
PBSD: Privileged Bayesian Self-Distillation for Long-Horizon Credit Assignment Reinforcement Learning via Self-Distillation
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4d30b6a0-8e11-456a-bc58-0883f2291ac4 · outbound
PBSD: Privileged Bayesian Self-Distillation for Long-Horizon Credit Assignment On-Policy Context Distillation for Language Models
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 450dffd7-2cd8-452e-9bbb-3b7aa4e0e614 · outbound
PBSD: Privileged Bayesian Self-Distillation for Long-Horizon Credit Assignment Self-distillation enables continual learning
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 45ea532e-4e9c-45a2-a9f5-3c661072cda0 · outbound
PBSD: Privileged Bayesian Self-Distillation for Long-Horizon Credit Assignment Privileged information distillation for language models
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2d906c2a-4658-45fd-99ba-3880d5d2aa2d · outbound
PBSD: Privileged Bayesian Self-Distillation for Long-Horizon Credit Assignment CRISP: Compressed Reasoning via Iterative Self-Policy Distillation
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fcd0fced-f996-43da-a381-533a4679ee28 · outbound
PBSD: Privileged Bayesian Self-Distillation for Long-Horizon Credit Assignment Mirothinker-1.7 & h1: Towards heavy-duty research agents via verification.arXiv preprint arXiv:2603.15726, 2026
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f6fd1f05-858b-4506-967e-3d4fcd0c2ded · outbound
PBSD: Privileged Bayesian Self-Distillation for Long-Horizon Credit Assignment MiroThinker: Pushing the Performance Boundaries of Open-Source Research Agents via Model, Context, and Interactive Scaling
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ea4bd47a-ef63-47f5-abe3-b9d1912d3f56 · outbound
PBSD: Privileged Bayesian Self-Distillation for Long-Horizon Credit Assignment Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2ad92bf4-9342-408e-bb14-2e64635209c3 · outbound
PBSD: Privileged Bayesian Self-Distillation for Long-Horizon Credit Assignment Deep- researcher: Scaling deep research via reinforcement learning in real-world environments
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e37dbe36-d3a8-4805-b5b3-099961d9c22c · outbound
PBSD: Privileged Bayesian Self-Distillation for Long-Horizon Credit Assignment Stabilizing moe reinforcement learning by aligning training and inference routers.arXiv e-prints, pages arXiv–2510, 2025
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f70d8dba-4c17-4acb-a0ff-3d1b11f825a1 · outbound
PBSD: Privileged Bayesian Self-Distillation for Long-Horizon Credit Assignment Openseeker: Democratizing frontier search agents by fully open-sourcing training data.arXiv preprint arXiv:2603.15594, 2026
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 797b6d42-dc69-4b74-8c2c-8de9f105a1d6 · outbound
PBSD: Privileged Bayesian Self-Distillation for Long-Horizon Credit Assignment BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bfae2e42-f81b-471e-862e-5d10667694f3 · outbound
PBSD: Privileged Bayesian Self-Distillation for Long-Horizon Credit Assignment Browsecomp-zh: Benchmarking web browsing ability of large language models in chinese
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 919d60fb-cd25-4d14-90cf-075ace980846 · outbound
PBSD: Privileged Bayesian Self-Distillation for Long-Horizon Credit Assignment Gaia: a benchmark for general ai assistants
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 154f2951-df48-4eeb-9675-884235253d8d · outbound
PBSD: Privileged Bayesian Self-Distillation for Long-Horizon Credit Assignment xbench: Tracking Agents Productivity Scaling with Profession-Aligned Real-World Evaluations
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c575b97d-a4a9-459a-b111-0819f8277cba · outbound
PBSD: Privileged Bayesian Self-Distillation for Long-Horizon Credit Assignment gpt-oss-120b & gpt-oss-20b Model Card
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4f27c5fc-cbe5-48ff-b64e-ab4e7b4beb19 · outbound
PBSD: Privileged Bayesian Self-Distillation for Long-Horizon Credit Assignment Qwen3 Technical Report
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8e5b875d-1eb9-4d53-b3e7-1a80f05fc2bb · outbound
PBSD: Privileged Bayesian Self-Distillation for Long-Horizon Credit Assignment Llamafactory: Unified efficient fine- tuning of 100+ language models
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d9bfc50b-5b94-4024-9a0f-5603a5456c53 · outbound
PBSD: Privileged Bayesian Self-Distillation for Long-Horizon Credit Assignment Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 937f6d50-f92b-4c9a-8b42-ef06584cd3f2 · outbound
PBSD: Privileged Bayesian Self-Distillation for Long-Horizon Credit Assignment param1":
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 78c16054-15cb-4cd2-b031-3895e2104ec9 · inbound
Fetch-then-Explore: Decoupling Selection from Extraction over a Persistent Workspace for Search Agents PBSD: Privileged Bayesian Self-Distillation for Long-Horizon Credit Assignment
Reference 78
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 38de8315-b88a-46a4-82dc-19a2ec9350d9 · inbound
Agentic Reinforcement Learning with Observation-Calibrated Self-Distillation PBSD: Privileged Bayesian Self-Distillation for Long-Horizon Credit Assignment
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.