Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-11T14:06:50.522833Z
Paper Citation Record · LEDGER
As of 5 August 2026, this Paper Citation Record lists 27 of 27 outbound references and 0 inbound Pith citation observations for arXiv:2607.04751.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-11T14:06:50.522833Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
27 of 27 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 50d9bc73-025b-4908-8b91-bf8db5a8b244 · outbound
Trust Region Policy Distillation Deepseek-v4: Towards highly efficient million-token context intelligence, 2026
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 45a6f1f5-e1ee-4487-81b7-3095a2c82701 · outbound
Trust Region Policy Distillation Revisiting On-Policy Distillation: Empirical Failure Modes and Simple Fixes
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a3834600-6e66-434f-bcfb-fc6935b88551 · outbound
Trust Region Policy Distillation MiniLLM: Knowledge distillation of large language models
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 94f8dc6f-a5d1-4b87-a35b-e4f0dd30cbb1 · outbound
Trust Region Policy Distillation DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 60f34f8a-e0a9-44c5-856c-7a84f9ac2310 · outbound
Trust Region Policy Distillation Stable On-Policy Distillation through Adaptive Target Reformulation
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 44d491eb-2600-47b8-b29b-17eb3d9f1767 · outbound
Trust Region Policy Distillation Approximately optimal approximate reinforcement learning
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d6a37485-07be-4e10-8dfc-d26184901118 · outbound
Trust Region Policy Distillation Scaling reasoning efficiently via relaxed on-policy distillation.arXiv preprint arXiv:2603.11137, 2026
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b2b9e8c0-55e7-4d50-9258-c60f58956d2b · outbound
Trust Region Policy Distillation Efficient memory management for large language model serving with pagedattention
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f7f7d7cd-9fac-4721-b148-8f9a5b4f4dc4 · outbound
Trust Region Policy Distillation Rethinking On-Policy Distillation of Large Language Models: Phenomenology, Mechanism, and Recipe
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5b251c66-e121-4082-8b9a-823cbd264d83 · outbound
Trust Region Policy Distillation On-policy distillation.Thinking Machines Lab: Con- nectionism, 2025
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d49b585c-a0c8-45de-9ed6-f5fa94f645d2 · outbound
Trust Region Policy Distillation Trust region policy optimization
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bf3b2550-5ff7-4784-b776-4a191fdde080 · outbound
Trust Region Policy Distillation Proximal Policy Optimization Algorithms
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation afd0ed7b-01ba-4327-9cab-59f833d905c2 · outbound
Trust Region Policy Distillation DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 63e02167-10e1-4612-9d5c-85f7d2d2ef8f · outbound
Trust Region Policy Distillation Self-Distillation Enables Continual Learning
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 65d21487-4e9b-434c-8227-b22a82babf2a · outbound
Trust Region Policy Distillation Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b153a3be-9986-448e-b601-1415b55cc3f0 · outbound
Trust Region Policy Distillation MiMo-V2-Flash Technical Report
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 36fe3aac-f7d1-47ab-885e-6d4e50b1a936 · outbound
Trust Region Policy Distillation Simple policy optimization
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 43792c86-26fd-407b-ad2f-bbd065be67eb · outbound
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2d2d46c9-a650-4c39-8ad9-1c94a96a044a · outbound
Trust Region Policy Distillation Nemotron-cascade 2: Post-training llms with cascade rl and multi-domain on-policy distillation.arXiv preprint arXiv:2603.19220, 2026
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a83808d3-167f-418c-a116-17a9d7b50b21 · outbound
Trust Region Policy Distillation DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 15a3054d-25f5-4980-90e9-a0396ab1ffb0 · outbound
Trust Region Policy Distillation GLM-5: from Vibe Coding to Agentic Engineering
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8f6a5ad7-ef3a-42bf-b02b-4a0d8abc3c07 · outbound
Trust Region Policy Distillation PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ded7804d-05a7-4db6-8cb3-26de4fa5ee7e · outbound
Trust Region Policy Distillation Sglang: Efficient execution of structured language model programs.Advances in neural information processing systems, 37: 62557–62583, 2024
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fba2cae8-6ce9-487d-9d2b-fadc178e8967 · outbound
Trust Region Policy Distillation Unresolved cited work
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5a8cda0c-cbe3-43ea-8c65-be5b750ea4b9 · outbound
Trust Region Policy Distillation Since limt→1+ w(t) = 0 , w(t)<0for allt∈(1,2)
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e00aaba0-1cd5-4ae5-ad0b-449741e1e412 · outbound
Trust Region Policy Distillation Unresolved cited work
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 50a270dc-2fab-4272-9fc5-41f210f48b29 · outbound
Trust Region Policy Distillation Since there is only one critical point, this unique stationary point corresponds to the global maximum of f(u)
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.