Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-02T09:32:16.230452Z
Paper Citation Record · LEDGER
As of 5 August 2026, this Paper Citation Record lists 19 of 19 outbound references and 1 inbound Pith citation observation for arXiv:2606.30420.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-02T09:32:16.230452Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-02T07:23:30.079906Z
A source-named dated measurement, never combined with another source.
Source: cited_works
19 of 19 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 1b286e9c-2af6-409b-83e6-3e0dc315ae09 · outbound
Experience Augmented Policy Optimization for LLM Reasoning Accelerating Large Language Model Decoding with Speculative Sampling
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c4049426-a120-4ed8-bbac-f7e342565507 · outbound
Experience Augmented Policy Optimization for LLM Reasoning Selective Expert Guidance for Effective and Diverse Exploration in Reinforcement Learning of LLMs
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7e20029b-a397-4cf4-a48f-87237f56c469 · outbound
Experience Augmented Policy Optimization for LLM Reasoning RePO: Replay-Enhanced Policy Optimization
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 28a77773-05bb-45a8-9677-2b533f386e3b · outbound
Experience Augmented Policy Optimization for LLM Reasoning DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 88ce81f3-b224-4549-a5a6-bd6f3475d16f · outbound
Experience Augmented Policy Optimization for LLM Reasoning Rethinking Visual Content Refinement in Low-Shot CLIP Adaptation
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ebc7516d-45b9-46d4-ae73-c842eadf7ec7 · outbound
Experience Augmented Policy Optimization for LLM Reasoning https://thinkingmachines.ai/blog/on-policy- distillation
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 46386856-93fc-4390-8536-0a77a634251d · outbound
Experience Augmented Policy Optimization for LLM Reasoning Proximal Policy Optimization Algorithms
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a216cc0c-c6e2-40fe-b94a-47990d50c1e2 · outbound
Experience Augmented Policy Optimization for LLM Reasoning Kimi K2: Open Agentic Intelligence
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6db0e7ed-3cb0-48b5-bb3d-71c657f378fe · outbound
Experience Augmented Policy Optimization for LLM Reasoning Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM Reasoning
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 42e7bb3d-c57d-4e99-977a-cb77b34051a2 · outbound
Experience Augmented Policy Optimization for LLM Reasoning Quantile advantage estimation for entropy-safe reasoning
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8845f0ac-26f2-429b-8ffa-f9e0d774395e · outbound
Experience Augmented Policy Optimization for LLM Reasoning Qwen2.5 Technical Report
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 45f2d38e-7cf3-41da-98eb-e44bb217ec08 · outbound
Experience Augmented Policy Optimization for LLM Reasoning DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fd9f26d3-69e3-426e-a660-13a15d4e4569 · outbound
Experience Augmented Policy Optimization for LLM Reasoning F., and Cheng, Y
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e9e37322-2c65-4c7c-9231-098e61e842ed · outbound
Experience Augmented Policy Optimization for LLM Reasoning RLEP: Reinforcement Learning with Experience Replay for LLM Reasoning
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2d8f3f81-f7d0-4d5a-a3a9-023ea96beba8 · outbound
Experience Augmented Policy Optimization for LLM Reasoning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 2017
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 762ab9cd-ba14-4cd8-b8a8-54b961580eeb · outbound
Experience Augmented Policy Optimization for LLM Reasoning Reasoning with Exploration: An Entropy Perspective
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d05ecd30-59cf-4ae0-81dc-2c38046ca202 · outbound
Experience Augmented Policy Optimization for LLM Reasoning AdaViP: Aligning Multi-modal LLMs via Adaptive Vision-enhanced Preference Optimization
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a348c0b9-b62c-4b34-9f49-d751d50e7444 · outbound
Experience Augmented Policy Optimization for LLM Reasoning Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cf5cb84c-5b23-49cd-9758-8fb274e39f43 · outbound
Experience Augmented Policy Optimization for LLM Reasoning OpenAI o1 System Card
Reference 2026
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 79e6168e-577f-4be5-ba8a-bb1cfab2dcca · inbound
ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples Experience Augmented Policy Optimization for LLM Reasoning
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.