Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-06-28T23:25:30.618655Z
Paper Citation Record · LEDGER
As of 23 August 2026, this Paper Citation Record lists 14 of 14 outbound references and 1 inbound Pith citation observation for arXiv:2605.30859.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-06-28T23:25:30.618655Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-01T06:48:16.678850Z
A source-named dated measurement, never combined with another source.
Source: cited_works
14 of 14 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation fa931942-a04f-4c4d-b214-8fc82ba12925 · outbound
DARTS: Distribution-Aware Active Rollout Trajectory Shaping for Accelerating LLM Reinforcement Learning Training Verifiers to Solve Math Word Problems
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 8226b003-ce25-4d5a-acca-88de5bf100e2 · outbound
DARTS: Distribution-Aware Active Rollout Trajectory Shaping for Accelerating LLM Reinforcement Learning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 5e9395cb-3585-40f2-b4a6-27a8ef09b296 · outbound
DARTS: Distribution-Aware Active Rollout Trajectory Shaping for Accelerating LLM Reinforcement Learning AsyncFlow: An Asynchronous Streaming RL Framework for Efficient LLM Post-Training
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation fed96b4c-3f39-42f9-a99e-59ac890accd3 · outbound
DARTS: Distribution-Aware Active Rollout Trajectory Shaping for Accelerating LLM Reinforcement Learning OpenAI o1 System Card
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 03eaa09e-d3c0-47ac-a641-c5585ce9c440 · outbound
DARTS: Distribution-Aware Active Rollout Trajectory Shaping for Accelerating LLM Reinforcement Learning Let’s verify step by step
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 727266ea-1d75-4775-b660-44b8b1e5b9ac · outbound
DARTS: Distribution-Aware Active Rollout Trajectory Shaping for Accelerating LLM Reinforcement Learning Part ii: Roll flash–accelerating rlvr and agentic training with asynchrony
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 60dff0bd-e376-47ee-b820-35dbb75640bd · outbound
DARTS: Distribution-Aware Active Rollout Trajectory Shaping for Accelerating LLM Reinforcement Learning Proximal Policy Optimization Algorithms
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation d654de1b-55a5-4e19-93be-47b8172e5883 · outbound
DARTS: Distribution-Aware Active Rollout Trajectory Shaping for Accelerating LLM Reinforcement Learning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 8efc664a-946f-4c94-aa80-9ce234bc3154 · outbound
DARTS: Distribution-Aware Active Rollout Trajectory Shaping for Accelerating LLM Reinforcement Learning Kimi K2: Open Agentic Intelligence
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation c2ea108f-23fc-4e1a-9d94-96e3e760afc8 · outbound
DARTS: Distribution-Aware Active Rollout Trajectory Shaping for Accelerating LLM Reinforcement Learning Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 907a62d6-9a38-4266-9f58-e8fbc157a8d2 · outbound
DARTS: Distribution-Aware Active Rollout Trajectory Shaping for Accelerating LLM Reinforcement Learning Qwen3 Technical Report
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 313d458a-1224-4b4d-8c51-f82bee97f89a · outbound
DARTS: Distribution-Aware Active Rollout Trajectory Shaping for Accelerating LLM Reinforcement Learning StreamRL: Scalable, Heterogeneous, and Elastic RL for LLMs with Disaggregated Stream Generation
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 6a9e9a85-78f6-444e-a77b-19276ff35ef2 · outbound
DARTS: Distribution-Aware Active Rollout Trajectory Shaping for Accelerating LLM Reinforcement Learning DB-GPT: Large language model meets database
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation fb1f1dd5-bc9c-4bbe-a916-ecc785a7e39b · outbound
DARTS: Distribution-Aware Active Rollout Trajectory Shaping for Accelerating LLM Reinforcement Learning This trend confirms that our method effectively mitigates the verbose long-tail issue early on, greatly enhancing training efficiency
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 730dca30-1ac0-4e30-ae2b-17eb3ba55269 · inbound
QLPO: Quadrant-weighted Sampling for Length-aware Policy Optimization DARTS: Distribution-Aware Active Rollout Trajectory Shaping for Accelerating LLM Reinforcement Learning
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.