Pith. sign in

Paper Citation Record · LEDGER

DARTS: Distribution-Aware Active Rollout Trajectory Shaping for Accelerating LLM Reinforcement Learning

As of 23 August 2026, this Paper Citation Record lists 14 of 14 outbound references and 1 inbound Pith citation observation for arXiv:2605.30859.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.30859 v1

Coverage vector

measured 14 of 14 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-28T23:25:30.618655Z

measured 15 of 15 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-01T06:48:16.678850Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

14 of 14 outbound references displayed

  • verified exact11
  • verified fuzzy0
  • unresolved2
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation fa931942-a04f-4c4d-b214-8fc82ba12925 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

DARTS: Distribution-Aware Active Rollout Trajectory Shaping for Accelerating LLM Reinforcement Learning Training Verifiers to Solve Math Word Problems

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-06-28T23:42:50.088583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-28T23:25:30.618655Z digest=sha256:eca01697832c78e28614bde95e6b833e3d3d312d3920af6a40e7723d3d016a9a

Observation 8226b003-ce25-4d5a-acca-88de5bf100e2 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

DARTS: Distribution-Aware Active Rollout Trajectory Shaping for Accelerating LLM Reinforcement Learning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-06-28T23:42:50.083973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-28T23:25:30.618655Z digest=sha256:af33db4a4e0bcb7183b1ca337402a84182cc141bd5ea7cb0175522ad2e8260f3

Observation 5e9395cb-3585-40f2-b4a6-27a8ef09b296 · outbound

This paper cites AsyncFlow: An Asynchronous Streaming RL Framework for Efficient LLM Post-Training.

DARTS: Distribution-Aware Active Rollout Trajectory Shaping for Accelerating LLM Reinforcement Learning AsyncFlow: An Asynchronous Streaming RL Framework for Efficient LLM Post-Training

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-06-28T23:42:50.076818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-28T23:25:30.618655Z digest=sha256:3b3e5bc611a18b6184f811ceddc19bf80e3156cc9ab90d30fe216204dc329e90

Observation fed96b4c-3f39-42f9-a99e-59ac890accd3 · outbound

This paper cites OpenAI o1 System Card.

DARTS: Distribution-Aware Active Rollout Trajectory Shaping for Accelerating LLM Reinforcement Learning OpenAI o1 System Card

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-06-28T23:42:50.078985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-28T23:25:30.618655Z digest=sha256:365a9cdce8896af08f5e8f525226da366047625baada9404a6103f7929257074

Observation 03eaa09e-d3c0-47ac-a641-c5585ce9c440 · outbound

This paper cites Let’s verify step by step.

DARTS: Distribution-Aware Active Rollout Trajectory Shaping for Accelerating LLM Reinforcement Learning Let’s verify step by step

Reference 5

Resolution
unresolved
no resolver link, observed 2026-06-28T23:25:30.618655Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T23:25:30.618655Z digest=sha256:eeb76cd3411c2f52caaf583b33c1138b3ec691fc020c1f39062d60371f8fd508

Observation 727266ea-1d75-4775-b660-44b8b1e5b9ac · outbound

This paper cites Part ii: Roll flash–accelerating rlvr and agentic training with asynchrony.

DARTS: Distribution-Aware Active Rollout Trajectory Shaping for Accelerating LLM Reinforcement Learning Part ii: Roll flash–accelerating rlvr and agentic training with asynchrony

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-06-28T23:42:50.081473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-28T23:25:30.618655Z digest=sha256:8a644299c79c73975ed3933519320758b90e9acb0fba1a1d1c7941e5df04fc7a

Observation 60dff0bd-e376-47ee-b820-35dbb75640bd · outbound

This paper cites Proximal Policy Optimization Algorithms.

DARTS: Distribution-Aware Active Rollout Trajectory Shaping for Accelerating LLM Reinforcement Learning Proximal Policy Optimization Algorithms

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-06-28T23:42:50.086265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-28T23:25:30.618655Z digest=sha256:473206b4587fdad11040085cc1a6b824741bdb1364f957f176fa8ebd2ee9bb68

Observation d654de1b-55a5-4e19-93be-47b8172e5883 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

DARTS: Distribution-Aware Active Rollout Trajectory Shaping for Accelerating LLM Reinforcement Learning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-06-28T23:42:50.090793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-28T23:25:30.618655Z digest=sha256:f673ba36d56917e2a3c754200ac701888e02d406c1ce70f3f59b66c12d3f5614

Observation 8efc664a-946f-4c94-aa80-9ce234bc3154 · outbound

This paper cites Kimi K2: Open Agentic Intelligence.

DARTS: Distribution-Aware Active Rollout Trajectory Shaping for Accelerating LLM Reinforcement Learning Kimi K2: Open Agentic Intelligence

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-06-28T23:42:50.093065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-28T23:25:30.618655Z digest=sha256:1165d94b5ba58dc827ed388f2ffc9ffd9ba69bee32bab8aa5c59abfdbd8cfa22

Observation c2ea108f-23fc-4e1a-9d94-96e3e760afc8 · outbound

This paper cites Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement.

DARTS: Distribution-Aware Active Rollout Trajectory Shaping for Accelerating LLM Reinforcement Learning Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-06-28T23:42:50.096061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-28T23:25:30.618655Z digest=sha256:7e312a5abe9cf5ab636f2c7e71db99fe25563d2ef83e8f9df415128235d3a6db

Observation 907a62d6-9a38-4266-9f58-e8fbc157a8d2 · outbound

This paper cites Qwen3 Technical Report.

DARTS: Distribution-Aware Active Rollout Trajectory Shaping for Accelerating LLM Reinforcement Learning Qwen3 Technical Report

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-06-28T23:42:50.098836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-28T23:25:30.618655Z digest=sha256:9397c89ba7e2a0faa0755391ebf7f57155d4fcaa4028ab5f56ce40a50980220a

Observation 313d458a-1224-4b4d-8c51-f82bee97f89a · outbound

This paper cites StreamRL: Scalable, Heterogeneous, and Elastic RL for LLMs with Disaggregated Stream Generation.

DARTS: Distribution-Aware Active Rollout Trajectory Shaping for Accelerating LLM Reinforcement Learning StreamRL: Scalable, Heterogeneous, and Elastic RL for LLMs with Disaggregated Stream Generation

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-06-28T23:42:50.074123Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-28T23:25:30.618655Z digest=sha256:018563155734ab72c34ec936d20473e988b5253c45269cac2fcd0bc431de019f

Observation 6a9e9a85-78f6-444e-a77b-19276ff35ef2 · outbound

This paper cites DB-GPT: Large language model meets database.

DARTS: Distribution-Aware Active Rollout Trajectory Shaping for Accelerating LLM Reinforcement Learning DB-GPT: Large language model meets database

Reference 13

Resolution
metadata mismatch
doi, observed 2026-06-28T23:32:46.860427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-28T23:25:30.618655Z digest=sha256:b63857803cb3a6d43d1ca590b8a228917d841b45a874503703f24c982d3d26a0

Observation fb1f1dd5-bc9c-4bbe-a916-ecc785a7e39b · outbound

This paper cites This trend confirms that our method effectively mitigates the verbose long-tail issue early on, greatly enhancing training efficiency.

DARTS: Distribution-Aware Active Rollout Trajectory Shaping for Accelerating LLM Reinforcement Learning This trend confirms that our method effectively mitigates the verbose long-tail issue early on, greatly enhancing training efficiency

Reference 14

Resolution
unresolved
no resolver link, observed 2026-06-28T23:25:30.618655Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T23:25:30.618655Z digest=sha256:bb6417cdfb024a4a484e8dae684d91c58917368b1e33fbc2362b3e4b0e336541

Pith citing papers

Observation 730dca30-1ac0-4e30-ae2b-17eb3ba55269 · inbound

QLPO: Quadrant-weighted Sampling for Length-aware Policy Optimization cites this paper.

QLPO: Quadrant-weighted Sampling for Length-aware Policy Optimization DARTS: Distribution-Aware Active Rollout Trajectory Shaping for Accelerating LLM Reinforcement Learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-01T06:48:16.678850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T06:48:16.678850Z digest=sha256:9093a7e122582e4572d699f463e890e5008e8a14481f36ed43d57266a1ae5f67