Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-22T08:01:05.911650Z
Paper Citation Record · LEDGER
As of 5 August 2026, this Paper Citation Record lists 21 of 21 outbound references and 1 inbound Pith citation observation for arXiv:2605.22156.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-22T08:01:05.911650Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-02T07:23:30.215870Z
A source-named dated measurement, never combined with another source.
Source: cited_works
21 of 21 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 44906f4b-b4bc-4132-b2a4-7acccd1d3ac0 · outbound
One-Way Policy Optimization for Self-Evolving LLMs MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation d9315f31-6061-459b-b04f-c92e4489b73b · outbound
One-Way Policy Optimization for Self-Evolving LLMs Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation d3ec8921-d40b-4556-8241-8312b01460a5 · outbound
One-Way Policy Optimization for Self-Evolving LLMs Reinforced Self-Training (ReST) for Language Modeling
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation d87eb0ca-b935-41bc-b11c-5effbd7a2150 · outbound
One-Way Policy Optimization for Self-Evolving LLMs DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 838b9c2f-661a-4e92-b41e-787ffd25f5aa · outbound
One-Way Policy Optimization for Self-Evolving LLMs A sober look at progress in language model reasoning: Pitfalls and paths to repro- ducibility.arXiv preprint arXiv:2504.07086
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation a1765dc9-b917-4675-bccc-74c6cc5f96e3 · outbound
One-Way Policy Optimization for Self-Evolving LLMs On the direction of rlvr updates for llm reasoning: Identification and exploitation
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 86465566-b956-45ce-89f9-5f4a0b3e4965 · outbound
One-Way Policy Optimization for Self-Evolving LLMs Qwen2.5-Coder Technical Report
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b755f437-3838-4efd-b328-b6a832c6c774 · outbound
One-Way Policy Optimization for Self-Evolving LLMs OpenAI o1 System Card
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation e7c4a0ce-b494-4da4-b36a-755779ab3bd0 · outbound
One-Way Policy Optimization for Self-Evolving LLMs On-policy distillation.Thinking Machines Lab: Con- nectionism
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 9554dc8e-6f30-4035-a952-7b05cd74a3ba · outbound
One-Way Policy Optimization for Self-Evolving LLMs Fipo: Eliciting deep reasoning with future-kl influenced policy optimization
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation d09a2380-f0d0-4713-ae77-44f835805bdd · outbound
One-Way Policy Optimization for Self-Evolving LLMs Sparse but critical: A token-level analysis of distributional shifts in rlvr fine-tuning of llms
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 6f27696a-61bc-4463-8705-cf33591c635b · outbound
One-Way Policy Optimization for Self-Evolving LLMs Proximal Policy Optimization Algorithms
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 5f32f5c9-9b31-4470-8148-6501ad6412ca · outbound
One-Way Policy Optimization for Self-Evolving LLMs DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 07623a88-bd20-4eef-8c06-94a642055627 · outbound
One-Way Policy Optimization for Self-Evolving LLMs Kimi K2: Open Agentic Intelligence
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 657bde64-5f1e-40e0-be6b-cb2fb25d9541 · outbound
One-Way Policy Optimization for Self-Evolving LLMs MiMo-V2-Flash Technical Report
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b3a484d2-b832-4b9e-baa0-f76a8cebdf82 · outbound
One-Way Policy Optimization for Self-Evolving LLMs KDRL: Post-Training Reasoning LLMs via Unified Knowledge Distillation and Reinforcement Learning
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation a5e60214-0376-4e5e-a286-0c3200dd1fce · outbound
One-Way Policy Optimization for Self-Evolving LLMs Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 3d9a9f6e-4919-4b1c-8675-77baf533b850 · outbound
One-Way Policy Optimization for Self-Evolving LLMs Qwen3 Technical Report
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation e4eb10b9-7510-470a-8ac1-d280943eb846 · outbound
One-Way Policy Optimization for Self-Evolving LLMs DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation e90d2e06-902c-4836-b0ff-8ea552de5d1c · outbound
One-Way Policy Optimization for Self-Evolving LLMs Group Sequence Policy Optimization
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 9fa3fdf8-75ef-443a-a160-cd62699f2212 · outbound
One-Way Policy Optimization for Self-Evolving LLMs For the training dataset, we utilize dapo-math-17kacross all main experiments
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 9b5031bc-8e80-4294-bedd-c73a5983a727 · inbound
ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples One-Way Policy Optimization for Self-Evolving LLMs
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.