Pith. sign in

Paper Citation Record · LEDGER

One-Way Policy Optimization for Self-Evolving LLMs

As of 5 August 2026, this Paper Citation Record lists 21 of 21 outbound references and 1 inbound Pith citation observation for arXiv:2605.22156.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.22156 v1

Coverage vector

measured 21 of 21 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-22T08:01:05.911650Z

measured 22 of 22 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T07:23:30.215870Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

21 of 21 outbound references displayed

  • verified exact19
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 44906f4b-b4bc-4132-b2a4-7acccd1d3ac0 · outbound

This paper cites MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention.

One-Way Policy Optimization for Self-Evolving LLMs MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-22T08:01:15.588609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T08:01:05.911650Z digest=sha256:04716317b0e861129525970c38d7ee7dfcd1b0f18997e6ac57e2bd033ce83d7b

Observation d9315f31-6061-459b-b04f-c92e4489b73b · outbound

This paper cites Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities.

One-Way Policy Optimization for Self-Evolving LLMs Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-22T08:01:15.504835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T08:01:05.911650Z digest=sha256:3aa8c498de8b3c468c1f05c58eb51b8e3ebb0ba642ee81525e1b2985f4ddf832

Observation d3ec8921-d40b-4556-8241-8312b01460a5 · outbound

This paper cites Reinforced Self-Training (ReST) for Language Modeling.

One-Way Policy Optimization for Self-Evolving LLMs Reinforced Self-Training (ReST) for Language Modeling

Reference 3

Resolution
metadata mismatch
local_arxiv, observed 2026-05-22T08:01:15.500354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T08:01:05.911650Z digest=sha256:b8d4b8a42a1e3ff670323dd4abee4a5a9d1b73332bfc1a4c9dca2caf08e478d3

Observation d87eb0ca-b935-41bc-b11c-5effbd7a2150 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

One-Way Policy Optimization for Self-Evolving LLMs DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-22T08:01:15.532477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T08:01:05.911650Z digest=sha256:43cb5ea9b289938ad837e2e4ea2f43f87052c59170f6bceba2c073672ae04cdb

Observation 838b9c2f-661a-4e92-b41e-787ffd25f5aa · outbound

This paper cites A sober look at progress in language model reasoning: Pitfalls and paths to repro- ducibility.arXiv preprint arXiv:2504.07086.

One-Way Policy Optimization for Self-Evolving LLMs A sober look at progress in language model reasoning: Pitfalls and paths to repro- ducibility.arXiv preprint arXiv:2504.07086

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-22T08:01:15.496215Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T08:01:05.911650Z digest=sha256:f94f4beda0f05037527881b7ba1eda2d78e43bc91287fa55b17e537493742e2c

Observation a1765dc9-b917-4675-bccc-74c6cc5f96e3 · outbound

This paper cites On the direction of rlvr updates for llm reasoning: Identification and exploitation.

One-Way Policy Optimization for Self-Evolving LLMs On the direction of rlvr updates for llm reasoning: Identification and exploitation

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-22T08:01:15.487068Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T08:01:05.911650Z digest=sha256:5e6537b5e3ab6dad8878b0169bfb9fa30285114f64f2caf5e2a64983ec82f257

Observation 86465566-b956-45ce-89f9-5f4a0b3e4965 · outbound

This paper cites Qwen2.5-Coder Technical Report.

One-Way Policy Optimization for Self-Evolving LLMs Qwen2.5-Coder Technical Report

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-22T08:01:15.559609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T08:01:05.911650Z digest=sha256:61ae2c1f7c529aea35d06efc5de26600cc6821efad1d8d1c78c5884ab91f7848

Observation b755f437-3838-4efd-b328-b6a832c6c774 · outbound

This paper cites OpenAI o1 System Card.

One-Way Policy Optimization for Self-Evolving LLMs OpenAI o1 System Card

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-22T08:01:15.491515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T08:01:05.911650Z digest=sha256:cfce95f1989c7a0ab8205d912cc8b1cdf466faf0b81625ce1fadac65efb3e88a

Observation e7c4a0ce-b494-4da4-b36a-755779ab3bd0 · outbound

This paper cites On-policy distillation.Thinking Machines Lab: Con- nectionism.

One-Way Policy Optimization for Self-Evolving LLMs On-policy distillation.Thinking Machines Lab: Con- nectionism

Reference 9

Resolution
verified exact
doi, observed 2026-05-22T08:01:15.044250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T08:01:05.911650Z digest=sha256:c0efc1efe77735bf1d3dd728717264b362b64fe629055ad459888910389cfc2e

Observation 9554dc8e-6f30-4035-a952-7b05cd74a3ba · outbound

This paper cites Fipo: Eliciting deep reasoning with future-kl influenced policy optimization.

One-Way Policy Optimization for Self-Evolving LLMs Fipo: Eliciting deep reasoning with future-kl influenced policy optimization

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-22T08:01:15.541515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T08:01:05.911650Z digest=sha256:a830b29f3ef4d4f31cd550934882854191faae29180cff5a8f644efad318fd85

Observation d09a2380-f0d0-4713-ae77-44f835805bdd · outbound

This paper cites Sparse but critical: A token-level analysis of distributional shifts in rlvr fine-tuning of llms.

One-Way Policy Optimization for Self-Evolving LLMs Sparse but critical: A token-level analysis of distributional shifts in rlvr fine-tuning of llms

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-22T08:01:15.519492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T08:01:05.911650Z digest=sha256:5925b786a9b46c0cbf6d92dfc5c917c282728bea9f073dd2c47c4df878d97a17

Observation 6f27696a-61bc-4463-8705-cf33591c635b · outbound

This paper cites Proximal Policy Optimization Algorithms.

One-Way Policy Optimization for Self-Evolving LLMs Proximal Policy Optimization Algorithms

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-22T08:01:15.554340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T08:01:05.911650Z digest=sha256:94f84a31b841b3e56d2acb89b6470919fdab39616964807f372fc0d18ca5a7d3

Observation 5f32f5c9-9b31-4470-8148-6501ad6412ca · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

One-Way Policy Optimization for Self-Evolving LLMs DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-05-22T08:01:15.523399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T08:01:05.911650Z digest=sha256:d71d594166f1c371dc45376fbb9e2e7f2ff75a0e31e98ceab2b910c9b02df57f

Observation 07623a88-bd20-4eef-8c06-94a642055627 · outbound

This paper cites Kimi K2: Open Agentic Intelligence.

One-Way Policy Optimization for Self-Evolving LLMs Kimi K2: Open Agentic Intelligence

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-05-22T08:01:15.545821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T08:01:05.911650Z digest=sha256:d0a634ac23f513c359ffb4877aec437605e2e7f764ab1d1272fdf789931f438a

Observation 657bde64-5f1e-40e0-be6b-cb2fb25d9541 · outbound

This paper cites MiMo-V2-Flash Technical Report.

One-Way Policy Optimization for Self-Evolving LLMs MiMo-V2-Flash Technical Report

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-22T08:01:15.513978Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T08:01:05.911650Z digest=sha256:551dba8198450bbe6d40c5bca7a2dd54d41dd0694621fa29a7edac8d858fbc3b

Observation b3a484d2-b832-4b9e-baa0-f76a8cebdf82 · outbound

This paper cites KDRL: Post-Training Reasoning LLMs via Unified Knowledge Distillation and Reinforcement Learning.

One-Way Policy Optimization for Self-Evolving LLMs KDRL: Post-Training Reasoning LLMs via Unified Knowledge Distillation and Reinforcement Learning

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-22T08:01:15.509689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T08:01:05.911650Z digest=sha256:1f74da4d64f4807f8ec2bbff50328b5f51c88d7f7de156f8f389c16c9c613787

Observation a5e60214-0376-4e5e-a286-0c3200dd1fce · outbound

This paper cites Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement.

One-Way Policy Optimization for Self-Evolving LLMs Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-05-22T08:01:15.536711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T08:01:05.911650Z digest=sha256:2cb1a2f1e04c4add458dbb2a0184f43a2c6fb2234ab2fe3e2daae65fcc6b6e68

Observation 3d9a9f6e-4919-4b1c-8675-77baf533b850 · outbound

This paper cites Qwen3 Technical Report.

One-Way Policy Optimization for Self-Evolving LLMs Qwen3 Technical Report

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-05-22T08:01:15.582276Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T08:01:05.911650Z digest=sha256:f3b08dd17afd4ee46f755390e04025a71c51150a3de62975ab74d821af54559d

Observation e4eb10b9-7510-470a-8ac1-d280943eb846 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

One-Way Policy Optimization for Self-Evolving LLMs DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-05-22T08:01:15.527668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T08:01:05.911650Z digest=sha256:d5e10381040b8cf83a2099d688a98bce1a6717f9ba76db0a32946fe29221cbd8

Observation e90d2e06-902c-4836-b0ff-8ea552de5d1c · outbound

This paper cites Group Sequence Policy Optimization.

One-Way Policy Optimization for Self-Evolving LLMs Group Sequence Policy Optimization

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-05-22T08:01:15.592991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T08:01:05.911650Z digest=sha256:ef5c5bfa7ec75e42d0fd2149cf4cfbc21ee595bae239557e1fa635f640a440e1

Observation 9fa3fdf8-75ef-443a-a160-cd62699f2212 · outbound

This paper cites For the training dataset, we utilize dapo-math-17kacross all main experiments.

One-Way Policy Optimization for Self-Evolving LLMs For the training dataset, we utilize dapo-math-17kacross all main experiments

Reference 21

Resolution
malformed identifier
arxiv_id, observed 2026-05-22T08:01:15.549942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T08:01:05.911650Z digest=sha256:370d808c7ba1496fcd7e7c984511b5cd82d7a841df26ffd04e11249744e18cf7

Pith citing papers

Observation 9b5031bc-8e80-4294-bedd-c73a5983a727 · inbound

ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples cites this paper.

ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples One-Way Policy Optimization for Self-Evolving LLMs

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-02T07:23:30.215870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T07:23:30.215870Z digest=sha256:4ff8b299fefd96f258105874e74576bcf5d917b59542d1ca58ad5b3e1d5f327c