Pith. sign in

Paper Citation Record · LEDGER

Experience Augmented Policy Optimization for LLM Reasoning

As of 5 August 2026, this Paper Citation Record lists 19 of 19 outbound references and 1 inbound Pith citation observation for arXiv:2606.30420.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2606.30420 v2

Coverage vector

measured 19 of 19 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-02T09:32:16.230452Z

measured 20 of 20 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T07:23:30.079906Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

19 of 19 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved18
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1b286e9c-2af6-409b-83e6-3e0dc315ae09 · outbound

This paper cites Accelerating Large Language Model Decoding with Speculative Sampling.

Experience Augmented Policy Optimization for LLM Reasoning Accelerating Large Language Model Decoding with Speculative Sampling

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-02T09:32:14.708381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T09:32:14.708381Z digest=sha256:0d82afcc4b67cd9e9fc5ca2068583ab157a4f9d33135b48dfe1db5b56b3c0a28

Observation c4049426-a120-4ed8-bbac-f7e342565507 · outbound

This paper cites Selective Expert Guidance for Effective and Diverse Exploration in Reinforcement Learning of LLMs.

Experience Augmented Policy Optimization for LLM Reasoning Selective Expert Guidance for Effective and Diverse Exploration in Reinforcement Learning of LLMs

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-02T09:32:14.999192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T09:32:14.999192Z digest=sha256:06cc79fa28012c9a913bb3c6d8ef4446f59bc2b417df8242077727b9aaa1e62b

Observation 7e20029b-a397-4cf4-a48f-87237f56c469 · outbound

This paper cites RePO: Replay-Enhanced Policy Optimization.

Experience Augmented Policy Optimization for LLM Reasoning RePO: Replay-Enhanced Policy Optimization

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-02T09:32:15.097798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T09:32:15.097798Z digest=sha256:07b15ecf51ca0a7479936000bf47686d758924af090bd48963646ebfe3db69fc

Observation 28a77773-05bb-45a8-9677-2b533f386e3b · outbound

This paper cites DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models.

Experience Augmented Policy Optimization for LLM Reasoning DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-02T09:32:15.163530Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T09:32:15.163530Z digest=sha256:91a9189775caf1472e981dda5a1e6684e3847d399f3c221190fc133e87c1b699

Observation 88ce81f3-b224-4549-a5a6-bd6f3475d16f · outbound

This paper cites Rethinking Visual Content Refinement in Low-Shot CLIP Adaptation.

Experience Augmented Policy Optimization for LLM Reasoning Rethinking Visual Content Refinement in Low-Shot CLIP Adaptation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-02T09:32:15.225984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T09:32:15.225984Z digest=sha256:c4c91b5dbd3bf5f2c22d71c379e11e53a01d37751634ed686ed1e83530cfdf72

Observation ebc7516d-45b9-46d4-ae73-c842eadf7ec7 · outbound

This paper cites https://thinkingmachines.ai/blog/on-policy- distillation.

Experience Augmented Policy Optimization for LLM Reasoning https://thinkingmachines.ai/blog/on-policy- distillation

Reference 10

Resolution
malformed identifier
no resolver link, observed 2026-08-02T09:32:15.408939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T09:32:15.408939Z digest=sha256:c16ec6e5da9d00375a127efe1c0ff7d43801df6e5a5cefc22081173c91e34b77

Observation 46386856-93fc-4390-8536-0a77a634251d · outbound

This paper cites Proximal Policy Optimization Algorithms.

Experience Augmented Policy Optimization for LLM Reasoning Proximal Policy Optimization Algorithms

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-02T09:32:15.469553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T09:32:15.469553Z digest=sha256:26c710394e020a70503498feceff3fa2ae006e6ac63740279df18a2146ac3b8a

Observation a216cc0c-c6e2-40fe-b94a-47990d50c1e2 · outbound

This paper cites Kimi K2: Open Agentic Intelligence.

Experience Augmented Policy Optimization for LLM Reasoning Kimi K2: Open Agentic Intelligence

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-02T09:32:15.610023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T09:32:15.610023Z digest=sha256:c42de1a17c89acdc4a0f70a456b6653fc3df8dbf8f4d2be41a231fa2b8f49578

Observation 6db0e7ed-3cb0-48b5-bb3d-71c657f378fe · outbound

This paper cites Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM Reasoning.

Experience Augmented Policy Optimization for LLM Reasoning Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM Reasoning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-02T09:32:15.699226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T09:32:15.699226Z digest=sha256:d6240b3ca824cb4b81995d23466dc0fc478f5958b849f7d56f2c2b7474a38d95

Observation 42e7bb3d-c57d-4e99-977a-cb77b34051a2 · outbound

This paper cites Quantile advantage estimation for entropy-safe reasoning.

Experience Augmented Policy Optimization for LLM Reasoning Quantile advantage estimation for entropy-safe reasoning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-02T09:32:15.760120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T09:32:15.760120Z digest=sha256:539d2a574560508e2610d79e90f0e16e1a536a42b0ef36ac85d9c307bbf22e30

Observation 8845f0ac-26f2-429b-8ffa-f9e0d774395e · outbound

This paper cites Qwen2.5 Technical Report.

Experience Augmented Policy Optimization for LLM Reasoning Qwen2.5 Technical Report

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-02T09:32:15.855961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T09:32:15.855961Z digest=sha256:a721ce80c3a8340659dab09824e927f093e721c59c8b25b672fa4bd21e2bfa91

Observation 45f2d38e-7cf3-41da-98eb-e44bb217ec08 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

Experience Augmented Policy Optimization for LLM Reasoning DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-02T09:32:15.949222Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T09:32:15.949222Z digest=sha256:5aadbd6aad74b3aae61269e8d960ab6a5409a9e6602219e846893aa279b446ec

Observation fd9f26d3-69e3-426e-a660-13a15d4e4569 · outbound

This paper cites F., and Cheng, Y.

Experience Augmented Policy Optimization for LLM Reasoning F., and Cheng, Y

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-02T09:32:16.137794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T09:32:16.137794Z digest=sha256:320c146327fb9fa7bba44daffc33fa801511edcc92c2d61191966d075ba66d9c

Observation e9e37322-2c65-4c7c-9231-098e61e842ed · outbound

This paper cites RLEP: Reinforcement Learning with Experience Replay for LLM Reasoning.

Experience Augmented Policy Optimization for LLM Reasoning RLEP: Reinforcement Learning with Experience Replay for LLM Reasoning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-02T09:32:16.230452Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T09:32:16.230452Z digest=sha256:9148ef51133f0fd1c28af00ee453c63f44551ca83405974b83fc6c33bdbc9079

Observation 2d8f3f81-f7d0-4d5a-a3a9-023ea96beba8 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Experience Augmented Policy Optimization for LLM Reasoning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-02T09:32:15.540371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T09:32:15.540371Z digest=sha256:0cf59cb2d6b3dd1368d7ea489da7ecbee20633a348715a3066bf926296718aac

Observation 762ab9cd-ba14-4cd8-b8a8-54b961580eeb · outbound

This paper cites Reasoning with Exploration: An Entropy Perspective.

Experience Augmented Policy Optimization for LLM Reasoning Reasoning with Exploration: An Entropy Perspective

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-02T09:32:14.770209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T09:32:14.770209Z digest=sha256:695ad2ebc0380ae17eb57e3be71d7d70ea91684e98e58d7c872c3b9ddc8055ea

Observation d05ecd30-59cf-4ae0-81dc-2c38046ca202 · outbound

This paper cites AdaViP: Aligning Multi-modal LLMs via Adaptive Vision-enhanced Preference Optimization.

Experience Augmented Policy Optimization for LLM Reasoning AdaViP: Aligning Multi-modal LLMs via Adaptive Vision-enhanced Preference Optimization

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-02T09:32:15.287630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T09:32:15.287630Z digest=sha256:bb4e263003d5d305ee1fb2c9959efe39c6b2c6cb24d97141dbc8aa6058ed887c

Observation a348c0b9-b62c-4b34-9f49-d751d50e7444 · outbound

This paper cites Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities.

Experience Augmented Policy Optimization for LLM Reasoning Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-02T09:32:14.841013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T09:32:14.841013Z digest=sha256:b5ade2493df45206d9abb9e75e8920a1b218dba784a064362c1e05a039ee846f

Observation cf5cb84c-5b23-49cd-9758-8fb274e39f43 · outbound

This paper cites OpenAI o1 System Card.

Experience Augmented Policy Optimization for LLM Reasoning OpenAI o1 System Card

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-02T09:32:14.933068Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T09:32:14.933068Z digest=sha256:9b901bcb7198851a40766f9a0f7fc3ffc12d94fc74e3052c72979f3b3043fe43

Pith citing papers

Observation 79e6168e-577f-4be5-ba8a-bb1cfab2dcca · inbound

ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples cites this paper.

ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples Experience Augmented Policy Optimization for LLM Reasoning

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-02T07:23:30.079906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T07:23:30.079906Z digest=sha256:eaf6019037c7eda414e36b27d49b99d3b7d91ab5d423985e3959706336c30997