Pith. sign in

Paper Citation Record · LEDGER

Accelerating RL for LLM Reasoning with Optimal Advantage Regression

As of 11 August 2026, this Paper Citation Record lists 13 of 13 outbound references and 3 inbound Pith citation observations for arXiv:2505.20686.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.20686 v1

Coverage vector

measured 13 of 13 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:55:57.728078Z

measured 16 of 16 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T06:47:16.744697Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-18T08:36:07.284835Z

Reference resolution

13 of 13 outbound references displayed

  • verified exact0
  • verified fuzzy7
  • unresolved6
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e81a2ab3-e7be-4a5c-b437-cf9bcd404833 · outbound

This paper cites an unresolved cited work.

Accelerating RL for LLM Reasoning with Optimal Advantage Regression Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:56:00.295544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T13:55:56.898779Z digest=sha256:424e882cde9d98e1693177ee7aec1178243b3b58ba4d008ca94e438746c68361

Observation 66aec46c-8ef6-45f2-8d3c-ebff8e96057e · outbound

This paper cites So,a2 = (−2)2 = 4.

Accelerating RL for LLM Reasoning with Optimal Advantage Regression So,a2 = (−2)2 = 4

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:55:59.902129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T13:55:57.019407Z digest=sha256:730c6714d1369389a69cffd4f9d96705ce83a97590c50e1e685d3925454f2c17

Observation d84a8ba9-54ec-48b2-9408-3179fe2abd2e · outbound

This paper cites an unresolved cited work.

Accelerating RL for LLM Reasoning with Optimal Advantage Regression Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:55:58.688420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T13:55:57.481389Z digest=sha256:68e51138a50167b6ccbed1d0e53ba241b78a74f00eb95786caec6421381ad671

Observation 325b6916-991d-49d0-b391-836cad9fdfb5 · outbound

This paper cites So,1 b2 = 1 32 = 1 9.

Accelerating RL for LLM Reasoning with Optimal Advantage Regression So,1 b2 = 1 32 = 1 9

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:56:00.109577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T13:55:56.962513Z digest=sha256:b6d5b77ae957842ff877705d0393e6f0dba739d088fe6e132b56e3a4e6ee12d5

Observation 5dc3bda0-7723-4f97-9ff4-08425c39f29a · outbound

This paper cites Since3≤b≤5 , the minimum value ofb is 3.

Accelerating RL for LLM Reasoning with Optimal Advantage Regression Since3≤b≤5 , the minimum value ofb is 3

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:55:59.679835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T13:55:57.100364Z digest=sha256:4d30d0c6179c578c8a58dce1722745226be71cd22b86fd57aa8e6d32d0ca5b9f

Observation 4af27d99-0b7b-4316-86b3-b1c4283560c3 · outbound

This paper cites Since−6≤a≤ −2, the minimum value ofa2 is (−2)2 = 4.

Accelerating RL for LLM Reasoning with Optimal Advantage Regression Since−6≤a≤ −2, the minimum value ofa2 is (−2)2 = 4

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:55:59.386030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T13:55:57.174578Z digest=sha256:0d36a05f0b9dc45fa3a3ffce40eaf5951a204ca593fb761c75c8fa892fb7665c

Observation 26adfdfb-51e0-40ad-97db-ff2df1d042ab · outbound

This paper cites an unresolved cited work.

Accelerating RL for LLM Reasoning with Optimal Advantage Regression Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:55:59.161666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T13:55:57.283784Z digest=sha256:b66e254e60ca57f629f4f9ed6e689ac2d0078a82b34c9d3fd31816fcebbf4911

Observation 42880f05-e596-424c-9dd5-7c7c1c519283 · outbound

This paper cites an unresolved cited work.

Accelerating RL for LLM Reasoning with Optimal Advantage Regression Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:55:58.888017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T13:55:57.381757Z digest=sha256:f7852fa37edf909ed9a377e8201b945350e1ef411fa4a8e7f5d29d5e4d64c986

Observation 8c2c594f-2fde-4bcc-9a59-d248968b7d93 · outbound

This paper cites From these calculations, we see that the greatest possible value is indeed achieved whena=−2andb= 3, giving us: −35 9.

Accelerating RL for LLM Reasoning with Optimal Advantage Regression From these calculations, we see that the greatest possible value is indeed achieved whena=−2andb= 3, giving us: −35 9

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:55:58.506243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T13:55:57.546525Z digest=sha256:1da0d344d54fc07c741fafa379c3ae640117672c0b7d60dcfe9ffb34e702e542

Observation f814d688-6ba3-4f94-b94b-90668a4e6144 · outbound

This paper cites a+c= 1−1 = 03.

Accelerating RL for LLM Reasoning with Optimal Advantage Regression a+c= 1−1 = 03

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:55:58.263134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T13:55:57.653874Z digest=sha256:d020050dbd367a7ce8a4e67538346ac07f349df78a00f442773f91b75aa327a0

Observation 5ea03061-2779-4809-8854-e3be1289dffc · outbound

This paper cites X y exp(⟨θT+1 , ϕ(x, y)⟩)P y′ exp(⟨θT+1 , ϕ(x, y′)⟩) − exp(⟨θ⋆, ϕ(x, y)⟩)P y′ exp(⟨θ⋆, ϕ(x, y′)⟩) # ≤Ex.

Accelerating RL for LLM Reasoning with Optimal Advantage Regression X y exp(⟨θT+1 , ϕ(x, y)⟩)P y′ exp(⟨θT+1 , ϕ(x, y′)⟩) − exp(⟨θ⋆, ϕ(x, y)⟩)P y′ exp(⟨θ⋆, ϕ(x, y′)⟩) # ≤Ex

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:55:58.018712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T13:55:57.728078Z digest=sha256:427932b8222ecea6cb8b76f6cbacbd4f10f6379427233a22bc6ae671534fd7b7

Observation a25873b9-4bbb-48a7-b9d8-600f23a38b99 · outbound

This paper cites Echo Chamber: RL Post-training Amplifies Behaviors Learned in Pretraining.

Accelerating RL for LLM Reasoning with Optimal Advantage Regression Echo Chamber: RL Post-training Amplifies Behaviors Learned in Pretraining

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T13:55:56.838631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:55:56.838631Z digest=sha256:21540aaf62827ad27d1b56042c0cb74caaa2880df473e28f0517ae5f65dbeef8

Observation 7031781b-ba7e-4080-88d5-327e559e5f6e · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Accelerating RL for LLM Reasoning with Optimal Advantage Regression DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-07T13:55:56.722789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:55:56.722789Z digest=sha256:ffa64589e4f47f8c4f1fdf86a1e5d8cd49cc7ba4439e1eff1fbae972ce422ede

Pith citing papers

Observation 6349f898-48a5-4df7-a4d1-6431d7e0e02f · inbound

On the optimization dynamics of RLVR: Gradient gap and step size thresholds cites this paper.

On the optimization dynamics of RLVR: Gradient gap and step size thresholds Accelerating RL for LLM Reasoning with Optimal Advantage Regression

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-18T08:36:07.288018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-18T08:34:36.543874Z digest=sha256:bde1728befae4984bf3b9464698535a23a7d8e79448eeba45478b2f4e439e47b

Observation b6ff4adb-ef38-43aa-be2c-f7875be47739 · inbound

Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards cites this paper.

Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards Accelerating RL for LLM Reasoning with Optimal Advantage Regression

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-03T19:38:19.340781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:38:19.340781Z digest=sha256:097fde846ddae7bb40f77a3d86a7315e09f1e1a87ce24d3b0233d43ad8a95b69

Observation 62febf29-27b9-4adc-8355-6fec57b6f8b8 · inbound

Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards cites this paper.

Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards Accelerating RL for LLM Reasoning with Optimal Advantage Regression

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:16.744697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:16.744697Z digest=sha256:576eebc851488f5d845ae352aad6240936f08c6e0a4c879cd8599b09eb5806ad