Pith. sign in

Paper Citation Record · LEDGER

Accelerating RL for LLM Reasoning with Optimal Advantage Regression

As of 11 August 2026, this Paper Citation Record lists 13 of 13 outbound references and 3 inbound Pith citation observations for arXiv:2505.20686.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.20686 v1

Coverage vector

measured 13 of 13 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:55:57.728078Z

measured 16 of 16 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T06:47:16.744697Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-18T08:36:07.284835Z

Reference resolution

13 of 13 outbound references displayed

  • verified exact0
  • verified fuzzy7
  • unresolved6
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e81a2ab3-e7be-4a5c-b437-cf9bcd404833 · outbound

This paper cites an unresolved cited work.

Accelerating RL for LLM Reasoning with Optimal Advantage Regression Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:56:00.295544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:55:56.898779Z digest=sha256:94c050bd745915bdabb81a941eeb7d35f2dd3b6d3d7e45ec9f4c3253cdc50fcf

Observation 66aec46c-8ef6-45f2-8d3c-ebff8e96057e · outbound

This paper cites So,a2 = (−2)2 = 4.

Accelerating RL for LLM Reasoning with Optimal Advantage Regression So,a2 = (−2)2 = 4

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:55:59.902129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:55:57.019407Z digest=sha256:c521a275b5a55736a5daa535cf89df32fa403683fcff2396122fff511b38aa25

Observation d84a8ba9-54ec-48b2-9408-3179fe2abd2e · outbound

This paper cites an unresolved cited work.

Accelerating RL for LLM Reasoning with Optimal Advantage Regression Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:55:58.688420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:55:57.481389Z digest=sha256:7cc244084302bfa16e04e9b0f1f619558c9218843aac359328ed6d87ef6cf492

Observation 325b6916-991d-49d0-b391-836cad9fdfb5 · outbound

This paper cites So,1 b2 = 1 32 = 1 9.

Accelerating RL for LLM Reasoning with Optimal Advantage Regression So,1 b2 = 1 32 = 1 9

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:56:00.109577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:55:56.962513Z digest=sha256:7eb29b7eec272209f9062ea9f6709a9db9c5d7714dee7cfbe14d0d4ce7b45ca2

Observation 5dc3bda0-7723-4f97-9ff4-08425c39f29a · outbound

This paper cites Since3≤b≤5 , the minimum value ofb is 3.

Accelerating RL for LLM Reasoning with Optimal Advantage Regression Since3≤b≤5 , the minimum value ofb is 3

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:55:59.679835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:55:57.100364Z digest=sha256:0674f056410019b43590700e3765290d258c910ad4fde621ecb84b682410250c

Observation 4af27d99-0b7b-4316-86b3-b1c4283560c3 · outbound

This paper cites Since−6≤a≤ −2, the minimum value ofa2 is (−2)2 = 4.

Accelerating RL for LLM Reasoning with Optimal Advantage Regression Since−6≤a≤ −2, the minimum value ofa2 is (−2)2 = 4

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:55:59.386030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:55:57.174578Z digest=sha256:66f354e5be58c62c562829f1f5328bb2348f4f762e6ad49940d23b24944f4e9d

Observation 26adfdfb-51e0-40ad-97db-ff2df1d042ab · outbound

This paper cites an unresolved cited work.

Accelerating RL for LLM Reasoning with Optimal Advantage Regression Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:55:59.161666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:55:57.283784Z digest=sha256:2706ea11b696732a0705a0ca11543563e80e5a60d5a7c34d4fb014f331667796

Observation 42880f05-e596-424c-9dd5-7c7c1c519283 · outbound

This paper cites an unresolved cited work.

Accelerating RL for LLM Reasoning with Optimal Advantage Regression Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:55:58.888017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:55:57.381757Z digest=sha256:d98a55234b421ce2cf1ce59de1b066bd68b9b9db24f32d982525f6a5b3608935

Observation 8c2c594f-2fde-4bcc-9a59-d248968b7d93 · outbound

This paper cites From these calculations, we see that the greatest possible value is indeed achieved whena=−2andb= 3, giving us: −35 9.

Accelerating RL for LLM Reasoning with Optimal Advantage Regression From these calculations, we see that the greatest possible value is indeed achieved whena=−2andb= 3, giving us: −35 9

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:55:58.506243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:55:57.546525Z digest=sha256:114483f80f7d7d4ec7a4659e9b4fe8e86cf4fc49ae7b73070847f88af184f179

Observation f814d688-6ba3-4f94-b94b-90668a4e6144 · outbound

This paper cites a+c= 1−1 = 03.

Accelerating RL for LLM Reasoning with Optimal Advantage Regression a+c= 1−1 = 03

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:55:58.263134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:55:57.653874Z digest=sha256:008ef911091a4de6f27a825fca44ff5b338587ea6811f8559e2d07ac76fa1881

Observation 5ea03061-2779-4809-8854-e3be1289dffc · outbound

This paper cites X y exp(⟨θT+1 , ϕ(x, y)⟩)P y′ exp(⟨θT+1 , ϕ(x, y′)⟩) − exp(⟨θ⋆, ϕ(x, y)⟩)P y′ exp(⟨θ⋆, ϕ(x, y′)⟩) # ≤Ex.

Accelerating RL for LLM Reasoning with Optimal Advantage Regression X y exp(⟨θT+1 , ϕ(x, y)⟩)P y′ exp(⟨θT+1 , ϕ(x, y′)⟩) − exp(⟨θ⋆, ϕ(x, y)⟩)P y′ exp(⟨θ⋆, ϕ(x, y′)⟩) # ≤Ex

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:55:58.018712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:55:57.728078Z digest=sha256:090d8e6738fd1a923e0cbfc39b14fba0125652d99ef4854446de33098a280cd7

Observation a25873b9-4bbb-48a7-b9d8-600f23a38b99 · outbound

This paper cites Echo Chamber: RL Post-training Amplifies Behaviors Learned in Pretraining.

Accelerating RL for LLM Reasoning with Optimal Advantage Regression Echo Chamber: RL Post-training Amplifies Behaviors Learned in Pretraining

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T13:55:56.838631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:55:56.838631Z digest=sha256:21540aaf62827ad27d1b56042c0cb74caaa2880df473e28f0517ae5f65dbeef8

Observation 7031781b-ba7e-4080-88d5-327e559e5f6e · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Accelerating RL for LLM Reasoning with Optimal Advantage Regression DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-07T13:55:56.722789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:55:56.722789Z digest=sha256:ffa64589e4f47f8c4f1fdf86a1e5d8cd49cc7ba4439e1eff1fbae972ce422ede

Pith citing papers

Observation 6349f898-48a5-4df7-a4d1-6431d7e0e02f · inbound

On the optimization dynamics of RLVR: Gradient gap and step size thresholds cites this paper.

On the optimization dynamics of RLVR: Gradient gap and step size thresholds Accelerating RL for LLM Reasoning with Optimal Advantage Regression

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-18T08:36:07.288018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-18T08:34:36.543874Z digest=sha256:bda60022d251ec1ed5e343a9fe250f48b8592bf629db3d7b85898364604a2b65

Observation b6ff4adb-ef38-43aa-be2c-f7875be47739 · inbound

Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards cites this paper.

Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards Accelerating RL for LLM Reasoning with Optimal Advantage Regression

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-03T19:38:19.340781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:38:19.340781Z digest=sha256:097fde846ddae7bb40f77a3d86a7315e09f1e1a87ce24d3b0233d43ad8a95b69

Observation 62febf29-27b9-4adc-8355-6fec57b6f8b8 · inbound

Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards cites this paper.

Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards Accelerating RL for LLM Reasoning with Optimal Advantage Regression

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:16.744697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:16.744697Z digest=sha256:576eebc851488f5d845ae352aad6240936f08c6e0a4c879cd8599b09eb5806ad