Pith. sign in

Paper Citation Record · LEDGER

Implicit Actor Critic Coupling via a Supervised Learning Framework for RLVR

As of 16 August 2026, this Paper Citation Record lists 25 of 25 outbound references and 0 inbound Pith citation observations for arXiv:2509.02522.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.02522 v3

Coverage vector

measured 25 of 25 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T16:45:52.887305Z

measured 25 of 25 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

25 of 25 outbound references displayed

  • verified exact0
  • verified fuzzy3
  • unresolved21
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9a760af3-2a8c-4e3d-859e-b9fa0da584e9 · outbound

This paper cites Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs.

Implicit Actor Critic Coupling via a Supervised Learning Framework for RLVR Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T16:45:52.790406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:45:52.790406Z digest=sha256:f58b16141bdef3889e80932a0181c31598e88f5da518433a753299fe4d0502a8

Observation 4e41b69c-3b9d-4edd-bcb6-530852def8de · outbound

This paper cites Process Reinforcement through Implicit Rewards.

Implicit Actor Critic Coupling via a Supervised Learning Framework for RLVR Process Reinforcement through Implicit Rewards

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T16:45:52.804111Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:45:52.804111Z digest=sha256:973e298c8069281c7ef64b93d85a46435bf2edc2d68b4f7638f810553d2b6fd0

Observation 934bb97d-16b7-4811-afd0-f5fbb08a86e6 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Implicit Actor Critic Coupling via a Supervised Learning Framework for RLVR DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T16:45:52.808234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:45:52.808234Z digest=sha256:a0f0d4dd20a67b31de886857f00148a1ecfb9afa78136c084ca0a1dc33311061

Observation 768f506d-3948-4e7f-bb26-68c93c54b937 · outbound

This paper cites REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization.

Implicit Actor Critic Coupling via a Supervised Learning Framework for RLVR REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T16:45:52.812942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:45:52.812942Z digest=sha256:afb29b48b5dffd333833fa222d1aee65b607908a0857d347e27c80be843c0df3

Observation 6af9588c-b213-4a22-a3e6-78da12653897 · outbound

This paper cites OpenAI o1 System Card.

Implicit Actor Critic Coupling via a Supervised Learning Framework for RLVR OpenAI o1 System Card

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T16:45:52.817433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:45:52.817433Z digest=sha256:c6861c6ee0c9cc509f94e4651328d03742c039d9b6ec39fedafe16ba0e64a58c

Observation 7031dc5b-c461-444e-a4c3-f9889dbcefcb · outbound

This paper cites ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models.

Implicit Actor Critic Coupling via a Supervised Learning Framework for RLVR ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T16:45:52.833521Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:45:52.833521Z digest=sha256:d962d8c714c25481b8975357d16584788df43f1bc2c71971e83041c8f6dfb7d7

Observation 9cb7f441-157c-410b-abca-fe9d2c731113 · outbound

This paper cites S$^2$R: Teaching LLMs to Self-verify and Self-correct via Reinforcement Learning.

Implicit Actor Critic Coupling via a Supervised Learning Framework for RLVR S$^2$R: Teaching LLMs to Self-verify and Self-correct via Reinforcement Learning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T16:45:52.837578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:45:52.837578Z digest=sha256:c12382a066deaac5ea5ccd7aa10abd56ed9647c2af3ca1891e5068e50b237896

Observation 6f514de6-b79c-4917-a74c-c10e0cf8eec0 · outbound

This paper cites Qwen2.5 Technical Report.

Implicit Actor Critic Coupling via a Supervised Learning Framework for RLVR Qwen2.5 Technical Report

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T16:45:52.841825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:45:52.841825Z digest=sha256:87679abac02a2214646d0a44bed8afde8eb21a421175ec5aa9909744c53b903e

Observation 1265fb87-c9d8-41ed-9062-e64f80621dd1 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Implicit Actor Critic Coupling via a Supervised Learning Framework for RLVR Proximal Policy Optimization Algorithms

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T16:45:52.845717Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:45:52.845717Z digest=sha256:6406cf0a2c4d1f713c269b6acff54024e45b7c03120fc88f2c9b2752213e37cc

Observation 2e52fec5-6cef-4b89-bf40-61ad9bc8284f · outbound

This paper cites HybridFlow: A Flexible and Efficient RLHF Framework.

Implicit Actor Critic Coupling via a Supervised Learning Framework for RLVR HybridFlow: A Flexible and Efficient RLHF Framework

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T16:45:52.853063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:45:52.853063Z digest=sha256:521ca1487821bec0e47b419eca4884b5eb8d82cf3e5b2abebcdceb0aa8792c8b

Observation d8f95830-1523-4e2f-b223-1e20f2851fe1 · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

Implicit Actor Critic Coupling via a Supervised Learning Framework for RLVR Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T16:45:52.856659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:45:52.856659Z digest=sha256:b45b8334f2f3f55c2c2e1292db8389a417775ad967fb20062e097df838ae370f

Observation 4e04719c-3cbf-4f4d-8cd9-ca2ab3fdb627 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

Implicit Actor Critic Coupling via a Supervised Learning Framework for RLVR DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T16:45:52.860369Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:45:52.860369Z digest=sha256:4103363bf8dc893f2cd6b40673c25ce8cd08dd666fc2c8a5c614a5dd14b23130

Observation 1fc4e330-70d5-43b0-801c-1c79e25d8c84 · outbound

This paper cites VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks.

Implicit Actor Critic Coupling via a Supervised Learning Framework for RLVR VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T16:45:52.864217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:45:52.864217Z digest=sha256:cec793f8d571c33fe4096304c11b1405df804d5dff35bb6cbaf28dcfc9edbea3

Observation 93ce47fd-f543-4faa-a1b0-0f6923e9e6ca · outbound

This paper cites Learning to Reason without External Rewards.

Implicit Actor Critic Coupling via a Supervised Learning Framework for RLVR Learning to Reason without External Rewards

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T16:45:52.868179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:45:52.868179Z digest=sha256:1b7a1b3f3d723fc5fd2ee11d845c2dc96fdf50a985862ae847bff3d1e8543691

Observation 93124b95-4a0d-43c4-bd28-d6a928a984d3 · outbound

This paper cites TTRL: Test-Time Reinforcement Learning.

Implicit Actor Critic Coupling via a Supervised Learning Framework for RLVR TTRL: Test-Time Reinforcement Learning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T16:45:52.872126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:45:52.872126Z digest=sha256:8393ce75fd171fa0e9fe78399388adb836b92e3ac0fac65fca16daf6cef826ad

Observation cdbee296-90bd-44a3-b8d5-efadb5dd55dc · outbound

This paper cites enumerate.

Implicit Actor Critic Coupling via a Supervised Learning Framework for RLVR enumerate

Reference 22

Resolution
malformed identifier
raw_fallback, observed 2026-08-15T16:45:53.178524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T16:45:52.875823Z digest=sha256:7cd3994ffcfe00908f11fb420ae7e6588e11af588a2d580d7910bc76450b7021

Observation b8f6a080-9342-4718-a2d2-efeb014351ca · outbound

This paper cites Underlined numbers indicate the second best.

Implicit Actor Critic Coupling via a Supervised Learning Framework for RLVR Underlined numbers indicate the second best

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:45:53.153625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T16:45:52.883696Z digest=sha256:ac2f91d3de7e42d1428c0d2c62b7c28abbe2d9600eee4c4385ecde2c79138968

Observation 74585305-0a13-4ba2-b65c-9952860ec6b8 · outbound

This paper cites Underlined numbers indicate the second best.

Implicit Actor Critic Coupling via a Supervised Learning Framework for RLVR Underlined numbers indicate the second best

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:45:53.141723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T16:45:52.887305Z digest=sha256:dceaf9d089d3fdd5a25da68fc2a0034f03f29b90b0db684367cd75f466ffdff5

Observation 5f5adc9b-572e-43e6-b146-a9ba472389d6 · outbound

This paper cites Underlined numbers indicate the second best.

Implicit Actor Critic Coupling via a Supervised Learning Framework for RLVR Underlined numbers indicate the second best

Reference 500

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:45:53.166668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T16:45:52.880121Z digest=sha256:2512daf66fe3e7b9a2156ac1f7b18f1e21e65216fbf19de24aa623f994254fdb

Observation f599a921-8fb9-4d05-b1db-86aea98e227b · outbound

This paper cites Wouter Kool, Herke van Hoof, and Max Welling.

Implicit Actor Critic Coupling via a Supervised Learning Framework for RLVR Wouter Kool, Herke van Hoof, and Max Welling

Reference 2001

Resolution
unresolved
no resolver link, observed 2026-08-15T16:45:52.825621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:45:52.825621Z digest=sha256:29abac4683aaea48b3f69f2fb6a59dae7ba9f2ef2a44555ca0d8a02d144d1a2c

Observation 4df44f2e-6bf6-4d89-90b8-86d56281d834 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Implicit Actor Critic Coupling via a Supervised Learning Framework for RLVR DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-15T16:45:52.849575Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:45:52.849575Z digest=sha256:be32354d91b49394202d07b1381f59456b7e7ff14649ab0fb66f6e9a8f4fe800

Observation c0491cae-4d10-4eea-988a-cd8bb5fc6090 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

Implicit Actor Critic Coupling via a Supervised Learning Framework for RLVR Evaluating Large Language Models Trained on Code

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-15T16:45:52.795338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:45:52.795338Z digest=sha256:d1d547da1f91c57120e4ef6abf89ca5cd9b96fa4e5ed22cf2e169419c430f9c5

Observation f4e868fc-b52f-4577-9da7-f280c4629213 · outbound

This paper cites Tulu 3: Pushing Frontiers in Open Language Model Post-Training.

Implicit Actor Critic Coupling via a Supervised Learning Framework for RLVR Tulu 3: Pushing Frontiers in Open Language Model Post-Training

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-15T16:45:52.829268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:45:52.829268Z digest=sha256:1b61f341a27c8f35e116832f1d58e5958bbf3bd3e760b8cdb8392ea835100433

Observation b07b515e-c171-45cd-8835-0d1c08964fb0 · outbound

This paper cites VinePPO: Refining Credit Assignment in RL Training of LLMs.

Implicit Actor Critic Coupling via a Supervised Learning Framework for RLVR VinePPO: Refining Credit Assignment in RL Training of LLMs

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-15T16:45:52.821880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:45:52.821880Z digest=sha256:f5dbaf4f65f46c488453d47116b124b55f7378067f44297cca56b74677146013

Observation bae7135e-7c4d-4b15-86c9-6aab06d8faeb · outbound

This paper cites Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities.

Implicit Actor Critic Coupling via a Supervised Learning Framework for RLVR Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-15T16:45:52.799417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:45:52.799417Z digest=sha256:8adc20aad0f2c2f95287d21a7a5a3e676b71d346a1e1ea2cae60f37630da8092

Pith citing papers

No inbound Pith citation observations are available.