Pith. sign in

Paper Citation Record · LEDGER

Self-rewarding correction for mathematical reasoning

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 19 inbound Pith citation observations for arXiv:2502.19613.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.19613 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 19 of 19 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 19 of 19 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T00:41:48.404875Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 97f62b28-d998-499f-aa58-9eb0923c1a72 · inbound

Scaling Test-time Compute for LLM Agents cites this paper.

Scaling Test-time Compute for LLM Agents Self-rewarding correction for mathematical reasoning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T00:41:48.404875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:41:48.404875Z digest=sha256:64391ad8cc92bde80bb5ef78a4339456c78b999a10ef5f3875603e33066a68da

Observation cc1bd704-6917-4eda-a7cf-32a282e1a413 · inbound

Beyond Correctness: Harmonizing Process and Outcome Rewards through RL Training cites this paper.

Beyond Correctness: Harmonizing Process and Outcome Rewards through RL Training Self-rewarding correction for mathematical reasoning

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-21T22:40:43.171530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-21T22:38:57.833414Z digest=sha256:b67196c703982d0a77bc14cf020926abd1132c22da1ef6445a9d11fd71fb0b6f

Observation 06743f5c-b56f-4ab0-a5ae-fcede9ccce86 · inbound

LightReasoner: Can Small Language Models Teach Large Language Models Reasoning? cites this paper.

LightReasoner: Can Small Language Models Teach Large Language Models Reasoning? Self-rewarding correction for mathematical reasoning

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-22T13:01:34.019544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T12:59:31.011283Z digest=sha256:550d49e11e5e313624e7924a7bdda07ff5216ee2feee919668027aa2dd107ac8

Observation 5830c2c4-5a28-44d0-9fea-33295906121a · inbound

Breaking the Self-Confirming Loop: Diagnosing and Mitigating Systemic Reward Bias in Self-Rewarding RL cites this paper.

Breaking the Self-Confirming Loop: Diagnosing and Mitigating Systemic Reward Bias in Self-Rewarding RL Self-rewarding correction for mathematical reasoning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-04T10:44:31.460549Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:44:31.460549Z digest=sha256:447b4806ebbc50ec0f57382ab4b2113c268e5ffdc12c7101eabdeea733277652

Observation b896a909-8106-4d06-ac64-1436ef0624b9 · inbound

CPMobius: Iterative Coach-Player Reasoning for Data-Free Reinforcement Learning cites this paper.

CPMobius: Iterative Coach-Player Reasoning for Data-Free Reinforcement Learning Self-rewarding correction for mathematical reasoning

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-03T05:14:21.533431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T05:14:21.533431Z digest=sha256:70e1327458fd98072751e3f6b8437c033f36c27aaa64e9b044e1abb83e06b0ad

Observation 81f621f4-3604-428e-8cc2-e0bf2c72abc3 · inbound

rePIRL: Learn PRM with Inverse RL for LLM Reasoning cites this paper.

rePIRL: Learn PRM with Inverse RL for LLM Reasoning Self-rewarding correction for mathematical reasoning

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-21T13:14:10.976923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T13:13:13.293921Z digest=sha256:e41ab39e1d81411c365c37d9f29a4d93de2fd38168a75083a5bbd22ace9c0c61

Observation dbfc3b72-9643-43e4-85a6-9b411c13cfd7 · inbound

rePIRL: Learn PRM with Inverse RL for LLM Reasoning cites this paper.

rePIRL: Learn PRM with Inverse RL for LLM Reasoning Self-rewarding correction for mathematical reasoning

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-03T03:33:45.150810Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:33:45.150810Z digest=sha256:c6e6edbd7cfbf60526b7105b39d8c109880dba41a320ceab06da0be0d0442da9

Observation 45e07ccf-0df1-4722-b143-7e48f81aada3 · inbound

Can LLMs Learn to Reason Robustly under Noisy Supervision? cites this paper.

Can LLMs Learn to Reason Robustly under Noisy Supervision? Self-rewarding correction for mathematical reasoning

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-13T17:08:01.264453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T16:58:42.129870Z digest=sha256:ca703c043647a7a8a91cb143f0726238caefa2331d250a9104670d9d3dafad40

Observation 6c7cb4b3-0047-45bd-a6ed-f07c5cd35174 · inbound

Internalizing Outcome Supervision into Process Supervision: A New Paradigm for Reinforcement Learning for Reasoning cites this paper.

Internalizing Outcome Supervision into Process Supervision: A New Paradigm for Reinforcement Learning for Reasoning Self-rewarding correction for mathematical reasoning

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-10T06:26:27.829461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T06:13:09.898530Z digest=sha256:f1e84c4e27a478d0c88110868e0646e212fbffff5aef52a3f1bab5b92ce25df2

Observation cf61d227-58ea-4435-9385-6261c2420a7c · inbound

ACE: Self-Evolving LLM Coding Framework via Adversarial Unit Test Generation and Preference Optimization cites this paper.

ACE: Self-Evolving LLM Coding Framework via Adversarial Unit Test Generation and Preference Optimization Self-rewarding correction for mathematical reasoning

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-21T00:59:19.336113Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T00:55:40.784295Z digest=sha256:51f7413d3f2766e175bed1af75f8f132e66a38aae90ffd51ad99ae433f6cf69b

Observation f414539e-896d-4e51-bb05-6e65dc61007e · inbound

ACE: Self-Evolving LLM Coding Framework via Adversarial Unit Test Generation and Preference Optimization cites this paper.

ACE: Self-Evolving LLM Coding Framework via Adversarial Unit Test Generation and Preference Optimization Self-rewarding correction for mathematical reasoning

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-22T10:14:47.165809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T10:14:00.478000Z digest=sha256:a1557b5601a0da5f60596a59f64aa209e8569ae159b2a349e795c772dda2d16f

Observation 0dd7aacd-e04f-4881-a417-eea4ab73c190 · inbound

Guarded Repair for Harm-Aware Post-hoc Replacement of LLM Mathematical Reasoning cites this paper.

Guarded Repair for Harm-Aware Post-hoc Replacement of LLM Mathematical Reasoning Self-rewarding correction for mathematical reasoning

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-06-30T13:44:40.809593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-06-30T13:38:48.617204Z digest=sha256:f889526ba2a119e18b80df14bc995548ea903af06ec6e831920224ff7d6923a7

Observation 03fe794f-2c38-44d0-961d-2ae2f83409a2 · inbound

Trust Region On-Policy Distillation cites this paper.

Trust Region On-Policy Distillation Self-rewarding correction for mathematical reasoning

Reference 162

Resolution
verified exact
arxiv_id, observed 2026-07-01T20:56:13.435749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-06-28T17:38:50.313305Z digest=sha256:5ef91c568ff9c8652aa179c39db7bbe54235c24f90a5404514d69776d3b9fd5c

Observation c9f7a4fa-0632-4437-b449-b34408d22a14 · inbound

The Periodic Table of LLM Reasoning: A Structured Survey of Reasoning Paradigms, Methods, and Failure Modes cites this paper.

The Periodic Table of LLM Reasoning: A Structured Survey of Reasoning Paradigms, Methods, and Failure Modes Self-rewarding correction for mathematical reasoning

Reference 274

Resolution
verified exact
arxiv_id, observed 2026-06-27T13:00:55.970436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-06-27T12:59:51.091008Z digest=sha256:ad8d7893eb836218b61c72c882de86f1d75bebe890dadcacaa6814975b3c7057

Observation 86fff995-f14c-4448-b088-6cff9d7586e5 · inbound

ReSum: Synergizing LLM Reasoning and Summarization with Reinforcement Learning cites this paper.

ReSum: Synergizing LLM Reasoning and Summarization with Reinforcement Learning Self-rewarding correction for mathematical reasoning

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-07-03T15:08:33.388470Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T06:39:34.199607Z digest=sha256:f1e41a920b1bfe089d0d86ce4b4940f02f53ef18c397c4d050f3ed93ba191995

Observation deeadaf6-7404-490b-a02c-e0291d1527f5 · inbound

ReSum: Synergizing LLM Reasoning and Summarization with Reinforcement Learning cites this paper.

ReSum: Synergizing LLM Reasoning and Summarization with Reinforcement Learning Self-rewarding correction for mathematical reasoning

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-03T02:12:24.010055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:12:24.010055Z digest=sha256:2ddaa357144636f8324b062d1805a1b98f2f15693f01d795bf759c245aeda720

Observation 2cf7f342-4619-40d3-8df7-e5cc461b49d7 · inbound

Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning cites this paper.

Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning Self-rewarding correction for mathematical reasoning

Reference 238

Resolution
verified exact
arxiv_id, observed 2026-07-04T07:59:40.132829Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-26T12:15:08.304150Z digest=sha256:7acafa6bd8da1358c1b5928c1d83ad5ca4603cbda201f8f76c21ef2e5582682a

Observation af7e3537-efcf-4951-ba7a-d7583bc4b8f7 · inbound

Cognitive Episodes in LLM Reasoning Traces Enable Interpretable Human Item Difficulty Prediction cites this paper.

Cognitive Episodes in LLM Reasoning Traces Enable Interpretable Human Item Difficulty Prediction Self-rewarding correction for mathematical reasoning

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-07-01T17:15:51.588940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-29T03:58:32.372896Z digest=sha256:07f40471e575b3edb2a14bc5b79b69f16ce4257c1e67eff513f9aa7f61195050

Observation 8428ec32-f496-4353-9265-0eb5a8b61781 · inbound

Cognitive Episodes in LLM Reasoning Traces Enable Interpretable Human Item Difficulty Prediction cites this paper.

Cognitive Episodes in LLM Reasoning Traces Enable Interpretable Human Item Difficulty Prediction Self-rewarding correction for mathematical reasoning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-07-14T17:12:24.565155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T17:12:24.565155Z digest=sha256:f46de7812adbec5ec4987bfb4c24e7c6e3e26eb9230c90a6b74fff06d292b28f