Pith. sign in

Paper Citation Record · LEDGER

S$^2$R: Teaching LLMs to Self-verify and Self-correct via Reinforcement Learning

As of 20 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 19 inbound Pith citation observations for arXiv:2502.12853.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.12853 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 19 of 19 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 19 of 19 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T12:02:43.487638Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-21T06:54:01.056295Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 3c58cff3-de1e-4c63-b657-f717887b1988 · inbound

From System 1 to System 2: A Survey of Reasoning Large Language Models cites this paper.

From System 1 to System 2: A Survey of Reasoning Large Language Models S$^2$R: Teaching LLMs to Self-verify and Self-correct via Reinforcement Learning

Reference 200

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:36:24.296155Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-13T01:36:23.845366Z digest=sha256:d012ffc8e82d94a877911ae3c8b7db2d874df704ec824c18308f0ef9f7e61b84

Observation 7fa75887-ab90-4ed2-9802-75f90428f9d3 · inbound

Generative AI Act II: Test Time Scaling Drives Cognition Engineering cites this paper.

Generative AI Act II: Test Time Scaling Drives Cognition Engineering S$^2$R: Teaching LLMs to Self-verify and Self-correct via Reinforcement Learning

Reference 222

Resolution
unresolved
no resolver link, observed 2026-08-16T12:02:43.487638Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:02:43.487638Z digest=sha256:4cea949c8bdce937894f8075e2a8bea10b5f967fa21fa65d0694457ed63bda90

Observation 27a71eaf-1c26-4cbc-b64f-d6301da3b8f1 · inbound

SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning cites this paper.

SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning S$^2$R: Teaching LLMs to Self-verify and Self-correct via Reinforcement Learning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-16T06:04:14.624752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T06:04:14.624752Z digest=sha256:350b91748af34cf00bac3ede19ccd66ad46bc390e2a42297da80be014fc41cda

Observation 21f50789-98d9-4c5f-ba51-f3255275429c · inbound

100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models cites this paper.

100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models S$^2$R: Teaching LLMs to Self-verify and Self-correct via Reinforcement Learning

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-16T04:43:44.306712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:43:44.306712Z digest=sha256:69ead4018358e69f482d4fe946179863639e781a0879ba870cc47f375be0cd46

Observation 93e86eb5-d40a-4b50-8c8c-b65249a2b758 · inbound

Training-Free Reasoning and Reflection in MLLMs cites this paper.

Training-Free Reasoning and Reflection in MLLMs S$^2$R: Teaching LLMs to Self-verify and Self-correct via Reinforcement Learning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T15:13:51.984499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:13:51.984499Z digest=sha256:9037cb1b6b9c4803b72999cc3a7582bab6b8c31fc588dd3c05993dfcbe0ca60c

Observation 3c54ff20-26cc-4bae-93be-5128e3dd5c00 · inbound

LeTS: Learning to Think-and-Search via Process-and-Outcome Reward Hybridization cites this paper.

LeTS: Learning to Think-and-Search via Process-and-Outcome Reward Hybridization S$^2$R: Teaching LLMs to Self-verify and Self-correct via Reinforcement Learning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T14:51:05.696234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:51:05.696234Z digest=sha256:61f78a2cbbdcd8aedf46b2f2478c1ef188d0d1551dc8542d301fadab7057c2ed

Observation 2d5e7b6f-4e1c-4faf-9f00-b351615363dd · inbound

Boosting LLM Reasoning via Spontaneous Self-Correction cites this paper.

Boosting LLM Reasoning via Spontaneous Self-Correction S$^2$R: Teaching LLMs to Self-verify and Self-correct via Reinforcement Learning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T05:51:30.691200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:51:30.691200Z digest=sha256:9094ccb2fa5e0b26c18e305c8518dbc9dbb54afc102839d24a28e65e97097d07

Observation e812ee5e-535c-40df-9a12-ddc265680545 · inbound

PAG: Multi-Turn Reinforced LLM Self-Correction with Policy as Generative Verifier cites this paper.

PAG: Multi-Turn Reinforced LLM Self-Correction with Policy as Generative Verifier S$^2$R: Teaching LLMs to Self-verify and Self-correct via Reinforcement Learning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T04:39:34.398292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:39:34.398292Z digest=sha256:b1538148c60b152ff69a0e77fd1f94d1a01f915eb330f92f2e9f5d9887655992

Observation 31143d57-8403-4307-9567-3d65736b6bca · inbound

CogDual: Enhancing Dual Cognition of LLMs via Reinforcement Learning with Implicit Rule-Based Rewards cites this paper.

CogDual: Enhancing Dual Cognition of LLMs via Reinforcement Learning with Implicit Rule-Based Rewards S$^2$R: Teaching LLMs to Self-verify and Self-correct via Reinforcement Learning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T15:02:26.932292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:02:26.932292Z digest=sha256:1619405d9467870c2de649b1223e3244dc4d5e396e41e8954015bae7a9d5d22d

Observation 9cb7f441-157c-410b-abca-fe9d2c731113 · inbound

Implicit Actor Critic Coupling via a Supervised Learning Framework for RLVR cites this paper.

Implicit Actor Critic Coupling via a Supervised Learning Framework for RLVR S$^2$R: Teaching LLMs to Self-verify and Self-correct via Reinforcement Learning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T16:45:52.837578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:45:52.837578Z digest=sha256:b5921207fedaaba8752b2153def7f7459d5a97d93f16da9d139d43c1ce2118f9

Observation 686afce3-f261-4af1-b3ef-9204c8fdb512 · inbound

Self-Reflective Generation at Test Time cites this paper.

Self-Reflective Generation at Test Time S$^2$R: Teaching LLMs to Self-verify and Self-correct via Reinforcement Learning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T12:41:42.687257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:41:42.687257Z digest=sha256:8184934859bfbf7f5069b44a6d21dad206a0511a0777b4b7884335099369c8fe

Observation 391e202b-2abc-490a-af25-b77342bd120f · inbound

Towards Sparse Video Understanding and Reasoning cites this paper.

Towards Sparse Video Understanding and Reasoning S$^2$R: Teaching LLMs to Self-verify and Self-correct via Reinforcement Learning

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-02T23:33:12.494748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:33:12.494748Z digest=sha256:bf7f5eead22ff7cabff53e9ec8dd2f1401d9a6a1c58698cfd1f27aeecd365f7a

Observation aa851a64-26fc-49e8-a29b-ea8e5102c6e2 · inbound

SVSR: A Self-Verification and Self-Rectification Paradigm for Multimodal Reasoning cites this paper.

SVSR: A Self-Verification and Self-Rectification Paradigm for Multimodal Reasoning S$^2$R: Teaching LLMs to Self-verify and Self-correct via Reinforcement Learning

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:31:03.849486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-10T15:59:04.802241Z digest=sha256:914811fdad0ef6e390591ce721630dfb6bdea76a4504a2d93bd4d41dc740714e

Observation f2f9c1d9-9e73-4d04-9ba7-3a2762563404 · inbound

SVSR: A Self-Verification and Self-Rectification Paradigm for Multimodal Reasoning cites this paper.

SVSR: A Self-Verification and Self-Rectification Paradigm for Multimodal Reasoning S$^2$R: Teaching LLMs to Self-verify and Self-correct via Reinforcement Learning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-07-12T22:47:33.917420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:47:33.917420Z digest=sha256:968a3b03685b6bbc3ebdc6573ade84aa67905e242fa9671918690ffd9c903b3f

Observation e91f14a7-9f89-43fd-94ad-3a06bde8597e · inbound

Multi-modal Reasoning with LLMs for Visual Semantic Arithmetic cites this paper.

Multi-modal Reasoning with LLMs for Visual Semantic Arithmetic S$^2$R: Teaching LLMs to Self-verify and Self-correct via Reinforcement Learning

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-11T12:56:10.337177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-10T02:35:17.215879Z digest=sha256:4b7228383d971e02bd46a85cc6201e2505d0a0ff2881a2c2612372c947e32e23

Observation 90a1b760-73b2-47f8-8e9b-649ec3c0717e · inbound

Internalizing Outcome Supervision into Process Supervision: A New Paradigm for Reinforcement Learning for Reasoning cites this paper.

Internalizing Outcome Supervision into Process Supervision: A New Paradigm for Reinforcement Learning for Reasoning S$^2$R: Teaching LLMs to Self-verify and Self-correct via Reinforcement Learning

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-10T06:31:31.255048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-10T06:13:09.898530Z digest=sha256:8ddabd3b554ea07889faa26de8e7d6d06b85f1ee125474b581d32216ab418cf3

Observation 68b6e2e8-4366-47fe-95da-e734b1f24516 · inbound

The Hidden Signal of Verifier Strictness: Controlling and Improving Step-Wise Verification via Selective Latent Steering cites this paper.

The Hidden Signal of Verifier Strictness: Controlling and Improving Step-Wise Verification via Selective Latent Steering S$^2$R: Teaching LLMs to Self-verify and Self-correct via Reinforcement Learning

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-21T06:54:01.058137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-21T06:51:56.556213Z digest=sha256:c362037f71bd0a970b3001cbf7a85dd6a7ef4d2303e0a8fa5ba908ea6cce45a8

Observation b5942e9d-b718-4427-8708-efc52af0e119 · inbound

Spectral Origins of the Self-Correction Blind Spot in Autoregressive Generation cites this paper.

Spectral Origins of the Self-Correction Blind Spot in Autoregressive Generation S$^2$R: Teaching LLMs to Self-verify and Self-correct via Reinforcement Learning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-07-14T15:30:23.485228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T15:30:23.485228Z digest=sha256:f859a80d4310dbfe97152210e418b7de76989e15aaeea04bb4dcb788bc495650

Observation 687c8283-b00e-4b9b-bef2-1f2890cdce7f · inbound

Reinforcing Step-level Reasoning for Effective Self-Correction in LLMs cites this paper.

Reinforcing Step-level Reasoning for Effective Self-Correction in LLMs S$^2$R: Teaching LLMs to Self-verify and Self-correct via Reinforcement Learning

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-16T00:39:40.272768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:39:40.272768Z digest=sha256:bffcf6b939087f8a2271c926b98abaefe0ae0dcbd6e29a83fc4713a35c83c248