Pith. sign in

Paper Citation Record · LEDGER

S$^2$R: Teaching LLMs to Self-verify and Self-correct via Reinforcement Learning

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 12 inbound Pith citation observations for arXiv:2502.12853.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.12853 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 12 of 12 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 12 of 12 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:51:30.691200Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-21T06:54:01.056295Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 3c58cff3-de1e-4c63-b657-f717887b1988 · inbound

From System 1 to System 2: A Survey of Reasoning Large Language Models cites this paper.

From System 1 to System 2: A Survey of Reasoning Large Language Models S$^2$R: Teaching LLMs to Self-verify and Self-correct via Reinforcement Learning

Reference 200

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:36:24.296155Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T01:36:23.845366Z digest=sha256:731c49ee3b43d0993d73a9b6fae454362a6fcd2ee230fff19688470d7f587084

Observation 2d5e7b6f-4e1c-4faf-9f00-b351615363dd · inbound

Boosting LLM Reasoning via Spontaneous Self-Correction cites this paper.

Boosting LLM Reasoning via Spontaneous Self-Correction S$^2$R: Teaching LLMs to Self-verify and Self-correct via Reinforcement Learning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T05:51:30.691200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:51:30.691200Z digest=sha256:992318ce305b230d566c01eb8f834c904cd08018d238c2586f941351b7f99c66

Observation e812ee5e-535c-40df-9a12-ddc265680545 · inbound

PAG: Multi-Turn Reinforced LLM Self-Correction with Policy as Generative Verifier cites this paper.

PAG: Multi-Turn Reinforced LLM Self-Correction with Policy as Generative Verifier S$^2$R: Teaching LLMs to Self-verify and Self-correct via Reinforcement Learning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T04:39:34.398292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:39:34.398292Z digest=sha256:3f1c9056cca95a13c2134f686a2c6d664ffe2280c77fda76cc42f21c10b23726

Observation 31143d57-8403-4307-9567-3d65736b6bca · inbound

CogDual: Enhancing Dual Cognition of LLMs via Reinforcement Learning with Implicit Rule-Based Rewards cites this paper.

CogDual: Enhancing Dual Cognition of LLMs via Reinforcement Learning with Implicit Rule-Based Rewards S$^2$R: Teaching LLMs to Self-verify and Self-correct via Reinforcement Learning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T15:02:26.932292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:02:26.932292Z digest=sha256:78be6183f2a70fe0a7ba0e53e51a955498d7aebcbb25c2a14a609f1ebd1448ff

Observation 686afce3-f261-4af1-b3ef-9204c8fdb512 · inbound

Self-Reflective Generation at Test Time cites this paper.

Self-Reflective Generation at Test Time S$^2$R: Teaching LLMs to Self-verify and Self-correct via Reinforcement Learning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T12:41:42.687257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:41:42.687257Z digest=sha256:e7b6cf5694dbbf71aa1420d74a5b6526b7f3373ae2ed0506cc897ecea928e37d

Observation 391e202b-2abc-490a-af25-b77342bd120f · inbound

Towards Sparse Video Understanding and Reasoning cites this paper.

Towards Sparse Video Understanding and Reasoning S$^2$R: Teaching LLMs to Self-verify and Self-correct via Reinforcement Learning

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-02T23:33:12.494748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:33:12.494748Z digest=sha256:4e4b0b0bc1897abf608cbefe0b8b7450762d38b53755ec6d01293ad5a11fa278

Observation aa851a64-26fc-49e8-a29b-ea8e5102c6e2 · inbound

SVSR: A Self-Verification and Self-Rectification Paradigm for Multimodal Reasoning cites this paper.

SVSR: A Self-Verification and Self-Rectification Paradigm for Multimodal Reasoning S$^2$R: Teaching LLMs to Self-verify and Self-correct via Reinforcement Learning

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:31:03.849486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T15:59:04.802241Z digest=sha256:c63078410a18c666c87038b918487df21f9469fc16c305d2d3be9d900177d773

Observation f2f9c1d9-9e73-4d04-9ba7-3a2762563404 · inbound

SVSR: A Self-Verification and Self-Rectification Paradigm for Multimodal Reasoning cites this paper.

SVSR: A Self-Verification and Self-Rectification Paradigm for Multimodal Reasoning S$^2$R: Teaching LLMs to Self-verify and Self-correct via Reinforcement Learning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-07-12T22:47:33.917420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:47:33.917420Z digest=sha256:c289d0e9976d14efdc2d20a9a2212d3a438e594c9dd06cf8bb219ec293945659

Observation e91f14a7-9f89-43fd-94ad-3a06bde8597e · inbound

Multi-modal Reasoning with LLMs for Visual Semantic Arithmetic cites this paper.

Multi-modal Reasoning with LLMs for Visual Semantic Arithmetic S$^2$R: Teaching LLMs to Self-verify and Self-correct via Reinforcement Learning

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-11T12:56:10.337177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T02:35:17.215879Z digest=sha256:f4b09f177585200dab0928823782d8dcf06e44a7dc60513fef9eb9264ea7e6c5

Observation 90a1b760-73b2-47f8-8e9b-649ec3c0717e · inbound

Internalizing Outcome Supervision into Process Supervision: A New Paradigm for Reinforcement Learning for Reasoning cites this paper.

Internalizing Outcome Supervision into Process Supervision: A New Paradigm for Reinforcement Learning for Reasoning S$^2$R: Teaching LLMs to Self-verify and Self-correct via Reinforcement Learning

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-10T06:31:31.255048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T06:13:09.898530Z digest=sha256:4527ff837d1239695e4997367604c0f26c12613b76bd2e64809d31587a743ac3

Observation 68b6e2e8-4366-47fe-95da-e734b1f24516 · inbound

The Hidden Signal of Verifier Strictness: Controlling and Improving Step-Wise Verification via Selective Latent Steering cites this paper.

The Hidden Signal of Verifier Strictness: Controlling and Improving Step-Wise Verification via Selective Latent Steering S$^2$R: Teaching LLMs to Self-verify and Self-correct via Reinforcement Learning

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-21T06:54:01.058137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T06:51:56.556213Z digest=sha256:082e75c35cbe5eb551b2bc3d3f7a3384215180940e71ea347048298349fc9a89

Observation b5942e9d-b718-4427-8708-efc52af0e119 · inbound

Spectral Origins of the Self-Correction Blind Spot in Autoregressive Generation cites this paper.

Spectral Origins of the Self-Correction Blind Spot in Autoregressive Generation S$^2$R: Teaching LLMs to Self-verify and Self-correct via Reinforcement Learning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-07-14T15:30:23.485228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T15:30:23.485228Z digest=sha256:eb262fa564e00155f1ee5a34149a2fd6ca0923161a7a7bbb9190c05ed0d205c3