Pith. sign in

Paper Citation Record · LEDGER

When Can LLMs Actually Correct Their Own Mistakes? A Critical Survey of Self-Correction of LLMs

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 12 inbound Pith citation observations for arXiv:2406.01297.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2406.01297 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 12 of 12 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 12 of 12 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:16:03.807732Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T03:59:33.348485Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 49214aca-bbe1-43d1-881b-4a95d2a9ccfc · inbound

Training Language Models to Self-Correct via Reinforcement Learning cites this paper.

Training Language Models to Self-Correct via Reinforcement Learning When Can LLMs Actually Correct Their Own Mistakes? A Critical Survey of Self-Correction of LLMs

Reference 61

Resolution
verified exact
arxiv_id, observed 2026-05-17T12:04:10.509116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-17T12:04:10.210508Z digest=sha256:065c0d8458259a0f2cd1facb1aa285b681e931b860bda0119bd46ad4435dcb22

Observation a58320bf-2d93-4227-abbf-f42fedf6c163 · inbound

Can LLMs $\textit{understand}$ Math? -- Exploring the Pitfalls in Mathematical Reasoning cites this paper.

Can LLMs $\textit{understand}$ Math? -- Exploring the Pitfalls in Mathematical Reasoning When Can LLMs Actually Correct Their Own Mistakes? A Critical Survey of Self-Correction of LLMs

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T15:16:03.807732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:16:03.807732Z digest=sha256:83a6d31d45a9e627e98e8147b5b57d2d86497de2da195cb373ccaf8488f2e086

Observation 4ccd1d5c-745c-4c72-a8c4-a1c8588a3896 · inbound

An Empirical Study of LLM-as-a-Judge: How Design Choices Impact Evaluation Reliability cites this paper.

An Empirical Study of LLM-as-a-Judge: How Design Choices Impact Evaluation Reliability When Can LLMs Actually Correct Their Own Mistakes? A Critical Survey of Self-Correction of LLMs

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T00:31:35.477675Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:31:35.477675Z digest=sha256:b92d3b3cad05eb06148c0934e36ad2bfe5d623c96f8fa7bed8c10db465b20568

Observation 2a698216-253f-4934-b8a6-d711f574b54f · inbound

Reinforcement Fine-Tuning Naturally Mitigates Forgetting in Continual Post-Training cites this paper.

Reinforcement Fine-Tuning Naturally Mitigates Forgetting in Continual Post-Training When Can LLMs Actually Correct Their Own Mistakes? A Critical Survey of Self-Correction of LLMs

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T19:31:27.394988Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:31:27.394988Z digest=sha256:e58ca59709f8f6bccfda769aae908147bf3b31efbd51beb608373d752ce1ee01

Observation 57a8272d-7e70-4836-b9c1-cf6b5f51b41c · inbound

Diagnosing Failures in Large Language Models' Answers: Integrating Error Attribution into Evaluation Framework cites this paper.

Diagnosing Failures in Large Language Models' Answers: Integrating Error Attribution into Evaluation Framework When Can LLMs Actually Correct Their Own Mistakes? A Critical Survey of Self-Correction of LLMs

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T18:23:29.536585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:23:29.536585Z digest=sha256:466944e7248f0c892538560a73d1544cb57f70ce26f832d748bda5dc07ad36b8

Observation ea6d09f7-4d09-4c89-865f-0c382b07344a · inbound

SC-Captioner: Improving Image Captioning with Self-Correction by Reinforcement Learning cites this paper.

SC-Captioner: Improving Image Captioning with Self-Correction by Reinforcement Learning When Can LLMs Actually Correct Their Own Mistakes? A Critical Survey of Self-Correction of LLMs

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-05T22:58:35.361067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:58:35.361067Z digest=sha256:a27ab56ab12dcfe596b269baf60dedc4d9819f707a515ce950549fa5f4c628fe

Observation 0d395e60-0ecc-4e15-8e5f-615f52c8605f · inbound

Agentic AI for Commercial Insurance Underwriting with Adversarial Self-Critique cites this paper.

Agentic AI for Commercial Insurance Underwriting with Adversarial Self-Critique When Can LLMs Actually Correct Their Own Mistakes? A Critical Survey of Self-Correction of LLMs

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-03T09:09:56.146465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:09:56.146465Z digest=sha256:fe2de955f78a5eee925d9595d8678cef377995c456141c8d4bea34e53a1679bf

Observation 9bab5262-b7e7-4cc0-aab5-77e258f9fd13 · inbound

FVRuleLearner: Operator-Level Reasoning Tree (Op-Tree)-Based Rules Learning for Formal Verification cites this paper.

FVRuleLearner: Operator-Level Reasoning Tree (Op-Tree)-Based Rules Learning for Formal Verification When Can LLMs Actually Correct Their Own Mistakes? A Critical Survey of Self-Correction of LLMs

Reference 45

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T14:35:55.916505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-15T14:31:36.233977Z digest=sha256:f26b1fe5772f951a8e684ea48d9348e3c356143fe63bd266742800c26a7b8ad3

Observation 947aeea8-eda5-46ad-a249-78ca3b490dec · inbound

FVRuleLearner: Operator-Level Reasoning Tree (Op-Tree)-Based Rules Learning for Formal Verification cites this paper.

FVRuleLearner: Operator-Level Reasoning Tree (Op-Tree)-Based Rules Learning for Formal Verification When Can LLMs Actually Correct Their Own Mistakes? A Critical Survey of Self-Correction of LLMs

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-02T18:43:47.090032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:43:47.090032Z digest=sha256:d2bd98afe7d0c14f355d8e28e35507c16171494f430f57b19a784fe333669231

Observation 2b771e12-cf91-4116-9dd0-8ca7af3cd7b6 · inbound

Self-Preference Is Weak or Absent in Verifiable Instruction-Following Revision: A Four-Model Test Under Genuine Authorship cites this paper.

Self-Preference Is Weak or Absent in Verifiable Instruction-Following Revision: A Four-Model Test Under Genuine Authorship When Can LLMs Actually Correct Their Own Mistakes? A Critical Survey of Self-Correction of LLMs

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T03:59:33.350376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-26T17:26:22.687010Z digest=sha256:87ca8bed8e786311868080ea27576b331c87dd86e87fbd3b02194907f33b4792

Observation eb4fb963-1868-4733-af0d-eaa11a3e8fa9 · inbound

LLM-as-a-Judge Scores Are Unreliable Optimization Signals in Closed-Loop Table Recognition cites this paper.

LLM-as-a-Judge Scores Are Unreliable Optimization Signals in Closed-Loop Table Recognition When Can LLMs Actually Correct Their Own Mistakes? A Critical Survey of Self-Correction of LLMs

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-02T05:30:00.881954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T05:30:00.881954Z digest=sha256:4538a34d9fd0d5d537732dc443a87e0a9fc6a5ce486eff7cef6de70ebaf4ab8b

Observation 7a3fa1e7-b0e4-43ca-862e-76b7460d077c · inbound

Sample More, Reflect Less: Self-Refine and Reflexion Lose to Repeated Sampling at Equal Token Cost, from 1.5B to 7B cites this paper.

Sample More, Reflect Less: Self-Refine and Reflexion Lose to Repeated Sampling at Equal Token Cost, from 1.5B to 7B When Can LLMs Actually Correct Their Own Mistakes? A Critical Survey of Self-Correction of LLMs

Reference 12

Resolution
unresolved
no resolver link, observed 2026-07-31T03:18:59.984759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T03:18:59.984759Z digest=sha256:9506e11613c53a05a096f3034c5708513824258713cf132026613784749c41ac