Pith. sign in

Paper Citation Record · LEDGER

Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models

As of 7 August 2026, this Paper Citation Record lists 60 of 60 outbound references and 4 inbound Pith citation observations for arXiv:2507.02778.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.02778 v3

Coverage vector

measured 60 of 60 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T20:29:03.805548Z

measured 64 of 64 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-03T21:20:00.041277Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T21:28:58.386085Z

Reference resolution

60 of 60 outbound references displayed

  • verified exact0
  • verified fuzzy15
  • unresolved45
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 19a388be-1abf-4a33-ba82-7b35e40fd82f · outbound

This paper cites GPT-4 Technical Report.

Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T20:28:55.823384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:28:55.823384Z digest=sha256:bdf600a5e9bbdf725f39810ec333c1d5351350048af2e9f353e49f465b4b383a

Observation 73f0cd9a-0b1b-408e-8072-9d94b04998e2 · outbound

This paper cites The claude 3 model family: Opus, sonnet, haiku, Mar 2024.

Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models The claude 3 model family: Opus, sonnet, haiku, Mar 2024

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:29:08.781161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T20:28:55.944606Z digest=sha256:3025982ffc81acb24b9b4d0cdab37a32df3512e1d2ec64574b7d3a9fcd1f329e

Observation 50bbfc04-ce8c-4b4b-b5b0-3fff3e8b3f5a · outbound

This paper cites Gemini 2.5: Pushing the frontier with advanced reasoning, multimodality, long context, and next generation agentic capabilities., June 2025.

Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models Gemini 2.5: Pushing the frontier with advanced reasoning, multimodality, long context, and next generation agentic capabilities., June 2025

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:29:08.531479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T20:28:56.097504Z digest=sha256:1c1ad316ea790da524c76b63b035376944396f09c13a97ba8753fddb06dc1769

Observation ae1989c8-73a6-49a0-8fc4-2d16699ff84d · outbound

This paper cites Qwen3 Technical Report.

Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models Qwen3 Technical Report

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T20:28:56.227672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:28:56.227672Z digest=sha256:a0a8f6530b3be77859e622d0dfca0274100f57da6327829168a69ae44fb4bf6f

Observation 586457e5-ab3b-4b4a-9464-8e20506298a9 · outbound

This paper cites The llama 4 herd: The beginning of a new era of natively multimodal ai innovation, Apr 2025.

Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models The llama 4 herd: The beginning of a new era of natively multimodal ai innovation, Apr 2025

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:29:08.248924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T20:28:56.333821Z digest=sha256:8370a31e980ad09314f9ed60933d767629a6dba6eb1b60b82383a94f4bc87ef5

Observation 91ab01f4-557c-433b-b672-94dbef010bc0 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T20:28:56.531856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:28:56.531856Z digest=sha256:44eddd690547449b12eaacce63ae713ad9ff1172c1c3c9bcef9d93aced5566f2

Observation f38e80ac-bd0f-42de-b022-b8b7c6500e47 · outbound

This paper cites On faithfulness and factuality in abstractive summarization.

Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models On faithfulness and factuality in abstractive summarization

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T20:28:56.646944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:28:56.646944Z digest=sha256:0323ba82fd3790475a427b73696e3ea2385b04345c49ff480f56ecafa90f968f

Observation 77338fbe-e90f-4bc5-bba8-76b029fc8ea4 · outbound

This paper cites A survey on hallucination in large language models: Principles, taxonomy, challenges, and open questions.

Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models A survey on hallucination in large language models: Principles, taxonomy, challenges, and open questions

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T20:28:56.783467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:28:56.783467Z digest=sha256:5c94dad506ef5959d8c2de17f3f7352a725076a3882015cd6cb69fdb13ffd09f

Observation 098af2f9-edf9-49c1-a735-ebf863240b3d · outbound

This paper cites Do, Yan Xu, and Pascale Fung.

Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models Do, Yan Xu, and Pascale Fung

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T20:28:56.946027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:28:56.946027Z digest=sha256:824c5bb47fd1ea20f5f81fbfa0116ffaf94cb908eaffcb67b95dca46dded8240

Observation caa6fbd6-6cbc-4f13-b88a-1bcb810818e9 · outbound

This paper cites Large language models can be easily distracted by irrelevant context.

Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models Large language models can be easily distracted by irrelevant context

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T20:28:57.071428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:28:57.071428Z digest=sha256:a9ea0184110478786610723ebbc8e18b3fe1caf62b003a570f7e97e4d8acdc39

Observation d87bcd25-bf69-48cd-b7fa-c14ce7548102 · outbound

This paper cites Alice in Wonderland: Simple Tasks Showing Complete Reasoning Breakdown in State-Of-the-Art Large Language Models.

Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models Alice in Wonderland: Simple Tasks Showing Complete Reasoning Breakdown in State-Of-the-Art Large Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T20:28:57.217863Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:28:57.217863Z digest=sha256:fca6b7b853d597527fcbc5f39900110800fd738d53f8786194dc98d0ad83f3fd

Observation fae702c4-d4da-4075-8c5e-a0238f612dfe · outbound

This paper cites Reflexion: language agents with verbal reinforcement learning.

Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models Reflexion: language agents with verbal reinforcement learning

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:29:08.005937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T20:28:57.332058Z digest=sha256:7531474b8321e89b393cf6ee555418e0ad4e9208f7be5674ddc05dd04793b486

Observation 1cecc2d2-7db8-48a5-94cf-cff9ceed325b · outbound

This paper cites Self-refine: Iterative refinement with self-feedback.

Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models Self-refine: Iterative refinement with self-feedback

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T20:28:57.429687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:28:57.429687Z digest=sha256:cacd72258b0e19cfbc66a23f45164cd6155fcc29f6ebb43bfb4660b16478256d

Observation 0952229d-8e2e-4192-9c73-59e1a4b6dcff · outbound

This paper cites Language models can solve computer tasks.

Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models Language models can solve computer tasks

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:29:07.724534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T20:28:57.560434Z digest=sha256:3204ea306da5f071dad9d59733a673fc6f1f897785c71c601c41f26f81c006c9

Observation ee98ddf2-8acf-4313-9d25-4c6955fe54ee · outbound

This paper cites When can LLM s actually correct their own mistakes? a critical survey of self-correction of LLM s.

Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models When can LLM s actually correct their own mistakes? a critical survey of self-correction of LLM s

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T20:28:57.717223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:28:57.717223Z digest=sha256:4f2830a93c6552df60376725b5e297b2dd01a6a70c5381a21395ce3fa223d746

Observation 7cb640df-e27c-4882-80c9-919eeba976dc · outbound

This paper cites Large Language Models Cannot Self-Correct Reasoning Yet.

Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models Large Language Models Cannot Self-Correct Reasoning Yet

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T20:28:57.838761Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:28:57.838761Z digest=sha256:ae0df6bd87395211f474b56e313df0c88c1ff66244645f9641ed17b3ac92d9eb

Observation 8923266a-b790-4707-a354-bce566c6368d · outbound

This paper cites LLM s cannot find reasoning errors, but can correct them given the error location.

Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models LLM s cannot find reasoning errors, but can correct them given the error location

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T20:28:57.998432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:28:57.998432Z digest=sha256:1c1f4372530925dbe12cc850675b8f57bd32411522460a60d7e7eed87af88153

Observation ed5bf506-a601-41b5-a800-544053498431 · outbound

This paper cites Evaluating LLM s at detecting errors in LLM responses.

Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models Evaluating LLM s at detecting errors in LLM responses

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:29:07.454677Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T20:28:58.142786Z digest=sha256:4832861b4be4f83f1471d2b250838fc0bb06f103ee5fb95f4bf11b08e2d4da87

Observation 5f1f55b3-3870-468a-9f67-1af9ce35a562 · outbound

This paper cites Training language models to self-correct via reinforcement learning.

Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models Training language models to self-correct via reinforcement learning

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:29:07.214068Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T20:28:58.292638Z digest=sha256:2078769dc19447d48e2982e9746f6ea6244fb69a39d9f6594971c2a725682d1d

Observation efddf2d7-1f69-4927-bb0a-f4f3896e609c · outbound

This paper cites Jailbroken: how does llm safety training fail? In Proceedings of the 37th International Conference on Neural Information Processing Systems, NIPS '23, Red Hook, NY, USA, 2023.

Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models Jailbroken: how does llm safety training fail? In Proceedings of the 37th International Conference on Neural Information Processing Systems, NIPS '23, Red Hook, NY, USA, 2023

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:29:06.943034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T20:28:58.454749Z digest=sha256:9453588348f0992c8565b6832535f7d510bf8e9dfdf623dbc34cf8ffbdff93a3

Observation 553309ad-0b0f-4417-b8b1-f864b3625a67 · outbound

This paper cites Formalizing and benchmarking prompt injection attacks and defenses.

Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models Formalizing and benchmarking prompt injection attacks and defenses

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:29:06.592654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T20:28:58.558097Z digest=sha256:434370fd6888258efd531626ff1e39343caec3019b8be4fd84029f9bd134dd81

Observation 7caeb070-8c3f-4b21-a32e-ab3575a5ea7d · outbound

This paper cites Measuring Faithfulness in Chain-of-Thought Reasoning.

Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models Measuring Faithfulness in Chain-of-Thought Reasoning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T20:28:58.704896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:28:58.704896Z digest=sha256:bbb68cad7e465828e9ba0fa270a369c2c65a3f17da0190e5763ac2782029d0d2

Observation 9aced9d1-1d08-48f8-87f4-1e5b8e6a1a5b · outbound

This paper cites an unresolved cited work.

Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-06T20:29:06.302848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T20:28:58.869328Z digest=sha256:24d235d111dffd7349476165d674c53328607795b047c1b23770d121eafb8908

Observation 6e96adf5-08a3-472a-8d93-9315cfd851d5 · outbound

This paper cites Scaling LLM test-time compute optimally can be more effective than scaling parameters for reasoning.

Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models Scaling LLM test-time compute optimally can be more effective than scaling parameters for reasoning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T20:28:59.074923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:28:59.074923Z digest=sha256:ee85035a99e96adea9494c3c18da278e2d8994a560901db6df4c1012337af919

Observation a61249e0-c3ef-455b-b871-6c8f7c79c70c · outbound

This paper cites s1: Simple test-time scaling.

Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models s1: Simple test-time scaling

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T20:28:59.270560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:28:59.270560Z digest=sha256:6caee72957f19e16b84c37747bc599323b820a19f7aebe16a65dfc6115008b12

Observation f0954cbd-2c14-4b6b-8801-899f2a583ed1 · outbound

This paper cites Benchmarking cognitive biases in large language models as evaluators.

Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models Benchmarking cognitive biases in large language models as evaluators

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T20:28:59.400211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:28:59.400211Z digest=sha256:0b9d454d297717dc470b57fe3c180e7838c337f37ada022c4687d96a0a061b64

Observation 78c11968-cc45-47a7-b086-2d93729aa853 · outbound

This paper cites Cognitive bias in decision-making with LLM s.

Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models Cognitive bias in decision-making with LLM s

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T20:28:59.570749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:28:59.570749Z digest=sha256:b33d2b0b0e65c1a64315adda3535d5625588f9006356822ca144fbde9b3a5c25

Observation bc521204-141d-47ff-be90-5f34ec347f6e · outbound

This paper cites Capturing failures of large language models via human cognitive biases.

Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models Capturing failures of large language models via human cognitive biases

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:29:06.008527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T20:28:59.745163Z digest=sha256:e6f39f8900bc1aeac31fb4dc7454fb470506805cfadf4a852f565e26d403ae81

Observation 42e7bac0-e04d-4130-8fc7-375e883a226c · outbound

This paper cites Lin, and Lee Ross.

Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models Lin, and Lee Ross

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T20:28:59.906202Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:28:59.906202Z digest=sha256:39db37328ff10f6c621b9ac35925227ff7420cf5fe54c9e22620aca2d2f59c96

Observation 2c893750-911c-4989-a3ab-7264afd41c4f · outbound

This paper cites ProcessBench: Identifying Process Errors in Mathematical Reasoning.

Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models ProcessBench: Identifying Process Errors in Mathematical Reasoning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:00.072960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:29:00.072960Z digest=sha256:6c40740dac64b1d12f9e7121727457d2d4bec21aa455abb6f839ff18654cfdd2

Observation af15ffc1-4b55-4d7b-893f-55462028cbe3 · outbound

This paper cites PRMBench: A Fine-grained and Challenging Benchmark for Process-Level Reward Models.

Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models PRMBench: A Fine-grained and Challenging Benchmark for Process-Level Reward Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:00.154021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:29:00.154021Z digest=sha256:796ca3232038a9f82563c37aa1b4bb00e0dc2a3983bfe9b44d38221e855c5da3

Observation 242c2d42-a346-4983-a5f9-e9fb052d06a4 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models Training Verifiers to Solve Math Word Problems

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:00.206673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:29:00.206673Z digest=sha256:e00f64d94833b3feb4222a0c0c6925707ceed22421381a2c103d25ca92f31c11

Observation 0a682beb-9d8b-4ffc-8185-1c377d0db050 · outbound

This paper cites Introducing gpt-4.1 in the api, Apr 2025.

Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models Introducing gpt-4.1 in the api, Apr 2025

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:29:05.741704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T20:29:00.270483Z digest=sha256:3b793d0487e2a1f3f1d05bc173458f29b2f617caf9d4ee3c998b8cb7d748c8e3

Observation 547769ab-98dd-42f4-8e5d-f44fecc985c4 · outbound

This paper cites Let's verify step by step.

Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models Let's verify step by step

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:00.355018Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:29:00.355018Z digest=sha256:5f1801ddcfab59147716300d4d17bd0d744ca63f5403f96c21211341dd939131

Observation cdac7ea2-3ff3-405a-9392-f1aa43da387b · outbound

This paper cites Measuring mathematical problem solving with the MATH dataset.

Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models Measuring mathematical problem solving with the MATH dataset

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:00.438344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:29:00.438344Z digest=sha256:deddd798fee44866379c30824ac6f5c0e0805e23dc1371d7d9602259fc73306e

Observation d1a30e71-b262-458f-8efa-c396237e35b6 · outbound

This paper cites Transformers: State-of-the-art natural language processing.

Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models Transformers: State-of-the-art natural language processing

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:00.521903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:29:00.521903Z digest=sha256:73095f91931d447dfa80c9924c787f5093814b7c96296b73750bed83e17fa8dc

Observation 98372d2d-19f9-45e5-8238-be050bbd90ed · outbound

This paper cites DeepSeek-V3 Technical Report.

Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models DeepSeek-V3 Technical Report

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:00.633907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:29:00.633907Z digest=sha256:a96a5a97578bcc774f60e50bf4afb1f7138591920d032475269375877fe87f26

Observation 5e52510a-4855-4949-9e1d-a50fd817f0f4 · outbound

This paper cites Qwen2.5 Technical Report.

Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models Qwen2.5 Technical Report

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:00.711530Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:29:00.711530Z digest=sha256:3d70ae7e32c8bac72ea2fa3f229959351ed69dc54f24a9d1cc4d26762f5bd77f

Observation a5b12f90-c1f5-4320-988b-69bde06197c9 · outbound

This paper cites Llama 3.3, Dec 2024.

Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models Llama 3.3, Dec 2024

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:29:05.454473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T20:29:00.790776Z digest=sha256:4f32c97880dde1a435cb3023556d0bd1eec3ec58b2db289cebfcc160f1f9715b

Observation 070d8f16-3cc9-4caa-ba86-afcde95ccfc4 · outbound

This paper cites Phi-4 Technical Report.

Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models Phi-4 Technical Report

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:00.903257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:29:00.903257Z digest=sha256:6f284ae9831fbad5b01caea62863abf4d90ced14e54667d5e96193da0004a3a8

Observation d8898eac-0793-4381-a9b0-2df54c6e3a79 · outbound

This paper cites Qwen2 Technical Report.

Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models Qwen2 Technical Report

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:01.073515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:29:01.073515Z digest=sha256:319c80c6d0511cf68a29dd29f83d87892360ccdd7c5d2312ad2d739e86500f19

Observation 598889cd-4a56-43db-8cd4-7e4e3b4c55bc · outbound

This paper cites The Llama 3 Herd of Models.

Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models The Llama 3 Herd of Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:01.235737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:29:01.235737Z digest=sha256:2cc8d7bc77512073b46e234323883ec16b7c6a08d8166c9e9c3023f8b98b01b1

Observation 9c5dc13c-19b8-4c72-8e21-1ba13f577e88 · outbound

This paper cites Mistral small 3, Jan 2025.

Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models Mistral small 3, Jan 2025

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:29:05.122636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T20:29:01.386989Z digest=sha256:8394cfe3db6bc5e7de71e363c5012e974b329090ab34933267f60b99e69b44c4

Observation 91a1453a-6c9b-4155-aebf-55fa2131cd9f · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:01.501958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:29:01.501958Z digest=sha256:26bfd41365fdf1e5a831d9f5b9e921720d2ab9adc4e1c5b6eefa2987195836f1

Observation 8a78e84b-9451-4307-9a7a-d78784925c50 · outbound

This paper cites Impact of pretraining term frequencies on few-shot numerical reasoning.

Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models Impact of pretraining term frequencies on few-shot numerical reasoning

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:01.660462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:29:01.660462Z digest=sha256:1875b538a9da3e07c1391c6f618cf25cfb3e45601d2a7b3e8801c25f41da3001

Observation a8943ef9-1315-4cd0-8398-f439535fbb8e · outbound

This paper cites Smith, Sarah Wiegreffe, and Yanai Elazar.

Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models Smith, Sarah Wiegreffe, and Yanai Elazar

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:29:04.782194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T20:29:01.845245Z digest=sha256:5182b2067b1d27b6926a2eca05363d71a89966b204d131ff6c4bc20db5253115

Observation e5ca97be-2dab-4115-a7b7-171ee449086e · outbound

This paper cites o pf, Yannic Kilcher, Dimitri von R \.

Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models o pf, Yannic Kilcher, Dimitri von R \

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:01.991836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:29:01.991836Z digest=sha256:9c9351daa39af124c4b1038f96161907881e342a03a9e9b7f952b080256f225a

Observation 9472acb6-eb08-412e-a62b-0d21af73a4f9 · outbound

This paper cites Openhermes 2.5: An open dataset of synthetic data for generalist llm assistants, 2023.

Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models Openhermes 2.5: An open dataset of synthetic data for generalist llm assistants, 2023

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:02.140460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:29:02.140460Z digest=sha256:afd76aca1f101d77c37e2790123c988f02e4de8f538e08e4439da04cd0b5396e

Observation c317eade-a7e4-48ad-8654-5687149260d8 · outbound

This paper cites Infinity Instruct: Scaling Instruction Selection and Synthesis to Enhance Language Models.

Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models Infinity Instruct: Scaling Instruction Selection and Synthesis to Enhance Language Models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:02.244807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:29:02.244807Z digest=sha256:711fb994ffd591129ce73eb60a6142b77aa79c2f5c6c99cc969e2e989bbafb2a

Observation 61ed6acb-e56a-49e0-8e0a-49c8c23c40b5 · outbound

This paper cites Ultrafeedback: Boosting language models with high-quality feedback, 2024.

Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models Ultrafeedback: Boosting language models with high-quality feedback, 2024

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:29:04.501080Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T20:29:02.356141Z digest=sha256:c3d22ca3f5f3ed7d961ca0dba85ea3559d6ffdbca622631c3198e488e93a49a1

Observation d42f5cd2-00f6-4eeb-98c0-831a16acdeb7 · outbound

This paper cites Hwang, Jiangjiang Yang, Ronan Le Bras, Oyvind Tafjord, Christopher Wilhelm, Luca Soldaini, Noah A.

Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models Hwang, Jiangjiang Yang, Ronan Le Bras, Oyvind Tafjord, Christopher Wilhelm, Luca Soldaini, Noah A

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:02.487929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:29:02.487929Z digest=sha256:7d4a2e108fe96be94f0f126067508ecb9880a62c99da081849535c26910e6ae6

Observation fa7a74bf-6fcc-4198-83eb-b128ade148e9 · outbound

This paper cites Open r1: A fully open reproduction of deepseek-r1, January 2025.

Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models Open r1: A fully open reproduction of deepseek-r1, January 2025

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:02.599732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:29:02.599732Z digest=sha256:e3023db1bda9928ae2fc817338e54972da7a75489e5658cf8cfd7a7601c8dbc2

Observation fab9b386-0834-49e7-9db5-dc727f9131ea · outbound

This paper cites OpenThoughts: Data Recipes for Reasoning Models.

Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models OpenThoughts: Data Recipes for Reasoning Models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:02.783626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:29:02.783626Z digest=sha256:e33eed2e474efda4bf3c86439f5dbafeba851ca76a0754d5fedce3c346535630

Observation 432e34ca-6d12-4311-9df2-c0977733d550 · outbound

This paper cites Training language models to follow instructions with human feedback.

Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models Training language models to follow instructions with human feedback

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:02.938680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:29:02.938680Z digest=sha256:dfa6bb9057b25478f898f5078c1577e3a710ab5edbf14c794826c919942c9557

Observation 5a580e71-de61-4a51-819e-81f6b955611e · outbound

This paper cites Learning From Mistakes Makes LLM Better Reasoner.

Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models Learning From Mistakes Makes LLM Better Reasoner

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:03.095630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:29:03.095630Z digest=sha256:bf24694ae0056f9294c2bf5d81284b42ef3e8e123af0bc6cc95a285e1bf421f7

Observation f256587f-fd18-4918-be5b-1e797c5c666e · outbound

This paper cites Critique Fine-Tuning: Learning to Critique is More Effective than Learning to Imitate.

Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models Critique Fine-Tuning: Learning to Critique is More Effective than Learning to Imitate

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:03.216236Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:29:03.216236Z digest=sha256:544cc1133426d934c517dc56345cc9ab158dc3ac4546225ed20aef6332b6904a

Observation 4eb03d92-b6a5-4658-90b2-9dd56e085279 · outbound

This paper cites The effect of sampling temperature on problem solving in large language models.

Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models The effect of sampling temperature on problem solving in large language models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:03.409054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:29:03.409054Z digest=sha256:4ce87ea3d257e0bae16a50ee6787167ac2a9157f921b57615121725d27185ba4

Observation d70d4393-6219-408a-8892-8d96da99a79c · outbound

This paper cites @esa (Ref.

Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models @esa (Ref

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:03.530733Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:29:03.530733Z digest=sha256:b542cdb8ef763d8788a05418c90670b695635d5a9e8e7e0573a9f17742d6cf7f

Observation a4614c9e-01f7-4332-adb6-eec84c16fca5 · outbound

This paper cites an unresolved cited work.

Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models Unresolved cited work

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:03.675873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:29:03.675873Z digest=sha256:998591081411da8a8ef5d48353593fbb58a796738adf4ee5f14c33f9663e5315

Observation e8dda0b1-32fe-40b6-bebf-104c5fd66015 · outbound

This paper cites after incorrect reasoning or answer to prompt LLMs to self-correct, without finetuning. We observe significant reductions in the blind spot after appending ``Wait.

Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models after incorrect reasoning or answer to prompt LLMs to self-correct, without finetuning. We observe significant reductions in the blind spot after appending ``Wait

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:03.805548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:29:03.805548Z digest=sha256:947ec1df66288c40fa42a837119fa0e4291304823f42096ecf856e86c2279df1

Pith citing papers

Observation 3547e909-d01a-42bc-8fbe-10036c058fef · inbound

ReFlect: An Effective Harness System for Complex Long-Horizon LLM Reasoning cites this paper.

ReFlect: An Effective Harness System for Complex Long-Horizon LLM Reasoning Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models

Reference 28

Resolution
metadata mismatch
arxiv_id, observed 2026-08-04T02:29:36.375484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-08T11:35:25.204350Z digest=sha256:c02bb7ece9bc5dabb2b9bb792b15b0415493ed1121920f7820812a39862cf055

Observation 96d7ffc6-6dae-45a3-989d-37e0f37d0836 · inbound

ProCrit: Self-Elicited Multi-Perspective Reasoning with Critic-Guided Revision for Multimodal Sarcasm Detection cites this paper.

ProCrit: Self-Elicited Multi-Perspective Reasoning with Critic-Guided Revision for Multimodal Sarcasm Detection Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-08-04T02:29:36.375484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-21T02:35:37.724212Z digest=sha256:0782dd64adbb782b598fa26c8a36e070a615abf56fceaac1df91e65137c8da42

Observation 3a91d3a0-6f38-4726-acc8-06b923d8d159 · inbound

Mixture of Debaters: Learn to Debate at Architectural Level in Multi-Agent Reasoning cites this paper.

Mixture of Debaters: Learn to Debate at Architectural Level in Multi-Agent Reasoning Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-08-04T02:29:36.375484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T07:11:02.464556Z digest=sha256:92f905919db9965d3fd1ddcdc1efb75d019e1b209e5396945f95d047b2e9aca4

Observation eaa9921f-da88-4f3c-9369-e2f5a7f177a6 · inbound

ESC: Emotional Self-Correction for Reliable Vision-Language Models cites this paper.

ESC: Emotional Self-Correction for Reliable Vision-Language Models Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models

Reference 85

Resolution
metadata mismatch
arxiv_id, observed 2026-08-04T02:29:36.375484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-03T21:20:00.041277Z digest=sha256:2ec9a0a140f5fc2c6f43e96e168baa137bc806d746f115d5b5adf70be6f99595