Pith. sign in

Paper Citation Record · LEDGER

Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models

As of 22 August 2026, this Paper Citation Record lists 60 of 60 outbound references and 4 inbound Pith citation observations for arXiv:2507.02778.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.02778 v3

Coverage vector

measured 60 of 60 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T20:29:03.805548Z

measured 64 of 64 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-03T21:20:00.041277Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T21:28:58.386085Z

Reference resolution

60 of 60 outbound references displayed

  • verified exact0
  • verified fuzzy15
  • unresolved45
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 19a388be-1abf-4a33-ba82-7b35e40fd82f · outbound

This paper cites GPT-4 Technical Report.

Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T20:28:55.823384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:28:55.823384Z digest=sha256:bbf547e464fca9c807a0c1a85f0307da08e5786187a217b9602bcd486beee5df

Observation 73f0cd9a-0b1b-408e-8072-9d94b04998e2 · outbound

This paper cites The claude 3 model family: Opus, sonnet, haiku, Mar 2024.

Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models The claude 3 model family: Opus, sonnet, haiku, Mar 2024

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:29:08.781161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-06T20:28:55.944606Z digest=sha256:ad3702a5c63a233e07e0d28d5a406c459059c9876a85738f459c8b61dae5221d

Observation 50bbfc04-ce8c-4b4b-b5b0-3fff3e8b3f5a · outbound

This paper cites Gemini 2.5: Pushing the frontier with advanced reasoning, multimodality, long context, and next generation agentic capabilities., June 2025.

Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models Gemini 2.5: Pushing the frontier with advanced reasoning, multimodality, long context, and next generation agentic capabilities., June 2025

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:29:08.531479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-06T20:28:56.097504Z digest=sha256:cfc29db509e432fb48043715305cd8aabc5b3b147f7be9d4e89eb30f7cbf61c5

Observation ae1989c8-73a6-49a0-8fc4-2d16699ff84d · outbound

This paper cites Qwen3 Technical Report.

Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models Qwen3 Technical Report

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T20:28:56.227672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:28:56.227672Z digest=sha256:55790898f08c6801a698014f6fa5797e702ee7b215c6a9d6debda0951b340838

Observation 586457e5-ab3b-4b4a-9464-8e20506298a9 · outbound

This paper cites The llama 4 herd: The beginning of a new era of natively multimodal ai innovation, Apr 2025.

Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models The llama 4 herd: The beginning of a new era of natively multimodal ai innovation, Apr 2025

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:29:08.248924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-06T20:28:56.333821Z digest=sha256:faf5b315ec5f9a04a9f2fc33a0c3279d8b0eeeaab47792bc051a15234a20a7e9

Observation 91ab01f4-557c-433b-b672-94dbef010bc0 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T20:28:56.531856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:28:56.531856Z digest=sha256:5754d1247bfd8390c2724625f67a57e4cae19b37bc7a03ddb0a74fbd8b114c7d

Observation f38e80ac-bd0f-42de-b022-b8b7c6500e47 · outbound

This paper cites On faithfulness and factuality in abstractive summarization.

Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models On faithfulness and factuality in abstractive summarization

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T20:28:56.646944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:28:56.646944Z digest=sha256:1892214022659174e7bf3ec2dd8ddb8a54a3a4ebc69cda442a1b79fc727d46ca

Observation 77338fbe-e90f-4bc5-bba8-76b029fc8ea4 · outbound

This paper cites A survey on hallucination in large language models: Principles, taxonomy, challenges, and open questions.

Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models A survey on hallucination in large language models: Principles, taxonomy, challenges, and open questions

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T20:28:56.783467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:28:56.783467Z digest=sha256:944e09de86c5f5d14e8f2e0a3633dca2c154af1141666ff3901a1f7fcc0cb693

Observation 098af2f9-edf9-49c1-a735-ebf863240b3d · outbound

This paper cites Do, Yan Xu, and Pascale Fung.

Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models Do, Yan Xu, and Pascale Fung

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T20:28:56.946027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:28:56.946027Z digest=sha256:2b25f40f4c41930c516157d5dbfcede9d45fe27ff239987f4154ee7a1ad92676

Observation caa6fbd6-6cbc-4f13-b88a-1bcb810818e9 · outbound

This paper cites Large language models can be easily distracted by irrelevant context.

Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models Large language models can be easily distracted by irrelevant context

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T20:28:57.071428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:28:57.071428Z digest=sha256:d5b416c79e6bbaa9dbb9c7bfdae1c830a02df119ffdc66f7c054e240fcdcf0ac

Observation d87bcd25-bf69-48cd-b7fa-c14ce7548102 · outbound

This paper cites Alice in Wonderland: Simple Tasks Showing Complete Reasoning Breakdown in State-Of-the-Art Large Language Models.

Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models Alice in Wonderland: Simple Tasks Showing Complete Reasoning Breakdown in State-Of-the-Art Large Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T20:28:57.217863Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:28:57.217863Z digest=sha256:379e4dd0394e567dbafdd0ee118b532abc6eb6b030d1c0dfd747fc1b5fcfe8e5

Observation fae702c4-d4da-4075-8c5e-a0238f612dfe · outbound

This paper cites Reflexion: language agents with verbal reinforcement learning.

Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models Reflexion: language agents with verbal reinforcement learning

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:29:08.005937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-06T20:28:57.332058Z digest=sha256:6babaa6d6d0d17dbe80d7845ed0d8db1c78eb1d63d02f69a0ca777cc171fccee

Observation 1cecc2d2-7db8-48a5-94cf-cff9ceed325b · outbound

This paper cites Self-refine: Iterative refinement with self-feedback.

Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models Self-refine: Iterative refinement with self-feedback

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T20:28:57.429687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:28:57.429687Z digest=sha256:3b8cebfc33b7f2b930b0ade237febe1316addaea35e91b29577d73037bdf5471

Observation 0952229d-8e2e-4192-9c73-59e1a4b6dcff · outbound

This paper cites Language models can solve computer tasks.

Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models Language models can solve computer tasks

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:29:07.724534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-06T20:28:57.560434Z digest=sha256:298751b69abc57410760cf39e2afd1fdae67c6b65102088a405832acc507395c

Observation ee98ddf2-8acf-4313-9d25-4c6955fe54ee · outbound

This paper cites When can LLM s actually correct their own mistakes? a critical survey of self-correction of LLM s.

Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models When can LLM s actually correct their own mistakes? a critical survey of self-correction of LLM s

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T20:28:57.717223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:28:57.717223Z digest=sha256:c52ca747a347d4c72c8fc5e1e82c3895eb642038cd1371d23a629df59c6142ec

Observation 7cb640df-e27c-4882-80c9-919eeba976dc · outbound

This paper cites Large Language Models Cannot Self-Correct Reasoning Yet.

Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models Large Language Models Cannot Self-Correct Reasoning Yet

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T20:28:57.838761Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:28:57.838761Z digest=sha256:4b65fb404b9d2575a4b89c276362acb6e74c1037a6664398e1f279411c0c24a8

Observation 8923266a-b790-4707-a354-bce566c6368d · outbound

This paper cites LLM s cannot find reasoning errors, but can correct them given the error location.

Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models LLM s cannot find reasoning errors, but can correct them given the error location

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T20:28:57.998432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:28:57.998432Z digest=sha256:21f90ee0745a73fc3251af000c9e9ed93f24e7bd3d48c25e8900224e92a0e5bb

Observation ed5bf506-a601-41b5-a800-544053498431 · outbound

This paper cites Evaluating LLM s at detecting errors in LLM responses.

Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models Evaluating LLM s at detecting errors in LLM responses

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:29:07.454677Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-06T20:28:58.142786Z digest=sha256:054da93ae1f16901139ad731dc0ac2b600af179ce811adaebb70c82e8334da83

Observation 5f1f55b3-3870-468a-9f67-1af9ce35a562 · outbound

This paper cites Training language models to self-correct via reinforcement learning.

Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models Training language models to self-correct via reinforcement learning

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:29:07.214068Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-06T20:28:58.292638Z digest=sha256:e471bbbe28727d2930759c5e5a536202e491a42f19a11ac643fccffd6b43a3a8

Observation efddf2d7-1f69-4927-bb0a-f4f3896e609c · outbound

This paper cites Jailbroken: how does llm safety training fail? In Proceedings of the 37th International Conference on Neural Information Processing Systems, NIPS '23, Red Hook, NY, USA, 2023.

Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models Jailbroken: how does llm safety training fail? In Proceedings of the 37th International Conference on Neural Information Processing Systems, NIPS '23, Red Hook, NY, USA, 2023

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:29:06.943034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-06T20:28:58.454749Z digest=sha256:f46d5a87714e8066ed6bc5c996bed53c23c08fdf61fcdf0f59385b37adca2e76

Observation 553309ad-0b0f-4417-b8b1-f864b3625a67 · outbound

This paper cites Formalizing and benchmarking prompt injection attacks and defenses.

Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models Formalizing and benchmarking prompt injection attacks and defenses

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:29:06.592654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-06T20:28:58.558097Z digest=sha256:4fc4fc0c033aaa8b791664b558beb466b1f618d4212f8a97c86a651a620f5e91

Observation 7caeb070-8c3f-4b21-a32e-ab3575a5ea7d · outbound

This paper cites Measuring Faithfulness in Chain-of-Thought Reasoning.

Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models Measuring Faithfulness in Chain-of-Thought Reasoning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T20:28:58.704896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:28:58.704896Z digest=sha256:2bce10f411d103af9924b2a8c0cdbfd83de10354c485dbaee3576d4bb2e07550

Observation 9aced9d1-1d08-48f8-87f4-1e5b8e6a1a5b · outbound

This paper cites an unresolved cited work.

Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-06T20:29:06.302848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-06T20:28:58.869328Z digest=sha256:71051adf6bc7915dd16dfb3eb684d20256a8ef56e5dafa4acb98b763a295824c

Observation 6e96adf5-08a3-472a-8d93-9315cfd851d5 · outbound

This paper cites Scaling LLM test-time compute optimally can be more effective than scaling parameters for reasoning.

Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models Scaling LLM test-time compute optimally can be more effective than scaling parameters for reasoning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T20:28:59.074923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:28:59.074923Z digest=sha256:963826393b0b9d2f8bffb287843714631ff0803b4cbdc31a4e4d8bba33800747

Observation a61249e0-c3ef-455b-b871-6c8f7c79c70c · outbound

This paper cites s1: Simple test-time scaling.

Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models s1: Simple test-time scaling

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T20:28:59.270560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:28:59.270560Z digest=sha256:b6e9f949bb4fbbcd4902fd58f9ea004cfea23c6d89a02c6737f6b142da6205b3

Observation f0954cbd-2c14-4b6b-8801-899f2a583ed1 · outbound

This paper cites Benchmarking cognitive biases in large language models as evaluators.

Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models Benchmarking cognitive biases in large language models as evaluators

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T20:28:59.400211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:28:59.400211Z digest=sha256:4541ce8627d70e1a65bee91f8c8876c1d825c3e31b315caa7568cbfcd4ec9e46

Observation 78c11968-cc45-47a7-b086-2d93729aa853 · outbound

This paper cites Cognitive bias in decision-making with LLM s.

Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models Cognitive bias in decision-making with LLM s

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T20:28:59.570749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:28:59.570749Z digest=sha256:f655fd421577869e4e242f53a018a8f228b9050f624d4276d7b8ca7b28e2ad6b

Observation bc521204-141d-47ff-be90-5f34ec347f6e · outbound

This paper cites Capturing failures of large language models via human cognitive biases.

Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models Capturing failures of large language models via human cognitive biases

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:29:06.008527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-06T20:28:59.745163Z digest=sha256:a24574d75d0bca20e9cab4a2ee44d4bf58608d85e77cb6acbd63bb87a7061535

Observation 42e7bac0-e04d-4130-8fc7-375e883a226c · outbound

This paper cites Lin, and Lee Ross.

Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models Lin, and Lee Ross

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T20:28:59.906202Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:28:59.906202Z digest=sha256:7118ae5eae7e4d04fcec7b48e8c567be457c7c79ac042c45b4753c8fda67bd2c

Observation 2c893750-911c-4989-a3ab-7264afd41c4f · outbound

This paper cites ProcessBench: Identifying Process Errors in Mathematical Reasoning.

Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models ProcessBench: Identifying Process Errors in Mathematical Reasoning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:00.072960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:29:00.072960Z digest=sha256:62a0e10aebb91a07421bce7beda2fd76641cc3676af3b849a749e720aa6aac2f

Observation af15ffc1-4b55-4d7b-893f-55462028cbe3 · outbound

This paper cites PRMBench: A Fine-grained and Challenging Benchmark for Process-Level Reward Models.

Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models PRMBench: A Fine-grained and Challenging Benchmark for Process-Level Reward Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:00.154021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:29:00.154021Z digest=sha256:e3ad510380d54bb6b5fa2aa158739f77e541a7bf35f492b0a64523aff420599a

Observation 242c2d42-a346-4983-a5f9-e9fb052d06a4 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models Training Verifiers to Solve Math Word Problems

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:00.206673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:29:00.206673Z digest=sha256:2ebfd5a8cdfcca1c0a50bdfb519f52438cf7839af5bbdd3836986204d4c1a14c

Observation 0a682beb-9d8b-4ffc-8185-1c377d0db050 · outbound

This paper cites Introducing gpt-4.1 in the api, Apr 2025.

Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models Introducing gpt-4.1 in the api, Apr 2025

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:29:05.741704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-06T20:29:00.270483Z digest=sha256:147a613794b20dca6a52c38e0bc1135c0607e42c8c0099880d723ac730f699c3

Observation 547769ab-98dd-42f4-8e5d-f44fecc985c4 · outbound

This paper cites Let's verify step by step.

Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models Let's verify step by step

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:00.355018Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:29:00.355018Z digest=sha256:ba2b920ac7ccfca8e58029f3126d7760e135d42653595c6780f583d1058e95f2

Observation cdac7ea2-3ff3-405a-9392-f1aa43da387b · outbound

This paper cites Measuring mathematical problem solving with the MATH dataset.

Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models Measuring mathematical problem solving with the MATH dataset

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:00.438344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:29:00.438344Z digest=sha256:8335618ce10a57c8aed644c208fcd84adb433d6d91fc74e8557477493beb8a22

Observation d1a30e71-b262-458f-8efa-c396237e35b6 · outbound

This paper cites Transformers: State-of-the-art natural language processing.

Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models Transformers: State-of-the-art natural language processing

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:00.521903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:29:00.521903Z digest=sha256:8eb7498e00fb756dc86d44982264b9ac1c2775b28d0d0b309f515b73c116ea61

Observation 98372d2d-19f9-45e5-8238-be050bbd90ed · outbound

This paper cites DeepSeek-V3 Technical Report.

Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models DeepSeek-V3 Technical Report

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:00.633907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:29:00.633907Z digest=sha256:989f4f6060466c4e76af89539e4584e0e81b1d0c749ec94a6504e2fdf5ba6919

Observation 5e52510a-4855-4949-9e1d-a50fd817f0f4 · outbound

This paper cites Qwen2.5 Technical Report.

Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models Qwen2.5 Technical Report

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:00.711530Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:29:00.711530Z digest=sha256:5c875eab8350ecbe3d8f49bc9fce8312f1316cce9d88503ede90827a4db1eddd

Observation a5b12f90-c1f5-4320-988b-69bde06197c9 · outbound

This paper cites Llama 3.3, Dec 2024.

Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models Llama 3.3, Dec 2024

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:29:05.454473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-06T20:29:00.790776Z digest=sha256:f189d185263ad2b3989ce68697f893d07019ea20ef7e716bc964c1f044724ec4

Observation 070d8f16-3cc9-4caa-ba86-afcde95ccfc4 · outbound

This paper cites Phi-4 Technical Report.

Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models Phi-4 Technical Report

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:00.903257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:29:00.903257Z digest=sha256:509d61163ce63ef75116ef0cc758881d20544f735789aeefad8d5426e677e990

Observation d8898eac-0793-4381-a9b0-2df54c6e3a79 · outbound

This paper cites Qwen2 Technical Report.

Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models Qwen2 Technical Report

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:01.073515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:29:01.073515Z digest=sha256:6a1be75d1bc62e3d4f149f5ec5ab4988bf8290d716788226fb00e6ac75e19271

Observation 598889cd-4a56-43db-8cd4-7e4e3b4c55bc · outbound

This paper cites The Llama 3 Herd of Models.

Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models The Llama 3 Herd of Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:01.235737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:29:01.235737Z digest=sha256:6fda4892bd8de9ba247e8af4d82356364e5d942c19310ae5b5853a4e91ac9fd5

Observation 9c5dc13c-19b8-4c72-8e21-1ba13f577e88 · outbound

This paper cites Mistral small 3, Jan 2025.

Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models Mistral small 3, Jan 2025

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:29:05.122636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-06T20:29:01.386989Z digest=sha256:62ea3b3a48bc3389222e02ab3148b579f3a36d1cab854b842b36467ccf9ef97c

Observation 91a1453a-6c9b-4155-aebf-55fa2131cd9f · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:01.501958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:29:01.501958Z digest=sha256:1c7c7a9b721c5ad7bc55f9afb9326e6ca4e12d7dc2802327f9866cc10d5992f1

Observation 8a78e84b-9451-4307-9a7a-d78784925c50 · outbound

This paper cites Impact of pretraining term frequencies on few-shot numerical reasoning.

Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models Impact of pretraining term frequencies on few-shot numerical reasoning

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:01.660462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:29:01.660462Z digest=sha256:0c19f5759284514523f7dfae2ad91e2acf7b170819bdf0f2aceecd3c1a150620

Observation a8943ef9-1315-4cd0-8398-f439535fbb8e · outbound

This paper cites Smith, Sarah Wiegreffe, and Yanai Elazar.

Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models Smith, Sarah Wiegreffe, and Yanai Elazar

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:29:04.782194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-06T20:29:01.845245Z digest=sha256:e6f70a48ee54bde4b5191931523ab358ab144ecbbba9e6f2cafdff0e3c61e9e7

Observation e5ca97be-2dab-4115-a7b7-171ee449086e · outbound

This paper cites o pf, Yannic Kilcher, Dimitri von R \.

Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models o pf, Yannic Kilcher, Dimitri von R \

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:01.991836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:29:01.991836Z digest=sha256:b7cfb8927b1be299b4514a4ad3715be31ae92bc01515e3e5b1b7d77a417d6928

Observation 9472acb6-eb08-412e-a62b-0d21af73a4f9 · outbound

This paper cites Openhermes 2.5: An open dataset of synthetic data for generalist llm assistants, 2023.

Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models Openhermes 2.5: An open dataset of synthetic data for generalist llm assistants, 2023

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:02.140460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:29:02.140460Z digest=sha256:7a42abfea76d8dff4eb4ff3fb9dd70050c02d9c212e770b404c996a9d0889a04

Observation c317eade-a7e4-48ad-8654-5687149260d8 · outbound

This paper cites Infinity Instruct: Scaling Instruction Selection and Synthesis to Enhance Language Models.

Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models Infinity Instruct: Scaling Instruction Selection and Synthesis to Enhance Language Models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:02.244807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:29:02.244807Z digest=sha256:d383c404ebb435f5ee98c22f619a7f25be671a5324417f2e5478e6f05682d26a

Observation 61ed6acb-e56a-49e0-8e0a-49c8c23c40b5 · outbound

This paper cites Ultrafeedback: Boosting language models with high-quality feedback, 2024.

Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models Ultrafeedback: Boosting language models with high-quality feedback, 2024

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:29:04.501080Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-06T20:29:02.356141Z digest=sha256:ebba9e36c5bb838a0a5e593088271bea4d18e21de1c0ed236270da5ae43620ea

Observation d42f5cd2-00f6-4eeb-98c0-831a16acdeb7 · outbound

This paper cites Hwang, Jiangjiang Yang, Ronan Le Bras, Oyvind Tafjord, Christopher Wilhelm, Luca Soldaini, Noah A.

Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models Hwang, Jiangjiang Yang, Ronan Le Bras, Oyvind Tafjord, Christopher Wilhelm, Luca Soldaini, Noah A

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:02.487929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:29:02.487929Z digest=sha256:a14a41cafad77bb1f4bc084cb6c62fd61b040cb9b85189108394309df3b4036c

Observation fa7a74bf-6fcc-4198-83eb-b128ade148e9 · outbound

This paper cites Open r1: A fully open reproduction of deepseek-r1, January 2025.

Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models Open r1: A fully open reproduction of deepseek-r1, January 2025

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:02.599732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:29:02.599732Z digest=sha256:9620cdd9a90b08b69cb4a31aca26c00575150e3503678ea6df51a03c9bd3df23

Observation fab9b386-0834-49e7-9db5-dc727f9131ea · outbound

This paper cites OpenThoughts: Data Recipes for Reasoning Models.

Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models OpenThoughts: Data Recipes for Reasoning Models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:02.783626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:29:02.783626Z digest=sha256:ee1c481f1d55447238ad5604a0a5e5ea18210394c4aea1f7b82c21a70ded1ef9

Observation 432e34ca-6d12-4311-9df2-c0977733d550 · outbound

This paper cites Training language models to follow instructions with human feedback.

Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models Training language models to follow instructions with human feedback

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:02.938680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:29:02.938680Z digest=sha256:e86f7707028fc272e3d460eae1b58485bf07ec3a3f9c6e49a19d887069d31a91

Observation 5a580e71-de61-4a51-819e-81f6b955611e · outbound

This paper cites Learning From Mistakes Makes LLM Better Reasoner.

Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models Learning From Mistakes Makes LLM Better Reasoner

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:03.095630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:29:03.095630Z digest=sha256:1e960695d8e11e76a50e5ed83c22e8df5030b990a1815d9aaa15ab3ae5fbe18e

Observation f256587f-fd18-4918-be5b-1e797c5c666e · outbound

This paper cites Critique Fine-Tuning: Learning to Critique is More Effective than Learning to Imitate.

Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models Critique Fine-Tuning: Learning to Critique is More Effective than Learning to Imitate

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:03.216236Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:29:03.216236Z digest=sha256:f8ef9bd50ce22d1b6aa7d4550fcd7844c5ead00685af489bb4fd5c728b224ad6

Observation 4eb03d92-b6a5-4658-90b2-9dd56e085279 · outbound

This paper cites The effect of sampling temperature on problem solving in large language models.

Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models The effect of sampling temperature on problem solving in large language models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:03.409054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:29:03.409054Z digest=sha256:38164ab32f6958466407a258ad1c4572c0dcd6b8b4c3abd61a7c73dce41a6258

Observation d70d4393-6219-408a-8892-8d96da99a79c · outbound

This paper cites @esa (Ref.

Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models @esa (Ref

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:03.530733Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:29:03.530733Z digest=sha256:f091af353bdcbd7a9e7599c1e72010aa02748bdd03193ad43ce0957962165026

Observation a4614c9e-01f7-4332-adb6-eec84c16fca5 · outbound

This paper cites an unresolved cited work.

Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models Unresolved cited work

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:03.675873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:29:03.675873Z digest=sha256:76defd574751a224541f956fa6f301b38841f2e3eb4a9df69f8ce9ea95a871aa

Observation e8dda0b1-32fe-40b6-bebf-104c5fd66015 · outbound

This paper cites after incorrect reasoning or answer to prompt LLMs to self-correct, without finetuning. We observe significant reductions in the blind spot after appending ``Wait.

Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models after incorrect reasoning or answer to prompt LLMs to self-correct, without finetuning. We observe significant reductions in the blind spot after appending ``Wait

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:03.805548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:29:03.805548Z digest=sha256:aa0374c2fe16a386313696907900c2ae6521ea6d46c77f8ab7f9e84e0ec03426

Pith citing papers

Observation 3547e909-d01a-42bc-8fbe-10036c058fef · inbound

ReFlect: An Effective Harness System for Complex Long-Horizon LLM Reasoning cites this paper.

ReFlect: An Effective Harness System for Complex Long-Horizon LLM Reasoning Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models

Reference 28

Resolution
metadata mismatch
arxiv_id, observed 2026-08-04T02:29:36.375484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-08T11:35:25.204350Z digest=sha256:2dfe9052a1503929e731895e7147201f77cfb4b232efec2d84205fe797770a46

Observation 96d7ffc6-6dae-45a3-989d-37e0f37d0836 · inbound

ProCrit: Self-Elicited Multi-Perspective Reasoning with Critic-Guided Revision for Multimodal Sarcasm Detection cites this paper.

ProCrit: Self-Elicited Multi-Perspective Reasoning with Critic-Guided Revision for Multimodal Sarcasm Detection Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-08-04T02:29:36.375484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T02:35:37.724212Z digest=sha256:ead514e3edf8cb4782b8d5a06d1145113fc1f0c8a70868f9aec7b788fdc4f2d4

Observation 3a91d3a0-6f38-4726-acc8-06b923d8d159 · inbound

Mixture of Debaters: Learn to Debate at Architectural Level in Multi-Agent Reasoning cites this paper.

Mixture of Debaters: Learn to Debate at Architectural Level in Multi-Agent Reasoning Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-08-04T02:29:36.375484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-30T07:11:02.464556Z digest=sha256:4b67088fd0f82fbf346d550a60d9ad690df97e649bcf8e46028d6c7485b7eaf8

Observation eaa9921f-da88-4f3c-9369-e2f5a7f177a6 · inbound

ESC: Emotional Self-Correction for Reliable Vision-Language Models cites this paper.

ESC: Emotional Self-Correction for Reliable Vision-Language Models Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models

Reference 85

Resolution
metadata mismatch
arxiv_id, observed 2026-08-04T02:29:36.375484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-03T21:20:00.041277Z digest=sha256:8436638395cd6df2f195221d4f7d5b7416bdac3689406d48f8ff7e84a4c94357