Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T20:29:03.805548Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 60 of 60 outbound references and 4 inbound Pith citation observations for arXiv:2507.02778.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T20:29:03.805548Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-07-03T21:20:00.041277Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-03T21:28:58.386085Z
60 of 60 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 19a388be-1abf-4a33-ba82-7b35e40fd82f · outbound
Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models GPT-4 Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 73f0cd9a-0b1b-408e-8072-9d94b04998e2 · outbound
Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models The claude 3 model family: Opus, sonnet, haiku, Mar 2024
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 50bbfc04-ce8c-4b4b-b5b0-3fff3e8b3f5a · outbound
Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models Gemini 2.5: Pushing the frontier with advanced reasoning, multimodality, long context, and next generation agentic capabilities., June 2025
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ae1989c8-73a6-49a0-8fc4-2d16699ff84d · outbound
Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models Qwen3 Technical Report
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 586457e5-ab3b-4b4a-9464-8e20506298a9 · outbound
Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models The llama 4 herd: The beginning of a new era of natively multimodal ai innovation, Apr 2025
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 91ab01f4-557c-433b-b672-94dbef010bc0 · outbound
Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f38e80ac-bd0f-42de-b022-b8b7c6500e47 · outbound
Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models On faithfulness and factuality in abstractive summarization
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 77338fbe-e90f-4bc5-bba8-76b029fc8ea4 · outbound
Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models A survey on hallucination in large language models: Principles, taxonomy, challenges, and open questions
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 098af2f9-edf9-49c1-a735-ebf863240b3d · outbound
Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models Do, Yan Xu, and Pascale Fung
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation caa6fbd6-6cbc-4f13-b88a-1bcb810818e9 · outbound
Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models Large language models can be easily distracted by irrelevant context
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d87bcd25-bf69-48cd-b7fa-c14ce7548102 · outbound
Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models Alice in Wonderland: Simple Tasks Showing Complete Reasoning Breakdown in State-Of-the-Art Large Language Models
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fae702c4-d4da-4075-8c5e-a0238f612dfe · outbound
Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models Reflexion: language agents with verbal reinforcement learning
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1cecc2d2-7db8-48a5-94cf-cff9ceed325b · outbound
Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models Self-refine: Iterative refinement with self-feedback
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0952229d-8e2e-4192-9c73-59e1a4b6dcff · outbound
Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models Language models can solve computer tasks
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ee98ddf2-8acf-4313-9d25-4c6955fe54ee · outbound
Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models When can LLM s actually correct their own mistakes? a critical survey of self-correction of LLM s
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7cb640df-e27c-4882-80c9-919eeba976dc · outbound
Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models Large Language Models Cannot Self-Correct Reasoning Yet
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8923266a-b790-4707-a354-bce566c6368d · outbound
Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models LLM s cannot find reasoning errors, but can correct them given the error location
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ed5bf506-a601-41b5-a800-544053498431 · outbound
Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models Evaluating LLM s at detecting errors in LLM responses
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5f1f55b3-3870-468a-9f67-1af9ce35a562 · outbound
Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models Training language models to self-correct via reinforcement learning
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation efddf2d7-1f69-4927-bb0a-f4f3896e609c · outbound
Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models Jailbroken: how does llm safety training fail? In Proceedings of the 37th International Conference on Neural Information Processing Systems, NIPS '23, Red Hook, NY, USA, 2023
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 553309ad-0b0f-4417-b8b1-f864b3625a67 · outbound
Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models Formalizing and benchmarking prompt injection attacks and defenses
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7caeb070-8c3f-4b21-a32e-ab3575a5ea7d · outbound
Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models Measuring Faithfulness in Chain-of-Thought Reasoning
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9aced9d1-1d08-48f8-87f4-1e5b8e6a1a5b · outbound
Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models Unresolved cited work
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6e96adf5-08a3-472a-8d93-9315cfd851d5 · outbound
Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models Scaling LLM test-time compute optimally can be more effective than scaling parameters for reasoning
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a61249e0-c3ef-455b-b871-6c8f7c79c70c · outbound
Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models s1: Simple test-time scaling
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f0954cbd-2c14-4b6b-8801-899f2a583ed1 · outbound
Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models Benchmarking cognitive biases in large language models as evaluators
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 78c11968-cc45-47a7-b086-2d93729aa853 · outbound
Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models Cognitive bias in decision-making with LLM s
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bc521204-141d-47ff-be90-5f34ec347f6e · outbound
Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models Capturing failures of large language models via human cognitive biases
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 42e7bac0-e04d-4130-8fc7-375e883a226c · outbound
Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models Lin, and Lee Ross
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2c893750-911c-4989-a3ab-7264afd41c4f · outbound
Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models ProcessBench: Identifying Process Errors in Mathematical Reasoning
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation af15ffc1-4b55-4d7b-893f-55462028cbe3 · outbound
Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models PRMBench: A Fine-grained and Challenging Benchmark for Process-Level Reward Models
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 242c2d42-a346-4983-a5f9-e9fb052d06a4 · outbound
Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models Training Verifiers to Solve Math Word Problems
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0a682beb-9d8b-4ffc-8185-1c377d0db050 · outbound
Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models Introducing gpt-4.1 in the api, Apr 2025
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 547769ab-98dd-42f4-8e5d-f44fecc985c4 · outbound
Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models Let's verify step by step
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cdac7ea2-3ff3-405a-9392-f1aa43da387b · outbound
Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models Measuring mathematical problem solving with the MATH dataset
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d1a30e71-b262-458f-8efa-c396237e35b6 · outbound
Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models Transformers: State-of-the-art natural language processing
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 98372d2d-19f9-45e5-8238-be050bbd90ed · outbound
Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models DeepSeek-V3 Technical Report
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5e52510a-4855-4949-9e1d-a50fd817f0f4 · outbound
Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models Qwen2.5 Technical Report
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a5b12f90-c1f5-4320-988b-69bde06197c9 · outbound
Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models Llama 3.3, Dec 2024
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 070d8f16-3cc9-4caa-ba86-afcde95ccfc4 · outbound
Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models Phi-4 Technical Report
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d8898eac-0793-4381-a9b0-2df54c6e3a79 · outbound
Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models Qwen2 Technical Report
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 598889cd-4a56-43db-8cd4-7e4e3b4c55bc · outbound
Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models The Llama 3 Herd of Models
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9c5dc13c-19b8-4c72-8e21-1ba13f577e88 · outbound
Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models Mistral small 3, Jan 2025
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 91a1453a-6c9b-4155-aebf-55fa2131cd9f · outbound
Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8a78e84b-9451-4307-9a7a-d78784925c50 · outbound
Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models Impact of pretraining term frequencies on few-shot numerical reasoning
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a8943ef9-1315-4cd0-8398-f439535fbb8e · outbound
Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models Smith, Sarah Wiegreffe, and Yanai Elazar
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e5ca97be-2dab-4115-a7b7-171ee449086e · outbound
Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models o pf, Yannic Kilcher, Dimitri von R \
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9472acb6-eb08-412e-a62b-0d21af73a4f9 · outbound
Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models Openhermes 2.5: An open dataset of synthetic data for generalist llm assistants, 2023
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c317eade-a7e4-48ad-8654-5687149260d8 · outbound
Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models Infinity Instruct: Scaling Instruction Selection and Synthesis to Enhance Language Models
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 61ed6acb-e56a-49e0-8e0a-49c8c23c40b5 · outbound
Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models Ultrafeedback: Boosting language models with high-quality feedback, 2024
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d42f5cd2-00f6-4eeb-98c0-831a16acdeb7 · outbound
Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models Hwang, Jiangjiang Yang, Ronan Le Bras, Oyvind Tafjord, Christopher Wilhelm, Luca Soldaini, Noah A
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fa7a74bf-6fcc-4198-83eb-b128ade148e9 · outbound
Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models Open r1: A fully open reproduction of deepseek-r1, January 2025
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fab9b386-0834-49e7-9db5-dc727f9131ea · outbound
Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models OpenThoughts: Data Recipes for Reasoning Models
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 432e34ca-6d12-4311-9df2-c0977733d550 · outbound
Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models Training language models to follow instructions with human feedback
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5a580e71-de61-4a51-819e-81f6b955611e · outbound
Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models Learning From Mistakes Makes LLM Better Reasoner
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f256587f-fd18-4918-be5b-1e797c5c666e · outbound
Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models Critique Fine-Tuning: Learning to Critique is More Effective than Learning to Imitate
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4eb03d92-b6a5-4658-90b2-9dd56e085279 · outbound
Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models The effect of sampling temperature on problem solving in large language models
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d70d4393-6219-408a-8892-8d96da99a79c · outbound
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a4614c9e-01f7-4332-adb6-eec84c16fca5 · outbound
Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models Unresolved cited work
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e8dda0b1-32fe-40b6-bebf-104c5fd66015 · outbound
Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models after incorrect reasoning or answer to prompt LLMs to self-correct, without finetuning. We observe significant reductions in the blind spot after appending ``Wait
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3547e909-d01a-42bc-8fbe-10036c058fef · inbound
ReFlect: An Effective Harness System for Complex Long-Horizon LLM Reasoning Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 96d7ffc6-6dae-45a3-989d-37e0f37d0836 · inbound
ProCrit: Self-Elicited Multi-Perspective Reasoning with Critic-Guided Revision for Multimodal Sarcasm Detection Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3a91d3a0-6f38-4726-acc8-06b923d8d159 · inbound
Mixture of Debaters: Learn to Debate at Architectural Level in Multi-Agent Reasoning Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation eaa9921f-da88-4f3c-9369-e2f5a7f177a6 · inbound
ESC: Emotional Self-Correction for Reliable Vision-Language Models Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models
Reference 85
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.