Pith. sign in

Paper Citation Record · LEDGER

GPT-4 Doesn't Know It's Wrong: An Analysis of Iterative Prompting for Reasoning Problems

As of 20 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 14 inbound Pith citation observations for arXiv:2310.12397.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2310.12397 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 14 of 14 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 14 of 14 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T11:33:53.531866Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

11
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 9551cc34-97a7-407a-a49b-39a58d7ec3ec · inbound

VeCoGen: Automating Generation of Formally Verified C Code with Large Language Models cites this paper.

VeCoGen: Automating Generation of Formally Verified C Code with Large Language Models GPT-4 Doesn't Know It's Wrong: An Analysis of Iterative Prompting for Reasoning Problems

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T10:24:29.727385Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:24:29.727385Z digest=sha256:5fe9fdea27afaef7fe9c7e0d9f6a39378b1856ab66424d3594ada6cab540f08b

Observation e7e492e3-3c34-4abf-aa8f-9aedd1c7967e · inbound

Optimization is Better than Generation: Optimizing Commit Message Leveraging Human-written Commit Message cites this paper.

Optimization is Better than Generation: Optimizing Commit Message Leveraging Human-written Commit Message GPT-4 Doesn't Know It's Wrong: An Analysis of Iterative Prompting for Reasoning Problems

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-10T19:41:51.210974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:41:51.210974Z digest=sha256:ab86eddd4e69cd2bd04659adbe7851a65a9d838b75f312b202d7d75568442040

Observation 2af25100-fe67-44ff-bece-e9f9a4e841fe · inbound

S$^2$-MAD: Breaking the Token Barrier to Enhance Multi-Agent Debate Efficiency cites this paper.

S$^2$-MAD: Breaking the Token Barrier to Enhance Multi-Agent Debate Efficiency GPT-4 Doesn't Know It's Wrong: An Analysis of Iterative Prompting for Reasoning Problems

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-08T21:33:37.395529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T21:33:37.395529Z digest=sha256:501b848021f273f438fbe387f0b9a41bc0f6cc88debb6e16a55453c8d24ce651

Observation 5d064cd7-dda9-4876-b173-ef880458fcdb · inbound

Training Language Models for Social Deduction with Multi-Agent Reinforcement Learning cites this paper.

Training Language Models for Social Deduction with Multi-Agent Reinforcement Learning GPT-4 Doesn't Know It's Wrong: An Analysis of Iterative Prompting for Reasoning Problems

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-08T16:57:01.690429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:57:01.690429Z digest=sha256:c9fb0ae9652bdcf810f5b046a2f3fe3585458c00f506977264eb7ea246419006

Observation 536088e7-b448-45e6-9de9-e38a6752366d · inbound

Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators cites this paper.

Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators GPT-4 Doesn't Know It's Wrong: An Analysis of Iterative Prompting for Reasoning Problems

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-16T11:33:53.531866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:33:53.531866Z digest=sha256:ca36fb2e61f38db33bf5542e609c707fb381c5e04291fe89eac4707d698d4a69

Observation c30d184e-8fd5-44b6-89be-c49c8358f46d · inbound

Veracity Bias and Beyond: Uncovering LLMs' Hidden Beliefs in Problem-Solving Reasoning cites this paper.

Veracity Bias and Beyond: Uncovering LLMs' Hidden Beliefs in Problem-Solving Reasoning GPT-4 Doesn't Know It's Wrong: An Analysis of Iterative Prompting for Reasoning Problems

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T15:12:13.428065Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:12:13.428065Z digest=sha256:a447e171b698bf47770c202bf8c6d3ef189d72902b609162051383cb86a6c1e1

Observation 0d1cd1f9-3767-42b5-80d3-572571960587 · inbound

Exchange of Perspective Prompting Enhances Reasoning in Large Language Models cites this paper.

Exchange of Perspective Prompting Enhances Reasoning in Large Language Models GPT-4 Doesn't Know It's Wrong: An Analysis of Iterative Prompting for Reasoning Problems

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:00.039133Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:04:00.039133Z digest=sha256:9ea251005bd8aa8627f6afe72499504132945a6d69776972047b2553b0b26ecb

Observation e490b109-c17f-4cd5-a8c2-c1fc2c6ca889 · inbound

It's Not That Simple. An Analysis of Simple Test-Time Scaling cites this paper.

It's Not That Simple. An Analysis of Simple Test-Time Scaling GPT-4 Doesn't Know It's Wrong: An Analysis of Iterative Prompting for Reasoning Problems

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T16:09:05.621221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:09:05.621221Z digest=sha256:c3157599881ba278ce686d6fab8cab8fb5576cdac172b647b02b33e56ea5baaf

Observation c2544d17-6930-4b8d-88d5-f83cfb1cdac9 · inbound

Lightweight Language Models are Prone to Reasoning Errors for Complex Computational Phenotyping Tasks cites this paper.

Lightweight Language Models are Prone to Reasoning Errors for Complex Computational Phenotyping Tasks GPT-4 Doesn't Know It's Wrong: An Analysis of Iterative Prompting for Reasoning Problems

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T11:06:07.559969Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:06:07.559969Z digest=sha256:9aa2c463ad94e1f2d55bae4a02ee8ef5799b19943de66b78b2a599565d90ef75

Observation 56886bbe-0b59-4a90-8aae-f3d38fa406a8 · inbound

OPT-BENCH: Evaluating the Iterative Self-Optimization of LLM Agents in Large-Scale Search Spaces cites this paper.

OPT-BENCH: Evaluating the Iterative Self-Optimization of LLM Agents in Large-Scale Search Spaces GPT-4 Doesn't Know It's Wrong: An Analysis of Iterative Prompting for Reasoning Problems

Reference 51

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T03:01:18.605389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-12T02:57:15.521594Z digest=sha256:cfd38422874d5fa6cd0dc1b724098bf6d0b15443f4a3631b610505d094f34f31

Observation cf2eba2c-f2d4-4fdf-897d-fe6536776d99 · inbound

Roll Out and Roll Back: Diffusion LLMs are Their Own Efficiency Teachers cites this paper.

Roll Out and Roll Back: Diffusion LLMs are Their Own Efficiency Teachers GPT-4 Doesn't Know It's Wrong: An Analysis of Iterative Prompting for Reasoning Problems

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-19T20:42:46.111972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-19T20:41:48.063871Z digest=sha256:af2217fb80d175094881209647a56055d638d9beab18830c2406a7e967c9e846

Observation 40139c4b-32fe-4dde-99d6-4ca72fa8a5ae · inbound

The Sword, Shield, and Achilles' Heel: Characterizing the Linguistic Inductive Bias of Large Language Models for Spatial Reasoning in Navigation Planning cites this paper.

The Sword, Shield, and Achilles' Heel: Characterizing the Linguistic Inductive Bias of Large Language Models for Spatial Reasoning in Navigation Planning GPT-4 Doesn't Know It's Wrong: An Analysis of Iterative Prompting for Reasoning Problems

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-07-01T19:36:08.375155Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-28T22:23:24.536258Z digest=sha256:bb9e7f0337425caa1e17623bedbd97758140958924795d8045f64e11b402307c

Observation 06a3d703-496c-4a31-aaef-27d37cbbcfb3 · inbound

The Self-Correction Illusion: Role Relabeling Gates Explicit Error Flagging in Large Language Models cites this paper.

The Self-Correction Illusion: Role Relabeling Gates Explicit Error Flagging in Large Language Models GPT-4 Doesn't Know It's Wrong: An Analysis of Iterative Prompting for Reasoning Problems

Reference 34

Resolution
metadata mismatch
arxiv_id, observed 2026-06-28T01:31:29.251945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-06-28T01:25:07.890796Z digest=sha256:8b96be4d050b9bac687e474702b9a6e69deb28cf261b7728e435d89d2f13fc79

Observation 3384582e-a352-49d0-a3ef-65f30119565d · inbound

DICE: Entropy-Regularized Equilibrium Selection for Stable Multi-Agent LLM Coordination cites this paper.

DICE: Entropy-Regularized Equilibrium Selection for Stable Multi-Agent LLM Coordination GPT-4 Doesn't Know It's Wrong: An Analysis of Iterative Prompting for Reasoning Problems

Reference 87

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T21:07:23.497033Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-06-27T19:58:32.016341Z digest=sha256:f3ecfcce5013cc1871981b1af18759bf38865f7a60679d81932e87bec5d0d513