Pith. sign in

Paper Citation Record · LEDGER

Can Large Language Models Really Improve by Self-critiquing Their Own Plans?

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 10 inbound Pith citation observations for arXiv:2310.08118.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2310.08118 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 10 of 10 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 10 of 10 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:04:00.034064Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T11:05:42.217787Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 0eaa7b4b-6f74-45f9-a1be-e544b5295761 · inbound

LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods cites this paper.

LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods Can Large Language Models Really Improve by Self-critiquing Their Own Plans?

Reference 232

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:08:35.544552Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T23:08:34.312466Z digest=sha256:f1ef03480e590c8af30cae3b2736d8173d59bfb3649f3c2a3d640e01abb2e403

Observation e4360466-1e79-4357-9803-c21c70600514 · inbound

Exchange of Perspective Prompting Enhances Reasoning in Large Language Models cites this paper.

Exchange of Perspective Prompting Enhances Reasoning in Large Language Models Can Large Language Models Really Improve by Self-critiquing Their Own Plans?

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:00.034064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:04:00.034064Z digest=sha256:48ed83623bebac49fb47c8c3ea458e6589880a5d69532db5e69c7903a9d423ec

Observation 1839f655-5ea1-40b2-87e4-919933da9d28 · inbound

Twill: Scheduling Compound AI Systems on Heterogeneous Mobile Edge Platforms cites this paper.

Twill: Scheduling Compound AI Systems on Heterogeneous Mobile Edge Platforms Can Large Language Models Really Improve by Self-critiquing Their Own Plans?

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T21:20:29.395654Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:20:29.395654Z digest=sha256:fcb9939621751ebb0de34929ad4124e01f6ac0608c62eda14fdf717866633089

Observation fc1db1c5-5c14-49db-8e7c-7a23c02e4b59 · inbound

Deep sequence models tend to memorize geometrically; it is unclear why cites this paper.

Deep sequence models tend to memorize geometrically; it is unclear why Can Large Language Models Really Improve by Self-critiquing Their Own Plans?

Reference 174

Resolution
verified exact
arxiv_id, observed 2026-05-21T20:40:36.287151Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-21T20:38:18.005002Z digest=sha256:e64a5c9954ba90e6399297fdceb43518074240ff3bdc570ff0c064e024064a02

Observation 62d2848f-8c45-4134-a8c1-ca9a5060a426 · inbound

End-to-end PDDL Planning with Hardcoded and Dynamic Agents cites this paper.

End-to-end PDDL Planning with Hardcoded and Dynamic Agents Can Large Language Models Really Improve by Self-critiquing Their Own Plans?

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-16T23:38:41.621289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-16T23:34:20.411940Z digest=sha256:bb32c12c41e76b365fda3ece05284384de1695513924c7b11e823f954e85a5f2

Observation 5218f3c7-8cc0-4bda-93fa-e2492323df5e · inbound

Stop the Flip-Flop: Context-Preserving Verification for Fast Revocable Diffusion Decoding cites this paper.

Stop the Flip-Flop: Context-Preserving Verification for Fast Revocable Diffusion Decoding Can Large Language Models Really Improve by Self-critiquing Their Own Plans?

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-03T04:06:11.908439Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:06:11.908439Z digest=sha256:fe493ca35dd750e68c620289b4274832bfe41c46d51b166d89608cf996450ad0

Observation 5258a308-80a1-47b3-9b9d-f41dc493e3a2 · inbound

OPT-BENCH: Evaluating the Iterative Self-Optimization of LLM Agents in Large-Scale Search Spaces cites this paper.

OPT-BENCH: Evaluating the Iterative Self-Optimization of LLM Agents in Large-Scale Search Spaces Can Large Language Models Really Improve by Self-critiquing Their Own Plans?

Reference 48

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T03:01:18.655422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-12T02:57:15.521594Z digest=sha256:ec84512d2a4098cdaf6dbc305aa087633de32a820703acbf7f6f9820b1f3da76

Observation 8ae7bdbd-a4bb-46d7-9beb-345465abd922 · inbound

Teaching Large Language Models When Not to Know: Learning Temporal Critique for Ex-Ante Reasoning cites this paper.

Teaching Large Language Models When Not to Know: Learning Temporal Critique for Ex-Ante Reasoning Can Large Language Models Really Improve by Self-critiquing Their Own Plans?

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-06-30T20:55:04.650977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T20:46:31.368628Z digest=sha256:3bf877b63c9a7465086d719ebf776f258e3e7040a5b292f9bea0b1451397955d

Observation 67c87912-15ca-415d-864e-66c22e0171fe · inbound

Roll Out and Roll Back: Diffusion LLMs are Their Own Efficiency Teachers cites this paper.

Roll Out and Roll Back: Diffusion LLMs are Their Own Efficiency Teachers Can Large Language Models Really Improve by Self-critiquing Their Own Plans?

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-19T20:42:46.127304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-19T20:41:48.063871Z digest=sha256:f239486d590c047f1eaff7d1738c2d7da254cd1980fd0ee6b6ed3f6cfb3a3387

Observation b6a978cc-5345-4543-bba1-919efe96c1e3 · inbound

Falsification, Not Exposure: An Internally Preregistered Placebo-Controlled Decomposition of Self-Repair Feedback in Frozen Small Code Models cites this paper.

Falsification, Not Exposure: An Internally Preregistered Placebo-Controlled Decomposition of Self-Repair Feedback in Frozen Small Code Models Can Large Language Models Really Improve by Self-critiquing Their Own Plans?

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-07-01T11:05:42.219249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-07-01T04:44:56.520156Z digest=sha256:f21bd807417f7229d35932f44ba3293f62ae40beaf371558e49a3073f6b9336d