Pith. sign in

Paper Citation Record · LEDGER

Can Large Language Models Really Improve by Self-critiquing Their Own Plans?

As of 23 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 18 inbound Pith citation observations for arXiv:2310.08118.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2310.08118 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 18 of 18 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 18 of 18 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T11:50:27.842026Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T11:05:42.217787Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 9eb97eb7-a5de-41db-8a88-c0eb63cba666 · inbound

One STEP at a time: Language Agents are Stepwise Planners cites this paper.

One STEP at a time: Language Agents are Stepwise Planners Can Large Language Models Really Improve by Self-critiquing Their Own Plans?

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-12T21:40:44.335667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T21:40:44.335667Z digest=sha256:50a335465581e0e937cab50ee70f562eb4a3c4909fa638527059d40bc95a85c8

Observation 84114d30-47a4-4217-bee0-c88cde2b66d3 · inbound

Marco-o1: Towards Open Reasoning Models for Open-Ended Solutions cites this paper.

Marco-o1: Towards Open Reasoning Models for Open-Ended Solutions Can Large Language Models Really Improve by Self-critiquing Their Own Plans?

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T15:17:55.128453Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T15:17:55.128453Z digest=sha256:9e17658ca33b39dddcbdd846f3c4030e0910eb946f00e719fe25f92fe980f86e

Observation 0eaa7b4b-6f74-45f9-a1be-e544b5295761 · inbound

LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods cites this paper.

LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods Can Large Language Models Really Improve by Self-critiquing Their Own Plans?

Reference 232

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:08:35.544552Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-11T23:08:34.312466Z digest=sha256:c1e0fdbdfe875fb9aa97a7d54c025c83e6a75f047dbd1d5269c2b5e93c34c4f2

Observation 77224a3e-b853-4f40-8b59-9c60c20210fc · inbound

EvoWiki: Evaluating LLMs on Evolving Knowledge cites this paper.

EvoWiki: Evaluating LLMs on Evolving Knowledge Can Large Language Models Really Improve by Self-critiquing Their Own Plans?

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T13:04:20.209828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:04:20.209828Z digest=sha256:41aaa646c0739c560af0870cda85a5a139f70b2122c41d113a616b1e74ebc9bd

Observation 3cc974f2-a2c4-4171-8590-00b868f03bd0 · inbound

Rethinking Mixture-of-Agents: Is Mixing Different Large Language Models Beneficial? cites this paper.

Rethinking Mixture-of-Agents: Is Mixing Different Large Language Models Beneficial? Can Large Language Models Really Improve by Self-critiquing Their Own Plans?

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-09T18:11:28.174002Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T18:11:28.174002Z digest=sha256:f16cbf3353ea4f9ff431b8e8faa911347eb851dd1aff5c40416c42d6dc73efcf

Observation 1d07930e-325d-4ba0-a1bd-68fc220040b9 · inbound

S$^2$-MAD: Breaking the Token Barrier to Enhance Multi-Agent Debate Efficiency cites this paper.

S$^2$-MAD: Breaking the Token Barrier to Enhance Multi-Agent Debate Efficiency Can Large Language Models Really Improve by Self-critiquing Their Own Plans?

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-08T21:33:37.640825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T21:33:37.640825Z digest=sha256:b044e4fbb6b085dece0886e0b5ccddda7fd362ec7bd60a7d126642003d9407ba

Observation 5e1110ee-275e-4c93-8f5c-966c12c56193 · inbound

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey cites this paper.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Can Large Language Models Really Improve by Self-critiquing Their Own Plans?

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:27.842026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:27.842026Z digest=sha256:72f120894adf0313508429e9fbcdd95be2d36b7b3be7505440d19fc03ab0b0b1

Observation 954c0c8f-311b-4be5-9824-a8c4696ce09c · inbound

Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators cites this paper.

Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators Can Large Language Models Really Improve by Self-critiquing Their Own Plans?

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-16T11:33:53.540998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:33:53.540998Z digest=sha256:f85523d2c59693355f6e4990b469528751955dca07b18573fe76ba9ae6eff855

Observation e4360466-1e79-4357-9803-c21c70600514 · inbound

Exchange of Perspective Prompting Enhances Reasoning in Large Language Models cites this paper.

Exchange of Perspective Prompting Enhances Reasoning in Large Language Models Can Large Language Models Really Improve by Self-critiquing Their Own Plans?

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:00.034064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:04:00.034064Z digest=sha256:7b88d97ef25a89f05de8879adcbaeb1c344f06ea5eb154a590ed1cdfb595d334

Observation 1839f655-5ea1-40b2-87e4-919933da9d28 · inbound

Twill: Scheduling Compound AI Systems on Heterogeneous Mobile Edge Platforms cites this paper.

Twill: Scheduling Compound AI Systems on Heterogeneous Mobile Edge Platforms Can Large Language Models Really Improve by Self-critiquing Their Own Plans?

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T21:20:29.395654Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:20:29.395654Z digest=sha256:81787c9682dba471e3492ca949648658dbffa9a38239ac6d5014d3ac97cbfa06

Observation fc1db1c5-5c14-49db-8e7c-7a23c02e4b59 · inbound

Deep sequence models tend to memorize geometrically; it is unclear why cites this paper.

Deep sequence models tend to memorize geometrically; it is unclear why Can Large Language Models Really Improve by Self-critiquing Their Own Plans?

Reference 174

Resolution
verified exact
arxiv_id, observed 2026-05-21T20:40:36.287151Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-21T20:38:18.005002Z digest=sha256:e1095bc98cbcbcf2297878bac040916cef736c680ec764ae8c3ad727fae3cba0

Observation 62d2848f-8c45-4134-a8c1-ca9a5060a426 · inbound

End-to-end PDDL Planning with Hardcoded and Dynamic Agents cites this paper.

End-to-end PDDL Planning with Hardcoded and Dynamic Agents Can Large Language Models Really Improve by Self-critiquing Their Own Plans?

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-16T23:38:41.621289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T23:34:20.411940Z digest=sha256:378c489f0e976fb1348efef8196aeed6b1d1f6022f24f433851eb87565f1654d

Observation 5218f3c7-8cc0-4bda-93fa-e2492323df5e · inbound

Stop the Flip-Flop: Context-Preserving Verification for Fast Revocable Diffusion Decoding cites this paper.

Stop the Flip-Flop: Context-Preserving Verification for Fast Revocable Diffusion Decoding Can Large Language Models Really Improve by Self-critiquing Their Own Plans?

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-03T04:06:11.908439Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:06:11.908439Z digest=sha256:648a14066335362cf171c5a113e119063af725d0a883207a4d23e6ee3b74d040

Observation 5258a308-80a1-47b3-9b9d-f41dc493e3a2 · inbound

OPT-BENCH: Evaluating the Iterative Self-Optimization of LLM Agents in Large-Scale Search Spaces cites this paper.

OPT-BENCH: Evaluating the Iterative Self-Optimization of LLM Agents in Large-Scale Search Spaces Can Large Language Models Really Improve by Self-critiquing Their Own Plans?

Reference 48

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T03:01:18.655422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-12T02:57:15.521594Z digest=sha256:6ce0bc6c74537a03b7661ff2f8d0ff0859349081e84317193862e67d7b1e440a

Observation 8ae7bdbd-a4bb-46d7-9beb-345465abd922 · inbound

Teaching Large Language Models When Not to Know: Learning Temporal Critique for Ex-Ante Reasoning cites this paper.

Teaching Large Language Models When Not to Know: Learning Temporal Critique for Ex-Ante Reasoning Can Large Language Models Really Improve by Self-critiquing Their Own Plans?

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-06-30T20:55:04.650977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-30T20:46:31.368628Z digest=sha256:236f7f67fabe601421ff7b44c0ab69023f75672a2d42ac6730c3e74cec9cb23c

Observation 67c87912-15ca-415d-864e-66c22e0171fe · inbound

Roll Out and Roll Back: Diffusion LLMs are Their Own Efficiency Teachers cites this paper.

Roll Out and Roll Back: Diffusion LLMs are Their Own Efficiency Teachers Can Large Language Models Really Improve by Self-critiquing Their Own Plans?

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-19T20:42:46.127304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-19T20:41:48.063871Z digest=sha256:890275ed5b9032b8796fb469217e8ed314bef5c474d4df44a288dba34181bd7b

Observation b6a978cc-5345-4543-bba1-919efe96c1e3 · inbound

Falsification, Not Exposure: An Internally Preregistered Placebo-Controlled Decomposition of Self-Repair Feedback in Frozen Small Code Models cites this paper.

Falsification, Not Exposure: An Internally Preregistered Placebo-Controlled Decomposition of Self-Repair Feedback in Frozen Small Code Models Can Large Language Models Really Improve by Self-critiquing Their Own Plans?

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-07-01T11:05:42.219249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-07-01T04:44:56.520156Z digest=sha256:2609f58e48522fd555b0f1f2200254e666c9d30ae07d5d2971e4c300048b27f9

Observation aa0e2dae-a3e0-4897-ad3b-cf91d50542d9 · inbound

Test-Time Augmentation for LLMs: When Input Diversity Beats Output Diversity at Matched Compute cites this paper.

Test-Time Augmentation for LLMs: When Input Diversity Beats Output Diversity at Matched Compute Can Large Language Models Really Improve by Self-critiquing Their Own Plans?

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T19:01:29.612699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:01:29.612699Z digest=sha256:94aa6d5549e5d6a244af3f1d8c53a959d9a28e0bdfa0b257ce82cda670120eef