Pith. sign in

Paper Citation Record · LEDGER

PlanBench: An Extensible Benchmark for Evaluating Large Language Models on Planning and Reasoning about Change

As of 19 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 28 inbound Pith citation observations for arXiv:2206.10498.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2206.10498 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 28 of 28 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 28 of 28 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T10:12:00.512507Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

50
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 46c28c21-9af8-4463-b3f1-f41756e9b448 · inbound

LLM+P: Empowering Large Language Models with Optimal Planning Proficiency cites this paper.

LLM+P: Empowering Large Language Models with Optimal Planning Proficiency PlanBench: An Extensible Benchmark for Evaluating Large Language Models on Planning and Reasoning about Change

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-14T18:36:18.607331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-14T18:36:18.342052Z digest=sha256:f65781596378a9d16f8701617b9995715ea55121ac137ebf754a19b5df9993e7

Observation bbaa211d-562c-4d18-983e-97e9def07777 · inbound

Distilling Step-by-Step! Outperforming Larger Language Models with Less Training Data and Smaller Model Sizes cites this paper.

Distilling Step-by-Step! Outperforming Larger Language Models with Less Training Data and Smaller Model Sizes PlanBench: An Extensible Benchmark for Evaluating Large Language Models on Planning and Reasoning about Change

Reference 103

Resolution
verified exact
arxiv_id, observed 2026-05-21T20:50:09.482935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-21T20:50:09.265838Z digest=sha256:ee6bfddbbcd76c28ae934f3455099bb96ab424328d5366e3b64cf7b23daf61cd

Observation 2c3f695a-f259-42a7-aab1-b7dac58cbd74 · inbound

Reasoning with Language Model is Planning with World Model cites this paper.

Reasoning with Language Model is Planning with World Model PlanBench: An Extensible Benchmark for Evaluating Large Language Models on Planning and Reasoning about Change

Reference 73

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T01:49:28.870191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-17T01:49:28.796581Z digest=sha256:5b0e32b0a9e4f465edc62f1ff3ed5431238f06d2ecf9c7835d61e3d3ef0ce714

Observation a17288dc-56c7-4a4c-98f2-39360714a03e · inbound

Cognitive Architectures for Language Agents cites this paper.

Cognitive Architectures for Language Agents PlanBench: An Extensible Benchmark for Evaluating Large Language Models on Planning and Reasoning about Change

Reference 80

Resolution
verified exact
arxiv_id, observed 2026-05-16T19:33:44.195442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-16T19:33:44.146134Z digest=sha256:2d33b9e4825b3f327552db0d500e5a59a27b4c90a947793a1b4fa670c109e166

Observation 458aa1fb-88d4-41bd-ba59-2b94e36d04d1 · inbound

CodeMind: Evaluating Large Language Models for Code Reasoning cites this paper.

CodeMind: Evaluating Large Language Models for Code Reasoning PlanBench: An Extensible Benchmark for Evaluating Large Language Models on Planning and Reasoning about Change

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-24T03:55:59.764519Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-24T03:53:55.964755Z digest=sha256:4f0686ea592e2251816d0492f251e5be2c5e6a9fe449f304b7cd5689aa79bb67

Observation 72e0e72d-42d8-4193-93b0-f6eec1c4488e · inbound

GSM-Symbolic: Understanding the Limitations of Mathematical Reasoning in Large Language Models cites this paper.

GSM-Symbolic: Understanding the Limitations of Mathematical Reasoning in Large Language Models PlanBench: An Extensible Benchmark for Evaluating Large Language Models on Planning and Reasoning about Change

Reference 96

Resolution
verified exact
arxiv_id, observed 2026-05-15T00:42:11.987226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-15T00:42:11.891829Z digest=sha256:81eee572da158ba3e2ee1d8adad4b3b3acabbae01a057eebacab2bb83fcf0dc3

Observation 3e779026-5705-4a77-815d-9cbe0a9710eb · inbound

A Review on Generative AI Models for Synthetic Medical Text, Time Series, and Longitudinal Data cites this paper.

A Review on Generative AI Models for Synthetic Medical Text, Time Series, and Longitudinal Data PlanBench: An Extensible Benchmark for Evaluating Large Language Models on Planning and Reasoning about Change

Reference 103

Resolution
unresolved
no resolver link, observed 2026-08-12T17:47:27.958065Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:47:27.958065Z digest=sha256:8ceacca99d1d3dac086767699028c8c65d794f3c1643a01852749419ad87d089

Observation fdac3854-9208-4f7e-b8a9-0ce2b26796fe · inbound

Evaluating LLM Reasoning in the Operations Research Domain with ORQA cites this paper.

Evaluating LLM Reasoning in the Operations Research Domain with ORQA PlanBench: An Extensible Benchmark for Evaluating Large Language Models on Planning and Reasoning about Change

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-11T06:01:13.407994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T06:01:13.407994Z digest=sha256:0e6821f707b856925a54e04ca387205077ab00a7a8b5c8d5cb4344bea6bd785f

Observation 6c337865-8896-49b5-8a4f-99a5a4cdd197 · inbound

A Tool for In-depth Analysis of Code Execution Reasoning of Large Language Models cites this paper.

A Tool for In-depth Analysis of Code Execution Reasoning of Large Language Models PlanBench: An Extensible Benchmark for Evaluating Large Language Models on Planning and Reasoning about Change

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-09T23:21:56.512878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T23:21:56.512878Z digest=sha256:04b636733a74437b32e092176f6b5904563edae0f1bfd73ceb5af109ed0a7a16

Observation fafff5a4-8a2a-4493-9f1c-9db9facc9e43 · inbound

Successor-Generator Planning with LLM-generated Heuristics cites this paper.

Successor-Generator Planning with LLM-generated Heuristics PlanBench: An Extensible Benchmark for Evaluating Large Language Models on Planning and Reasoning about Change

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-09T22:34:42.282181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T22:34:42.282181Z digest=sha256:bb6fd020d6e0e30c75779b1b27424c6dcd2942d7c4e8f1390ef5078aef2dd9c5

Observation 974da3df-3e6e-49d1-a9d6-e115a0e0bc69 · inbound

Lost in Cultural Translation: Do LLMs Struggle with Math Across Cultural Contexts? cites this paper.

Lost in Cultural Translation: Do LLMs Struggle with Math Across Cultural Contexts? PlanBench: An Extensible Benchmark for Evaluating Large Language Models on Planning and Reasoning about Change

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-22T22:27:12.365359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-22T22:27:02.059459Z digest=sha256:516926e1f674f2986bca96ad69cf4c312a0d9f68ed06132219113ade3f10d1f1

Observation 96af8fdd-1d7b-41c9-be48-323a75670dcb · inbound

Toward Generalizable Evaluation in the LLM Era: A Survey Beyond Benchmarks cites this paper.

Toward Generalizable Evaluation in the LLM Era: A Survey Beyond Benchmarks PlanBench: An Extensible Benchmark for Evaluating Large Language Models on Planning and Reasoning about Change

Reference 107

Resolution
unresolved
no resolver link, observed 2026-08-16T10:12:00.507239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:12:00.507239Z digest=sha256:ecbdadf70b21ee43c4c44abcdd156d0b5df62b931eee0103f8fc8522dd35177c

Observation 70585823-cd63-4605-b721-20a1a7f3b1ec · inbound

Toward Generalizable Evaluation in the LLM Era: A Survey Beyond Benchmarks cites this paper.

Toward Generalizable Evaluation in the LLM Era: A Survey Beyond Benchmarks PlanBench: An Extensible Benchmark for Evaluating Large Language Models on Planning and Reasoning about Change

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-16T10:12:00.512507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:12:00.512507Z digest=sha256:77554fafefbe1e901f2d399238364ccb50f48ee2559c0bdf7e54a42e22e52e7d

Observation 8588f832-6e54-4e86-a15a-c5f113590284 · inbound

The Illusion of Thinking: Understanding the Strengths and Limitations of Reasoning Models via the Lens of Problem Complexity cites this paper.

The Illusion of Thinking: Understanding the Strengths and Limitations of Reasoning Models via the Lens of Problem Complexity PlanBench: An Extensible Benchmark for Evaluating Large Language Models on Planning and Reasoning about Change

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-15T16:10:31.560405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-15T16:10:31.440921Z digest=sha256:f3f4dcc5540bcac997903ddce986e914906ecf6083f40653ff8eb631e293b90e

Observation 3f61024a-0ca6-4b6f-a3fe-7ec72afc538f · inbound

GenPlanX. Generation of Plans and Execution cites this paper.

GenPlanX. Generation of Plans and Execution PlanBench: An Extensible Benchmark for Evaluating Large Language Models on Planning and Reasoning about Change

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T04:18:26.474007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:18:26.474007Z digest=sha256:beecb7bba2e8a6356783e59883424cf30e174187b41cc0a2da06ac30f6c92552

Observation d2e22f05-25ae-4671-b40e-d5b1c81925dc · inbound

Application of LLMs to Multi-Robot Path Planning and Task Allocation cites this paper.

Application of LLMs to Multi-Robot Path Planning and Task Allocation PlanBench: An Extensible Benchmark for Evaluating Large Language Models on Planning and Reasoning about Change

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T18:48:17.855773Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:48:17.855773Z digest=sha256:8d6e6418b7c3277815887c0f8d5d17ad5e6e6ad065d111a3d6cb8fa6cdba43eb

Observation 1f25c7e5-5586-4b32-85a8-31548ce6f10d · inbound

ISO-Bench: Benchmarking Multimodal Causal Reasoning in Visual-Language Models through Procedural Plans cites this paper.

ISO-Bench: Benchmarking Multimodal Causal Reasoning in Visual-Language Models through Procedural Plans PlanBench: An Extensible Benchmark for Evaluating Large Language Models on Planning and Reasoning about Change

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T11:08:10.452864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T11:08:10.452864Z digest=sha256:bdd2e1d3991929c89aa92ebd9c8cae0465624fb3551326a8d4d8251d1493b9f5

Observation d73f5b39-ae97-4245-838c-a31c63948ff2 · inbound

Assessing Coherency and Consistency of Code Execution Reasoning by Large Language Models cites this paper.

Assessing Coherency and Consistency of Code Execution Reasoning by Large Language Models PlanBench: An Extensible Benchmark for Evaluating Large Language Models on Planning and Reasoning about Change

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-18T06:00:57.198770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-18T05:59:00.400429Z digest=sha256:175e32983e5164525e07bd631113f80a79c1b7afd396f1a7348b8fbf9c538451

Observation 1dab2a04-e5c0-4f11-a94c-aa669c1db994 · inbound

SAT: Sequential Agent Tuning for Coordinator Free Plug and Play Multi-LLM Training with Monotonic Improvement Guarantees cites this paper.

SAT: Sequential Agent Tuning for Coordinator Free Plug and Play Multi-LLM Training with Monotonic Improvement Guarantees PlanBench: An Extensible Benchmark for Evaluating Large Language Models on Planning and Reasoning about Change

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-10T09:38:42.239711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-10T09:34:33.271537Z digest=sha256:64dd82ee072ffc4cef40ecdf3a0ee65fcb962198b469ba72fb0eadd92f2ae9a8

Observation 95108b2b-5ec8-4fbd-8306-a6b0d4747c0b · inbound

OPT-BENCH: Evaluating the Iterative Self-Optimization of LLM Agents in Large-Scale Search Spaces cites this paper.

OPT-BENCH: Evaluating the Iterative Self-Optimization of LLM Agents in Large-Scale Search Spaces PlanBench: An Extensible Benchmark for Evaluating Large Language Models on Planning and Reasoning about Change

Reference 50

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T03:01:18.575710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-12T02:57:15.521594Z digest=sha256:20ad50a3794a628e0f05674b0f2e7703cdcc96d8af45ce79cd64ee4cb83b30e7

Observation 55b9f9ec-cad0-4ceb-9d3d-29102fa8297e · inbound

Consistency as a Testable Property: Statistical Methods to Evaluate AI Agent Reliability cites this paper.

Consistency as a Testable Property: Statistical Methods to Evaluate AI Agent Reliability PlanBench: An Extensible Benchmark for Evaluating Large Language Models on Planning and Reasoning about Change

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-12T04:41:21.825043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-12T04:41:15.286881Z digest=sha256:b72e6a67f86881df32b1112e6d43a780e0a2b041611fb2ea9cfda5f555403be3

Observation d460edfd-ebd1-4424-a809-5fcbf51d7b1e · inbound

Zero-Shot Goal Recognition with Large Language Models cites this paper.

Zero-Shot Goal Recognition with Large Language Models PlanBench: An Extensible Benchmark for Evaluating Large Language Models on Planning and Reasoning about Change

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-19T16:12:39.235918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-19T16:09:12.822056Z digest=sha256:c6e76d7bd18649f8615660f79d48cef2d52d09003e015257c32dd459705f3f5a

Observation 7eb88bfb-b7f7-4f44-b3ce-5b81605174dc · inbound

HyperGuide: Hyperbolic Guidance for Efficient Multi-Step Reasoning in Large Language Models cites this paper.

HyperGuide: Hyperbolic Guidance for Efficient Multi-Step Reasoning in Large Language Models PlanBench: An Extensible Benchmark for Evaluating Large Language Models on Planning and Reasoning about Change

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-04T05:05:38.669864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T05:05:38.669864Z digest=sha256:88b50eb0c7dfd4dee2e3c67500964f2f464507d54fe01cdb2d2698e4f9ff3693

Observation 1ff5f9a1-93ce-4990-9f19-baf398000339 · inbound

Managing Uncertainty in LLM-Generated Procedural Knowledge for Virtual Laboratory Planning cites this paper.

Managing Uncertainty in LLM-Generated Procedural Knowledge for Virtual Laboratory Planning PlanBench: An Extensible Benchmark for Evaluating Large Language Models on Planning and Reasoning about Change

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-06-29T21:23:58.747294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-29T21:21:45.412616Z digest=sha256:552160d2cf7424d6e9887001d7fefd051d9e8fcd463f3196da271b46a18190fa

Observation 5e40099d-66c1-4566-bf85-85ca3420a0c6 · inbound

REPOT: Recoverable Program-of-Thought via Checkpoint Repair cites this paper.

REPOT: Recoverable Program-of-Thought via Checkpoint Repair PlanBench: An Extensible Benchmark for Evaluating Large Language Models on Planning and Reasoning about Change

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-06-29T06:23:08.706412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-06-29T06:22:46.440678Z digest=sha256:822e3be3c6c8f6638b7e1c3edc1030d9433bb784a5a2d81654f13547b4320e20

Observation ce0c0d53-c2c9-4b7d-b831-feb1191cd8a1 · inbound

Lost in Aggregation: A Multi-Scale Diagnostic Benchmark for LLM Spatial Navigation cites this paper.

Lost in Aggregation: A Multi-Scale Diagnostic Benchmark for LLM Spatial Navigation PlanBench: An Extensible Benchmark for Evaluating Large Language Models on Planning and Reasoning about Change

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-06-26T10:39:18.328036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-26T10:37:24.946718Z digest=sha256:f5ac0067cc662125ce15fb3c0c68c1c0bcfbda3f2c7e00189813dcf09bf0195e

Observation 5d6a6729-a5ca-414c-aa49-15f52a989fe1 · inbound

Search, Fail, Recover: A Training Framework for Correction-Aware Reasoning cites this paper.

Search, Fail, Recover: A Training Framework for Correction-Aware Reasoning PlanBench: An Extensible Benchmark for Evaluating Large Language Models on Planning and Reasoning about Change

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-07-09T09:26:08.694841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-09T09:20:08.212837Z digest=sha256:fa40ac01d755797f8730a56cb882f20b47237c86ac27dfe097c6d34f634ba810

Observation cbeebfab-0187-479c-a309-613c04f069b1 · inbound

LatticeMind: A Conflict-Aware Memory Primitive for Multi-Agent Systems cites this paper.

LatticeMind: A Conflict-Aware Memory Primitive for Multi-Agent Systems PlanBench: An Extensible Benchmark for Evaluating Large Language Models on Planning and Reasoning about Change

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T00:19:26.639704Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T00:19:26.639704Z digest=sha256:8a75c50f9962bd3b3e122a6163eacc244848569c11b38c1ddb3055bb700cf7c3