Pith. sign in

Paper Citation Record · LEDGER

LLMs Still Can't Plan; Can LRMs? A Preliminary Evaluation of OpenAI's o1 on PlanBench

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 15 inbound Pith citation observations for arXiv:2409.13373.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2409.13373 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 15 of 15 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 15 of 15 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:17:11.933468Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T20:46:14.059699Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 0ef2c66e-f8bd-463a-af41-24e46fee40ee · inbound

Towards Large Reasoning Models: A Survey of Reinforced Reasoning with Large Language Models cites this paper.

Towards Large Reasoning Models: A Survey of Reinforced Reasoning with Large Language Models LLMs Still Can't Plan; Can LRMs? A Preliminary Evaluation of OpenAI's o1 on PlanBench

Reference 146

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T21:20:59.312729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T21:20:59.128986Z digest=sha256:bbecfbe590a3abbe8ce6de1ef2968d1a96365d1dc902713a4a124de677cb2cd4

Observation 8db6e529-9719-4a96-84a7-908ef14ccbbb · inbound

A Comprehensive Survey of Agents for Computer Use: Foundations, Challenges, and Future Directions cites this paper.

A Comprehensive Survey of Agents for Computer Use: Foundations, Challenges, and Future Directions LLMs Still Can't Plan; Can LRMs? A Preliminary Evaluation of OpenAI's o1 on PlanBench

Reference 157

Resolution
verified exact
arxiv_id, observed 2026-05-23T05:02:35.955850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-23T04:59:36.994758Z digest=sha256:e32586ef068bc6e985bd11f09f5455b3ce35f64c6ada6076d4e8c78c8307e87d

Observation 22eb4399-b9e0-4c8c-bcac-f6660bdcc979 · inbound

The Road to Generalizable Neuro-Symbolic Learning Should be Paved with Foundation Models cites this paper.

The Road to Generalizable Neuro-Symbolic Learning Should be Paved with Foundation Models LLMs Still Can't Plan; Can LRMs? A Preliminary Evaluation of OpenAI's o1 on PlanBench

Reference 2015

Resolution
unresolved
no resolver link, observed 2026-08-07T12:17:11.933468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:17:11.933468Z digest=sha256:a284b591c94a86080b9b82a4daa3bc53fcfd2c6f6d54c08515c2b989bbe7b711

Observation 16a779d1-6d61-4191-9907-fea88806053e · inbound

LogiPlan: A Structured Benchmark for Logical Planning and Relational Reasoning in LLMs cites this paper.

LogiPlan: A Structured Benchmark for Logical Planning and Relational Reasoning in LLMs LLMs Still Can't Plan; Can LRMs? A Preliminary Evaluation of OpenAI's o1 on PlanBench

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T04:28:46.939636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:28:46.939636Z digest=sha256:82b3065129972d969f02a34d5af540154b6c34d2c0c203bc5a73e54de2af8dbb

Observation 0eea2786-3587-46b6-8b5f-be4c58072569 · inbound

Large Reasoning Models are not thinking straight: on the unreliability of thinking trajectories cites this paper.

Large Reasoning Models are not thinking straight: on the unreliability of thinking trajectories LLMs Still Can't Plan; Can LRMs? A Preliminary Evaluation of OpenAI's o1 on PlanBench

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T21:13:13.017038Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:13:13.017038Z digest=sha256:0757abce2d8d6b72e3404da8b73fb29a778e3e89646e0f0283309510055bf39b

Observation fcc2a004-06f6-47e3-a5c9-d9af1404e752 · inbound

Can LLM-Reasoning Models Replace Classical Planning? A Benchmark Study cites this paper.

Can LLM-Reasoning Models Replace Classical Planning? A Benchmark Study LLMs Still Can't Plan; Can LRMs? A Preliminary Evaluation of OpenAI's o1 on PlanBench

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T10:45:59.558754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T10:45:59.558754Z digest=sha256:0218865502418677d121f3057eb0a29fd64cbe9988ab429d039460488dffe659

Observation db498e8f-a55a-414b-8b03-363333ea7d8e · inbound

Beyond Memorization: Extending Reasoning Depth with Recurrence, Memory and Test-Time Compute Scaling cites this paper.

Beyond Memorization: Extending Reasoning Depth with Recurrence, Memory and Test-Time Compute Scaling LLMs Still Can't Plan; Can LRMs? A Preliminary Evaluation of OpenAI's o1 on PlanBench

Reference 66

Resolution
verified exact
arxiv_id, observed 2026-05-18T20:51:50.864328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T20:49:26.966293Z digest=sha256:e1f93ffb13576c3df8f9685ef9c8d5c04d6e92969774c8d0f97a2affd1a5cf33

Observation ca01ad81-91f9-4a27-a92f-8287cd1823b2 · inbound

CoreThink: A Symbolic Reasoning Layer to reason over Long Horizon Tasks with LLMs cites this paper.

CoreThink: A Symbolic Reasoning Layer to reason over Long Horizon Tasks with LLMs LLMs Still Can't Plan; Can LRMs? A Preliminary Evaluation of OpenAI's o1 on PlanBench

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-05T13:05:33.736041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:05:33.736041Z digest=sha256:f7119249991e92843e5ad6ded45c0e1bf85db5805693d02bfbee1cdf95c2fcda

Observation 7a34e7a8-7f36-43df-ac75-2086d5ec205c · inbound

OPT-Engine: Benchmarking the Limits of LLMs in Optimization Modeling via Complexity Scaling cites this paper.

OPT-Engine: Benchmarking the Limits of LLMs in Optimization Modeling via Complexity Scaling LLMs Still Can't Plan; Can LRMs? A Preliminary Evaluation of OpenAI's o1 on PlanBench

Reference 32

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T16:18:05.361987Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T16:17:43.055124Z digest=sha256:5b0f41f9088ea1bbedf2eb972d0a29f0b721d64e26ab3cbd45e1bd1c43a21181

Observation 2f2f336b-d722-4d80-9679-9e5b1989b44a · inbound

LLM-WikiRace Benchmark: How Far Can LLMs Plan over Real-World Knowledge Graphs? cites this paper.

LLM-WikiRace Benchmark: How Far Can LLMs Plan over Real-World Knowledge Graphs? LLMs Still Can't Plan; Can LRMs? A Preliminary Evaluation of OpenAI's o1 on PlanBench

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-02T22:26:38.525975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T22:26:38.525975Z digest=sha256:9b96e93f0cec133f3eccd31afd1b9284467b83e9b00988a7b52fe8abac661cac

Observation 271c43f6-fbee-40df-ac67-99d40c02221d · inbound

SYMBOLIZER: Symbolic Model-free Task Planning with VLMs cites this paper.

SYMBOLIZER: Symbolic Model-free Task Planning with VLMs LLMs Still Can't Plan; Can LRMs? A Preliminary Evaluation of OpenAI's o1 on PlanBench

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T09:43:49.597853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T05:10:38.939375Z digest=sha256:253b6d017dc810069d6ebcff9316155cb4432d0f451163945a568850efc3b969

Observation d62170c4-c495-43ea-99d1-c0f787ed6409 · inbound

Robust Asynchronous Planning via Auto-Formalization cites this paper.

Robust Asynchronous Planning via Auto-Formalization LLMs Still Can't Plan; Can LRMs? A Preliminary Evaluation of OpenAI's o1 on PlanBench

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T20:46:14.061633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T17:43:24.820773Z digest=sha256:cdf28b155223393d054d78619ccccb0438e670ca00bdf7c1dfd0679e6f9729e4

Observation 2ae09f3e-bb85-4f3c-8a70-a6d10b6d9e6b · inbound

Embark Now: User Demand Oriented Framework for Multi-day Urban Travel Itinerary Planning cites this paper.

Embark Now: User Demand Oriented Framework for Multi-day Urban Travel Itinerary Planning LLMs Still Can't Plan; Can LRMs? A Preliminary Evaluation of OpenAI's o1 on PlanBench

Reference 32

Resolution
unresolved
no resolver link, observed 2026-07-14T10:13:55.891370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T10:13:55.891370Z digest=sha256:5ecd11abdb0f6608dcca858fb65b4197256e35a7eeb01847ccb8e4f007252f78

Observation f1378f71-7fa8-4b2a-a07d-ac12e0427fd7 · inbound

SymStep: Symbolic Step Verification for Logical Reasoning cites this paper.

SymStep: Symbolic Step Verification for Logical Reasoning LLMs Still Can't Plan; Can LRMs? A Preliminary Evaluation of OpenAI's o1 on PlanBench

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-01T03:48:35.008851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:48:35.008851Z digest=sha256:727901e8d569f177bc5378c7677bbdb6f675e4678d3417f8e36f93e8d8dd1986

Observation 3d53ffc1-8162-43c9-9076-006a108fe105 · inbound

Confidently Wrong: Exception Chain Collapse in Frontier LLM Rule Evaluation cites this paper.

Confidently Wrong: Exception Chain Collapse in Frontier LLM Rule Evaluation LLMs Still Can't Plan; Can LRMs? A Preliminary Evaluation of OpenAI's o1 on PlanBench

Reference 12

Resolution
unresolved
no resolver link, observed 2026-07-30T23:41:20.670945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T23:41:20.670945Z digest=sha256:a33df7e30e804aca9ecad49347d65664e2e5377c6786c1c9c55d9f4846d41a12