Pith. sign in

Paper Citation Record · LEDGER

PLANET: A Collection of Benchmarks for Evaluating LLMs' Planning Capabilities

As of 17 August 2026, this Paper Citation Record lists 73 of 73 outbound references and 1 inbound Pith citation observation for arXiv:2504.14773.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.14773 v1

Coverage vector

measured 73 of 73 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T11:42:30.874258Z

measured 74 of 74 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T22:26:36.732238Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

73 of 73 outbound references displayed

  • verified exact5
  • verified fuzzy4
  • unresolved64
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 437f4928-54b3-4d1b-bbf3-338c7770a7a8 · outbound

This paper cites write newline.

PLANET: A Collection of Benchmarks for Evaluating LLMs' Planning Capabilities write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-16T11:42:30.566702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:42:30.566702Z digest=sha256:dda0b045c972801fc0b04118c97fc5ae6343a8ff97d10e09bd714a22f9d938d9

Observation 800a59d7-ca20-4f31-bdac-1ff3fc66d388 · outbound

This paper cites Can Large Language Models be Good Path Planners? A Benchmark and Investigation on Spatial-temporal Reasoning.

PLANET: A Collection of Benchmarks for Evaluating LLMs' Planning Capabilities Can Large Language Models be Good Path Planners? A Benchmark and Investigation on Spatial-temporal Reasoning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-16T11:42:30.572023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:42:30.572023Z digest=sha256:3e292076d8b01bd79c4ba7995a08038c6a834c474e458d6293228cd71c05912f

Observation f709a02c-cd01-4374-a0c9-91ac323d5e7f · outbound

This paper cites WorkArena++: Towards Compositional Planning and Reasoning-based Common Knowledge Work Tasks.

PLANET: A Collection of Benchmarks for Evaluating LLMs' Planning Capabilities WorkArena++: Towards Compositional Planning and Reasoning-based Common Knowledge Work Tasks

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-16T11:42:30.576526Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:42:30.576526Z digest=sha256:c2987729322a1ae6bc02d8de03b267d7e0cda20674eca0eb8b33f427c1575271

Observation 97692f16-a116-4764-a58a-102f8bf982aa · outbound

This paper cites Windows Agent Arena: Evaluating Multi-Modal OS Agents at Scale.

PLANET: A Collection of Benchmarks for Evaluating LLMs' Planning Capabilities Windows Agent Arena: Evaluating Multi-Modal OS Agents at Scale

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-16T11:42:30.581189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:42:30.581189Z digest=sha256:7716aead247d06bef420b37b46eb27ecccdac110d31ce61293e7ee558938afaf

Observation 68bbc0d6-bcd5-4a8f-9e74-dffb0a2f6aed · outbound

This paper cites PlaSma: Making Small Language Models Better Procedural Knowledge Models for (Counterfactual) Planning.

PLANET: A Collection of Benchmarks for Evaluating LLMs' Planning Capabilities PlaSma: Making Small Language Models Better Procedural Knowledge Models for (Counterfactual) Planning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-16T11:42:30.585566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:42:30.585566Z digest=sha256:3328396fdb8b783ab82472a49e04b04e04f86c60551191c67d5af76acda542fb

Observation 05665e30-b543-419c-8c3a-a6438eb26726 · outbound

This paper cites PARTNR: A Benchmark for Planning and Reasoning in Embodied Multi-agent Tasks.

PLANET: A Collection of Benchmarks for Evaluating LLMs' Planning Capabilities PARTNR: A Benchmark for Planning and Reasoning in Embodied Multi-agent Tasks

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-16T11:42:30.589774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:42:30.589774Z digest=sha256:b5dc16373183fdbf2ec34e14eeb39570b11018d306073a6d20488970b51e6cee

Observation ed1af03c-e461-4663-a38e-530ff0dc4c76 · outbound

This paper cites Put Your Money Where Your Mouth Is: Evaluating Strategic Planning and Execution of LLM Agents in an Auction Arena.

PLANET: A Collection of Benchmarks for Evaluating LLMs' Planning Capabilities Put Your Money Where Your Mouth Is: Evaluating Strategic Planning and Execution of LLM Agents in an Auction Arena

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-16T11:42:30.594239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:42:30.594239Z digest=sha256:11090789f0c11cffcb84f219de468dfe888d768c1f6460367dc31c1225820c81

Observation b46779be-7e7d-4d18-9cdc-967121141d92 · outbound

This paper cites EgoPlan-Bench: Benchmarking Multimodal Large Language Models for Human-Level Planning.

PLANET: A Collection of Benchmarks for Evaluating LLMs' Planning Capabilities EgoPlan-Bench: Benchmarking Multimodal Large Language Models for Human-Level Planning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-16T11:42:30.598769Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:42:30.598769Z digest=sha256:d186c3aee065b9a38ff3527f4607c2f152f8001e6f2fa491e589ad004858f159

Observation 86f944a4-8e43-4c60-ae28-b7830bfac183 · outbound

This paper cites Do Large Language Models have Problem-Solving Capability under Incomplete Information Scenarios?.

PLANET: A Collection of Benchmarks for Evaluating LLMs' Planning Capabilities Do Large Language Models have Problem-Solving Capability under Incomplete Information Scenarios?

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-08-16T11:42:31.753342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T11:42:30.602664Z digest=sha256:17f7e904a4e94e8e0f935809ac5423857089253ea30659adf39a026c8fc6a0f9

Observation b7983657-5c1b-4998-817b-1220bf106969 · outbound

This paper cites Plancraft: an evaluation dataset for planning with LLM agents.

PLANET: A Collection of Benchmarks for Evaluating LLMs' Planning Capabilities Plancraft: an evaluation dataset for planning with LLM agents

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-16T11:42:30.606642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:42:30.606642Z digest=sha256:919f52eb08a8d800cb4188cf841e3736ffa19697bbe3cdc60d1e8808dab0b607

Observation 092da20e-953e-445f-8088-43b099aa739e · outbound

This paper cites Mind2Web: Towards a Generalist Agent for the Web.

PLANET: A Collection of Benchmarks for Evaluating LLMs' Planning Capabilities Mind2Web: Towards a Generalist Agent for the Web

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-16T11:42:30.610641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:42:30.610641Z digest=sha256:5b7f7071b53286a410956e0e1db37ee55184d53fa6b3bac71c02afacd11a43be

Observation 04eed57b-eacd-4bb5-813e-9aa0ca54f0c9 · outbound

This paper cites GTBench: Uncovering the Strategic Reasoning Limitations of LLMs via Game-Theoretic Evaluations.

PLANET: A Collection of Benchmarks for Evaluating LLMs' Planning Capabilities GTBench: Uncovering the Strategic Reasoning Limitations of LLMs via Game-Theoretic Evaluations

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-16T11:42:30.614767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:42:30.614767Z digest=sha256:ff97d7cbbc8896fd59df5171cf99969ba9b25d54dff93d240194135d14a9d37f

Observation fda76345-77e4-4320-b51e-1bcb93bba2bf · outbound

This paper cites MineDojo: Building Open-Ended Embodied Agents with Internet-Scale Knowledge.

PLANET: A Collection of Benchmarks for Evaluating LLMs' Planning Capabilities MineDojo: Building Open-Ended Embodied Agents with Internet-Scale Knowledge

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-16T11:42:30.618903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:42:30.618903Z digest=sha256:035a241b68a1749c53872d17946428357b629ba8f3096edb69ce8212bd6cae2f

Observation d1faf3ff-6695-493b-8e3e-5f3445e96331 · outbound

This paper cites REALM-Bench: A Benchmark for Evaluating Multi-Agent Systems on Real-world, Dynamic Planning and Scheduling Tasks.

PLANET: A Collection of Benchmarks for Evaluating LLMs' Planning Capabilities REALM-Bench: A Benchmark for Evaluating Multi-Agent Systems on Real-world, Dynamic Planning and Scheduling Tasks

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-16T11:42:30.623040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:42:30.623040Z digest=sha256:88c29979b16824ea6ed61809cd00f11a6216272d1c01fafcea2c2996e496252b

Observation 7759bf5f-4757-4694-a544-cc100b26917a · outbound

This paper cites Robotouille: An Asynchronous Planning Benchmark for LLM Agents.

PLANET: A Collection of Benchmarks for Evaluating LLMs' Planning Capabilities Robotouille: An Asynchronous Planning Benchmark for LLM Agents

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-16T11:42:30.627423Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:42:30.627423Z digest=sha256:32397e2851288c5173aa6a1eda871e1f7d5be787fd00a0d1b0420b3cf07541a8

Observation ed876cd4-b325-4b55-bc97-896ac88ed323 · outbound

This paper cites an unresolved cited work.

PLANET: A Collection of Benchmarks for Evaluating LLMs' Planning Capabilities Unresolved cited work

Reference 16

Resolution
verified exact
doi, observed 2026-08-16T11:42:30.948573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T11:42:30.631643Z digest=sha256:b1eda87ae37f0615940c719ca463cc482ab6ffca5327bcd3f3f142271f309345

Observation 2e143f0e-a9fb-4317-a8c2-e54b4661ceb3 · outbound

This paper cites Benchmarking the Spectrum of Agent Capabilities.

PLANET: A Collection of Benchmarks for Evaluating LLMs' Planning Capabilities Benchmarking the Spectrum of Agent Capabilities

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-16T11:42:30.635685Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:42:30.635685Z digest=sha256:8ceff075e71cf699598cfbbbf6217cd0c9e1b6af5f9c706ee0d4a50505778fc9

Observation 9dbba190-bc17-427e-b823-07caf2aa6e31 · outbound

This paper cites Reasoning with Language Model is Planning with World Model.

PLANET: A Collection of Benchmarks for Evaluating LLMs' Planning Capabilities Reasoning with Language Model is Planning with World Model

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-16T11:42:30.639672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:42:30.639672Z digest=sha256:a8093bed7550647223cebdf914c7ff71a2b2fa7506769bbe81c80b498b2d0eb8

Observation afcca4b6-03c2-4bfc-be1f-dea03a5b5865 · outbound

This paper cites Planning Anything with Rigor: General-Purpose Zero-Shot Planning with LLM-based Formalized Programming.

PLANET: A Collection of Benchmarks for Evaluating LLMs' Planning Capabilities Planning Anything with Rigor: General-Purpose Zero-Shot Planning with LLM-based Formalized Programming

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-16T11:42:30.643822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:42:30.643822Z digest=sha256:724b9ec923e0c9d8846a6569b72411b51fb6c150589692942e8e906a892d732c

Observation 9b0de9c2-5572-4871-b4f5-5653cda51224 · outbound

This paper cites Hoffmann and S.

PLANET: A Collection of Benchmarks for Evaluating LLMs' Planning Capabilities Hoffmann and S

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-16T11:42:30.648012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:42:30.648012Z digest=sha256:25ef459d4000786589d7dec2211bca69f142b270979885a831c329a94d718847

Observation 5acfda5a-ea43-4cd1-ad6c-79bdee617df5 · outbound

This paper cites Game-theoretic LLM: Agent Workflow for Negotiation Games.

PLANET: A Collection of Benchmarks for Evaluating LLMs' Planning Capabilities Game-theoretic LLM: Agent Workflow for Negotiation Games

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-16T11:42:30.652939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:42:30.652939Z digest=sha256:a4b40cd424932fece5bb0a23d797d33cdc0d5ef323988bb2d3150ff6f7c094d5

Observation becebb25-ba08-416d-8851-3b465549bbc4 · outbound

This paper cites How Far Are We on the Decision-Making of LLMs? Evaluating LLMs' Gaming Ability in Multi-Agent Environments.

PLANET: A Collection of Benchmarks for Evaluating LLMs' Planning Capabilities How Far Are We on the Decision-Making of LLMs? Evaluating LLMs' Gaming Ability in Multi-Agent Environments

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-16T11:42:30.657427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:42:30.657427Z digest=sha256:b332c6ec1a6de4c26716c69c1e7403c515325a95e22e105c37e02c3720fcf11d

Observation 5eac89ff-f8ae-49e7-8d6b-c343b7e19715 · outbound

This paper cites CRMArena: Understanding the Capacity of LLM Agents to Perform Professional CRM Tasks in Realistic Environments.

PLANET: A Collection of Benchmarks for Evaluating LLMs' Planning Capabilities CRMArena: Understanding the Capacity of LLM Agents to Perform Professional CRM Tasks in Realistic Environments

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-16T11:42:30.661831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:42:30.661831Z digest=sha256:bbf4f253164f0736e1b888e6cd915bd2ee9edd5ab5762d62dd85b520912b4419

Observation 9dc7b07c-eb70-4f03-a556-6dc64688e083 · outbound

This paper cites Understanding the planning of LLM agents: A survey.

PLANET: A Collection of Benchmarks for Evaluating LLMs' Planning Capabilities Understanding the planning of LLM agents: A survey

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-16T11:42:30.665906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:42:30.665906Z digest=sha256:95d6b2ab68bb8ba982e15f2430867f70c26d428a6d04d44207e29f6c0edec239

Observation caccdbe1-0d3e-4b59-b5fd-97c5c77d6fbd · outbound

This paper cites McNamara, and Deming Chen.

PLANET: A Collection of Benchmarks for Evaluating LLMs' Planning Capabilities McNamara, and Deming Chen

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-16T11:42:30.669993Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:42:30.669993Z digest=sha256:a4b83bcee98c3e579612918bc5ed60342c35953ee30afdb58549db20a6e0d02c

Observation 4b81b7d5-62d8-4374-8a41-03c8649e1a1b · outbound

This paper cites SWE-bench: Can Language Models Resolve Real-World GitHub Issues?.

PLANET: A Collection of Benchmarks for Evaluating LLMs' Planning Capabilities SWE-bench: Can Language Models Resolve Real-World GitHub Issues?

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-16T11:42:30.673917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:42:30.673917Z digest=sha256:9d0cb520d2ee5035a11545ad5de46cbe1626c4bdb910ea503f2519400820903b

Observation 9f602155-bf47-426c-8f98-f3ba1cad0d82 · outbound

This paper cites To the Globe (TTG): Towards Language-Driven Guaranteed Travel Planning.

PLANET: A Collection of Benchmarks for Evaluating LLMs' Planning Capabilities To the Globe (TTG): Towards Language-Driven Guaranteed Travel Planning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-16T11:42:30.677816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:42:30.677816Z digest=sha256:871c3e28d8dd25b2048a86fac959bcf751ed076c7e2463e440b343d6ebb823f8

Observation 00bbfa85-def3-45a1-b19c-575aecf3ae68 · outbound

This paper cites Towards a foundation for evaluating ai planners.

PLANET: A Collection of Benchmarks for Evaluating LLMs' Planning Capabilities Towards a foundation for evaluating ai planners

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:42:31.930089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T11:42:30.681821Z digest=sha256:ae3d1d79f200e258959d27aa5c2b34525606a8074bdfbf5c8186935d0d36244d

Observation de3cb5ef-3fa9-44f9-8d63-15a385916f7e · outbound

This paper cites VisualWebArena: Evaluating Multimodal Agents on Realistic Visual Web Tasks.

PLANET: A Collection of Benchmarks for Evaluating LLMs' Planning Capabilities VisualWebArena: Evaluating Multimodal Agents on Realistic Visual Web Tasks

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-16T11:42:30.686117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:42:30.686117Z digest=sha256:9d88a7f37efabf37e3b3676820da36a3b2bef0a87a6505b28a142709b7de9896

Observation effb16c8-f817-4a4c-8669-f96e58f0fb2d · outbound

This paper cites Beyond A*: Better Planning with Transformers via Search Dynamics Bootstrapping.

PLANET: A Collection of Benchmarks for Evaluating LLMs' Planning Capabilities Beyond A*: Better Planning with Transformers via Search Dynamics Bootstrapping

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-16T11:42:30.690136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:42:30.690136Z digest=sha256:5884395134bbe44f90a68e6838f8268b5d1abe7025edd04a8e01106df08bfa59

Observation 8d8762b8-1d70-4513-b343-7e784854a888 · outbound

This paper cites LASP: Surveying the State-of-the-Art in Large Language Model-Assisted AI Planning.

PLANET: A Collection of Benchmarks for Evaluating LLMs' Planning Capabilities LASP: Surveying the State-of-the-Art in Large Language Model-Assisted AI Planning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-16T11:42:30.695376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:42:30.695376Z digest=sha256:3d7d50691b86d83f252aab974ec626f73af285057edeca7c5602ffb8b576a8e9

Observation ad2e18da-06fb-4a7d-9881-341f6efbf7d9 · outbound

This paper cites AgentBench: Evaluating LLMs as Agents.

PLANET: A Collection of Benchmarks for Evaluating LLMs' Planning Capabilities AgentBench: Evaluating LLMs as Agents

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-16T11:42:30.699686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:42:30.699686Z digest=sha256:8584522cd6103e450e69d50578ecb8f7c65824dbf8995e49091f17dca63689a9

Observation 2342f4ae-9a70-4c27-aa32-a65f467b1c02 · outbound

This paper cites AgentBoard: An Analytical Evaluation Board of Multi-turn LLM Agents.

PLANET: A Collection of Benchmarks for Evaluating LLMs' Planning Capabilities AgentBoard: An Analytical Evaluation Board of Multi-turn LLM Agents

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-16T11:42:30.703522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:42:30.703522Z digest=sha256:dc568bd2862e93b5b8a20dd5af2be37b128e1b62e2f887d2704269a7f1601d83

Observation bd41abc8-7cfd-40e0-9b37-0378cec5168e · outbound

This paper cites GAIA: a benchmark for General AI Assistants.

PLANET: A Collection of Benchmarks for Evaluating LLMs' Planning Capabilities GAIA: a benchmark for General AI Assistants

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-16T11:42:30.707355Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:42:30.707355Z digest=sha256:7c49573ed66cd6a6007d0cc8c7710bac9b7fa5358e42fba4738f9a445714ec6a

Observation ef09f6f1-af60-48e0-b07e-59c5772b4654 · outbound

This paper cites LLaMAR: Long-Horizon Planning for Multi-Agent Robots in Partially Observable Environments.

PLANET: A Collection of Benchmarks for Evaluating LLMs' Planning Capabilities LLaMAR: Long-Horizon Planning for Multi-Agent Robots in Partially Observable Environments

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-16T11:42:30.712090Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:42:30.712090Z digest=sha256:0c8ad56fed1d300455ff29eae6bcd6ba2c1dab082b4fafb8c5e3aedf32e14bff

Observation 27054b1d-5257-4aca-b532-c1dd74c79b9d · outbound

This paper cites WorldAPIs: The World Is Worth How Many APIs? A Thought Experiment.

PLANET: A Collection of Benchmarks for Evaluating LLMs' Planning Capabilities WorldAPIs: The World Is Worth How Many APIs? A Thought Experiment

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-16T11:42:30.716603Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:42:30.716603Z digest=sha256:5a3368c306e8cc77c9fae0c2c6a82923bddf98f203b2232d262e928f182432ed

Observation 02e6f43b-b23f-4025-b65a-6f953d5a0063 · outbound

This paper cites TEACh: Task-driven Embodied Agents that Chat.

PLANET: A Collection of Benchmarks for Evaluating LLMs' Planning Capabilities TEACh: Task-driven Embodied Agents that Chat

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-16T11:42:30.720846Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:42:30.720846Z digest=sha256:c709d2dc70575d3f0ba8d55d60a9cfae30ae40f6a4206af6dd6876bb26e87ced

Observation 1270d246-03f5-49a0-971f-c7c87a248c62 · outbound

This paper cites VirtualHome: Simulating Household Activities via Programs.

PLANET: A Collection of Benchmarks for Evaluating LLMs' Planning Capabilities VirtualHome: Simulating Household Activities via Programs

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-16T11:42:30.725194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:42:30.725194Z digest=sha256:0264d2975cb2a34da26f77547e5a3098dc209e689e404123805060656590ccf8

Observation 4df2c259-59b7-45bc-8d13-36a89e34f829 · outbound

This paper cites Artificial I ntelligence: A modern approach.

PLANET: A Collection of Benchmarks for Evaluating LLMs' Planning Capabilities Artificial I ntelligence: A modern approach

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:42:31.916816Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T11:42:30.729516Z digest=sha256:fe22bab78bfe97fdab6b46af7c6c235dd3b59c1acbbd35967991c901362afbd5

Observation bcccf922-b9f6-4e6f-85b1-70b95ae322d1 · outbound

This paper cites an unresolved cited work.

PLANET: A Collection of Benchmarks for Evaluating LLMs' Planning Capabilities Unresolved cited work

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-16T11:42:30.733203Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:42:30.733203Z digest=sha256:f50f6faaf76295e80518799057e1629d7741976cb4d8813be56ed03c65a41e05

Observation fe945321-7320-42a0-9d3b-304a3af33e46 · outbound

This paper cites Language Models Are Greedy Reasoners: A Systematic Formal Analysis of Chain-of-Thought.

PLANET: A Collection of Benchmarks for Evaluating LLMs' Planning Capabilities Language Models Are Greedy Reasoners: A Systematic Formal Analysis of Chain-of-Thought

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-16T11:42:30.737180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:42:30.737180Z digest=sha256:6d4a7fbf6c89f2c8441c2c7c944fed9a035f72e02c270d63d9321a2318a0489b

Observation 0ac9f86e-7e86-43ca-871f-adcbb3cef6f7 · outbound

This paper cites Reflexion: Language agents with verbal reinforcement learning.

PLANET: A Collection of Benchmarks for Evaluating LLMs' Planning Capabilities Reflexion: Language agents with verbal reinforcement learning

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-16T11:42:30.741392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:42:30.741392Z digest=sha256:b628fb724d5a89c4b8e39358ae5278659538d75b05b0f8b35a1c213106a27fca

Observation 1cf8db43-aeec-4025-be1f-3cd0991f6f64 · outbound

This paper cites ALFRED: A Benchmark for Interpreting Grounded Instructions for Everyday Tasks.

PLANET: A Collection of Benchmarks for Evaluating LLMs' Planning Capabilities ALFRED: A Benchmark for Interpreting Grounded Instructions for Everyday Tasks

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-16T11:42:30.745222Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:42:30.745222Z digest=sha256:a86bb4393c3a441b1034be14c30047a88a928a680da6f9e2602a261dc5d35d04

Observation f079fb56-3404-4631-b8de-a847da033a08 · outbound

This paper cites ALFWorld: Aligning Text and Embodied Environments for Interactive Learning.

PLANET: A Collection of Benchmarks for Evaluating LLMs' Planning Capabilities ALFWorld: Aligning Text and Embodied Environments for Interactive Learning

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-16T11:42:30.753570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:42:30.753570Z digest=sha256:13ef4b62fc77b29ac54ab70fcbb0c512931e35f260f49bd109800e300f4670ff

Observation eb34491c-4abb-4cea-8aa6-c7fc90b09bdf · outbound

This paper cites Dualformer: Controllable Fast and Slow Thinking by Learning with Randomized Reasoning Traces.

PLANET: A Collection of Benchmarks for Evaluating LLMs' Planning Capabilities Dualformer: Controllable Fast and Slow Thinking by Learning with Randomized Reasoning Traces

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-16T11:42:30.757325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:42:30.757325Z digest=sha256:6c64a7126daa0cac0371603120e686ebf1941cd3521c8572c53049aeb2611042

Observation 22f22cfb-abaf-4fd9-bac9-82bbb914c959 · outbound

This paper cites A ct P lan-1 K : Benchmarking the procedural planning ability of visual language models in household activities.

PLANET: A Collection of Benchmarks for Evaluating LLMs' Planning Capabilities A ct P lan-1 K : Benchmarking the procedural planning ability of visual language models in household activities

Reference 47

Resolution
verified exact
doi, observed 2026-08-16T11:42:30.918517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T11:42:30.761635Z digest=sha256:64492be7884e78d29355dcbb2cddbaa082831e9807d5546717ed5b492560d99d

Observation af7c1c86-79ce-42d9-b11a-a338691e9d9e · outbound

This paper cites Stop Overthinking: A Survey on Efficient Reasoning for Large Language Models.

PLANET: A Collection of Benchmarks for Evaluating LLMs' Planning Capabilities Stop Overthinking: A Survey on Efficient Reasoning for Large Language Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-16T11:42:30.765562Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:42:30.765562Z digest=sha256:3cd211c0db84545bde4230b8912f75636f0d00fdc4b9c5e69152e7b96ab325d4

Observation b289fbf6-e96b-4180-87ea-a1f0e017baae · outbound

This paper cites Planbench: An extensible benchmark for evaluating large language models on planning and reasoning about change.

PLANET: A Collection of Benchmarks for Evaluating LLMs' Planning Capabilities Planbench: An extensible benchmark for evaluating large language models on planning and reasoning about change

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:42:31.895495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T11:42:30.770331Z digest=sha256:9954f32a35add9deb7501d503ba7855f5972a5aa4ee974afe98c7d0c416df187

Observation e19cb867-08ee-43b7-82b8-c8e95f7d22aa · outbound

This paper cites TMGBench: A Systematic Game Benchmark for Evaluating Strategic Reasoning Abilities of LLMs.

PLANET: A Collection of Benchmarks for Evaluating LLMs' Planning Capabilities TMGBench: A Systematic Game Benchmark for Evaluating Strategic Reasoning Abilities of LLMs

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-16T11:42:30.774151Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:42:30.774151Z digest=sha256:80b76f52b5116706bf3d97874f5bdd3d19d45c0a2a9c26d397b69161de30642a

Observation 610b3a32-17dd-4587-ba1b-dca90f4805c7 · outbound

This paper cites ByteSized32: A Corpus and Challenge Task for Generating Task-Specific World Models Expressed as Text Games.

PLANET: A Collection of Benchmarks for Evaluating LLMs' Planning Capabilities ByteSized32: A Corpus and Challenge Task for Generating Task-Specific World Models Expressed as Text Games

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-16T11:42:30.778162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:42:30.778162Z digest=sha256:31d38ab553c26d733265ecaff67f96e735e17153e1e0d6570cc7d3c1d022a0af

Observation a41e4c9f-9017-4045-b2a3-23ae75f6fa59 · outbound

This paper cites PlanGenLLMs: A Modern Survey of LLM Planning Capabilities.

PLANET: A Collection of Benchmarks for Evaluating LLMs' Planning Capabilities PlanGenLLMs: A Modern Survey of LLM Planning Capabilities

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-16T11:42:30.782744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:42:30.782744Z digest=sha256:ec3180db86686a77715d499f95c3fd5c3361c3121cba588f72ba32ad932ee305

Observation 8e5c67e6-ed17-4d1c-be77-e217a9e33843 · outbound

This paper cites SmartPlay: A Benchmark for LLMs as Intelligent Agents.

PLANET: A Collection of Benchmarks for Evaluating LLMs' Planning Capabilities SmartPlay: A Benchmark for LLMs as Intelligent Agents

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-16T11:42:30.786968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:42:30.786968Z digest=sha256:8aecd2b1d463010aec83b9d0d342c32db719cd01a3896280a5a3c7bb1084b87f

Observation 2d1a69f6-afe7-4a6a-b2a9-260deb31ece3 · outbound

This paper cites Haste Makes Waste: Evaluating Planning Abilities of LLMs for Efficient and Feasible Multitasking with Time Constraints Between Actions.

PLANET: A Collection of Benchmarks for Evaluating LLMs' Planning Capabilities Haste Makes Waste: Evaluating Planning Abilities of LLMs for Efficient and Feasible Multitasking with Time Constraints Between Actions

Reference 54

Resolution
verified exact
local_arxiv, observed 2026-08-16T11:42:31.218150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T11:42:30.791026Z digest=sha256:a1c74cfca18f3e971d8a3800ec43375c193c1fcebba4f198e4eb15684841889d

Observation 877a497b-2c2c-4508-900b-19af58227af9 · outbound

This paper cites AgentGym: Evolving Large Language Model-based Agents across Diverse Environments.

PLANET: A Collection of Benchmarks for Evaluating LLMs' Planning Capabilities AgentGym: Evolving Large Language Model-based Agents across Diverse Environments

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-16T11:42:30.795512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:42:30.795512Z digest=sha256:68a1867820765ea57398b2bec16cc5852352ce5f28933c5687842d1ab7b85f3d

Observation 613c7175-759a-412e-aa26-ba6f541983ef · outbound

This paper cites TravelPlanner: A Benchmark for Real-World Planning with Language Agents.

PLANET: A Collection of Benchmarks for Evaluating LLMs' Planning Capabilities TravelPlanner: A Benchmark for Real-World Planning with Language Agents

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-16T11:42:30.799753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:42:30.799753Z digest=sha256:b79fcf25df66076e66b754c163ae70a5903d1ed8f3bb355e85d2011b8a87e769

Observation cb4d9f63-921a-487d-9260-2e52dce4171e · outbound

This paper cites OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments.

PLANET: A Collection of Benchmarks for Evaluating LLMs' Planning Capabilities OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-16T11:42:30.803851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:42:30.803851Z digest=sha256:32c5585aefd3622584d7b3cba6fe2a6d509992e7cd9ff24006ed2bf1cf27d67c

Observation fd382599-bed7-4f6b-a92c-8a76ead1bad3 · outbound

This paper cites TheAgentCompany: Benchmarking LLM Agents on Consequential Real World Tasks.

PLANET: A Collection of Benchmarks for Evaluating LLMs' Planning Capabilities TheAgentCompany: Benchmarking LLM Agents on Consequential Real World Tasks

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-16T11:42:30.807885Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:42:30.807885Z digest=sha256:3c011827ee13e2af0f5af1c9a019bfc3d4ab9d4912732e34820e896b32f0d8f3

Observation 78ba1a21-c0b7-420e-bce6-d5f59afc08da · outbound

This paper cites Code to Think, Think to Code: A Survey on Code-Enhanced Reasoning and Reasoning-Driven Code Intelligence in LLMs.

PLANET: A Collection of Benchmarks for Evaluating LLMs' Planning Capabilities Code to Think, Think to Code: A Survey on Code-Enhanced Reasoning and Reasoning-Driven Code Intelligence in LLMs

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-16T11:42:30.812017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:42:30.812017Z digest=sha256:63408602d841d299b1b785a17fcbcc8fbd5f609ce4e2676a872115a6a422c953

Observation 9026ff36-f3bd-4e07-9890-70b3a0c8aa81 · outbound

This paper cites EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents.

PLANET: A Collection of Benchmarks for Evaluating LLMs' Planning Capabilities EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-16T11:42:30.816111Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:42:30.816111Z digest=sha256:27dcc07683477f8a8f394c7e42d9a41bb1a300987fbef6e4ff5b7426fd1e7deb

Observation 5791e06e-b1ed-4ce3-a82e-256a228f2f7b · outbound

This paper cites WebShop: Towards Scalable Real-World Web Interaction with Grounded Language Agents.

PLANET: A Collection of Benchmarks for Evaluating LLMs' Planning Capabilities WebShop: Towards Scalable Real-World Web Interaction with Grounded Language Agents

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-16T11:42:30.820259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:42:30.820259Z digest=sha256:7299aaf0434d27deee9a7c966972f51387d7d2c5507b9de48165266e1f4a2ea7

Observation 93bb9eaf-4fdd-4a47-a684-007cfe67a7ac · outbound

This paper cites Tree of Thoughts: Deliberate Problem Solving with Large Language Models.

PLANET: A Collection of Benchmarks for Evaluating LLMs' Planning Capabilities Tree of Thoughts: Deliberate Problem Solving with Large Language Models

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-16T11:42:30.824579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:42:30.824579Z digest=sha256:cde19e8886d8b37abc9b2d801fdd51d6e957f6a1288136ca912f486a26363012

Observation 39454e2b-b2b6-4842-b595-c65d6329bd54 · outbound

This paper cites Safeagentbench: A benchmark for safe task planning of embodied llm agents, 2025.

PLANET: A Collection of Benchmarks for Evaluating LLMs' Planning Capabilities Safeagentbench: A benchmark for safe task planning of embodied llm agents, 2025

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-16T11:42:30.828802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:42:30.828802Z digest=sha256:0c75fe4bb8bc6625bbe89d6ec0dfaaf56404cbc872ab21c24f943fb2adb8cc59

Observation 8f6941aa-c503-45bf-a047-3b1ca55e54de · outbound

This paper cites AssistantBench: Can Web Agents Solve Realistic and Time-Consuming Tasks?.

PLANET: A Collection of Benchmarks for Evaluating LLMs' Planning Capabilities AssistantBench: Can Web Agents Solve Realistic and Time-Consuming Tasks?

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-16T11:42:30.832699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:42:30.832699Z digest=sha256:11646fa449455d6815a7c9a58346058a35ded09c6abfb7581b04cb6c7faf707e

Observation d2cc9214-7e4e-45a2-af0e-8a05d3bb772c · outbound

This paper cites TaskLAMA: Probing the Complex Task Understanding of Language Models.

PLANET: A Collection of Benchmarks for Evaluating LLMs' Planning Capabilities TaskLAMA: Probing the Complex Task Understanding of Language Models

Reference 65

Resolution
verified exact
local_arxiv, observed 2026-08-16T11:42:31.021415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T11:42:30.836841Z digest=sha256:916ec9f10c38c480d0ffe00d5a60c9b5757abfff6bbe02cca8af7d3e2bbd97d5

Observation f4f9304a-ad9b-4ed1-9354-fa45d1a1189a · outbound

This paper cites Distilling Script Knowledge from Large Language Models for Constrained Language Planning.

PLANET: A Collection of Benchmarks for Evaluating LLMs' Planning Capabilities Distilling Script Knowledge from Large Language Models for Constrained Language Planning

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-16T11:42:30.840903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:42:30.840903Z digest=sha256:c13166e04d4fc48978b49ff7dada522aaa7ce45f2d9d4fcdb5653b356c7fedd1

Observation 25ca4793-3531-4e10-9260-ff078c74fba3 · outbound

This paper cites Learning to decompose and organize complex tasks.

PLANET: A Collection of Benchmarks for Evaluating LLMs' Planning Capabilities Learning to decompose and organize complex tasks

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:42:31.882453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T11:42:30.845008Z digest=sha256:8ed98f9c466ddea263e446cbc81505b052d8b94c1041a9658a7b35d551a9c3f9

Observation 4f66b0b4-75bf-49d0-832a-49936e34aa8a · outbound

This paper cites T ime A rena: Shaping efficient multitasking language agents in a time-aware simulation.

PLANET: A Collection of Benchmarks for Evaluating LLMs' Planning Capabilities T ime A rena: Shaping efficient multitasking language agents in a time-aware simulation

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-16T11:42:30.849101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:42:30.849101Z digest=sha256:8aaf65c48603e8a37c2e38d85c3c0bd7b512cdf4ff02552f34a048981f1b712a

Observation 90c3fc94-0feb-4a78-8540-f09d38be8100 · outbound

This paper cites NATURAL PLAN: Benchmarking LLMs on Natural Language Planning.

PLANET: A Collection of Benchmarks for Evaluating LLMs' Planning Capabilities NATURAL PLAN: Benchmarking LLMs on Natural Language Planning

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-16T11:42:30.853368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:42:30.853368Z digest=sha256:cf3098af8f81a5440a7cd1578ab8fb3a0d917c2de45742f0def358d36a525f08

Observation 9fc1a769-5486-4c94-ab34-06c58aa9bd11 · outbound

This paper cites Language Agent Tree Search Unifies Reasoning Acting and Planning in Language Models.

PLANET: A Collection of Benchmarks for Evaluating LLMs' Planning Capabilities Language Agent Tree Search Unifies Reasoning Acting and Planning in Language Models

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-16T11:42:30.857299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:42:30.857299Z digest=sha256:90fd07600c5c7be82bf1204a46f22c1421e542dd62f81770408e4a1f8c026297

Observation 269952f8-09ba-4982-b8e6-7cdcbd61e686 · outbound

This paper cites WebArena: A Realistic Web Environment for Building Autonomous Agents.

PLANET: A Collection of Benchmarks for Evaluating LLMs' Planning Capabilities WebArena: A Realistic Web Environment for Building Autonomous Agents

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-16T11:42:30.861679Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:42:30.861679Z digest=sha256:a6e155354579b978d567074db5e569999de6f1d9fd3dbf53ef51995fec426276

Observation 68d164e5-a1e6-407a-8c32-9b1b54c4621e · outbound

This paper cites @esa (Ref.

PLANET: A Collection of Benchmarks for Evaluating LLMs' Planning Capabilities @esa (Ref

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-16T11:42:30.865939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:42:30.865939Z digest=sha256:ca9f99d5539b94c7a37949557246abb9039cf8364536991998e30b8d5c7e4b43

Observation cb914a45-a361-4faf-b6cc-62f0213c3d43 · outbound

This paper cites an unresolved cited work.

PLANET: A Collection of Benchmarks for Evaluating LLMs' Planning Capabilities Unresolved cited work

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-16T11:42:30.870250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:42:30.870250Z digest=sha256:2a3db59a193e62cf31770d8bf4ff09ab0e3e967f46592da17aeb0649b866a158

Observation a6de5176-4b8d-417c-8ef4-607aa6fe8af9 · outbound

This paper cites an unresolved cited work.

PLANET: A Collection of Benchmarks for Evaluating LLMs' Planning Capabilities Unresolved cited work

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-16T11:42:30.874258Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:42:30.874258Z digest=sha256:398def51ca1c7e112c8297e8f3db08f9889e8f8034a97621795bcb23d4e230e4

Pith citing papers

Observation 950f6c28-c90a-4404-93b6-ad96e4efce4c · inbound

LLM-WikiRace Benchmark: How Far Can LLMs Plan over Real-World Knowledge Graphs? cites this paper.

LLM-WikiRace Benchmark: How Far Can LLMs Plan over Real-World Knowledge Graphs? PLANET: A Collection of Benchmarks for Evaluating LLMs' Planning Capabilities

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-02T22:26:36.732238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T22:26:36.732238Z digest=sha256:4228de3062f83e74dcdf4101591ef37bd6baee63b6c16e178f2a8a9c95ec18b8