Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-16T11:42:30.874258Z
Paper Citation Record · LEDGER
As of 17 August 2026, this Paper Citation Record lists 73 of 73 outbound references and 1 inbound Pith citation observation for arXiv:2504.14773.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-16T11:42:30.874258Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-02T22:26:36.732238Z
A source-named dated measurement, never combined with another source.
Source: cited_works
73 of 73 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 437f4928-54b3-4d1b-bbf3-338c7770a7a8 · outbound
PLANET: A Collection of Benchmarks for Evaluating LLMs' Planning Capabilities write newline
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 800a59d7-ca20-4f31-bdac-1ff3fc66d388 · outbound
PLANET: A Collection of Benchmarks for Evaluating LLMs' Planning Capabilities Can Large Language Models be Good Path Planners? A Benchmark and Investigation on Spatial-temporal Reasoning
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f709a02c-cd01-4374-a0c9-91ac323d5e7f · outbound
PLANET: A Collection of Benchmarks for Evaluating LLMs' Planning Capabilities WorkArena++: Towards Compositional Planning and Reasoning-based Common Knowledge Work Tasks
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 97692f16-a116-4764-a58a-102f8bf982aa · outbound
PLANET: A Collection of Benchmarks for Evaluating LLMs' Planning Capabilities Windows Agent Arena: Evaluating Multi-Modal OS Agents at Scale
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 68bbc0d6-bcd5-4a8f-9e74-dffb0a2f6aed · outbound
PLANET: A Collection of Benchmarks for Evaluating LLMs' Planning Capabilities PlaSma: Making Small Language Models Better Procedural Knowledge Models for (Counterfactual) Planning
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 05665e30-b543-419c-8c3a-a6438eb26726 · outbound
PLANET: A Collection of Benchmarks for Evaluating LLMs' Planning Capabilities PARTNR: A Benchmark for Planning and Reasoning in Embodied Multi-agent Tasks
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ed1af03c-e461-4663-a38e-530ff0dc4c76 · outbound
PLANET: A Collection of Benchmarks for Evaluating LLMs' Planning Capabilities Put Your Money Where Your Mouth Is: Evaluating Strategic Planning and Execution of LLM Agents in an Auction Arena
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b46779be-7e7d-4d18-9cdc-967121141d92 · outbound
PLANET: A Collection of Benchmarks for Evaluating LLMs' Planning Capabilities EgoPlan-Bench: Benchmarking Multimodal Large Language Models for Human-Level Planning
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 86f944a4-8e43-4c60-ae28-b7830bfac183 · outbound
PLANET: A Collection of Benchmarks for Evaluating LLMs' Planning Capabilities Do Large Language Models have Problem-Solving Capability under Incomplete Information Scenarios?
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation b7983657-5c1b-4998-817b-1220bf106969 · outbound
PLANET: A Collection of Benchmarks for Evaluating LLMs' Planning Capabilities Plancraft: an evaluation dataset for planning with LLM agents
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 092da20e-953e-445f-8088-43b099aa739e · outbound
PLANET: A Collection of Benchmarks for Evaluating LLMs' Planning Capabilities Mind2Web: Towards a Generalist Agent for the Web
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 04eed57b-eacd-4bb5-813e-9aa0ca54f0c9 · outbound
PLANET: A Collection of Benchmarks for Evaluating LLMs' Planning Capabilities GTBench: Uncovering the Strategic Reasoning Limitations of LLMs via Game-Theoretic Evaluations
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fda76345-77e4-4320-b51e-1bcb93bba2bf · outbound
PLANET: A Collection of Benchmarks for Evaluating LLMs' Planning Capabilities MineDojo: Building Open-Ended Embodied Agents with Internet-Scale Knowledge
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d1faf3ff-6695-493b-8e3e-5f3445e96331 · outbound
PLANET: A Collection of Benchmarks for Evaluating LLMs' Planning Capabilities REALM-Bench: A Benchmark for Evaluating Multi-Agent Systems on Real-world, Dynamic Planning and Scheduling Tasks
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7759bf5f-4757-4694-a544-cc100b26917a · outbound
PLANET: A Collection of Benchmarks for Evaluating LLMs' Planning Capabilities Robotouille: An Asynchronous Planning Benchmark for LLM Agents
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ed876cd4-b325-4b55-bc97-896ac88ed323 · outbound
PLANET: A Collection of Benchmarks for Evaluating LLMs' Planning Capabilities Unresolved cited work
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 2e143f0e-a9fb-4317-a8c2-e54b4661ceb3 · outbound
PLANET: A Collection of Benchmarks for Evaluating LLMs' Planning Capabilities Benchmarking the Spectrum of Agent Capabilities
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9dbba190-bc17-427e-b823-07caf2aa6e31 · outbound
PLANET: A Collection of Benchmarks for Evaluating LLMs' Planning Capabilities Reasoning with Language Model is Planning with World Model
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation afcca4b6-03c2-4bfc-be1f-dea03a5b5865 · outbound
PLANET: A Collection of Benchmarks for Evaluating LLMs' Planning Capabilities Planning Anything with Rigor: General-Purpose Zero-Shot Planning with LLM-based Formalized Programming
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9b0de9c2-5572-4871-b4f5-5653cda51224 · outbound
PLANET: A Collection of Benchmarks for Evaluating LLMs' Planning Capabilities Hoffmann and S
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5acfda5a-ea43-4cd1-ad6c-79bdee617df5 · outbound
PLANET: A Collection of Benchmarks for Evaluating LLMs' Planning Capabilities Game-theoretic LLM: Agent Workflow for Negotiation Games
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation becebb25-ba08-416d-8851-3b465549bbc4 · outbound
PLANET: A Collection of Benchmarks for Evaluating LLMs' Planning Capabilities How Far Are We on the Decision-Making of LLMs? Evaluating LLMs' Gaming Ability in Multi-Agent Environments
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5eac89ff-f8ae-49e7-8d6b-c343b7e19715 · outbound
PLANET: A Collection of Benchmarks for Evaluating LLMs' Planning Capabilities CRMArena: Understanding the Capacity of LLM Agents to Perform Professional CRM Tasks in Realistic Environments
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9dc7b07c-eb70-4f03-a556-6dc64688e083 · outbound
PLANET: A Collection of Benchmarks for Evaluating LLMs' Planning Capabilities Understanding the planning of LLM agents: A survey
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation caccdbe1-0d3e-4b59-b5fd-97c5c77d6fbd · outbound
PLANET: A Collection of Benchmarks for Evaluating LLMs' Planning Capabilities McNamara, and Deming Chen
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4b81b7d5-62d8-4374-8a41-03c8649e1a1b · outbound
PLANET: A Collection of Benchmarks for Evaluating LLMs' Planning Capabilities SWE-bench: Can Language Models Resolve Real-World GitHub Issues?
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9f602155-bf47-426c-8f98-f3ba1cad0d82 · outbound
PLANET: A Collection of Benchmarks for Evaluating LLMs' Planning Capabilities To the Globe (TTG): Towards Language-Driven Guaranteed Travel Planning
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 00bbfa85-def3-45a1-b19c-575aecf3ae68 · outbound
PLANET: A Collection of Benchmarks for Evaluating LLMs' Planning Capabilities Towards a foundation for evaluating ai planners
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation de3cb5ef-3fa9-44f9-8d63-15a385916f7e · outbound
PLANET: A Collection of Benchmarks for Evaluating LLMs' Planning Capabilities VisualWebArena: Evaluating Multimodal Agents on Realistic Visual Web Tasks
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation effb16c8-f817-4a4c-8669-f96e58f0fb2d · outbound
PLANET: A Collection of Benchmarks for Evaluating LLMs' Planning Capabilities Beyond A*: Better Planning with Transformers via Search Dynamics Bootstrapping
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8d8762b8-1d70-4513-b343-7e784854a888 · outbound
PLANET: A Collection of Benchmarks for Evaluating LLMs' Planning Capabilities LASP: Surveying the State-of-the-Art in Large Language Model-Assisted AI Planning
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ad2e18da-06fb-4a7d-9881-341f6efbf7d9 · outbound
PLANET: A Collection of Benchmarks for Evaluating LLMs' Planning Capabilities AgentBench: Evaluating LLMs as Agents
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2342f4ae-9a70-4c27-aa32-a65f467b1c02 · outbound
PLANET: A Collection of Benchmarks for Evaluating LLMs' Planning Capabilities AgentBoard: An Analytical Evaluation Board of Multi-turn LLM Agents
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bd41abc8-7cfd-40e0-9b37-0378cec5168e · outbound
PLANET: A Collection of Benchmarks for Evaluating LLMs' Planning Capabilities GAIA: a benchmark for General AI Assistants
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ef09f6f1-af60-48e0-b07e-59c5772b4654 · outbound
PLANET: A Collection of Benchmarks for Evaluating LLMs' Planning Capabilities LLaMAR: Long-Horizon Planning for Multi-Agent Robots in Partially Observable Environments
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 27054b1d-5257-4aca-b532-c1dd74c79b9d · outbound
PLANET: A Collection of Benchmarks for Evaluating LLMs' Planning Capabilities WorldAPIs: The World Is Worth How Many APIs? A Thought Experiment
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 02e6f43b-b23f-4025-b65a-6f953d5a0063 · outbound
PLANET: A Collection of Benchmarks for Evaluating LLMs' Planning Capabilities TEACh: Task-driven Embodied Agents that Chat
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1270d246-03f5-49a0-971f-c7c87a248c62 · outbound
PLANET: A Collection of Benchmarks for Evaluating LLMs' Planning Capabilities VirtualHome: Simulating Household Activities via Programs
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4df2c259-59b7-45bc-8d13-36a89e34f829 · outbound
PLANET: A Collection of Benchmarks for Evaluating LLMs' Planning Capabilities Artificial I ntelligence: A modern approach
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation bcccf922-b9f6-4e6f-85b1-70b95ae322d1 · outbound
PLANET: A Collection of Benchmarks for Evaluating LLMs' Planning Capabilities Unresolved cited work
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fe945321-7320-42a0-9d3b-304a3af33e46 · outbound
PLANET: A Collection of Benchmarks for Evaluating LLMs' Planning Capabilities Language Models Are Greedy Reasoners: A Systematic Formal Analysis of Chain-of-Thought
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0ac9f86e-7e86-43ca-871f-adcbb3cef6f7 · outbound
PLANET: A Collection of Benchmarks for Evaluating LLMs' Planning Capabilities Reflexion: Language agents with verbal reinforcement learning
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1cf8db43-aeec-4025-be1f-3cd0991f6f64 · outbound
PLANET: A Collection of Benchmarks for Evaluating LLMs' Planning Capabilities ALFRED: A Benchmark for Interpreting Grounded Instructions for Everyday Tasks
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f079fb56-3404-4631-b8de-a847da033a08 · outbound
PLANET: A Collection of Benchmarks for Evaluating LLMs' Planning Capabilities ALFWorld: Aligning Text and Embodied Environments for Interactive Learning
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eb34491c-4abb-4cea-8aa6-c7fc90b09bdf · outbound
PLANET: A Collection of Benchmarks for Evaluating LLMs' Planning Capabilities Dualformer: Controllable Fast and Slow Thinking by Learning with Randomized Reasoning Traces
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 22f22cfb-abaf-4fd9-bac9-82bbb914c959 · outbound
PLANET: A Collection of Benchmarks for Evaluating LLMs' Planning Capabilities A ct P lan-1 K : Benchmarking the procedural planning ability of visual language models in household activities
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation af7c1c86-79ce-42d9-b11a-a338691e9d9e · outbound
PLANET: A Collection of Benchmarks for Evaluating LLMs' Planning Capabilities Stop Overthinking: A Survey on Efficient Reasoning for Large Language Models
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b289fbf6-e96b-4180-87ea-a1f0e017baae · outbound
PLANET: A Collection of Benchmarks for Evaluating LLMs' Planning Capabilities Planbench: An extensible benchmark for evaluating large language models on planning and reasoning about change
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation e19cb867-08ee-43b7-82b8-c8e95f7d22aa · outbound
PLANET: A Collection of Benchmarks for Evaluating LLMs' Planning Capabilities TMGBench: A Systematic Game Benchmark for Evaluating Strategic Reasoning Abilities of LLMs
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 610b3a32-17dd-4587-ba1b-dca90f4805c7 · outbound
PLANET: A Collection of Benchmarks for Evaluating LLMs' Planning Capabilities ByteSized32: A Corpus and Challenge Task for Generating Task-Specific World Models Expressed as Text Games
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a41e4c9f-9017-4045-b2a3-23ae75f6fa59 · outbound
PLANET: A Collection of Benchmarks for Evaluating LLMs' Planning Capabilities PlanGenLLMs: A Modern Survey of LLM Planning Capabilities
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8e5c67e6-ed17-4d1c-be77-e217a9e33843 · outbound
PLANET: A Collection of Benchmarks for Evaluating LLMs' Planning Capabilities SmartPlay: A Benchmark for LLMs as Intelligent Agents
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2d1a69f6-afe7-4a6a-b2a9-260deb31ece3 · outbound
PLANET: A Collection of Benchmarks for Evaluating LLMs' Planning Capabilities Haste Makes Waste: Evaluating Planning Abilities of LLMs for Efficient and Feasible Multitasking with Time Constraints Between Actions
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 877a497b-2c2c-4508-900b-19af58227af9 · outbound
PLANET: A Collection of Benchmarks for Evaluating LLMs' Planning Capabilities AgentGym: Evolving Large Language Model-based Agents across Diverse Environments
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 613c7175-759a-412e-aa26-ba6f541983ef · outbound
PLANET: A Collection of Benchmarks for Evaluating LLMs' Planning Capabilities TravelPlanner: A Benchmark for Real-World Planning with Language Agents
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cb4d9f63-921a-487d-9260-2e52dce4171e · outbound
PLANET: A Collection of Benchmarks for Evaluating LLMs' Planning Capabilities OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fd382599-bed7-4f6b-a92c-8a76ead1bad3 · outbound
PLANET: A Collection of Benchmarks for Evaluating LLMs' Planning Capabilities TheAgentCompany: Benchmarking LLM Agents on Consequential Real World Tasks
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 78ba1a21-c0b7-420e-bce6-d5f59afc08da · outbound
PLANET: A Collection of Benchmarks for Evaluating LLMs' Planning Capabilities Code to Think, Think to Code: A Survey on Code-Enhanced Reasoning and Reasoning-Driven Code Intelligence in LLMs
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9026ff36-f3bd-4e07-9890-70b3a0c8aa81 · outbound
PLANET: A Collection of Benchmarks for Evaluating LLMs' Planning Capabilities EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5791e06e-b1ed-4ce3-a82e-256a228f2f7b · outbound
PLANET: A Collection of Benchmarks for Evaluating LLMs' Planning Capabilities WebShop: Towards Scalable Real-World Web Interaction with Grounded Language Agents
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 93bb9eaf-4fdd-4a47-a684-007cfe67a7ac · outbound
PLANET: A Collection of Benchmarks for Evaluating LLMs' Planning Capabilities Tree of Thoughts: Deliberate Problem Solving with Large Language Models
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 39454e2b-b2b6-4842-b595-c65d6329bd54 · outbound
PLANET: A Collection of Benchmarks for Evaluating LLMs' Planning Capabilities Safeagentbench: A benchmark for safe task planning of embodied llm agents, 2025
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8f6941aa-c503-45bf-a047-3b1ca55e54de · outbound
PLANET: A Collection of Benchmarks for Evaluating LLMs' Planning Capabilities AssistantBench: Can Web Agents Solve Realistic and Time-Consuming Tasks?
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d2cc9214-7e4e-45a2-af0e-8a05d3bb772c · outbound
PLANET: A Collection of Benchmarks for Evaluating LLMs' Planning Capabilities TaskLAMA: Probing the Complex Task Understanding of Language Models
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation f4f9304a-ad9b-4ed1-9354-fa45d1a1189a · outbound
PLANET: A Collection of Benchmarks for Evaluating LLMs' Planning Capabilities Distilling Script Knowledge from Large Language Models for Constrained Language Planning
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 25ca4793-3531-4e10-9260-ff078c74fba3 · outbound
PLANET: A Collection of Benchmarks for Evaluating LLMs' Planning Capabilities Learning to decompose and organize complex tasks
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 4f66b0b4-75bf-49d0-832a-49936e34aa8a · outbound
PLANET: A Collection of Benchmarks for Evaluating LLMs' Planning Capabilities T ime A rena: Shaping efficient multitasking language agents in a time-aware simulation
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 90c3fc94-0feb-4a78-8540-f09d38be8100 · outbound
PLANET: A Collection of Benchmarks for Evaluating LLMs' Planning Capabilities NATURAL PLAN: Benchmarking LLMs on Natural Language Planning
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9fc1a769-5486-4c94-ab34-06c58aa9bd11 · outbound
PLANET: A Collection of Benchmarks for Evaluating LLMs' Planning Capabilities Language Agent Tree Search Unifies Reasoning Acting and Planning in Language Models
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 269952f8-09ba-4982-b8e6-7cdcbd61e686 · outbound
PLANET: A Collection of Benchmarks for Evaluating LLMs' Planning Capabilities WebArena: A Realistic Web Environment for Building Autonomous Agents
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 68d164e5-a1e6-407a-8c32-9b1b54c4621e · outbound
PLANET: A Collection of Benchmarks for Evaluating LLMs' Planning Capabilities @esa (Ref
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cb914a45-a361-4faf-b6cc-62f0213c3d43 · outbound
PLANET: A Collection of Benchmarks for Evaluating LLMs' Planning Capabilities Unresolved cited work
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a6de5176-4b8d-417c-8ef4-607aa6fe8af9 · outbound
PLANET: A Collection of Benchmarks for Evaluating LLMs' Planning Capabilities Unresolved cited work
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 950f6c28-c90a-4404-93b6-ad96e4efce4c · inbound
LLM-WikiRace Benchmark: How Far Can LLMs Plan over Real-World Knowledge Graphs? PLANET: A Collection of Benchmarks for Evaluating LLMs' Planning Capabilities
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.