Pith. sign in

Paper Citation Record · LEDGER

Plancraft: an evaluation dataset for planning with LLM agents

As of 11 August 2026, this Paper Citation Record lists 50 of 50 outbound references and 9 inbound Pith citation observations for arXiv:2412.21033.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.21033 v2

Coverage vector

measured 50 of 50 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T23:21:35.262071Z

measured 59 of 59 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 9 of 9 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:11:54.275568Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-12T05:51:24.590137Z

Reference resolution

50 of 50 outbound references displayed

  • verified exact1
  • verified fuzzy11
  • unresolved38
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c6899f4e-af3b-4b1c-9a30-0f0ebda2d7fd · outbound

This paper cites Qwen2.5-VL Technical Report.

Plancraft: an evaluation dataset for planning with LLM agents Qwen2.5-VL Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:34.962279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:34.962279Z digest=sha256:baedda048f2ae15ccf7e22d3fdf2634fb704fc28f16d7afe0503910667c8dc26

Observation 83a6e7a3-bdb3-4633-8b20-6b3202ab33c4 · outbound

This paper cites Video pretraining (vpt): Learning to act by watching unlabeled online videos.

Plancraft: an evaluation dataset for planning with LLM agents Video pretraining (vpt): Learning to act by watching unlabeled online videos

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:21:36.127871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T23:21:34.968716Z digest=sha256:e693b3a55ef86ce851fca53b6fcfe8bcd353d1c48b62da53e68d825bfeff547c

Observation ffdfbd1e-9f5e-424c-ae64-f3751cd43797 · outbound

This paper cites Baby AI : First steps towards grounded language learning with a human in the loop.

Plancraft: an evaluation dataset for planning with LLM agents Baby AI : First steps towards grounded language learning with a human in the loop

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:21:36.110266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T23:21:34.974536Z digest=sha256:0219a4fd52f7c8bbf9ec2bb6bfc83dfecd7a408e59a648fa3ceb1ffa46526c23

Observation cb9af781-5bdc-42d0-b432-8c2ee8eca8ab · outbound

This paper cites Mind2web: Towards a generalist agent for the web, 2023.

Plancraft: an evaluation dataset for planning with LLM agents Mind2web: Towards a generalist agent for the web, 2023

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:34.980883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:34.980883Z digest=sha256:f304969c58bb143d302a5cc85429371c415a95b0a71d146ec8ff31ebf1e5d5c2

Observation bcaafdb6-5669-4b4d-b95e-669cdd2d95a4 · outbound

This paper cites The Llama 3 Herd of Models.

Plancraft: an evaluation dataset for planning with LLM agents The Llama 3 Herd of Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:34.987102Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:34.987102Z digest=sha256:ec4a03bff5df867526fd4ed94a95682cdc864b5163d21d59901d04d0cc1fbdb2

Observation 95ddbc39-c668-4e58-a37b-218ed7e79878 · outbound

This paper cites Minedojo: Building open-ended embodied agents with internet-scale knowledge.

Plancraft: an evaluation dataset for planning with LLM agents Minedojo: Building open-ended embodied agents with internet-scale knowledge

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:34.992805Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:34.992805Z digest=sha256:30e340b9700653f1b213734edbe2a93693fb5a53880c09308c202dac5d1da280

Observation 38d2ea49-d04b-4549-9ab2-085fda3ae948 · outbound

This paper cites MineRL: A Large-Scale Dataset of Minecraft Demonstrations.

Plancraft: an evaluation dataset for planning with LLM agents MineRL: A Large-Scale Dataset of Minecraft Demonstrations

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:34.998601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:34.998601Z digest=sha256:9d130e636ef8e133341c8c7199ea470cc9c8333445b241084210d7d2e611bd0d

Observation 9ef094eb-3095-4944-b64a-279ead6ec346 · outbound

This paper cites Exploring the capacity of pretrained language models for reasoning about actions and change.

Plancraft: an evaluation dataset for planning with LLM agents Exploring the capacity of pretrained language models for reasoning about actions and change

Reference 8

Resolution
verified exact
doi, observed 2026-08-10T23:21:35.379056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T23:21:35.004691Z digest=sha256:f988951503a6be3af6199bc0ab09ce9a5ca5cf8b83cca0612738b3d8eb6f9110

Observation 69b3124f-1bbe-4667-b15a-6e8e7662b8c8 · outbound

This paper cites The fast downward planning system.

Plancraft: an evaluation dataset for planning with LLM agents The fast downward planning system

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:35.011784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:35.011784Z digest=sha256:f360746fe3b4faa6153321044f6490ce692308af9608ca0c051176c928b373be

Observation 8db0f65c-f67f-42e7-b71b-22a5e3cfb518 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

Plancraft: an evaluation dataset for planning with LLM agents LoRA: Low-Rank Adaptation of Large Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:35.020699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:35.020699Z digest=sha256:eb7937dcb5182b7095185c8b314afa1e950190424bf4ee050739294e4e6b3361

Observation 339902b5-51f6-4cc1-b3e5-ead8166a9eaa · outbound

This paper cites Inner Monologue: Embodied Reasoning through Planning with Language Models.

Plancraft: an evaluation dataset for planning with LLM agents Inner Monologue: Embodied Reasoning through Planning with Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:35.026785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:35.026785Z digest=sha256:0c307f6770268ed8a0f10e72d42fda16b95b1ff2669ca4255b4828433ee69932

Observation f78bec82-42db-4698-88be-c9715c2f0f62 · outbound

This paper cites The malmo platform for artificial intelligence experimentation.

Plancraft: an evaluation dataset for planning with LLM agents The malmo platform for artificial intelligence experimentation

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:35.033683Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:35.033683Z digest=sha256:22988d246911d39428fad3c732bdf7bce85cf533344e48af27df43d3a46c4841

Observation 334b2164-3887-47e6-86c4-be78344f1904 · outbound

This paper cites Position: LLM s can’t plan, but can help planning in LLM -modulo frameworks.

Plancraft: an evaluation dataset for planning with LLM agents Position: LLM s can’t plan, but can help planning in LLM -modulo frameworks

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:21:36.043705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T23:21:35.039172Z digest=sha256:6d7dcd326771fe435f03eae38c29258f69080e6acf798a787121f55e21fc1c84

Observation 6d9561c3-a4ab-45b1-a64c-9e52a305d5e1 · outbound

This paper cites ACPB ench hard: Unrestrained reasoning about action, change, and planning.

Plancraft: an evaluation dataset for planning with LLM agents ACPB ench hard: Unrestrained reasoning about action, change, and planning

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:21:36.022450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T23:21:35.044518Z digest=sha256:5c56abebbbcad24e31277fc6a82c5d58f063518803b1bbd876a8acb59bb2da57

Observation 078b8185-56bd-4951-89f5-cc1b9010685b · outbound

This paper cites Retrieval-augmented generation for knowledge-intensive nlp tasks.

Plancraft: an evaluation dataset for planning with LLM agents Retrieval-augmented generation for knowledge-intensive nlp tasks

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:21:36.004799Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T23:21:35.050085Z digest=sha256:95ce4d617447fa42bd70eb8da4f16e9e70e8cfb1faecba526cdefb3774a39f4a

Observation c072f7f4-658f-4ecb-99b7-25cf3fda2706 · outbound

This paper cites Benchmarking Detection Transfer Learning with Vision Transformers.

Plancraft: an evaluation dataset for planning with LLM agents Benchmarking Detection Transfer Learning with Vision Transformers

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:35.055671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:35.055671Z digest=sha256:f6a82b978fa83b0810a8abbab659d8d84e442d92855d410807e4c9ba4139606d

Observation 555f29c1-caf5-44a2-a03e-2c87edf7aac0 · outbound

This paper cites Reinforcement Learning on Web Interfaces Using Workflow-Guided Exploration.

Plancraft: an evaluation dataset for planning with LLM agents Reinforcement Learning on Web Interfaces Using Workflow-Guided Exploration

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:35.061524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:35.061524Z digest=sha256:9046a72ae34487949bfc317cc545c90074f3e696543d07b5ceca824113235dbc

Observation a246f7b0-38bf-4ab1-b99a-10f34eb346cf · outbound

This paper cites AgentBench: Evaluating LLMs as Agents.

Plancraft: an evaluation dataset for planning with LLM agents AgentBench: Evaluating LLMs as Agents

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:35.068214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:35.068214Z digest=sha256:deeff01a5acf6a5d1fa32eeca33a8a5a09933f3b6c49dc07353c3bfe538a56ec

Observation 962a9973-7623-474f-9fe5-6d44f7cd8606 · outbound

This paper cites Howe, Craig A.

Plancraft: an evaluation dataset for planning with LLM agents Howe, Craig A

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:35.073575Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:35.073575Z digest=sha256:9144493b820aaa372b8221d1650df7297436eced089b9cb848438af8d5970192

Observation baf260aa-6547-40dc-8624-10217e91f32f · outbound

This paper cites GAIA: a benchmark for General AI Assistants.

Plancraft: an evaluation dataset for planning with LLM agents GAIA: a benchmark for General AI Assistants

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:35.079301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:35.079301Z digest=sha256:882b435a0f8ea796ff375fb1b6a55a92cebe61155f4f8dfc47faf8ad45606553

Observation 50d5c23b-1742-4ff0-a903-716ee56be25e · outbound

This paper cites WebGPT: Browser-assisted question-answering with human feedback.

Plancraft: an evaluation dataset for planning with LLM agents WebGPT: Browser-assisted question-answering with human feedback

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:35.085746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:35.085746Z digest=sha256:4d78c89b644a19bf517d7aa34c8cba5938f1827ee0c788fd5f2ca9dd5dd7fe22

Observation bdb18fd8-7ebc-4a8c-a102-06938619207b · outbound

This paper cites GPT-4 Technical Report.

Plancraft: an evaluation dataset for planning with LLM agents GPT-4 Technical Report

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:35.096695Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:35.096695Z digest=sha256:04f28e6079d9be8b0037544638838fa41eab452c6e16e3244e75546f845b4fde

Observation c2ca0f0e-10e2-4c33-9437-895551410b4e · outbound

This paper cites TALM: Tool Augmented Language Models.

Plancraft: an evaluation dataset for planning with LLM agents TALM: Tool Augmented Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:35.103520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:35.103520Z digest=sha256:ab19777e14c03e9466d0fbc460cdb91cb806ce7d71fa8392e05122b6ce8a3eb5

Observation 2865db62-cda1-4847-8197-ae4521bf585c · outbound

This paper cites Gorilla: Large Language Model Connected with Massive APIs.

Plancraft: an evaluation dataset for planning with LLM agents Gorilla: Large Language Model Connected with Massive APIs

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:35.109381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:35.109381Z digest=sha256:63ea80159f6d6a6ca0566dcad233eff32dbc8b6320c464720c20c3cc0678d6e8

Observation c53212ee-6663-47e8-883b-bfc414e84533 · outbound

This paper cites The Multi-Agent Reinforcement Learning in Malm\"O (MARL\"O) Competition.

Plancraft: an evaluation dataset for planning with LLM agents The Multi-Agent Reinforcement Learning in Malm\"O (MARL\"O) Competition

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:35.114523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:35.114523Z digest=sha256:923dd550894e0cd0a6e3afc292599b69779032dcbac29c3ab2b8e511a222d088

Observation 82ceb89d-cca3-41c6-a076-5178223b2734 · outbound

This paper cites ADaPT: As-Needed Decomposition and Planning with Language Models.

Plancraft: an evaluation dataset for planning with LLM agents ADaPT: As-Needed Decomposition and Planning with Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:35.120847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:35.120847Z digest=sha256:1e8b70af41fcd581fd0b46d47380fc89f1e7f2eb7b1f3c307f128731eec423d9

Observation f1152107-ef53-4331-a093-fcc67dabf01a · outbound

This paper cites Virtualhome: Simulating household activities via programs.

Plancraft: an evaluation dataset for planning with LLM agents Virtualhome: Simulating household activities via programs

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:35.126921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:35.126921Z digest=sha256:9d1ae678e7055546858cace3bdb19e3dd7a4d33cf9ab8b98fcfb326b31781b50

Observation 2aff9efb-81e9-47f5-8a39-58823c8b1925 · outbound

This paper cites Toolformer: Language models can teach themselves to use tools.

Plancraft: an evaluation dataset for planning with LLM agents Toolformer: Language models can teach themselves to use tools

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:21:35.964322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T23:21:35.133256Z digest=sha256:9133e6c4275dc0e84c0bd58891e98a412b145c64ac164d962f9495de7b4639b4

Observation 116c7831-baaf-4084-a3dd-af5a46b4c68a · outbound

This paper cites Narasimhan, and Shunyu Yao.

Plancraft: an evaluation dataset for planning with LLM agents Narasimhan, and Shunyu Yao

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:21:35.946960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T23:21:35.139959Z digest=sha256:781b32120e5aff8007c76116df55d752f4471c63302fb6a09c71c8a566258353

Observation c12d3bed-6c15-4c79-a0ec-4a930b70fe5d · outbound

This paper cites ALFRED: A Benchmark for Interpreting Grounded Instructions for Everyday Tasks.

Plancraft: an evaluation dataset for planning with LLM agents ALFRED: A Benchmark for Interpreting Grounded Instructions for Everyday Tasks

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:35.147072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:35.147072Z digest=sha256:b6967e1dea325040fc1d565b5619dc3331e041aeae20959f85704b2d560ad753

Observation 9cbccd52-50dc-427f-916f-7ec7d2b8b666 · outbound

This paper cites ALFWorld: Aligning Text and Embodied Environments for Interactive Learning.

Plancraft: an evaluation dataset for planning with LLM agents ALFWorld: Aligning Text and Embodied Environments for Interactive Learning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:35.152677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:35.152677Z digest=sha256:5da5c539db6c04101670f45aedb82500dd2f4601e7b1014d92e561faf046b5a2

Observation 055f60ee-d2ce-4e03-9aa1-de7187c82301 · outbound

This paper cites Gemma 3 Technical Report.

Plancraft: an evaluation dataset for planning with LLM agents Gemma 3 Technical Report

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:35.158607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:35.158607Z digest=sha256:20ea324dd07464bd4ec8db34ca885f0810b4a8b67e30f8736a60ef3ae53a1950

Observation 081fc9bc-764e-4035-9d14-32dde4d12d35 · outbound

This paper cites Planbench: An extensible benchmark for evaluating large language models on planning and reasoning about change.

Plancraft: an evaluation dataset for planning with LLM agents Planbench: An extensible benchmark for evaluating large language models on planning and reasoning about change

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:21:35.926300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T23:21:35.165118Z digest=sha256:42188bd95152184ecec3685a6202aa0d83049f399a710ffc32b85a3ef84fb6d8

Observation 91aa0402-1940-4531-8be3-14587ccb4094 · outbound

This paper cites Voyager: An Open-Ended Embodied Agent with Large Language Models.

Plancraft: an evaluation dataset for planning with LLM agents Voyager: An Open-Ended Embodied Agent with Large Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:35.170924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:35.170924Z digest=sha256:25f2d12aa43d06b155fe14ebc1924c1482cdf89a6f4c6e9fe6138ea8fbee7b98

Observation a4f0be2a-6056-470e-9f2c-4e36370ee0a7 · outbound

This paper cites ScienceWorld: Is your Agent Smarter than a 5th Grader?.

Plancraft: an evaluation dataset for planning with LLM agents ScienceWorld: Is your Agent Smarter than a 5th Grader?

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:35.176667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:35.176667Z digest=sha256:53bbd2cdf667f52d9b2bba15d669bcd8341d12a945c29acf32392327ffd76820

Observation da4e836e-958f-492f-ac7b-beeb7ee13d4f · outbound

This paper cites Le, Ed H.

Plancraft: an evaluation dataset for planning with LLM agents Le, Ed H

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:21:35.908331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T23:21:35.182272Z digest=sha256:f01ec3aec7ee4a2ab59066527d875396a4c7b5d601c785f856b4647ea122639f

Observation f02204f5-6697-41a8-9983-636805a31499 · outbound

This paper cites JARVIS-1: Open-World Multi-task Agents with Memory-Augmented Multimodal Language Models.

Plancraft: an evaluation dataset for planning with LLM agents JARVIS-1: Open-World Multi-task Agents with Memory-Augmented Multimodal Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:35.187048Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:35.187048Z digest=sha256:17804411a2212ac566ccc9a7b4d8d235f7ccbbad9a19fa64deb98c0790ac39de

Observation a9b14b8f-aa06-4a74-8cd8-af1181c01011 · outbound

This paper cites Describe, Explain, Plan and Select: Interactive Planning with Large Language Models Enables Open-World Multi-Task Agents.

Plancraft: an evaluation dataset for planning with LLM agents Describe, Explain, Plan and Select: Interactive Planning with Large Language Models Enables Open-World Multi-Task Agents

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:35.192867Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:35.192867Z digest=sha256:34f21a6b6b5b4bc6f147aed09be2268fc59610313bf0f0055805f116d101e7cd

Observation 46b5cc90-0735-4706-a9e7-014ee022779f · outbound

This paper cites The Rise and Potential of Large Language Model Based Agents: A Survey.

Plancraft: an evaluation dataset for planning with LLM agents The Rise and Potential of Large Language Model Based Agents: A Survey

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:35.198539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:35.198539Z digest=sha256:f44a63bcc0a53337e60b9c26a8faa58d98020e24e1f21de3532878989bc208d2

Observation ad43041a-4c60-44c8-a1c1-03b4c1f24634 · outbound

This paper cites Agentgym: Evolving large language model-based agents across diverse environments, 2024.

Plancraft: an evaluation dataset for planning with LLM agents Agentgym: Evolving large language model-based agents across diverse environments, 2024

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:35.203606Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:35.203606Z digest=sha256:2f578d1cf5ca98b6c2fea3920faa940e7af4fb908783488f5509fbcbff98b57c

Observation 622c2cef-fc5a-47e2-8f75-c69add1c58cb · outbound

This paper cites T ravel P lanner: A benchmark for real-world planning with language agents.

Plancraft: an evaluation dataset for planning with LLM agents T ravel P lanner: A benchmark for real-world planning with language agents

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:21:35.876732Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T23:21:35.209907Z digest=sha256:04e6e4116dda652b9a28ed34ab5ce3c4e9a575a18dcd45896cd5e9c369f19989

Observation ab05e995-3cba-4319-ad51-5ce72c237bac · outbound

This paper cites Qwen3 Technical Report.

Plancraft: an evaluation dataset for planning with LLM agents Qwen3 Technical Report

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:35.214636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:35.214636Z digest=sha256:71c83d65521a98fb3eb260fb50a51b1ee257f113370804d9c6f6b7f767b0837c

Observation 4a114be2-1e77-43af-a2bc-2cda882c73f4 · outbound

This paper cites GPT4Tools: Teaching Large Language Model to Use Tools via Self-instruction.

Plancraft: an evaluation dataset for planning with LLM agents GPT4Tools: Teaching Large Language Model to Use Tools via Self-instruction

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:35.221340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:35.221340Z digest=sha256:01091fc24def0b0b8bd2aaf6300b4f6d6f8280a9a3e5f88e7508b420830df069

Observation 743ffc22-c821-4684-8ded-a37454626c15 · outbound

This paper cites Webshop: Towards scalable real-world web interaction with grounded language agents.

Plancraft: an evaluation dataset for planning with LLM agents Webshop: Towards scalable real-world web interaction with grounded language agents

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:35.227225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:35.227225Z digest=sha256:735444365d4a9885cd797685e44b47432746a1b3cb42501d0dfeab4a78fc4d12

Observation cd52a878-8b99-4885-82ed-2115bc8ae527 · outbound

This paper cites Tree of thoughts: Deliberate problem solving with large language models.

Plancraft: an evaluation dataset for planning with LLM agents Tree of thoughts: Deliberate problem solving with large language models

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:21:35.838013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T23:21:35.232695Z digest=sha256:4db4eb051449bd3023f56786bbe1c72d7a6b24e19355768c79863c74449285f7

Observation 9fd08eaf-e80b-4421-b36c-ec56d382f42e · outbound

This paper cites ReAct : Synergizing reasoning and acting in language models.

Plancraft: an evaluation dataset for planning with LLM agents ReAct : Synergizing reasoning and acting in language models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:35.238152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:35.238152Z digest=sha256:fdc65a317177a0cb7c6519b7663b845cba74e0bfd5323d020d463ec6b8e9a5da

Observation 284316ec-9da3-40f3-bcb9-955f90fb8d78 · outbound

This paper cites WebArena: A Realistic Web Environment for Building Autonomous Agents.

Plancraft: an evaluation dataset for planning with LLM agents WebArena: A Realistic Web Environment for Building Autonomous Agents

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:35.243148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:35.243148Z digest=sha256:b03f284c6eb621e60bd46e8ddded2496b0e56d3d9c6a6cb7aac9415016e6c902

Observation 591c7327-c25b-4752-87d7-325f51abd00a · outbound

This paper cites @esa (Ref.

Plancraft: an evaluation dataset for planning with LLM agents @esa (Ref

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:35.248761Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:35.248761Z digest=sha256:bc879061b336fa8b798c7b38203fa4b1d5b4c7a2d02b0a3966b076a6a9497820

Observation b796486c-0b2f-4657-948b-394d834a896a · outbound

This paper cites an unresolved cited work.

Plancraft: an evaluation dataset for planning with LLM agents Unresolved cited work

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:35.255755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:35.255755Z digest=sha256:4a4218a665aee50db36421714ad0bddac5a9fd7d1ee3f2f3caf6a205902e5246

Observation 96bc7824-9550-4c8d-b846-9b683a9cd7a0 · outbound

This paper cites an unresolved cited work.

Plancraft: an evaluation dataset for planning with LLM agents Unresolved cited work

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:35.262071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:35.262071Z digest=sha256:5e126261c9f75d036bf600d20904cda05c2c4b4e10129bfa402196724bb74906

Pith citing papers

Observation abff24fd-b60f-419c-a4d9-d0d3e9d60792 · inbound

Large Language Models for Planning: A Comprehensive and Systematic Survey cites this paper.

Large Language Models for Planning: A Comprehensive and Systematic Survey Plancraft: an evaluation dataset for planning with LLM agents

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T14:11:54.275568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:11:54.275568Z digest=sha256:d093184550c574dd1f6f99dbda6c42bb1fa2e9eb85e030951fa00e017de5defe

Observation ff96bea1-04fc-4465-bc89-20013780566c · inbound

Memory in the Age of AI Agents cites this paper.

Memory in the Age of AI Agents Plancraft: an evaluation dataset for planning with LLM agents

Reference 294

Resolution
verified exact
arxiv_id, observed 2026-05-11T18:18:20.482073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-11T18:18:19.911342Z digest=sha256:e17564ee8200b453eadecb63b5ac92996b23a87f6baaf8a87254630b348d9427

Observation c672ca42-ac22-4414-974b-925c85ff5c05 · inbound

LLM-WikiRace Benchmark: How Far Can LLMs Plan over Real-World Knowledge Graphs? cites this paper.

LLM-WikiRace Benchmark: How Far Can LLMs Plan over Real-World Knowledge Graphs? Plancraft: an evaluation dataset for planning with LLM agents

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-02T22:26:35.019522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T22:26:35.019522Z digest=sha256:dd2a4892e4ed3d341b4510b95fec42f69f309abbaac76b87768db5a9762b7f0e

Observation bf88a171-f6f1-45dc-8b52-bb38ba31f581 · inbound

Mastering PokeGym: Graph-Guided Multimodal Evolution at Test Time cites this paper.

Mastering PokeGym: Graph-Guided Multimodal Evolution at Test Time Plancraft: an evaluation dataset for planning with LLM agents

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:41:00.539463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T17:59:48.877783Z digest=sha256:9111794a2ac4f4d6f601aa33c403491e4d9f525acc7efe92fb5d0db71ba957e8

Observation cdc6800e-bb63-4bf2-a048-6976b23682f7 · inbound

Mastering PokeGym: Graph-Guided Multimodal Evolution at Test Time cites this paper.

Mastering PokeGym: Graph-Guided Multimodal Evolution at Test Time Plancraft: an evaluation dataset for planning with LLM agents

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T05:32:20.258624Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T05:32:20.258624Z digest=sha256:b2f71b3f2a368f2ac75be63903270fa91648709063750c1f9d345ce428ed2550

Observation 9acd0518-47c9-4290-bf37-7442fdd9cd96 · inbound

TacoMAS: Test-Time Co-Evolution of Topology and Capability in LLM-based Multi-Agent Systems cites this paper.

TacoMAS: Test-Time Co-Evolution of Topology and Capability in LLM-based Multi-Agent Systems Plancraft: an evaluation dataset for planning with LLM agents

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-12T05:51:24.593178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-12T04:53:54.754878Z digest=sha256:d615fcff8db17248f6c583e60070f5a84c2e5dbace59bbff5afa2c2da743e9d6

Observation 89f93662-bec2-4466-8eec-8de955794d7a · inbound

Object-Centric Environment Modeling for Agentic Tasks cites this paper.

Object-Centric Environment Modeling for Agentic Tasks Plancraft: an evaluation dataset for planning with LLM agents

Reference 21

Resolution
unresolved
no resolver link, observed 2026-07-12T06:39:21.435864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-12T06:39:21.435864Z digest=sha256:5b8501fb76f7c37654cf07f22627ff3b5dd574f38c1bc21bccad0a7a6b85f058

Observation f7a7032f-82e0-41ce-8f6d-08bafddaeae7 · inbound

Spinning Straw into Gold: Relabeling LLM Agent Trajectories in Hindsight for Successful Demonstrations cites this paper.

Spinning Straw into Gold: Relabeling LLM Agent Trajectories in Hindsight for Successful Demonstrations Plancraft: an evaluation dataset for planning with LLM agents

Reference 41

Resolution
unresolved
no resolver link, observed 2026-07-11T20:45:51.403759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T20:45:51.403759Z digest=sha256:4001fc13b254c432225783276ba062cd52820a5ce4badd124cfc3f5ee7ad8043

Observation 19a1ebac-1b72-4a5f-b1d8-97e85909ecdc · inbound

AgentRadio: Passive Awareness for Long-Horizon Multi-Agent Collaboration cites this paper.

AgentRadio: Passive Awareness for Long-Horizon Multi-Agent Collaboration Plancraft: an evaluation dataset for planning with LLM agents

Reference 4

Resolution
unresolved
no resolver link, observed 2026-07-31T07:38:10.235062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T07:38:10.235062Z digest=sha256:6d93dda80f449235ff27353f461395f23cfe4d3eb10febe6b03092b6e29f8750