{"id":"a7d10bb2-0579-45c4-a787-daf9d28b4716","arxiv_id":"2504.14773","paper_version":1,"verdict":"REJECT","confidence":"MODERATE","novelty_score":2.0,"correctness_risk":"high","formal_verification":"none","parameter_count":0,"one_line_summary":"A survey that catalogs and categorizes existing planning benchmarks for LLM agents and lists open gaps, without introducing new benchmarks, data, or experimental results.","lead":"This paper surveys benchmarks for testing how well large language models plan, grouping them into seven categories from embodied robots to web navigation and scheduling. It points out gaps like missing uncertainty and long-horizon tasks, but offers no new benchmarks, experiments, or concrete recommendations.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Section 5 is partly a foreign DualFormer passage with 'xxxxxx' placeholders, so the claimed seven-category overview is not coherent; the abstract's algorithm-to-benchmark recommendations are also never delivered.","rationale":"The reader's rejection is justified, but my emphasis is slightly different. The reader's weakest assumption is that a hand-picked selection without inclusion criteria cannot support the survey's recommendations. I agree that the selection method is undisclosed, but the more direct and load-bearing problem is internal: Section 5 contains a substantial foreign passage, including placeholder text, that breaks the coherence of one of the seven claimed categories. This is visible in the manuscript itself and does not depend on any assumption about what should have been included. I also checked the abstract's promise of recommendations: the body contains no explicit algorithm-to-benchmark mapping, so the central utility claim is not met by the text as written. This is an internal mismatch rather than a disagreement with field consensus. I do not attribute the corruption to bad faith; the appropriate response is a clean revision with the foreign passage removed, a stated inclusion protocol, and a concrete recommendations table. On the current text, rejection is the correct verdict. I mark partial agreement with the reader because the reader identified the selection-method problem and also noted the corrupted section, but the corruption deserves to be treated as the primary load-bearing failure.","tokens_in":16321,"tokens_out":4528,"duration_ms":44254,"concrete_test":"Run a textual audit: (1) identify every 'xxxxxx' run and every paragraph in Section 5 that is not about a named planning benchmark; (2) extract from Sections 9 and 10 every sentence containing 'recommend', 'appropriate', or 'should use' and tabulate each stated benchmark-to-algorithm pairing. If the foreign DualFormer passage can be removed without leaving Section 5 with a coherent game/puzzle survey, the seven-category coverage claim is unsupported; if the pairing table is empty, the abstract's recommendation promise is unfulfilled.","verdict_should_be":"REJECT","load_bearing_attack":"The central claim is that PLANET provides a structured, comprehensive overview of planning benchmarks and recommends the most appropriate benchmarks for various algorithms. That claim requires two things: coherent coverage of all seven announced categories, and an actual mapping from algorithms to recommended benchmarks. The first fails in Section 5 ('Planning in Games and Puzzles'): after describing SmartPlay, AucArena, GAMA-Bench, Plancraft, and PPNL, the text veers into a foreign passage about DualFormer and A* search, containing Figures 3.1 and 3.2, '3 RANDOMIZED STRATEGIC TRACE PRUNING', 'stochastic masking Mark's paper U2D2', a long block of 'xxxxxx' placeholders, and a '4 EXPERIMENTS' list. This is not a harmless formatting issue: one of the seven categories is partially occupied by unrelated content, so a reader cannot determine which game/puzzle benchmarks the survey actually covers. The second fails as well: the abstract promises that the paper 'recommends the most appropriate benchmarks for various algorithms,' but no section delivers a benchmark-to-algorithm mapping; Section 10 only restates the survey's intention. Additionally, Sections 1 and 10 provide no inclusion criteria, search protocol, or completeness check, so the four gaps in Section 9 are derived from an invisible sample rather than a transparent corpus. Together these problems undermine both the comprehensiveness and the practical utility that the paper advertises.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper is a survey of benchmarks for evaluating LLMs' planning capabilities. It claims to organize existing benchmarks into several categories, identify commonly used testbeds, and recommend appropriate benchmarks for different algorithms. The body describes a number of well-known benchmarks across embodied environments, web navigation, scheduling, games and puzzles, task automation, text-based reasoning, and agentic benchmarks, and it lists four gaps in current benchmark design: simple world models, fragile long-horizon planning, lack of uncertainty handling, and limited multimodal support. The abstract also advertises that the paper 'recommends the most appropriate benchmarks for various algorithms.' As submitted, however, the manuscript contains a corrupted foreign passage in Section 5 with 'xxxxxx' placeholders and unrelated DualFormer material, the abstract promises a five-category scheme while the introduction lists seven, and no section actually delivers the promised benchmark-to-algorithm recommendations. The survey also lacks any stated inclusion criteria or search protocol, so the selection of benchmarks is not verifiable.","tokens_in":16703,"tokens_out":3796,"duration_ms":34343,"significance":"If the paper were cleaned and properly scoped, it could serve as a useful entry point for researchers seeking an overview of popular LLM planning benchmarks. The descriptions of individual benchmarks are, for the most part, accurate and give helpful pointers to the primary sources. The gap discussion in Section 9 also identifies plausible directions for future work. However, the central claims of 'comprehensive understanding' and of recommending benchmarks for various algorithms are not supported as written. The corrupted Section 5 makes one of the seven announced categories unreadable, the absence of a methodology makes the coverage claims non-transparent, and the promised recommendations are absent. These are load-bearing issues for a survey whose stated purpose is to help readers select benchmarks. Because the problems are addressable in revision, I do not treat them as irreparable, but they are substantial.","major_comments":[{"comment":"Section 5 is not a coherent overview: after the descriptions of SmartPlay, AucArena, GAMA-Bench, Plancraft, and PPNL, the text is interrupted by a foreign passage titled '3 RANDOMIZED STRATEGIC TRACE PRUNING' and '4 EXPERIMENTS' that discusses DualFormer, includes Figures 3.1 and 3.2, a line reading 'stochastic masking Mark's paper U2D2', and a long block of 'xxxxxx' placeholders. A reader cannot determine which game-and-puzzle benchmarks the survey actually covers, and the section does not provide the announced overview of this category. This directly undermines the paper's claim of a comprehensive, seven-category survey.","section":"Section 5 (Planning in Games and Puzzles)"},{"comment":"The abstract states that the paper 'recommends the most appropriate benchmarks for various algorithms' and that benchmarks are categorized into five groups (embodied environments, web navigation, scheduling, games and puzzles, and everyday task automation), but Section 1 lists seven categories, adding text-based reasoning and planning as a subtask in agentic benchmarks. More importantly, no section of the paper delivers the promised mapping from algorithms or agent capabilities to recommended benchmarks; Section 10 only restates the survey's aim. The authors should either add the missing recommendation mapping or revise the abstract to reflect what the paper actually provides, and they must resolve the five-versus-seven category inconsistency.","section":"Abstract and Sections 1 and 10"},{"comment":"The survey gives no inclusion criteria, no search protocol, no time window, and no completeness check for its benchmark selection. The four gaps identified in Section 9 (static world models, long-horizon fragility, lack of uncertainty, limited multimodality) are therefore derived from an invisible sample rather than a transparent corpus. Since the paper advertises a 'comprehensive understanding' of planning benchmarks, the absence of a stated methodology is load-bearing: without it, the selection is an unverifiable convenience sample and the gap analysis cannot be reproduced or trusted.","section":"Sections 1 and 10 (methodology)"}],"minor_comments":[{"comment":"The heading 'Planning in Embodied Environments' is followed by an orphaned line 'TextWorld Embodied' that appears to be a leftover artifact from a figure or sidebar; it should be removed or integrated.","section":"Section 2 (header)"},{"comment":"In the MDP description, 'a reward functionS×A→ R' lacks spacing, and the sentence 'The ultimate goal of an MDP is to develop a policy, denoted as at = pϕ(a|st), focuses on identifying the optimal action...' is grammatically awkward. Please rewrite this passage.","section":"Section 1 (formal definition)"},{"comment":"The sentence 'with a max step limit of 15 steps' is redundant; 'max' and 'limit' convey the same constraint. Please simplify to 'with a maximum of 15 steps' or similar.","section":"Section 3 (OSWorld)"},{"comment":"The text says 'An illustration of this can be seen in Figure 5' when referring to Blocksworld, but Figure 5 appears much later and is primarily about RAP; the cross-reference should be fixed or the figure should be placed with the Blocksworld discussion.","section":"Section 2 (cross-reference)"},{"comment":"The corrupted passage contains 'Figure 3.1' and 'Figure 3.2' labels, which do not match the paper's figure numbering. If the passage is removed, these labels will disappear; if retained, they must be renumbered and integrated.","section":"Section 5 (figure numbering)"}],"recommendation":"major_revision","confidential_remarks":"The manuscript has serious presentation and content problems, but they appear fixable: the corrupted Section 5 can be excised or replaced, the abstract can be aligned with the body, and the authors can either add a methodology section describing how benchmarks were selected or explicitly reframe the paper as a selective overview rather than a comprehensive one. Because the central contribution is a survey, the missing recommendations and missing methodology are the main obstacles to publication. I recommend major revision rather than rejection, provided the authors address these issues in full."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Hi [Colleague],\n\nQuick take: this survey of LLM planning benchmarks is not ready for publication. One of its seven categories, games and puzzles, is partly occupied by a foreign passage about DualFormer and randomized trace pruning, complete with 'xxxxxx' placeholders and a figure caption that just says 'Caption'. That alone is a critical integrity issue. On top of that, the abstract promises that the paper 'recommends the most appropriate benchmarks for various algorithms,' but no section ever delivers such a mapping; Section 10 only restates the intention. The stress-test concern holds up.\n\nWhat the paper does well: the descriptions of individual benchmarks are mostly accurate and concise. A junior researcher looking for a starting point—ALFWorld, WebArena, TravelPlanner, PlanBench, and so on—would come away with a reasonable map of the area. The four gaps flagged in Section 9 (static world models, long-horizon fragility, lack of uncertainty, limited multimodality) are reasonable, even if not new.\n\nWhere it falls short beyond the corruption: the selection of benchmarks is unsystematic. There are no inclusion criteria, search protocol, or completeness check, so the claim of providing a 'comprehensive understanding' is unearned. The categorization into seven groups is subjective, and the paper's own prior surveys (LASP, PlanGenLLMs) cover much of the same ground; the incremental value over those is modest. The self-citation itself is harmless in a review context, but the overlap means this needs to justify its existence more than it does.\n\nThe central argument, as stated, is not supported: the promised benchmark-to-algorithm recommendations are absent, and the corrupted section breaks the coherence of the seven-category overview. The paper could be rehabilitated with a clean rewrite, a methodology section, and a decision to either deliver the recommendations or drop that claim.\n\nFor peer review: I would not send this out for formal review in its current state; it should be rejected or sent back to the authors for major revision before any referee sees it. If cleaned up, it could become a useful orientation piece for newcomers, but that is the audience it serves. I wouldn't cite it in my own work.\n\nBest,\n[You]","headline":"A useful benchmark map for newcomers, but the corrupted Section 5 and missing promised recommendations make this survey unpublishable as is.","tokens_in":17094,"tokens_out":2734,"would_cite":false,"duration_ms":24141,"reading_group":"no","serious_thinker":"no","would_accept_peer_review":false},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper organizes LLM planning benchmarks into seven categories and identifies four gaps in how planning is tested.","keywords":["planning benchmarks","LLM agents","agentic AI","benchmark survey","world models","long-horizon tasks","planning under uncertainty","multimodal agents"],"falsifier":"A systematic enumeration of planning-related benchmark papers from the same period that finds a substantial cluster fitting none of the seven categories, for example benchmarks built around formal PDDL planning domains with dynamic, stochastic elements, or a pre-existing multimodal benchmark with dynamic world models that the survey omits, would show the taxonomy and the four-gap analysis to be incomplete.","tokens_in":16118,"feed_emoji":"🧩","tokens_out":4757,"duration_ms":39383,"temperature":0.7,"pith_summary":"The paper tries to bring order to the growing menagerie of benchmarks for evaluating whether large language models can plan. It organizes recent benchmarks into seven categories, from embodied household environments to web navigation, scheduling, games and puzzles, everyday task automation, text-based reasoning, and planning as a subtask of general agentic benchmarks. On this basis it identifies four gaps: static world models, fragile long-horizon planning, little planning under uncertainty, and scant multimodal support. The intended payoff is a map that helps researchers choose the right testbed for a planning algorithm and shows where new benchmarks are needed.","feed_headline":"Planning benchmarks: 7 categories, 4 gaps","feed_subtitle":"A survey maps embodied, web, scheduling, games, task, text, and agentic testbeds—and where testing falls short.","key_machinery":"The organizing device is the paper's working definition of planning: explicit state modeling, outcome reasoning, goal orientation, and constructing sequences or policies under constraints, formalized through a Markov decision process with states $S$, actions $A$, transition model $p_\\theta(s_{t+1}\\mid s_t, a_t)$, reward $r_\\theta$, and policy $a_t = p_\\phi(a\\mid s_t)$. This definition is the filter that selects benchmarks for the seven categories and the lens through which the four gaps are derived; those gaps are static world models, long-horizon fragility, lack of uncertainty, and limited multimodality.","core_discovery":"The paper claims that the field lacks a comprehensive understanding of planning benchmarks, and that a principled way to define planning, as tasks with explicit state modeling, outcome reasoning, goal orientation, and sequences or policies within constraints, yields a seven-way taxonomy of available testbeds. Surveying those testbeds, it argues that common benchmarks make planning too easy by relying on static, fully observable worlds, so LLMs can succeed by pattern matching rather than building and revising world models; that long-horizon plans are fragile because agents lack state tracking and error recovery; that uncertainty and partial information are under-tested; and that text-only evaluation bypasses the visual grounding needed for multimodal agents. The paper's recommendation is that future benchmark development should target dynamic environments, long horizons, uncertainty, and multimodality.","pith_inferences":["The four gaps suggest a concrete re-ranking test: adding a stochastic or partially observable variant of an existing benchmark, say TravelPlanner with flight delays, would likely separate planners that rebuild state estimates from those that rely on static context.","Because planning spans games, web use, and scheduling, a single benchmark can exercise several capabilities at once; future design could treat planning as a compositional dimension rather than a task family.","If static world models inflate apparent planning ability, then model rankings from existing leaderboards are probably environment-specific, and transferring them to partially observable deployments would be unreliable."],"forward_implications":["Researchers choosing a testbed can use the seven-category map to match a benchmark to the planning capability they want to isolate, such as constraint satisfaction in scheduling or long-horizon execution in web navigation.","If the four gaps are real, new benchmarks that stress dynamic world models, long horizons, uncertainty, and multimodality would better expose whether LLMs plan or pattern-match.","Text-only benchmarks may overstate LLM planning ability relative to multimodal settings, since visual grounding is largely bypassed in current suites.","The MDP-based definition implies that benchmarks evaluating planning should report state transitions and goal conditions explicitly, so that plan validity can be checked mechanically."],"supporting_citations":[{"why":"Anchors the embodied-environment category by pairing textual world models with PDDL-style household actions.","marker":"Shridhar et al. 2021"},{"why":"Supplies the core planning-capability probes, including plan generation, verification, and replanning on Blocksworld and Logistics.","marker":"Valmeekam et al. 2023"},{"why":"Provides the long-horizon web navigation evidence, including the large gap between GPT-4 agents and human success.","marker":"Zhou et al. 2024b"},{"why":"Defines the scheduling category through constraint-heavy travel itinerary planning with tool use.","marker":"Xie et al. 2024a"},{"why":"Shows how natural-language scheduling tasks expose fragility as constraint complexity grows.","marker":"Zheng et al. 2024"},{"why":"Supports the claim that LLMs struggle to select correct proof steps, treating reasoning as planning.","marker":"Saparov & He 2023"},{"why":"Exemplifies planning as a subtask within general agentic benchmarks spanning many environments.","marker":"Liu et al. 2023"},{"why":"Provides the classical Blocksworld substrate that PlanBench and other planning tests build on.","marker":"Gupta & Nau 1992"}],"fun_headline_variants":["Planning benchmarks too easy: static worlds, no uncertainty","Survey: planning testbeds lack dynamic and visual challenges","Seven-way taxonomy reveals four planning benchmark gaps","Planning tests miss long-horizon and multimodal demands"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The survey's utility rests on its hand-picked selection of benchmarks being representative enough to ground its recommendations, but it states no inclusion criteria and performs no systematic search, so an unrepresentative selection would weaken the category map and the gap analysis.","fun_headline_variants_meta":{"raw":{"variants":["Planning benchmarks too easy: static worlds, no uncertainty","Survey: planning testbeds lack dynamic and visual challenges","Seven-way taxonomy reveals four planning benchmark gaps","Planning tests miss long-horizon and multimodal demands"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000196,"raw_usage":{"total_tokens":1306,"prompt_tokens":834,"completion_tokens":472,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":450,"completion_tokens_details":{"reasoning_tokens":410}},"tokens_in":450,"tokens_out":472,"duration_ms":4812,"temperature":1.0,"reasoning_tokens":410,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T11:39:45.319960+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A systematic enumeration of planning-related benchmark papers from the same period that finds a substantial cluster fitting none of the seven categories, for example benchmarks built around formal PDDL planning domains with dynamic, stochastic elements, or a pre-existing multimodal benchmark with dynamic world models that the survey omits, would show the taxonomy and the four-gap analysis to be incomplete.","supporting_citations":[{"cited_title":"Planbench: An extensible benchmark for evaluating large language models on planning and reasoning about change","cited_arxiv_id":null,"evidence_quote":"Supplies the core planning-capability probes, including plan generation, verification, and replanning on Blocksworld and Logistics."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the classical Blocksworld substrate that PlanBench and other planning tests build on."}],"review_version":1}