{"id":"ef7ec14a-a30a-420f-9c9b-4126c32ec20e","arxiv_id":"2508.03232","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"CookBench is a new long-horizon cooking benchmark with 14,394 bilingual instructions and over 4,500 embodied cooking tasks that current AI models cannot yet solve autonomously.","lead":"This paper introduces CookBench, a simulated cooking benchmark with 120-step tasks that tests an AI's ability to understand cooking requests and then plan physical actions. It measures how well current language models can handle long, realistic cooking workflows and identifies where they fail, such as navigation and spatial reasoning.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 4,446 multi-dish tasks are never executed in the embodied Stage 2 study, leaving the central long-horizon challenge unvalidated for the bulk of the benchmark.","rationale":"Reading in good faith, the intention-recognition stage is well supported: the dataset is substantial (14,394 bilingual instructions), the accuracy metric is strict and clearly defined, and the model sweep is meaningful. The decoupled two-stage framework is a useful design choice. The problem is in the embodied stage, which is the heart of the long-horizon claim. Two gaps remain even if the automated scorer is accepted as valid for single-dish tasks. First, the 4,446 multi-dish tasks, which constitute the vast majority of the benchmark and the source of its claimed multi-task scheduling difficulty, are never run in Stage 2. Second, the single-dish HITL results are not a clean model evaluation because humans supply navigation corrections and decide when to terminate. The reader's weakest_assumption about the five-point scoring is related but narrower: it assumes the scorer is valid, whereas the more fundamental issue is that for multi-dish tasks there is no evidence the scorer exists and works, or even that the tasks can be completed and submitted in the simulator. The paper's own 'feasibility study' framing reduces but does not eliminate the gap between the proposal and the strong abstract-level claim. I therefore retain the CONDITIONAL verdict: accept the benchmark proposal on the strength of Stage 1 and the environment, but require multi-dish feasibility and scorer validation before treating CookBench as a rigorous long-horizon benchmark. This is not a rejection; it is a request for the evidence that the headline contribution requires.","tokens_in":12147,"tokens_out":5454,"duration_ms":72259,"concrete_test":"Run the same HITL protocol on a stratified sample of 50 multi-dish tasks spanning 4/5/6 dishes and both batch and sequential order patterns. For each task, record the model-generated plan, the number and type of human interventions (limited to the eight navigation/camera commands), the number of planning-loop terminations, and the final automated completion score. Also run a human-only condition on the same tasks and an inter-rater check where three independent cooks re-score 20 fixed single-dish final states. If multi-dish tasks cannot reach a submitted final dish, or if HITL scores do not separate from a no-planning control, the benchmark's central multi-dish claim fails.","verdict_should_be":"UNCHANGED","load_bearing_attack":"CookBench's strongest claim is that it is a valid, knowledge-driven benchmark with average task length around 120 steps, and that leading models fail on it. For that claim to hold, the full task suite must be runnable in the simulator and the five-point automated score must reflect cooking success. The paper provides direct evidence only for the 131 single-dish tasks in Table 3 ('Experiments and Analysis' → 'Stage 2: A Human-in-the-Loop Feasibility Study'). The 4,446 multi-dish combinations are enumerated and described (batch vs. sequential orders, 4/5/6 dishes), but they are never executed, submitted, or scored in Stage 2. This matters because multi-dish tasks are precisely where the claimed long-horizon, multi-task scheduling, and resource-contention challenges live; a benchmark that validates only single-dish tasks cannot support the headline complexity claim. The single-dish evidence is also confounded: HITL scores of 0.32–0.75 are produced with human-corrected navigation and camera commands, human termination decisions, and no no-model baseline or intervention-count ablation. The claim that 'current leading models exhibit severe limitations' is therefore a qualitative observation from a feasibility study, not a rigorous evaluation, even though the abstract and contributions list present it as a finding. The paper is candid that Stage 2 is a feasibility study, but the central benchmark claim still leans on unvalidated multi-dish execution and scoring.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"CookBench is introduced as a two-stage, long-horizon embodied planning benchmark for cooking. Stage 1 is a bilingual (Chinese/English) intention-recognition dataset of approximately 14,394 instructions spanning single-dish and multi-dish orders; Stage 2 is an embodied interaction task suite of 131 single-dish and 4,446 multi-dish cooking scenarios built on the Unity-based Cooking Simulator, with a 35-action atomic action library, a knowledge base of 131 recipes, and an automated five-point completion scoring system. The paper evaluates several LLMs on Stage 1, showing strong single-dish performance but large drops on multi-dish instructions, and runs a human-in-the-loop (HITL) feasibility study on Stage 2 using GPT-4.1 and GPT-4o with human-corrected navigation, reporting scores between 0.32 and 0.75 on the single-dish tasks. The central claims are that CookBench has an average task length of 120 steps, that it stresses long-horizon planning, multi-task scheduling, and physical commonsense reasoning, and that current leading models show severe limitations on this challenge.","tokens_in":12436,"tokens_out":3904,"duration_ms":48921,"significance":"If the benchmark is delivered and validated as described, CookBench would fill a real gap: existing embodied planning benchmarks mostly have short, linear task flows and coarse action abstractions, whereas cooking naturally combines long horizons, irreversible state changes, resource contention, and domain knowledge. The Stage 1 dataset is substantial, deliberately structured, and the reported experiments provide quantitative evidence that multi-dish intent recognition is hard for current LLMs. The fine-grained atomic action library, the three-dimensional semantic map, and the plan to open-source the benchmark are also concrete strengths. The evaluation outcomes are independent measurements of external models, so there is no circularity between benchmark construction and model results. The significance is conditional, however, on validation of the embodied Stage 2: the full multi-dish suite is never executed, and the automated scoring system is not validated against independent judgments of cooking success.","major_comments":[{"comment":"The paper claims to evaluate 131 single-dish and 4,446 multi-dish tasks, but Table 3 and the entire Stage 2 study report results only for the 131 single-dish tasks. The multi-dish combinations, which are precisely where the long-horizon, multi-task scheduling, and resource-contention claims live, are never executed, submitted, or scored. This is a load-bearing gap: without at least a demonstrative subset of multi-dish tasks being run in the simulator and scored, the headline claim that CookBench is a validated testbed for long-horizon multi-dish planning is unsupported. The authors should run a representative sample of 4-, 5-, and 6-dish tasks (e.g., a few across difficulty levels and scheduling patterns) and report completion scores, completion times, and failure modes.","section":"Stage 2: A Human-in-the-Loop Feasibility Study (Table 3)"},{"comment":"The automated five-point completion score is asserted to assess ingredient correctness, cooking techniques, and final component states, but its validity is never established. The only evidence offered is the 'Human' row in Table 3, where two volunteers score 5 on all tasks; this assumes rather than demonstrates that the metric aligns with cooking success. Without validation, all Stage 2 scores are unreliable. The authors should validate the scorer against independent human ratings, report inter-annotator agreement, or run controlled ablations showing that the score detects missing ingredients, wrong techniques, and incorrect final states.","section":"The CookBench Environment and Task Framework (evaluation metrics paragraph)"},{"comment":"The HITL results are confounded by extensive human assistance: navigation and camera commands are manually corrected, the process is manually terminated when the agent loops, and there is no no-model baseline or intervention-count ablation. The reported scores of 0.32–0.75 therefore measure a mixed human-model system rather than the planning capability of the models, making the conclusion that 'current leading models exhibit severe limitations' a qualitative observation rather than a rigorous finding. The authors should quantify the number and type of human interventions, include a human-only and a no-intervention model baseline, and report variance across tasks and repeated runs.","section":"Stage 2: A Human-in-the-Loop Feasibility Study (Experimental Setup and Figure 5)"},{"comment":"The central quantitative claim that CookBench has an 'average task length of 120 steps' is presented without derivation. It is unclear whether a step is a high-level sub-task, an atomic action from the 35-action library, or an API call including navigation increments, and whether the average is over single-dish tasks, multi-dish tasks, or both. Because this number is the basis for the long-horizon comparison in Figure 1 and for the first stated contribution, the authors should define the step unit and report the distribution of task lengths separately for single-dish and multi-dish suites.","section":"Figure 1 and Contributions (§1)"}],"minor_comments":[{"comment":"There are several typos and ungrammatical phrases, including 'domain knowledgee', 'recognization', and 'for the a future'; these should be corrected.","section":"Introduction"},{"comment":"The figure contains incomplete or ungrammatical labels, such as 'Interacte' and 'API function', which should be fixed for clarity.","section":"Figure 2"},{"comment":"The 'Average' column is not defined; it is unclear whether it is a simple mean over all conditions or a weighted mean, and this ambiguity makes cross-model comparisons harder to interpret.","section":"Table 2"},{"comment":"The 'Human' row is based on only two volunteers, and the paper does not report their level of cooking expertise or any inter-rater statistics; this should be stated explicitly.","section":"Table 3"},{"comment":"The abstract and contribution list describe the Stage 2 results as an 'in-depth analysis' and 'bottleneck analysis' of leading models, while the body repeatedly and honestly calls the study a feasibility study; the language should be aligned so that the exploratory nature of the HITL evaluation is not overstated.","section":"Abstract and Contributions"},{"comment":"The model names 'Gpt-4.1' and 'gpt-4.1' are used inconsistently; standard capitalization should be used throughout.","section":"Experiments and Analysis"}],"recommendation":"major_revision","confidential_remarks":"The paper is honest about the feasibility nature of the Stage 2 study, and the Stage 1 evaluation is a solid, non-circular contribution. However, the benchmark paper's central claims about long-horizon complexity and the validity of the automated scoring rest on evidence that is currently limited to 131 single-dish tasks and an unvalidated scorer. I believe this is fixable within the manuscript's scope by adding a small multi-dish execution study, validating the scoring metric, and quantifying HITL interventions. The overclaiming in the abstract and contribution list should also be tempered or supported by the new evidence."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"CookBench is a real attempt to fill a gap: existing embodied planning benchmarks are short-horizon (ALFRED ~12 steps) and coarse; CookBench offers ~120-step tasks, fine-grained spatial actions, continuous state changes, and bilingual instructions, all in a high-fidelity cooking simulator. The two-stage decoupling of intent recognition from embodied execution is sensible, and the intent-recognition dataset (14,394 bilingual instructions, with single/multi-dish, interference, ambiguity) is a solid contribution on its own. The Stage 1 experiments are legitimate: they test a range of LLMs, report strict accuracy, and show that multi-dish instructions drop accuracy dramatically. That is real empirical content.\n\nThe soft spots are in Stage 2. The stress-test note is right: the 4,446 multi-dish tasks -- precisely where the long-horizon scheduling and resource-contention contributions live -- are never executed. All Stage 2 evidence is from the 131 single-dish tasks, and that evidence has confounds: human-corrected navigation and camera commands, human termination decisions, no no-model baseline, no intervention count. So the claim that 'leading models exhibit severe limitations' on long-horizon embodied planning is a qualitative observation from a feasibility study, not a measured result. The paper says as much ('exploratory evaluation', 'feasibility study'), so it is not deceptive, but the abstract and contribution list lean on it more than the body supports. Also, the automated five-point scoring is asserted but not validated: no correlation with human ratings (beyond perfect expert scores), no analysis of inter-rater reliability, no sensitivity check. That matters because the entire benchmark's usefulness depends on that score being a reliable proxy for cooking success.\n\nWho is this for? Researchers building embodied planning benchmarks or evaluating long-horizon agents. They will want the code and data; the paper says 'will be open-sourced' but links nothing, so verification is impossible. My recommendation: send it to peer review, but the review must insist on (1) executing at least a sample of the multi-dish tasks, ideally without human navigation corrections, (2) validating the automated score, and (3) releasing artifacts before acceptance. The benchmark design is good enough that a revised version could be a standard testbed. If the artifacts never materialize, treat it as a position paper.","headline":"CookBench has a genuinely useful intent-recognition dataset and a well-aimed long-horizon task design, but the embodied results are only a single-dish feasibility study with human corrections, so the headline claims outrun the evidence.","tokens_in":12929,"tokens_out":3024,"would_cite":false,"duration_ms":30409,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"CookBench claims that long-horizon embodied planning in a kitchen, with tasks averaging 120 steps, is now measurable, and that current large language and vision-language models pass intent parsing but fail at autonomous physical execution.","keywords":["embodied planning","long-horizon tasks","cooking benchmark","intention recognition","vision-language models","human-in-the-loop evaluation","physical commonsense","simulation environment"],"falsifier":"Run a blind comparison in which experienced cooks rate finished simulator dishes on the same five-point scale and compare their ratings to the automated scores, and also test a scripted policy that completes dishes by rote without long-horizon planning; if the rote policy earns passing scores or the human ratings diverge widely from the automated ones, the claim that CookBench measures long-horizon planning would be refuted.","tokens_in":11999,"feed_emoji":"🍳","tokens_out":6764,"duration_ms":77748,"temperature":0.7,"pith_summary":"CookBench is a proposal for a benchmark that makes long-horizon embodied planning measurable in a kitchen: the average task demands around 120 high-level actions, an order of magnitude beyond mainstream benchmarks such as ALFRED. The paper's aim is to establish that this task length, combined with fine-grained spatial action primitives, exposes capabilities that short-horizon benchmarks do not, and that the two-stage decomposition of intention recognition and embodied interaction is the right way to study them. It reports that leading language and vision-language models handle simple intent parsing well but deteriorate on multi-dish and ambiguous instructions, and that no current model can autonomously complete the embodied stage; all results from that stage come from a human-in-the-loop feasibility study. If the benchmark holds up, it gives the community a reproducible way to separate goal understanding from physical execution and to measure progress on memory, scheduling, and physical commonsense.","feed_headline":"Cooking benchmark stretches embodied AI to 120-step tasks","feed_subtitle":"First tests show models parse recipes well but fail at physical, long-horizon execution.","key_machinery":"The two-layer task framework is the load-bearing design: Stage 1 turns natural-language instructions into a structured list of target dishes, and Stage 2 turns that list into sequences of atomic actions in the simulator. The atomic action library is the central interface, keeping the spatial operational detail the planner needs, such as which object to move and how far, while hiding low-level robotic control. The automated evaluation, which scores ingredient correctness, technique, and final state on a five-point scale and returns textual error feedback, is what makes the benchmark reusable without human judges.","core_discovery":"The paper's central discovery is that long-horizon cooking tasks of about 120 steps can be operationalized as a benchmark: a Unity-based simulator with 131 recipes, a knowledge base of ingredients and tools, and an atomic action library of roughly 35 functions such as pick up, cut, pour, flip, and start heating. Against this testbed, the authors find that modern LLMs and VLMs are competent at the first stage, with the best models exceeding 90 percent accuracy on simple English instructions, but their accuracy falls to around 50 percent on six-dish instructions, and in the embodied stage they fail at navigation, visual state verification, and precise spatial manipulation, requiring human correction for even simple dishes. The intended takeaway is that CookBench is a valid and challenging evaluation platform that isolates the bottleneck of physical execution from the already-hard problem of intent recognition.","pith_inferences":["(Inference) The spatial-level action abstraction could generalize to other long-horizon domains where the planner must choose positions and orientations without controlling a robot's joints, such as assembly or surgery simulation.","(Inference) The sharp accuracy drop between single-dish and multi-dish instructions suggests that multi-goal parsing and ordering, not dish vocabulary, is the current bottleneck in language grounding; a targeted dataset on disambiguation might improve embodied agents before any simulator training.","(Inference) If the automated score is shown to correlate with independent human ratings, CookBench could serve as a general physical-commonsense probe rather than only a cooking benchmark.","(Inference) A direct next test is to couple the VLM perception module with a small navigation policy that handles corner-recovery and distance alignment; the paper's failure list predicts such a hybrid would solve the easiest dishes without human help."],"forward_implications":["A task length of about 120 steps makes CookBench roughly ten times longer than ALFRED's typical 12-step tasks, so it can stress-test long-term memory and multi-task scheduling.","The two-stage decomposition means a model's failure can be localized: an agent that knows the recipe but cannot turn it into motion will score low only in Stage 2.","The reported intent-recognition results put a concrete upper bound on current models: above 90 percent on simple English instructions but near 50 percent on six-dish Chinese orders, so language grounding is far from solved.","The automated five-point score with textual feedback gives a reproducible metric for comparing future embodied planners on the same dishes.","Because no model could finish the embodied stage autonomously, the paper's human-in-the-loop results define the current frontier and set a concrete target for fully automated systems."],"supporting_citations":[{"why":"ALFRED is the main comparison baseline for short-horizon embodied instruction following, providing the 12-step task-length contrast.","marker":"Shridhar et al. 2020a"},{"why":"VirtualHome supplies the household-activity simulator comparison and the 349-object environment that CookBench contrasts.","marker":"Puig et al. 2018"},{"why":"ET-Plan-Bench is the closest longer-horizon task-level planning benchmark that CookBench extends beyond.","marker":"Zhang et al. 2024"},{"why":"EmbodiedBench is one of the compared multimodal embodied benchmarks and a reference point for task length and interaction granularity.","marker":"Yang et al. 2025c"},{"why":"ALFWorld provides the text-based planning comparison and a 6-step task-length data point used in the benchmark comparison.","marker":"Shridhar et al. 2020b"},{"why":"GPT-4 is the base model used as the high-level planner in the human-in-the-loop embodied evaluation.","marker":"Achiam et al. 2023"},{"why":"Gemini 2.5 is one of the top-scoring closed-source models in the intent-recognition experiments.","marker":"Comanici et al. 2025"},{"why":"DeepSeek-R1 is the best-performing open-source model in the intent-recognition results.","marker":"Guo et al. 2025"}],"fun_headline_variants":["Cooking benchmark reveals AI's weak physical skills","AI understands recipes but can't physically cook them","CookBench: 120-step cooking tasks stump AI agents","Embodied AI fails at physical cooking in new benchmark","Cooking benchmark: intent easy, execution hard for AI"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the simulator's automated five-point score, which assesses ingredient correctness, technique, and final state, is a trustworthy measure of whether a dish was actually cooked well; if that scoring does not track real cooking success, the benchmark's results lose their meaning.","fun_headline_variants_meta":{"raw":{"variants":["Cooking benchmark reveals AI's weak physical skills","AI understands recipes but can't physically cook them","CookBench: 120-step cooking tasks stump AI agents","Embodied AI fails at physical cooking in new benchmark","Cooking benchmark: intent easy, execution hard for AI"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000613,"raw_usage":{"total_tokens":2870,"prompt_tokens":983,"completion_tokens":1887,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":599,"completion_tokens_details":{"reasoning_tokens":1811}},"tokens_in":599,"tokens_out":1887,"duration_ms":15577,"temperature":1.0,"reasoning_tokens":1811,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T04:34:02.591115+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run a blind comparison in which experienced cooks rate finished simulator dishes on the same five-point scale and compare their ratings to the automated scores, and also test a scripted policy that completes dishes by rote without long-horizon planning; if the rote policy earns passing scores or the human ratings diverge widely from the automated ones, the claim that CookBench measures long-horizon planning would be refuted.","supporting_citations":[],"review_version":1}