{"id":"0e070af6-e6d0-4f25-a145-65b3e0d920ba","arxiv_id":"2508.13530","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"CrafterDojo provides CrafterVPT, CrafterCLIP, and CrafterSteve-1 models plus datasets and benchmarks, making Crafter a usable testbed for general-purpose embodied agents.","lead":"This paper introduces CrafterDojo, a collection of pretrained models and toolkits that bring Minecraft-style foundation modeling to the lightweight Crafter game environment. The aim is to let embodied-agent researchers prototype quickly without Minecraft's heavy engineering overhead.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The suite's value rests on the released models actually providing useful behavior priors and instruction-following in Crafter; the abstract provides no quantitative evidence and the corrupted full text blocks verification.","rationale":"I read the abstract and the corrupted full-text extraction in good faith. The paper is an infrastructure contribution: it provides models, datasets, and code. The central claim depends on these artifacts being usable and beneficial for embodied-agent research in Crafter. I could not identify an internal logical inconsistency from the abstract alone. The most fragile premise is empirical: the models must genuinely work, especially CrafterSteve-1 for instruction following. The reader's weakest_assumption also mentioned model performance, so we partially agree, though I focus on the instruction-following model as the entry point that would falsify the 'unlock' claim. Because the full text is unreadable, I cannot verify whether the paper already contains such evaluations. The current UNVERDICTED verdict is appropriate, and my stress-test does not move it: if the concrete test passes, the claims are likely supported; if it fails, the central claim weakens substantially.","tokens_in":9792,"tokens_out":4844,"duration_ms":51891,"concrete_test":"Download the released CrafterSteve-1 checkpoint and the CrafterCaption dataset from the linked codebase. Generate a held-out set of 50–100 natural-language instructions covering all 22 achievements, excluding any prompts that appear in CrafterCaption. Run the model in the Crafter environment with a greedy decoding policy and measure per-instruction success rate. Compare against a random-action baseline and a PPO agent trained without CrafterVPT or CrafterSteve-1 for the same number of environment interactions. If CrafterSteve-1 does not exceed both baselines by a meaningful margin (e.g., success rate at least double the baseline), then the claim that the suite provides instruction-following foundation models is not established and the verdict should be revised.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that CrafterDojo unlocks Crafter as a lightweight testbed for general-purpose embodied agent research. For this to hold, the three introduced models must be functional building blocks: CrafterVPT must provide a behavior prior that improves learning, CrafterCLIP must provide useful vision-language grounding, and CrafterSteve-1 must follow instructions. None of these are supported by the abstract, which lists components but gives no evaluation numbers, baselines, ablations, or qualitative examples. The supplied full-text extraction is corrupted and appears to contain material from a different paper, so the promised benchmark evaluations cannot be located. The single most load-bearing unverified premise is therefore that CrafterSteve-1 (or the pipeline built on it) actually outperforms a simple scripted agent or a from-scratch RL policy on held-out Crafter instructions. If the instruction-following model is at chance on unseen instructions, the 'foundation model' framing overstates what is delivered, and the suite reduces to datasets and tools rather than a working testbed.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript presents CrafterDojo, a proposed open-source suite for the Crafter environment, comprising three models (CrafterVPT for behavior priors, CrafterCLIP for vision-language grounding, and CrafterSteve-1 for instruction following) plus two dataset-generation toolkits (CrafterPlay and CrafterCaption), reference agents, and benchmark evaluations. The abstract claims these components unlock Crafter as a lightweight, Minecraft-like testbed for embodied-agent research. The supplied full text is corrupted and unreadable, so the experimental evidence behind these claims cannot be inspected.","tokens_in":10001,"tokens_out":3752,"duration_ms":39002,"significance":"If the released models and benchmarks deliver what is promised, CrafterDojo would be a useful community resource: pretrained behavior, vision-language, and instruction-following models do not currently exist for Crafter, and the environment's speed and low overhead make it a credible prototyping setting. The open-source codebase and datasets would support reproducible research. However, the significance is conditional: the abstract reports no quantitative results, and the corruption of the supplied full text prevents verification of the models' actual performance, the benchmark protocols, and the claimed suitability of Crafter as a general-purpose embodied-agent testbed.","major_comments":[{"comment":"The abstract lists three models and two toolkits but reports no quantitative results, baselines, or ablations for any of them; the central claim that CrafterVPT, CrafterCLIP, and CrafterSteve-1 are usable foundation models is therefore unsupported. Please add headline metrics, at minimum task-success rates against a scripted agent and a from-scratch reinforcement-learning baseline on held-out instructions.","section":"Abstract"},{"comment":"The submitted full text is unreadable mojibake and includes material from an unrelated cond-mat paper (arXiv:2508.13519v1); no experimental protocol, training details, dataset statistics, or benchmark tables can be located. This blocks verification of every load-bearing claim in the paper, and a complete, readable manuscript must be provided before the work can be evaluated.","section":"Full text (as supplied)"},{"comment":"The premise that Crafter is a faithful proxy for Minecraft and that progress on Crafter transfers to general-purpose embodied agents is asserted rather than argued; there is no analysis of which Minecraft-like capabilities CrafterDojo exercises, nor any transfer experiment to a more complex environment. Please either add such an analysis or experiment, or soften the 'testbed for general-purpose embodied agents' claim to match the evidence.","section":"Abstract / contribution"},{"comment":"The abstract mentions generating behavior and caption datasets but does not specify how the benchmark evaluation tasks are constructed; if the same captioner used for training also generates the evaluation tasks, the instruction-following results could partly reflect captioner idiosyncrasies rather than general embodied instruction following. Please specify the task-generation protocol, including any human validation or held-out task distribution.","section":"Benchmark methodology"}],"minor_comments":[{"comment":"The term 'foundation model' is used in at least three different senses (behavior prior, vision-language encoder, instruction-following policy); the abstract would benefit from a brief clarification of each model's role.","section":"Abstract"},{"comment":"To support the 'open-source codebase' claim, the abstract or introduction should include the repository URL or model-release identifiers.","section":"Abstract"},{"comment":"The manuscript should state the licenses under which the models and datasets are released, since licensing affects the suite's usability as a community resource.","section":"General"}],"recommendation":"major_revision","confidential_remarks":"The full text supplied for review is corrupted and appears to be extracted from a different arXiv paper (2508.13519). Before any decision, the editor should obtain a clean copy of the original submission. Based on the abstract alone, the manuscript does not meet the evidentiary standard for acceptance, but the shortcomings are fixable if the readable text contains the missing evaluations."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: the paper does something genuinely useful — it ports three proven foundation-model recipes (VPT, CLIP, Steve-1) to Crafter and bundles datasets, benchmarks, and an open-source codebase. That gives the embodied-agent community a lightweight testbed for prototyping generalist agents. But the abstract reports no quantitative results, and the full-text extraction I was given is corrupt (mojibake, with fragments from an unrelated cond-mat paper), so I can't verify whether the benchmark evaluations exist or whether the instruction-following model, CrafterSteve-1, beats a scripted baseline. That is the load-bearing question. If the model is near chance on held-out instructions, the 'foundation model' label oversells what's delivered, and the package reduces to datasets and tools.\n\nWhat's good: the adaptation is sensible and the authors don't claim algorithmic novelty. Crafter is a reasonable lightweight alternative to Minecraft for prototyping, and the field has lacked ready-made behavior priors, caption data, and instruction-following checkpoints for it. The toolkit components (CrafterPlay, CrafterCaption, reference agents, code) are concrete deliverables that people can use even if the headline claims are modest.\n\nSoft spots, in proportion: the main one is the missing evidence, which the abstract makes worse. A second, more conceptual issue is the 'Minecraft-like testbed' framing. Crafter is a 2D grid world with dictionary observations; transfer of insights from it to Minecraft is not automatic. The paper's value would increase if it either showed transfer to MineDojo/Minecraft or explicitly framed Crafter as a standalone substrate. Third, there is a risk of evaluation circularity if the captioner used to generate instruction tasks is the same one being evaluated; the abstract does not rule that out.\n\nIf the full text is intact and the benchmarks show a real edge over from-scratch RL or scripted agents, this is a solid venue paper. If the numbers are missing, it is an infrastructure note. I'd still send it to peer review: the artifacts deserve a critical evaluation, and the community would benefit from knowing whether the behavior prior actually speeds up learning. Get the real PDF before making a final call.","headline":"Useful Crafter infrastructure with the right ambitions, but the lack of any numbers in the abstract makes it impossible to judge the central claim.","tokens_in":10461,"tokens_out":2811,"would_cite":false,"duration_ms":27488,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"CrafterDojo claims to unlock Crafter as a lightweight, prototyping-friendly, Minecraft-like testbed for general-purpose embodied agent research by providing pretrained foundation models, datasets, and benchmark toolkits.","keywords":["CrafterDojo","Crafter","foundation models","embodied agents","instruction following","vision-language grounding","behavior priors","open-ended environments"],"falsifier":"Train a batch of agents with CrafterDojo in Crafter and the same architectures directly in Minecraft on matched suites of instruction-following and exploration tasks; if strong Crafter performance consistently fails to correspond to or transfer to Minecraft performance, the central premise that Crafter captures the relevant challenges is falsified.","tokens_in":9645,"feed_emoji":"🤖","tokens_out":7197,"duration_ms":62075,"temperature":0.7,"pith_summary":"This paper argues that the Crafter environment, a lightweight Minecraft-like game, has been underused in embodied AI research because it lacks the foundation models that made Minecraft a productive testbed. To close that gap, it presents CrafterDojo, a suite of three pretrained models — CrafterVPT for behavior priors, CrafterCLIP for vision-language grounding, and CrafterSteve-1 for instruction following — along with datasets, reference agents, and benchmark evaluations. The paper's central claim is that this suite unlocks Crafter as a fast, prototyping-friendly testbed for general-purpose embodied agents, letting researchers iterate without the slow speed and engineering overhead of Minecraft.","feed_headline":"Crafter gets a foundation-model suite for fast agent prototyping","feed_subtitle":"Three pretrained models bring behavior, vision-language, and instruction-following skills to the Minecraft-like game.","key_machinery":"The suite's machinery is a trio of pretrained models and two dataset toolkits. CrafterVPT learns a behavior prior from unlabeled Crafter gameplay, giving agents a repertoire of sensible actions; CrafterCLIP aligns visual frames with caption text, supplying vision-language grounding; and CrafterSteve-1 combines these to follow natural-language instructions during play. CrafterPlay and CrafterCaption generate the behavior and caption datasets used to train and evaluate the models, while the attached benchmarks and reference agent implementations provide a standard way to measure progress. The whole stack carries the argument by supplying the missing infrastructure that previously made Minecraft the default choice for this type of research.","core_discovery":"The central claim is that CrafterDojo provides everything needed to treat Crafter as a general-purpose embodied-agent testbed: behavior priors that let agents act from visual observation, vision-language grounding that ties perception to language, instruction-following agents that convert natural-language commands into action, and the data-generation and evaluation tools to train and measure such agents. The claim is that these components, which previously existed only for the full Minecraft environment, transfer their design to Crafter without loss of relevance, making the lightweight environment a viable substrate for prototyping and benchmarking open-ended agents.","pith_inferences":["If the Crafter proxy holds, the same 'foundation-model enablement' recipe could be applied to other lightweight game-like environments, creating a pattern for fast-prototyping testbeds beyond Crafter.","The CrafterPlay and CrafterCaption datasets may become useful for offline RL and imitation-learning research that currently depends on expensive human demonstration or Minecraft-scale data.","A direct head-to-head: training identical agents in Crafter and Minecraft on matched tasks would let the community measure how much of Minecraft's challenge survives in Crafter, and adjust the proxy accordingly.","CrafterDojo could serve as a low-cost sanity-check layer before scaling methods to Minecraft or other high-fidelity simulators, reducing wasted large-scale runs."],"forward_implications":["Researchers can prototype new embodied-agent methods in Crafter in hours instead of the days usually spent on Minecraft infrastructure.","The same VPT/CLIP/Steve-style pipeline becomes reproducible and measurable in a lightweight setting, enabling controlled studies of these methods.","Standardized benchmarks and reference agents allow fair comparison of instruction-following and open-ended exploration agents.","The released datasets and codebase reduce the data-collection and engineering burden for embodied-AI labs with limited resources."],"supporting_citations":[],"fun_headline_variants":["CrafterDojo brings foundation models to lightweight agent prototyping","New foundation-model suite turns Crafter into an agent testbed","CrafterDojo unlocks Crafter for general-purpose agent research","Three pretrained models make Crafter a fast agent testbed"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The claim rests on Crafter being a sufficiently faithful stand-in for Minecraft's complexity — that advances made in Crafter inform or transfer to general-purpose embodied agents; if the simplified mechanics skip the capabilities that actually matter, the suite's value as a testbed collapses.","fun_headline_variants_meta":{"raw":{"variants":["CrafterDojo brings foundation models to lightweight agent prototyping","New foundation-model suite turns Crafter into an agent testbed","CrafterDojo unlocks Crafter for general-purpose agent research","Three pretrained models make Crafter a fast agent testbed"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000778,"raw_usage":{"total_tokens":3381,"prompt_tokens":826,"completion_tokens":2555,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":442,"completion_tokens_details":{"reasoning_tokens":2485}},"tokens_in":442,"tokens_out":2555,"duration_ms":20549,"temperature":1.0,"reasoning_tokens":2485,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T17:11:28.057611+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Train a batch of agents with CrafterDojo in Crafter and the same architectures directly in Minecraft on matched suites of instruction-following and exploration tasks; if strong Crafter performance consistently fails to correspond to or transfer to Minecraft performance, the central premise that Crafter captures the relevant challenges is falsified.","supporting_citations":[],"review_version":1}