Pith. sign in

REVIEW 5 cited by

MCU: An Evaluation Framework for Open-Ended Game Agents

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2310.08367 v4 pith:XPZMQSM6 submitted 2023-10-12 cs.AI cs.CLcs.CVcs.LG

classification cs.AIcs.CLcs.CVcs.LG
keywords agentsevaluationopen-endedtasksframeworkcapablediverseenvironments
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Developing AI agents capable of interacting with open-world environments to solve diverse tasks is a compelling challenge. However, evaluating such open-ended agents remains difficult, with current benchmarks facing scalability limitations. To address this, we introduce Minecraft Universe (MCU), a comprehensive evaluation framework set within the open-world video game Minecraft. MCU incorporates three key components: (1) an expanding collection of 3,452 composable atomic tasks that encompasses 11 major categories and 41 subcategories of challenges; (2) a task composition mechanism capable of generating infinite diverse tasks with varying difficulty; and (3) a general evaluation framework that achieves 91.5\% alignment with human ratings for open-ended task assessment. Empirical results reveal that even state-of-the-art foundation agents struggle with the increasing diversity and complexity of tasks. These findings highlight the necessity of MCU as a robust benchmark to drive progress in AI agent development within open-ended environments. Our evaluation code and scripts are available at https://github.com/CraftJarvis/MCU.

Discussion (0). Sign in to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Training One Model to Master Cross-Level Agentic Actions via Reinforcement Learning

    cs.LG 2025-12 conditional novelty 6.0 of 10

    CrossAgent learns step-level action-interface selection via a three-stage SFT + single-turn GRPO + multi-turn GRPO pipeline, reporting 54.6% mean success on 800+ Minecraft tasks after RL on only 30 tasks.

  2. ANNIE: Be Careful of Your Robots

    cs.AI 2025-09 conditional novelty 6.0 of 10

    The authors build a safety-centered benchmark and attack method that induces vision-language-action robot policies to violate ISO-based safety rules in a majority of tested episodes.

  3. Scalable Multi-Task Reinforcement Learning for Generalizable Spatial Intelligence in Visuomotor Agents

    cs.RO 2025-07 conditional novelty 6.0 of 10

    RL post-training on 100,000 synthesized cross-view Minecraft tasks raises interaction success from 7% to 28% and transfers zero-shot to DMLab, Unreal, and a real robot.

  4. EmbRACE-3K: Embodied Reasoning and Action in Complex Environments

    cs.CV 2025-07 conditional novelty 6.0 of 10

    EmbRACE-3K provides a photorealistic, closed-loop embodied benchmark with stepwise reasoning annotations, and fine-tuning Qwen2.5-VL on it improves embodied task success in simulation.

  5. Large Language Model Powered Intelligent Urban Agents: Concepts, Capabilities, and Applications

    cs.MA 2025-07 conditional novelty 4.0 of 10

    The paper defines urban LLM agents, surveys their sensing, memory, reasoning, execution, and learning workflows, and organizes their applications across planning, transportation, environment, safety, and society.

Pith tools