Pith. sign in

GameCraft-Bench: Can Agents Build Playable Games End-to-End in a Real Game Engine?

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it
abstract

Game generation is an emerging application of coding agents, requiring models to transform natural-language specifications into playable interactive systems. Unlike traditional coding tasks, game generation takes place within a game engine, where scripts, scenes, assets, rendering, and runtime interactions must jointly produce coherent gameplay. We formalize end-to-end game generation as the problem of producing a complete game artifact that realizes a specification through observable player-game interaction in a target environment. We argue that evaluating this setting requires three desiderata: Engine Grounding, Artifact Completeness, and Interactive Verification. We propose an interaction-grounded evaluation framework that assesses executable gameplay through replayed demonstrations and rubric-guided multimodal judging. We instantiate this framework as GameCraft-Bench, a benchmark comprising 140 Godot tasks across 15 game families. Evaluations of frontier coding agents show that end-to-end game generation remains highly challenging: the strongest agent achieves only 41.46%, and most agents score below 40%. Further analysis reveals that while agents often implement recognizable mechanics, they struggle to deliver complete games with sufficient content, functional visual feedback, and coherent presentation. See https://tongxuluo.github.io/gamecraft-bench-website for demos, code, and data.

fields

cs.AI 1

years

2026 1

verdicts

CONDITIONAL 1

representative citing papers

DashArena: Benchmarking LLMs on Interactive Analytic Dashboard Generation

cs.AI · 2026-08-11 · conditional · novelty 7.0

DashArena evaluates LLM-generated interactive dashboards by replaying each model's own interaction walkthrough in a browser and judging the evidence with a distilled VLM judge, showing that current frontier models frequently generate non-functional or analytically weak dashboards.

citing papers explorer

Showing 1 of 1 citing paper.

  • DashArena: Benchmarking LLMs on Interactive Analytic Dashboard Generation cs.AI · 2026-08-11 · conditional · none · ref 69 · internal anchor

    DashArena evaluates LLM-generated interactive dashboards by replaying each model's own interaction walkthrough in a browser and judging the evidence with a distilled VLM judge, showing that current frontier models frequently generate non-functional or analytically weak dashboards.