Orak is a foundational benchmark providing training data, interfaces, and evaluation tools for LLM agents across diverse video game genres.
game over
4 Pith papers cite this work, alongside 9 external citations. Polarity classification is still indexing.
representative citing papers
Frontier vision-language models complete only 0.48% of VideoGameBench and 1.6% of its paused Lite version, a new real-time benchmark of 10 1990s games with raw visuals and minimal scaffolding.
OmniGameArena is a unified UE5 benchmark with 12 games and the IDC harness for cold-start scores and improvement dynamics of VLM agents.
Proof-of-concept shows fine-tuned small language models achieve adequate quality for real-time game content generation in a scoped RPG loop via retry-until-success and LLM-as-judge evaluation.
citing papers explorer
-
Orak: A Foundational Benchmark for Training and Evaluating LLM Agents on Diverse Video Games
Orak is a foundational benchmark providing training data, interfaces, and evaluation tools for LLM agents across diverse video game genres.
-
VideoGameBench: Can Vision-Language Models complete popular video games?
Frontier vision-language models complete only 0.48% of VideoGameBench and 1.6% of its paused Lite version, a new real-time benchmark of 10 1990s games with raw visuals and minimal scaffolding.
-
OmniGameArena: A Unified UE5 Benchmark for VLM Game Agents with Improvement Dynamics
OmniGameArena is a unified UE5 benchmark with 12 games and the IDC harness for cold-start scores and improvement dynamics of VLM agents.
-
High-quality generation of dynamic game content via small language models: A proof of concept
Proof-of-concept shows fine-tuned small language models achieve adequate quality for real-time game content generation in a scoped RPG loop via retry-until-success and LLM-as-judge evaluation.