Pith. sign in

hub

Balrog: Bench- marking agentic llm and vlm reasoning on games

19 Pith papers cite this work, alongside 2 external citations. Polarity classification is still indexing.

19 Pith papers citing it
2 external citations · external index

hub tools

citation-role summary

background 3 dataset 1

citation-polarity summary

years

2026 17 2025 2

polarities

background 3 unclear 1

representative citing papers

RTSGameBench: An RTS Benchmark for Strategic Reasoning by Vision-Language Models

cs.AI · 2026-06-17 · unverdicted · novelty 7.0

RTSGameBench is a new extensible benchmark for VLMs using diverse RTS matchups, diagnostic mini-games targeting individual competencies, and a self-evolving query-to-game generator, with results showing poor VLM performance on tight coordination and large-scale tasks.

Hierarchical Behaviour Spaces

cs.AI · 2026-04-27 · unverdicted · novelty 6.0

Hierarchical Behaviour Spaces uses linear combinations of reward functions to induce expressive behavior spaces in hierarchical RL, yielding strong performance on NetHack primarily through better exploration rather than long-term planning.

Resonant Minds: Closed-Loop Social Avatars with Theory of Mind

cs.CV · 2026-06-04 · unverdicted · novelty 5.0

A dual-agent closed-loop system integrates Theory of Mind reasoning with multimodal video generation to create social avatars that outperform full-information baselines on dialogue quality under information asymmetry.

citing papers explorer

Showing 19 of 19 citing papers.