Pith. sign in

REVIEW 4 major objections 6 minor 88 references

TextAtari: 100K Frames Game Playing with Language Agents

T0 review · 4 major / 6 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read TextAtari turns Atari games into text and finds language agents fall below 10% of human performance on most tasks.

desk verdict Useful benchmark resource, but the 100k-step claim isn't supported by the 1k-step experiments. read the letter →

arxiv 2506.04098 v2 pith:AX234YOR submitted 2025-06-04 cs.CL cs.AIcs.LG

classification cs.CLcs.AIcs.LG
keywords TextAtarilong-horizondecision-makinglanguageagentsgamesbenchmarklargemodelschain-of-thoughtpromptingreflection
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper introduces TextAtari, a benchmark that renders classic Atari games as textual state descriptions so that large language models can play them through language alone, with planning horizons targeted up to 100,000 steps. The authors evaluate three open-source 7–8B models under three prompting frameworks and four knowledge-injection scenarios, and report that current language agents achieve less than 10% of human-level performance in over 90% of tested configurations. The central message is that these agents lack the sustained state tracking, strategic planning, and decision consistency needed for very long-horizon tasks, and that injecting game manuals or expert demonstrations helps far more than chain-of-thought or reflection prompting. A sympathetic reader would take this as a call for new memory and reasoning mechanisms, and as a standardized testbed to measure progress toward human-like extended planning.

What carries the argument

The central object is the TextAtari environment pipeline, built on the AtariARI wrapper, which turns the 128-byte RAM state of an Atari emulator into structured, human-readable textual observations (entity positions, scores, lives) without any visual input. Around this core, the paper builds four scenario conditions that vary only in auxiliary knowledge: Basic (no extra information), Obscured (domain nouns replaced with 'item'), Manual Augmentation (a game manual excerpt prepended), and Reference-based (an expert PPO trajectory subsampled into state-action pairs). These conditions, combined with the three prompting frameworks, isolate whether performance differences come from injected knowledge or from reasoning structure. The 100,000-step horizon is the benchmark's stated target, though the reported experiments use 1,000 interaction steps per episode for cost reasons.

What would settle it

Run the same three agent configurations (Basic, CoT, Reflection) on a subset of TextAtari games with a true 100,000-step horizon and compare per-game normalized scores against the 1,000-step results; the central claim would be disconfirmed if the performance gap relative to humans does not grow with horizon length, or if any configuration exceeds 10% of human performance on a majority of tasks at the longer horizon.

Watch

Extended reading notes

Core claim

The paper claims that language agents, despite strong short-context reasoning, fail dramatically when asked to maintain coherent decision-making over long horizons in text-rendered Atari games. Across nearly 100 tasks built from 23 Atari games, all three tested models (Qwen2.5-7B, Gemma-7B, Llama3.1-8B) and all three agent frameworks (zero-shot, chain-of-thought, reflection) scored below 10% of human performance in more than 90% of scenarios, with only two isolated cases approaching or exceeding human level. The most consistent performance gains came from prior knowledge injection: game manuals and expert demonstrations each produced average improvements exceeding 100%, whereas chain-of-thought and reflection showed inconsistent and often negligible benefits. The authors interpret this as evidence that current language models lack the cognitive mechanisms for extended reasoning, and they position TextAtari as a benchmark that fills the near-empty niche of 100,000-step evaluation.

Load-bearing premise

The entire 'long-horizon failure' conclusion rests on the assumption that 1,000 interaction steps behave like 100,000 steps, an extrapolation the paper itself flags as untested in its appendix.

Editorial extensions

If this is right

  • If the benchmark's conclusions hold, the community gets a standardized, reproducible protocol for measuring long-horizon language-agent capability, with public baselines that future work can compare against directly.
  • Knowledge injection, not stronger prompting, is the most reliable lever for improving language-agent gameplay, suggesting that external grounding (manuals, demonstrations) should be a default component in agent design.
  • The failure of chain-of-thought and reflection to help consistently implies that explicit reasoning alone does not fix state tracking or planning in these settings, shifting attention toward memory architectures and training objectives.
  • The horizon-gap analysis, which found that fewer than 2.5% of 163 surveyed benchmarks reach 100,000 steps, would justify prioritizing new ultra-long-horizon evaluation environments.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The 1,000-step evaluation protocol may not capture the challenges that appear only after tens of thousands of steps, so the paper's strongest long-horizon conclusions are an extrapolation the data does not yet fully support; running a subset of games to 10,000 or 100,000 steps would test whether the performance gap widens, stays flat, or even narrows.
  • The Obscured scenario, which strips away game-specific nouns, could be repurposed as a diagnostic for whether a model reasons from spatial and numerical structure rather than from memorized lexical patterns, potentially separating genuine planning from semantic priors.
  • The finding that expert trajectories help most suggests that in-context imitation, rather than self-critique, is the effective mode for these models; future work could test whether combining trajectory priming with lightweight memory of past episodes outperforms each alone.
  • If the 100,000-step target is taken seriously, the benchmark's cost of roughly 820,000 GPU-minutes will force the community to develop cheaper proxies (e.g., curriculum horizons or learned state abstractions) to make progress measurable at all.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 6 minor

Summary. The paper introduces TextAtari, a benchmark that converts Atari 2600 game states into textual observations via the AtariARI wrapper and proposes four prompting scenarios (Basic, Obscured, Manual Augmentation, Reference-based) to study how prior knowledge affects language-agent game playing. The authors evaluate Qwen2.5-7B, Gemma-7B, and Llama3.1-8B with zero-shot, chain-of-thought, and reflection agents on 23 games, reporting that LLM agents fall below 10% of human performance in over 90% of scenarios and that knowledge injection improves performance more than reasoning-enhancement techniques. The paper claims to target horizons of up to 100,000 steps, but all experiments use a 1,000-step horizon, and no human baseline scores are reported.

Significance. If the central claims were fully supported, TextAtari would be a valuable standardized testbed for long-horizon language-agent evaluation, and the four-scenario design is a genuinely useful controlled comparison. The paper contributes a large catalog of existing benchmarks, an open-source implementation, and a reproducible protocol outline. However, the headline long-horizon claim (100,000 steps) is not demonstrated by any experiment, and the human-relative performance claims are unverifiable without reported human baselines or raw scores. The qualitative finding that injecting manuals or expert trajectories helps more than CoT/reflection is plausible but currently rests on relative-difference plots that are not backed by numeric tables or error bars. These gaps are load-bearing because the paper's novelty and title rest on the 100K-frame claim and on the quantitative comparison to human players.

major comments (4)
  1. [Section 3.1 and Appendix E] All experiments in the paper use a horizon of 1,000 interaction steps, yet the title, abstract, and Section 2 state that TextAtari evaluates 'very long-horizon decision-making tasks spanning up to 100,000 steps.' Appendix E explicitly concedes that the reduced horizon 'may not fully capture the challenges of extremely long-horizon reasoning.' Since the 100K-step capability is the paper's defining novelty and the stated motivation for the benchmark, the absence of any experiment at or near that horizon leaves the central claim unsupported. At minimum, the authors should either report results at longer horizons (e.g., a subset of games at 10K or 100K steps) or reformulate the title and abstract to describe the actual 1,000-step evaluation.
  2. [Section 3.3, Appendix E] The headline quantitative result that language agents fall 'below 10% of human capability' in over 90% of scenarios is not verifiable from the paper: no human performance scores are reported, the human reference is not defined (ALE human-normalized scores, a new human study, or the PPO 'average human' threshold treated as human performance), and no absolute scores or confidence intervals are given for any model/scenario combination. The bar charts in Appendix E show only relative differences and normalized averages, so the reader cannot check the claimed 10% figure or the 'average improvement exceeding 100%' statement. The authors should include a table of mean scores with standard errors alongside the human baseline for every game.
  3. [Section 2.3, Limitations] The token-consumption report is internally inconsistent. Section 2.3 states that 'the most token-intensive games consumed between 40k-50k tokens per decision step across all models,' which would imply roughly 45 million tokens per 1,000-step run and tens of billions of tokens for the full evaluation, consistent with the 'billions of tokens' phrase. However, Section 2.2 says each observation is length-limited to fit within the context window, and Gemma-7B and Qwen2.5-7B have 8K/32K contexts, so a 40K-50K token per-step prompt cannot fit without truncation for those models. If the intended figure is 40-50k tokens per episode (i.e., 40-50 tokens per step), then the claimed 'billions of tokens' is overestimated by three orders of magnitude. Please clarify the unit and report per-model token counts.
  4. [Section 3.1, Section 2.2] The claim that observed performance differences across scenarios 'can be attributed to the injected knowledge rather than disparities in prompt size or compute budget' is not justified by the protocol. Because augmentation increases prompt length, the sliding-window policy discards different amounts of history for different scenarios, so the effective information available to the model differs beyond the intended knowledge injection. A model in the Manual or Reference-based condition may lose older observations sooner than in the Basic condition, confounding the comparison. The authors should control for prompt size (e.g., equalizing token budgets or reporting the effective history length per condition) before drawing conclusions about the relative benefit of knowledge injection.
minor comments (6)
  1. [Abstract, Section 2.1] The abstract says 'nearly 100 distinct tasks' whereas Section 2.1 reports 23 classic Atari games; please reconcile these numbers.
  2. [Appendix E] Appendix E uses 'RL Trajectory' for the 'Reference-based' scenario and 'Reflexion_last'/'Reflexion_max' for the Reflection agent; please unify terminology with the main text to avoid confusion.
  3. [Table 3, SmartPlay row] The horizon entry '100, 100000, 200, 5' for SmartPlay appears to be a typo; please correct it.
  4. [Appendix E, all figures] Please add a table of mean scores and standard deviations per game/model/scenario/agent; the relative-difference bar charts alone are insufficient for reproducibility and for verifying the paper's quantitative claims.
  5. [Section 2.2] The paper mentions a 5 Hz observation throttle; specify the frame-skipping scheme and the number of frames per interaction step so that the relation between 'frames' and 'steps' in the title is clear.
  6. [Section 2.3] Typo: 'across millions steps' should be 'across millions of steps.'

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the reported results are measured rollouts of fixed LLM agents, not fitted predictions; the 100K-step framing is an unsupported scope claim rather than a circular derivation.

full rationale

The paper contains no formal derivation chain, fitted parameters, or machine-checked theorem whose conclusion is equivalent to its inputs. The central empirical claims are direct measurements of LLM agents interacting with the TextAtari environment under fixed prompts, so they are not circular by construction. The Reference-based scenario supplies a PPO-generated expert trajectory as an input condition; the observed benefit of demonstrations is an empirical outcome of that condition, not a quantity derived from the trajectory itself. The 100,000-step claim is stated in the title, abstract, and Section 2, but Section 3.1 fixes the evaluation horizon at 1,000 interaction steps, and Appendix E explicitly concedes that this reduced horizon 'may not fully capture the challenges of extremely long-horizon reasoning.' That is a validity and generalizability limitation, not a circularity: the benchmark infrastructure may support longer episodes, but the experiments do not demonstrate the claimed long-horizon regime. Similarly, the comparison to human capability relies on the stated PPO training target of 'at least average human performance' without in-paper human baselines; this is missing evidence, not a self-referential reduction. The paper does not lean on load-bearing self-citations; AtariARI and the surveyed benchmarks are external prior work. No specific equation or fitted value is reused as a prediction, and no claimed result is equivalent to its own input by definition. Therefore the appropriate circularity score is 0.

Assumptions & free parameters 3 free parameters · 3 assumptions · 0 invented entities

The central results rest on unverified baselines (human and PPO 'expert' performance), the assumption that AtariARI text preserves actionable game state, the assumption that context trimming leaves comparisons unbiased, and design choices like the 1,000-step horizon and 5 Hz throttle. No free parameters are fitted in a mathematical model; the listed parameters are hand-chosen evaluation settings that materially affect what the benchmark measures.

free parameters (3)
  • evaluation horizon = 1000 steps
    All experiments use 1,000 steps, not the 100,000 steps claimed in the title and abstract; this cost-saving choice may not capture long-horizon effects.
  • observation frequency = ~5 Hz
    Throttled observation generation; this hand-chosen rate reduces decisions per real second and may affect game performance.
  • reference trajectory subsampling = every 10th state-action pair, 400-token block
    Arbitrary subsampling choices for expert demonstrations are likely to affect Reference-based results.
assumptions (3)
  • domain assumption AtariARI RAM annotations preserve sufficient game state for competent decision-making.
    Section 2.2 assumes RAM-derived labels capture key information; if essential visual details are lost, the benchmark measures representation degradation rather than reasoning.
  • domain assumption The PPO controllers used to generate reference trajectories reach at least average human performance.
    Section 2.2 Reference-based scenario states this, but no human or PPO scores are reported, so the 'expert' status of demonstrations is unverified.
  • domain assumption Sliding-window context trimming does not distort comparisons across scenarios.
    Section 3.1 discards oldest messages when the prompt approaches the context limit; the paper asserts differences are due to injected knowledge, not prompt size, but this is not tested.

how reviews work

0 comments
Cite this review

Pith. "Pith review of TextAtari: 100K Frames Game Playing with Language Agents." pith.science (2026). https://pith.science/paper/AX234YOR

@misc{pith2026250604098,
  author       = {Pith},
  title        = {Pith review of: TextAtari: 100K Frames Game Playing with Language Agents},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/AX234YOR}},
  note         = {Machine review of arXiv:2506.04098}
}
read the original abstract

We present TextAtari, a benchmark for evaluating language agents on very long-horizon decision-making tasks spanning up to 100,000 steps. By translating the visual state representations of classic Atari games into rich textual descriptions, TextAtari creates a challenging test bed that bridges sequential decision-making with natural language processing. The benchmark includes nearly 100 distinct tasks with varying complexity, action spaces, and planning horizons, all rendered as text through an unsupervised representation learning framework (AtariARI). We evaluate three open-source large language models (Qwen2.5-7B, Gemma-7B, and Llama3.1-8B) across three agent frameworks (zero-shot, few-shot chain-of-thought, and reflection reasoning) to assess how different forms of prior knowledge affect performance on these long-horizon challenges. Four scenarios-Basic, Obscured, Manual Augmentation, and Reference-based-investigate the impact of semantic understanding, instruction comprehension, and expert demonstrations on agent decision-making. Our results reveal significant performance gaps between language agents and human players in extensive planning tasks, highlighting challenges in sequential reasoning, state tracking, and strategic planning across tens of thousands of steps. TextAtari provides standardized evaluation protocols, baseline implementations, and a framework for advancing research at the intersection of language models and planning. Our code is available at https://github.com/Lww007/Text-Atari-Agents.

Figures

Figures reproduced from arXiv: 2506.04098 by the authors.

Figure 1
Figure 1. Statistics of tasks and horizons. Abstract We present TextAtari, a comprehensive benchmark for evaluating language agents on very long-horizon decision-making tasks spanning up to 100, 000 steps. By translating the visual state representations of classic Atari games into rich tex￾tual descriptions, TextAtari creates a challenging test bed that bridges sequential decision-making with natural language processing. Our … view at source ↗
Figure 2
Figure 2. Screenshot from Kwa et al. (2025). The AI community has developed numerous benchmarks for evaluating sequential decision￾making, spanning web interfaces, desktop soft￾ware, games, and embodied environments (Tan et al., 2024). However, our comprehensive anal￾ysis of 163 existing sequential decision bench￾marks (see Appendix for the full list) reveals a critical limitation: most operate on remarkably short horizons. W… view at source ↗
Figure 3
Figure 3. Horizon statistics. In this work, we introduce TextAtari, a comprehensive benchmark for evaluating language agents on very long￾horizon decision-making tasks spanning up to 100, 000 steps. TextAtari transforms the visual states of clas￾sic Atari games into rich textual descriptions using an unsupervised representation learning framework (Atari￾ARI) (Anand et al., 2019), creating a challenging testbed that bridges se… view at source ↗
Figures from the paper (36 more)
Figure 4
Figure 4. Figure 4: Environment verbalization with AtariARI. [PITH_FULL_IMAGE:figures/full_fig_p005_4.png]
Figure 5
Figure 5. Figure 5: Computational costs. Token consumption (lower right figure) further emphasizes the scale of these experiments. The most token-intensive games consumed between 40k-50k tokens per decision step across all models, with LLaMA3.1-8B consistently using more tokens than its c…
Figure 6
Figure 6. Figure 6: Relative perfor￾mances. Chain-of-Thought. The CoT agent extends the Basic template by demanding explicit step-by-step reasoning. A leading system in￾struction frames the model as an expert Atari player and mandates a JSON reply with two keys–"thought process" and "acti…
Figure 7
Figure 7. Figure 7: Selected performance comparison. See Appendix for more details. [PITH_FULL_IMAGE:figures/full_fig_p009_7.png]
Figure 8
Figure 8. Figure 8: Performance comparison of Qwen2.5-7B with Naive agent across different scenarios. Each [PITH_FULL_IMAGE:figures/full_fig_p019_8.png]
Figure 9
Figure 9. Figure 9: Performance comparison of Qwen2.5-7B with Chain-of-Thought (CoT) agent across [PITH_FULL_IMAGE:figures/full_fig_p019_9.png]
Figure 10
Figure 10. Figure 10: Performance comparison of Qwen2.5-7B with Reflexion_last agent (using the most recent [PITH_FULL_IMAGE:figures/full_fig_p020_10.png]
Figure 11
Figure 11. Figure 11: Performance comparison of Qwen2.5-7B with Reflexion_max agent (using the best [PITH_FULL_IMAGE:figures/full_fig_p020_11.png]
Figure 12
Figure 12. Figure 12: Performance comparison of Llama3.1-8B with Naive agent across different scenarios. [PITH_FULL_IMAGE:figures/full_fig_p021_12.png]
Figure 13
Figure 13. Figure 13: Performance comparison of Llama3.1-8B with Chain-of-Thought (CoT) agent across [PITH_FULL_IMAGE:figures/full_fig_p021_13.png]
Figure 14
Figure 14. Figure 14: Performance comparison of Llama3.1-8B with Reflexion_last agent (using the most recent [PITH_FULL_IMAGE:figures/full_fig_p022_14.png]
Figure 15
Figure 15. Figure 15: Performance comparison of Llama3.1-8B with Reflexion_max agent (using the best [PITH_FULL_IMAGE:figures/full_fig_p022_15.png]
Figure 16
Figure 16. Figure 16: Performance comparison of Gemma-7B with Naive agent across different scenarios. Each [PITH_FULL_IMAGE:figures/full_fig_p023_16.png]
Figure 17
Figure 17. Figure 17: Performance comparison of Gemma-7B with Chain-of-Thought (CoT) agent across [PITH_FULL_IMAGE:figures/full_fig_p023_17.png]
Figure 18
Figure 18. Figure 18: Performance comparison of Gemma-7B with Reflexion_last agent (using the most recent [PITH_FULL_IMAGE:figures/full_fig_p024_18.png]
Figure 19
Figure 19. Figure 19: Performance comparison of Gemma-7B with Reflexion_max agent (using the best [PITH_FULL_IMAGE:figures/full_fig_p024_19.png]
Figure 20
Figure 20. Figure 20: Performance comparison of Qwen2.5-7B in the Basic scenario across different agent types. [PITH_FULL_IMAGE:figures/full_fig_p025_20.png]
Figure 21
Figure 21. Figure 21: Performance comparison of Qwen2.5-7B in the Obscured scenario across different agent [PITH_FULL_IMAGE:figures/full_fig_p025_21.png]
Figure 22
Figure 22. Figure 22: Performance comparison of Qwen2.5-7B in the Game Manual scenario across different [PITH_FULL_IMAGE:figures/full_fig_p026_22.png]
Figure 23
Figure 23. Figure 23: Performance comparison of Qwen2.5-7B in the RL Trajectory scenario across different [PITH_FULL_IMAGE:figures/full_fig_p026_23.png]
Figure 24
Figure 24. Figure 24: Performance comparison of Llama3.1-8B in the Basic scenario across different agent types. [PITH_FULL_IMAGE:figures/full_fig_p027_24.png]
Figure 25
Figure 25. Figure 25: Performance comparison of Llama3.1-8B in the Obscured scenario across different agent [PITH_FULL_IMAGE:figures/full_fig_p027_25.png]
Figure 26
Figure 26. Figure 26: Performance comparison of Llama3.1-8B in the Game Manual scenario across different [PITH_FULL_IMAGE:figures/full_fig_p028_26.png]
Figure 27
Figure 27. Figure 27: Performance comparison of Llama3.1-8B in the RL Trajectory scenario across different [PITH_FULL_IMAGE:figures/full_fig_p028_27.png]
Figure 28
Figure 28. Figure 28: Performance comparison of Gemma-7B in the Basic scenario across different agent types. [PITH_FULL_IMAGE:figures/full_fig_p029_28.png]
Figure 29
Figure 29. Figure 29: Performance comparison of Gemma-7B in the Obscured scenario across different agent [PITH_FULL_IMAGE:figures/full_fig_p029_29.png]
Figure 30
Figure 30. Figure 30: Performance comparison of Gemma-7B in the Game Manual scenario across different [PITH_FULL_IMAGE:figures/full_fig_p036_30.png]
Figure 31
Figure 31. Figure 31: Performance comparison of Gemma-7B in the RL Trajectory scenario across different [PITH_FULL_IMAGE:figures/full_fig_p037_31.png]
Figure 32
Figure 32. Figure 32: Cross-model performance comparison in the Basic scenario using Naive agent (top row) [PITH_FULL_IMAGE:figures/full_fig_p037_32.png]
Figure 33
Figure 33. Figure 33: Cross-model performance comparison in the Basic scenario using Reflexion_last agent [PITH_FULL_IMAGE:figures/full_fig_p038_33.png]
Figure 34
Figure 34. Figure 34: Cross-model performance comparison in the Obscured scenario using Naive agent (top [PITH_FULL_IMAGE:figures/full_fig_p038_34.png]
Figure 35
Figure 35. Figure 35: Cross-model performance comparison in the Obscured scenario using Reflexion_last agent [PITH_FULL_IMAGE:figures/full_fig_p039_35.png]
Figure 36
Figure 36. Figure 36: Cross-model performance comparison in the Game Manual scenario using Naive agent [PITH_FULL_IMAGE:figures/full_fig_p039_36.png]
Figure 37
Figure 37. Figure 37: Cross-model performance comparison in the Game Manual scenario using Reflexion_last [PITH_FULL_IMAGE:figures/full_fig_p040_37.png]
Figure 38
Figure 38. Figure 38: Cross-model performance comparison in the RL Trajectory scenario using Naive agent [PITH_FULL_IMAGE:figures/full_fig_p040_38.png]
Figure 39
Figure 39. Figure 39: Cross-model performance comparison in the RL Trajectory scenario using Reflexion_last [PITH_FULL_IMAGE:figures/full_fig_p041_39.png]

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

88 extracted references · 26 canonical work pages

  1. [1]

    J., and Yao, Z

    Aghzal, M., Plaku, E., Stein, G. J., and Yao, Z. A survey on large language models for automated planning.arXiv preprint arXiv:2502.12435,

  2. [3]

    O., and Clune, J

    Ecoffet, A., Huizinga, J., Lehman, J., Stanley, K. O., and Clune, J. Go-explore: a new approach for hard-exploration problems.arXiv preprint arXiv:1901.10995,

  3. [4]

    Hallmonster

    Game Detailed Description Venture An exploration game where players control Winky, an adventurer navigating through a multi-room dungeon to collect treasures. Each room contains different monsters guarding treasure, requiring specific strategies to overcome. Players view the dungeon layout from an overhead perspective but transition to a zoomed-in view wh...

  4. [5]

    Openai o1 system card.arXiv preprint arXiv:2412.16720,

    Jaech, A., Kalai, A., Lerer, A., Richardson, A., El-Kishky, A., Low, A., Helyar, A., Madry, A., Beutel, A., Carney, A., et al. Openai o1 system card.arXiv preprint arXiv:2412.16720,

  5. [7]

    Measuring ai ability to complete long tasks.arXiv preprint arXiv:2503.14499,

    Kwa, T., West, B., Becker, J., Deng, A., Garcia, K., Hasin, M., Jawhar, S., Kinniment, M., Rush, N., V on Arx, S., et al. Measuring ai ability to complete long tasks.arXiv preprint arXiv:2503.14499,

  6. [9]

    A survey of temporal credit assignment in deep reinforcement learning.arXiv preprint arXiv:2312.01072,

    Pignatelli, E., Ferret, J., Geist, M., Mesnard, T., van Hasselt, H., Pietquin, O., and Toni, L. A survey of temporal credit assignment in deep reinforcement learning.arXiv preprint arXiv:2312.01072,

  7. [11]

    Dichotomy of control: Separating what you can control from what you cannot.arXiv preprint arXiv:2210.13435,

    Yang, M., Schuurmans, D., Abbeel, P., and Nachum, O. Dichotomy of control: Separating what you can control from what you cannot.arXiv preprint arXiv:2210.13435,

  8. [12]

    Video as the new language for real-world decision making.arXiv preprint arXiv:2402.17139,

    Yang, S., Walker, J., Parker-Holder, J., Du, Y ., Bruce, J., Barreto, A., Abbeel, P., and Schuurmans, D. Video as the new language for real-world decision making.arXiv preprint arXiv:2402.17139,

Show all 88 references
  1. [13]

    horizon gap

    10 A TextAtari Supplementary Material • Section B: Related Work • Section C: Task Details • Section D: Prompt Engineering • Section E: Missing Results • Section F: Border Impact B Related Work This section provides a comprehensive analysis of existing sequential decision-makin...

  2. [14]

    Software 50 2000 single-agent46 HCAST (Rein et al.,

  3. [15]

    Software, Web 5, 50 200 single-agent55 TheAgentCompany (Xu et al., 2024a) Software, Web 10, 50 200 single-agent like StarCraft II (Ma et al., 2024), Red Dead Redemption II (Tan et al., 2024), and Minecraft variants (MineDojo (Fan et al., 2022), Mars (Tang et al., 2024), MineLa...

  4. [16]

    Video Game 2000 5000 multi-agent96 APIBench (Peng et al.,

  5. [18]

    Web 50 2000 single-agent108 METAGUI (Sun et al.,

  6. [19]

    13 Benchmarks like MLE-Bench (Chan et al., 2024), RE-Bench (Wijk et al., 2024), and DISCOVERY- WORLD (Jansen et al.,

    similarly restricts itself to basic grid-world scenarios with limited objects and interactions, creating artificially simplified planning problems. 13 Benchmarks like MLE-Bench (Chan et al., 2024), RE-Bench (Wijk et al., 2024), and DISCOVERY- WORLD (Jansen et al.,

  7. [20]

    ghost,” “paddle,

    limit themselves to specialized domains (machine learning experiments and scientific discovery) that contain substantial human annotation and guidance. These embed- ded hints and structured exploration spaces implicitly simplify the planning challenge compared to TextAtari’s m...

  8. [21]

    Software 5 10 single-agent 136 AsyncHow (Lin et al., 2024a) Text Game 10 2000 single-agent 137 VirtualHome (Puig et al.,

  9. [22]

    Web 20 2000 single-agent 140 HandMeThat (Wan et al.,

  10. [23]

    Software, Web 5 2000 single-agent 149 ToolLens (Qu et al.,

  11. [24]

    Software 2000 100 single-agent 154 RE-Bench (Wijk et al.,

  12. [26]

    Agashe, S., Fan, Y ., Reyna, A., and Wang, X. E. Llm-coordination: evaluating and analyzing multi-agent coordination abilities in large language models.arXiv preprint arXiv:2310.03903,

  13. [27]

    Werewolf arena: A case study in llm evaluation via social deduction.arXiv preprint arXiv:2407.13943,

    Bailis, S., Friedhoff, J., and Chen, F. Werewolf arena: A case study in llm evaluation via social deduction.arXiv preprint arXiv:2407.13943,

  14. [28]

    K., et al

    Bonatti, R., Zhao, D., Dupont, D., Abdali, S., Li, Y ., Lu, Y ., Wagle, J., Koishida, K., Bucker, A., Jang, L. K., et al. Windows agent arena: Evaluating multi-modal os agents at scale. InNeurIPS 2024 Workshop on Open-World Agents,

  15. [29]

    A3: Android agent arena for mobile gui agents.arXiv preprint arXiv:2501.01149,

    Chai, Y ., Li, H., Zhang, J., Liu, L., Liu, G., Wang, G., Ren, S., Huang, S., and Li, H. A3: Android agent arena for mobile gui agents.arXiv preprint arXiv:2501.01149,

  16. [30]

    clembench: Using game play to evaluate chat-optimized language models as conversational agents

    Chalamalasetti, K., Götze, J., Hakimov, S., Madureira, B., Sadler, P., and Schlangen, D. clembench: Using game play to evaluate chat-optimized language models as conversational agents. InPro- ceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, ...

  17. [31]

    S., Chowdhury, N., Jaffe, O., Aung, J., Sherburn, D., Mays, E., Starace, G., Liu, K., Maksin, L., Patwardhan, T., et al

    Chan, J. S., Chowdhury, N., Jaffe, O., Aung, J., Sherburn, D., Mays, E., Starace, G., Liu, K., Maksin, L., Patwardhan, T., et al. Mle-bench: Evaluating machine learning agents on machine learning engineering.arXiv preprint arXiv:2410.07095,

  18. [32]

    D., Desai, R., Hlavac, M., Karashchuk, V ., Krantz, J., Mottaghi, R., Parashar, P., et al

    Chang, M., Chhablani, G., Clegg, A., Cote, M. D., Desai, R., Hlavac, M., Karashchuk, V ., Krantz, J., Mottaghi, R., Parashar, P., et al. Partnr: A benchmark for planning and reasoning in embodied multi-agent tasks.arXiv preprint arXiv:2411.00081,

  19. [33]

    Alchemy: A quantum chemistry dataset for benchmarking ai models.arXiv preprint arXiv:1906.09427,

    Chen, G., Chen, P., Hsieh, C.-Y ., Lee, C.-K., Liao, B., Liao, R., Liu, W., Qiu, J., Sun, Q., Tang, J., et al. Alchemy: A quantum chemistry dataset for benchmarking ai models.arXiv preprint arXiv:1906.09427,

  20. [35]

    Cloos, N., Jens, M., Naim, M., Kuo, Y .-L., Cases, I., Barbu, A., and Cueva, C. J. Baba is ai: Break the rules to beat the benchmark. InICML 2024 Workshop on LLMs and Cognition,

  21. [36]

    M., and Yadav, A

    Costarelli, A., Allen, M., Hauksson, R., Sodunke, G., Hariharan, S., Cheng, C., Li, W., Clymer, J. M., and Yadav, A. Gamebench: Evaluating strategic reasoning abilities of llm agents. InLanguage Gamification-NeurIPS 2024 Workshop,

  22. [37]

    Textworld: A learning environment for text-based games

    Côté, M.-A., Kádár, A., Yuan, X., Kybartas, B., Barnes, T., Fine, E., Moore, J., Hausknecht, M., El Asri, L., Adada, M., et al. Textworld: A learning environment for text-based games. In Computer Games: 7th Workshop, CGW 2018, Held in Conjunction with the 27th International Co...

  23. [38]

    R., Veselovsky, V ., Josifoski, M., Peyrard, M., Bosselut, A., Kosinski, M., and West, R

    Davidson, T. R., Veselovsky, V ., Josifoski, M., Peyrard, M., Bosselut, A., Kosinski, M., and West, R. Evaluating language model agency through negotiations. InICLR 2024,

  24. [39]

    Villageragent: A graph-based multi-agent framework for coordinating complex task dependencies in minecraft

    Dong, Y ., Zhu, X., Pan, Z., Zhu, L., and Yang, Y . Villageragent: A graph-based multi-agent framework for coordinating complex task dependencies in minecraft. InFindings of the Association for Computational Linguistics ACL 2024, pp. 16290–16314,

  25. [40]

    Assistgui: Task-oriented desktop graphical user interface automation.arXiv preprint arXiv:2312.13108,

    Gao, D., Ji, L., Bai, Z., Ouyang, M., Li, P., Mao, D., Wu, Q., Zhang, W., Wang, P., Guo, X., et al. Assistgui: Task-oriented desktop graphical user interface automation.arXiv preprint arXiv:2312.13108,

  26. [41]

    Mindagent: Emergent gaming interaction

    Gong, R., Huang, Q., Ma, X., Noda, Y ., Durante, Z., Zheng, Z., Terzopoulos, D., Fei-Fei, L., Gao, J., and V o, H. Mindagent: Emergent gaming interaction. InFindings of the Association for Computational Linguistics: NAACL 2024, pp. 3154–3183,

  27. [42]

    S., Sunkara, N., and Choudhury, S

    Gonzalez-Pumariega, G., Yean, L. S., Sunkara, N., and Choudhury, S. Robotouille: An asynchronous planning benchmark for llm agents.arXiv preprint arXiv:2502.05227,

  28. [43]

    Textarena.arXiv preprint arXiv:2504.11442,

    Guertler, L., Cheng, B., Yu, S., Liu, B., Choshen, L., and Tan, C. Textarena.arXiv preprint arXiv:2504.11442,

  29. [44]

    Stabletoolbench-mirrorapi: Modeling tool environments as mirrors of 7,000+ real-world apis.arXiv preprint arXiv:2503.20527,

    Guo, Z., Cheng, S., Niu, Y ., Wang, H., Zhou, S., Huang, W., and Liu, Y . Stabletoolbench-mirrorapi: Modeling tool environments as mirrors of 7,000+ real-world apis.arXiv preprint arXiv:2503.20527,

  30. [45]

    Hill, W., Liu, I., Koch, A. D. M., Harvey, D., Kumar, N., Konidaris, G., and James, S. Mineplanner: A benchmark for long-horizon planning in large minecraft worlds.arXiv preprint arXiv:2312.12891,

  31. [46]

    Gamearena: Evaluating llm reasoning through live computer games.arXiv preprint arXiv:2412.06394,

    Hu, L., Li, Q., Xie, A., Jiang, N., Stoica, I., Jin, H., and Zhang, H. Gamearena: Evaluating llm reasoning through live computer games.arXiv preprint arXiv:2412.06394,

  32. [47]

    Game-theoretic llm: Agent workflow for negotiation games.arXiv preprint arXiv:2411.05990,

    Hua, W., Liu, O., Li, L., Amayuelas, A., Chen, J., Jiang, L., Jin, M., Fan, L., Sun, F., Wang, W., et al. Game-theoretic llm: Agent workflow for negotiation games.arXiv preprint arXiv:2411.05990,

  33. [48]

    Mlagentbench: Evaluating language agents on machine learning experimentation

    Huang, Q., V ora, J., Liang, P., and Leskovec, J. Mlagentbench: Evaluating language agents on machine learning experimentation. InInternational Conference on Machine Learning, pp. 20271– 20309. PMLR, 2024a. Huang, S., Zhong, W., Lu, J., Zhu, Q., Gao, J., Liu, W., Hou, Y ., Zen...

  34. [50]

    Elt-bench: An end-to-end benchmark for evaluating ai agents on elt pipelines.arXiv preprint arXiv:2504.04808,

    Jin, T., Zhu, Y ., and Kang, D. Elt-bench: An end-to-end benchmark for evaluating ai agents on elt pipelines.arXiv preprint arXiv:2504.04808,

  35. [51]

    L., and Jin, C

    Karten, S., Nguyen, A. L., and Jin, C. Pok \’echamp: an expert-level minimax language agent.arXiv preprint arXiv:2503.04094,

  36. [52]

    Benchmarking mobile device control agents across diverse configurations.arXiv preprint arXiv:2404.16660,

    Lee, J., Min, T., An, M., Hahm, D., Lee, H., Kim, C., and Lee, K. Benchmarking mobile device control agents across diverse configurations.arXiv preprint arXiv:2404.16660,

  37. [53]

    Sheetcopilot: Bringing software productivity to the next level through large language models.Advances in Neural Information Processing Systems, 36:4952–4984, 2023a

    Li, H., Su, J., Chen, Y ., Li, Q., and ZHANG, Z.-X. Sheetcopilot: Bringing software productivity to the next level through large language models.Advances in Neural Information Processing Systems, 36:4952–4984, 2023a. Li, H., Cao, Y ., Yu, Y ., Javaji, S. R., Deng, Z., He, Y .,...

  38. [54]

    Avalonbench: Evaluating llms playing the game of avalon

    45 Light, J., Cai, M., Shen, S., and Hu, Z. Avalonbench: Evaluating llms playing the game of avalon. In NeurIPS 2023 Foundation Models for Decision Making Workshop,

  39. [55]

    Visescape: A benchmark for evaluating exploration-driven decision-making in virtual escape rooms.arXiv preprint arXiv:2503.14427,

    Lim, S., Kim, S., Yu, J., Lee, S., Chung, J., and Yu, Y . Visescape: A benchmark for evaluating exploration-driven decision-making in virtual escape rooms.arXiv preprint arXiv:2503.14427,

  40. [56]

    Robocasa: Large-scale simulation of everyday tasks for generalist robots

    Nasiriany, S., Maddukuri, A., Zhang, L., Parikh, A., Lo, A., Joshi, A., Mandlekar, A., and Zhu, Y . Robocasa: Large-scale simulation of everyday tasks for generalist robots. InRSS 2024 Workshop: Data Generation for Robotics,

  41. [57]

    Mlgym: A new framework and benchmark for advancing ai research agents.arXiv preprint arXiv:2502.14499,

    46 Nathani, D., Madaan, L., Roberts, N., Bashlykov, N., Menon, A., Moens, V ., Budhiraja, A., Magka, D., V orotilov, V ., Chaurasia, G., et al. Mlgym: A new framework and benchmark for advancing ai research agents.arXiv preprint arXiv:2502.14499,

  42. [58]

    Balrog: Benchmarking agentic llm and vlm reasoning on games

    Paglieri, D., Cupiał, B., Coward, S., Piterbarg, U., Wolczyk, M., Khan, A., Pignatelli, E., Kuci´nski, Ł., Pinto, L., Fergus, R., et al. Balrog: Benchmarking agentic llm and vlm reasoning on games. arXiv preprint arXiv:2411.13543,

  43. [59]

    Webcanvas: Benchmarking web agents in online environments

    Pan, Y ., Kong, D., Zhou, S., Cui, C., Leng, Y ., Jiang, B., Liu, H., Shang, Y ., Zhou, S., Wu, T., et al. Webcanvas: Benchmarking web agents in online environments. InAgentic Markets Workshop at ICML 2024,

  44. [60]

    S., Poelitz, C., Baral, C., Roy, S., Chakravarthy, R., Van Durme, B., and Nouri, E

    Payan, J., Mishra, S., Singh, M., Negreanu, C. S., Poelitz, C., Baral, C., Roy, S., Chakravarthy, R., Van Durme, B., and Nouri, E. Instructexcel: A benchmark for natural language instruction in excel. InThe 2023 Conference on Empirical Methods in Natural Language Processing,

  45. [61]

    Escapebench: Pushing language models to think outside the box.arXiv preprint arXiv:2412.13549,

    Qian, C., Han, P., Luo, Q., He, B., Chen, X., Zhang, Y ., Du, H., Yao, J., Yang, X., Zhang, D., et al. Escapebench: Pushing language models to think outside the box.arXiv preprint arXiv:2412.13549,

  46. [62]

    Towards completeness- oriented tool retrieval for large language models

    Qu, C., Dai, S., Wei, X., Cai, H., Wang, S., Yin, D., Xu, J., and Wen, J.-R. Towards completeness- oriented tool retrieval for large language models. InProceedings of the 33rd ACM International Conference on Information and Knowledge Management, pp. 1930–1940,

  47. [63]

    Android in the wild: A large-scale dataset for android device control, 2023.URL https://arxiv

    Rawles, C., Li, A., Rodriguez, D., Riva, O., and Lillicrap, T. Android in the wild: A large-scale dataset for android device control, 2023.URL https://arxiv. org/abs/2307.10088,

  48. [64]

    Androidworld: A dynamic benchmarking environment for autonomous agents.arXiv preprint arXiv:2405.14573,

    Rawles, C., Clinckemaillie, S., Chang, Y ., Waltz, J., Lau, G., Fair, M., Li, A., Bishop, W., Li, W., Campbell-Ajala, F., et al. Androidworld: A dynamic benchmarking environment for autonomous agents.arXiv preprint arXiv:2405.14573,

  49. [65]

    Hcast: Human-calibrated autonomy software tasks.arXiv preprint arXiv:2503.17354,

    Rein, D., Becker, J., Deng, A., Nix, S., Canal, C., O’Connel, D., Arnott, P., Bloom, R., Broadley, T., Garcia, K., et al. Hcast: Human-calibrated autonomy software tasks.arXiv preprint arXiv:2503.17354,

  50. [66]

    Vgrp- bench: Visual grid reasoning puzzle benchmark for large vision-language models.arXiv preprint arXiv:2503.23064,

    Ren, Y ., Tertikas, K., Maiti, S., Han, J., Zhang, T., Süsstrunk, S., and Kokkinos, F. Vgrp- bench: Visual grid reasoning puzzle benchmark for large vision-language models.arXiv preprint arXiv:2503.23064,

  51. [67]

    Benchmarking llms’ swarm intelligence.arXiv preprint arXiv:2505.04364,

    47 Ruan, K., Huang, M., Wen, J.-R., and Sun, H. Benchmarking llms’ swarm intelligence.arXiv preprint arXiv:2505.04364,

  52. [68]

    Alfred: A benchmark for interpreting grounded instructions for everyday tasks

    Shridhar, M., Thomason, J., Gordon, D., Bisk, Y ., Han, W., Mottaghi, R., Zettlemoyer, L., and Fox, D. Alfred: A benchmark for interpreting grounded instructions for everyday tasks. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 10740–...

  53. [69]

    Collab-overcooked: Benchmark- ing and evaluating large language models as collaborative agents.arXiv preprint arXiv:2502.20073,

    Sun, H., Zhang, S., Ren, L., Xu, H., Fu, H., Yuan, C., and Wang, X. Collab-overcooked: Benchmark- ing and evaluating large language models as collaborative agents.arXiv preprint arXiv:2502.20073,

  54. [70]

    Meta-gui: Towards multi-modal conver- sational agents on mobile gui

    Sun, L., Chen, X., Chen, L., Dai, T., Zhu, Z., and Yu, K. Meta-gui: Towards multi-modal conver- sational agents on mobile gui. InProceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pp. 6699–6712,

  55. [71]

    Towards general computer control: A multimodal agent for red dead redemption ii as a case study

    Tan, W., Ding, Z., Zhang, W., Li, B., Zhou, B., Yue, J., Xia, H., Jiang, J., Zheng, L., Xu, X., et al. Towards general computer control: A multimodal agent for red dead redemption ii as a case study. InICLR 2024 Workshop on Large Language Model (LLM) Agents,

  56. [72]

    Toolalpaca: Generalized tool learning for language models with 3000 simulated cases.arXiv preprint arXiv:2306.05301,

    Tang, Q., Deng, Z., Lin, H., Han, X., Liang, Q., Cao, B., and Sun, L. Toolalpaca: Generalized tool learning for language models with 3000 simulated cases.arXiv preprint arXiv:2306.05301,

  57. [73]

    Dsgbench: A diverse strategic game benchmark for evaluating llm-based agents in complex decision-making environments.arXiv preprint arXiv:2503.06047,

    Tang, W., Zhou, Y ., Xu, E., Cheng, K., Li, M., and Xiao, L. Dsgbench: A diverse strategic game benchmark for evaluating llm-based agents in complex decision-making environments.arXiv preprint arXiv:2503.06047,

  58. [74]

    J., Kang, J., Wu, W., Christianos, F., Greenlee, F., Toulis, A., and Pur- torab, M

    Thomas, G., Chan, A. J., Kang, J., Wu, W., Christianos, F., Greenlee, F., Toulis, A., and Pur- torab, M. Webgames: Challenging general-purpose web-browsing ai agents.arXiv preprint arXiv:2502.18356,

  59. [75]

    Learning to speak and act in a fantasy text adventure game

    48 Urbanek, J., Fan, A., Karamcheti, S., Jain, S., Humeau, S., Dinan, E., Rocktäschel, T., Kiela, D., Szlam, A., and Weston, J. Learning to speak and act in a fantasy text adventure game. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing...

  60. [76]

    G., Talukdar, P., and Narayanan, S

    Venkatesh, S. G., Talukdar, P., and Narayanan, S. Ugif: Ui grounded instruction following.arXiv preprint arXiv:2211.07615,

  61. [77]

    Gta: a benchmark for general tool agents

    Wang, J., Zerun, M., Li, Y ., Zhang, S., Chen, C., Chen, K., and Le, X. Gta: a benchmark for general tool agents. InThe Thirty-eight Conference on Neural Information Processing Systems Datasets and Benchmarks Track, 2024a. Wang, L., Deng, Y ., Zha, Y ., Mao, G., Wang, Q., Min,...

  62. [78]

    Battleagentbench: A benchmark for evaluating cooperation and competition capabilities of language models in multi-agent systems.arXiv preprint arXiv:2408.15971, 2024c

    Wang, W., Zhang, D., Feng, T., Wang, B., and Tang, J. Battleagentbench: A benchmark for evaluating cooperation and competition capabilities of language models in multi-agent systems.arXiv preprint arXiv:2408.15971, 2024c. Wang, X., Li, D., Zhao, Y ., Wang, H., et al. Metatool:...

  63. [79]

    Re-bench: Evaluating frontier ai r&d capabilities of language model agents against human experts.arXiv preprint arXiv:2411.15114,

    Wijk, H., Lin, T., Becker, J., Jawhar, S., Parikh, N., Broadley, T., Chan, L., Chen, M., Clymer, J., Dhyani, J., et al. Re-bench: Evaluating frontier ai r&d capabilities of language model agents against human experts.arXiv preprint arXiv:2411.15114,

  64. [80]

    Deciphering digital detectives: Understanding llm behaviors and capabilities in multi-agent mystery games

    Wu, D., Shi, H., Sun, Z., and Liu, B. Deciphering digital detectives: Understanding llm behaviors and capabilities in multi-agent mystery games. InFindings of the Association for Computational Linguistics ACL 2024, pp. 8225–8291, 2024a. 49 Wu, Y ., Tang, X., Mitchell, T., and ...

  65. [81]

    F., Song, Y ., Li, B., Tang, Y ., Jain, K., Bao, M., Wang, Z

    Xu, F. F., Song, Y ., Li, B., Tang, Y ., Jain, K., Bao, M., Wang, Z. Z., Zhou, X., Guo, Z., Cao, M., et al. Theagentcompany: benchmarking llm agents on consequential real world tasks.arXiv preprint arXiv:2412.14161, 2024a. Xu, K., Kordi, Y ., Nayak, T., Asija, A., Wang, Y ., S...

  66. [82]

    Spin-bench: How well do llms plan strategically and reason socially?arXiv preprint arXiv:2503.12349,

    Yao, J., Wang, K., Hsieh, R., Zhou, H., Zou, T., Cheng, Z., Wang, Z., and Viswanath, P. Spin-bench: How well do llms plan strategically and reason socially?arXiv preprint arXiv:2503.12349,

  67. [83]

    τ-bench: A benchmark for tool-agent-user interaction in real-world domains.arXiv preprint arXiv:2406.12045,

    Yao, S., Shinn, N., Razavi, P., and Narasimhan, K. τ-bench: A benchmark for tool-agent-user interaction in real-world domains.arXiv preprint arXiv:2406.12045,

  68. [84]

    Multi-mission tool bench: Assessing the robustness of llm based agents through related and dynamic missions.arXiv preprint arXiv:2504.02623,

    Yu, P., Yang, Y ., Li, J., Zhang, Z., Wang, H., Feng, X., and Zhang, F. Multi-mission tool bench: Assessing the robustness of llm based agents through related and dynamic missions.arXiv preprint arXiv:2504.02623,

  69. [85]

    M., Feghali, C

    Yuan, X., Moss, M. M., Feghali, C. E., Singh, C., Moldavskaya, D., MacPhee, D., Caccia, L., Pereira, M., Kim, M., Sordoni, A., et al. debug-gym: A text-based environment for interactive debugging. arXiv preprint arXiv:2503.21557,

  70. [86]

    Mobile-env: Building qualified evaluation benchmarks for llm-gui interaction.arXiv preprint arXiv:2305.08144,

    Zhang, D., Shen, Z., Xie, R., Zhang, S., Xie, T., Zhao, Z., Chen, S., Chen, L., Xu, H., Cao, R., et al. Mobile-env: Building qualified evaluation benchmarks for llm-gui interaction.arXiv preprint arXiv:2305.08144,

  71. [87]

    Ing-vp: Mllms cannot play easy vision-based games yet.arXiv preprint arXiv:2410.06555, 2024a

    Zhang, H., Guo, H., Guo, S., Cao, M., Huang, W., Liu, J., and Zhang, G. Ing-vp: Mllms cannot play easy vision-based games yet.arXiv preprint arXiv:2410.06555, 2024a. Zhang, L., Wang, S., Jia, X., Zheng, Z., Yan, Y ., Gao, L., Li, Y ., and Xu, M. Llamatouch: A faithful and scal...

  72. [88]

    J., Yan, R., Yao, Y ., and Wang, L

    Zheng, X., Li, L., Yang, Z., Yu, P., Wang, A. J., Yan, R., Yao, Y ., and Wang, L. V-mage: A game evaluation framework for assessing visual-centric capabilities in multimodal large language models. arXiv preprint arXiv:2504.06148,

  73. [2015]

    Cradle: Empowering foundation agents towards general computer control.arXiv preprint arXiv:2403.03186,

    Tan, W., Zhang, W., Xu, X., Xia, H., Ding, Z., Li, B., Zhou, B., Yue, J., Jiang, J., Li, Y ., et al. Cradle: Empowering foundation agents towards general computer control.arXiv preprint arXiv:2403.03186,

  74. [2017]

    From system 1 to system 2: A survey of reasoning large language models.arXiv preprint arXiv:2502.17419,

    Li, Z.-Z., Zhang, D., Zhang, M.-L., Zhang, J., Liu, Z., Yao, Y ., Xu, H., Zheng, J., Wang, P.-J., Chen, X., et al. From system 1 to system 2: A survey of reasoning large language models.arXiv preprint arXiv:2502.17419,

  75. [2019]

    Llmarena: Assessing capabilities of large language models in dynamic multi-agent environments

    42 Chen, J., Hu, X., Liu, S., Huang, S., Tu, W.-W., He, Z., and Wen, L. Llmarena: Assessing capabilities of large language models in dynamic multi-agent environments. InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Pape...

  76. [2020]

    K., Li, Y ., Ding, C., Lin, J., Liang, P

    Jang, L. K., Li, Y ., Ding, C., Lin, J., Liang, P. P., Zhao, D., Bonatti, R., and Koishida, K. Vide- owebarena: Evaluating long context multimodal agents with video understanding web tasks. In NeurIPS 2024 Workshop on Open-World Agents,

  77. [2022]

    Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning.arXiv preprint arXiv:2501.12948,

    Guo, D., Yang, D., Zhang, H., Song, J., Zhang, R., Xu, R., Zhu, Q., Ma, S., Wang, P., Bi, X., et al. Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning.arXiv preprint arXiv:2501.12948,

  78. [2023]

    Web 10 2000 single-agent107 WEBLINX (Lu et al.,

  79. [2024]

    Kang, L., Zhao, Z., Hsu, D., and Lee, W. S. On the empirical complexity of reasoning and planning in llms.arXiv preprint arXiv:2404.11041,

  80. [2025]

    Chen, Q., Qin, L., Liu, J., Peng, D., Guan, J., Wang, P., Hu, M., Zhou, Y ., Gao, T., and Che, W

    URLhttps://www.anthropic.com/claude/sonnet. Chen, Q., Qin, L., Liu, J., Peng, D., Guan, J., Wang, P., Hu, M., Zhou, Y ., Gao, T., and Che, W. Towards reasoning era: A survey of long chain-of-thought for reasoning large language models. arXiv preprint arXiv:2503.09567,

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.