A hand-curated 67-task multilingual long-horizon agent benchmark finds large domain, language, and harness-driven performance gaps for current LLM agents.
Agentbench: Evaluating llms as agents
2 Pith papers cite this work. Polarity classification is still indexing.
fields
cs.AI 2years
2026 2representative citing papers
Interactive evaluation of AI must be reframed as a distinct paradigm that maps interaction trajectories to judgments on process, recoverability, coordination, robustness, and system performance, supported by a two-axis taxonomy and design principles.
citing papers explorer
-
PolyWorkBench: Benchmarking Multilingual Long-Horizon LLM Agents
A hand-curated 67-task multilingual long-horizon agent benchmark finds large domain, language, and harness-driven performance gaps for current LLM agents.
-
Interactive Evaluation Requires a Design Science
Interactive evaluation of AI must be reframed as a distinct paradigm that maps interaction trajectories to judgments on process, recoverability, coordination, robustness, and system performance, supported by a two-axis taxonomy and design principles.