Pith. sign in

Agentbench: Evaluating llms as agents

2 Pith papers cite this work. Polarity classification is still indexing.

2 Pith papers citing it

fields

cs.AI 2

years

2026 2

representative citing papers

Interactive Evaluation Requires a Design Science

cs.AI · 2026-05-18 · unverdicted · novelty 5.0

Interactive evaluation of AI must be reframed as a distinct paradigm that maps interaction trajectories to judgments on process, recoverability, coordination, robustness, and system performance, supported by a two-axis taxonomy and design principles.

citing papers explorer

Showing 2 of 2 citing papers.

  • PolyWorkBench: Benchmarking Multilingual Long-Horizon LLM Agents cs.AI · 2026-07-07 · conditional · none · ref 10 · 2 links

    A hand-curated 67-task multilingual long-horizon agent benchmark finds large domain, language, and harness-driven performance gaps for current LLM agents.

  • Interactive Evaluation Requires a Design Science cs.AI · 2026-05-18 · unverdicted · none · ref 36

    Interactive evaluation of AI must be reframed as a distinct paradigm that maps interaction trajectories to judgments on process, recoverability, coordination, robustness, and system performance, supported by a two-axis taxonomy and design principles.