Agentic CLEAR automates multi-level evaluation of LLM agents, generating textual insights at system, trace, and node granularity that align with human annotations and predict task success.
Gonzalez and Ion Stoica , booktitle=
2 Pith papers cite this work. Polarity classification is still indexing.
2
Pith papers citing it
years
2026 2verdicts
UNVERDICTED 2representative citing papers
ClawArena-Team is a new execution-based benchmark that isolates and scores an LLM's subagent orchestration skill via a Subagent-Management Score multiplying correctness by privilege and modality factors.
citing papers explorer
-
Agentic CLEAR: Automating Multi-Level Evaluation of LLM Agents
Agentic CLEAR automates multi-level evaluation of LLM agents, generating textual insights at system, trace, and node granularity that align with human annotations and predict task success.
-
ClawArena-Team: Benchmarking Subagent Orchestration and Dynamic Workflows in Language-Model Agents
ClawArena-Team is a new execution-based benchmark that isolates and scores an LLM's subagent orchestration skill via a Subagent-Management Score multiplying correctness by privilege and modality factors.