TUA-Bench provides 120 manually designed terminal tasks across five families with execution-based scoring; the top agent reaches 65.8% success.
Setupbench: Assessing software engineering agents’ ability to bootstrap development environments
4 Pith papers cite this work. Polarity classification is still indexing.
years
2026 4verdicts
UNVERDICTED 4representative citing papers
CLAWAUDIT applies a STRIDE-derived taxonomy and 47 Semgrep plus 30 CodeQL rules to local LLM agent code, lifting recall on held-out OpenClaw advisories from 21.7% and 13.8% baselines to 66.8% and 75.1%.
DeployBench is a new benchmark of 51 research-artifact deployment tasks where four LLMs with OpenHands achieve 7.8-51% pass rates, with failures mostly from agents stopping after weaker self-checks than the paper requires.
BootstrapAgent distills repository bootstrapping heuristics into a persistent .bootstrap contract via multi-agent evidence extraction, Docker verification, and trace-driven repair, reporting 92.9% success and efficiency gains on three benchmarks.
citing papers explorer
-
TUA-Bench: A Benchmark for General-Purpose Terminal-Use Agents
TUA-Bench provides 120 manually designed terminal tasks across five families with execution-based scoring; the top agent reaches 65.8% success.
-
Local LLM Agents as Vulnerable Runtimes:A Source-Code Audit of the Agent Runtime Layer
CLAWAUDIT applies a STRIDE-derived taxonomy and 47 Semgrep plus 30 CodeQL rules to local LLM agent code, lifting recall on held-out OpenClaw advisories from 21.7% and 13.8% baselines to 66.8% and 75.1%.
-
DeployBench: Benchmarking LLM Agents for Research Artifact Deployment
DeployBench is a new benchmark of 51 research-artifact deployment tasks where four LLMs with OpenHands achieve 7.8-51% pass rates, with failures mostly from agents stopping after weaker self-checks than the paper requires.
-
BootstrapAgent: Distilling Repository Setup into Reusable Agent Knowledge
BootstrapAgent distills repository bootstrapping heuristics into a persistent .bootstrap contract via multi-agent evidence extraction, Docker verification, and trace-driven repair, reporting 92.9% success and efficiency gains on three benchmarks.