Pith. sign in

hub Mixed citations

ClawBench: Can AI Agents Complete Everyday Online Tasks?

Mixed citation behavior. Most common role is background (60%).

15 Pith papers citing it
Background 60% of classified citations
abstract

AI agents may be able to automate your inbox, but can they automate other routine aspects of your life? Everyday online tasks offer a realistic yet unsolved testbed for evaluating the next generation of AI agents. To this end, we introduce ClawBench, an evaluation framework of 153 simple tasks that people need to accomplish regularly in their lives and work, spanning 144 live platforms across 15 categories, from completing purchases and booking appointments to submitting job applications. These tasks require demanding capabilities beyond existing benchmarks, such as obtaining relevant information from user-provided documents, navigating multi-step workflows across diverse platforms, and write-heavy operations like filling in many detailed forms correctly. Unlike existing benchmarks that evaluate agents in offline sandboxes with static pages, ClawBench operates on production websites, preserving the full complexity, dynamic nature, and challenges of real-world web interaction. A lightweight interception layer captures and blocks only the final submission request, ensuring safe evaluation without real-world side effects. Our evaluations of 7 frontier models show that both proprietary and open-source models can complete only a small portion of these tasks. For example, Claude Sonnet 4.6 achieves only 33.3%. Progress on ClawBench brings us closer to AI agents that can function as reliable general-purpose assistants.

hub tools

citation-role summary

background 3 dataset 2

citation-polarity summary

years

2026 15

representative citing papers

MacAgentBench: Benchmarking AI Agents on Real-World macOS Desktop

cs.AI · 2026-06-21 · unverdicted · novelty 7.0

MacAgentBench is a new benchmark for macOS AI agents with 676 tasks, deterministic multi-checkpoint evaluation, and tests across frameworks showing skill libraries drive performance more than framework design.

SING: Synthetic Intention Graph for Scalable Active Tool Discovery in LLM Agents

cs.CL · 2026-06-15 · unverdicted · novelty 6.0

SING builds an intention-tool graph linking user intentions, tool capabilities, and collaboration patterns to enable dynamic retrieval, improving Global Recall@5 by up to 59.8% and success rate by up to 28.9% on three benchmarks while cutting schema exposure by 99.8%.

NeuroClaw Technical Report

cs.CV · 2026-04-27 · unverdicted · novelty 6.0

NeuroClaw is a domain-specialized multi-agent framework with NeuroBench benchmark that improves executability and reproducibility for multimodal neuroimaging research.

citing papers explorer

Showing 15 of 15 citing papers.