Pith. sign in

hub

Holistic agent leaderboard: The missing infrastructure for ai agent evaluation

22 Pith papers cite this work. Polarity classification is still indexing.

22 Pith papers citing it

hub tools

citation-role summary

background 2 method 1

citation-polarity summary

years

2026 22

representative citing papers

StaminaBench: Stress-Testing Coding Agents over 100 Interaction Turns

cs.SE · 2026-06-17 · unverdicted · novelty 7.0

StaminaBench evaluates coding agents over 100 procedurally generated change requests to a REST API, finding that tested models fail within 5-6 turns without feedback but improve up to 12x with test feedback and good harnesses.

AI Coding Agents Can Reproduce Social Science Findings

cs.CL · 2026-06-09 · conditional · novelty 7.0

A new benchmark shows AI coding agents reproduce many social science findings from provided materials, outperforming prior agent benchmarks, while highlighting prompt sensitivity.

Agent Meltdowns: The Road to Hell Is Paved with Helpful Agents

cs.CL · 2026-05-18 · unverdicted · novelty 7.0

The paper defines accidental meltdowns as unsafe agent behavior triggered by benign errors and reports that such meltdowns occur in 64.7% of evaluated rollouts across GPT, Grok, and Gemini agents.

The Agentic Web Requires New Normative Infrastructure

cs.CY · 2026-06-09 · conditional · novelty 6.0 · 2 refs

The web's anti-bot regime should be replaced by a framework that presumptively lets user-authorized AI agents act for their principals, requires platforms to disclose access policies, and permits agent blocking only when proportionate to concrete harms.

MarketBench: Evaluating AI Agents as Market Participants

cs.AI · 2026-04-26 · unverdicted · novelty 6.0

LLMs show poor calibration in predicting task success and token use on software engineering benchmarks, causing market auctions to underperform compared to perfect information scenarios, with limited improvement from added context.

Stop Comparing LLM Agents Without Disclosing the Harness

cs.AI · 2026-05-07 · unverdicted · novelty 4.0

The Binding Constraint Thesis states that harness configuration governs performance variance more than model choice in long-horizon agent tasks, leading to misattribution in evaluations.

ClinQueryAgent: A Conversational Agent for Population Health Management

cs.IR · 2026-04-13 · unverdicted · novelty 4.0

The paper introduces ClinQueryAgent, a conversational agent that converts natural language queries into database queries for population health management while keeping patient data secure, and reports its use by 128 staff across 15 NHS practices covering 148,319 patients.

citing papers explorer

Showing 22 of 22 citing papers.