AgentTether repairs 69% of initially failed LLM agent tasks on τ-bench by combining graph-guided root-cause diagnosis, cross-iteration repair memory, and guarded runtime intervention, improving over blind retry by 26 percentage points.
Swe-smith: Scaling data for software engineering agents,
2 Pith papers cite this work. Polarity classification is still indexing.
2
Pith papers citing it
fields
cs.SE 2years
2026 2representative citing papers
LLVM-Bench supplies 423 validated LLVM issues and LLVM-Gym automates evaluation, showing LLMs are limited but an ensemble reaches 21.99% resolution.
citing papers explorer
-
AgentTether: Graph-Guided Diagnosis and Runtime Intervention for Reliable LLM Agent Operation
AgentTether repairs 69% of initially failed LLM agent tasks on τ-bench by combining graph-guided root-cause diagnosis, cross-iteration repair memory, and guarded runtime intervention, improving over blind retry by 26 percentage points.
-
LLVM-Bench: Benchmarking and Advancing Large Language Models for LLVM Compiler Issue Resolution
LLVM-Bench supplies 423 validated LLVM issues and LLVM-Gym automates evaluation, showing LLMs are limited but an ensemble reaches 21.99% resolution.