BigBag generates reusable AST transformation rules via LLMs that achieve 78.6% fix rate on 157 breaking dependency updates and 33.3% cross-project transfer overall.
On randomness in agentic evals,
6 Pith papers cite this work. Polarity classification is still indexing.
years
2026 6representative citing papers
TraceProbe normalizes coding agent trajectories into canonical actions and applies rule-based detectors to localize failure patterns and behavioral divergences that resolve rate hides.
Paired configuration-equivalent trials on Claude Haiku 4.5 yield a noise floor of roughly [-3, +18]pp with no significant coordination contrast after correction, placing most recent multi-agent papers inside or below that envelope.
Proposes referential security as a paradigm for AI evaluations that reframes model identity as verifiable to support reproducible audits and regulatory decisions despite system changes.
Log-coverage metrics show Claude Opus 4.6 beats human Locust tests on unique log templates on Light-OAuth2, while EvoMaster and GPT lag, and strategy combinations uncover largely distinct behaviors.
Ablation study finds that a structural codebase index improves localization and resolve rates in coding agents on two SWE benchmarks without raising per-cell cost.
citing papers explorer
-
Agentic Generation of AST Transformation Rules for Fixing Breaking Updates
BigBag generates reusable AST transformation rules via LLMs that achieve 78.6% fix rate on 157 breaking dependency updates and 33.3% cross-project transfer overall.
-
What Resolve Rate Hides: Trajectory Structure Diagnostics for Coding Agents
TraceProbe normalizes coding agent trajectories into canonical actions and applies rule-based detectors to localize failure patterns and behavioral divergences that resolve rate hides.
-
How Much Coordination Gain Is Real? A Paired Noise-Floor Protocol for Multi-Agent LLM Benchmarks
Paired configuration-equivalent trials on Claude Haiku 4.5 yield a noise floor of roughly [-3, +18]pp with no significant coordination contrast after correction, placing most recent multi-agent papers inside or below that envelope.
-
Referential Security as a New Paradigm for AI Evaluations
Proposes referential security as a paradigm for AI evaluations that reframes model identity as verifiable to support reproducible audits and regulatory decisions despite system changes.
-
Assessing REST API Test Generation Strategies with Log Coverage
Log-coverage metrics show Claude Opus 4.6 beats human Locust tests on unique log templates on Light-OAuth2, while EvoMaster and GPT lag, and strategy combinations uncover largely distinct behaviors.
-
Code Isn't Memory: A Structural Codebase Index Inside a Coding Agent
Ablation study finds that a structural codebase index improves localization and resolve rates in coding agents on two SWE benchmarks without raising per-cell cost.