Coding-agent performance is workload- and framework-dependent, and raw speedup is an unsafe score because agents exploit benchmark-specific shortcuts.
and Nadgir, Nitya and Narayanan, Arvind , journal=
2 Pith papers cite this work. Polarity classification is still indexing.
2
Pith papers citing it
years
2026 2representative citing papers
Compiling agentic workflows into LLM weights creates subterranean agents with near-frontier quality at two orders of magnitude less cost, validated empirically on travel booking, Zoom support, and insurance claims tasks.
citing papers explorer
-
PERFOPT-Bench: Evaluating Coding Agents on Software Performance Optimization
Coding-agent performance is workload- and framework-dependent, and raw speedup is an unsafe score because agents exploit benchmark-specific shortcuts.
-
Compiling Agentic Workflows into LLM Weights: Near-Frontier Quality at Two Orders of Magnitude Less Cost
Compiling agentic workflows into LLM weights creates subterranean agents with near-frontier quality at two orders of magnitude less cost, validated empirically on travel booking, Zoom support, and insurance claims tasks.