SWE-Chain provides 155 chained version transitions and 1,660 requirements across 9 Python packages, where frontier agents resolve 44.8% of tasks on average and struggle to preserve functionality across releases.
NLPerturbator: Study- ing the robustness of code LLMs to natural language variations.arXiv preprint arXiv:2406.19783
3 Pith papers cite this work. Polarity classification is still indexing.
years
2026 3verdicts
UNVERDICTED 3representative citing papers
Paraphrase sensitivity in Lean 4 autoformalization is dominated by code-generation failures that differ between undergraduate and Olympiad datasets across multiple models.
Any deterministic prompt filter for code LLMs has a provable mutual-information lower bound of at least 0.84 nats on HumanEval and 1.20 nats on MBPP under pass-only acceptance, with no tested filter achieving zero proxy-axis leakage.
citing papers explorer
-
SWE-Chain: Benchmarking Coding Agents on Chained Release-Level Package Upgrades
SWE-Chain provides 155 chained version transitions and 1,660 requirements across 9 Python packages, where frontier agents resolve 44.8% of tasks on average and struggle to preserve functionality across releases.
-
Characterizing Paraphrase-Induced Failures in Lean 4 Autoformalization
Paraphrase sensitivity in Lean 4 autoformalization is dominated by code-generation failures that differ between undergraduate and Olympiad datasets across multiple models.
-
The Security Budget of Code-LLM Prompt Hardening: Provable Limits Under Pass-Only Acceptance
Any deterministic prompt filter for code LLMs has a provable mutual-information lower bound of at least 0.84 nats on HumanEval and 1.20 nats on MBPP under pass-only acceptance, with no tested filter achieving zero proxy-axis leakage.