An LLM agent can improve itself at test time by rewriting its surrounding executable harness from unlabeled traces, using only proxy signals and a frozen model.
Title resolution pending
3 Pith papers cite this work. Polarity classification is still indexing.
years
2026 3representative citing papers
WM-SAR greedily grows a compact connected repair region by marginal residual-spectral relief so that fixing it stabilizes subsequent agent rollouts better than symptom scans under tight token budgets.
Viverra generates C code from text descriptions together with assertions that are verified by model checkers, and a user study with over 400 participants shows the verified assertions improve code comprehension.
citing papers explorer
-
TTHE: Test-Time Harness Evolution
An LLM agent can improve itself at test time by rewriting its surrounding executable harness from unlabeled traces, using only proxy signals and a frozen model.
-
Repair the Amplifier, Not the Symptom: Stable World-Model Correction for Agent Rollouts
WM-SAR greedily grows a compact connected repair region by marginal residual-spectral relief so that fixing it stabilizes subsequent agent rollouts better than symptom scans under tight token budgets.
-
Viverra: Text-to-Code with Guarantees
Viverra generates C code from text descriptions together with assertions that are verified by model checkers, and a user study with over 400 participants shows the verified assertions improve code comprehension.