Across five 100-task agent streams, sequential experience improves normalized reward by 16.9% in 14 of 15 model-domain combinations, but explicit skill maintenance matches pure in-context learning (0.602 vs 0.605) and weaker models build larger, less reusable skill pools.
- If the question wording is broad (`had`,`liquidity`, `available`) and context explicitly states a combined figure, prefer the context-stated combined amount
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.AI 1years
2026 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
ContinualSkillBench: Can LLM Agents Truly Evolve Their Capabilities?
Across five 100-task agent streams, sequential experience improves normalized reward by 16.9% in 14 of 15 model-domain combinations, but explicit skill maintenance matches pure in-context learning (0.602 vs 0.605) and weaker models build larger, less reusable skill pools.