A benchmark with matched persistence on/off controls shows personal agents improve unevenly from retained experience, and a new Hermes+ framework strengthens evidence on update tasks, though its overall gain is within run-to-run noise.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2026 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
PAST-Bench: Benchmarking the Foundations of Recursive Self-Improvement in Personal Agents
A benchmark with matched persistence on/off controls shows personal agents improve unevenly from retained experience, and a new Hermes+ framework strengthens evidence on update tasks, though its overall gain is within run-to-run noise.