Training a 9B LLM agent with reinforcement learning to actively navigate a structured multi-granularity memory pyramid yields competitive performance on memory-intensive benchmarks while preserving non-memory capabilities.
We use LongMemEval only as a held-out test benchmark to evaluate out-of-distribution generalization, with 100 test questions in total
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.AI 1years
2026 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
From Passive Retrieval to Active Memory Navigation: Learning to Use Memory as a Structured Action Space
Training a 9B LLM agent with reinforcement learning to actively navigate a structured multi-granularity memory pyramid yields competitive performance on memory-intensive benchmarks while preserving non-memory capabilities.