MemGym unifies agent gyms into a memory benchmark with isolated scoring across tool-use, research, coding, and computer-use regimes plus a lightweight reward model for tractable coding evaluation.
O’Brien, Carrie J
2 Pith papers cite this work. Polarity classification is still indexing.
2
Pith papers citing it
fields
cs.CL 2years
2026 2verdicts
UNVERDICTED 2representative citing papers
CWL enables an LLM agent to complete 89 sequential tasks over 80 million tokens with no accuracy loss versus isolated per-task sessions by using dependency-aware episode eviction rather than summarization or recency truncation.
citing papers explorer
-
MemGym: a Long-Horizon Memory Environment for LLM Agents
MemGym unifies agent gyms into a memory benchmark with isolated scoring across tool-use, research, coding, and computer-use regimes plus a lightweight reward model for tractable coding evaluation.
-
Beyond Compaction: Structured Context Eviction for Long-Horizon Agents
CWL enables an LLM agent to complete 89 sequential tasks over 80 million tokens with no accuracy loss versus isolated per-task sessions by using dependency-aware episode eviction rather than summarization or recency truncation.