{"paper":{"title":"MemoryBench: A Benchmark for Memory and Continual Learning in LLM Systems","license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","headline":"Existing benchmarks fall short for testing LLM memory and continual learning from user feedback.","cross_cats":["cs.AI","cs.IR"],"primary_cat":"cs.LG","authors_text":"Changyue Wang, Jianming Long, Qingyao Ai, Weihang Su, Yichen Tang, Yiqun Liu","submitted_at":"2025-10-20T08:16:12Z","abstract_excerpt":"Scaling up data, parameters, and test-time computation has been the mainstream methods to improve LLM systems (LLMsys), but their upper bounds are almost reached due to the gradual depletion of high-quality data and marginal gains obtained from larger computational resource consumption. Inspired by the abilities of human and traditional AI systems in learning from practice, constructing memory and continual learning frameworks for LLMsys has become an important and popular research direction in recent literature. Yet, existing benchmarks for LLM memory often focus on evaluating the system on h"},"claims":{"count":4,"items":[{"kind":"strongest_claim","text":"Experiments show that the effectiveness and efficiency of state-of-the-art baselines are far from satisfying.","source":"verdict.strongest_claim","status":"machine_extracted","claim_id":"C1","attestation":"unclaimed"},{"kind":"weakest_assumption","text":"The proposed user feedback simulation framework produces interactions that are representative of real user behavior in deployed LLM services.","source":"verdict.weakest_assumption","status":"machine_extracted","claim_id":"C2","attestation":"unclaimed"},{"kind":"one_line_summary","text":"MemoryBench is a new multi-domain benchmark that simulates ongoing user feedback to evaluate continual learning in LLM systems, finding that state-of-the-art memory methods are ineffective and inefficient.","source":"verdict.one_line_summary","status":"machine_extracted","claim_id":"C3","attestation":"unclaimed"},{"kind":"headline","text":"Existing benchmarks fall short for testing LLM memory and continual learning from user feedback.","source":"verdict.pith_extraction.headline","status":"machine_extracted","claim_id":"C4","attestation":"unclaimed"}],"snapshot_sha256":"ea45e55a376dde9e2f52c8277ba43a49af4dd23e2f3bed48b6740007274f53c5"},"source":{"id":"2510.17281","kind":"arxiv","version":7},"verdict":{"id":"bd2b4ec2-4de8-4eba-abe0-0c941d19151a","model_set":{"reader":"grok-4.3"},"created_at":"2026-05-18T06:16:57.254358Z","strongest_claim":"Experiments show that the effectiveness and efficiency of state-of-the-art baselines are far from satisfying.","one_line_summary":"MemoryBench is a new multi-domain benchmark that simulates ongoing user feedback to evaluate continual learning in LLM systems, finding that state-of-the-art memory methods are ineffective and inefficient.","pipeline_version":"pith-pipeline@v0.9.0","weakest_assumption":"The proposed user feedback simulation framework produces interactions that are representative of real user behavior in deployed LLM services.","pith_extraction_headline":"Existing benchmarks fall short for testing LLM memory and continual learning from user feedback."},"integrity":{"clean":true,"summary":{"advisory":0,"critical":0,"by_detector":{},"informational":0},"endpoint":"/pith/2510.17281/integrity.json","findings":[],"available":true,"detectors_run":[],"snapshot_sha256":"c28c3603d3b5d939e8dc4c7e95fa8dfce3d595e45f758748cecf8e644a296938"},"references":{"count":0,"sample":[],"resolved_work":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57","internal_anchors":0},"formal_canon":{"evidence_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"}