REVIEW 3 major objections 4 minor 34 references
Per-query mode selection over one shared memory substrate is a workable control surface for agent memory, and one deployed configuration scores 61–86% on three long-context benchmarks while explicitly disclaiming causal routing claims.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-01 13:27 UTC pith:GDSI437I
load-bearing objection Worth a serious look: the per-query routing architecture is real and the paper is unusually honest, but the LongMemEval headline sits on a scoring join that could be an artifact, and no public artifact exists to check it. the 3 major comments →
Supra Cognitive Modes: A Routed Architecture for Agent Memory
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The central claim is that a memory operating point can be chosen per query rather than fixed for an entire deployment. SCM defines modes as semantic labels, payloads as concrete retrieval-and-synthesis configurations, and procedures as execution families; several labels can share one initial payload, and runtime gates can revise the tier before execution. Source inspection verifies that this per-query control surface and the shared substrate exist, and completed runs on three benchmarks characterize one deployed configuration. The paper explicitly disclaims causal routing effects, statistical significance, efficiency gains, and component attribution: the reported numbers are benchmark-metric
What carries the argument
The mode–payload–procedure interface: a frozen semantic classifier (or an application-supplied explicit label) picks one of four modes, a payload map turns that label into retrieval flags, and runtime gates may revise the final tier before a procedure—direct fused retrieval, graph or iterative multi-hop handling, or stratified long-form synthesis—executes against the shared substrate. The substrate is the second load-bearing piece: one mandatory ingest path writes chunks, embeddings, extracted triples, and fact-version metadata so that all three procedure families read from the same store, with optional asynchronous enrichments added later as nullable signals.
Load-bearing premise
The reported scores depend on the scoring-join and preprocessing conventions matching the benchmarks' official definitions; the paper itself discloses that an alternative join on the same rows changed one benchmark score from 21.8% to 88.2%, and raw baseline rows are not available for independent re-joins.
What would settle it
Run the stored retrieval outputs through each benchmark's official scorer using the standard question-ID join and without the one-row category reclassification; if the result differs materially from the reported 84.87%, 68.61%, 61.49%, and 86.00% figures—or if the longitudinal-memory score falls back toward the disclosed 21.8% enumerate-join value—the characterization is an artifact of the evaluation pipeline rather than of the architecture.
If this is right
- If the interface works as implemented, agent-memory systems can expose a per-query selector to applications without rebuilding or duplicating the memory store.
- A single shared substrate can support direct, graph-capable, and long-form procedures, so ingest is paid once and query-time choices do not require a separate store per capability.
- The mode-conditioned failure strata—for example, recommendation and multi-hop conflict-resolution errors concentrate in particular semantic labels—give concrete targets for fixes rather than undifferentiated pipeline changes.
- The paper's stated next steps define what would turn descriptive scores into causal claims: persist final procedures, score the same answer whose timing is retained, and compare routed versus fixed policies on identical corpus snapshots.
Where Pith is reading between the lines
- A test the paper leaves implicit: hold the shared substrate and synthesis model fixed and vary only the mode selector against a fixed policy; without that control, the 61–86% figures cannot be attributed to routing.
- The disclosed scoring-join fragility—one benchmark's score swinging between 21.8% and 88.2% on the same rows depending on the join convention—suggests that published agent-memory benchmark numbers may be more sensitive to evaluation plumbing than to architecture, so a shared, auditable scorer harness would benefit the field.
- The per-query control surface likely generalizes beyond conversational memory to other heterogeneous memory workloads, such as retrieval over tool-use logs or evolving user profiles, wherever the same question-shape diversity appears.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper describes Supra Cognitive Modes (SCM), an agent-memory architecture that exposes a per-query mode–payload–procedure control interface over a shared ingest substrate. It reports a deployed configuration's accuracy on LoCoMo (84.87% factoid, 68.61% adversarial abstention, 81.22% mixed), MemoryAgentBench (61.49%, runs 61.72% and 61.26%), and LongMemEval (86.00%), together with route-conditioned failure strata. The manuscript is careful to bound these claims: it states that the results characterize one implemented configuration, that causal routing effects and statistical significance are outside the evidence, and it discloses missing raw baseline outputs, absent Stage-3 timing, and the use of judge-based metrics. The paper's central contributions are the interface/implementation claim and the bounded benchmark characterization.
Significance. If the implementation claim holds, the architecture is a useful and clearly described organizational contribution: it defines an explicit per-query selector over shared memory substrates, with frozen classifier, runtime gates, and procedure families all documented in source. The evaluation has genuine strengths: two MAB repetitions are reported explicitly, factoid and abstention cuts are separated, the one-row LoCoMo correction is disclosed, and the diagnostic failure strata in Table 9 provide reproducible engineering hypotheses. The paper also shows unusual discipline in refusing to claim causal or efficiency advantages. However, the numerical characterization is less secure than the architecture claim. The most serious issue is the scoring-join sensitivity for LongMemEval: the same retained rows reportedly score 21.8% under one join and 88.2% under another, and the raw rows are not public. Because the headline scores and Table 8 task-type numbers depend on this choice, the central empirical contribution is not independently verifiable in the current artifact. I regard the architecture claim as defensible and the benchmark characterization as needing substantial strengthening.
major comments (3)
- [Appendix A.4; Tables 4 and 8] The LongMemEval numeric characterization is determined by the scoring-join convention. The runbook reports the same retained rows scoring 21.8% under an enumerate-based join and 88.2% under a question-ID join; the paper adopts the question-ID join, yielding the 86.00% headline, the Table 8 task-type percentages, and the +51.88 temporal-reasoning difference. The raw Stage-3 outputs and the row-level mapping are not available, and Appendix B.4 states that no public artifact accompanies this version. No external party can check question-ID uniqueness, row duplication/dropping, or agreement with the official scorer. This is not a presentation issue: the paper's third stated contribution is a repository-traceable characterization, and that traceability is blocked at the join. Please release the raw rows and join mapping, or re-report LongMemEval with a conservative join, or remove the LongMem
- [§5.4, Tables 4 and 7] The one-row LoCoMo correction is applied asymmetrically. SCM factoid and mixed scores are corrected to the 1,540/446 fixture split, while the Mem0 and production-comparator values in Table 4 'retain the externally reported scorer outputs.' The corrected 84.87% factoid and 81.22% mixed values are therefore not on the same scoring basis as the 13.08% and 63.10% displayed next to them. Even if the reclassification itself is right, the cross-system comparison in Table 4 rests on an unverified assumption that the baselines would be unchanged by the same correction. The manuscript needs either to reprocess the baseline rows on the corrected split or to state explicitly that only the SCM column uses the corrected cut and refrain from displaying an apples-to-oranges comparison.
- [§4.5, §A.3, §A.4] The MAB aggregate depends on a post-hoc Detective-QA resynthesis step. The paper states that 'without this step, Detective-QA rows are judged against summary-style predictions rather than the benchmark's detective prompt,' but it does not establish that the resynthesis is part of the official MemoryAgentBench scoring protocol, and it does not quantify the effect of this step on the 61.49% headline. Because raw baseline rows are unavailable, the reader cannot determine how much of the SCM–baseline gap on MAB reflects this rescoring choice rather than the architecture. Please provide evidence that the resynthesis is canonical or report the sensitivity of the aggregate to this step.
minor comments (4)
- [§3.2] The dispatch pseudocode is informal; the functions classify, runtime_gates, and procedure_for are not fully specified. Since the paper's contribution is an interface, a more precise signature for mode_to_payload would help readers implement the abstraction independently.
- [Table 5] The column header 'MAB v3' and the note 'deployed v4 fast build' are ambiguous. If these are benchmark version identifiers, define them; if they are internal build names, say so explicitly to avoid confusion with the fixtures in Appendix D.
- [Figure 4 and Table 8] The task-type difference plot would be easier to interpret with n displayed per category, since the task-type denominators range from 30 to 133 and the displayed differences are not accompanied by any uncertainty.
- [Appendix A.4] The statement that score_locomo.py and score_longmem.py use hard-coded absolute roots is a useful disclosure, but the manuscript should also note whether the supplied aggregate numbers were produced before or after any path normalization, so a future artifact user can reproduce them.
Circularity Check
No significant circularity: benchmark scores are external measurements, the router was tuned off-test, and the disclosed judge-prompt concern does not reduce any reported number to its inputs.
full rationale
The paper's central claims are (1) implementation of a per-query mode-to-payload-routing interface over a shared substrate, verified by source inspection, and (2) descriptive benchmark characterization of one deployed configuration. Neither claim is derived from the other or from the reported scores. The router classifier was selected on a synthetic off-test set (Appendix C) and checked on a consensus-labeled development slice; it was not fit to LoCoMo, MAB, or LongMemEval outcomes. The benchmark numbers are measurements against external fixtures and fixed judge pipelines, not predictions derived from fitted parameters. The LongMemEval 21.8% vs. 88.2% join discrepancy (Appendix A.4) is a reproducibility/audit limitation, not a circularity: the paper discloses the convention it chose, and the score remains a measurement of external rows rather than an algebraic consequence of the join. The one-row LoCoMo correction (Section 5.4) reclassifies an existing output by fixture identity and is disclosed, so it is not a fitted input renamed as a prediction. The only self-identified circularity concern is in Section 4.5: 'The LoCoMo prompt originated with one compared system, creating a potential circularity; using the same prompt for all outputs standardizes scoring but does not eliminate that concern.' This is a judge-alignment and fairness caveat, not a derivation-equivalence: the judge prompt is applied identically to all systems and is not fitted to SCM outputs, so the 84.87% factoid score does not reduce by construction to the prompt's origin. The paper also repeatedly disclaims causal routing claims, significance, and efficiency gains, further avoiding the pattern of presenting fitted or self-referential quantities as independent predictions. No load-bearing self-citations or author-imported uniqueness theorems appear. Overall, the derivation chain is self-contained and externally referenced; score 0.
Axiom & Free-Parameter Ledger
free parameters (8)
- final retrieval depth top_k =
100
- RRF k_constant =
30
- BM25 k1 and b =
k1=1.5, b=0.75
- narrative rerank coefficients =
boost_factor=0.5, location_weight=0.5, min_freq=2
- classifier prompt candidate =
c5-balanced; 96.06% overall / 95.80% macro on 1,065 synthetic questions
- scoring join convention =
question_id join (vs enumerate-order: 21.8% vs 88.2% LongMem)
- eval concurrency =
12 in runbook; hard cap 24
- synthesis max tokens =
4096 summaries / 1024 recommendation lists / 600 other
axioms (6)
- domain assumption The three benchmark fixtures (LoCoMo n=1,986; MAB n=3,671; LongMemEval n=500) and their scorer semantics are authoritative inputs.
- domain assumption Benchmark LLM judges (gpt-4o-mini for LoCoMo factoids, gpt-4o for LongMemEval and MAB LRU) measure the quantity of interest.
- domain assumption The LoCoMo judge prompt from the Mem0 system is a fair scoring instrument for SCM and other baselines.
- domain assumption The frozen classifier's off-test calibration transfers to deployed query mixes.
- ad hoc to paper The one-row LoCoMo reclassification by fixture identity (1,540/446 split) is correct without a new model call.
- domain assumption The unnamed 'production comparator' scores are reliable despite absent raw rows.
read the original abstract
Agent-memory workloads mix direct factual lookup, relation-chain and current-state reasoning, and broad synthesis over long histories. We describe Supra Cognitive Modes (SCM), an architecture that maps explicit or automatically selected per-query modes to retrieval and synthesis payloads over one shared ingest substrate. A frozen semantic classifier and runtime gates dispatch queries among fused lexical and dense lookup, graph or iterative multi-hop handling, and stratified long-form synthesis. The substrate combines multi-granularity embeddings, extracted triples, fact-version metadata, and optional asynchronous enrichments. We characterize the deployed configuration on three benchmarks: Long-term Conversational Memory (LoCoMo; n = 1,986), MemoryAgentBench (MAB; n = 3,671), and LongMemEval (n = 500). The reference run records 84.87% on LoCoMo factoid categories and 68.61% on adversarial abstention, 61.49% on MAB across two repetitions, and 86.00% on LongMemEval. A repository-backed reproduction produces similar aggregate scores and supports task- and mode-conditioned failure analysis. Raw baseline outputs, aligned end-to-end timing for LoCoMo and LongMemEval, and complete token ledgers are unavailable; stored rows also omit some final runtime decisions. The results characterize one implemented routed configuration and its diagnostic failure patterns, while source inspection verifies the per-query control interface and shared-substrate design. Causal routing effects, efficiency gains, and statistical significance remain outside the available evidence.
Reference graph
Works this paper leans on
-
[1]
Proceedings of the International Conference on Learning Representations , year =
Asai, Akari and Wu, Zeqiu and Wang, Yizhong and Sil, Avirup and Hajishirzi, Hannaneh , title =. Proceedings of the International Conference on Learning Representations , year =
-
[2]
Proceedings of the Thirty-Second International Joint Conference on Artificial Intelligence , pages =
Cai, Borui and Xiang, Yong and Gao, Longxiang and Zhang, He and Li, Yunfeng and Li, Jianxin , title =. Proceedings of the Thirty-Second International Joint Conference on Artificial Intelligence , pages =. 2023 , url =
2023
-
[3]
arXiv preprint arXiv:2305.05176 , year =
Chen, Lingjiao and Zaharia, Matei and Zou, James , title =. arXiv preprint arXiv:2305.05176 , year =
-
[4]
arXiv preprint arXiv:2504.19413 , year =
Chhikara, Prateek and Khant, Dev and Aryan, Saket and Singh, Taranjeet and Yadav, Deshraj , title =. arXiv preprint arXiv:2504.19413 , year =
-
[5]
and Clarke, Charles L
Cormack, Gordon V. and Clarke, Charles L. A. and Buettcher, Stefan , title =. Proceedings of the 32nd International. 2009 , url =
2009
-
[6]
arXiv preprint arXiv:2404.16130 , year =
Edge, Darren and Trinh, Ha and Cheng, Newman and Bradley, Joshua and Chao, Alex and Mody, Apurva and Truitt, Steven and Metropolitansky, Dasha and Ness, Robert Osazuwa and Larson, Jonathan , title =. arXiv preprint arXiv:2404.16130 , year =
-
[7]
arXiv preprint arXiv:2510.18866 , year =
Fang, Jizhan and Deng, Xinle and Xu, Haoming and Jiang, Ziyan and Tang, Yuqi and Xu, Ziwen and Deng, Shumin and Yao, Yunzhi and Wang, Mengru and Qiao, Shuofei and Chen, Huajun and Zhang, Ningyu , title =. arXiv preprint arXiv:2510.18866 , year =
-
[8]
arXiv preprint arXiv:2507.05257 , year =
Hu, Yuanzhe and Wang, Yu and McAuley, Julian , title =. arXiv preprint arXiv:2507.05257 , year =
-
[9]
Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics , year =
Jeong, Soyeong and Baek, Jinheon and Cho, Sukmin and Hwang, Sung Ju and Park, Jong , title =. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics , year =
2024
-
[10]
Jiang, Zhengbao and Xu, Frank F. and Gao, Luyu and Sun, Zhiqing and Liu, Qian and Dwivedi-Yu, Jane and Yang, Yiming and Callan, Jamie and Neubig, Graham , title =. arXiv preprint arXiv:2305.06983 , year =
-
[11]
Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing , pages =
Karpukhin, Vladimir and Oguz, Barlas and Min, Sewon and Lewis, Patrick and Wu, Ledell and Edunov, Sergey and Chen, Danqi and Yih, Wen-tau , title =. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing , pages =. 2020 , url =
2020
-
[12]
Agent Memory and Architecture , year =
-
[13]
Advances in Neural Information Processing Systems , volume =
Lewis, Patrick and Perez, Ethan and Piktus, Aleksandra and Petroni, Fabio and Karpukhin, Vladimir and Goyal, Naman and Kuttler, Heinrich and Lewis, Mike and Yih, Wen-tau and Rocktaschel, Tim and Riedel, Sebastian and Kiela, Douwe , title =. Advances in Neural Information Processing Systems , volume =. 2020 , url =
2020
-
[14]
arXiv preprint arXiv:2507.03724 , year =
Li, Zhiyu and Xi, Chenyang and Li, Chunyu and Chen, Ding and others , title =. arXiv preprint arXiv:2507.03724 , year =
-
[15]
Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics , year =
Maharana, Adyasha and Lee, Dong-Ho and Tulyakov, Sergey and Bansal, Mohit and Barbieri, Francesco and Fang, Yuwei , title =. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics , year =
-
[16]
arXiv preprint arXiv:2606.06448 , year =
Omri, Yasmine and Gan, Ziyu and Broveak, Zachary and Geens, Robin and He, Zexue and Pentland, Alex and Verhelst, Marian and Weissman, Tsachy and Tambe, Thierry , title =. arXiv preprint arXiv:2606.06448 , year =
-
[17]
Ong, Isaac and Almahairi, Amjad and Wu, Vincent and Chiang, Wei-Lin and Wu, Tianhao and Gonzalez, Joseph E. and Kadous, M. Waleed and Stoica, Ion , title =. arXiv preprint arXiv:2406.18665 , year =
-
[18]
and Lin, Kevin and Wooders, Sarah and Gonzalez, Joseph E
Packer, Charles and Fang, Vivian and Patil, Shishir G. and Lin, Kevin and Wooders, Sarah and Gonzalez, Joseph E. , title =. arXiv preprint arXiv:2310.08560 , year =
-
[19]
Park, Joon Sung and O'Brien, Joseph C. and Cai, Carrie J. and Morris, Meredith Ringel and Liang, Percy and Bernstein, Michael S. , title =. arXiv preprint arXiv:2304.03442 , year =
-
[20]
arXiv preprint arXiv:2501.13956 , year =
Rasmussen, Preston and Paliychuk, Pavlo and Beauvais, Travis and Ryan, Jack and Chalef, Daniel , title =. arXiv preprint arXiv:2501.13956 , year =
-
[21]
, title =
Sarthi, Parth and Abdullah, Salman and Tuli, Aditi and Khanna, Shubh and Goldie, Anna and Manning, Christopher D. , title =. Proceedings of the International Conference on Learning Representations , year =
-
[22]
arXiv preprint arXiv:2303.11366 , year =
Shinn, Noah and Cassano, Federico and Berman, Edward and Gopinath, Ashwin and Narasimhan, Karthik and Yao, Shunyu , title =. arXiv preprint arXiv:2303.11366 , year =
-
[23]
and Yao, Shunyu and Narasimhan, Karthik and Griffiths, Thomas L
Sumers, Theodore R. and Yao, Shunyu and Narasimhan, Karthik and Griffiths, Thomas L. , title =. arXiv preprint arXiv:2309.02427 , year =
-
[24]
arXiv preprint arXiv:2506.21605 , year =
Tan, Haoran and Zhang, Zeyu and Ma, Chen and Chen, Xu and Dai, Quanyu and Dong, Zhenhua , title =. arXiv preprint arXiv:2506.21605 , year =
-
[25]
arXiv preprint arXiv:2510.27246 , year =
Tavakoli, Mohammad and Salemi, Alireza and Ye, Carrie and Abdalla, Mohamed and Zamani, Hamed and Mitchell, J Ross , title =. arXiv preprint arXiv:2510.27246 , year =
-
[26]
Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics , year =
Trivedi, Harsh and Balasubramanian, Niranjan and Khot, Tushar and Sabharwal, Ashish , title =. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics , year =
-
[27]
arXiv preprint arXiv:2507.07957 , year =
Wang, Yu and Chen, Xi , title =. arXiv preprint arXiv:2507.07957 , year =
-
[28]
arXiv preprint arXiv:2305.16291 , year =
Wang, Guanzhi and Xie, Yuqi and Jiang, Yunfan and Mandlekar, Ajay and Xiao, Chaowei and Zhu, Yuke and Fan, Linxi and Anandkumar, Anima , title =. arXiv preprint arXiv:2305.16291 , year =
-
[29]
arXiv preprint arXiv:2410.10813 , year =
Wu, Di and Wang, Hongwei and Yu, Wenhao and Zhang, Yuwei and Chang, Kai-Wei and Yu, Dong , title =. arXiv preprint arXiv:2410.10813 , year =
-
[30]
arXiv preprint arXiv:2605.12493 , year =
Wu, Di and Ji, Zixiang and Kawatkar, Asmi and Kwan, Bryan and Gu, Jia-Chen and Peng, Nanyun and Chang, Kai-Wei , title =. arXiv preprint arXiv:2605.12493 , year =
-
[31]
arXiv preprint arXiv:2502.12110 , year =
Xu, Wujiang and Liang, Zujie and Mei, Kai and Gao, Hang and Tan, Juntao and Zhang, Yongfeng , title =. arXiv preprint arXiv:2502.12110 , year =
-
[32]
and Zhang, Hao and Gonzalez, Joseph E
Zheng, Lianmin and Chiang, Wei-Lin and Sheng, Ying and Zhuang, Siyuan and Wu, Zhanghao and Zhuang, Yonghao and Lin, Zi and Li, Zhuohan and Li, Dacheng and Xing, Eric P. and Zhang, Hao and Gonzalez, Joseph E. and Stoica, Ion , title =. Advances in Neural Information Processing Systems , year =
-
[33]
arXiv preprint arXiv:2602.22769 , year =
Zhao, Yujie and Yuan, Boqin and Huang, Junbo and Yuan, Haocheng and Yu, Zhongming and Xu, Haozhou and Hu, Lanxiang and Shankarampeta, Abhilash and Huang, Zimeng and Ni, Wentao and Tian, Yuandong and Zhao, Jishen , title =. arXiv preprint arXiv:2602.22769 , year =
-
[34]
Proceedings of the AAAI Conference on Artificial Intelligence , volume =
Zhong, Wanjun and Guo, Lianghong and Gao, Qiqi and He, Ye and Wang, Yanlin , title =. Proceedings of the AAAI Conference on Artificial Intelligence , volume =. 2024 , url =
2024
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.