Pith. sign in

REVIEW 3 major objections 5 minor 66 references

Long-video agents get better by learning how to search, not by replaying old answers.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-31 16:07 UTC pith:VXAMVP5Q

load-bearing objection Clean systems win on query-only experience reuse for long-video agents; isolation is designed, not audited, but the controlled gains still hold. the 3 major comments →

arxiv 2607.28156 v1 pith:VXAMVP5Q submitted 2026-07-30 cs.CL

RRM: Experience-Driven Reflective Retrieval Memory for Long-Horizon Multimodal Reasoning

classification cs.CL
keywords reflective retrieval memorylong-horizon multimodal reasoningprocedural experiencequery-only reuseentity-centric memory graphlong-video understandingonline lifecycle management
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Long-horizon multimodal agents usually stockpile facts from long videos but still fail when retrieval goes wrong and stay stuck repeating bad search habits. This paper argues the missing piece is reusable procedural knowledge about how to search: what evidence a question needs, which query patterns work, and how to correct failed retrieval. Reflective Retrieval Memory (RRM) keeps a separate experience store distilled from past successful and failed task trajectories, turns selected experiences into query-level guidance only, and still answers solely from newly retrieved current-video facts. A lifecycle manager prunes and merges experiences by reuse, feedback, and time so the store stays useful. On three long-video benchmarks the method beats strong memory agents while using fewer retrieval rounds, supporting the claim that learning how to retrieve improves both accuracy and efficiency.

Core claim

The paper establishes that augmenting an entity-centric multimodal memory graph with reflective experience memory—procedural retrieval knowledge distilled from historical trajectories and reused only as query-level control, with answers conditioned only on newly retrieved current-video evidence—consistently improves long-horizon multimodal reasoning, raising scores over the shared-backbone baseline by 9.1, 5.8, and 7.4 points on M3-Bench-Robot, M3-Bench-Web, and Video-MME-Long while cutting average retrieval rounds.

What carries the argument

Reflective Retrieval Memory (RRM): a third memory layer of structured successful/failure procedural records, gated by Online Query Reflection and anomaly-triggered query-only reuse, plus lifecycle management by usage, reuse feedback, and temporal decay, so experience steers queries without entering the answer context.

Load-bearing premise

The load-bearing premise is that LLM-extracted “how to search” records, after stripping answers and entities and blocking same-video reuse, transfer as pure retrieval control without leftover task-specific leakage under the mini-batch delayed-feedback protocol.

What would settle it

Run the same benchmarks with a leakage audit: if query-only reuse still injects source-video entities or answers into later queries, or if gains vanish when experience is drawn only from truly disjoint videos with no ground-truth-aided failure extraction, the central claim fails.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Long-video agents should store search procedures separately from video facts and never feed historical experience into the answer generator.
  • Failed retrieval trajectories are useful mainly as corrective query strategies, not as full episodes to replay.
  • Within-task anomaly detection plus cross-task experience can raise accuracy while reducing retrieval rounds.
  • Experience stores need online lifecycle control (merge, decay, prune) or they accumulate noise that hurts later tasks.
  • The same query-only reuse pattern can be stacked on other factual memory graphs without redesigning what is stored from the current video.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • If procedural retrieval memory generalizes, streaming robot and web agents could improve mid-deployment without enlarging the video context window.
  • Stricter automated leakage tests (entity overlap between experience text and current-video answers) would be a natural next measurement the paper leaves open.
  • The success/failure split suggests asymmetric trust policies may matter in other agent memory systems beyond video QA.
  • Reducing retrieval rounds at higher accuracy implies experience reuse could cut inference cost in multi-round tool-using agents.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper proposes Reflective Retrieval Memory (RRM), which augments an entity-centric multimodal factual memory graph (episodic + semantic, following M3-Agent) with a third layer of reflective experience memory. From historical task trajectories, RRM distills structured procedural retrieval records (evidence requirements, query patterns, failure modes, adjustment strategies) into separate success and failure banks, reuses them only as query-level control signals when anomaly triggers fire, and keeps answer generation conditioned solely on newly retrieved current-video evidence. Online Query Reflection provides within-task repair; a lifecycle mechanism merges, decays, and prunes experiences. Under a mini-batch online protocol with delayed ground-truth feedback and cross-video isolation, RRM reports gains of 9.1, 5.8, and 7.4 points over M3-Agent on M3-Bench-Robot, M3-Bench-Web, and Video-MME-Long, with fewer average retrieval rounds, supported by progressive ablations and a prompt-injection vs. query-only comparison.

Significance. The work targets a genuine bottleneck in long-horizon multimodal agents: not only what to store, but how to improve retrieval strategy across tasks without contaminating current-video reasoning with historical facts. The separation of procedural experience from factual memory, the query-only reuse constraint, and the controlled comparison against the shared M3-Agent backbone are clear design strengths. Consistent accuracy gains across three benchmarks, complementary ablations (OQR, experience memory, lifecycle), the query-only vs. prompt-level result (Fig. 3a), and simultaneous reductions in retrieval rounds make a credible empirical case that retrieval-control learning is useful. If the isolation premise holds under stronger audits, the framework is a transferable template for memory-augmented multimodal agents beyond the specific backbone.

major comments (3)
  1. [Three-Layer Memory Architecture; Triggered Query-Only Experience Reuse; Fig. 3(a)] Three-Layer Memory Architecture and Triggered Query-Only Experience Reuse: the central attribution—that gains come from transferable procedural retrieval knowledge rather than residual task-specific cues—is enforced by design (answer/entity stripping, cross-video exclusion, query-only injection) and indirectly supported by Fig. 3(a), but never directly measured. There is no audit of extracted records for lingering entities/events/answers, no retrieval-log analysis of whether experience-guided queries still surface source-task content, and no controlled same-domain vs. cross-domain or poisoned-experience test. Without at least one such check, the 9.1/5.8/7.4 gains cannot be cleanly attributed to pure search-control transfer versus a stronger generic retrieval controller. A compact qualitative sample of records plus a simple leakage or transfer split would make the claim load-bearing rathe
  2. [Failure-Aware Experience Modeling; Evaluation Protocol; Table 2] Failure-Aware Experience Modeling and the mini-batch protocol: failure experiences are extracted with access to y*_i and ŷ_i after the batch is scored. That is consistent with the stated delayed-feedback protocol, but it means the failure bank depends on ground-truth labels that are unavailable in true label-free deployment. The paper should either (i) report a success-only / no-GT ablation on the same tables, or (ii) clearly scope failure-bank updates as an offline or weakly supervised mode and show that main gains survive without M−. As written, Table 2’s “Reflective Experience Memory” conflates two regimes with different supervision assumptions.
  3. [Implementation Details; Online Memory Lifecycle Management; Eq. (4)] Implementation Details / free parameters: several load-bearing knobs (mini-batch size 64 with same-video co-batching, max five Search–Answer rounds, at most one success and one failure experience per step, Dedup_L, lifecycle thresholds for merge/prune/cold-start/decay) are fixed without sensitivity analysis. Because experience write timing and anomaly-triggered reuse both depend on these choices, at least a short sensitivity or stability check on batch size and experience cap (e.g., on one benchmark) is needed to show that the headline margins are not brittle to the chosen online schedule.
minor comments (5)
  1. [Abstract; Introduction] Abstract and early pages show systematic missing word spaces (e.g., “mostmethodsemphasize,” “tasktrajectories,” “ReflectiveRetrievalMemory”). Clean the camera-ready text; if this is a PDF extraction artifact, verify the source.
  2. [Figure 2] Figure 2 is central but dense; a short caption walkthrough of the A/B/C path (OQR vs. experience reuse vs. lifecycle write-back) would help readers separate within-task repair from cross-task transfer.
  3. [Table 1] Table 1: NS-Mem has no Video-MME-Long entry; state explicitly whether the method is inapplicable or was not run, to avoid an incomplete SOTA row.
  4. [Reflective Retrieval Control; Eq. (4)] Eqs. (1)–(4) are clear, but Dedup_L and K^base_i,t are only lightly defined in-text; point to the appendix schema or give a one-line definition in the main text.
  5. [Related Work] Related Work correctly positions RRM against Reflexion/ExpeL/AWM/ReMe; a single sentence on why multimodal long-video retrieval makes prompt-level experience injection especially risky (vs. text-only agents) would sharpen the contrast already tested in Fig. 3(a).

Circularity Check

0 steps flagged

No significant circularity: empirical systems gains on external benchmarks, not a self-referential derivation.

full rationale

RRM is an empirical agent/memory systems paper. Its central claims are measured accuracy and retrieval-round improvements versus M3-Agent and other baselines on M3-Bench-Robot, M3-Bench-Web, and Video-MME-Long under a stated mini-batch online protocol with delayed feedback. There is no first-principles derivation in which a quantity is defined from the target and then reported as a prediction; no fitted constant is renamed as an out-of-sample forecast; and no load-bearing uniqueness or ansatz is imported via overlapping-author self-citation. Building on M3-Agent’s factual graph and Search–Answer controller is ordinary baseline reuse, not circular reduction. Online extraction of procedural experiences (including failure records that see y* after prediction is fixed) and lifecycle updates from reuse feedback are standard test-time adaptation loops; they do not make the reported benchmark deltas true by construction. Concerns about unmeasured residual leakage of task-specific content into “procedural” records are validity/isolation issues, not circularity under this rubric. Score 0; steps empty.

Axiom & Free-Parameter Ledger

5 free parameters · 5 axioms · 3 invented entities

Load-bearing commitments are architectural and evaluation-protocol choices, not physical axioms. The central claim rests on M3-Agent’s factual graph and Search–Answer controller as given infrastructure; on LLM extractors/selectors producing transferable procedural records; on anomaly heuristics triggering reuse; and on a mini-batch delayed-GT protocol that defines when experience may influence later tasks. Free parameters are operational hyperparameters (batch size, retrieval budget, experience caps, lifecycle decay/merge/prune rules) rather than fitted scientific constants.

free parameters (5)
  • mini-batch size (64 questions; same-video co-batching) = 64
    Defines the delayed-feedback horizon and prevents within-video experience leakage by construction; chosen operationally, not derived.
  • max Search–Answer rounds = 5
    Caps retrieval budget; affects both accuracy and reported efficiency comparisons.
  • experience retrieval cap per step = 1 success + 1 failure
    At most one successful and one failure experience consulted per retrieval step; controls guidance strength.
  • Dedup_L candidate query cap = at most L (value in appendix)
    Limits supplemental queries after merging OQR, experience focus, and baseline keys; L is a capacity knob.
  • lifecycle thresholds (usage, reuse feedback, temporal decay, cold-start protection, merge/prune)
    Utility estimation and pruning rules that keep experience memory compact; hand-set hyperparameters claimed in appendix.
axioms (5)
  • domain assumption M3-Agent entity-centric episodic/semantic graph and Search–Answer controller are an adequate factual retrieval backbone; RRM only augments query construction.
    Stated in Preliminary and Method; all main gains are interpreted relative to this inherited stack.
  • ad hoc to paper Procedural retrieval knowledge can be separated from task-specific facts by LLM extraction that strips answers, option labels, and source entities.
    Core design claim in Reflective Experience Memory and Failure-Aware Experience Modeling; not independently validated with a leakage metric.
  • ad hoc to paper Anomaly triggers (no valid evidence, near-duplicate queries, no new evidence across rounds) are sufficient signals to invoke OQR and experience reuse.
    Online Query Reflection / Failure-State Detector; rule-based, not learned.
  • domain assumption Mini-batch online adaptation with delayed ground-truth and LLM binary success labels is a valid evaluation of forward-only test-time learning (citing Evo-Memory-style protocol).
    Evaluation Protocol section; experiences affect only later mini-batches.
  • domain assumption Failure experiences are less reliable than success experiences and must be selected asymmetrically under stricter applicability match.
    Failure-Aware Experience Modeling; motivates dual banks M+ / M−.
invented entities (3)
  • Reflective experience memory (M_ref = {M+, M−}) no independent evidence
    purpose: Store cross-task procedural retrieval strategies separately from current-video factual memory.
    Third memory layer introduced by the paper; contents are LLM-distilled structured records, not a new physical object. Independent evidence is only downstream benchmark gains.
  • Experience-derived retrieval focus f_ref_i,t no independent evidence
    purpose: Neutral query-level control signal produced from selected experience without historical answers/entities.
    Interface that enforces query-only reuse in equation (4)’s supplemental query set.
  • Online Query Reflection (OQR) local repair query no independent evidence
    purpose: Within-task anomaly repair before/without cross-task experience.
    Component A in ablations; rule-triggered LLM rewrite of the current search trajectory.

pith-pipeline@v1.2.0-daily-grok45 · 17099 in / 3978 out tokens · 79261 ms · 2026-07-31T16:07:27.069829+00:00 · methodology

0 comments
read the original abstract

Existing multimodal long-term memory agents use external memory to overcome the limited context available for long videos. However, most methods emphasize what to store rather than how stored memory should be retrieved. When retrieval becomes inaccurate or repeatedly fails to obtain useful evidence, existing agents lack mechanisms to diagnose failures from previous task trajectories and adapt future search strategies.We introduce Reflective Retrieval Memory (RRM), a reflective memory framework for long-horizon multimodal reasoning. RRM augments an entity-centric multimodal memory graph with reflective experience memory, which distills transferable procedural retrieval knowledge from historical task trajectories. Unlike episodic and semantic memories that preserve factual evidence from the current video, reflective experience memory captures reusable search strategies across tasks. RRM converts retrieved experiences into query-level guidance, while answer generation remains conditioned only on factual evidence newly retrieved from the current video. A lifecycle management mechanism further regulates experience memory through usage frequency, reuse feedback, and temporal decay, thereby reducing redundancy and noise. RRM consistently outperforms previous state-of-the-art approaches on M3-Bench-Robot, M3-Bench-Web, and Video-MME-Long, demonstrating the effectiveness of reflective retrieval memory for long-horizon multimodal reasoning.

Figures

Figures reproduced from arXiv: 2607.28156 by Bochao Zou, Jingxiang Fan, Junbao Zhuo.

Figure 1
Figure 1. Figure 1: Motivation and retrieval workflow of RRM. Procedural retrieval experience distilled from prior successful and failed [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: Overall architecture of RRM: (a) current-video factual memory, (b) reflective retrieval control with Online Query [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Experience reuse effectiveness and retrieval effi [PITH_FULL_IMAGE:figures/full_fig_p007_3.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

66 extracted references · 4 linked inside Pith

  1. [1]

    Xu, Jin and Guo, Zhifang and He, Jinzheng and Hu, Hangrui and He, Ting and Bai, Shuai and Chen, Keqin and Wang, Jialin and Fan, Yang and Dang, Kai and Zhang, Bin and Wang, Xiong and Chu, Yunfei and Lin, Junyang , journal=

  2. [2]

    Bai, Shuai and Chen, Keqin and Liu, Xuejing and Wang, Jialin and Ge, Wenbin and Song, Sibo and Dang, Kai and Wang, Peng and Wang, Shijie and Tang, Jun and Zhong, Humen and Zhu, Yuanzhi and Yang, Mingkun and Li, Zhaohai and Wan, Jianqiang and Wang, Pengfei and Ding, Wei and Fu, Zheren and Xu, Yiheng and Ye, Jiabo and Zhang, Xi and Xie, Tianbao and Cheng, Z...

  3. [3]

    Reid, Machel and Savinov, Nikolay and Teplyashin, Denis and Lepikhin, Dmitry and Lillicrap, Timothy and Alayrac, Jean-Baptiste and Soricut, Radu and Lazaridou, Angeliki and Firat, Orhan and Schrittwieser, Julian and others , journal=

  4. [4]

    2024 , howpublished=

  5. [5]

    Song, Enxin and Chai, Wenhao and Wang, Guanhong and Zhang, Yucheng and Zhou, Haoyang and Wu, Feiyang and Chi, Haozhe and Guo, Xun and Ye, Tian and Zhang, Yanting and Lu, Yan and Hwang, Jenq-Neng and Wang, Gaoang , booktitle=

  6. [6]

    He, Bo and Li, Hengduo and Jang, Young Kyun and Jia, Menglin and Cao, Xuefei and Shah, Ashish and Shrivastava, Abhinav and Lim, Ser-Nam , booktitle=

  7. [7]

    2024 , doi=

    Wang, Xiaohan and Zhang, Yuhui and Zohar, Orr and Yeung-Levy, Serena , booktitle=. 2024 , doi=

  8. [8]

    Long, Lin and He, Yichen and Ye, Wentao and Pan, Yiyuan and Lin, Yuan and Li, Hang and Zhao, Junbo and Li, Wei , booktitle=

  9. [9]

    Yeo, Woongyeong and Kim, Kangsan and Yoon, Jaehong and Hwang, Sung Ju , booktitle=

  10. [10]

    Zhang, Haoji and Wang, Yiqin and Tang, Yansong and Liu, Yong and Feng, Jiashi and Jin, Xiaojie , booktitle=

  11. [11]

    Huang, Zhenpeng and Li, Xinhao and Li, Jiaqi and Wang, Jing and Zeng, Xiangyu and Liang, Cheng and Wu, Tao and Chen, Xi and Li, Liang and Wang, Limin , booktitle=

  12. [12]

    Yao, Linli and Li, Yicheng and Wei, Yuancheng and Li, Lei and Ren, Shuhuai and Liu, Yuanxin and Ouyang, Kun and Wang, Lean and Li, Shicheng and Li, Sida and Kong, Lingpeng and Liu, Qi and Zhang, Yuanxing and Sun, Xu , booktitle=

  13. [13]

    Zeng, Xiangyu and Qiu, Kefan and Zhang, Qingyu and Li, Xinhao and Wang, Jing and Li, Jiaxin and Yan, Ziang and Tian, Kun and Tian, Meng and Zhao, Xinhai and Wang, Yi and Wang, Limin , booktitle=

  14. [14]

    Jiang, Rongjie and Wang, Jianwei and Zhao, Gengda and Luo, Chengyang and Wang, Kai and Zhang, Wenjie , journal=

  15. [15]

    2026 , doi=

    Wang, Junxi and Sun, Te and Zhu, Jiayi and Li, Junxian and Xu, Haowen and Wen, Zichen and Hu, Xuming and Li, Zhiyu and Zhang, Linfeng , booktitle=. 2026 , doi=

  16. [16]

    2025 , doi=

    Fu, Chaoyou and Dai, Yuhan and Luo, Yongdong and Li, Lei and Ren, Shuhuai and Zhang, Renrui and Wang, Zihan and Zhou, Chenyu and Shen, Yunhang and Zhang, Mengdan and Chen, Peixian and Li, Yanwei and Lin, Shaohui and Zhao, Sirui and Li, Ke and Xu, Tong and Zheng, Xiawu and Chen, Enhong and Shan, Caifeng and He, Ran and Sun, Xing , booktitle=. 2025 , doi=

  17. [17]

    Shinn, Noah and Cassano, Federico and Gopinath, Ashwin and Narasimhan, Karthik and Yao, Shunyu , booktitle=

  18. [18]

    2024 , doi=

    Zhao, Andrew and Huang, Daniel and Xu, Quentin and Lin, Matthieu and Liu, Yong-Jin and Huang, Gao , booktitle=. 2024 , doi=

  19. [19]

    2025 , publisher=

    Wang, Zora Zhiruo and Mao, Jiayuan and Fried, Daniel and Neubig, Graham , booktitle=. 2025 , publisher=

  20. [20]

    2026 , doi=

    Cao, Zouying and Deng, Jiaji and Yu, Li and Zhou, Weikang and Liu, Zhaoyang and Ding, Bolin and Zhao, Hai , booktitle=. 2026 , doi=

  21. [21]

    Ma, Ziyu and Gou, Chenhui and Shi, Hengcan and Sun, Bin and Li, Shutao and Rezatofighi, Hamid and Cai, Jianfei , booktitle=

  22. [22]

    Kim, Junho and Kim, Hyunjun and Lee, Hosu and Ro, Yong Man , booktitle=

  23. [23]

    Diko, Anxhelo and Wang, Tinghuai and Swaileh, Wassim and Sun, Shiyan and Patras, Ioannis , booktitle=

  24. [24]

    Zhang, Yanzhao and Li, Mingxin and Long, Dingkun and Zhang, Xin and Lin, Huan and Yang, Baosong and Xie, Pengjun and Yang, An and Liu, Dayiheng and Lin, Junyang and Huang, Fei and Zhou, Jingren , journal=

  25. [25]

    Proceedings of the 40th International Conference on Machine Learning , series =

    Large Language Models Can Be Easily Distracted by Irrelevant Context , author =. Proceedings of the 40th International Conference on Machine Learning , series =. 2023 , publisher =

  26. [26]

    Findings of the Association for Computational Linguistics: NAACL 2024 , pages =

    Why So Gullible? Enhancing the Robustness of Retrieval-Augmented Models against Counterfactual Noise , author =. Findings of the Association for Computational Linguistics: NAACL 2024 , pages =. 2024 , publisher =

  27. [27]

    Transactions of the Association for Computational Linguistics , volume =

    Lost in the Middle: How Language Models Use Long Contexts , author =. Transactions of the Association for Computational Linguistics , volume =. 2024 , publisher =

  28. [28]

    2024 , url =

    Hsieh, Cheng-Ping and Sun, Simeng and Kriman, Samuel and Acharya, Shantanu and Rekesh, Dima and Jia, Fei and Ginsburg, Boris , booktitle =. 2024 , url =

  29. [29]

    2024 , doi =

    Wu, Haoning and Li, Dongxu and Chen, Bei and Li, Junnan , booktitle =. 2024 , doi =

  30. [31]

    Communication, Simulation, and Intelligent Agents: Implications of Personal Intelligent Machines for Medical Education

    Clancey, William J. Communication, Simulation, and Intelligent Agents: Implications of Personal Intelligent Machines for Medical Education. Proceedings of the Eighth International Joint Conference on Artificial Intelligence (IJCAI-83)

  31. [32]

    Classification Problem Solving

    Clancey, William J. Classification Problem Solving. Proceedings of the Fourth National Conference on Artificial Intelligence

  32. [33]

    , title =

    Robinson, Arthur L. , title =. 1980 , doi =. https://science.sciencemag.org/content/208/4447/1019.full.pdf , journal =

  33. [34]

    New Ways to Make Microcircuits Smaller---Duplicate Entry

    Robinson, Arthur L. New Ways to Make Microcircuits Smaller---Duplicate Entry. Science

  34. [35]

    Clancey and Glenn Rennels , abstract =

    Diane Warner Hasling and William J. Clancey and Glenn Rennels , abstract =. Strategic explanations for a diagnostic consultation system , journal =. 1984 , issn =. doi:https://doi.org/10.1016/S0020-7373(84)80003-6 , url =

  35. [36]

    and Rennels, Glenn R

    Hasling, Diane Warner and Clancey, William J. and Rennels, Glenn R. and Test, Thomas. Strategic Explanations in Consultation---Duplicate. The International Journal of Man-Machine Studies

  36. [37]

    Poligon: A System for Parallel Problem Solving

    Rice, James. Poligon: A System for Parallel Problem Solving

  37. [38]

    Transfer of Rule-Based Expertise through a Tutorial Dialogue

    Clancey, William J. Transfer of Rule-Based Expertise through a Tutorial Dialogue

  38. [39]

    The Engineering of Qualitative Models

    Clancey, William J. The Engineering of Qualitative Models

  39. [40]

    2023 , eprint=

    Attention Is All You Need , author=. 2023 , eprint=

  40. [41]

    Pluto: The 'Other' Red Planet

    NASA. Pluto: The 'Other' Red Planet

  41. [42]

    Bai, S.; Chen, K.; Liu, X.; Wang, J.; Ge, W.; Song, S.; Dang, K.; Wang, P.; Wang, S.; Tang, J.; Zhong, H.; Zhu, Y.; Yang, M.; Li, Z.; Wan, J.; Wang, P.; Ding, W.; Fu, Z.; Xu, Y.; Ye, J.; Zhang, X.; Xie, T.; Cheng, Z.; Zhang, H.; Yang, Z.; Xu, H.; and Lin, J. 2025. Qwen2.5-VL Technical Report . arXiv preprint arXiv:2502.13923

  42. [43]

    Cao, Z.; Deng, J.; Yu, L.; Zhou, W.; Liu, Z.; Ding, B.; and Zhao, H. 2026. Remember Me, Refine Me: A Dynamic Procedural Memory Framework for Experience-Driven Agent Evolution . In Findings of the Association for Computational Linguistics: ACL 2026, 16803--16822

  43. [44]

    Fu, C.; Dai, Y.; Luo, Y.; Li, L.; Ren, S.; Zhang, R.; Wang, Z.; Zhou, C.; Shen, Y.; Zhang, M.; Chen, P.; Li, Y.; Lin, S.; Zhao, S.; Li, K.; Xu, T.; Zheng, X.; Chen, E.; Shan, C.; He, R.; and Sun, X. 2025. Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-Modal LLMs in Video Analysis . In Proceedings of the IEEE/CVF Conference on Comput...

  44. [45]

    K.; Jia, M.; Cao, X.; Shah, A.; Shrivastava, A.; and Lim, S.-N

    He, B.; Li, H.; Jang, Y. K.; Jia, M.; Cao, X.; Shah, A.; Shrivastava, A.; and Lim, S.-N. 2024. MA-LMM: Memory-Augmented Large Multimodal Model for Long-Term Video Understanding . In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 13504--13514

  45. [46]

    Hong, G.; Kim, J.; Kang, J.; Myaeng, S.-H.; and Whang, J. J. 2024. Why So Gullible? Enhancing the Robustness of Retrieval-Augmented Models against Counterfactual Noise. In Findings of the Association for Computational Linguistics: NAACL 2024, 2474--2495. Association for Computational Linguistics

  46. [47]

    Hsieh, C.-P.; Sun, S.; Kriman, S.; Acharya, S.; Rekesh, D.; Jia, F.; and Ginsburg, B. 2024. RULER : What's the Real Context Size of Your Long-Context Language Models? In First Conference on Language Modeling

  47. [48]

    Huang, Z.; Li, X.; Li, J.; Wang, J.; Zeng, X.; Liang, C.; Wu, T.; Chen, X.; Li, L.; and Wang, L. 2025. Online Video Understanding: OVBench and VideoChat-Online . In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 3328--3338

  48. [49]

    Jiang, R.; Wang, J.; Zhao, G.; Luo, C.; Wang, K.; and Zhang, W. 2026. Advancing Multimodal Agent Reasoning with Long-Term Neuro-Symbolic Memory . arXiv preprint arXiv:2603.15280

  49. [50]

    F.; Lin, K.; Hewitt, J.; Paranjape, A.; Bevilacqua, M.; Petroni, F.; and Liang, P

    Liu, N. F.; Lin, K.; Hewitt, J.; Paranjape, A.; Bevilacqua, M.; Petroni, F.; and Liang, P. 2024. Lost in the Middle: How Language Models Use Long Contexts. Transactions of the Association for Computational Linguistics, 12: 157--173

  50. [51]

    Long, L.; He, Y.; Ye, W.; Pan, Y.; Lin, Y.; Li, H.; Zhao, J.; and Li, W. 2026. Seeing, Listening, Remembering, and Reasoning: A Multimodal Agent with Long-Term Memory . In International Conference on Learning Representations

  51. [52]

    OpenAI . 2024. GPT-4o System Card . Technical report

  52. [53]

    Reid, M.; Savinov, N.; Teplyashin, D.; Lepikhin, D.; Lillicrap, T.; Alayrac, J.-B.; Soricut, R.; Lazaridou, A.; Firat, O.; Schrittwieser, J.; et al. 2024. Gemini 1.5: Unlocking Multimodal Understanding Across Millions of Tokens of Context . arXiv preprint arXiv:2403.05530

  53. [54]

    H.; Sch \"a rli, N.; and Zhou, D

    Shi, F.; Chen, X.; Misra, K.; Scales, N.; Dohan, D.; Chi, E. H.; Sch \"a rli, N.; and Zhou, D. 2023. Large Language Models Can Be Easily Distracted by Irrelevant Context. In Proceedings of the 40th International Conference on Machine Learning, volume 202 of Proceedings of Machine Learning Research, 31210--31227. PMLR

  54. [55]

    Shinn, N.; Cassano, F.; Gopinath, A.; Narasimhan, K.; and Yao, S. 2023. Reflexion: Language Agents with Verbal Reinforcement Learning . In Advances in Neural Information Processing Systems, volume 36

  55. [56]

    Song, E.; Chai, W.; Wang, G.; Zhang, Y.; Zhou, H.; Wu, F.; Chi, H.; Guo, X.; Ye, T.; Zhang, Y.; Lu, Y.; Hwang, J.-N.; and Wang, G. 2024. MovieChat: From Dense Token to Sparse Memory for Long Video Understanding . In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 18221--18232

  56. [57]

    Wang, J.; Sun, T.; Zhu, J.; Li, J.; Xu, H.; Wen, Z.; Hu, X.; Li, Z.; and Zhang, L. 2026. StreamMeCo: Long-Term Agent Memory Compression for Efficient Streaming Video Understanding . In Findings of the Association for Computational Linguistics: ACL 2026, 13234--13251

  57. [58]

    Wang, X.; Zhang, Y.; Zohar, O.; and Yeung-Levy, S. 2024. VideoAgent: Long-Form Video Understanding with Large Language Model as Agent . In Proceedings of the European Conference on Computer Vision

  58. [59]

    Z.; Mao, J.; Fried, D.; and Neubig, G

    Wang, Z. Z.; Mao, J.; Fried, D.; and Neubig, G. 2025. Agent Workflow Memory . In Proceedings of the 42nd International Conference on Machine Learning, volume 267 of Proceedings of Machine Learning Research, 63897--63911. PMLR

  59. [60]

    H.; Wang, C.; Chen, S.; Pereira, F.; Kang, W.-C.; and Cheng, D

    Wei, T.; Sachdeva, N.; Coleman, B.; He, Z.; Bei, Y.; Ning, X.; Ai, M.; Li, Y.; He, J.; Chi, E. H.; Wang, C.; Chen, S.; Pereira, F.; Kang, W.-C.; and Cheng, D. Z. 2025. Evo-Memory: Benchmarking LLM Agent Test-time Learning with Self-Evolving Memory. arXiv preprint arXiv:2511.20857

  60. [61]

    Wu, H.; Li, D.; Chen, B.; and Li, J. 2024. LongVideoBench : A Benchmark for Long-Context Interleaved Video-Language Understanding. In Advances in Neural Information Processing Systems, volume 37

  61. [62]

    Xu, J.; Guo, Z.; He, J.; Hu, H.; He, T.; Bai, S.; Chen, K.; Wang, J.; Fan, Y.; Dang, K.; Zhang, B.; Wang, X.; Chu, Y.; and Lin, J. 2025. Qwen2.5-Omni Technical Report . arXiv preprint arXiv:2503.20215

  62. [63]

    Yao, L.; Li, Y.; Wei, Y.; Li, L.; Ren, S.; Liu, Y.; Ouyang, K.; Wang, L.; Li, S.; Li, S.; Kong, L.; Liu, Q.; Zhang, Y.; and Sun, X. 2025. TimeChat-Online: 80\ In Proceedings of the 33rd ACM International Conference on Multimedia, 10807--10816

  63. [64]

    Yeo, W.; Kim, K.; Yoon, J.; and Hwang, S. J. 2026. WorldMM: Dynamic Multimodal Memory Agent for Long Video Reasoning . In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 25599--25609

  64. [65]

    Zeng, X.; Qiu, K.; Zhang, Q.; Li, X.; Wang, J.; Li, J.; Yan, Z.; Tian, K.; Tian, M.; Zhao, X.; Wang, Y.; and Wang, L. 2025. StreamForest: Efficient Online Video Understanding with Persistent Event Memory . In Advances in Neural Information Processing Systems, volume 38

  65. [66]

    Zhang, H.; Wang, Y.; Tang, Y.; Liu, Y.; Feng, J.; and Jin, X. 2025. Flash-VStream: Efficient Real-Time Understanding for Long Video Streams . In Proceedings of the IEEE/CVF International Conference on Computer Vision, 21059--21069

  66. [67]

    Zhao, A.; Huang, D.; Xu, Q.; Lin, M.; Liu, Y.-J.; and Huang, G. 2024. ExpeL: LLM Agents Are Experiential Learners . In Proceedings of the AAAI Conference on Artificial Intelligence, volume 38, 19632--19642