REVIEW 3 major objections 5 minor 66 references
Long-video agents get better by learning how to search, not by replaying old answers.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-31 16:07 UTC pith:VXAMVP5Q
load-bearing objection Clean systems win on query-only experience reuse for long-video agents; isolation is designed, not audited, but the controlled gains still hold. the 3 major comments →
RRM: Experience-Driven Reflective Retrieval Memory for Long-Horizon Multimodal Reasoning
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The paper establishes that augmenting an entity-centric multimodal memory graph with reflective experience memory—procedural retrieval knowledge distilled from historical trajectories and reused only as query-level control, with answers conditioned only on newly retrieved current-video evidence—consistently improves long-horizon multimodal reasoning, raising scores over the shared-backbone baseline by 9.1, 5.8, and 7.4 points on M3-Bench-Robot, M3-Bench-Web, and Video-MME-Long while cutting average retrieval rounds.
What carries the argument
Reflective Retrieval Memory (RRM): a third memory layer of structured successful/failure procedural records, gated by Online Query Reflection and anomaly-triggered query-only reuse, plus lifecycle management by usage, reuse feedback, and temporal decay, so experience steers queries without entering the answer context.
Load-bearing premise
The load-bearing premise is that LLM-extracted “how to search” records, after stripping answers and entities and blocking same-video reuse, transfer as pure retrieval control without leftover task-specific leakage under the mini-batch delayed-feedback protocol.
What would settle it
Run the same benchmarks with a leakage audit: if query-only reuse still injects source-video entities or answers into later queries, or if gains vanish when experience is drawn only from truly disjoint videos with no ground-truth-aided failure extraction, the central claim fails.
If this is right
- Long-video agents should store search procedures separately from video facts and never feed historical experience into the answer generator.
- Failed retrieval trajectories are useful mainly as corrective query strategies, not as full episodes to replay.
- Within-task anomaly detection plus cross-task experience can raise accuracy while reducing retrieval rounds.
- Experience stores need online lifecycle control (merge, decay, prune) or they accumulate noise that hurts later tasks.
- The same query-only reuse pattern can be stacked on other factual memory graphs without redesigning what is stored from the current video.
Where Pith is reading between the lines
- If procedural retrieval memory generalizes, streaming robot and web agents could improve mid-deployment without enlarging the video context window.
- Stricter automated leakage tests (entity overlap between experience text and current-video answers) would be a natural next measurement the paper leaves open.
- The success/failure split suggests asymmetric trust policies may matter in other agent memory systems beyond video QA.
- Reducing retrieval rounds at higher accuracy implies experience reuse could cut inference cost in multi-round tool-using agents.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes Reflective Retrieval Memory (RRM), which augments an entity-centric multimodal factual memory graph (episodic + semantic, following M3-Agent) with a third layer of reflective experience memory. From historical task trajectories, RRM distills structured procedural retrieval records (evidence requirements, query patterns, failure modes, adjustment strategies) into separate success and failure banks, reuses them only as query-level control signals when anomaly triggers fire, and keeps answer generation conditioned solely on newly retrieved current-video evidence. Online Query Reflection provides within-task repair; a lifecycle mechanism merges, decays, and prunes experiences. Under a mini-batch online protocol with delayed ground-truth feedback and cross-video isolation, RRM reports gains of 9.1, 5.8, and 7.4 points over M3-Agent on M3-Bench-Robot, M3-Bench-Web, and Video-MME-Long, with fewer average retrieval rounds, supported by progressive ablations and a prompt-injection vs. query-only comparison.
Significance. The work targets a genuine bottleneck in long-horizon multimodal agents: not only what to store, but how to improve retrieval strategy across tasks without contaminating current-video reasoning with historical facts. The separation of procedural experience from factual memory, the query-only reuse constraint, and the controlled comparison against the shared M3-Agent backbone are clear design strengths. Consistent accuracy gains across three benchmarks, complementary ablations (OQR, experience memory, lifecycle), the query-only vs. prompt-level result (Fig. 3a), and simultaneous reductions in retrieval rounds make a credible empirical case that retrieval-control learning is useful. If the isolation premise holds under stronger audits, the framework is a transferable template for memory-augmented multimodal agents beyond the specific backbone.
major comments (3)
- [Three-Layer Memory Architecture; Triggered Query-Only Experience Reuse; Fig. 3(a)] Three-Layer Memory Architecture and Triggered Query-Only Experience Reuse: the central attribution—that gains come from transferable procedural retrieval knowledge rather than residual task-specific cues—is enforced by design (answer/entity stripping, cross-video exclusion, query-only injection) and indirectly supported by Fig. 3(a), but never directly measured. There is no audit of extracted records for lingering entities/events/answers, no retrieval-log analysis of whether experience-guided queries still surface source-task content, and no controlled same-domain vs. cross-domain or poisoned-experience test. Without at least one such check, the 9.1/5.8/7.4 gains cannot be cleanly attributed to pure search-control transfer versus a stronger generic retrieval controller. A compact qualitative sample of records plus a simple leakage or transfer split would make the claim load-bearing rathe
- [Failure-Aware Experience Modeling; Evaluation Protocol; Table 2] Failure-Aware Experience Modeling and the mini-batch protocol: failure experiences are extracted with access to y*_i and ŷ_i after the batch is scored. That is consistent with the stated delayed-feedback protocol, but it means the failure bank depends on ground-truth labels that are unavailable in true label-free deployment. The paper should either (i) report a success-only / no-GT ablation on the same tables, or (ii) clearly scope failure-bank updates as an offline or weakly supervised mode and show that main gains survive without M−. As written, Table 2’s “Reflective Experience Memory” conflates two regimes with different supervision assumptions.
- [Implementation Details; Online Memory Lifecycle Management; Eq. (4)] Implementation Details / free parameters: several load-bearing knobs (mini-batch size 64 with same-video co-batching, max five Search–Answer rounds, at most one success and one failure experience per step, Dedup_L, lifecycle thresholds for merge/prune/cold-start/decay) are fixed without sensitivity analysis. Because experience write timing and anomaly-triggered reuse both depend on these choices, at least a short sensitivity or stability check on batch size and experience cap (e.g., on one benchmark) is needed to show that the headline margins are not brittle to the chosen online schedule.
minor comments (5)
- [Abstract; Introduction] Abstract and early pages show systematic missing word spaces (e.g., “mostmethodsemphasize,” “tasktrajectories,” “ReflectiveRetrievalMemory”). Clean the camera-ready text; if this is a PDF extraction artifact, verify the source.
- [Figure 2] Figure 2 is central but dense; a short caption walkthrough of the A/B/C path (OQR vs. experience reuse vs. lifecycle write-back) would help readers separate within-task repair from cross-task transfer.
- [Table 1] Table 1: NS-Mem has no Video-MME-Long entry; state explicitly whether the method is inapplicable or was not run, to avoid an incomplete SOTA row.
- [Reflective Retrieval Control; Eq. (4)] Eqs. (1)–(4) are clear, but Dedup_L and K^base_i,t are only lightly defined in-text; point to the appendix schema or give a one-line definition in the main text.
- [Related Work] Related Work correctly positions RRM against Reflexion/ExpeL/AWM/ReMe; a single sentence on why multimodal long-video retrieval makes prompt-level experience injection especially risky (vs. text-only agents) would sharpen the contrast already tested in Fig. 3(a).
Circularity Check
No significant circularity: empirical systems gains on external benchmarks, not a self-referential derivation.
full rationale
RRM is an empirical agent/memory systems paper. Its central claims are measured accuracy and retrieval-round improvements versus M3-Agent and other baselines on M3-Bench-Robot, M3-Bench-Web, and Video-MME-Long under a stated mini-batch online protocol with delayed feedback. There is no first-principles derivation in which a quantity is defined from the target and then reported as a prediction; no fitted constant is renamed as an out-of-sample forecast; and no load-bearing uniqueness or ansatz is imported via overlapping-author self-citation. Building on M3-Agent’s factual graph and Search–Answer controller is ordinary baseline reuse, not circular reduction. Online extraction of procedural experiences (including failure records that see y* after prediction is fixed) and lifecycle updates from reuse feedback are standard test-time adaptation loops; they do not make the reported benchmark deltas true by construction. Concerns about unmeasured residual leakage of task-specific content into “procedural” records are validity/isolation issues, not circularity under this rubric. Score 0; steps empty.
Axiom & Free-Parameter Ledger
free parameters (5)
- mini-batch size (64 questions; same-video co-batching) =
64
- max Search–Answer rounds =
5
- experience retrieval cap per step =
1 success + 1 failure
- Dedup_L candidate query cap =
at most L (value in appendix)
- lifecycle thresholds (usage, reuse feedback, temporal decay, cold-start protection, merge/prune)
axioms (5)
- domain assumption M3-Agent entity-centric episodic/semantic graph and Search–Answer controller are an adequate factual retrieval backbone; RRM only augments query construction.
- ad hoc to paper Procedural retrieval knowledge can be separated from task-specific facts by LLM extraction that strips answers, option labels, and source entities.
- ad hoc to paper Anomaly triggers (no valid evidence, near-duplicate queries, no new evidence across rounds) are sufficient signals to invoke OQR and experience reuse.
- domain assumption Mini-batch online adaptation with delayed ground-truth and LLM binary success labels is a valid evaluation of forward-only test-time learning (citing Evo-Memory-style protocol).
- domain assumption Failure experiences are less reliable than success experiences and must be selected asymmetrically under stricter applicability match.
invented entities (3)
-
Reflective experience memory (M_ref = {M+, M−})
no independent evidence
-
Experience-derived retrieval focus f_ref_i,t
no independent evidence
-
Online Query Reflection (OQR) local repair query
no independent evidence
read the original abstract
Existing multimodal long-term memory agents use external memory to overcome the limited context available for long videos. However, most methods emphasize what to store rather than how stored memory should be retrieved. When retrieval becomes inaccurate or repeatedly fails to obtain useful evidence, existing agents lack mechanisms to diagnose failures from previous task trajectories and adapt future search strategies.We introduce Reflective Retrieval Memory (RRM), a reflective memory framework for long-horizon multimodal reasoning. RRM augments an entity-centric multimodal memory graph with reflective experience memory, which distills transferable procedural retrieval knowledge from historical task trajectories. Unlike episodic and semantic memories that preserve factual evidence from the current video, reflective experience memory captures reusable search strategies across tasks. RRM converts retrieved experiences into query-level guidance, while answer generation remains conditioned only on factual evidence newly retrieved from the current video. A lifecycle management mechanism further regulates experience memory through usage frequency, reuse feedback, and temporal decay, thereby reducing redundancy and noise. RRM consistently outperforms previous state-of-the-art approaches on M3-Bench-Robot, M3-Bench-Web, and Video-MME-Long, demonstrating the effectiveness of reflective retrieval memory for long-horizon multimodal reasoning.
Figures
Reference graph
Works this paper leans on
-
[1]
Xu, Jin and Guo, Zhifang and He, Jinzheng and Hu, Hangrui and He, Ting and Bai, Shuai and Chen, Keqin and Wang, Jialin and Fan, Yang and Dang, Kai and Zhang, Bin and Wang, Xiong and Chu, Yunfei and Lin, Junyang , journal=
-
[2]
Bai, Shuai and Chen, Keqin and Liu, Xuejing and Wang, Jialin and Ge, Wenbin and Song, Sibo and Dang, Kai and Wang, Peng and Wang, Shijie and Tang, Jun and Zhong, Humen and Zhu, Yuanzhi and Yang, Mingkun and Li, Zhaohai and Wan, Jianqiang and Wang, Pengfei and Ding, Wei and Fu, Zheren and Xu, Yiheng and Ye, Jiabo and Zhang, Xi and Xie, Tianbao and Cheng, Z...
-
[3]
Reid, Machel and Savinov, Nikolay and Teplyashin, Denis and Lepikhin, Dmitry and Lillicrap, Timothy and Alayrac, Jean-Baptiste and Soricut, Radu and Lazaridou, Angeliki and Firat, Orhan and Schrittwieser, Julian and others , journal=
-
[4]
2024 , howpublished=
2024
-
[5]
Song, Enxin and Chai, Wenhao and Wang, Guanhong and Zhang, Yucheng and Zhou, Haoyang and Wu, Feiyang and Chi, Haozhe and Guo, Xun and Ye, Tian and Zhang, Yanting and Lu, Yan and Hwang, Jenq-Neng and Wang, Gaoang , booktitle=
-
[6]
He, Bo and Li, Hengduo and Jang, Young Kyun and Jia, Menglin and Cao, Xuefei and Shah, Ashish and Shrivastava, Abhinav and Lim, Ser-Nam , booktitle=
-
[7]
2024 , doi=
Wang, Xiaohan and Zhang, Yuhui and Zohar, Orr and Yeung-Levy, Serena , booktitle=. 2024 , doi=
2024
-
[8]
Long, Lin and He, Yichen and Ye, Wentao and Pan, Yiyuan and Lin, Yuan and Li, Hang and Zhao, Junbo and Li, Wei , booktitle=
-
[9]
Yeo, Woongyeong and Kim, Kangsan and Yoon, Jaehong and Hwang, Sung Ju , booktitle=
-
[10]
Zhang, Haoji and Wang, Yiqin and Tang, Yansong and Liu, Yong and Feng, Jiashi and Jin, Xiaojie , booktitle=
-
[11]
Huang, Zhenpeng and Li, Xinhao and Li, Jiaqi and Wang, Jing and Zeng, Xiangyu and Liang, Cheng and Wu, Tao and Chen, Xi and Li, Liang and Wang, Limin , booktitle=
-
[12]
Yao, Linli and Li, Yicheng and Wei, Yuancheng and Li, Lei and Ren, Shuhuai and Liu, Yuanxin and Ouyang, Kun and Wang, Lean and Li, Shicheng and Li, Sida and Kong, Lingpeng and Liu, Qi and Zhang, Yuanxing and Sun, Xu , booktitle=
-
[13]
Zeng, Xiangyu and Qiu, Kefan and Zhang, Qingyu and Li, Xinhao and Wang, Jing and Li, Jiaxin and Yan, Ziang and Tian, Kun and Tian, Meng and Zhao, Xinhai and Wang, Yi and Wang, Limin , booktitle=
-
[14]
Jiang, Rongjie and Wang, Jianwei and Zhao, Gengda and Luo, Chengyang and Wang, Kai and Zhang, Wenjie , journal=
-
[15]
2026 , doi=
Wang, Junxi and Sun, Te and Zhu, Jiayi and Li, Junxian and Xu, Haowen and Wen, Zichen and Hu, Xuming and Li, Zhiyu and Zhang, Linfeng , booktitle=. 2026 , doi=
2026
-
[16]
2025 , doi=
Fu, Chaoyou and Dai, Yuhan and Luo, Yongdong and Li, Lei and Ren, Shuhuai and Zhang, Renrui and Wang, Zihan and Zhou, Chenyu and Shen, Yunhang and Zhang, Mengdan and Chen, Peixian and Li, Yanwei and Lin, Shaohui and Zhao, Sirui and Li, Ke and Xu, Tong and Zheng, Xiawu and Chen, Enhong and Shan, Caifeng and He, Ran and Sun, Xing , booktitle=. 2025 , doi=
2025
-
[17]
Shinn, Noah and Cassano, Federico and Gopinath, Ashwin and Narasimhan, Karthik and Yao, Shunyu , booktitle=
-
[18]
2024 , doi=
Zhao, Andrew and Huang, Daniel and Xu, Quentin and Lin, Matthieu and Liu, Yong-Jin and Huang, Gao , booktitle=. 2024 , doi=
2024
-
[19]
2025 , publisher=
Wang, Zora Zhiruo and Mao, Jiayuan and Fried, Daniel and Neubig, Graham , booktitle=. 2025 , publisher=
2025
-
[20]
2026 , doi=
Cao, Zouying and Deng, Jiaji and Yu, Li and Zhou, Weikang and Liu, Zhaoyang and Ding, Bolin and Zhao, Hai , booktitle=. 2026 , doi=
2026
-
[21]
Ma, Ziyu and Gou, Chenhui and Shi, Hengcan and Sun, Bin and Li, Shutao and Rezatofighi, Hamid and Cai, Jianfei , booktitle=
-
[22]
Kim, Junho and Kim, Hyunjun and Lee, Hosu and Ro, Yong Man , booktitle=
-
[23]
Diko, Anxhelo and Wang, Tinghuai and Swaileh, Wassim and Sun, Shiyan and Patras, Ioannis , booktitle=
-
[24]
Zhang, Yanzhao and Li, Mingxin and Long, Dingkun and Zhang, Xin and Lin, Huan and Yang, Baosong and Xie, Pengjun and Yang, An and Liu, Dayiheng and Lin, Junyang and Huang, Fei and Zhou, Jingren , journal=
-
[25]
Proceedings of the 40th International Conference on Machine Learning , series =
Large Language Models Can Be Easily Distracted by Irrelevant Context , author =. Proceedings of the 40th International Conference on Machine Learning , series =. 2023 , publisher =
2023
-
[26]
Findings of the Association for Computational Linguistics: NAACL 2024 , pages =
Why So Gullible? Enhancing the Robustness of Retrieval-Augmented Models against Counterfactual Noise , author =. Findings of the Association for Computational Linguistics: NAACL 2024 , pages =. 2024 , publisher =
2024
-
[27]
Transactions of the Association for Computational Linguistics , volume =
Lost in the Middle: How Language Models Use Long Contexts , author =. Transactions of the Association for Computational Linguistics , volume =. 2024 , publisher =
2024
-
[28]
2024 , url =
Hsieh, Cheng-Ping and Sun, Simeng and Kriman, Samuel and Acharya, Shantanu and Rekesh, Dima and Jia, Fei and Ginsburg, Boris , booktitle =. 2024 , url =
2024
-
[29]
2024 , doi =
Wu, Haoning and Li, Dongxu and Chen, Bei and Li, Junnan , booktitle =. 2024 , doi =
2024
-
[31]
Communication, Simulation, and Intelligent Agents: Implications of Personal Intelligent Machines for Medical Education
Clancey, William J. Communication, Simulation, and Intelligent Agents: Implications of Personal Intelligent Machines for Medical Education. Proceedings of the Eighth International Joint Conference on Artificial Intelligence (IJCAI-83)
-
[32]
Classification Problem Solving
Clancey, William J. Classification Problem Solving. Proceedings of the Fourth National Conference on Artificial Intelligence
-
[33]
, title =
Robinson, Arthur L. , title =. 1980 , doi =. https://science.sciencemag.org/content/208/4447/1019.full.pdf , journal =
1980
-
[34]
New Ways to Make Microcircuits Smaller---Duplicate Entry
Robinson, Arthur L. New Ways to Make Microcircuits Smaller---Duplicate Entry. Science
-
[35]
Clancey and Glenn Rennels , abstract =
Diane Warner Hasling and William J. Clancey and Glenn Rennels , abstract =. Strategic explanations for a diagnostic consultation system , journal =. 1984 , issn =. doi:https://doi.org/10.1016/S0020-7373(84)80003-6 , url =
-
[36]
and Rennels, Glenn R
Hasling, Diane Warner and Clancey, William J. and Rennels, Glenn R. and Test, Thomas. Strategic Explanations in Consultation---Duplicate. The International Journal of Man-Machine Studies
-
[37]
Poligon: A System for Parallel Problem Solving
Rice, James. Poligon: A System for Parallel Problem Solving
-
[38]
Transfer of Rule-Based Expertise through a Tutorial Dialogue
Clancey, William J. Transfer of Rule-Based Expertise through a Tutorial Dialogue
-
[39]
The Engineering of Qualitative Models
Clancey, William J. The Engineering of Qualitative Models
-
[40]
2023 , eprint=
Attention Is All You Need , author=. 2023 , eprint=
2023
-
[41]
Pluto: The 'Other' Red Planet
NASA. Pluto: The 'Other' Red Planet
-
[42]
Bai, S.; Chen, K.; Liu, X.; Wang, J.; Ge, W.; Song, S.; Dang, K.; Wang, P.; Wang, S.; Tang, J.; Zhong, H.; Zhu, Y.; Yang, M.; Li, Z.; Wan, J.; Wang, P.; Ding, W.; Fu, Z.; Xu, Y.; Ye, J.; Zhang, X.; Xie, T.; Cheng, Z.; Zhang, H.; Yang, Z.; Xu, H.; and Lin, J. 2025. Qwen2.5-VL Technical Report . arXiv preprint arXiv:2502.13923
Pith/arXiv arXiv 2025
-
[43]
Cao, Z.; Deng, J.; Yu, L.; Zhou, W.; Liu, Z.; Ding, B.; and Zhao, H. 2026. Remember Me, Refine Me: A Dynamic Procedural Memory Framework for Experience-Driven Agent Evolution . In Findings of the Association for Computational Linguistics: ACL 2026, 16803--16822
2026
-
[44]
Fu, C.; Dai, Y.; Luo, Y.; Li, L.; Ren, S.; Zhang, R.; Wang, Z.; Zhou, C.; Shen, Y.; Zhang, M.; Chen, P.; Li, Y.; Lin, S.; Zhao, S.; Li, K.; Xu, T.; Zheng, X.; Chen, E.; Shan, C.; He, R.; and Sun, X. 2025. Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-Modal LLMs in Video Analysis . In Proceedings of the IEEE/CVF Conference on Comput...
2025
-
[45]
K.; Jia, M.; Cao, X.; Shah, A.; Shrivastava, A.; and Lim, S.-N
He, B.; Li, H.; Jang, Y. K.; Jia, M.; Cao, X.; Shah, A.; Shrivastava, A.; and Lim, S.-N. 2024. MA-LMM: Memory-Augmented Large Multimodal Model for Long-Term Video Understanding . In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 13504--13514
2024
-
[46]
Hong, G.; Kim, J.; Kang, J.; Myaeng, S.-H.; and Whang, J. J. 2024. Why So Gullible? Enhancing the Robustness of Retrieval-Augmented Models against Counterfactual Noise. In Findings of the Association for Computational Linguistics: NAACL 2024, 2474--2495. Association for Computational Linguistics
2024
-
[47]
Hsieh, C.-P.; Sun, S.; Kriman, S.; Acharya, S.; Rekesh, D.; Jia, F.; and Ginsburg, B. 2024. RULER : What's the Real Context Size of Your Long-Context Language Models? In First Conference on Language Modeling
2024
-
[48]
Huang, Z.; Li, X.; Li, J.; Wang, J.; Zeng, X.; Liang, C.; Wu, T.; Chen, X.; Li, L.; and Wang, L. 2025. Online Video Understanding: OVBench and VideoChat-Online . In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 3328--3338
2025
-
[49]
Jiang, R.; Wang, J.; Zhao, G.; Luo, C.; Wang, K.; and Zhang, W. 2026. Advancing Multimodal Agent Reasoning with Long-Term Neuro-Symbolic Memory . arXiv preprint arXiv:2603.15280
arXiv 2026
-
[50]
F.; Lin, K.; Hewitt, J.; Paranjape, A.; Bevilacqua, M.; Petroni, F.; and Liang, P
Liu, N. F.; Lin, K.; Hewitt, J.; Paranjape, A.; Bevilacqua, M.; Petroni, F.; and Liang, P. 2024. Lost in the Middle: How Language Models Use Long Contexts. Transactions of the Association for Computational Linguistics, 12: 157--173
2024
-
[51]
Long, L.; He, Y.; Ye, W.; Pan, Y.; Lin, Y.; Li, H.; Zhao, J.; and Li, W. 2026. Seeing, Listening, Remembering, and Reasoning: A Multimodal Agent with Long-Term Memory . In International Conference on Learning Representations
2026
-
[52]
OpenAI . 2024. GPT-4o System Card . Technical report
2024
-
[53]
Reid, M.; Savinov, N.; Teplyashin, D.; Lepikhin, D.; Lillicrap, T.; Alayrac, J.-B.; Soricut, R.; Lazaridou, A.; Firat, O.; Schrittwieser, J.; et al. 2024. Gemini 1.5: Unlocking Multimodal Understanding Across Millions of Tokens of Context . arXiv preprint arXiv:2403.05530
Pith/arXiv arXiv 2024
-
[54]
H.; Sch \"a rli, N.; and Zhou, D
Shi, F.; Chen, X.; Misra, K.; Scales, N.; Dohan, D.; Chi, E. H.; Sch \"a rli, N.; and Zhou, D. 2023. Large Language Models Can Be Easily Distracted by Irrelevant Context. In Proceedings of the 40th International Conference on Machine Learning, volume 202 of Proceedings of Machine Learning Research, 31210--31227. PMLR
2023
-
[55]
Shinn, N.; Cassano, F.; Gopinath, A.; Narasimhan, K.; and Yao, S. 2023. Reflexion: Language Agents with Verbal Reinforcement Learning . In Advances in Neural Information Processing Systems, volume 36
2023
-
[56]
Song, E.; Chai, W.; Wang, G.; Zhang, Y.; Zhou, H.; Wu, F.; Chi, H.; Guo, X.; Ye, T.; Zhang, Y.; Lu, Y.; Hwang, J.-N.; and Wang, G. 2024. MovieChat: From Dense Token to Sparse Memory for Long Video Understanding . In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 18221--18232
2024
-
[57]
Wang, J.; Sun, T.; Zhu, J.; Li, J.; Xu, H.; Wen, Z.; Hu, X.; Li, Z.; and Zhang, L. 2026. StreamMeCo: Long-Term Agent Memory Compression for Efficient Streaming Video Understanding . In Findings of the Association for Computational Linguistics: ACL 2026, 13234--13251
2026
-
[58]
Wang, X.; Zhang, Y.; Zohar, O.; and Yeung-Levy, S. 2024. VideoAgent: Long-Form Video Understanding with Large Language Model as Agent . In Proceedings of the European Conference on Computer Vision
2024
-
[59]
Z.; Mao, J.; Fried, D.; and Neubig, G
Wang, Z. Z.; Mao, J.; Fried, D.; and Neubig, G. 2025. Agent Workflow Memory . In Proceedings of the 42nd International Conference on Machine Learning, volume 267 of Proceedings of Machine Learning Research, 63897--63911. PMLR
2025
-
[60]
H.; Wang, C.; Chen, S.; Pereira, F.; Kang, W.-C.; and Cheng, D
Wei, T.; Sachdeva, N.; Coleman, B.; He, Z.; Bei, Y.; Ning, X.; Ai, M.; Li, Y.; He, J.; Chi, E. H.; Wang, C.; Chen, S.; Pereira, F.; Kang, W.-C.; and Cheng, D. Z. 2025. Evo-Memory: Benchmarking LLM Agent Test-time Learning with Self-Evolving Memory. arXiv preprint arXiv:2511.20857
Pith/arXiv arXiv 2025
-
[61]
Wu, H.; Li, D.; Chen, B.; and Li, J. 2024. LongVideoBench : A Benchmark for Long-Context Interleaved Video-Language Understanding. In Advances in Neural Information Processing Systems, volume 37
2024
-
[62]
Xu, J.; Guo, Z.; He, J.; Hu, H.; He, T.; Bai, S.; Chen, K.; Wang, J.; Fan, Y.; Dang, K.; Zhang, B.; Wang, X.; Chu, Y.; and Lin, J. 2025. Qwen2.5-Omni Technical Report . arXiv preprint arXiv:2503.20215
Pith/arXiv arXiv 2025
-
[63]
Yao, L.; Li, Y.; Wei, Y.; Li, L.; Ren, S.; Liu, Y.; Ouyang, K.; Wang, L.; Li, S.; Li, S.; Kong, L.; Liu, Q.; Zhang, Y.; and Sun, X. 2025. TimeChat-Online: 80\ In Proceedings of the 33rd ACM International Conference on Multimedia, 10807--10816
2025
-
[64]
Yeo, W.; Kim, K.; Yoon, J.; and Hwang, S. J. 2026. WorldMM: Dynamic Multimodal Memory Agent for Long Video Reasoning . In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 25599--25609
2026
-
[65]
Zeng, X.; Qiu, K.; Zhang, Q.; Li, X.; Wang, J.; Li, J.; Yan, Z.; Tian, K.; Tian, M.; Zhao, X.; Wang, Y.; and Wang, L. 2025. StreamForest: Efficient Online Video Understanding with Persistent Event Memory . In Advances in Neural Information Processing Systems, volume 38
2025
-
[66]
Zhang, H.; Wang, Y.; Tang, Y.; Liu, Y.; Feng, J.; and Jin, X. 2025. Flash-VStream: Efficient Real-Time Understanding for Long Video Streams . In Proceedings of the IEEE/CVF International Conference on Computer Vision, 21059--21069
2025
-
[67]
Zhao, A.; Huang, D.; Xu, Q.; Lin, M.; Liu, Y.-J.; and Huang, G. 2024. ExpeL: LLM Agents Are Experiential Learners . In Proceedings of the AAAI Conference on Artificial Intelligence, volume 38, 19632--19642
2024
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.