REVIEW 3 major objections 5 minor 48 references
MIND cuts memory-injection attack success on LLM agents by about half by keeping only the intent–behavior signal and discarding trajectory clutter.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-31 17:44 UTC pith:7TUIAAJI
load-bearing objection Solid agent-memory defense paper: real ASR cuts at near-zero latency tax, but the detector is supervised on known poisons and only two attack families. the 3 major comments →
MIND: Lightweight and Effective Memory Injection Defense for LLM Agents via Intent-Aware Information Bottleneck
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
Benign and poisoned agent trajectories are distinguishable by the relationship between the initial user intent and later turn behavior; an intent-aware information bottleneck can extract compact representations of that relationship, filter out task-irrelevant and repetitive multi-turn content, and let a small classifier reject poisoned memories while preserving task performance and inference speed.
What carries the argument
Intent-aware Information Bottleneck (IB): a variational compressor on the concatenated initial-intent and turn hidden states that minimizes mutual information with the raw input while maximizing (via margin alignment) information about the benign/poisoned label relative to the intent anchor, feeding a multi-hyperplane decision region.
Load-bearing premise
Last-token hidden states of the pair [initial user intent, current turn]—even when taken from a frozen proxy model for closed-source agents—carry attack-discriminative geometry that a supervised bottleneck trained on labeled trajectories will still see on new domains and new attacks.
What would settle it
Train MIND only on the paper’s QA/EHR trajectories, then measure ASR-r, ASR-a, and accuracy on ReAct-StrategyQA and MMLU under AgentPoison and MINJA; if mean ASR-r/ASR-a do not fall by roughly half versus no defense while accuracy and latency stay comparable to the undefended agent, the central claim fails.
If this is right
- Memory defense for long-horizon agents can be cast as intent-conditioned denoising rather than full-trajectory encoding or per-turn LLM auditing.
- On ReAct-StrategyQA, mean ASR-r and ASR-a drop by about 55% relative to no defense across four backbones while average accuracy and latency match the undefended agent.
- Defense latency stays near the undefended baseline and is substantially lower than LLM-auditor and A-MemGuard-style methods.
- Cross-domain transfer from QA/EHR training trajectories to StrategyQA and MMLU is feasible for this detector under the tested attacks.
- Write-time and retrieve-time filtering can both use the same compact intent–behavior score without re-invoking a large judge model.
Where Pith is reading between the lines
- If intent–turn geometry is the real signal, similar bottlenecks might flag other gradual steering attacks (prompt drift, tool-policy hijacks) without storing full histories.
- Proxy hidden states from an open 8B model working for closed-source agents suggests defense can sit outside the agent API, which matters for production stacks that cannot expose internals.
- When benign and poisoned classes become separable only after IB compression, representation-level monitoring may outperform content-only filters on attacks that look locally harmless.
- A natural next measurement is whether the same IB detector holds when the attacker optimizes directly against the bottleneck latents rather than against the agent’s answers.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes MIND, a lightweight defense against memory-injection attacks on memory-augmented LLM agents. Motivated by preliminary evidence that benign and poisoned multi-turn trajectories differ in how turn representations relate to the initial user intent (§4.1, Fig. 2), MIND extracts last-token hidden states of the intent–turn pair, compresses them with a variational intent-aware Information Bottleneck (Eqs. 4–9), and classifies the denoised latents with a multi-hyperplane (convex polytope) decision head (Eqs. 10–11). Training uses supervised labels that mark whether the retrieved set intersects a known poison set M_adv (Eq. 3). Across four agent backbones on ReAct-StrategyQA and MMLU, and against AgentPoison and MINJA, MIND reports large average reductions in ASR-r/ASR-a relative to No Defense while keeping task accuracy and episode latency near the undefended baseline and below LLM-auditing defenses (Table 1). Ablations (Table 2), hyperparameter sweeps (Table 3), and staged t-SNE (Fig. 4) support the contribution of the IB and multi-hyperplane components.
Significance. Memory injection is a timely and practically important threat for long-horizon LLM agents. If the reported security–utility trade-off holds under realistic defender knowledge, MIND would be a useful contribution: it targets cross-turn intent–behavior deviation rather than single-record content, avoids repeated LLM auditing, and demonstrates competitive accuracy and latency on multiple backbones including closed-source models. Strengths include a clear problem formulation (§3), multi-backbone evaluation with standard baselines (Table 1), component ablations, and efficiency comparisons against A-MemGuard and LLM Auditor. The work is incremental relative to existing trajectory monitors and record-level filters, but the intent-anchored IB framing and polytope head are a coherent lightweight design that the community can build on.
major comments (3)
- [§3.2, §4.2 Eq. (3), Table 1] §3.2 states that the defender knows neither M_adv nor the attack strategy, yet §4.2 Eq. (3) defines training labels li from exact intersection of the retrieved set with a known M_adv, and D is built from QA/EHR trajectories under author-controlled poisons. This is ordinary supervised learning, but it is a stronger training assumption than the stated threat model. The manuscript should explicitly separate (i) offline supervised training with labeled poisoned trajectories from (ii) attack-agnostic deployment, and should report at least one mismatched-attack or held-out-attack protocol (train on one of AgentPoison/MINJA, test on the other; or train without the test trigger/target family). Without that, the headline mean 55.4%/55.3% ASR-r/ASR-a cuts on ReAct-StrategyQA (abstract, Table 1) may overstate robustness under the paper’s own defender knowledge.
- [§5.1, §4.3 Eqs. (10)–(11)] Evaluation covers only AgentPoison and MINJA (§5.1). There is no adaptive attacker who optimizes injected records so that [h0; ht] (or the Llama-3.1-8B proxy states used for closed-source agents) remain inside the learned benign polytope of Eqs. (10)–(11). Given that the decision boundary is a low-capacity multi-hyperplane classifier on a fixed frozen encoder, adaptive evasion is a load-bearing risk for the central security claim. A minimal adaptive or transfer attack study—or a clear limitation statement with a concrete negative result—belongs in the main evaluation, not only as future work.
- [§4.2 Eqs. (4), (9); Table 2] The IB ‘informativeness’ term max I(Z;Y) is replaced by a label-conditioned margin alignment of μ(x) to μ(h0) (Eq. 9), not by a standard supervised IB decoder p(y|z). That surrogate is reasonable but is an ad-hoc axiom: it assumes that intent–behavior geometric deviation is the right Y-signal for poison detection. The paper should either (a) justify why this surrogate tracks I(Z;Y) more tightly than a simple classification head on z, or (b) ablate against a plain VIB+classifier without L_align (beyond the coarser ‘w/o IB’ row in Table 2). Table 2’s w/o IB and w/o Multi-hyperplane variants still leave ASR-a near 51%, so the quantitative necessity of the intent-alignment surrogate for the claimed gains needs a sharper control.
minor comments (5)
- [Table 1] Table 1: AV Filter is dashed for closed-source backbones (expected) but also shows weak open-source numbers; a one-sentence note on why attention-variance filtering fails under these multi-turn memory attacks would help readers.
- [§5.1, Table 3] §5.1 Implementation: latent dimension d, K, margins m and m_a, and encoder depth are free parameters (Eqs. 8–12) but only β and λ are swept in Table 3. Report default K, d, and m in the main text or a compact appendix table.
- [Figure 4] Fig. 1 and Fig. 3 captions are helpful; ensure Fig. 4’s three panels share axis scales or state that scales differ, so ‘enlarged margin’ is not a plotting artifact.
- [Abstract, §1] Minor prose issues: spacing typos such as ‘memoryinjectionattacks’, ‘frominformationredundancy’, and ‘ReAct-StrategyQA, MIND reduces mean ASR-r’ appear in the abstract/intro PDF text; clean these in production.
- [§2] Related Work should briefly position MIND against concurrent trajectory guardrails (e.g., SafeHarbor-style hierarchical memory guards cited as Liu et al. 2026) so novelty of the intent-aware IB is sharper.
Circularity Check
No circular derivation: MIND is standard supervised VIB + empirical evaluation; ASR cuts are measured, not forced by construction.
full rationale
The paper is an empirical systems/defense paper, not a first-principles derivation. The load-bearing chain is: (i) preliminary t-SNE/attention observations motivate separability of intent–turn geometry under poison (§4.1); (ii) a standard variational IB + margin alignment + multi-hyperplane hinge objective is trained with explicit poison labels li from known M_adv on QA/EHR trajectories (Eqs. 3–12); (iii) the trained filter is evaluated on held-out StrategyQA/MMLU under AgentPoison and MINJA, reporting measured ACC/ASR/latency (Table 1). Nothing in that chain equates the reported mean 55.4%/55.3% ASR reductions to a fitted constant or to a quantity defined in terms of itself. Training labels that require knowing M_adv at train time are ordinary supervised learning and a threat-model/generalization concern, not algebraic circularity. Citations for IB (Tishby; Alemi), contrastive margins (Hadsell), and polytope machines (Kantchelian) are external and not author-uniqueness theorems. No self-citation is load-bearing for the central claim. No ansatz is smuggled in as a uniqueness result. Score 0; steps empty.
Axiom & Free-Parameter Ledger
free parameters (5)
- β (IB / KL weight in Eq. 12) =
1e-3 (default)
- λ (intent-alignment weight in Eq. 12) =
0.2 (default)
- alignment margin m_a (Eq. 9)
- classifier margin m and K hyperplanes (Eq. 10–11)
- latent dimension d and encoder architecture
axioms (5)
- standard math Variational IB upper bound with diagonal Gaussian posterior and standard normal prior is a valid training surrogate for min I(Z;X) (Eq. 5–8).
- ad hoc to paper Maximizing a supervised margin alignment of μ(x) to μ(h0) is a valid surrogate for max I(Z;Y) (Eq. 9).
- domain assumption Attacker can inject at most Δ records that are retrieved and drive the agent to l_adv (threat model Eq. 1); defender only filters R and W.
- domain assumption Last-token hidden state of <observation> (and Llama-3.1-8B proxy for closed agents) is a sufficient turn representation for detection.
- standard math Multi-hyperplane (convex polytope) boundaries are expressive enough for diverse poison patterns (Kantchelian et al. 2014).
invented entities (2)
-
Intent–behavior latent z via intent-aware IB encoder E
no independent evidence
-
MIND multi-hyperplane memory filter g_φ
no independent evidence
read the original abstract
Memory-augmented LLM-based agents are vulnerable to memory injection attacks: Agents may retrieve poisoned memory from attackers, which diverts their behavior from initial user intent and finally causes task failure. However, existing defense mechanisms either incur high computational cost or suffer from information redundancy in multi-turn contexts. To address these challenges, we propose Memory Intent-Aware Neural Denoising(MIND), a lightweight defense framework for memory injection attack. Our preliminary analysis reveals that benign and poisoned trajectories exhibit distinguishable relationships between the initial user intent and subsequent behavior. Building on this observation, MIND employs an intent-aware Information Bottleneck(IB) to extract compact intent--behavior representations from the initial intent and turn-level behavior. The IB preserves intent-relevant cross-turn attack signals while filtering task-irrelevant and repetitive information, and a lightweight detector identifies malicious memories from the resulting representations. As such, MIND mitigates information redundancy in multi-turn contexts while avoiding the overhead of repeated LLM auditing. Extensive experiments show that MIND reduces attack success rates while preserving task accuracy and inference efficiency. Notably, on ReAct-StrategyQA, MIND reduces mean ASR-r and ASR-a by 55.4% and 55.3%, respectively, while matching the undefended agent in average accuracy and latency.
Figures
Reference graph
Works this paper leans on
-
[1]
O'Brien and Carrie J
Joon Sung Park and Joseph C. O'Brien and Carrie J. Cai and Meredith Ringel Morris and Percy Liang and Michael S. Bernstein , title =. Proceedings of the 36th Annual ACM Symposium on User Interface Software and Technology , year =
-
[2]
Patil and Ion Stoica and Joseph E
Charles Packer and Sarah Wooders and Kevin Lin and Vivian Fang and Shishir G. Patil and Ion Stoica and Joseph E. Gonzalez , title =. 2310.08560 , archivePrefix =
-
[3]
Proceedings of the AAAI Conference on Artificial Intelligence , volume =
Wanjun Zhong and Lianghong Guo and Qiqi Gao and He Ye and Yanlin Wang , title =. Proceedings of the AAAI Conference on Artificial Intelligence , volume =
-
[4]
Advances in Neural Information Processing Systems , volume =
Noah Shinn and Federico Cassano and Ashwin Gopinath and Karthik Narasimhan and Shunyu Yao , title =. Advances in Neural Information Processing Systems , volume =
-
[5]
Proceedings of the AAAI Conference on Artificial Intelligence , volume =
Andrew Zhao and Daniel Huang and Quentin Xu and Matthieu Lin and Yong-Jin Liu and Gao Huang , title =. Proceedings of the AAAI Conference on Artificial Intelligence , volume =
-
[6]
Guanzhi Wang and Yuqi Xie and Yunfan Jiang and Ajay Mandlekar and Chaowei Xiao and Yuke Zhu and Linxi Fan and Anima Anandkumar , title =. 2305.16291 , archivePrefix =
-
[7]
Zora Zhiruo Wang and Jiayuan Mao and Daniel Fried and Graham Neubig , title =. 2409.07429 , archivePrefix =
-
[8]
Wujiang Xu and Zujie Liang and Kai Mei and Hang Gao and Juntao Tan and Yongfeng Zhang , title =. 2502.12110 , archivePrefix =
-
[9]
Prateek Chhikara and Dev Khant and Saket Aryan and Taranjeet Singh and Deshraj Yadav , title =. 2504.19413 , archivePrefix =
-
[10]
Retrieval-Augmented Generation for Knowledge-Intensive
Patrick Lewis and Ethan Perez and Aleksandra Piktus and Fabio Petroni and Vladimir Karpukhin and Naman Goyal and Heinrich K. Retrieval-Augmented Generation for Knowledge-Intensive. Advances in Neural Information Processing Systems , volume =
-
[11]
International Conference on Machine Learning , year =
Kelvin Guu and Kenton Lee and Zora Tung and Panupong Pasupat and Ming-Wei Chang , title =. International Conference on Machine Learning , year =
-
[12]
Transactions of the Association for Computational Linguistics , volume =
Mor Geva and Daniel Khashabi and Elad Segal and Tushar Khot and Dan Roth and Jonathan Berant , title =. Transactions of the Association for Computational Linguistics , volume =
-
[13]
Advances in Neural Information Processing Systems , volume =
Zhaorun Chen and Zhen Xiang and Chaowei Xiao and Dawn Song and Bo Li , title =. Advances in Neural Information Processing Systems , volume =
-
[14]
34th USENIX Security Symposium (USENIX Security 25) , year =
Wei Zou and Runpeng Geng and Binghui Wang and Jinyuan Jia , title =. 34th USENIX Security Symposium (USENIX Security 25) , year =
-
[15]
Xiang , title =
Shen Dong and Shaochen Xu and Pengfei He and Yige Li and Jiliang Tang and Tianming Liu and Hui Liu and Zhen J. Xiang , title =. Advances in Neural Information Processing Systems , volume =
-
[16]
Zhenlin Xu and Xiaogang Zhu and Yu Yao and Minhui Xue and Yiliao Song , title =. 2603.15125 , archivePrefix =
-
[17]
The Thirteenth International Conference on Learning Representations , year =
Hanrong Zhang and Jingyuan Huang and Kai Mei and Yifei Yao and Zhenting Wang and Chenlu Zhan and Hongwei Wang and Yongfeng Zhang , title =. The Thirteenth International Conference on Learning Representations , year =
-
[18]
Qianshan Wei and Tengchao Yang and Yaochen Wang and Xinfeng Li and Lijun Li and Zhenfei Yin and Yi Zhan and Thorsten Holz and Zhiqiang Lin and XiaoFeng Wang , title =. 2510.02373 , archivePrefix =
-
[19]
Chong Xiang and Tong Wu and Zexuan Zhong and David Wagner and Danqi Chen and Prateek Mittal , title =. 2405.15556 , archivePrefix =
-
[20]
Hakan Inan and Kartikeya Upasani and Jianfeng Chi and Rashi Rungta and Krithika Iyer and Yuning Mao and Michael Tontchev and Qing Hu and Brian Fuller and Davide Testuggine and Madian Khabsa , title =. 2312.06674 , archivePrefix =
-
[21]
Victor Sanh and Lysandre Debut and Julien Chaumond and Thomas Wolf , title =. 1910.01108 , archivePrefix =
Pith/arXiv arXiv 1910
-
[22]
Gabriel Alon and Michael Kamfonas , title =. 2308.14132 , archivePrefix =
-
[23]
Ciyan Ouyang and Rui Hou , title =. 2605.14421 , archivePrefix =
-
[24]
Philippe Laban and Hiroaki Hayashi and Yingbo Zhou and Jennifer Neville , title =. 2505.06120 , archivePrefix =
-
[25]
Zhiheng Xi and Wenxiang Chen and Xin Guo and Wei He and Yiwen Ding and Boyang Hong and Ming Zhang and Junzhe Wang and Senjie Jin and Enyu Zhou and Rui Zheng and Xiaoran Fan and Xiao Wang and Limao Xiong and Yuhao Zhou and Weiran Wang and Changhao Jiang and Yicheng Zou and Xiangyang Liu and Zhangyue Yin and Shihan Dou and Rongxiang Weng and Wensen Cheng an...
-
[26]
Jimenez and Alexander Wettig and Kilian Lieret and Shunyu Yao and Karthik Narasimhan and Ofir Press , title =
John Yang and Carlos E. Jimenez and Alexander Wettig and Kilian Lieret and Shunyu Yao and Karthik Narasimhan and Ofir Press , title =. Advances in Neural Information Processing Systems , volume =
-
[27]
Chris Lu and Cong Lu and Robert Tjarko Lange and Jakob Foerster and Jeff Clune and David Ha , title =. 2408.06292 , archivePrefix =
-
[28]
Long Ouyang and Jeff Wu and Xu Jiang and Diogo Almeida and Carroll L. Wainwright and Pamela Mishkin and Chong Zhang and Sandhini Agarwal and Katarina Slama and Alex Ray and John Schulman and Jacob Hilton and Fraser Kelton and Luke Miller and Maddie Simens and Amanda Askell and Peter Welinder and Paul Christiano and Jan Leike and Ryan Lowe , title =. Advan...
-
[29]
Yueh-Han Chen and Nitish Joshi and Yulin Chen and Maksym Andriushchenko and Rico Angell and He He , title =. 2506.10949 , archivePrefix =
-
[30]
Sarthak Choudhary and Nils Palumbo and Ashish Hooda and Krishnamurthy Dj Dvijotham and Somesh Jha , title =. 2506.04390 , archivePrefix =
-
[31]
2019 , eprint=
Information bottleneck through variational glasses , author=. 2019 , eprint=
2019
-
[32]
2024 , eprint=
A Survey on Universal Approximation Theorems , author=. 2024 , eprint=
2024
-
[33]
and Joseph, Anthony D
Kantchelian, Alex and Tschantz, Michael Carl and Huang, Ling and Bartlett, Peter L. and Joseph, Anthony D. and Tygar, J. D. , booktitle =. Large-Margin Convex Polytope Machine , volume =
-
[34]
Proceedings of the 47th International
Shitao Xiao and Zheng Liu and Peitian Zhang and Niklas Muennighoff and Defu Lian and Jian-Yun Nie , title =. Proceedings of the 47th International. doi:10.1145/3626772.3657878 , url =
-
[35]
Deep Research: A Survey of Autonomous Research Agents , author=. 2025 , journal=. 2508.12752 , archivePrefix=
Pith/arXiv arXiv 2025
-
[36]
2025 , url=
Tinghao Xie and Xiangyu Qi and Yi Zeng and Yangsibo Huang and Udari Madhushani Sehwag and Kaixuan Huang and Luxi He and Boyi Wei and Dacheng Li and Ying Sheng and Ruoxi Jia and Bo Li and Kai Li and Danqi Chen and Peter Henderson and Prateek Mittal , booktitle=. 2025 , url=
2025
-
[37]
2026 , eprint=
SafeHarbor: Defining Precise Decision Boundaries via Hierarchical Memory-Augmented Guardrail for LLM Agent Safety , author=. 2026 , eprint=
2026
-
[38]
The Ninth International Conference on Learning Representations , year=
Measuring Massive Multitask Language Understanding , author=. The Ninth International Conference on Learning Representations , year=
-
[39]
Chi and Nathanael Scharli and Denny Zhou , title=
Freda Shi and Xinyun Chen and Kanishka Misra and Nathan Scales and David Dohan and Ed H. Chi and Nathanael Scharli and Denny Zhou , title=. Proceedings of the 40th International Conference on Machine Learning , volume=. 2023 , url=
2023
-
[40]
Liu and Kevin Lin and John Hewitt and Ashwin Paranjape and Michele Bevilacqua and Fabio Petroni and Percy Liang , title=
Nelson F. Liu and Kevin Lin and John Hewitt and Ashwin Paranjape and Michele Bevilacqua and Fabio Petroni and Percy Liang , title=. Transactions of the Association for Computational Linguistics , volume=. 2024 , doi=
2024
-
[41]
The Eleventh International Conference on Learning Representations , year=
Shunyu Yao and Jeffrey Zhao and Dian Yu and Nan Du and Izhak Shafran and Karthik Narasimhan and Yuan Cao , title=. The Eleventh International Conference on Learning Representations , year=
-
[42]
Pereira and William Bialek , title=
Naftali Tishby and Fernando C. Pereira and William Bialek , title=. Proceedings of the 37th Annual Allerton Conference on Communication, Control, and Computing , pages=. 1999 , url=
1999
-
[43]
Alemi and Ian Fischer and Joshua V
Alexander A. Alemi and Ian Fischer and Joshua V. Dillon and Kevin Murphy , title=. The Fifth International Conference on Learning Representations , year=
-
[44]
Kingma and Max Welling , title=
Diederik P. Kingma and Max Welling , title=. The Second International Conference on Learning Representations , year=
-
[45]
2006 IEEE Computer Society Conference on Computer Vision and Pattern Recognition , volume=
Raia Hadsell and Sumit Chopra and Yann LeCun , title=. 2006 IEEE Computer Society Conference on Computer Vision and Pattern Recognition , volume=. 2006 , doi=
2006
-
[46]
Journal of Machine Learning Research , volume=
Laurens van der Maaten and Geoffrey Hinton , title=. Journal of Machine Learning Research , volume=. 2008 , url=
2008
-
[47]
2024 , eprint=
Abhimanyu Dubey and others , title=. 2024 , eprint=
2024
-
[48]
2025 , eprint=
An Yang and others , title=. 2025 , eprint=
2025
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.