REVIEW 3 minor 132 references
Security and Privacy in Retrieval-Augmented Generation: Architectures, Threats, Defenses, and Future Directions for Building Trustworthy Systems
T0 review · 0 major / 3 minor · reviewed 2026-06-25 · grok-4.3
Pith's one-line read Retrieval-augmented generation systems introduce security and privacy risks through their retrieval mechanisms that go beyond those of standard language models.
desk verdict A competent survey that organizes RAG security threats and defenses but adds no new techniques or findings. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
A unified taxonomy of threat surfaces spanning the retrieval, context construction, and generation stages.
What would settle it
Identification of a major privacy or security vulnerability in an operational RAG system whose attack surface falls outside the taxonomy categories of retrieval, context construction, and generation.
Extended reading notes
Core claim
Integrating retrieval pipelines in RAG systems exposes sensitive information through retrieval indices, query logs, context construction, or federated updates, while adversarial manipulation of knowledge bases can undermine trust in generated outputs, requiring a unified taxonomy of threat surfaces across retrieval, context construction, and generation stages together with analysis of attacks and defenses in centralized, on-device, federated, and hybrid paradigms.
Load-bearing premise
The presented unified taxonomy of threat surfaces and listed attack classes adequately covers the primary risks across all RAG paradigms.
Editorial extensions
If this is right
- Defenses must address retrieval-specific surfaces such as index poisoning and query log leakage in addition to generation-stage threats.
- Architectural choices for RAG deployments carry measurable privacy-utility trade-offs that vary by paradigm.
- Attacks including membership inference and collusion can compromise both data confidentiality and output integrity.
- Hybrid and federated RAG setups require additional protections against gradient leakage and inter-party collusion.
Reading between the lines
- Standard LLM safety layers will need explicit extensions to cover the retrieval stage if the taxonomy holds.
- Empirical measurement of real-world RAG data exposure events could test whether the listed attack classes are exhaustive.
- Cryptographic techniques applied at the index level might reduce leakage without full redesign of retrieval pipelines.
- The taxonomy could guide standardized auditing protocols for RAG systems in regulated domains.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript is a survey on security and privacy in Retrieval-Augmented Generation (RAG) systems. It examines risks across centralized, on-device (Micro-RAG), federated, and hybrid deployment paradigms; introduces a unified taxonomy of threat surfaces spanning retrieval, context construction, and generation stages; catalogs attack classes including membership inference, index inference, poisoning, gradient leakage, and collusion; reviews architectural, algorithmic, and cryptographic defenses with attention to privacy-utility trade-offs; and identifies open research challenges for trustworthy RAG systems.
Significance. If the coverage and taxonomy hold, the survey provides a useful organizing framework for an emerging area where RAG adoption introduces privacy and security risks beyond standard LLM threats. Systematizing attacks and defenses across multiple paradigms, along with explicit discussion of deployment considerations, can help researchers identify gaps and practitioners evaluate trade-offs. The work is a literature survey with no original theorems, experiments, or machine-checked proofs.
minor comments (3)
- Abstract: the claim of a 'unified taxonomy' would be strengthened by a brief statement of the criteria used to unify prior taxonomies or by a forward reference to the section where the unification is demonstrated.
- The manuscript would benefit from an explicit table or figure that maps each attack class to the deployment paradigms (centralized, on-device, federated, hybrid) in which it has been studied or is applicable.
- Section headings and subsection numbering should be checked for consistency with the taxonomy stages (retrieval, context construction, generation) to improve navigation.
Simulated Author's Rebuttal
We thank the referee for the positive summary and assessment of our survey, as well as the recommendation for minor revision. No specific major comments were provided in the report, so we have no individual points to address. We remain available to incorporate any minor editorial suggestions if requested.
Circularity Check
No significant circularity
full rationale
This is a literature survey paper with no original derivations, equations, predictions, or formal claims that could reduce to their own inputs. The unified taxonomy and attack/defense catalog are syntheses of external literature across RAG paradigms; the coverage statement is definitional to the survey genre rather than a testable proposition derived from fitted parameters or self-citations. No load-bearing steps match any enumerated circularity pattern.
Assumptions & free parameters
Cite this review
Pith. "Pith review of Security and Privacy in Retrieval-Augmented Generation: Architectures, Threats, Defenses, and Future Directions for Building Trustworthy Systems." pith.science (2026). https://pith.science/paper/SCWYTAMU
@misc{pith2026260625533,
author = {Pith},
title = {Pith review of: Security and Privacy in Retrieval-Augmented Generation: Architectures, Threats, Defenses, and Future Directions for Building Trustworthy Systems},
year = {2026},
howpublished = {\url{https://pith.science/paper/SCWYTAMU}},
note = {Machine review of arXiv:2606.25533}
}
read the original abstract
Retrieval-Augmented Generation (RAG) has emerged as a dominant paradigm for enhancing large language models with external knowledge. By coupling retrieval mechanisms with generative models, RAG systems improve factual grounding and adaptability across domains. However, integrating retrieval pipelines introduces new security and privacy risks that extend beyond conventional language modeling threats. Sensitive information may be exposed through retrieval indices, query logs, context construction, or federated updates, while adversarial manipulation of knowledge bases can undermine trust in generated outputs. This survey provides a comprehensive examination of privacy and security challenges across RAG systems deployed in centralized, on-device (Micro-RAG), federated, and hybrid paradigms. We present a unified taxonomy of threat surfaces spanning the retrieval, context construction, and generation stages and systematically analyze attack classes, including membership inference, index inference, poisoning, gradient leakage, and collusion. We further review architectural, algorithmic, and cryptographic defenses, highlighting privacy-utility trade-offs and deployment considerations. Finally, we outline open research challenges toward building trustworthy, secure, and resilient RAG systems for real-world applications.
Figures
Figures from the paper (6 more)
Reference graph
Works this paper leans on
-
[1]
A Comprehensive Overview of Large Language Models.ACM Transactions on Intelligent Systems and Technology, 16(5):1–72, 2025
Humza Naveed, Asad Ullah Khan, Shi Qiu, Muhammad Saqib, Saeed Anwar, Muhammad Usman, Naveed Akhtar, Nick Barnes, and Ajmal Mian. A Comprehensive Overview of Large Language Models.ACM Transactions on Intelligent Systems and Technology, 16(5):1–72, 2025
2025
-
[2]
Large language models: a survey of their development, capabilities, and applications.Knowledge and Information Systems, 67(3):2967–3022, 2025
Yadagiri Annepaka and Partha Pakray. Large language models: a survey of their development, capabilities, and applications.Knowledge and Information Systems, 67(3):2967–3022, 2025
2025
-
[3]
A review of prominent paradigms for LLM-based agents: Tool use, planning (including RAG), and feedback learning
Xinzhe Li. A review of prominent paradigms for LLM-based agents: Tool use, planning (including RAG), and feedback learning. In Owen Rambow, Leo Wanner, Marianna Apidianaki, Hend Al-Khalifa, Barbara Di Eugenio, and Steven Schockaert, editors,Proceedings of the 31st International Conference on Computational Linguistics, pages 9760–9779, Abu Dhabi, UAE, Janu...
2025
-
[4]
Know your RAG: Dataset taxonomy and generation strategies for evaluating RAG systems
Rafael Teixeira de Lima, Shubham Gupta, Cesar Berrospi Ramis, Lokesh Mishra, Michele Dolfi, Peter Staar, and Panagiotis Vagenas. Know your RAG: Dataset taxonomy and generation strategies for evaluating RAG systems. In Owen Rambow, Leo Wanner, Marianna Apidianaki, Hend Al-Khalifa, Barbara Di Eugenio, Steven Schockaert, Kareem Darwish, and Apoorv Agarwal, e...
2025
-
[5]
MeMemo: On-device Retrieval Augmentation for Private and Personalized Text Generation
Zijie J Wang and Duen Horng Chau. MeMemo: On-device Retrieval Augmentation for Private and Personalized Text Generation. InProceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval, pages 2765–2770, 2024
2024
-
[6]
Val Andrei Fajardo, David B Emerson, Amandeep Singh, Veronica Chatrath, Marcelo Lotif, Ravi Theja, Alex Cheung, and Izuki Matsuba. Fedrag: A framework for fine-tuning retrieval-augmented generation systems.arXiv preprint arXiv:2506.09200, 2025
-
[7]
A Survey on RAG Meeting LLMs: Towards Retrieval-Augmented Large Language Models
Wenqi Fan, Yujuan Ding, Liangbo Ning, Shijie Wang, Hengyun Li, Dawei Yin, Tat-Seng Chua, and Qing Li. A Survey on RAG Meeting LLMs: Towards Retrieval-Augmented Large Language Models. InProceedings of the 30th ACM SIGKDD conference on knowledge discovery and data mining, pages 6491–6501, 2024
2024
-
[8]
Pratik Sharma and Saishab Bhattarai. A Review on Retrieval-Augmented Generation: Architectures, Research Challenges, and Emerging Frontiers.Journal of Future Artificial Intelligence and Technologies, 2(4):616–628, 2026
2026
Show all 132 references
-
[9]
Yangning Li, Weizhi Zhang, Yuyao Yang, Wei-Chieh Huang, Yaozu Wu, Junyu Luo, Yuanchen Bei, Henry Peng Zou, Xiao Luo, Yusheng Zhao, Chunkit Chan, Yankai Chen, Zhongfen Deng, Yinghui Li, Hai-Tao Zheng, Dongyuan Li, Renhe Jiang, Ming Zhang, Yangqiu Song, and Philip S. Yu. A surve...
2025
-
[10]
CRAG - Comprehensive RAG Benchmark
Xiao Yang, Kai Sun, Hao Xin, Yushi Sun, Nikita Bhalla, Xiangsen Chen, Sajal Choudhary, Rongze Daniel Gui, Ziran Will Jiang, Ziyu Jiang, Lingkun Kong, Brian Moran, Jiaqi Wang, Yifan Ethan Xu, An Yan, Chenyu Yang, Eting Yuan, Hanwen Zha, Nan Tang, Lei Chen, Nicolas Scheffer, Yue...
2024
-
[11]
Communication- Efficient Learning of Deep Networks from Decentralized Data
Brendan McMahan, Eider Moore, Daniel Ramage, Seth Hampson, and Blaise Aguera y Arcas. Communication- Efficient Learning of Deep Networks from Decentralized Data. In Aarti Singh and Jerry Zhu, editors,Proceedings of the 20th International Conference on Artificial Intelligence a...
2017
-
[12]
Federated Learning: Challenges, Methods, and Future Directions.IEEE Signal Processing Magazine, 37(3):50–60, 2020
Tian Li, Anit Kumar Sahu, Ameet Talwalkar, and Virginia Smith. Federated Learning: Challenges, Methods, and Future Directions.IEEE Signal Processing Magazine, 37(3):50–60, 2020
2020
-
[13]
Nguyen, Ming Ding, Pubudu N
Dinh C. Nguyen, Ming Ding, Pubudu N. Pathirana, Aruna Seneviratne, Jun Li, and H. Vincent Poor. Federated Learning for Internet of Things: A Comprehensive Survey.IEEE Communications Surveys & Tutorials, 23(3):1622–1658, 2021
2021
-
[14]
Sybil-aware adaptive defence framework for robust federated learning.Pervasive and Mobile Computing, page 102157, 2025
Dnyanesh Khedekar, Tanmaya Mahapatra, and Amitesh Singh Rajput. Sybil-aware adaptive defence framework for robust federated learning.Pervasive and Mobile Computing, page 102157, 2025
2025
-
[15]
Efficient Federated Search for Retrieval-Augmented Generation
Rachid Guerraoui, Anne-Marie Kermarrec, Diana Petrescu, Rafael Pires, Mathis Randl, and Martijn de V os. Efficient Federated Search for Retrieval-Augmented Generation. InProceedings of the 5th Workshop on Machine Learning and Systems, pages 74–81, 2025. 22 APREPRINT- JUNE25, 2026
2025
-
[16]
Adversarial Attacks on Large Language Models: A Survey
Meera Al Kuwaiti and Heba Ismail. Adversarial Attacks on Large Language Models: A Survey. InProceedings of Eighth International Conference on Information System Design and Intelligent Applications, pages 529–547. Springer, 2025
2025
-
[17]
SoK: Taxonomy and Evaluation of Prompt Security in Large Language Models.arXiv preprint arXiv:2510.15476, 2025
Hanbin Hong, Shuya Feng, Nima Naderloui, Shenao Yan, Jingyu Zhang, Biying Liu, Ali Arastehfard, Heqing Huang, and Yuan Hong. SoK: Taxonomy and Evaluation of Prompt Security in Large Language Models.arXiv preprint arXiv:2510.15476, 2025
2025
-
[18]
The good and the bad: Exploring privacy issues in retrieval-augmented generation (RAG)
Shenglai Zeng, Jiankun Zhang, Pengfei He, Yiding Liu, Yue Xing, Han Xu, Jie Ren, Yi Chang, Shuaiqiang Wang, Dawei Yin, and Jiliang Tang. The good and the bad: Exploring privacy issues in retrieval-augmented generation (RAG). In Lun-Wei Ku, Andre Martins, and Vivek Srikumar, ed...
2024
-
[19]
Surveying the RAG Attack Surface and Defenses: Protecting Sensitive Company Data
Lynn V onderhaar, Daniel Machado, and Omar Ochoa. Surveying the RAG Attack Surface and Defenses: Protecting Sensitive Company Data. In2025 IEEE International Conference on Artificial Intelligence Testing (AITest), pages 69–76, 2025
2025
-
[20]
SafeRAG: Benchmarking security in retrieval-augmented generation of large language model
Xun Liang, Simin Niu, Zhiyu Li, Sensen Zhang, Hanyu Wang, Feiyu Xiong, Zhaoxin Fan, Bo Tang, Jihao Zhao, Jiawei Yang, Shichao Song, and Mengwei Wang. SafeRAG: Benchmarking security in retrieval-augmented generation of large language model. In Wanxiang Che, Joyce Nabende, Ekate...
2025
-
[21]
Retrieval-Augmented Generation: A Survey of Security Challenges and Countermeasures
Chao Wang, Haonan Li, Weijian Song, and Yiyang Lin. Retrieval-Augmented Generation: A Survey of Security Challenges and Countermeasures. In2025 11th IEEE International Conference on Privacy Computing and Data Security (PCDS), pages 210–217, 2025
2025
-
[22]
{PoisonedRAG}: Knowledge corruption attacks to {Retrieval-Augmented} generation of large language models
Wei Zou, Runpeng Geng, Binghui Wang, and Jinyuan Jia. {PoisonedRAG}: Knowledge corruption attacks to {Retrieval-Augmented} generation of large language models. In34th USENIX Security Symposium (USENIX Security 25), pages 3827–3844, 2025
2025
-
[23]
Traceback of Poisoning Attacks to Retrieval-Augmented Generation
Baolei Zhang, Haoran Xin, Minghong Fang, Zhuqing Liu, Biao Yi, Tong Li, and Zheli Liu. Traceback of Poisoning Attacks to Retrieval-Augmented Generation. InProceedings of the ACM on Web Conference 2025, pages 2085–2097, 2025
2025
-
[24]
BadRAG: Identifying Vulnerabilities in Retrieval Augmented Generation of Large Language Models.arXiv preprint arXiv:2406.00083, 2024
Jiaqi Xue, Mengxin Zheng, Yebowen Hu, Fei Liu, Xun Chen, and Qian Lou. BadRAG: Identifying Vulnerabilities in Retrieval Augmented Generation of Large Language Models.arXiv preprint arXiv:2406.00083, 2024
2024
-
[25]
Luan, Siran Wang, and Yuntao Wang
Yuan Chang, Tom H. Luan, Siran Wang, and Yuntao Wang. SafeRAG: Secure Cloud-Based Retrieval-Augmented Generation for LLM-Empowered V oice Assistants.IEEE Transactions on Network Science and Engineering, 13:6211–6224, 2026
2026
-
[26]
Privacy-Aware RAG-Enabled LLMs for Collaborative AI in Organizations
Georgios Fragkos, Bradley Marx, Sasha Safonov, Robert Manley, Winnie Patta, and Shelby Hiens. Privacy-Aware RAG-Enabled LLMs for Collaborative AI in Organizations. In2025 Cyber Awareness and Research Symposium (CARS), pages 1–8, 2025
2025
-
[27]
Trusted Execution Environments: Properties, Applications, and Challenges.IEEE Security & Privacy, 18(2):56–60, 2020
Patrick Jauernig, Ahmad-Reza Sadeghi, and Emmanuel Stapf. Trusted Execution Environments: Properties, Applications, and Challenges.IEEE Security & Privacy, 18(2):56–60, 2020
2020
-
[28]
Ai on the edge: a comprehensive review
Weixing Su, Linfeng Li, Fang Liu, Maowei He, and Xiaodan Liang. Ai on the edge: a comprehensive review. Artificial Intelligence Review, 55(8):6125–6183, 2022
2022
-
[29]
A Survey of AI Inference Technologies for On-Device Systems.IEEE Internet of Things Journal, 12(24):51927–51950, 2025
Wenzhu Wang, Ke Li, Bin Ji, Xiaodong Liu, Jie Yu, and Qingbo Wu. A Survey of AI Inference Technologies for On-Device Systems.IEEE Internet of Things Journal, 12(24):51927–51950, 2025
2025
-
[30]
Towards On-device Personalization: Cloud-device Collaborative Data Augmentation for Efficient On-device Language Model.ACM Transactions on Intelligent Systems and Technology, 2025
Zhaofeng Zhong, Wei Yuan, Liang Qu, Tong Chen, Hao Wang, Xiangyu Zhao, and Hongzhi Yin. Towards On-device Personalization: Cloud-device Collaborative Data Augmentation for Efficient On-device Language Model.ACM Transactions on Intelligent Systems and Technology, 2025
2025
-
[31]
Towards Trustworthy Retrieval Augmented Generation for Large Language Models: A Survey.arXiv preprint arXiv:2502.06872, 2025
Bo Ni, Zheyuan Liu, Leyao Wang, Yongjia Lei, Yuying Zhao, Xueqi Cheng, Qingkai Zeng, Luna Dong, Yinglong Xia, Krishnaram Kenthapadi, et al. Towards Trustworthy Retrieval Augmented Generation for Large Language Models: A Survey.arXiv preprint arXiv:2502.06872, 2025
2025
-
[32]
Federated retrieval-augmented generation: A systematic mapping study
Abhijit Chakraborty, Chahana Dahal, and Vivek Gupta. Federated retrieval-augmented generation: A systematic mapping study. In Christos Christodoulopoulos, Tanmoy Chakraborty, Carolyn Rose, and Violet Peng, editors, Findings of the Association for Computational Linguistics: EMN...
2025
-
[33]
Chien Van Nguyen, Xuan Shen, Ryan Aponte, Yu Xia, Samyadeep Basu, Zhengmian Hu, Jian Chen, Mihir Parmar, Sasidhar Kunapuli, Joe Barrow3, Junda Wu, Ashish Singh, Yu Wang, Jiuxiang Gu, Nesreen K. Ahmed, Nedim Lipka, Ruiyi Zhang, Xiang Chen, Tong Yu, Sungchul Kim, Hanieh Deilamsa...
2025
-
[34]
The Language Model Revolution: LLM and SLM Analysis
Zeynep Örpek, Bü¸ sra Tural, and Zeynep Destan. The Language Model Revolution: LLM and SLM Analysis. In 2024 8th International Artificial Intelligence and Data Processing Symposium (IDAP), pages 1–4, 2024
2024
-
[35]
Edge ai: A comprehensive survey of technologies, applications, and challenges
Vasuki Shankar. Edge ai: A comprehensive survey of technologies, applications, and challenges. In2024 1st International Conference on Advanced Computing and Emerging Technologies (ACET), pages 1–6, 2024
2024
-
[36]
Federated Learning and RAG Integration: A Scalable Approach for Medical Large Language Models
Jincheol Jung, Hongju Jeong, and Eui-Nam Huh. Federated Learning and RAG Integration: A Scalable Approach for Medical Large Language Models. In2025 International Conference on Artificial Intelligence in Information and Communication (ICAIIC), pages 0968–0973, 2025
2025
-
[37]
A Survey of Retrieval-Augmented Generation (RAG) for Large Language Models
Yusong Ma, Hongxuan Nie, Chao Chen, Jiujie Zhang, Jiali Jiang, Bisheng Wang, and Yuqin Xia. A Survey of Retrieval-Augmented Generation (RAG) for Large Language Models. In2025 International Conference on Trustworthy Big Data and Artificial Intelligence (ICTBAI), pages 7–13, 2025
2025
-
[38]
Graph-Based Approaches and Functionalities in Retrieval-Augmented Generation: A Comprehensive Survey.ACM Computing Surveys, 2025
Zulun Zhu, Tiancheng Huang, Kai Wang, Junda Ye, Xinghe Chen, and Siqiang Luo. Graph-Based Approaches and Functionalities in Retrieval-Augmented Generation: A Comprehensive Survey.ACM Computing Surveys, 2025
2025
-
[39]
A systematic literature review of retrieval-augmented generation: Techniques, metrics, and challenges.arXiv preprint arXiv:2508.06401, 2025
Andrew Brown, Muhammad Roman, and Barry Devereux. A systematic literature review of retrieval-augmented generation: Techniques, metrics, and challenges.arXiv preprint arXiv:2508.06401, 2025
2025
-
[40]
Backdoored Retrievers for Prompt Injection Attacks on Retrieval Augmented Generation of Large Language Models.arXiv preprint arXiv:2410.14479, 2024
Cody Clop and Yannick Teglia. Backdoored Retrievers for Prompt Injection Attacks on Retrieval Augmented Generation of Large Language Models.arXiv preprint arXiv:2410.14479, 2024
2024
-
[41]
Saidakhror Gulyamov, Said Gulyamov, Andrey Rodionov, Rustam Khursanov, Kambariddin Mekhmonov, Djakhongir Babaev, and Akmaljon Rakhimjonov. Prompt injection attacks in large language models and ai agent systems: A comprehensive review of vulnerabilities, attack vectors, and def...
2026
-
[42]
The Hidden Threat in Plain Text: Attacking RAG Data Loaders
Alberto Castagnaro, Umberto Salviati, Mauro Conti, Luca Pajola, and Simeone Pizzi. The Hidden Threat in Plain Text: Attacking RAG Data Loaders. InProceedings of the 18th ACM Workshop on Artificial Intelligence and Security, pages 170–181, 2025
2025
-
[43]
RAG LLMs are not safer: A safety analysis of retrieval-augmented generation for large language models
Bang An, Shiyue Zhang, and Mark Dredze. RAG LLMs are not safer: A safety analysis of retrieval-augmented generation for large language models. In Luis Chiruzzo, Alan Ritter, and Lu Wang, editors,Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the A...
2025
-
[45]
ALIGNMENT FAKING IN LARGE LANGUAGE MODELS.arXiv preprint arXiv:2412.14093, 2024
Ryan Greenblatt, Carson Denison, Benjamin Wright, Fabien Roger, Monte MacDiarmid, Sam Marks, Johannes Treutlein, Tim Belonax, Jack Chen, David Duvenaud, et al. ALIGNMENT FAKING IN LARGE LANGUAGE MODELS.arXiv preprint arXiv:2412.14093, 2024
2024 arXiv
-
[46]
Visual contextual attack: Jailbreaking MLLMs with image-driven context injection
Miao Ziqi, Yi Ding, Lijun Li, and Jing Shao. Visual contextual attack: Jailbreaking MLLMs with image-driven context injection. In Christos Christodoulopoulos, Tanmoy Chakraborty, Carolyn Rose, and Violet Peng, editors, Proceedings of the 2025 Conference on Empirical Methods in...
2025
-
[47]
Shaping the safety boundaries: Understanding and defending against jailbreaks in large language models
Lang Gao, Jiahui Geng, Xiangliang Zhang, Preslav Nakov, and Xiuying Chen. Shaping the safety boundaries: Understanding and defending against jailbreaks in large language models. In Wanxiang Che, Joyce Nabende, Ekaterina Shutova, and Mohammad Taher Pilehvar, editors,Proceedings...
-
[48]
Association for Computational Linguistics
-
[49]
Retrieval Poisoning Attacks Based on Prompt Injections into Retrieval-Augmented Generation Systems that Store Generated Responses
Yegor Anichkov, Victor Popov, and Sergey Bolovtsov. Retrieval Poisoning Attacks Based on Prompt Injections into Retrieval-Augmented Generation Systems that Store Generated Responses. InInternational Conference on Distributed Computer and Communication Networks, pages 417–429. ...
2024
-
[50]
Retrieval-augmented defense: Adaptive and controllable jailbreak prevention for large language models.arXiv preprint arXiv:2508.16406, 2025
Guangyu Yang, Jinghong Chen, Jingbiao Mei, Weizhe Lin, and Bill Byrne. Retrieval-augmented defense: Adaptive and controllable jailbreak prevention for large language models.arXiv preprint arXiv:2508.16406, 2025. 24 APREPRINT- JUNE25, 2026
2025
-
[51]
LatentPoison - Adversarial Attacks On The Latent Space.arXiv preprint arXiv:1711.02879, 2017
Antonia Creswell, Anil A Bharath, and Biswa Sengupta. LatentPoison - Adversarial Attacks On The Latent Space.arXiv preprint arXiv:1711.02879, 2017
2017 arXiv
-
[52]
Black-box Adversarial Attacks against Dense Retrieval Models: A Multi-view Contrastive Learning Method
Yu-An Liu, Ruqing Zhang, Jiafeng Guo, Maarten de Rijke, Wei Chen, Yixing Fan, and Xueqi Cheng. Black-box Adversarial Attacks against Dense Retrieval Models: A Multi-view Contrastive Learning Method. InProceedings of the 32nd ACM International Conference on Information and Know...
2023
-
[53]
Mask-based Membership Inference Attacks for Retrieval- Augmented Generation
Mingrui Liu, Sixiao Zhang, and Cheng Long. Mask-based Membership Inference Attacks for Retrieval- Augmented Generation. InProceedings of the ACM on Web Conference 2025, pages 2894–2907, 2025
2025
-
[54]
Flippedrag: Black-box opinion manipulation adversarial attacks to retrieval-augmented generation models
Zhuo Chen, Yuyang Gong, Jiawei Liu, Miaokun Chen, Haotan Liu, Qikai Cheng, Fan Zhang, Wei Lu, and Xi- aozhong Liu. Flippedrag: Black-box opinion manipulation adversarial attacks to retrieval-augmented generation models. InProceedings of the 2025 ACM SIGSAC Conference on Comput...
2025
-
[55]
Membership Inference Attacks Against Machine Learning Models
Reza Shokri, Marco Stronati, Congzheng Song, and Vitaly Shmatikov. Membership Inference Attacks Against Machine Learning Models. In2017 IEEE Symposium on Security and Privacy (SP), pages 3–18, 2017
2017
-
[56]
Practical Poisoning Attacks against Retrieval-Augmented Generation.arXiv preprint arXiv:2504.03957, 2025
Baolei Zhang, Yuxi Chen, Zhuqing Liu, Lihai Nie, Tong Li, Zheli Liu, and Minghong Fang. Practical Poisoning Attacks against Retrieval-Augmented Generation.arXiv preprint arXiv:2504.03957, 2025
2025
-
[57]
Drift Detection in Text Data with Document Embeddings
Robert Feldhans, Adrian Wilke, Stefan Heindorf, Mohammad Hossein Shaker, Barbara Hammer, Axel-Cyrille Ngonga Ngomo, and Eyke Hüllermeier. Drift Detection in Text Data with Document Embeddings. InIn- ternational Conference on Intelligent Data Engineering and Automated Learning,...
2021
-
[58]
Enhancing adversarial resilience in semantic caching for secure retrieval augmented generation systems.Scientific Reports, 16(1):5936, 2026
Mohanad Afiffy, Mohamed Waleed Fakhr, and Fahima A Maghraby. Enhancing adversarial resilience in semantic caching for secure retrieval augmented generation systems.Scientific Reports, 16(1):5936, 2026
2026
-
[59]
Towards More Robust Retrieval- Augmented Generation: Evaluating RAG Under Adversarial Poisoning Attacks.arXiv preprint arXiv:2412.16708, 2024
Jinyan Su, Jin Peng Zhou, Zhengxin Zhang, Preslav Nakov, and Claire Cardie. Towards More Robust Retrieval- Augmented Generation: Evaluating RAG Under Adversarial Poisoning Attacks.arXiv preprint arXiv:2412.16708, 2024
2024
-
[60]
DeRAG: Black-box Adversarial Attacks on Multiple Retrieval-Augmented Generation Applications via Prompt Injection.arXiv preprint arXiv:2507.15042, 2025
Jerry Wang and Fang Yu. DeRAG: Black-box Adversarial Attacks on Multiple Retrieval-Augmented Generation Applications via Prompt Injection.arXiv preprint arXiv:2507.15042, 2025
2025
-
[61]
Liu, Kevin Lin, John Hewitt, Ashwin Paranjape, Michele Bevilacqua, Fabio Petroni, and Percy Liang
Nelson F. Liu, Kevin Lin, John Hewitt, Ashwin Paranjape, Michele Bevilacqua, Fabio Petroni, and Percy Liang. Lost in the middle: How language models use long contexts.Transactions of the Association for Computational Linguistics, 12:157–173, 2024
2024
-
[62]
Out-of-context and out-of-scope: Manip- ulating large language models through minimal instruction set modifications.PLoS One, 21(2):e0341558, 2026
Monty-Maximilian Zühlke, Daniel Kudenko, and Wolfgang Nejdl. Out-of-context and out-of-scope: Manip- ulating large language models through minimal instruction set modifications.PLoS One, 21(2):e0341558, 2026
2026
-
[63]
Chi, Nathanael Schärli, and Denny Zhou
Freda Shi, Xinyun Chen, Kanishka Misra, Nathan Scales, David Dohan, Ed H. Chi, Nathanael Schärli, and Denny Zhou. Large Language Models Can Be Easily Distracted by Irrelevant Context. In Andreas Krause, Emma Brunskill, Kyunghyun Cho, Barbara Engelhardt, Sivan Sabato, and Jonat...
2023
-
[64]
Knowledge conflicts for LLMs: A survey
Rongwu Xu, Zehan Qi, Zhijiang Guo, Cunxiang Wang, Hongru Wang, Yue Zhang, and Wei Xu. Knowledge conflicts for LLMs: A survey. In Yaser Al-Onaizan, Mohit Bansal, and Yun-Nung Chen, editors,Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, p...
2024
-
[65]
When not to trust language models: Investigating effectiveness of parametric and non-parametric memories
Alex Mallen, Akari Asai, Victor Zhong, Rajarshi Das, Daniel Khashabi, and Hannaneh Hajishirzi. When not to trust language models: Investigating effectiveness of parametric and non-parametric memories. In Anna Rogers, Jordan Boyd-Graber, and Naoaki Okazaki, editors,Proceedings ...
2023
-
[66]
Adversarial and Multilingual Threats in Retrieval-Augmented Generation: From Prompt Injection to Model Exploitation
Basma ElSaify and Mohamed Baderelden. Adversarial and Multilingual Threats in Retrieval-Augmented Generation: From Prompt Injection to Model Exploitation. In2025 2nd International Generative AI and Computational Language Modelling Conference (GACLM), pages 155–162, 2025
2025
-
[67]
Enhancing noise robustness of retrieval-augmented language models with adaptive adversarial training
Feiteng Fang, Yuelin Bai, Shiwen Ni, Min Yang, Xiaojun Chen, and Ruifeng Xu. Enhancing noise robustness of retrieval-augmented language models with adaptive adversarial training. In Lun-Wei Ku, Andre Martins, and Vivek Srikumar, editors,Proceedings of the 62nd Annual Meeting o...
2026
-
[68]
Information Leakage in Embedding Models
Congzheng Song and Ananth Raghunathan. Information Leakage in Embedding Models. InProceedings of the 2020 ACM SIGSAC conference on computer and communications security, pages 377–390, 2020
2020
-
[69]
Reverse engineering convolutional neural networks through side-channel information leaks
Weizhe Hua, Zhiru Zhang, and G Edward Suh. Reverse engineering convolutional neural networks through side-channel information leaks. InProceedings of the 55th Annual Design Automation Conference, pages 1–6, 2018
2018
-
[70]
Deep learning model inversion attacks and defenses: a comprehensive survey.Artificial Intelligence Review, 58(8):242, 2025
Wencheng Yang, Song Wang, Di Wu, Taotao Cai, Yanming Zhu, Shicheng Wei, Yiying Zhang, Xu Yang, Zhaohui Tang, and Yan Li. Deep learning model inversion attacks and defenses: a comprehensive survey.Artificial Intelligence Review, 58(8):242, 2025
2025
-
[71]
Federated retrieval augmented generation for multi-product question answering
Parshin Shojaee, Sai Sree Harsha, Dan Luo, Akash Maharaj, Tong Yu, and Yunyao Li. Federated retrieval augmented generation for multi-product question answering. In Owen Rambow, Leo Wanner, Marianna Apidianaki, Hend Al-Khalifa, Barbara Di Eugenio, Steven Schockaert, Kareem Darw...
2025
-
[72]
The Limitations of Federated Learning in Sybil Settings
Clement Fung, Chris JM Yoon, and Ivan Beschastnikh. The Limitations of Federated Learning in Sybil Settings. In23rd International symposium on research in attacks, intrusions and defenses (RAID 2020), pages 301–316, 2020
2020
-
[73]
Federated Retrieval- Augmented Generation-Based LLM for Enhanced Cyber Threat Detection in the Internet-of-Energy.IEEE Network, 40(1):13–19, 2026
Tianxing Fu, Jia Hu, Geyong Min, Sunder Ali Khowaja, Keshav Singh, and Kapal Dev. Federated Retrieval- Augmented Generation-Based LLM for Enhanced Cyber Threat Detection in the Internet-of-Energy.IEEE Network, 40(1):13–19, 2026
2026
-
[74]
Sybil Attacks and Defense on Differential Privacy based Federated Learning
Yupeng Jiang, Yong Li, Yipeng Zhou, and Xi Zheng. Sybil Attacks and Defense on Differential Privacy based Federated Learning. In2021 IEEE 20th International Conference on Trust, Security and Privacy in Computing and Communications (TrustCom), pages 355–362, 2021
2021
-
[75]
A Survey of Federated Learning: Advances in Architecture, Synchronization, and Security Threats.Computers, Materials & Continua, 86(3), 2026
Faisal Mahmud, Fahim Mahmud, and Rashedur M Rahman. A Survey of Federated Learning: Advances in Architecture, Synchronization, and Security Threats.Computers, Materials & Continua, 86(3), 2026
2026
-
[76]
Privacy protection in RAG: A novel method and evaluation framework.Information Processing & Management, 63(3):104505, 2026
Yuan Zhang, Jionghan Wu, Rui Li, Tong Zhang, Yujie Song, Chuanyi Li, Shangqi Wang, Hao Shen, Jiao Yin, Jidong Ge, and Bin Luo. Privacy protection in RAG: A novel method and evaluation framework.Information Processing & Management, 63(3):104505, 2026
2026
-
[77]
Efficient Byzantine-Robust and Privacy-Preserving Federated Learning on Compressive Domain.IEEE Internet of Things Journal, 11(4):7116–7127, 2024
Guiqiang Hu, Hongwei Li, Wenshu Fan, and Yushu Zhang. Efficient Byzantine-Robust and Privacy-Preserving Federated Learning on Compressive Domain.IEEE Internet of Things Journal, 11(4):7116–7127, 2024
2024
-
[78]
Fali Wang, Zhiwei Zhang, Xianren Zhang, Zongyu Wu, Tzuhao Mo, Qiuhao Lu, Wanjing Wang, Rui Li, Junjie Xu, Xianfeng Tang, et al. A Comprehensive Survey of Small Language Models in the Era of Large Language Models: Techniques, Enhancements, Applications, Collaboration with LLMs,...
2025
-
[79]
RAG-Guardrails Integration for AI Content Control
Rakesh More. RAG-Guardrails Integration for AI Content Control. InProceedings of the 2025 18th International Conference on Computer Science and Information Technology, pages 260–268, 2025
2025
-
[80]
Innovative Guardrails for Generative AI: Designing an Intelligent Filter for Safe and Responsible LLM Deployment.Applied Sciences, 15(13):7298, 2025
Olga Shvetsova, Danila Katalshov, and Sang-Kon Lee. Innovative Guardrails for Generative AI: Designing an Intelligent Filter for Safe and Responsible LLM Deployment.Applied Sciences, 15(13):7298, 2025
2025
-
[81]
Guardrails for Large Language Models: A Review of Techniques and Challenges.J Artif Intell Mach Learn & Data Sci, 3(1):2504–2512, 2025
Syed Arham Akheel. Guardrails for Large Language Models: A Review of Techniques and Challenges.J Artif Intell Mach Learn & Data Sci, 3(1):2504–2512, 2025
2025
-
[82]
Anonymization Techniques for Privacy Preserving Data Publishing: A Comprehensive Survey.IEEE Access, 9:8512–8545, 2021
Abdul Majeed and Sungchang Lee. Anonymization Techniques for Privacy Preserving Data Publishing: A Comprehensive Survey.IEEE Access, 9:8512–8545, 2021
2021
-
[83]
BAG-RAG: Bidirectional Retrieval-Augmented Generation Based on Multi-Layer Semantic Graphs for Budget Auditing QA
Gaofeng Xu, Runzhe Wang, Guilin Qi, Xiaolong Ye, Yongrui Chen, Yuxin Zhang, Xinbang Dai, Yuan Meng, and Shenwen Zhong. BAG-RAG: Bidirectional Retrieval-Augmented Generation Based on Multi-Layer Semantic Graphs for Budget Auditing QA. InInternational Conference on Database Syst...
2025
-
[84]
LAIR: A Language For Automated Semantics- aware Text Sanitization based on Frame Semantics
Steffen Hedegaard, Søren Houen, and Jakob Grue Simonsen. LAIR: A Language For Automated Semantics- aware Text Sanitization based on Frame Semantics. In2009 IEEE International Conference on Semantic Computing, pages 47–52. IEEE, 2009
2009
-
[85]
Making pre-trained language models better few-shot learners
Tianyu Gao, Adam Fisch, and Danqi Chen. Making pre-trained language models better few-shot learners. In Chengqing Zong, Fei Xia, Wenjie Li, and Roberto Navigli, editors,Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th Interna...
2026
-
[86]
A Survey on Differential Privacy for Unstructured Data Content.ACM Computing Surveys (CSUR), 54(10s):1–28, 2022
Ying Zhao and Jinjun Chen. A Survey on Differential Privacy for Unstructured Data Content.ACM Computing Surveys (CSUR), 54(10s):1–28, 2022
2022
-
[87]
RAG with Differential Privacy
Nicolas Grislain. RAG with Differential Privacy. In2025 IEEE Conference on Artificial Intelligence (CAI), pages 847–852, 2025
2025
-
[88]
Textual Differential Privacy for Context-Aware Reasoning with Large Language Model
Junwei Yu, Jieyu Zhou, Yepeng Ding, Lingfeng Zhang, Yuheng Guo, and Hiroyuki Sato. Textual Differential Privacy for Context-Aware Reasoning with Large Language Model. In2024 IEEE 48th Annual Computers, Software, and Applications Conference (COMPSAC), pages 988–997, 2024
2024
-
[89]
DF-RAG: A Dual Federated Retrieval-Augmented Generation Framework for Collaborative Medical AI
Julian Garcia, Jiaqi Gong, Michal Zajac, and Andrew Hahn. DF-RAG: A Dual Federated Retrieval-Augmented Generation Framework for Collaborative Medical AI. InProceedings of the ACM/IEEE International Conference on Connected Health: Applications, Systems and Engineering Technolog...
2025
-
[90]
Software Architecture for Federated Retrieval-Augmented Clinical QA System Using IoT for Continuous Monitoring
Felix Negoit, ˘a and Andreea Ionela Dumachi. Software Architecture for Federated Retrieval-Augmented Clinical QA System Using IoT for Continuous Monitoring. In2025 International Semiconductor Conference (CAS), pages 381–384, 2025
2025
-
[91]
Optimizing Legal Information Access: Federated Search and RAG for Secure AI-Powered Legal Solutions
Flora Amato, Egidia Cirillo, Mattia Fonisto, and Alberto Moccardi. Optimizing Legal Information Access: Federated Search and RAG for Secure AI-Powered Legal Solutions. In2024 IEEE International Conference on Big Data (BigData), pages 7632–7639, 2024
2024
-
[92]
Privacy-Preserving Data Sharing with Personalized Encrypted Retrieval.Applied Sciences, 16(6):2771, 2026
Hongfei Song, Lianhai Wang, Shujiang Xu, Shuhui Zhang, Wei Shao, and Qizheng Wang. Privacy-Preserving Data Sharing with Personalized Encrypted Retrieval.Applied Sciences, 16(6):2771, 2026
2026
-
[93]
Leveraging Searchable Encryption through Homomorphic Encryption: A Comprehensive Analysis.Mathematics, 11(13):2948, 2023
Ivone Amorim and Ivan Costa. Leveraging Searchable Encryption through Homomorphic Encryption: A Comprehensive Analysis.Mathematics, 11(13):2948, 2023
2023
-
[94]
RemoteRAG: A privacy-preserving LLM cloud RAG service
Yihang Cheng, Lan Zhang, Junyang Wang, Mu Yuan, and Yunhao Yao. RemoteRAG: A privacy-preserving LLM cloud RAG service. In Wanxiang Che, Joyce Nabende, Ekaterina Shutova, and Mohammad Taher Pilehvar, editors,Findings of the Association for Computational Linguistics: ACL 2025, p...
2025
-
[95]
Confidential Computing Using Trusted Execution Environments.International Journal of AI, BigData, Computational and Management Studies, 4(2):97–110, 2023
Sunil Anasuri. Confidential Computing Using Trusted Execution Environments.International Journal of AI, BigData, Computational and Management Studies, 4(2):97–110, 2023
2023
-
[96]
Don’t forget private retrieval: distributed private similarity search for large language models
Guy Zyskind, Tobin South, and Alex Pentland. Don’t forget private retrieval: distributed private similarity search for large language models. In Ivan Habernal, Sepideh Ghanavati, Abhilasha Ravichander, Vijayanta Jain, Patricia Thaine, Timour Igamberdiev, Niloofar Mireshghallah...
2024
-
[97]
Kuniyasu Suzaki, Kenta Nakajima, Tsukasa Oi, and Akira Tsukamoto. TS-PERF: General Performance Measurement of Trusted Execution Environment and Rich Execution Environment on Intel SGX, ARM TrustZone, and RISC-V Keystone.IEEE Access, 9:133520–133530, 2021
2021
-
[98]
Secure Multi-Party Computation: Theory, practice and applications.Information Sciences, 476:357–372, 2019
Chuan Zhao, Shengnan Zhao, Minghao Zhao, Zhenxiang Chen, Chong-Zhi Gao, Hongwei Li, and Yu-an Tan. Secure Multi-Party Computation: Theory, practice and applications.Information Sciences, 476:357–372, 2019
2019
-
[99]
Privacy-preserving aggregation in federated learning: A survey.IEEE Transactions on Big Data, 2022
Ziyao Liu, Jiale Guo, Wenzhuo Yang, Jiani Fan, Kwok-Yan Lam, and Jun Zhao. Privacy-preserving aggregation in federated learning: A survey.IEEE Transactions on Big Data, 2022
2022
-
[100]
Secure multi-party computation (SMPC) protocols and privacy
Mosiur Rahaman, Varsha Arya, Sheila Mae Orozco, and Princy Pappachan. Secure multi-party computation (SMPC) protocols and privacy. InInnovations in Modern Cryptography, pages 193–218. IGI Global Scientific Publishing, 2024
2024
-
[101]
Secure Multi-Party Computation for Machine Learning: A Survey.IEEE Access, 12:53881–53899, 2024
Ian Zhou, Farzad Tofigh, Massimo Piccardi, Mehran Abolhasan, Daniel Franklin, and Justin Lipman. Secure Multi-Party Computation for Machine Learning: A Survey.IEEE Access, 12:53881–53899, 2024
2024
-
[102]
Preserving User Privacy in Retrieval Augmented Generation: A Novel Approach Using Local Placeholder Tagging
Thang Nguyen Xuan, Vinh Nguyen Thanh, Thuy Duong Nguyen Duy, Son Tran Huy Hoang, Gia Bao Nguyen, and Thao Nguyen Thi Ngoc. Preserving User Privacy in Retrieval Augmented Generation: A Novel Approach Using Local Placeholder Tagging. InInternational Conference on Responsible Art...
2024
-
[103]
Open-RAG: Enhanced retrieval augmented reasoning with open-source large language models
Shayekh Bin Islam, Md Asib Rahman, K S M Tozammel Hossain, Enamul Hoque, Shafiq Joty, and Md Rizwan Parvez. Open-RAG: Enhanced retrieval augmented reasoning with open-source large language models. In Yaser Al-Onaizan, Mohit Bansal, and Yun-Nung Chen, editors,Findings of the As...
2024
-
[104]
(accessed January 06, 2025)
MS MARCO Dataset.https://microsoft.github.io/msmarco/. (accessed January 06, 2025)
2025
-
[105]
Dai, Jakob Uszkoreit, Quoc Le, and Slav Petrov
Tom Kwiatkowski, Jennimaria Palomaki, Olivia Redfield, Michael Collins, Ankur Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Jacob Devlin, Kenton Lee, Kristina Toutanova, Llion Jones, Matthew Kelcey, Ming-Wei Chang, Andrew M. Dai, Jakob Uszkoreit, Quoc Le, and Slav...
2019
-
[106]
Resources for brewing beir: Reproducible reference models and statistical analyses
Ehsan Kamalloo, Nandan Thakur, Carlos Lassance, Xueguang Ma, Jheng-Hong Yang, and Jimmy Lin. Resources for brewing beir: Reproducible reference models and statistical analyses. InProceedings of the 47th International ACM SIGIR Conference on Research and Development in Informat...
2024
-
[107]
TriviaQA: A large scale distantly supervised challenge dataset for reading comprehension
Mandar Joshi, Eunsol Choi, Daniel Weld, and Luke Zettlemoyer. TriviaQA: A large scale distantly supervised challenge dataset for reading comprehension. In Regina Barzilay and Min-Yen Kan, editors,Proceedings of the 55th Annual Meeting of the Association for Computational Lingu...
2017
-
[108]
Zhilin Yang, Peng Qi, Saizheng Zhang, Yoshua Bengio, William Cohen, Ruslan Salakhutdinov, and Christopher D. Manning. HotpotQA: A dataset for diverse, explainable multi-hop question answering. In Ellen Riloff, David Chiang, Julia Hockenmaier, and Jun’ichi Tsujii, editors,Proce...
2018
-
[109]
The DRAGON benchmark for clinical NLP
Joeran S Bosma, Koen Dercksen, Luc Builtjes, Romain André, Christian Roest, Stefan J Fransen, Constant R Noordman, Mar Navarro-Padilla, Judith Lefkes, Natália Alves, et al. The DRAGON benchmark for clinical NLP. NPJ Digital Medicine, 8(1):289, 2025
2025
-
[110]
Information Retrieval in the Age of Generative AI: The RGB Model
Michele Garetto, Alessandro Cornacchia, Franco Galante, Emilio Leonardi, Alessandro Nordio, and Alberto Tarable. Information Retrieval in the Age of Generative AI: The RGB Model. InProceedings of the 48th International ACM SIGIR Conference on Research and Development in Inform...
2025
-
[111]
PubMedQA: A dataset for biomedical research question answering
Qiao Jin, Bhuwan Dhingra, Zhengping Liu, William Cohen, and Xinghua Lu. PubMedQA: A dataset for biomedical research question answering. In Kentaro Inui, Jing Jiang, Vincent Ng, and Xiaojun Wan, editors, Proceedings of the 2019 Conference on Empirical Methods in Natural Languag...
2019
-
[112]
What Disease Does This Patient Have? A Large-Scale Open Domain Question Answering Dataset from Medical Exams.Applied Sciences, 11(14):6421, 2021
Di Jin, Eileen Pan, Nassim Oufattole, Wei-Hung Weng, Hanyi Fang, and Peter Szolovits. What Disease Does This Patient Have? A Large-Scale Open Domain Question Answering Dataset from Medical Exams.Applied Sciences, 11(14):6421, 2021
2021
-
[113]
ChatDoctor: A Medical Chat Model Fine-Tuned on a Large Language Model Meta-AI (LLaMA) Using Medical Domain Knowledge.Cureus, 15(6), 2023
Yunxiang Li, Zihan Li, Kai Zhang, Ruilong Dan, Steve Jiang, and You Zhang. ChatDoctor: A Medical Chat Model Fine-Tuned on a Large Language Model Meta-AI (LLaMA) Using Medical Domain Knowledge.Cureus, 15(6), 2023
2023
-
[114]
(accessed April 13, 2026)
National Vulnerability Database.https://nvd.nist.gov/. (accessed April 13, 2026)
2026
-
[115]
LegalBench: A Collaboratively Built Benchmark for Measuring Legal Reasoning in Large Language Models
Neel Guha, Julian Nyarko, Daniel Ho, Christopher Ré, Adam Chilton, Aditya K, Alex Chohlas-Wood, Austin Peters, Brandon Waldon, Daniel Rockmore, Diego Zambrano, Dmitry Talisman, Enam Hoque, Faiz Surani, Frank Fagan, Galit Sarfaty, Gregory Dickinson, Haggai Porat, Jason Hegland,...
2023
-
[116]
LEAF: A Benchmark for Federated Settings.arXiv preprint arXiv:1812.01097, 2018
Sebastian Caldas, Sai Meher Karthik Duddu, Peter Wu, Tian Li, Jakub Koneˇcn`y, H Brendan McMahan, Virginia Smith, and Ameet Talwalkar. LEAF: A Benchmark for Federated Settings.arXiv preprint arXiv:1812.01097, 2018
2018
-
[117]
FedMatch: Federated Learning Over Heterogeneous Question Answering Data
Jiangui Chen, Ruqing Zhang, Jiafeng Guo, Yixing Fan, and Xueqi Cheng. FedMatch: Federated Learning Over Heterogeneous Question Answering Data. InProceedings of the 30th ACM International Conference on Information & Knowledge Management, pages 181–190, 2021
2021
-
[118]
PrivacyLens: Evaluating Privacy Norm Awareness of Language Models in Action
Yijia Shao, Tianshi Li, Weiyan Shi, Yanchen Liu, and Diyi Yang. PrivacyLens: Evaluating Privacy Norm Awareness of Language Models in Action. In A. Globerson, L. Mackey, D. Belgrave, A. Fan, U. Paquet, 28 APREPRINT- JUNE25, 2026 J. Tomczak, and C. Zhang, editors,Advances in Neu...
2026
-
[119]
Privacy protection in RAG: A novel method and evaluation framework.Information Processing & Management, 63(3):104505, 2026
Yuan Zhang, Jionghan Wu, Rui Li, Tong Zhang, Yujie Song, Chuanyi Li, Shangqi Wang, Hao Shen, Jiao Yin, Jidong Ge, et al. Privacy protection in RAG: A novel method and evaluation framework.Information Processing & Management, 63(3):104505, 2026
2026
-
[120]
Guardian Angel: A Secure and Efficient Retrieval-Augmented Generation Framework
Xi Fang, Liang Qiao, Jun Shi, and Hong An. Guardian Angel: A Secure and Efficient Retrieval-Augmented Generation Framework. In2025 5th International Conference on Artificial Intelligence and Industrial Technology Applications (AIITA), pages 1773–1777, 2025
2025
-
[121]
Kail Eszter, Tafferner-Gulyás Viktória, and Habil
Dobrovodsky Patrik, Balázsné Dr. Kail Eszter, Tafferner-Gulyás Viktória, and Habil. Fleiner Rita. Evaluation of Large Language Models Enhanced with Retrieval-Augmented Generation: A literature review. In2026 IEEE 24th World Symposium on Applied Machine Intelligence and Informa...
2026
-
[122]
Benchmarking of Retrieval Augmented Generation: A Comprehensive Systematic Literature Review on Evaluation Dimensions, Evaluation Metrics and Datasets
Simon Knollmeyer, O˘guz Caymazer, Leonid Koval, Muhammad Uzair Akmal, Saara Asif, Selvine George Math- ias, and Daniel Großmann. Benchmarking of Retrieval Augmented Generation: A Comprehensive Systematic Literature Review on Evaluation Dimensions, Evaluation Metrics and Datase...
2024
-
[123]
Agentic AI: a comprehensive survey of architectures, applications, and future directions.Artificial Intelligence Review, 59(1):11, 2025
Mohamad Abou Ali, Fadi Dornaika, and Jinan Charafeddine. Agentic AI: a comprehensive survey of architectures, applications, and future directions.Artificial Intelligence Review, 59(1):11, 2025
2025
-
[124]
Guided and Federated RAG: Architec- tural Models for Trustworthy AI in Data Spaces
Carlos Mario Braga, Manuel A Serrano, and Eduardo Fernández-Medina. Guided and Federated RAG: Architec- tural Models for Trustworthy AI in Data Spaces. InInternational Conference on Intelligent Data Engineering and Automated Learning, pages 363–374. Springer, 2025
2025
-
[125]
Summary of a haystack: A challenge to long-context LLMs and RAG systems
Philippe Laban, Alexander Fabbri, Caiming Xiong, and Chien-Sheng Wu. Summary of a haystack: A challenge to long-context LLMs and RAG systems. In Yaser Al-Onaizan, Mohit Bansal, and Yun-Nung Chen, editors, Proceedings of the 2024 Conference on Empirical Methods in Natural Langu...
2024
-
[126]
Inference scaling for bridging retrieval and augmented generation
Youngwon Lee, Seung-won Hwang, Daniel F Campos, Filip Grali´nski, Zhewei Yao, and Yuxiong He. Inference scaling for bridging retrieval and augmented generation. In Luis Chiruzzo, Alan Ritter, and Lu Wang, editors, Findings of the Association for Computational Linguistics: NAAC...
2025
-
[127]
Robust Neural Information Retrieval: An Adversarial and Out-of-Distribution Perspective.ACM Transactions on Information Systems, 44(1):1–48, 2025
Yu-An Liu, Ruqing Zhang, Jiafeng Guo, Maarten de Rijke, Yixing Fan, and Xueqi Cheng. Robust Neural Information Retrieval: An Adversarial and Out-of-Distribution Perspective.ACM Transactions on Information Systems, 44(1):1–48, 2025
2025
-
[128]
Efficient Vector-Multiplicative Privacy-Preserving Retrieval-Augmented Generation for Large Language Models.IEEE Transactions on Dependable and Secure Computing, pages 1–17, 2026
Jinhao Zhou and Jun Wu. Efficient Vector-Multiplicative Privacy-Preserving Retrieval-Augmented Generation for Large Language Models.IEEE Transactions on Dependable and Secure Computing, pages 1–17, 2026
2026
-
[129]
pFedRAG: A personalized federated retrieval- augmented generation system with depth-adaptive tiered embedding tuning
Hangyu He, Xin Yuan, Kai Wu, Ren Ping Liu, and Wei Ni. pFedRAG: A personalized federated retrieval- augmented generation system with depth-adaptive tiered embedding tuning. In Christos Christodoulopoulos, Tanmoy Chakraborty, Carolyn Rose, and Violet Peng, editors,Findings of t...
2025
-
[130]
Enhancing Security and Appli- cability of Local LLM-Based Document Retrieval Systems in Smart Grid Isolated Environments.Electronics, 14(17):3407, 2025
Kiho Lee, Sumi Yang, Jaeyeong Jeong, Yongjoon Lee, and Dongkyoo Shin. Enhancing Security and Appli- cability of Local LLM-Based Document Retrieval Systems in Smart Grid Isolated Environments.Electronics, 14(17):3407, 2025
2025
-
[131]
Trustworthy Recommender Systems.ACM Transactions on Intelligent Systems and Technology, 15(4):1–20, 2024
Shoujin Wang, Xiuzhen Zhang, Yan Wang, and Francesco Ricci. Trustworthy Recommender Systems.ACM Transactions on Intelligent Systems and Technology, 15(4):1–20, 2024
2024
-
[132]
Shambhu Adhikari. Large Language Models in Modern Data Engineering: A Systematic Review of Architectures, Use Cases, and Limitations.International Journal of Business & Computational Science, 2(1), 2025
2025
-
[133]
Large Language Models for Con- structing and Optimizing Machine Learning Workflows: A Survey.ACM Transactions on Software Engineering and Methodology, 2025
Yang Gu, Hengyu You, Jian Cao, Muran Yu, Haoran Fan, and Shiyou Qian. Large Language Models for Con- structing and Optimizing Machine Learning Workflows: A Survey.ACM Transactions on Software Engineering and Methodology, 2025. 29
2025
Reviewed June 25, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.