Pith. sign in

REVIEW 3 major objections 5 minor 12 cited by

Towards Trustworthy Retrieval Augmented Generation for Large Language Models: A Survey

T0 review · 3 major / 5 minor · reviewed 2026-08-08 · deepseek-v4-flash

Pith's one-line read A unified roadmap for trustworthy retrieval-augmented generation

desk verdict Useful six-pillar taxonomy for trustworthy RAG, but the 'comprehensive roadmap' claim outruns the reported search methodology; worth a serious referee after protocol reporting and mechanical fixes. read the letter →

arxiv 2502.06872 v1 pith:TFCULT3Z submitted 2025-02-08 cs.CL cs.AI

classification cs.CLcs.AI
keywords retrieval-augmentedgenerationtrustworthyAIlargelanguagemodelsreliabilityprivacysafetyfairnessexplainability
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This survey claims that the trustworthiness problems of retrieval-augmented generation (RAG) for large language models form a structured family of risks that deserves a single roadmap. It argues that existing surveys cover either RAG generally or trustworthy LLMs generally, but none maps the full trust landscape of RAG itself. The intended contribution is a unifying framework built on six aspects: reliability, privacy, safety, fairness, explainability, and accountability. For each aspect, the survey offers a taxonomy, a review of methods, evaluation protocols, and future directions. A sympathetic reader would take the central claim to be that this roadmap lets researchers position new work, compare solutions, and spot gaps instead of treating trustworthiness as a vague goal.

What carries the argument

The organizing device is a six-pillar trustworthiness taxonomy taken from NIST's framework and applied to the RAG pipeline: reliability, privacy, safety, fairness, explainability, and accountability. Each pillar is decomposed by module such as retrieval, generation, or both, with a table of representative works, evaluation metrics, datasets, and future directions attached. The taxonomy carries the argument by converting 'trustworthy RAG' from a diffuse goal into a set of reviewable categories, and the comparison table against prior surveys supplies the gap the paper says it fills.

What would settle it

A reader could re-run the Section 2.4 search protocol with explicit inclusion criteria and snowballing of references; if that replication surfaces a substantial body of trustworthy-RAG work published before October 2024 that falls outside the six taxonomies or is missing from the survey's sections, the roadmap's completeness claim fails.

Watch

Extended reading notes

Core claim

On its own terms, the paper's discovery is organizational: it finds that the scattered literature on RAG failures and fixes falls naturally into six aspects, with each aspect split by pipeline stage (retrieval, generation, or the whole system). It claims that previous RAG surveys and trustworthy-LLM surveys each illuminate only part of this space, so no existing review gives a researcher a way to see how a retrieval-poisoning attack, a privacy leak, and a bias problem relate. The survey further claims that each aspect already has enough published attacks, defenses, evaluation metrics, and datasets to be taxonomized and compared, and that doing so exposes concrete gaps, such as the near-absence of dedicated adversarial defenses for RAG and the absence of explainability methods aimed specifically at the retriever.

Load-bearing premise

The roadmap's completeness rests on the assumption that keyword searches of Google Scholar, ACM Digital Library, and arXiv up to October 2024, without reported inclusion or exclusion criteria or coverage checks, recovered a representative sample of all relevant trustworthy-RAG work.

Editorial extensions

If this is right

  • Researchers can place a new attack or defense on a specific pillar and pipeline stage, making results from different papers comparable in a way the field currently lacks.
  • The per-pillar evaluation reviews expose missing infrastructure, including the absence of a standard RAG robustness benchmark, of established privacy baselines, and of unified datasets for explainability in RAG.
  • The roadmap identifies underdeveloped fronts, such as dedicated adversarial defenses for the retriever and explainability of the retrieval stage, turning those into concrete research programs.
  • Downstream domains like healthcare, law, and education can use the framework to decide which trust dimension matters most for a given deployment, since the six pillars interact differently in each domain.
  • The future-direction lists, such as integrated uncertainty and robustness handling and unified retrieval-plus-generation watermarking, become explicit targets for follow-up work.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the coverage claim is right, the natural next step is to turn the taxonomy into a living benchmark suite that continuously re-classifies new RAG trust papers, which the survey itself does not provide.
  • The six-pillar split may understate interactions: a single retrieval poisoning attack can simultaneously degrade safety, privacy, fairness, and reliability, so a cross-pillar matrix might be a needed extension beyond the paper's independent sections.
  • A testable extension would apply the taxonomy to post-October 2024 work and to non-text RAG variants such as knowledge-graph and multimodal retrieval, checking whether the six categories still partition the space cleanly.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. This survey aims to provide a unified, comprehensive roadmap for trustworthiness in Retrieval-Augmented Generation (RAG) for large language models. It organizes the literature into six perspectives—reliability, privacy, safety, fairness, explainability, and accountability—and for each offers a taxonomy, representative methods, evaluation metrics and datasets, and future research directions. It also discusses downstream applications in healthcare, law, and education, and compares its scope with prior RAG and LLM-trustworthiness surveys.

Significance. If the roadmap holds up, the survey fills a real gap: prior RAG surveys touch trustworthiness only in passing, and LLM trustworthiness surveys do not account for the retrieval stage. The paper's structured taxonomies per perspective, its comparison table with existing surveys (Table 1), and the accompanying GitHub repository of references are useful assets for researchers positioning new work. The per-section treatment of evaluation metrics and datasets is particularly valuable. However, the central claim of comprehensiveness rests on a weak, under-reported literature search methodology, and several internal inconsistencies suggest the manuscript has not been carefully checked.

major comments (3)
  1. [Section 2.4] The paper collection methodology is described only as keyword searches on Google Scholar, ACM Digital Library, and arXiv with an October 2024 cutoff. No query strings, inclusion/exclusion criteria, screening counts, deduplication steps, or coverage validation are reported. Since the paper's central value proposition is a 'comprehensive roadmap' (Abstract; Section 2.3), this makes the completeness of the taxonomy unfalsifiable. Please provide a detailed search protocol, a PRISMA-style flow diagram with screening counts, and a coverage check (e.g., recall against the reference lists of recent RAG/trustworthiness surveys or a citation-based validation).
  2. [Section 7.1.1 and Section 4.1.2] The survey states that 'no dedicated research efforts have been made to explain retrieval within the context of RAG' (Section 7.1.1), that privacy research is 'in its infancy' (Section 4.1.2), and that adversarial defenses are 'rudimentary' (Section 5.4). These admissions are compatible with the survey's claims, but the 'comprehensive roadmap' framing requires the authors to state explicitly which parts of each taxonomy are descriptive of existing work and which are prescriptive research agendas. Without this distinction, a reader cannot tell whether a sparse section reflects a genuinely underdeveloped area or an incomplete literature search; the classification currently blurs these two cases.
  3. [Abstract] The abstract says the discussion is organized around 'five key perspectives' but immediately lists six: reliability, privacy, safety, fairness, explainability, and accountability. The paper also contains multiple broken cross-references (Figure ?? in Section 2.1 and Section 3.2.2; Table ?? in Section 2.3.1). For a survey whose contribution is organizational structure, these are not merely cosmetic: they suggest the 'comprehensive' claim has not been checked against the manuscript's own scope and presentation. Please correct the count and all unresolved references.
minor comments (5)
  1. [Section 7.4] The sentence 'Our work pioneers the integration of knowledge graphs (KGs) and large language models (LLMs)...' is a self-referential claim that is out of place in a survey. Please rephrase neutrally, e.g., 'Prior work [124] has explored...'.
  2. [References] Some references contain corrupted or malformed text: reference [45] has 'arXiv preprint arXiv:2024.05' instead of a valid arXiv ID, and references [190] and [216] contain stray LaTeX fragments ('@inproceedingscohen2018understanding...'). These should be cleaned.
  3. [Table 1] Table 1 uses symbols 'S!' and '%' extensively but never defines them. A legend or footnote explaining each symbol is needed.
  4. [Section 3.5] Typo: 'Exsting' should be 'Existing'. Also, Section 5.2 contains 'PoisedRAG' which should be 'PoisonedRAG' for consistency.
  5. [Section 8.5.1] The phrase 'For details, see Section 5 Robustness' is misleading because Section 5 addresses adversarial attacks on RAG, not watermark robustness against editing attacks. Either add a dedicated discussion or rephrase the cross-reference.

Circularity Check

0 steps flagged · score 2.0 of 10

No circular derivation: the survey's taxonomy is anchored in external standards and cited works, with only a minor non-load-bearing self-citation.

full rationale

This paper is a survey, so there is no fitted parameter or derived prediction whose value is forced by construction. The six-part trustworthiness taxonomy is taken from an external source: 'In 2022, the National Institute of Standards and Technology (NIST) published guidelines for trustworthy AI, defining trustworthiness from several perspectives [169]: Reliability, Privacy, Explainability, Fairness, Accountability, and Safety.' The section-level taxonomies (e.g., uncertainty vs. robust generalization, retrieval vs. generation privacy, targeted vs. jailbreak safety) are organizing categories imposed by the authors, not results derived from the cited papers. The only self-citation that affects content is in Section 7.4, where the authors call their own KG+LLM work pioneering: 'Our work pioneers the integration of knowledge graphs (KGs) and large language models (LLMs) to enhance retrieval faithfulness in Retrieval-Augmented Generation (RAG) systems [124].' This is a priority claim rather than a load-bearing premise: the survey's explainability taxonomy and its other sections do not depend on this claim, so it does not make the roadmap circular. Similarly, the weak reporting of the Section 2.4 paper-collection methodology (no query strings, screening counts, or coverage validation) is a completeness and recall limitation, not a self-referential derivation. Consequently, the central contribution is self-contained as a literature organization, and the circularity score is low.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

The survey contributes a taxonomy rather than a derivation; all load-bearing premises are organizational assumptions about completeness and category validity.

assumptions (3)
  • domain assumption The NIST six-pillar trustworthiness framework (reliability, privacy, fairness, explainability, accountability, safety) is a valid organizing lens for RAG systems.
    Adopted in Section 1 and used to structure all sections; no argument is given that these dimensions are complete for RAG.
  • domain assumption Reliability of RAG is adequately decomposed into uncertainty quantification and robust generalization.
    Section 3.1 explicitly drops adaptability based on prior work [176], without evaluating whether other reliability facets are needed.
  • domain assumption The keyword-search-based paper collection (Section 2.4) yields a representative and comprehensive set of trustworthy RAG papers.
    No screening counts, inclusion/exclusion criteria, or inter-rater reliability are reported; comprehensiveness is assumed.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Towards Trustworthy Retrieval Augmented Generation for Large Language Models: A Survey." pith.science (2026). https://pith.science/paper/TFCULT3Z

@misc{pith2026250206872,
  author       = {Pith},
  title        = {Pith review of: Towards Trustworthy Retrieval Augmented Generation for Large Language Models: A Survey},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/TFCULT3Z}},
  note         = {Machine review of arXiv:2502.06872}
}
read the original abstract

Retrieval-Augmented Generation (RAG) is an advanced technique designed to address the challenges of Artificial Intelligence-Generated Content (AIGC). By integrating context retrieval into content generation, RAG provides reliable and up-to-date external knowledge, reduces hallucinations, and ensures relevant context across a wide range of tasks. However, despite RAG's success and potential, recent studies have shown that the RAG paradigm also introduces new risks, including robustness issues, privacy concerns, adversarial attacks, and accountability issues. Addressing these risks is critical for future applications of RAG systems, as they directly impact their trustworthiness. Although various methods have been developed to improve the trustworthiness of RAG methods, there is a lack of a unified perspective and framework for research in this topic. Thus, in this paper, we aim to address this gap by providing a comprehensive roadmap for developing trustworthy RAG systems. We place our discussion around five key perspectives: reliability, privacy, safety, fairness, explainability, and accountability. For each perspective, we present a general framework and taxonomy, offering a structured approach to understanding the current challenges, evaluating existing solutions, and identifying promising future research directions. To encourage broader adoption and innovation, we also highlight the downstream applications where trustworthy RAG systems have a significant impact.

Figures

Figures reproduced from arXiv: 2502.06872 by the authors.

Figure 1
Figure 1. An overview of the key components and dimensions of Trustworthy Retrieval Augmented [PITH_FULL_IMAGE:figures/full_fig_p004_1.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 12 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. When Knowledge Changes: Metamorphic Testing of RAG Systems with Mutations

    cs.SE 2026-07 conditional novelty 6.5 of 10

    Eleven corpus mutations expose 4.9–10.2% metamorphic violations in RAG pipelines, with an oracle F1 of 0.927–1.000 versus at most 0.570 for RAGAS.

  2. HypRAG: Hyperbolic Dense Retrieval for Retrieval Augmented Generation

    cs.IR 2026-02 conditional novelty 6.0 of 10

    Hyperbolic embeddings with a geometry-aware pooling operator improve retrieval-augmented generation over Euclidean dense retrievers, with reported gains up to 29% on RAGBench.

  3. KERAG: Knowledge-Enhanced Retrieval-Augmented Generation for Advanced Question Answering

    cs.CL 2025-09 conditional novelty 6.0 of 10

    KERAG improves knowledge-graph question answering by retrieving broad entity neighborhoods instead of exact query paths and using a fine-tuned chain-of-thought summarizer.

  4. RAG Safety: Exploring Knowledge Poisoning Attacks to Retrieval-Augmented Generation

    cs.CR 2025-07 conditional novelty 6.0 of 10

    Inserting a small number of crafted triples into a knowledge graph can flip KG-RAG question answering toward attacker-chosen incorrect answers across four recent systems.

  5. MrM: Black-Box Membership Inference Attacks against Multimodal RAG Systems

    cs.CV 2025-06 conditional novelty 6.0 of 10

    MrM is a black-box membership inference attack on multimodal RAG systems that masks key objects in a target image and uses the system's ability to reconstruct them as a membership signal.

  6. MAP-Graph: Provenance-Aware Shared Memory for Multi-Agent Workflows

    cs.AI 2026-08 conditional novelty 5.0 of 10

    MAP-Graph combines permission filtering, ancestry-based trust scoring, and action-time gating for shared agent memory, reporting 94.96% task success, 72.70% exact accuracy, and 0% observed unauthorized access on a 2,7...

  7. The safety failures we are not instrumenting: a perspective on hidden safety-critical challenges in modern AI systems

    cs.CY 2026-07 conditional novelty 5.0 of 10

    AI safety should be measured by whether deployed systems keep errors visible, contestable, containable, and recoverable across five integrity layers, not only by whether individual model outputs look safe.

  8. Testing Retrieval-Augmented Generation Systems with Chunk Coverage

    cs.SE 2026-07 conditional novelty 5.0 of 10

    Chunk Coverage, a suite-level, oracle-independent measure of how much of a RAG corpus a test suite retrieves, speeds up coverage growth and earlier fault discovery in clinical and financial RAG systems.

  9. Healthier LLMs: Retrieval-Augmented Generation for Public Health Question Answering

    cs.CL 2026-07 conditional novelty 5.0 of 10

    Hybrid RAG over UK public health guidance sharply raises MCQA accuracy and free-form faithfulness, letting smaller open models match larger closed models without retrieval.

  10. A Comprehensive Survey in LLM(-Agent) Full Stack Safety: Data, Training and Deployment

    cs.CR 2025-04 conditional novelty 5.0 of 10

    A large collaborative survey organizes LLM and LLM-agent safety issues into a full-stack lifecycle framework from data preparation to deployment.

  11. Enhancing LLMs through human feedback: a journey towards self-improvement

    cs.IR 2026-07 unverdicted novelty 4.0 of 10

    An auxiliary feedback RAG continuously ingests classified human feedback to iteratively raise a primary RAG system’s answer accuracy and relevance.

  12. Effective Damage Data Generation by Fusing Imagery with Human Knowledge Using Vision-Language Models

    cs.CV 2025-08 reject novelty 4.0 of 10

    Prompting Gemini with damage-level definitions yields synthetic disaster imagery that a pre-trained classifier labels at F1 around 0.64, close to its score on real images, but the comparison lacks statistical support.

Reference graph

Works this paper leans on

244 extracted references · 11 canonical work pages · cited by 12 Pith papers

  1. [190]

    Angelina Wang, Alexander Liu, Ryan Zhang, Anat Kleiman, Leslie Kim, Dora Zhao, Iroha Shirai, Arvind Narayanan, and Olga Russakovsky. 2022. REVISE: A tool for measuring and mitigating bias in visual datasets. IJCV (2022)

  2. [216]

    Dalal, Jennifer L

    Cyril Zakka, Akash Chaurasia, Rohan Shad, Alex R. Dalal, Jennifer L. Kim, Michael Moor, Kevi@inproceedingscohen2018understanding, title=Understanding the representational power of neural retrieval models using NLP tasks, author=Cohen, Daniel and O’Connor, Brendan and Croft, W Bruce, booktitle=Proceedings of the 2018 ACM SIGIR International Conference on T...

  3. [1]

    Abdelrahman Abdallah, Bhaskar Piryani, and Adam Jatowt. 2023. Exploring the state of the art in legal QA systems. Journal of Big Data 10, 127 (2023). https://doi.org/10.1186/ s40537-023-00802-8

  4. [2]

    Sahar Abdelnabi and Mario Fritz. 2021. Adversarial watermarking transformer: Towards tracing text provenance with data hiding. In 2021 IEEE Symposium on Security and Privacy (SP). IEEE, 121–140

  5. [3]

    So, Jamie Hall, Noah Fiedel, Romal Thoppilan, Zi Yang, Apoorv Kulshreshtha, Gaurav Nemade, Yifeng Lu, and Quoc V

    Daniel Adiwardana, Minh-Thang Luong, David R. So, Jamie Hall, Noah Fiedel, Romal Thoppilan, Zi Yang, Apoorv Kulshreshtha, Gaurav Nemade, Yifeng Lu, and Quoc V . Le. 2020. Towards a Human-like Open-Domain Chatbot. InProceedings of the International Conference on Learning Representations (ICLR). arXiv:2001.09977 [cs.CL] https://arxiv.org/ abs/2001.09977

  6. [4]

    Wasi Ahmad, Jianfeng Chi, Yuan Tian, and Kai-Wei Chang. 2020. PolicyQA: A Reading Comprehension Dataset for Privacy Policies. InFindings of the Association for Computational Linguistics: EMNLP 2020, Trevor Cohn, Yulan He, and Yang Liu (Eds.). Association for Computational Linguistics, Online, 743–749. https://doi.org/10.18653/v1/2020. findings-emnlp.66

  7. [5]

    Microsoft Research AI4Science and Microsoft Azure Quantum. 2023. The Impact of Large Language Models on Scientific Discovery: a Preliminary Study using GPT-4. arXiv:2311.07361 [cs.CL] https://arxiv.org/abs/2311.07361

  8. [6]

    Rama Akkiraju, Anbang Xu, Deepak Bora, Tan Yu, Lu An, Vishal Seth, Aaditya Shukla, Pritam Gundecha, Hridhay Mehta, Ashwin Jha, Prithvi Raj, Abhinav Balasubramanian, Murali Maram, Guru Muthusamy, Shivakesh Reddy Annepally, Sidney Knowles, Min Du, Nick Burnett, Sean Javiya, Ashok Marannan, Mamta Kumari, Surbhi Jha, Ethan Dereszenski, Anupam Chakraborty, Sub...

Show all 244 references
  1. [7]

    M. H. Al-Rabeah and A. Lakizadeh. 2022. Prediction of drug-drug interaction events using graph neural networks based feature extraction. Scientific Reports 12 (2022), 15590. https: //doi.org/10.1038/s41598-022-19999-4 33

  2. [8]

    Nourah Alangari, Mohamed El Bachir Menai, Hassan Mathkour, and Ibrahim Almosallam

  3. [9]

    Mohammed Hazim Alkawaz, Ghazali Sulong, Tanzila Saba, Abdulaziz S Almazyad, and Am- jad Rehman. 2016. Concise analysis of current text automation and watermarking approaches. Security and Communication Networks 9, 18 (2016), 6365–6378

  4. [10]

    Avishek Anand, Procheta Sen, Sourav Saha, Manisha Verma, and Mandar Mitra. 2023. Explain- able information retrieval. In Proceedings of the 46th International ACM SIGIR Conference on Research and Development in Information Retrieval. 3448–3451

  5. [11]

    Angelopoulos, Stephen Bates, Emmanuel J

    Anastasios N. Angelopoulos, Stephen Bates, Emmanuel J. Candès, Michael I. Jordan, and Lihua Lei. 2022. Learn then Test: Calibrating Predictive Algorithms to Achieve Risk Control. arXiv:2110.01052 [cs.LG] https://arxiv.org/abs/2110.01052

  6. [12]

    Anthropic. 2024. Legal Summarization - Claude Use Case Guide. https://docs. anthropic.com/en/docs/about-claude/use-case-guides/legal-summarization Accessed: 2025-02-03

  7. [13]

    Alejandro Barredo Arrieta, Natalia Díaz-Rodríguez, Javier Del Ser, Adrien Bennetot, Si- ham Tabik, Alberto Barbado, Salvador García, Sergio Gil-López, Daniel Molina, Richard Benjamins, et al . 2020. Explainable Artificial Intelligence (XAI): Concepts, taxonomies, opportunities...

  8. [14]

    Mikhail J Atallah, Victor Raskin, Michael Crogan, Christian Hempelmann, Florian Ker- schbaum, Dina Mohamed, and Sanket Naik. 2001. Natural language watermarking: Design, analysis, and a proof-of-concept implementation. In Information Hiding: 4th International Workshop, IH 2001...

  9. [15]

    Payal Bajaj, Daniel Campos, Nick Craswell, Li Deng, Jianfeng Gao, Xiaodong Liu, Rangan Majumder, Andrew McNamara, Bhaskar Mitra, Tri Nguyen, Mir Rosenberg, Xia Song, Alina Stoica, Saurabh Tiwary, and Tong Wang. 2018. MS MARCO: A Human Generated MAchine Reading COmprehension Da...

  10. [16]

    Andrew Bell, Ian Solano-Kamaiko, Oded Nov, and Julia Stoyanovich. 2022. It’s just not that simple: an empirical study of the accuracy-explainability trade-off in machine learning for public policy. In Proceedings of the 2022 ACM conference on fairness, accountability, and tran...

  11. [17]

    Asma Ben Abacha, Yassine Mrabet, Mark Sharp, Travis R Goodwin, Sonya E Shooshan, and Dina Demner-Fushman. 2019. Bridging the Gap Between Consumers’ Medication Questions and Trusted Answers. Studies in Health Technology and Informatics 264 (August 21 2019), 25–29. https://doi.o...

  12. [18]

    Abeba Birhane and Vinay Uday Prabhu. 2021. Large image datasets: A pyrrhic win for computer vision?. In W ACV

  13. [19]

    Abeba Birhane, Vinay Uday Prabhu, and Emmanuel Kahembwe. 2021. Multimodal datasets: misogyny, pornography, and malignant stereotypes. arXiv preprint arXiv:2110.01963 (2021)

  14. [20]

    Nicholas Boucher, Ilia Shumailov, Ross Anderson, and Nicolas Papernot. 2022. Bad characters: Imperceptible nlp attacks. In 2022 IEEE Symposium on Security and Privacy (SP). IEEE, 1987–2004

  15. [21]

    CA Bowman and H Holzer. 2021. EMR Precharting Efficiency in Internal Medicine: A Scoping Review. J Med Educ Curric Dev 8 (Jul 2021), 23821205211032414. https: //doi.org/10.1177/23821205211032414

  16. [22]

    Jack T Brassil, Steven Low, Nicholas F Maxemchuk, and Lawrence O’Gorman. 1995. Elec- tronic marking and identification techniques to discourage document copying. IEEE Journal on Selected Areas in Communications 13, 8 (1995), 1495–1504. 34

  17. [23]

    Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-V oss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey ...

  18. [24]

    Gillian Cameron, David Cameron, Gavin Megaw, Raymond Bond, Maurice Mulvenna, Siobhan O’Neill, Cherie Armour, and Michael McTear. 2019. Assessing the Usability of a Chatbot for Mental Health Care. InInternet Science (Lecture Notes in Computer Science, V ol.11551), Tatiana Antip...

  19. [25]

    Centers for Disease Control and Prevention (CDC). 2023. Internet Use for Health Information and Communications Among Adults: United States, 2022. https://www.cdc.gov/nchs/ products/databriefs/db482.htm Accessed: 2025-01-29

  20. [26]

    Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde de Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, Alex Ray, Raul Puri, Gretchen Krueger, Michael Petrov, Heidy Khlaaf, Girish Sastry, Pamela Mishkin, Brooke Chan, Scott G...

  21. [27]

    Zhiyu Zoey Chen, Jing Ma, Xinlu Zhang, Nan Hao, An Yan, Armineh Nourbakhsh, Xianjun Yang, Julian McAuley, Linda Petzold, and William Yang Wang. 2024. A Survey on Large Language Models for Critical Societal Domains: Finance, Healthcare, and Law. arXiv preprint arXiv:2405.01769 ...

  22. [28]

    Jaekeol Choi, Euna Jung, Sungjun Lim, and Wonjong Rhee. 2022. Finding Inverse Document Frequency Information in BERT. arXiv preprint arXiv:2202.12191 (2022)

  23. [29]

    Miranda Christ, Sam Gunn, and Or Zamir. 2024. Undetectable watermarks for language models. In The Thirty Seventh Annual Conference on Learning Theory. PMLR, 1125–1139

  24. [30]

    Daniel Cohen, Brendan O’Connor, and W Bruce Croft. 2018. Understanding the represen- tational power of neural retrieval models using NLP tasks. In Proceedings of the 2018 ACM SIGIR International Conference on Theory of Information Retrieval. 67–74

  25. [31]

    Stav Cohen, Ron Bitton, and Ben Nassi. 2024. Unleashing Worms and Extracting Data: Escalating the Outcome of Attacks against RAG-based Inference in Scale and Severity Using Jailbreaking. arXiv preprint arXiv:2409.08045 (2024). https://doi.org/10.48550/ arXiv.2409.08045

  26. [32]

    Nick Craswell, Bhaskar Mitra, Emine Yilmaz, Daniel Campos, and Ellen M V oorhees. 2020. Overview of the TREC 2019 deep learning track. arXiv preprint arXiv:2003.07820 (2020)

  27. [33]

    Barnaby Crook, Maximilian Schlüter, and Timo Speith. 2023. Revisiting the performance- explainability trade-off in explainable artificial intelligence (XAI). In 2023 IEEE 31st International Requirements Engineering Conference Workshops (REW). IEEE, 316–324. 35

  28. [34]

    Florin Cuconasu, Giovanni Trappolini, Federico Siciliano, Simone Filice, Cesare Campagnano, Yoelle Maarek, Nicola Tonellotto, and Fabrizio Silvestri. 2024. The power of noise: Redefining retrieval for rag systems. In Proceedings of the 47th International ACM SIGIR Conference o...

  29. [35]

    Junyun Cui, Xiaoyu Shen, Feiping Nie, Zheng Wang, Jinglong Wang, and Yulong Chen

  30. [36]

    Eliot Dai, Tianhao Zhao, Hongfu Zhu, et al . 2024. A Comprehensive Survey on Trustworthy Graph Neural Networks: Privacy, Robustness, Fairness, and Explainability. Machine Intelligence Research 21 (2024), 1011–1061. https://doi.org/10.1007/ s11633-024-1510-8

  31. [37]

    Sagnik Dakshit. 2024. Faculty Perspectives on the Potential of RAG in Computer Science Higher Education. arXiv:2408.01462 [cs.CY] https://arxiv.org/abs/2408.01462

  32. [38]

    Boyi Deng, Wenjie Wang, Fengbin Zhu, Qifan Wang, and Fuli Feng. 2024. CrAM: Credibility- Aware Attention Modification in LLMs for Combating Misinformation in RAG.arXiv preprint arXiv:2406.11497 (2024)

  33. [39]

    Gelei Deng, Yi Liu, Kailong Wang, Yuekang Li, Tianwei Zhang, and Yang Liu. 2024. Pandora: Jailbreak GPTs by Retrieval Augmented Generation Poisoning. arXiv:2402.08416 [cs.CR] https://arxiv.org/abs/2402.08416

  34. [40]

    Finale Doshi-Velez, Mason Kortz, Ryan Budish, Chris Bavitz, Sam Gershman, David O’Brien, Kate Scott, Stuart Schieber, James Waldo, David Weinberger, et al. 2017. Accountability of AI under the law: The role of explanation. arXiv preprint arXiv:1711.01134 (2017)

  35. [41]

    Michael D Ekstrand, Graham McDonald, Amifa Raj, and Isaac Johnson. 2023. Overview of the TREC 2022 fair ranking track. arXiv preprint arXiv:2302.05558 (2023)

  36. [42]

    OpenAI et al. 2024. GPT-4 Technical Report. arXiv:2303.08774 [cs.CL] https://arxiv. org/abs/2303.08774

  37. [43]

    Jaiden Fairoze, Sanjam Garg, Somesh Jha, Saeed Mahloujifar, Mohammad Mahmoody, and Mingyuan Wang. 2023. Publicly-Detectable Watermarking for Language Models. Cryptology ePrint Archive, Paper 2023/1661. https://eprint.iacr.org/2023/1661

  38. [44]

    Wenqi Fan, Yujuan Ding, Liangbo Ning, Shijie Wang, Hengyun Li, Dawei Yin, Tat-Seng Chua, and Qing Li. 2024. A Survey On Rag Meeting LLMs: Towards retrieval-augmented large language models. In Proceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining....

  39. [45]

    Feiteng Fang, Yuelin Bai, Shiwen Ni, Min Yang, Xiaojun Chen, and Ruifeng Xu. 2024. Enhanc- ing Noise Robustness of Retrieval-Augmented Language Models with Adaptive Adversarial Training. arXiv preprint arXiv:2024.05 (2024). https://arxiv.org/abs/2024.05

  40. [46]

    Foulds, and Shimei Pan

    Philip Feldman, James R. Foulds, and Shimei Pan. 2024. RAGged Edges: The Double-Edged Sword of Retrieval-Augmented Chatbots. arXiv:2403.01193 [cs.CL] https://arxiv.org/ abs/2403.01193

  41. [47]

    Christiane Fellbaum. 1998. WordNet: An electronic lexical database. MIT Press google schola 2 (1998), 678–686

  42. [48]

    Qizhang Feng, Ninghao Liu, Fan Yang, Ruixiang Tang, Mengnan Du, and Xia Hu. 2023. Degree: Decomposition based explanation for graph neural networks. arXiv preprint arXiv:2305.12895 (2023)

  43. [49]

    Pierre Fernandez, Antoine Chaffin, Karim Tit, Vivien Chappelier, and Teddy Furon. 2023. Three bricks to consolidate watermarks for large language models. In 2023 IEEE International Workshop on Information Forensics and Security (WIFS). IEEE, 1–6. 36

  44. [50]

    Fernando Ferraretto, Thiago Laitz, Roberto Lotufo, and Rodrigo Nogueira. 2023. Exaranker: Explanation-augmented neural ranker. arXiv preprint arXiv:2301.10521 (2023)

  45. [51]

    Thibault Formal, Benjamin Piwowarski, and Stéphane Clinchant. 2021. A white box analysis of ColBERT. InAdvances in Information Retrieval: 43rd European Conference on IR Research, ECIR 2021, Virtual Event, March 28–April 1, 2021, Proceedings, Part II 43. Springer, 257– 263

  46. [52]

    Sorelle A Friedler, Carlos Scheidegger, Suresh Venkatasubramanian, Sonam Choudhary, Evan P Hamilton, and Derek Roth. 2019. A comparative study of fairness-enhancing interven- tions in machine learning. In Proceedings of the conference on fairness, accountability, and transparency

  47. [53]

    Yu Fu, Deyi Xiong, and Yue Dong. 2024. Watermarking conditional text generation for ai detection: Unveiling challenges and a semantic-aware watermark remedy. InProceedings of the AAAI Conference on Artificial Intelligence, V ol. 38. 18003–18011

  48. [54]

    Gallegos, Ryan A

    Isabel O. Gallegos, Ryan A. Rossi, Joe Barrow, Md Mehrab Tanjim, Sungchul Kim, Franck Dernoncourt, Tong Yu, Ruiyi Zhang, and Nesreen K. Ahmed. 2024. Bias and Fairness in Large Language Models: A Survey. Computational Linguistics 50, 3 (09 2024), 1097–1179. https://doi.org/10.1...

  49. [56]

    Ghaffari Laleh, D

    N. Ghaffari Laleh, D. Truhn, G. P. Veldhuizen, et al. 2022. Adversarial attacks and adversarial robustness in computational pathology.Nature Communications 13 (Sep 2022), 5711. https: //doi.org/10.1038/s41467-022-33266-0

  50. [57]

    Amirata Ghorbani, Abubakar Abid, and James Zou. 2019. Interpretation of neural networks is fragile. In Proceedings of the AAAI conference on artificial intelligence, V ol. 33. 3681–3688

  51. [58]

    Nicole Gillespie, Simon Lockey, Claire Curtis, Julian Pool, and Arian Akbari. 2023. Trust in Artificial Intelligence: A Global Study. Technical Report. The University of Queensland and KPMG Australia. https://doi.org/10.14264/00d3c94

  52. [59]

    Michael Glass, Gaetano Rossiello, Md Faisal Mahbub Chowdhury, Ankita Rajaram Naik, Pengshan Cai, and Alfio Gliozzo. 2022. Re2G: Retrieve, rerank, generate. arXiv preprint arXiv:2207.06300 (2022)

  53. [60]

    Kai Greshake, Sahar Abdelnabi, Shailesh Mishra, Christoph Endres, Thorsten Holz, and Mario Fritz. 2023. Not what you’ve signed up for: Compromising real-world LLM-integrated applications with indirect prompt injection. In Proceedings of the 16th ACM Workshop on Artificial Inte...

  54. [61]

    Chenchen Gu, Xiang Lisa Li, Percy Liang, and Tatsunori Hashimoto. 2023. On the learnability of watermarks for language models. arXiv preprint arXiv:2312.04469 (2023)

  55. [62]

    Chenchen Gu, Xiang Lisa Li, Percy Liang, and Tatsunori Hashimoto. 2024. On the Learnability of Watermarks for Language Models. In The Twelfth International Conference on Learning Representations. https://openreview.net/forum?id=9k0krNzvlV

  56. [63]

    Batu Guan, Yao Wan, Zhangqian Bi, Zheng Wang, Hongyu Zhang, Pan Zhou, and Lichao Sun. 2024. CodeIP: A Grammar-Guided Multi-Bit Watermark for Large Language Mod- els of Code. In Findings of the Association for Computational Linguistics: EMNLP 2024, Yaser Al-Onaizan, Mohit Bansa...

  57. [64]

    Seungju Han, Beomsu Kim, and Buru Chang. 2022. Measuring and Improving Seman- tic Diversity of Dialogue Generation. In Findings of the Association for Computational Linguistics: EMNLP 2022, Yoav Goldberg, Zornitsa Kozareva, and Yue Zhang (Eds.). Association for Computational L...

  58. [65]

    Moritz Hardt, Eric Price, and Nati Srebro. 2016. Equality of opportunity in supervised learning. Advances in neural information processing systems (2016)

  59. [66]

    Kai He, Rui Mao, Qika Lin, Yucheng Ruan, Xiang Lan, Mengling Feng, and Erik Cam- bria. 2025. A survey of large language models for healthcare: from data, technology, and applications to accountability and ethics. Information Fusion 118 (2025), 102963. https://doi.org/10.1016/j...

  60. [67]

    Wenchong He, Zhe Jiang, Tingsong Xiao, Zelin Xu, and Yukun Li. 2024. A Survey on Uncertainty Quantification Methods for Deep Learning. arXiv:2302.13425 [cs.LG] https: //arxiv.org/abs/2302.13425

  61. [68]

    Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. 2021. Measuring Massive Multitask Language Understanding. arXiv:2009.03300 [cs.CY] https://arxiv.org/abs/2009.03300

  62. [69]

    Giwon Hong, Jeonghwan Kim, Junmo Kang, Sung-Hyon Myaeng, and Joyce Jiyoung Whang

  63. [70]

    Abe Hou, Jingyu Zhang, Tianxing He, Yichen Wang, Yung-Sung Chuang, Hongwei Wang, Lingfeng Shen, Benjamin Van Durme, Daniel Khashabi, and Yulia Tsvetkov. 2024. SemStamp: A Semantic Watermark with Paraphrastic Robustness for Text Generation. In Proceedings of the 2024 Conference...

  64. [71]

    Abe Bohan Hou, Jingyu Zhang, Tianxing He, Yichen Wang, Yung-Sung Chuang, Hongwei Wang, Lingfeng Shen, Benjamin Van Durme, Daniel Khashabi, and Yulia Tsvetkov. 2023. Semstamp: A semantic watermark with paraphrastic robustness for text generation. arXiv preprint arXiv:2310.03991 (2023)

  65. [72]

    arXiv preprint arXiv:2305.01579 (2023)

    Why So Gullible? Enhancing the Robustness of Retrieval-Augmented Models against Counterfactual Noise. arXiv preprint arXiv:2305.01579 (2023)

  66. [73]

    Mengxuan Hu, Hongyi Wu, Zihan Guan, Ronghang Zhu, Dongliang Guo, Daiqing Qi, and Sheng Li. 2024. No Free Lunch: Retrieval-Augmented Generation Undermines Fairness in LLMs, Even for Vigilant Users. arXiv:2410.07589 [cs.IR] https://arxiv.org/abs/2410. 07589

  67. [74]

    Xiaowei Huang, Wenjie Ruan, Wei Huang, Gaojie Jin, Yi Dong, Changshun Wu, Saddek Bensalem, Ronghui Mu, Yi Qi, Xingyu Zhao, Kaiwen Cai, Yanghao Zhang, Sihao Wu, Peipei Xu, Dengyu Wu, Andre Freitas, and Mustafa A. Mustafa. 2023. A Survey of Safety and Trustworthiness of Large La...

  68. [75]

    Xinyi Hou, Yanjie Zhao, Yue Liu, Zhou Yang, Kailong Wang, Li Li, Xiapu Luo, David Lo, John Grundy, and Haoyu Wang. 2024. Large Language Models for Software Engineering: A Systematic Literature Review. ACM Trans. Softw. Eng. Methodol. 33, 8, Article 220 (Dec. 2024), 79 pages. h...

  69. [76]

    Muhammad Munwar Iqbal, Umair Khadam, Ki Jun Han, Jihun Han, and Sohail Jabbar. 2019. A robust digital watermarking algorithm for text document copyright protection based on feature coding. In 2019 15th International Wireless Communications & Mobile Computing Conference (IWCMC)...

  70. [77]

    Changyue Jiang, Xudong Pan, Geng Hong, Chenfu Bao, and Min Yang. 2024. RAG-Thief: Scalable Extraction of Private Data from Retrieval-Augmented Generation Applications with Agent-based Attacks. arXiv preprint arXiv:2411.14110 (2024). https://doi.org/10. 48550/arXiv.2411.14110 38

  71. [78]

    Mohamed Manzour Hussien, Angie Nataly Melo, Augusto Luis Ballardini, Carlota Salinas Maldonado, Rubén Izquierdo, and Miguel Ángel Sotelo. 2024. RAG-based Explainable Prediction of Road Users Behaviors for Automated Driving using Knowledge Graphs and Large Language Models. arXi...

  72. [79]

    Mandar Joshi, Eunsol Choi, Daniel Weld, and Luke Zettlemoyer. 2017. TriviaQA: A Large Scale Distantly Supervised Challenge Dataset for Reading Comprehension. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (V olume1: Long Papers), Reg...

  73. [80]

    Nikola Jovanovi´c, Robin Staab, Maximilian Baader, and Martin Vechev. 2024. Ward: Provable RAG Dataset Inference via LLM Watermarks. arXiv preprint arXiv:2410.03537 (2024)

  74. [81]

    Bowen Jin, Chulin Xie, Jiawei Zhang, Kashob Kumar Roy, Yu Zhang, Suhang Wang, Yu Meng, and Jiawei Han. 2024. Graph chain-of-thought: Augmenting large language models by reasoning on graphs. arXiv preprint arXiv:2404.07103 (2024)

  75. [82]

    To Eun Kim and Fernando Diaz. 2024. Towards Fair RAG: On the Impact of Fair Ranking in Retrieval-Augmented Generation. arXiv preprint arXiv:2409.11598 (2024)

  76. [83]

    John Kirchenbauer, Jonas Geiping, Yuxin Wen, Jonathan Katz, Ian Miers, and Tom Goldstein

  77. [84]

    Jaehyung Kim, Jaehyun Nam, Sangwoo Mo, Jongjin Park, Sang-Woo Lee, Minjoon Seo, Jung- Woo Ha, and Jinwoo Shin. 2024. SuRe: Summarizing Retrievals using Answer Candidates for Open-domain QA of LLMs. arXiv preprint arXiv:2404.13081 (2024)

  78. [85]

    John Kirchenbauer, Jonas Geiping, Yuxin Wen, Manli Shu, Khalid Saifullah, Kezhi Kong, Kasun Fernando, Aniruddha Saha, Micah Goldblum, and Tom Goldstein. 2024. On the Reli- ability of Watermarks for Large Language Models. InThe Twelfth International Conference on Learning Repre...

  79. [86]

    Bryan Klimt and Yiming Yang. 2004. The Enron Corpus: A New Dataset for Email Clas- sification Research. In Machine Learning: ECML 2004 (Lecture Notes in Computer Science, V ol.3201). Springer, 217–226. https://doi.org/10.1007/978-3-540-30115-8_22

  80. [87]

    In International Conference on Machine Learning

    A watermark for large language models. In International Conference on Machine Learning. PMLR, 17061–17084

  81. [88]

    John Kirchenbauer, Jonas Geiping, Yuxin Wen, Manli Shu, Khalid Saifullah, Kezhi Kong, Kasun Fernando, Aniruddha Saha, Micah Goldblum, and Tom Goldstein. 2023. On the reliability of watermarks for large language models. arXiv preprint arXiv:2306.04634 (2023)

  82. [89]

    Dai, Jakob Uszkoreit, Quoc Le, and Slav Petrov

    Tom Kwiatkowski, Jennimaria Palomaki, Olivia Redfield, Michael Collins, Ankur Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Jacob Devlin, Kenton Lee, Kristina Toutanova, Llion Jones, Matthew Kelcey, Ming-Wei Chang, Andrew M. Dai, Jakob Uszkoreit, Quoc Le, and Slav...

  83. [90]

    Jinqi Lai, Wensheng Gan, Jiayang Wu, Zhenlian Qi, and Philip S. Yu. 2023. Large Language Models in Law: A Survey. arXiv:2312.03718 [cs.CL] https://arxiv.org/abs/2312. 03718

  84. [91]

    Fanjie Kong, Shuai Yuan, Weituo Hao, and Ricardo Henao. 2023. Mitigating test-time bias for fair image retrieval. In Proceedings of the 37th International Conference on Neural Information Processing Systems (New Orleans, LA, USA)(NIPS ’23). Curran Associates Inc., Red Hook, NY...

  85. [92]

    Bhawesh Kumar, Charlie Lu, Gauri Gupta, Anil Palepu, David Bellamy, Ramesh Raskar, and Andrew Beam. 2023. Conformal Prediction with Large Language Models for Multi-Choice Question Answering. arXiv:2305.18404 [cs.CL] https://arxiv.org/abs/2305.18404

  86. [93]

    Alessandro Giaj Levra, Mauro Gatti, Roberto Mene, Dana Shiffer, Giorgio Costantino, Monica Solbiati, Raffaello Furlan, and Franca Dipaola. 2025. A large language model-based clinical decision support system for syncope recognition in the emergency department: A framework for c...

  87. [94]

    Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, et al. 2020. Retrieval- augmented generation for knowledge-intensive nlp tasks. Advances in Neural Information Processing ...

  88. [95]

    Gregory Kang Ruey Lau, Xinyuan Niu, Hieu Dao, Jiangwei Chen, Chuan-Sheng Foo, and Bryan Kian Hsiang Low. 2024. Waterfall: Framework for robust and scalable text watermark- ing. In ICML 2024 Workshop on Foundation Models in the Wild

  89. [96]

    Peter Lee, Sebastien Bubeck, and Joseph Petro. 2023. Benefits, Limits, and Risks of GPT-4 as an AI Chatbot for Medicine. New England Journal of Medicine 388, 13 (2023), 1233–1239. https://doi.org/10.1056/NEJMsr2214184 39

  90. [97]

    Jinming Li, Wentao Zhang, Tian Wang, Guanglei Xiong, Alan Lu, and Gerard Medioni. 2023. GPT4Rec: A Generative Framework for Personalized Recommendation and User Interests Interpretation. arXiv:2304.03879 [cs.IR] https://arxiv.org/abs/2304.03879

  91. [98]

    Shuo Li et al . 2023. TRAQ: Trustworthy Retrieval Augmented Question Answering via Conformal Prediction. arXiv preprint arXiv:2307.04642 (2023). https://arxiv.org/ abs/2307.04642

  92. [99]

    Bo Li, Peng Qi, Bo Liu, Shuai Di, Jingen Liu, Jiquan Pei, Jinfeng Yi, and Bowen Zhou

  93. [100]

    arXiv:2110.01167 [cs.AI] https: //arxiv.org/abs/2110.01167

    Trustworthy AI: From Principles to Practices. arXiv:2110.01167 [cs.AI] https: //arxiv.org/abs/2110.01167

  94. [101]

    Jierui Li, Lemao Liu, Huayang Li, Guanlin Li, Guoping Huang, and Shuming Shi. 2020. Eval- uating explanation methods for neural machine translation. arXiv preprint arXiv:2005.01672 (2020)

  95. [102]

    Jain, and Jiliang Tang

    Haochen Liu, Yiqi Wang, Wenqi Fan, Xiaorui Liu, Yaxin Li, Shaili Jain, Yunhao Liu, Anil K. Jain, and Jiliang Tang. 2021. Trustworthy AI: A Computational Perspective. arXiv:2107.06641 [cs.AI] https://arxiv.org/abs/2107.06641

  96. [103]

    Mingrui Liu, Sixiao Zhang, and Cheng Long. 2024. Mask-based Membership Inference Attacks for Retrieval-Augmented Generation. arXiv preprint arXiv:2410.20142 (2024). https://doi.org/10.48550/arXiv.2410.20142

  97. [104]

    Percy Liang, Rishi Bommasani, Tony Lee, Dimitris Tsipras, Dilara Soylu, Michihiro Yasunaga, Yian Zhang, Deepak Narayanan, Yuhuai Wu, Ananya Kumar, et al. 2022. Holistic evaluation of language models. arXiv preprint arXiv:2211.09110 (2022)

  98. [105]

    Chin-Yew Lin. 2004. Rouge: A package for automatic evaluation of summaries. In Text summarization branches out. 74–81

  99. [106]

    Aiwei Liu, Leyi Pan, Yijian Lu, Jingjing Li, Xuming Hu, Xi Zhang, Lijie Wen, Irwin King, Hui Xiong, and Philip Yu. 2024. A survey of text watermarking in the era of large language models. Comput. Surveys 57, 2 (2024), 1–36

  100. [107]

    Zheyuan Liu, Guangyao Dou, Zhaoxuan Tan, Yijun Tian, and Meng Jiang. 2024. Machine unlearning in generative ai: A survey. arXiv preprint arXiv:2407.20516 (2024)

  101. [108]

    Zheyuan Liu, Guangyao Dou, Zhaoxuan Tan, Yijun Tian, and Meng Jiang. 2024. Towards safer large language models through machine unlearning. arXiv preprint arXiv:2402.10058 (2024). 40

  102. [109]

    Shengjie Liu, Jing Wu, Jingyuan Bao, Wenyi Wang, Naira Hovakimyan, and Christo- pher G Healey. 2024. Towards a Robust Retrieval-Based Summarization System. arXiv:2403.19889 [cs.CL] https://arxiv.org/abs/2403.19889

  103. [110]

    Yang Liu et al. 2024. Trustworthy LLMs: A Survey and Guideline for Evaluating Large Language Models’ Alignment. arXiv preprint arXiv:2308.05374 (2024). https://arxiv. org/abs/2308.05374

  104. [111]

    Yi Liu, Gelei Deng, Yuekang Li, Kailong Wang, Tianwei Zhang, Yepang Liu, Haoyu Wang, Yan Zheng, and Yang Liu. 2023. Prompt injection attack against LLM-integrated applications. arXiv preprint arXiv:2306.05499 (2023)

  105. [112]

    Dong Lu, Tianyu Pang, Chao Du, Qian Liu, Xianjun Yang, and Min Lin. 2024. Test-time backdoor attacks on multimodal large language models. arXiv preprint arXiv:2402.08577 (2024)

  106. [113]

    Yijian Lu, Aiwei Liu, Dianzhi Yu, Jingjing Li, and Irwin King. 2024. An Entropy-based Text Watermarking Detection Method. arXiv preprint arXiv:2403.13485 (2024)

  107. [114]

    Zheyuan Liu, Chunhui Zhang, Yijun Tian, Erchi Zhang, Chao Huang, Yanfang Ye, and Chuxu Zhang. 2023. Fair graph representation learning via diverse mixture-of-experts. InProceedings of the ACM Web Conference 2023. 28–38

  108. [115]

    Quanyu Long, Yue Deng, LeiLei Gan, Wenya Wang, and Sinno Jialin Pan. 2024. Backdoor Attacks on Dense Passage Retrievers for Disseminating Misinformation. arXiv:2402.13532 [cs.CL] https://arxiv.org/abs/2402.13532

  109. [116]

    Antoine Louis, Gijs van Dijck, and Gerasimos Spanakis. 2023. Interpretable Long- Form Legal Question Answering with Retrieval-Augmented Large Language Models. arXiv:2309.17050 [cs.CL] https://arxiv.org/abs/2309.17050

  110. [117]

    J. Miao, C. Thongprayoon, S. Suppadungsuk, O. A. Garcia Valencia, and W. Cheungpasit- porn. 2024. Integrating Retrieval-Augmented Generation with Large Language Models in Nephrology: Advancing Practical Applications. Medicina (Kaunas) 60, 3 (Mar 2024), 445. https://doi.org/10....

  111. [118]

    Nighat Mir. 2014. Copyright for web content using invisible text watermarking. Computers in Human Behavior 30 (2014), 648–653

  112. [119]

    Linhao Luo, Yuan-Fang Li, Gholamreza Haffari, and Shirui Pan. 2024. Reasoning on Graphs: Faithful and Interpretable Large Language Model Reasoning. In International Conference on Learning Representations

  113. [120]

    Sean MacAvaney, Sergey Feldman, Nazli Goharian, Doug Downey, and Arman Cohan. 2022. ABNIRML: Analyzing the behavior of neural IR models. Transactions of the Association for Computational Linguistics 10 (2022), 224–239

  114. [121]

    Behrooz Mansouri and Ricardo Campos. 2023. FALQU: Finding Answers to Legal Questions. arXiv:2304.05611 [cs.IR] https://arxiv.org/abs/2304.05611

  115. [122]

    Meike Nauta, Jan Trienes, Shreyasi Pathak, Elisa Nguyen, Michelle Peters, Yasmin Schmitt, Jörg Schlötterer, Maurice Van Keulen, and Christin Seifert. 2023. From anecdotal evidence to quantitative evaluation methods: A systematic review on evaluating explainable ai. Comput. Sur...

  116. [123]

    Chee Ng and Yuen Fung. 2024. Educational Personalized Learning Path Planning with Large Language Models. arXiv:2407.11773 [cs.CL] https://arxiv.org/abs/2407.11773

  117. [124]

    Eric Mitchell, Yoonho Lee, Alexander Khazatsky, Christopher D Manning, and Chelsea Finn

  118. [125]

    In International Conference on Machine Learning

    Detectgpt: Zero-shot machine-generated text detection using probability curvature. In International Conference on Machine Learning. PMLR, 24950–24962

  119. [126]

    Piotr Molenda, Adian Liusie, and Mark JF Gales. 2024. WaterJudge: Quality-Detection Trade-off when Watermarking Large Language Models. arXiv preprint arXiv:2403.19548 (2024)

  120. [127]

    Elizabeth A Mullins, Adrian Portillo, Kristalys Ruiz-Rohena, and Aritran Piplai. 2024. En- hancing classroom teaching with LLMs and RAG. arXiv:2411.04341 [cs.LG] https: //arxiv.org/abs/2411.04341

  121. [128]

    Alicia Parrish, Angelica Chen, Nikita Nangia, Vishakh Padmakumar, Jason Phang, Jana Thompson, Phu Mon Htut, and Samuel R Bowman. 2021. BBQ: A hand-built bias benchmark for question answering. arXiv preprint arXiv:2110.08193 (2021)

  122. [129]

    Yuefeng Peng, Junda Wang, Hong Yu, and Amir Houmansadr. 2024. Data Extraction Attacks in Retrieval-Augmented Generation via Backdoors. arXiv preprint arXiv:2411.01705 (2024). https://doi.org/10.48550/arXiv.2411.01705

  123. [130]

    Bo Ni, Yu Wang, Lu Cheng, Erik Blasch, and Tyler Derr. 2025. Towards Trustworthy Knowledge Graph Reasoning: An Uncertainty Aware Perspective. In Proceedings of the AAAI Conference on Artificial Intelligence (AAAI). 41

  124. [131]

    S. Pal, M. Bhattacharya, M. A. Islam, and C. Chakraborty. 2023. ChatGPT or LLM in next- generation drug discovery and development: pharmaceutical and biotechnology companies can make use of the artificial intelligence-based device for a faster way of drug discovery and develop...

  125. [132]

    Leyi Pan, Aiwei Liu, Zhiwei He, Zitian Gao, Xuandong Zhao, Yijian Lu, Binglin Zhou, Shuliang Liu, Xuming Hu, Lijie Wen, et al. 2024. Markllm: An open-source toolkit for llm watermarking. arXiv preprint arXiv:2405.10051 (2024)

  126. [133]

    Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002. BLEU: a method for automatic evaluation of machine translation. In Proceedings of the 40th Annual Meeting on Association for Computational Linguistics (Philadelphia, Pennsylvania) (ACL ’02). Associa- tion for C...

  127. [134]

    Sayantan Polley. 2022. Towards Explainable Search in Legal Text. InEuropean Conference on Information Retrieval. Springer, 528–536

  128. [135]

    Lip Yee Por, KokSheik Wong, and Kok Onn Chee. 2012. UniSpaCh: A text-based data hiding method using Unicode space characters. Journal of Systems and Software 85, 5 (2012), 1075–1082

  129. [136]

    Mike Perkins. 2023. Academic Integrity considerations of AI Large Language Models in the post-pandemic era: ChatGPT and beyond. Journal of University Teaching and Learning Practice 20, 2 (2023)

  130. [137]

    Julien Piet, Chawin Sitawarin, Vivian Fang, Norman Mu, and David Wagner. 2023. Mark my words: Analyzing and evaluating language model watermarks. arXiv preprint arXiv:2312.00273 (2023)

  131. [138]

    Nicholas Pipitone and Ghita Houir Alami. 2024. LegalBench-RAG: A Benchmark for Retrieval- Augmented Generation in the Legal Domain. arXiv:2408.10343 [cs.AI] https://arxiv. org/abs/2408.10343

  132. [139]

    Gregory Plumb, Maruan Al-Shedivat, Ángel Alexander Cabrera, Adam Perer, Eric Xing, and Ameet Talwalkar. 2020. Regularizing black-box models for improved interpretability. Advances in Neural Information Processing Systems 33 (2020), 10526–10536

  133. [140]

    Navid Rekabsaz and Markus Schedl. 2020. Do neural ranking models intensify gender bias?. In Proceedings of the 43rd International ACM SIGIR Conference on Research and Development in Information Retrieval

  134. [141]

    Why should i trust you?

    Marco Tulio Ribeiro, Sameer Singh, and Carlos Guestrin. 2016. " Why should i trust you?" Explaining the predictions of any classifier. In Proceedings of the 22nd ACM SIGKDD international conference on knowledge discovery and data mining. 1135–1144

  135. [142]

    Yifan Qiao, Chenyan Xiong, Zhenghao Liu, and Zhiyuan Liu. 2019. Understanding the Behaviors of BERT in Ranking. arXiv preprint arXiv:1904.07531 (2019)

  136. [143]

    Jaakkola, and Regina Barzilay

    Victor Quach, Adam Fisch, Tal Schuster, Adam Yala, Jae Ho Sohn, Tommi S. Jaakkola, and Regina Barzilay. 2024. Conformal Language Modeling. arXiv:2306.10193 [cs.CL] https://arxiv.org/abs/2306.10193

  137. [144]

    P Rajpurkar. 2016. Squad: 100,000+ questions for machine comprehension of text. arXiv preprint arXiv:1606.05250 (2016)

  138. [145]

    Navid Rekabsaz, Simone Kopeinik, and Markus Schedl. 2021. Societal biases in retrieved contents: Measurement framework and adversarial mitigation of bert rankers. InProceedings of the 44th International ACM SIGIR Conference on Research and Development in Information Retrieval. 42

  139. [146]

    Gamble, Moein Shariatnia, and Bradley J

    Pouria Rouzrokh, Shahriar Faghani, Cooper U. Gamble, Moein Shariatnia, and Bradley J. Erickson. 2024. CONFLARE: CONFormal LArge language model REtrieval. arXiv preprint arXiv:2404.04287 (2024). https://arxiv.org/abs/2404.04287

  140. [147]

    Mark Roy, Baichuan Sun, Nihir Chadderwala, Derrick Choo, Mani Khanuja, and Frank Winkler. 2024. Use RAG for drug discovery with Amazon Bedrock Knowledge Bases. Amazon Web Services. https://aws.amazon.com/blogs/machine-learning/ use-rag-for-drug-discovery-with-amazon-bedrock-kn...

  141. [148]

    Stefano Giovanni Rizzo, Flavio Bertini, and Danilo Montesi. 2016. Content-preserving text wa- termarking through unicode homoglyph substitution. In Proceedings of the 20th International Database Engineering & Applications Symposium. 97–104

  142. [149]

    Marko Robnik-Šikonja and Marko Bohanec. 2018. Perturbation-based explanations of prediction models. Human and Machine Learning: Visible, Explainable, Trustworthy and Transparent (2018), 159–175

  143. [150]

    Rohani and C

    N. Rohani and C. Eslahchi. 2019. Drug-Drug Interaction Predicting by Neural Network Using Integrated Similarity. Scientific Reports 9 (2019), 13645. https://doi.org/10.1038/ s41598-019-50121-3

  144. [151]

    Joel Rorseth, Parke Godfrey, Lukasz Golab, Divesh Srivastava, and Jaroslaw Szlichta. 2024. RAGE Against the Machine: Retrieval-Augmented LLM Explanations. arXiv preprint arXiv:2405.13000 (2024)

  145. [152]

    Ramprasaath R Selvaraju, Michael Cogswell, Abhishek Das, Ramakrishna Vedantam, Devi Parikh, and Dhruv Batra. 2020. Grad-CAM: visual explanations from deep networks via gradient-based localization. International journal of computer vision 128 (2020), 336–359

  146. [153]

    Rajendra Acharya

    Silvia Seoni, Vicnesh Jahmunah, Massimo Salvi, Prabal Datta Barua, Filippo Molinari, and U. Rajendra Acharya. 2023. Application of uncertainty quantification to artificial intelligence in healthcare: A review of last decade (2013–2023). Computers in Biology and Medicine 165 (2...

  147. [154]

    Cynthia Rudin, Chaofan Chen, Zhi Chen, Haiyang Huang, Lesia Semenova, and Chudi Zhong

  148. [155]

    Statistic Surveys 16 (2022), 1–85

    Interpretable machine learning: Fundamental principles and 10 grand challenges. Statistic Surveys 16 (2022), 1–85

  149. [156]

    Alireza Salemi, Sheshera Mysore, Michael Bendersky, and Hamed Zamani. 2023. Lamp: When large language models meet personalization. arXiv preprint arXiv:2304.11406 (2023)

  150. [157]

    Ryoma Sato, Yuki Takezawa, Han Bao, Kenta Niwa, and Makoto Yamada. 2023. Embarrass- ingly simple text watermarks. arXiv preprint arXiv:2310.08920 (2023)

  151. [158]

    Johannes Schneider. 2024. Explainable Generative AI (GenXAI): a survey, conceptualization, and research agenda. Artificial Intelligence Review 57, 11 (2024), 289

  152. [159]

    Robik Shrestha, Yang Zou, Qiuyu Chen, Zhiheng Li, Yusheng Xie, and Siqi Deng. 2024. FairRAG: Fair human generation via fair retrieval augmentation. In Computer Vision and Pattern Recognition

  153. [160]

    Jaspreet Singh, Megha Khosla, Wang Zhenye, and Avishek Anand. 2021. Extracting per query valid explanations for blackbox learning-to-rank models. In Proceedings of the 2021 ACM SIGIR International Conference on Theory of Information Retrieval. 203–210

  154. [161]

    Glenn Shafer and Vladimir V ovk. 2007. A tutorial on conformal prediction. arXiv:0706.3188 [cs.LG] https://arxiv.org/abs/0706.3188

  155. [162]

    Erfan Shayegani, Md Abdullah Al Mamun, Yu Fu, Pedram Zaree, Yue Dong, and Nael Abu- Ghazaleh. 2023. Survey of Vulnerabilities in Large Language Models Revealed by Adversarial Attacks. arXiv:2310.10844 [cs.CL] https://arxiv.org/abs/2310.10844 43

  156. [163]

    Nice try, kiddo

    Emily Sheng, Kai-Wei Chang, Prem Natarajan, and Nanyun Peng. 2021. "Nice try, kiddo": Investigating Ad Hominems in Dialogue Responses. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technolog...

  157. [164]

    Emily Sheng, Kai-Wei Chang, Prem Natarajan, and Nanyun Peng. 2021. Societal Biases in Language Generation: Progress and Challenges. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural...

  158. [165]

    Pengfei He Yue Xing Yiding Liu Han Xu Jie Ren Shuaiqiang Wang Dawei Yin Yi Chang Jiliang Tang Shenglai Zeng, Jiankun Zhang. 2024. The Good and The Bad: Exploring Privacy Issues in Retrieval-Augmented Generation (RAG). ACL (2024). https://arxiv.org/abs/ 2402.16893

  159. [166]

    Jiashuo Sun, Chengjin Xu, Lumingyuan Tang, Saizhuo Wang, Chen Lin, Yeyun Gong, Heung- Yeung Shum, and Jian Guo. 2023. Think-on-Graph: Deep and Responsible Reasoning of Large Language Model with Knowledge Graph. In International Conference on Learning Representations. arXiv:230...

  160. [167]

    Zhensu Sun, Xiaoning Du, Fu Song, and Li Li. 2023. Codemark: Imperceptible watermarking for code datasets against neural code completion models. InProceedings of the 31st ACM Joint European Software Engineering Conference and Symposium on the Foundations of Software Engineerin...

  161. [168]

    Singhal, S

    K. Singhal, S. Azizi, T. Tu, et al. 2023. Large language models encode clinical knowledge. Nature 620 (Aug 2023), 172–180. https://doi.org/10.1038/s41586-023-06291-2

  162. [169]

    Sara Mahdavi, Joelle Barral, Dale Webster, Greg S

    Karan Singhal, Tao Tu, Juraj Gottweis, Rory Sayres, Ellery Wulczyn, Le Hou, Kevin Clark, Stephen Pfohl, Heather Cole-Lewis, Darlene Neal, Mike Schaekermann, Amy Wang, Mohamed Amin, Sami Lachgar, Philip Mansfield, Sushant Prakash, Bradley Green, Ewa Dominowska, Blaise Aguera y ...

  163. [170]

    Jiayuan Su, Jing Luo, Hongwei Wang, and Lu Cheng. 2024. API Is Enough: Conformal Prediction for Large Language Models Without Logit-Access. arXiv:2403.01216 [cs.CL] https://arxiv.org/abs/2403.01216

  164. [171]

    Viju Sudhi, Sinchana Ramakanth Bhat, Max Rudat, and Roman Teucher. 2024. RAG-Ex: A Generic Framework for Explaining Retrieval Augmented Generation. In Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval. 2776–2780

  165. [172]

    Ryuichi Sumida, Koji Inoue, and Tatsuya Kawahara. 2024. Should RAG Chatbots Forget Unimportant Conversations? Exploring Importance and Forgetting with Psychological Insights. arXiv:2409.12524 [cs.CL] https://arxiv.org/abs/2409.12524

  166. [173]

    Mercan Topkara, Umut Topkara, and Mikhail J Atallah. 2006. Words are not enough: sentence level natural language watermarking. In Proceedings of the 4th ACM international workshop on Contents protection and security. 37–46

  167. [174]

    Umut Topkara, Mercan Topkara, and Mikhail J Atallah. 2006. The hiding virtues of ambiguity: quantifiably resilient watermarking of natural language text through synonym substitutions. In Proceedings of the 8th workshop on Multimedia and security. 164–174

  168. [175]

    Zhensu Sun, Xiaoning Du, Fu Song, Mingze Ni, and Li Li. 2022. Coprotector: Protect open-source code against unauthorized training usage with data poisoning. In Proceedings of the ACM Web Conference 2022. 652–660. 44

  169. [176]

    Elham Tabassi. 2022. Trustworthy AI: Managing the Risks of Ar- tificial Intelligence. https://www.nist.gov/speech-testimony/ trustworthy-ai-managing-risks-artificial-intelligence Accessed: 2024- 07-01

  170. [177]

    Alon Talmor and Jonathan Berant. 2018. The web as a knowledge-base for answering complex questions. arXiv preprint arXiv:1803.06643 (2018)

  171. [178]

    Sule Tekkesinoglu and Lars Kunze. 2024. From Feature Importance to Natural Language Explanations Using LLMs with RAG. arXiv preprint arXiv:2407.20990 (2024)

  172. [180]

    Michael Völske, Alexander Bondarenko, Maik Fröbe, Benno Stein, Jaspreet Singh, Matthias Hagen, and Avishek Anand. 2021. Towards axiomatic explanations for neural ranking models. In Proceedings of the 2021 ACM SIGIR International Conference on Theory of Information Retrieval. 13–22

  173. [181]

    Yixin Wan, George Pu, Jiao Sun, Aparna Garimella, Kai-Wei Chang, and Nanyun Peng

  174. [182]

    Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, Aure- lien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lample. 2023. LLaMA: Open and Efficient Foundation ...

  175. [183]

    Dusenberry, Du Phan, Mark Collier, Jie Ren, Kehang Han, Zi Wang, Zelda Mariet, Huiyi Hu, Neil Band, Tim G

    Dustin Tran, Jeremiah Liu, Michael W. Dusenberry, Du Phan, Mark Collier, Jie Ren, Kehang Han, Zi Wang, Zelda Mariet, Huiyi Hu, Neil Band, Tim G. J. Rudner, Karan Singhal, Zachary Nado, Joost van Amersfoort, Andreas Kirsch, Rodolphe Jenatton, Nithum Thain, Honglin Yuan, Kelly B...

  176. [184]

    Shangqing Tu, Yuliang Sun, Yushi Bai, Jifan Yu, Lei Hou, and Juanzi Li. 2024. WaterBench: Towards Holistic Evaluation of Watermarks for Large Language Models. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (V olume1: Long Papers), Lu...

  177. [185]

    Manisha Verma and Debasis Ganguly. 2019. LIRME: locally interpretable ranking model explanation. In Proceedings of the 42nd International ACM SIGIR Conference on Research and Development in Information Retrieval. 1281–1284

  178. [186]

    Giulia Vilone and Luca Longo. 2021. Notions of explainability and evaluation approaches for explainable artificial intelligence. Information Fusion 76 (2021), 89–106

  179. [187]

    Ziqiu Wang, Jun Liu, Shengkai Zhang, and Yang Yang. 2024. Poisoned LangChain: Jailbreak LLMs by LangChain. arXiv:2406.18122 [cs.CL] https://arxiv.org/abs/2406.18122

  180. [188]

    Pierre Wargnier, Samuel Benveniste, Pierre Jouvelot, and Anne-Sophie Rigaud. 2018. Usability Assessment of Interaction Management Support in LOUISE, an ECA-based User Interface for Elders with Cognitive Impairment. Technology and Disability 30, 3 (2018), 105–126. https://doi.o...

  181. [189]

    Kelly is a Warm Person, Joseph is a Role Model

    "Kelly is a Warm Person, Joseph is a Role Model": Gender Biases in LLM-Generated Reference Letters. arXiv preprint arXiv:2310.09219 (2023). https://doi.org/10.48550/ arXiv.2310.09219 Accepted to EMNLP 2023 Findings. 45

  182. [191]

    Truong, Simran Arora, Manias Mazeika, Dan Hendrycks, Zinan Lin, Yu Cheng, Sanmi Koyejo, Dawn Song, and Bo Li

    Boxin Wang, Weixin Chen, Hengzhi Pei, Chulin Xie, Mintong Kang, Chenhui Zhang, Chejian Xu, Zidi Xiong, Ritik Dutta, Rylan Schaeffer, Sang T. Truong, Simran Arora, Manias Mazeika, Dan Hendrycks, Zinan Lin, Yu Cheng, Sanmi Koyejo, Dawn Song, and Bo Li. 2023. Decod- ingTrust: a c...

  183. [192]

    Yu, and Qingsong Wen

    Shen Wang, Tianlong Xu, Hang Li, Chaoli Zhang, Joleen Liang, Jiliang Tang, Philip S. Yu, and Qingsong Wen. 2024. Large Language Models for Education: A Survey and Outlook. arXiv:2403.18105 [cs.CL] https://arxiv.org/abs/2403.18105

  184. [193]

    Yu Wang, Nedim Lipka, Ryan A Rossi, Alexa Siu, Ruiyi Zhang, and Tyler Derr. 2024. Knowledge graph prompting for multi-document question answering. In Proceedings of the AAAI Conference on Artificial Intelligence, V ol. 38. 19206–19214

  185. [194]

    Yuan Wang, Xuyang Wu, Hsin-Tai Wu, Zhiqiang Tao, and Yi Fang. 2024. Do Large Language Models Rank Fairly? An Empirical Study on the Fairness of LLMs as Rankers. arXiv preprint arXiv:2404.03192 (2024)

  186. [195]

    Guangzhi Xiong, Qiao Jin, Zhiyong Lu, and Aidong Zhang. 2024. Benchmarking Retrieval- Augmented Generation for Medicine. arXiv:2402.13178 [cs.CL] https://arxiv.org/ abs/2402.13178

  187. [196]

    Hengyuan Xu, Liyao Xiang, Xingjun Ma, Borui Yang, and Baochun Li. 2024. Hufu: A Modality-Agnositc Watermarking System for Pre-Trained Transformers via Permutation Equivariance. arXiv preprint arXiv:2403.05842 (2024)

  188. [197]

    Alexander Wei, Nika Haghtalab, and Jacob Steinhardt. 2023. Jailbroken: how does LLM safety training fail?. In Proceedings of the 37th International Conference on Neural Information Processing Systems (New Orleans, LA, USA) (NIPS ’23). Curran Associates Inc., Red Hook, NY , USA...

  189. [199]

    Caesar Wu, Yuan-Fang Lib, and Pascal Bouvry. 2023. Survey of Trustworthy AI: A Meta Decision of AI. arXiv:2306.00380 [cs.AI] https://arxiv.org/abs/2306.00380

  190. [200]

    Xuyang Wu, Shuowei Li, Hsin-Tai Wu, Zhiqiang Tao, and Yi Fang. 2024. Does RAG Introduce Unfairness in LLMs? Evaluating Fairness in Retrieval-Augmented Generation Systems. arXiv preprint arXiv:2409.19804 (2024)

  191. [201]

    Yiquan Wu, Siying Zhou, Yifei Liu, Weiming Lu, Xiaozhong Liu, Yating Zhang, Changlong Sun, Fei Wu, and Kun Kuang. 2023. Precedent-Enhanced Legal Judgment Prediction with LLM and Domain-Model Collaboration. In Proceedings of the 2023 Conference on Empirical Methods in Natural L...

  192. [202]

    Chong Xiang, Tong Wu, Zexuan Zhong, David Wagner, Danqi Chen, and Prateek Mittal

  193. [203]

    Haoyan Yang, Zhitao Li, Yong Zhang, Jianzong Wang, Ning Cheng, Ming Li, and Jing Xiao

  194. [204]

    R. Yang, Y . Ning, E. Keppo, et al . 2025. Retrieval-augmented generation for generative artificial intelligence in health care. npj Health Systems 2 (2025), 2. https://doi.org/10. 1038/s44401-024-00004-1

  195. [205]

    Xi Yang, Jie Zhang, Kejiang Chen, Weiming Zhang, Zehua Ma, Feng Wang, and Nenghai Yu

  196. [206]

    Jing Xu, Arthur Szlam, and Jason Weston. 2022. Beyond Goldfish Memory: Long-Term Open-Domain Conversation. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (V olume1: Long Papers). Association for Computational Lin- guistics, 1163–1180...

  197. [207]

    Shicheng Xu, Liang Pang, Huawei Shen, and Xueqi Cheng. 2024. A Theory for Token-Level Harmonization in Retrieval-Augmented Generation. arXiv:2406.00944 [cs.CL] https: //arxiv.org/abs/2406.00944

  198. [208]

    Xiaojun Xu, Yuanshun Yao, and Yang Liu. 2024. Learning to Watermark LLM-generated Text via Reinforcement Learning. arXiv:2403.10553 [cs.LG] https://arxiv.org/abs/2403. 10553

  199. [209]

    Jiaqi Xue, Mengxin Zheng, Yebowen Hu, Fei Liu, Xun Chen, and Qian Lou. 2024. BadRAG: Identifying Vulnerabilities in Retrieval Augmented Generation of Large Language Models. arXiv preprint arXiv:2406.00083 (2024)

  200. [210]

    Jiaqi Xue, Mengxin Zheng, Ting Hua, Yilin Shen, Yepeng Liu, Ladislau Bölöni, and Qian Lou

  201. [211]

    Advances in Neural Information Processing Systems 36 (2024)

    TrojLLM: A black-box Trojan prompt attack on large language models. Advances in Neural Information Processing Systems 36 (2024)

  202. [212]

    Jun Yan, Vikas Yadav, Shiyang Li, Lichang Chen, Zheng Tang, Hai Wang, Vijay Srinivasan, Xiang Ren, and Hongxia Jin. 2023. Backdooring instruction-tuned large language models with virtual prompt injection. arXiv preprint arXiv:2307.16888 (2023)

  203. [213]

    KiYoon Yoo, Wonhyuk Ahn, and Nojun Kwak. 2024. Advancing Beyond Identification: Multi-bit Watermark for Large Language Models. In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (V o...

  204. [214]

    arXiv preprint arXiv:2310.18347 (2023)

    Prca: Fitting black-box large language models for retrieval question answering via pluggable reward-driven contextual adapter. arXiv preprint arXiv:2310.18347 (2023)

  205. [215]

    Zhenrui Yue, Huimin Zeng, Yimeng Lu, Lanyu Shang, Yang Zhang, and Dong Wang. 2024. Evidence-Driven Retrieval Augmented Response Generation for Online Misinformation. In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Lingui...

  206. [217]

    In Proceedings of the AAAI Conference on Artificial Intelligence, V ol

    Tracing text provenance via context-aware lexical substitution. In Proceedings of the AAAI Conference on Artificial Intelligence, V ol. 36. 11613–11621

  207. [218]

    Yifan Yao, Jinhao Duan, Kaidi Xu, Yuanfang Cai, Zhibo Sun, and Yue Zhang. 2024. A survey on large language model (LLM) security and privacy: The Good, The Bad, and The Ugly. High-Confidence Computing 4, 2 (June 2024), 100211. https://doi.org/10.1016/j. hcc.2024.100211

  208. [219]

    Wong, Emine Yilmaz, Shuming Shi, and Zhaopeng Tu

    Fanghua Ye, Mingming Yang, Jianhui Pang, Longyue Wang, Derek F. Wong, Emine Yilmaz, Shuming Shi, and Zhaopeng Tu. 2024. Benchmarking LLMs via Uncertainty Quantification. arXiv:2401.12794 [cs.CL] https://arxiv.org/abs/2401.12794

  209. [220]

    Antonio Jimeno Yepes, Yao You, Jan Milczek, Sebastian Laverde, and Renyu Li

  210. [221]

    arXiv:2402.05131 [cs.CL] https://arxiv.org/abs/2402.05131 47

    Financial Report Chunking for Effective Retrieval Augmented Generation. arXiv:2402.05131 [cs.CL] https://arxiv.org/abs/2402.05131 47

  211. [222]

    Wen-tau Yih, Matthew Richardson, Christopher Meek, Ming-Wei Chang, and Jina Suh

  212. [223]

    Yue Zhang, Yafu Li, Leyang Cui, Deng Cai, Lemao Liu, Tingchen Fu, Xinting Huang, Enbo Zhao, Yu Zhang, Yulong Chen, Longyue Wang, Anh Tuan Luu, Wei Bi, Freda Shi, and Shuming Shi. 2023. Siren’s Song in the AI Ocean: A Survey on Hallucination in Large Language Models. arXiv prep...

  213. [224]

    Zhitao Ying, Dylan Bourgeois, Jiaxuan You, Marinka Zitnik, and Jure Leskovec. 2019. Gnnex- plainer: Generating explanations for graph neural networks. Advances in neural information processing systems 32 (2019)

  214. [225]

    Kenji Yokotani, Gen Takagi, and Kobun Wakashima. 2018. Advantages of Virtual Agents over Clinical Psychologists during Comprehensive Mental Health Interviews Using a Mixed Methods Design. Computers in Human Behavior 85 (2018), 135–145. https://doi.org/ 10.1016/j.chb.2018.03.045

  215. [226]

    KiYoon Yoo, Wonhyuk Ahn, Jiho Jang, and Nojun Kwak. 2023. Robust Multi-bit Natural Lan- guage Watermarking through Invariant Features. InProceedings of the 61st Annual Meeting of the Association for Computational Linguistics (V olume1: Long Papers), Anna Rogers, Jordan Boyd-Gr...

  216. [227]

    Zongmeng Zhang, Yufeng Shi, Jinhua Zhu, Wengang Zhou, Xiang Qi, Peng Zhang, and Houqiang Li. 2024. Trustworthy alignment of retrieval-augmented large language models via reinforcement learning. In Proceedings of the 41st International Conference on Machine Learning (Vienna, Au...

  217. [228]

    Ori Yoran, Tomer Wolfson, Ori Ram, and Jonathan Berant. 2024. Making Retrieval-Augmented Language Models Robust to Irrelevant Context. arXiv:2310.01558 [cs.CL] https://arxiv. org/abs/2310.01558

  218. [229]

    Zexuan Zhong, Ziqing Huang, Alexander Wettig, and Danqi Chen. 2023. Poisoning Retrieval Corpora by Injecting Adversarial Passages. arXiv:2310.19156 [cs.CL] https://arxiv. org/abs/2310.19156

  219. [230]

    Yujia Zhou, Yan Liu, Xiaoxi Li, Jiajie Jin, Hongjin Qian, Zheng Liu, Chaozhuo Li, Zhicheng Dou, Tsung-Yi Ho, and Philip S Yu. 2024. Trustworthiness in Retrieval-Augmented Generation Systems: A Survey. arXiv preprint arXiv:2409.10102 (2024)

  220. [231]

    arXiv preprint arXiv:2312.10997 (2023)

    Almanac: Retrieval-Augmented Language Models for Clinical Medicine. arXiv preprint arXiv:2312.10997 (2023)

  221. [232]

    Rowan Zellers, Ari Holtzman, Yonatan Bisk, Ali Farhadi, and Yejin Choi. 2019. HellaSwag: Can a Machine Really Finish Your Sentence? ACL 2019 (2019). https://doi.org/10. 48550/arXiv.1905.07830 arXiv:1905.07830 [cs.CL]

  222. [233]

    Shenglai Zeng, Jiankun Zhang, Pengfei He, Jie Ren, Tianqi Zheng, Hanqing Lu, Han Xu, Hui Liu, Yue Xing, and Jiliang Tang. 2024. Mitigating the Privacy Issues in Retrieval- Augmented Generation (RAG) via Pure Synthetic Data. arXiv preprint arXiv:2406.14773 (2024). https://doi.o...

  223. [234]

    He Zhang, Bang Wu, Xingliang Yuan, Shirui Pan, Hanghang Tong, and Jian Pei. 2024. Trustworthy Graph Neural Networks: Aspects, Methods, and Trends. arXiv preprint arXiv:2205.07424 (2024). https://arxiv.org/pdf/2205.07424 48

  224. [235]

    Qin Zhang, Shangsi Chen, Dongkuan Xu, Qingqing Cao, Xiaojun Chen, Trevor Cohn, and Meng Fang. 2023. A Survey for Efficient Open Domain Question Answering. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (V olume1: Long Papers), Anna R...

  225. [236]

    Ruisi Zhang, Shehzeen Samarah Hussain, Paarth Neekhara, and Farinaz Koushanfar. 2024. REMARK-LLM: A robust and efficient watermarking framework for generative large language models. In 33rd USENIX Security Symposium (USENIX Security 24). 1813–1830

  226. [237]

    Yongfeng Zhang, Xu Chen, et al . 2020. Explainable recommendation: A survey and new perspectives. Foundations and Trends® in Information Retrieval 14, 1 (2020), 1–101

  227. [239]

    Yongfeng Zhang, Jiaxin Mao, and Qingyao Ai. 2019. Www’19 tutorial on explainable recom- mendation and search. In Companion Proceedings of The 2019 World Wide Web Conference. 1330–1331

  228. [240]

    Yongfeng Zhang, Yi Zhang, and Min Zhang. 2018. SIGIR 2018 workshop on explainable recommendation and search (EARS 2018). In The 41st International ACM SIGIR Conference on Research & Development in Information Retrieval. 1411–1413

  229. [241]

    Zijian Zhang, Koustav Rudra, and Avishek Anand. 2021. Explain and predict, and then predict again. In Proceedings of the 14th ACM international conference on web search and data mining. 418–426

  230. [243]

    Wayne Xin Zhao, Jing Liu, Ruiyang Ren, and Ji-Rong Wen. 2024. Dense Text Retrieval Based on Pretrained Language Models: A Survey. ACM Trans. Inf. Syst. 42, 4, Article 89 (Feb. 2024), 60 pages. https://doi.org/10.1145/3637870

  231. [246]

    Andy Zou, Zifan Wang, J Zico Kolter, and Matt Fredrikson. 2023. Universal and transferable adversarial attacks on aligned language models. arXiv preprint arXiv:2307.15043 (2023)

  232. [247]

    Wei Zou, Runpeng Geng, Binghui Wang, and Jinyuan Jia. 2024. PoisonedRAG: Knowl- edge Corruption Attacks to Retrieval-Augmented Generation of Large Language Models. arXiv:2402.07867 [cs.CR] https://arxiv.org/abs/2402.07867 49

  233. [2016]

    In Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics

    The Value of Semantic Parse Labeling for Knowledge Base Question Answering. In Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics. Association for Computational Linguistics, Berlin, Germany, 201–206

  234. [2022]

    arXiv:2204.04859 [cs.CL] https://arxiv.org/abs/2204.04859

    A Survey on Legal Judgment Prediction: Datasets, Metrics, Models and Challenges. arXiv:2204.04859 [cs.CL] https://arxiv.org/abs/2204.04859

  235. [2023]

    Exploring evaluation methods for interpretable machine learning: A survey.Information 14, 8 (2023), 469

  236. [2024]

    arXiv:2405.15556 [cs.LG] https://arxiv.org/abs/2405.15556

    Certifiably Robust RAG against Retrieval Corruption. arXiv:2405.15556 [cs.LG] https://arxiv.org/abs/2405.15556

Pith tools

Reviewed August 8, 2026 · model on record in the stance chip above.