REVIEW 4 major objections 6 minor 1 cited by
MIRAGE: Defending Long-Form RAG Against Misinformation Pollution
T0 review · 4 major / 6 minor · reviewed 2026-07-11 · grok-4.5
Pith's one-line read Cross-document claim consistency can filter or block polluted RAG evidence before generation, restoring long-form factuality.
desk verdict Solid systems paper: training-free claim-graph defense plus a reusable pollution protocol, with large multi-model gains that partly rest on synthetic edits and fallback blocking. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The Defended-Claims Gate over an NLI-based cross-document claim graph: claims become nodes with support and contradiction edges; a consistent multi-source subset is selected; retrieval is trusted only when the contradiction ratio is low and source diversity is high, otherwise generation is blocked and the model answers parametrically.
What would settle it
Construct a mixed or fully polluted retrieval set in which every polluted claim is lexically anchored, multi-source consistent, and not contradicted by counter-evidence probes, then check whether MIRAGE still passes the gate and whether VeriScore F1 falls back toward vanilla polluted RAG rather than No-RAG or clean levels.
Extended reading notes
Core claim
The paper claims that truthful evidence for long-form answers tends to be corroborated across independent sources, whereas polluted evidence induces structural inconsistency that an NLI claim graph can detect. By scoring claims for multi-source support, pruning contradictions, and gating generation on global consistency and source diversity, MIRAGE can either generate from a defended subset of verified claims or fall back to no-retrieval answering, restoring factuality under mixed and fully polluted retrieval while staying close to clean RAG when evidence is coherent.
Load-bearing premise
The method assumes that true claims usually agree across sources while polluted claims create detectable contradictions under NLI and a lexical-overlap pair filter, so coordinated multi-source misinformation or sparse paraphrase-heavy pollution can defeat the gate.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes MIRAGE, a training-free, model-agnostic wrapper for long-form RAG that extracts sentence-level claims from top-k retrieved passages, builds a cross-source NLI support/contradiction graph (with optional counter-evidence probes), greedily prunes to a defended claim set S⋆, and applies a Defended-Claims Gate (contradiction ratio r_contr and source diversity d_src; eqs. 5–6) that either conditions generation on verified claims or blocks retrieval and answers parametrically. The authors also introduce a minimal-edit pollution protocol with four families (Unambiguous, Conflicting, Misleading, Fabricated) to create matched Clean/Mixed/FullP regimes. Across four long-form QA benchmarks and five LLM backbones, vanilla RAG degrades sharply under pollution while MIRAGE largely restores VeriScore F1@k and outperforms many robust-RAG baselines (Tables 1–2), with ablations, gate pass/block rates, k-sensitivity, and NLI/VeriScore label audits.
Significance. If the results transfer beyond the synthetic pollution generator, this is a useful systems contribution: a modular, training-free evidence-adjudication layer that targets on-topic misinformation rather than mere irrelevance, with released code, datasets, and a controllable pollution protocol that the community can reuse. Strengths include breadth of evaluation (four datasets, commercial and open-weight models, many baselines), offline gate calibration without the generator, separation of VeriScore’s external evidence from RAG passages, component ablations (Table 5), and human agreement checks on NLI edges and VeriScore labels (Table 6). The work is significant primarily as an empirical defense and evaluation protocol rather than as a new theoretical principle of multi-source consistency.
major comments (4)
- [§4 Pollution Protocol; Limitations] §4 and App. B/K: The central mechanism claim—that truthful claims are multi-source consistent while polluted claims induce NLI-detectable structural inconsistency (§1, §3; eqs. 5–6)—is stress-tested almost exclusively under GPT-4o-mini minimal-edit rewrites that often flip atomic facts or invent particulars while preserving topic and lexical anchors. Coordinated multi-source falsehoods, shared framing across domains, or paraphrase/omission pollution that preserves surface agreement are acknowledged in Limitations but not measured. Without at least one such stress regime (or real web-pollution sample), Tables 1–2 support robustness to this generator more than to the broader misinformation setting the abstract invokes. Please add such an evaluation or substantially narrow the claim language.
- [Table 3; §6.1–6.2; Appendix I] Table 3 and §6.1–6.2: Under FullP the gate pass rate is ~0%, so reported FullP recovery is largely safer parametric fallback (plus the stronger no-evidence prompt noted in App. I), not successful adjudication of polluted evidence. MixP gains mix pruning of inconsistent claims with partial blocking (~48.6% pass). The main text currently presents both regimes as “restoring factuality” via the same pipeline. Please decompose results by gate decision (PASS vs BLOCK), report factuality conditional on pass, and state explicitly how much of FullP improvement is blocking versus claim pruning.
- [§3.4; Table 8; eqs. (5)–(6)] Table 8 / eqs. (5)–(6): The global configuration sets r_max = 0.98, so the contradiction-ratio criterion is nearly inactive unless almost all edges are contradictions; d_min = 2 then dominates the trust decision. This makes the gate’s advertised dual criteria asymmetric and raises the question whether “consistency” is doing the work claimed in §3.4. Please report (i) which criterion triggers blocks in Clean/MixP/FullP, (ii) sensitivity of Tables 1–2 to r_max and d_min, and (iii) justification for r_max = 0.98 beyond offline grid search on held-out graph statistics.
- [§3.2 Pair selection; Appendix J.2] §3.2 / App. J.2: Pair selection uses containment-style token overlap with τ_overlap = 0.62 before NLI. The diagnostic (Table 18) shows large NLI savings but also that paraphrastic or omission-based pollution can sparsify the graph and push the system toward conservative blocking. Given Misleading pollution is one of the four families, please quantify false-negative edge rates on Misleading/paraphrase pairs and show how often MixP failures are due to missed support/contradiction edges rather than gate thresholds.
minor comments (6)
- [Figure 1; Figure 3] Figure 1 and Figure 3 are helpful but low-resolution in the manuscript text; ensure vector figures and readable edge labels in the camera-ready version.
- [§3.2; §5.2; Table 4] Notation: ˜r(c) is introduced as max-normalized retrieval score; later tables use F1@k with k overloaded as both retrieval depth and VeriScore’s claim-count median. Disambiguate (e.g., F1@K vs. top-k).
- [§2 Related Work] Related Work (§2) groups many concurrent robust-RAG methods; a short table of assumptions (training-free? multi-source consistency? pollution type) would help readers place MIRAGE relative to AstuteRAG, TrustRAG, RobustRAG, and MADAM-RAG.
- [Appendix C.2; Table 10] Domain-trust priors (Table 10) assign Wikipedia 0.90 and forums 0.20. Even if secondary (α = 0.5), a one-paragraph sensitivity check with α = 0 (no domain prior) would strengthen the claim that multi-source agreement dominates.
- [Throughout] Minor prose: “FA V A” spacing artifacts and occasional hyphenation breaks (e.g., “MI-RAGE”) should be cleaned for production.
- [Appendix G] App. G runtime is on 30 samples with API cost only; wall-clock NLI latency per query would help practitioners assess the local compute trade-off claimed in Limitations.
Circularity Check
No significant circularity: MIRAGE is an empirical systems defense whose results do not reduce by construction to fitted inputs or self-citation.
full rationale
MIRAGE does not present a first-principles derivation or a fitted quantity renamed as a prediction. The core pipeline (claim extraction, cross-source NLI graph, greedy pruning of F(S), and the Defended-Claims Gate on r_contr and d_src) is a designed procedure whose success is measured empirically against external VeriScore labels. Gate thresholds are chosen offline from held-out graph statistics without the generator and are excluded from evaluation, so they are hyperparameters rather than circular fits. Claim-scoring weights (α=0.5, β=0.3) are fixed priority choices, not outcome-fitted parameters. Evaluation evidence for VeriScore is retrieved via Serper independently of the RAG passages, so the metric is not the same object as the defended graph. The concurrent Eletter et al. (2026) self-citation is non-load-bearing multimodal concurrent work. Author-defined pollution families and domain priors raise external-validity questions about whether synthetic edits induce the inconsistency the method assumes, but that is evaluation design risk, not a reduction of a claimed prediction to its inputs by construction. Under the stated circularity criteria, the paper is self-contained against its benchmarks.
Assumptions & free parameters
free parameters (6)
- r_max (max contradiction ratio)
- d_min (min source diversity)
- α, β claim-scoring weights
- τ_overlap lexical pair filter
- Graph caps (k, N, P_max, M, k_ctr, θ_support)
- Domain-trust prior table
assumptions (5)
- ad hoc to paper Truthful claims tend to be corroborated across independent sources, while misinformation induces contradictions detectable by NLI.
- domain assumption Off-the-shelf NLI (DeBERTa-large-MNLI etc.) is a sufficiently reliable adjudicator of support/contradiction between claims and counter-evidence.
- ad hoc to paper Minimal-edit GPT-4o-mini rewrites of clean passages are a realistic proxy for real-world retrieval pollution.
- domain assumption VeriScore F1@k with external Serper evidence measures long-form factuality comparably across clean and polluted generation contexts.
- domain assumption Standard dense retrieval and sentence segmentation preserve enough claim structure for graph construction without LLM claim decomposition.
invented entities (2)
-
MIRAGE defended-claims graph and Defended-Claims Gate
-
Four-family minimal-edit pollution protocol (Unambiguous, Conflicting, Misleading, Fabricated)
Cite this review
Pith. "Pith review of MIRAGE: Defending Long-Form RAG Against Misinformation Pollution." pith.science (2026). https://pith.science/paper/NMAQXCIT
@misc{pith2026260705069,
author = {Pith},
title = {Pith review of: MIRAGE: Defending Long-Form RAG Against Misinformation Pollution},
year = {2026},
howpublished = {\url{https://pith.science/paper/NMAQXCIT}},
note = {Machine review of arXiv:2607.05069}
}
read the original abstract
Retrieval-Augmented Generation (RAG) improves factuality by grounding LLMs in external evidence, but real-world retrieval is often polluted: semantically relevant passages may contain subtle misinformation, misleading framings, or fabrications. We introduce MIRAGE, a training-free, model-agnostic defense for long-form RAG. MIRAGE builds an NLI-based cross-document claim graph and applies a Defended-Claims Gate to either condition generation on a consistent, multi-source supported subset or to block retrieval and answer parametrically. We also release a minimal-edit pollution protocol spanning four perturbation families (Unambiguous, Conflicting, Misleading, Fabricated) to construct matched clean, mixed, and fully polluted evaluation regimes. Across four long-form QA benchmarks and multiple commercial and open-weight LLMs, pollution severely degrades vanilla RAG, while MIRAGE consistently restores factuality under mixed and fully polluted evidence and outperforms prior robust-RAG methods. Our implementation and datasets are available at https://github.com/SaadElDine/MIRAGE.
Figures
Figures from the paper (1 more)
Forward citations
Cited by 1 Pith paper
-
Trust Before Fusion: QIMG-7 and Source-Aware Resolution for Polluted Multimodal RAG
Under multimodal retrieval pollution, selective source-aware trust (Field-Selector SATR) beats unconditional fusion by ~12 balanced points on QIMG-7, driven mainly by text-reliability modeling.
Reference graph
Works this paper leans on
-
[1]
Retrieval-Augmented Generation for
Zhao, Penghao and Zhang, Hailin and Yu, Qinhan and Wang, Zhengren and Geng, Yunteng and Fu, Fangcheng and Yang, Ling and Zhang, Wentao and Jiang, Jie and Cui, Bin , journal =. Retrieval-Augmented Generation for. 2026 , publisher =
2026
-
[2]
2024 , publisher =
Yu, Yue and Ping, Wei and Liu, Zihan and Wang, Boxin and You, Jiaxuan and Zhang, Chao and Shoeybi, Mohammad and Catanzaro, Bryan , booktitle =. 2024 , publisher =
2024
-
[3]
Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , series =
When Not to Trust Language Models: Investigating Effectiveness of Parametric and Non-Parametric Memories , author =. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , series =. 2023 , address =
2023
-
[4]
ArXiv preprint , volume =
Reliable, Adaptable, and Attributable Language Models with Retrieval , author =. ArXiv preprint , volume =. 2024 , eprint =
2024
-
[5]
Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP) , series =
Dense Passage Retrieval for Open-Domain Question Answering , author =. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP) , series =. 2020 , address =
2020
-
[6]
Findings of the Association for Computational Linguistics: EMNLP 2023 , series =
How to Train Your Dragon: Diverse Augmentation Towards Generalizable Dense Retrieval , author =. Findings of the Association for Computational Linguistics: EMNLP 2023 , series =. 2023 , address =
2023
-
[7]
and Smith, Noah A
Kasai, Jungo and Sakaguchi, Keisuke and Takahashi, Yoichi and Le Bras, Ronan and Asai, Akari and Yu, Xinyan Velocity and Radev, Dragomir R. and Smith, Noah A. and Choi, Yejin and Inui, Kentaro , booktitle =. 2023 , publisher =
2023
-
[8]
Findings of the Association for Computational Linguistics: ACL 2024 , series =
Benchmarking Retrieval-Augmented Generation for Medicine , author =. Findings of the Association for Computational Linguistics: ACL 2024 , series =. 2024 , address =
2024
Show all 88 references
-
[9]
Findings of the Association for Computational Linguistics: EMNLP 2023 , series =
On the Risk of Misinformation Pollution with Large Language Models , author =. Findings of the Association for Computational Linguistics: EMNLP 2023 , series =. 2023 , address =
2023
-
[10]
Findings of the Association for Computational Linguistics: NAACL 2024 , series =
Why So Gullible? Enhancing the Robustness of Retrieval-Augmented Models against Counterfactual Noise , author =. Findings of the Association for Computational Linguistics: NAACL 2024 , series =. 2024 , address =
2024
-
[11]
Findings of the Association for Computational Linguistics: NAACL 2024 , series =
Adapting Fake News Detection to the Era of Large Language Models , author =. Findings of the Association for Computational Linguistics: NAACL 2024 , series =. 2024 , address =
2024
-
[12]
ArXiv preprint , volume =
Fake News Detectors are Biased Against Texts Generated by Large Language Models , author =. ArXiv preprint , volume =. 2023 , eprint =
2023
-
[13]
2020 , publisher =
Shu, Kai and Mahudeswaran, Deepak and Wang, Suhang and Lee, Dongwon and Liu, Huan , journal =. 2020 , publisher =
2020
-
[14]
Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , series =
Improving Factuality with Explicit Working Memory , author =. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , series =. 2025 , address =
2025
-
[15]
On the Structural Memory of
Zeng, Ruihong and Fang, Jinyuan and Liu, Siwei and Meng, Zaiqiao , journal =. On the Structural Memory of. 2024 , eprint =
2024
-
[16]
Proceedings of the International Conference on Learning Representations , series =
Self-Consistency Improves Chain of Thought Reasoning in Language Models , author =. Proceedings of the International Conference on Learning Representations , series =. 2023 , url =
2023
-
[17]
Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , series =
Malaviya, Chaitanya and Shaw, Peter and Chang, Ming-Wei and Lee, Kenton and Toutanova, Kristina , editor =. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , series =. 2023 , address =
2023
-
[18]
ArXiv preprint , volume =
Learning to Reason for Factuality , author =. ArXiv preprint , volume =. 2025 , eprint =
2025
-
[19]
Findings of the Association for Computational Linguistics: ACL 2024 , series =
Chain-of-Verification Reduces Hallucination in Large Language Models , author =. Findings of the Association for Computational Linguistics: ACL 2024 , series =. 2024 , address =
2024
-
[20]
Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , series =
Verify-and-Edit: A Knowledge-Enhanced Chain-of-Thought Framework , author =. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , series =. 2023 , address =
2023
-
[21]
Advances in Neural Information Processing Systems , series =
Long-form Factuality in Large Language Models , author =. Advances in Neural Information Processing Systems , series =. 2024 , publisher =
2024
-
[22]
Proceedings of the Conference on Language Modeling , series =
Fine-grained Hallucination Detection and Editing for Language Models , author =. Proceedings of the Conference on Language Modeling , series =. 2024 , url =
2024
-
[23]
and Hashimoto, Tatsunori B
Dubois, Yann and Li, Chen Xuechen and Taori, Rohan and Zhang, Tianyi and Gulrajani, Ishaan and Ba, Jimmy and Guestrin, Carlos and Liang, Percy S. and Hashimoto, Tatsunori B. , booktitle =. 2023 , publisher =
2023
-
[24]
Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing , series =
Min, Sewon and Krishna, Kalpesh and Lyu, Xinxi and Lewis, Mike and Yih, Wen-tau and Koh, Pang and Iyyer, Mohit and Zettlemoyer, Luke and Hajishirzi, Hannaneh , editor =. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing , series =. 2023 , address =
2023
-
[25]
ArXiv preprint , volume =
A Comprehensive Survey of Hallucination Mitigation Techniques in Large Language Models , author =. ArXiv preprint , volume =. 2024 , eprint =
2024
-
[26]
2023 , publisher =
Liu, Yi-Hsien and Han, Tianle and Ma, Siyuan and Zhang, Jia-Yu and Yang, Yuanyu and Tian, Jiaming and He, Haoyang and Li, Antong and He, Mengshen and Liu, Zheng and others , journal =. 2023 , publisher =
2023
-
[27]
Self-Alignment for Factuality: Mitigating Hallucinations in
Zhang, Xiaoying and Peng, Baolin and Tian, Ye and Zhou, Jingyan and Jin, Lifeng and Song, Linfeng and Mi, Haitao and Meng, Helen , editor =. Self-Alignment for Factuality: Mitigating Hallucinations in. Proceedings of the 62nd Annual Meeting of the Association for Computational...
2024
-
[28]
Findings of the Association for Computational Linguistics: EMNLP 2024 , series =
Wang, Yuxia and Reddy, Revanth Gangi and Mujahid, Zain Muhammad and Arora, Arnav and Rubashevskii, Aleksandr and Geng, Jiahui and Afzal, Osama Mohammed and Pan, Liangming and Borenstein, Nadav and Pillai, Aditya and Augenstein, Isabelle and Gurevych, Iryna and Nakov, Preslav ,...
2024
-
[29]
Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , series =
Fatahi Bayat, Farima and Zhang, Lechen and Munir, Sheza and Wang, Lu , editor =. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , series =. 2025 , address =
2025
-
[30]
Findings of the Association for Computational Linguistics: EMNLP 2024 , series =
Song, Yixiao and Kim, Yekyung and Iyyer, Mohit , editor =. Findings of the Association for Computational Linguistics: EMNLP 2024 , series =. 2024 , address =
2024
-
[31]
Length-Controlled
Dubois, Yann and Liang, Percy and Hashimoto, Tatsunori , booktitle =. Length-Controlled. 2024 , url =
2024
-
[32]
Retrieval-Augmented Generation for Knowledge-Intensive
Lewis, Patrick and Perez, Ethan and Piktus, Aleksandra and Petroni, Fabio and Karpukhin, Vladimir and Goyal, Naman and K. Retrieval-Augmented Generation for Knowledge-Intensive. Advances in Neural Information Processing Systems , series =. 2020 , publisher =
2020
-
[33]
Proceedings of the 37th International Conference on Machine Learning , series =
Retrieval Augmented Language Model Pre-Training , author =. Proceedings of the 37th International Conference on Machine Learning , series =. 2020 , publisher =
2020
-
[34]
ArXiv preprint , volume =
Retrieval-Augmented Generation for Large Language Models: A Survey , author =. ArXiv preprint , volume =. 2023 , eprint =
2023
-
[35]
Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing , series =
Active Retrieval Augmented Generation , author =. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing , series =. 2023 , address =
2023
-
[36]
Asai, Akari and Wu, Zeqiu and Wang, Yizhong and Sil, Avirup and Hajishirzi, Hannaneh , booktitle =. Self-. 2024 , url =
2024
-
[37]
Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , series =
Su, Weihang and Tang, Yichen and Ai, Qingyao and Wu, Zhijing and Liu, Yiqun , editor =. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , series =. 2024 , address =
2024
-
[38]
2025 , url =
Wei, Zhepei and Chen, Wei-Lin and Meng, Yu , booktitle =. 2025 , url =
2025
-
[39]
Certifiably Robust
Xiang, Chong and Wu, Tong and Zhong, Zexuan and Wagner, David and Chen, Danqi and Mittal, Prateek , booktitle =. Certifiably Robust. 2024 , url =
2024
-
[40]
Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics , series =
Lewis, Mike and Liu, Yinhan and Goyal, Naman and Ghazvininejad, Marjan and Mohamed, Abdelrahman and Levy, Omer and Stoyanov, Veselin and Zettlemoyer, Luke , editor =. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics , series =. 2020 , address =
2020
-
[41]
Worse than Zero-shot? A Fact-Checking Dataset for Evaluating the Robustness of
Zeng, Linda and Gupta, Rithwik and Motwani, Divij and Zhang, Yi and Yang, Diji , booktitle =. Worse than Zero-shot? A Fact-Checking Dataset for Evaluating the Robustness of. 2025 , url =
2025
-
[42]
Proceedings of the Conference on Language Modeling , series =
Retrieval-Augmented Generation with Conflicting Evidence , author =. Proceedings of the Conference on Language Modeling , series =. 2025 , url =
2025
-
[43]
, editor =
Wang, Fei and Wan, Xingchen and Sun, Ruoxi and Chen, Jiefeng and Arik, Sercan O. , editor =. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , series =. 2025 , address =
2025
-
[44]
Findings of the Association for Computational Linguistics: ACL 2023 , series =
Discovering Language Model Behaviors with Model-Written Evaluations , author =. Findings of the Association for Computational Linguistics: ACL 2023 , series =. 2023 , address =
2023
-
[45]
Proceedings of the International Conference on Learning Representations , series =
Making Retrieval-Augmented Language Models Robust to Irrelevant Context , author =. Proceedings of the International Conference on Learning Representations , series =. 2024 , url =
2024
-
[46]
ArXiv preprint , volume =
After Retrieval, Before Generation: Enhancing the Trustworthiness of Large Language Models in Retrieval-Augmented Generation , author =. ArXiv preprint , volume =. 2025 , eprint =
2025
-
[47]
2025 , eprint =
Zhou, Huichi and Lee, Kin-Hei and Zhan, Zhonghao and Chen, Yue and Li, Zhenhao and Wang, Zhaoyang and Haddadi, Hamed and Yilmaz, Emine , journal =. 2025 , eprint =
2025
-
[48]
Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers) , series =
Evidence-Driven Retrieval Augmented Response Generation for Online Misinformation , author =. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers) , series =. 2024...
2024
-
[49]
Grattafiori, Aaron and Dubey, Abhimanyu and Jauhri, Abhinav and Pandey, Abhinav and Kadian, Abhishek and Al-Dahle, Ahmad and Letman, Aiesha and Mathur, Akhil and Schelten, Alan and Vaughan, Alex and Yang, Amy and Fan, Angela and Goyal, Anirudh and Hartshorn, Anthony and Yang, ...
2024
-
[50]
Jiang, Albert Q. and Sablayrolles, Alexandre and Mensch, Arthur and Bamford, Chris and Chaplot, Devendra Singh and Casas, Diego de las and Bressand, Florian and Lengyel, Gianna and Lample, Guillaume and Saulnier, Lucile and Lavaud, Lelio Renard and Lachaux, Marie-Anne and Stoc...
2023
-
[51]
2021 , url =
He, Pengcheng and Liu, Xiaodong and Gao, Jianfeng and Chen, Weizhu , booktitle =. 2021 , url =
2021
-
[52]
2025 , eprint =
Yang, An and Li, Anfeng and Yang, Baosong and Zhang, Beichen and Hui, Binyuan and Zheng, Bo and Yu, Bowen and Gao, Chang and Huang, Chengen and Lv, Chenxu and Zheng, Chujie and Liu, Dayiheng and Zhou, Fan and Huang, Fei and Hu, Feng and Ge, Hao and Wei, Haoran and Lin, Huan an...
2025
-
[53]
2019 , eprint =
Liu, Yinhan and Ott, Myle and Goyal, Naman and Du, Jingfei and Joshi, Mandar and Chen, Danqi and Levy, Omer and Lewis, Mike and Zettlemoyer, Luke and Stoyanov, Veselin , journal =. 2019 , eprint =
2019
-
[54]
Proceedings of the 48th Annual Meeting of the Association for Computational Linguistics , series =
Soricut, Radu and Echihabi, Abdessamad , editor =. Proceedings of the 48th Annual Meeting of the Association for Computational Linguistics , series =. 2010 , address =
2010
-
[55]
Findings of the Association for Computational Linguistics: EMNLP 2025 , series =
Wan, Yingjia and Tan, Haochen and Zhu, Xiao and Zhou, Xinyu and Li, Zhiwei and Lv, Qingsong and Sun, Changxuan and Zeng, Jiaqi and Xu, Yi and Lu, Jianqiao and Liu, Yinhong and Guo, Zhijiang , editor =. Findings of the Association for Computational Linguistics: EMNLP 2025 , ser...
2025
-
[56]
Findings of the Association for Computational Linguistics: EMNLP 2025 , series =
Rajendhran, Rishanth and Zadeh, Amir and Sarte, Matthew and Li, Chuan and Iyyer, Mohit , editor =. Findings of the Association for Computational Linguistics: EMNLP 2025 , series =. 2025 , address =
2025
-
[57]
Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers) , series =
A Broad-Coverage Challenge Corpus for Sentence Understanding through Inference , author =. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers) , series =. 2018 , ...
2018
-
[58]
2026 , howpublished =
2026
-
[59]
Trust Before Fusion:
Eletter, Saadeldine and Aijaz, Owais and Nakov, Preslav , year =. Trust Before Fusion:
-
[60]
Akari Asai, Zeqiu Wu, Yizhong Wang, Avirup Sil, and Hannaneh Hajishirzi. 2024. https://openreview.net/forum?id=hSyW5go0v8 Self- RAG : Learning to retrieve, generate, and critique through self-reflection . In Proceedings of the International Conference on Learning Representatio...
2024
-
[61]
Mingda Chen, Yang Li, Karthik Padthe, Rulin Shao, Alicia Yi Sun, Luke Zettlemoyer, Gargi Ghosh, and Wen-tau Yih. 2025. https://doi.org/10.18653/v1/2025.acl-long.548 Improving factuality with explicit working memory . In Proceedings of the 63rd Annual Meeting of the Association...
2025 doi
-
[62]
Xinbang Dai, Huikang Hu, Yuncheng Hua, Jiaqi Li, Yongrui Chen, Rihui Jin, Nan Hu, and Guilin Qi. 2025. https://arxiv.org/abs/2505.17118 After retrieval, before generation: Enhancing the trustworthiness of large language models in retrieval-augmented generation . ArXiv preprint...
2025
-
[63]
Liang, and Tatsunori B
Yann Dubois, Chen Xuechen Li, Rohan Taori, Tianyi Zhang, Ishaan Gulrajani, Jimmy Ba, Carlos Guestrin, Percy S. Liang, and Tatsunori B. Hashimoto. 2023. https://proceedings.neurips.cc/paper_files/paper/2023/hash/5fc47800ee5b30b8777fdd30abcaaf3b-Abstract-Conference.html AlpacaFa...
2023
-
[64]
Saadeldine Eletter, Owais Aijaz, and Preslav Nakov. 2026. Trust before fusion: QIMG -7 and source-aware resolution for polluted multimodal RAG . Manuscript
2026
-
[65]
Aaron Grattafiori, Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Alex Vaughan, Amy Yang, Angela Fan, Anirudh Goyal, Anthony Hartshorn, Aobo Yang, Archi Mitra, Archie Sravankumar, Artem Korenev, Art...
2024 arXiv
-
[66]
Pengcheng He, Xiaodong Liu, Jianfeng Gao, and Weizhu Chen. 2021. https://openreview.net/forum?id=XPZIaotutsD DeBERTa : Decoding-enhanced BERT with disentangled attention . In Proceedings of the International Conference on Learning Representations, ICLR '21
2021
-
[67]
Albert Q. Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, Lelio Renard Lavaud, Marie-Anne Lachaux, Pierre Stock, Teven Le Scao, Thibaut Lavril, Thomas ...
2023 arXiv
-
[68]
Xu, Luyu Gao, Zhiqing Sun, Qian Liu, Jane Dwivedi-Yu, Yiming Yang, Jamie Callan, and Graham Neubig
Zhengbao Jiang, Frank F. Xu, Luyu Gao, Zhiqing Sun, Qian Liu, Jane Dwivedi-Yu, Yiming Yang, Jamie Callan, and Graham Neubig. 2023 b . https://doi.org/10.18653/v1/2023.emnlp-main.495 Active retrieval augmented generation . In Proceedings of the 2023 Conference on Empirical Meth...
2023 doi
-
[69]
Sewon Min, Kalpesh Krishna, Xinxi Lyu, Mike Lewis, Wen-tau Yih, Pang Koh, Mohit Iyyer, Luke Zettlemoyer, and Hannaneh Hajishirzi. 2023. https://doi.org/10.18653/v1/2023.emnlp-main.741 FA ct S core: Fine-grained atomic evaluation of factual precision in long form text generatio...
2023 doi
-
[70]
Abhika Mishra, Akari Asai, Vidhisha Balachandran, Yizhong Wang, Yulia Tsvetkov, Graham Neubig, and Hannaneh Hajishirzi. 2024. https://openreview.net/forum?id=dJMTn3QOWO Fine-grained hallucination detection and editing for language models . In Proceedings of the Conference on L...
2024
-
[71]
OpenAI . 2026. OpenAI API model documentation. https://developers.openai.com/api/docs/models. Accessed: 2026-06-10
2026
-
[72]
Yikang Pan, Liangming Pan, Wenhu Chen, Preslav Nakov, Min-Yen Kan, and William Wang. 2023. https://doi.org/10.18653/v1/2023.findings-emnlp.97 On the risk of misinformation pollution with large language models . In Findings of the Association for Computational Linguistics: EMNL...
2023 doi
-
[73]
Ethan Perez, Sam Ringer, Kamile Lukosiute, Karina Nguyen, Edwin Chen, Scott Heiner, Craig Pettit, Catherine Olsson, Sandipan Kundu, Saurav Kadavath, Andy Jones, Anna Chen, Ben Mann, Brian Israel, Bryan Seethor, Cameron McKinnon, Christopher Olah, Da Yan, Daniela Amodei, and 44...
2023 doi
-
[74]
Rishanth Rajendhran, Amir Zadeh, Matthew Sarte, Chuan Li, and Mohit Iyyer. 2025. https://doi.org/10.18653/v1/2025.findings-emnlp.491 V eri F ast S core: Speeding up long-form factuality evaluation . In Findings of the Association for Computational Linguistics: EMNLP 2025, EMNL...
2025 doi
-
[75]
Yixiao Song, Yekyung Kim, and Mohit Iyyer. 2024. https://doi.org/10.18653/v1/2024.findings-emnlp.552 V eri S core: Evaluating the factuality of verifiable claims in long-form text generation . In Findings of the Association for Computational Linguistics: EMNLP 2024, EMNLP '24,...
2024 doi
-
[76]
Radu Soricut and Abdessamad Echihabi. 2010. https://aclanthology.org/P10-1063/ TrustRank : Inducing trust in automatic translations via ranking . In Proceedings of the 48th Annual Meeting of the Association for Computational Linguistics, ACL '10, pages 612--621, Uppsala, Swede...
2010
-
[77]
Weihang Su, Yichen Tang, Qingyao Ai, Zhijing Wu, and Yiqun Liu. 2024. https://doi.org/10.18653/v1/2024.acl-long.702 DRAGIN : Dynamic retrieval augmented generation based on the real-time information needs of large language models . In Proceedings of the 62nd Annual Meeting of ...
2024 doi
-
[78]
Yingjia Wan, Haochen Tan, Xiao Zhu, Xinyu Zhou, Zhiwei Li, Qingsong Lv, Changxuan Sun, Jiaqi Zeng, Yi Xu, Jianqiao Lu, Yinhong Liu, and Zhijiang Guo. 2025. https://doi.org/10.18653/v1/2025.findings-emnlp.1295 F a S t F act: Faster, stronger long-form factuality evaluations in ...
2025 doi
-
[79]
Fei Wang, Xingchen Wan, Ruoxi Sun, Jiefeng Chen, and Sercan O. Arik. 2025 a . https://doi.org/10.18653/v1/2025.acl-long.1476 ASTUTE RAG : Overcoming imperfect retrieval augmentation and knowledge conflicts for large language models . In Proceedings of the 63rd Annual Meeting o...
2025 doi
-
[80]
Han Wang, Archiki Prasad, Elias Stengel-Eskin, and Mohit Bansal. 2025 b . https://openreview.net/forum?id=z1MHB2m3V9 Retrieval-augmented generation with conflicting evidence . In Proceedings of the Conference on Language Modeling, COLM '25
2025
-
[81]
Jerry Wei, Chengrun Yang, Xinying Song, Yifeng Lu, Nathan Hu, Jie Huang, Dustin Tran, Daiyi Peng, Ruibo Liu, Da Huang, Cosmo Du, and Quoc V. Le. 2024. https://proceedings.neurips.cc/paper_files/paper/2024/hash/937ae0e83eb08d2cb8627fe1def8c751-Abstract-Conference.html Long-form...
2024
-
[82]
Zhepei Wei, Wei-Lin Chen, and Yu Meng. 2025. https://openreview.net/forum?id=P1qhkp8gQT InstructRAG : Instructing retrieval-augmented generation via self-synthesized rationales . In Proceedings of the International Conference on Learning Representations, ICLR '25
2025
-
[83]
Chong Xiang, Tong Wu, Zexuan Zhong, David Wagner, Danqi Chen, and Prateek Mittal. 2024. https://openreview.net/forum?id=qsEeACAJjD Certifiably robust RAG against retrieval corruption . In ICML 2024 Workshop on Next Generation of AI Safety , ICML Workshop '24
2024
-
[84]
An Yang, Anfeng Li, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Gao, Chengen Huang, Chenxu Lv, Chujie Zheng, Dayiheng Liu, Fan Zhou, Fei Huang, Feng Hu, Hao Ge, Haoran Wei, Huan Lin, Jialong Tang, and 41 others. 2025. https://arxiv.org/abs/2505.09388 Qw...
2025 arXiv
-
[85]
Ori Yoran, Tomer Wolfson, Ori Ram, and Jonathan Berant. 2024. https://openreview.net/forum?id=ZS4m74kZpH Making retrieval-augmented language models robust to irrelevant context . In Proceedings of the International Conference on Learning Representations, ICLR '24
2024
-
[86]
Zhenrui Yue, Huimin Zeng, Yimeng Lu, Lanyu Shang, Yang Zhang, and Dong Wang. 2024. https://doi.org/10.18653/v1/2024.naacl-long.313 Evidence-driven retrieval augmented response generation for online misinformation . In Proceedings of the 2024 Conference of the North American Ch...
2024 doi
-
[87]
Linda Zeng, Rithwik Gupta, Divij Motwani, Yi Zhang, and Diji Yang. 2025. https://openreview.net/forum?id=R4MeWTeVej Worse than zero-shot? a fact-checking dataset for evaluating the robustness of RAG against misleading retrievals . In The Thirty-ninth Annual Conference on Neura...
2025
-
[88]
Huichi Zhou, Kin-Hei Lee, Zhonghao Zhan, Yue Chen, Zhenhao Li, Zhaoyang Wang, Hamed Haddadi, and Emine Yilmaz. 2025. https://arxiv.org/abs/2501.00879 TrustRAG : Enhancing robustness and trustworthiness in RAG . ArXiv preprint, arXiv:2501.00879
2025 arXiv
Reviewed July 11, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.