REVIEW 2 major objections 5 minor 79 references
A Survey on Proactive Defense Strategies Against Misinformation in Large Language Models
T0 review · 2 major / 5 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read A survey of the field argues that stopping LLM misinformation at the source—through trusted knowledge, self-correcting reasoning, and hardened inputs—outperforms traditional post-hoc detection by 42–63%.
desk verdict Useful survey taxonomy, but the headline 42–63% claim is contradicted by the paper's own limitations section. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the Three Pillars of Preventative Assurance—a taxonomy that organizes 127 defensive techniques into Knowledge Credibility, Inference Reliability, and Input Robustness. The framework does the work of turning a scattered literature into a single design space, and the paper's headline quantitative claim rests on a meta-analysis of 48 benchmark studies comparing proactive techniques against detection baselines. The pillars are meant to be complementary layers: internal knowledge consolidation (data curation, knowledge editing, RAG), reasoning-time reliability (decoding, alignment, adversarial training), and input-surface hardening (prompting and injection defenses).
What would settle it
Re-run the meta-analysis with a protocol that requires matched detection baselines and identical outcome definitions across all 48 studies; if the pooled improvement over detection falls below parity or is not statistically significant, the headline claim collapses.
Extended reading notes
Core claim
The central claim is that misinformation defense for LLMs should be reconceived as a continuum of preventative assurance rather than a detection problem. The paper defines three pillars: Knowledge Credibility fortifies what the model knows, through cleaner training data, knowledge editing, and retrieval-augmented grounding; Inference Reliability shapes how the model reasons, through contrastive decoding, self-verification, and factual alignment; Input Robustness protects the interface, through prompt optimization and countermeasures to injection attacks. The paper asserts that this proactive paradigm outperforms reactive detection by 42% to 63% in misinformation prevention, based on its comparative meta-analysis, while acknowledging tradeoffs in compute and generalization. The intended outcome is a research agenda for building 'self-vaccinating' LLMs that resist fabricating falsehoods in the first place.
Load-bearing premise
The 42–63% improvement figure assumes the 48 surveyed studies measure the same construct, 'misinformation prevention,' under comparable threat models and baselines—an assumption the paper itself flags as unmet.
Editorial extensions
If this is right
- Deployment budgets for misinformation safety should shift from detection systems toward preventive layers: training-data curation, knowledge editing pipelines, and retrieval verification.
- Model builders should expect a latency increase of 1.5–3× and a cross-domain generalization gap of 18–25% when adding proactive defenses, and plan for those costs.
- A unified evaluation standard for 'misinformation prevention' is needed; the current absence of shared benchmarks blocks credible comparison across the 48 studies.
- The three pillars should be co-designed rather than assembled as modular add-ons, because weaknesses in one layer (e.g., uncurated retrieval) can undermine the others.
- Knowledge-editing interfaces become a new attack surface: the survey notes BadEdit-style backdoors can be implanted with 15 poisoned samples, so proactive defense includes securing the editing pipeline itself.
Reading between the lines
- The paper's own limitation section concedes the lack of standardized benchmarks, so the 42–63% figure is best read as a directional upper bound rather than a precise effect size.
- If proactive defense matures, the competitive benchmark for misinformation systems will likely shift from detection accuracy to prevention rate under a fixed compute budget, making the latency tradeoff the key design variable.
- The three-pillar taxonomy suggests a natural ablative test—removing each pillar independently on a single model and benchmark—which the survey does not itself run and which would reveal the marginal contribution of each layer.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This manuscript is a survey of proactive defenses against LLM-generated misinformation. It organizes the literature around a proposed 'Three Pillars of Preventative Assurance' framework: Knowledge Credibility (datasets, knowledge editing, retrieval-augmented generation), Inference Reliability (decoding methods, factual alignment, adversarial training), and Input Robustness (prompting techniques, injection defenses). The authors report surveying 127 techniques and claim a 'comparative meta-analysis' of 48 benchmark studies, asserting that proactive strategies achieve 42-63% superiority over conventional detection methods in misinformation prevention (Abstract; Section 1). The paper also discusses limitations, future work, and open challenges.
Significance. If the quantitative superiority claim were supported, the paper would provide a strong practical argument for shifting from detection-based to proactive defenses, which would be a significant contribution to LLM safety research. The proposed three-pillar taxonomy is a reasonable organizing device for a survey. However, the central empirical claim is not supported by the manuscript: no meta-analysis protocol, study list, effect-size table, or error bars are provided, and the cited source (Wu et al., 2024b) is a general survey of retrieval-augmented generation rather than a comparative meta-analysis. Moreover, the manuscript's own Limitations section explicitly disclaims the ability to draw definitive comparative conclusions. As submitted, the paper's headline result is unverifiable and internally contradicted, so the survey's value is reduced to its taxonomy, which alone does not justify the claimed contributions.
major comments (2)
- [Section 1 (Contributions) and Abstract] The paper claims a 'rigorous comparative evaluation' via a 'meta-analysis of 48 benchmark studies' and states that 'state-of-the-art proactive strategies demonstrate 42-63% superiority over conventional detection in misinformation prevention (Wu et al., 2024b).' No meta-analysis protocol, inclusion criteria, list of the 48 studies, effect-size table, confidence intervals, or baseline definitions appear anywhere in the manuscript. The single citation, Wu et al. (2024b), is a survey of retrieval-augmented generation, not a comparative meta-analysis of proactive versus detection methods. The associated tradeoff figures ('1.5-3 × latency,' '18-25% cross-domain variance') are also presented without derivation or a citable comparison. The 42-63% claim is therefore not a demonstrated result and should not be presented as one.
- [Limitations (unnumbered section before References)] The Limitations section states: 'the lack of standardized benchmarks and evaluation metrics across studies limits the ability to draw definitive conclusions about the comparative effectiveness of proactive strategies.' This directly contradicts the Abstract's 'we demonstrate ... up to 63% improvement over conventional methods' and Section 1's 'demonstrate 42-63% superiority.' The manuscript's own text disclaims the central quantitative claim, creating an internal inconsistency that must be resolved. Because the comparative-effectiveness claim is the load-bearing contribution listed in Section 1, the paper as written cannot be accepted without either supplying the missing meta-analysis or substantially rewriting the abstract and contributions to withdraw the quantitative claim.
minor comments (5)
- [Section 6 (before the heading)] The line 'Here's the combined and slightly condensed version of the two sections:' appears in the manuscript text before Section 6; this is an editing artifact that should be removed.
- [Section 5.2 (Countering Injection Attacks)] The statement that the Rossi et al. taxonomy 'informs 83% of contemporary detection frameworks' lacks a source or a clear basis; please provide a citation or qualify the claim.
- [Section 3.2.1 (Knowledge Credibility through Retrieval-Augmented Architectures)] The RGB benchmark is said to show '35% improvement over baselines,' but the baselines are not named; please identify the comparison systems and clarify whether these numbers are author-reported.
- [Section 2.1 (Misinformation)] The term 'misinformation' is defined broadly as 'any content that deviates from factual accuracy,' while later text discusses adversarial attacks and disinformation; consider distinguishing misinformation from disinformation and defining 'LLM-generated misinformation' explicitly at first use.
- [Abstract and Section 1] The phrase 'Three Pillars of Preventative Assurance' is used as if it were a standard term, but no definition is given; please define the phrase or replace it with a more transparent description.
Circularity Check
No circular derivation found: the 42–63% superiority headline is imported from a non-overlapping external citation (Wu et al., 2024b) and contradicted by the paper's own Limitations section (a support problem, not a reduction); the only author-overlap citations (Liu et al., 2023; Liu et al., 2024a) are motivational and non-load-bearing.
full rationale
This survey contains no equations, fitted parameters, or quantities derived within the paper, so none of the construction-based circularity patterns (self-definitional, fitted-input-called-prediction, ansatz-via-citation, uniqueness-imported, renaming-known-result) apply. The central quantitative claim — 'proactive defense strategies offer up to 63% improvement over conventional methods in misinformation prevention' (Abstract) and 'state-of-the-art proactive strategies demonstrate 42-63% superiority over conventional detection in misinformation prevention (Wu et al., 2024b)' (Section 1) — is imported from an external citation whose author list (Shangyu Wu, Ying Xiong, ..., Nan Guan) does not overlap with the present authors. Because the figure is imported rather than derived, there is no derivation chain that could reduce to the paper's own inputs. The claimed 'meta-analysis of 48 benchmark studies' (Section 1, contribution 2) is asserted without a protocol, study list, or effect-size table, and the manuscript's own Limitation section states: 'the lack of standardized benchmarks and evaluation metrics across studies limits the ability to draw definitive conclusions about the comparative effectiveness of proactive strategies.' I flag that passage explicitly: it directly disclaims the Abstract and Section 1 headlines, but an unsupported or internally contradicted claim is an evidence/correctness problem, not a circular reduction, and per the hard rules I may not raise the circularity score on that basis alone. The only self-citations by overlapping authors are Liu et al., 2023 (38.7% false-negative rate and 39.8% error amplification, both in Section 1, where Xuming Hu is the survey's corresponding author and a co-author of that prior work) and Liu et al., 2024a (Section 2.1, by Aiwei Liu and Xuming Hu). Both are used as motivational context for the paradigm shift; neither feeds into the 42–63% superiority figure, which rests on Wu et al., 2024b. These are therefore minor, non-load-bearing self-citations, placing the paper at score 2 on the rubric. No further circular steps were identified.
Assumptions & free parameters
assumptions (3)
- domain assumption The percentages quoted from prior work are accurate and directly comparable across studies.
- domain assumption The 48 benchmark studies measure the same construct, misinformation prevention, under comparable threat models.
- domain assumption Proactive and reactive defense form a clean partition with a well-defined baseline.
invented entities (1)
-
Three Pillars of Preventative Assurance framework
Cite this review
Pith. "Pith review of A Survey on Proactive Defense Strategies Against Misinformation in Large Language Models." pith.science (2026). https://pith.science/paper/4PFU3XB5
@misc{pith2026250705288,
author = {Pith},
title = {Pith review of: A Survey on Proactive Defense Strategies Against Misinformation in Large Language Models},
year = {2026},
howpublished = {\url{https://pith.science/paper/4PFU3XB5}},
note = {Machine review of arXiv:2507.05288}
}
read the original abstract
The widespread deployment of large language models (LLMs) across critical domains has amplified the societal risks posed by algorithmically generated misinformation. Unlike traditional false content, LLM-generated misinformation can be self-reinforcing, highly plausible, and capable of rapid propagation across multiple languages, which traditional detection methods fail to mitigate effectively. This paper introduces a proactive defense paradigm, shifting from passive post hoc detection to anticipatory mitigation strategies. We propose a Three Pillars framework: (1) Knowledge Credibility, fortifying the integrity of training and deployed data; (2) Inference Reliability, embedding self-corrective mechanisms during reasoning; and (3) Input Robustness, enhancing the resilience of model interfaces against adversarial attacks. Through a comprehensive survey of existing techniques and a comparative meta-analysis, we demonstrate that proactive defense strategies offer up to 63\% improvement over conventional methods in misinformation prevention, despite non-trivial computational overhead and generalization challenges. We argue that future research should focus on co-designing robust knowledge foundations, reasoning certification, and attack-resistant interfaces to ensure LLMs can effectively counter misinformation across varied domains.
Figures
Reference graph
Works this paper leans on
-
[1]
Shawqi Al-Maliki, Adnan Qayyum, Hassan Ali, Mohamed Abdallah, Junaid Qadir, Dinh Thai Hoang, Dusit Niyato, and Ala Al-Fuqaha. 2024. Adversarial machine learning for social good: Reframing the adversary as an ally. IEEE Transactions on Artificial Intelligence
work page 2024
-
[2]
Ashutosh Bajpai, Aaryan Goyal, Atif Anwer, and Tanmoy Chakraborty. 2024. Temporally consistent factuality probing for large language models. arXiv preprint arXiv:2409.14065
work page Pith review arXiv 2024
-
[3]
Farima Fatahi Bayat, Kun Qian, Benjamin Han, Yisi Sang, Anton Belyi, Samira Khorshidi, Fei Wu, Ihab F Ilyas, and Yunyao Li. 2023. Fleek: Factual error detection and correction with evidence retrieved from external knowledge. arXiv preprint arXiv:2310.17119
work page Pith review arXiv 2023
-
[4]
Maciej Besta, Ales Kubicek, Roman Niggli, Robert Gerstenberger, Lucas Weitzendorf, Mingyuan Chi, Patrick Iff, Joanna Gajda, Piotr Nyczyk, J \"u rgen M \"u ller, et al. 2024. Multi-head rag: Solving multi-aspect problems with llms. arXiv preprint arXiv:2406.05085
arXiv 2024
-
[5]
Baolong Bi, Shenghua Liu, Lingrui Mei, Yiwei Wang, Pengliang Ji, and Xueqi Cheng. 2024. Decoding by contrasting knowledge: Enhancing llms' confidence on edited facts. arXiv preprint arXiv:2405.11613
arXiv 2024
-
[6]
Chung-Ching Chang, David Reitter, Renat Aksitov, and Yun-Hsuan Sung. 2023. Kl-divergence guided temperature sampling. arXiv preprint arXiv:2306.01286
arXiv 2023
-
[7]
Haw-Shiuan Chang, Nanyun Peng, Mohit Bansal, Anil Ramakrishna, and Tagyoung Chung. 2024. Real sampling: Boosting factuality and diversity of open-ended generation via asymptotic entropy. arXiv preprint arXiv:2406.07735
arXiv 2024
-
[8]
Canyu Chen and Kai Shu. 2024. Combating misinformation in the age of llms: Opportunities and challenges. AI Magazine, 45(3):354--368
2024
Show all 79 references
-
[9]
Canyu Chen, Haoran Wang, Matthew Shapiro, Yunyu Xiao, Fei Wang, and Kai Shu. 2022. Combating health misinformation in social media: Characterization, detection, intervention, and open issues. arXiv preprint arXiv:2211.05289
2022 arXiv
-
[10]
Jiawei Chen, Hongyu Lin, Xianpei Han, and Le Sun. 2024 a . Benchmarking large language models in retrieval-augmented generation. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 38, pages 17754--17762
2024
-
[11]
Yangyi Chen, Karan Sikka, Michael Cogswell, Heng Ji, and Ajay Divakaran. 2024 b . Dress: Instructing large vision-language models to align and interact with humans via natural language feedback. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recogniti...
2024
-
[12]
Daixuan Cheng, Shaohan Huang, Junyu Bi, Yuefeng Zhan, Jianfeng Liu, Yujing Wang, Hao Sun, Furu Wei, Denvy Deng, and Qi Zhang. 2023. Uprise: Universal prompt retrieval for improving zero-shot evaluation. arXiv preprint arXiv:2303.08518
2023 arXiv
-
[13]
Yung-Sung Chuang, Yujia Xie, Hongyin Luo, Yoon Kim, James Glass, and Pengcheng He. 2023. Dola: Decoding by contrasting layers improves factuality in large language models. arXiv preprint arXiv:2309.03883
2023 arXiv
-
[14]
Tianyu Cui, Yanling Wang, Chuanpu Fu, Yong Xiao, Sijia Li, Xinhao Deng, Yunpeng Liu, Qinglin Zhang, Ziyi Qiu, Peiyang Li, et al. 2024. Risk taxonomy, mitigation, and assessment benchmarks of large language model systems. arXiv preprint arXiv:2401.05778
2024 arXiv
-
[15]
Damai Dai, Li Dong, Yaru Hao, Zhifang Sui, Baobao Chang, and Furu Wei. 2021. Knowledge neurons in pretrained transformers. arXiv preprint arXiv:2104.08696
2021 arXiv
-
[16]
Souvik Das, Lifeng Jin, Linfeng Song, Haitao Mi, Baolin Peng, and Dong Yu. 2024. Entropy guided extrapolative decoding to improve factuality in large language models. arXiv preprint arXiv:2404.09338
2024 arXiv
-
[17]
Shehzaad Dhuliawala, Mojtaba Komeili, Jing Xu, Roberta Raileanu, Xian Li, Asli Celikyilmaz, and Jason Weston. 2023. Chain-of-verification reduces hallucination in large language models. arXiv preprint arXiv:2309.11495
2023 arXiv
-
[18]
Guanting Dong, Xiaoshuai Song, Yutao Zhu, Runqi Qiao, Zhicheng Dou, and Ji-Rong Wen. 2024. Toward general instruction-following alignment for retrieval-augmented generation. arXiv preprint arXiv:2410.09584
2024 arXiv
-
[19]
Feiteng Fang, Yuelin Bai, Shiwen Ni, Min Yang, Xiaojun Chen, and Ruifeng Xu. 2024. Enhancing noise robustness of retrieval-augmented language models with adaptive adversarial training. arXiv preprint arXiv:2405.20978
2024 arXiv
-
[20]
Bishwamittra Ghosh, Sarah Hasan, Naheed Anjum Arafat, and Arijit Khan. 2024. Logical consistency of large language models in fact-checking. arXiv preprint arXiv:2412.16100
2024 arXiv
-
[21]
Kelvin Guu, Kenton Lee, Zora Tung, Panupong Pasupat, and Mingwei Chang. 2020. Retrieval augmented language model pre-training. In International conference on machine learning, pages 3929--3938. PMLR
2020
-
[22]
Peter Hase, Mohit Bansal, Been Kim, and Asma Ghandeharioun. 2024. Does localization inform editing? surprising differences in causality-based localization vs. knowledge editing in language models. Advances in Neural Information Processing Systems, 36
2024
-
[23]
Jason Hoelscher-Obermaier, Julia Persson, Esben Kran, Ioannis Konstas, and Fazl Barez. 2023. Detecting edit failures in large language models: An improved specificity benchmark. arXiv preprint arXiv:2305.17553
2023 arXiv
-
[24]
Jiaming Ji, Mickel Liu, Josef Dai, Xuehai Pan, Chi Zhang, Ce Bian, Boyuan Chen, Ruiyang Sun, Yizhou Wang, and Yaodong Yang. 2024. Beavertails: Towards improved safety alignment of llm via a human-preference dataset. Advances in Neural Information Processing Systems, 36
2024
-
[25]
Albert Q Jiang, Alexandre Sablayrolles, Antoine Roux, Arthur Mensch, Blanche Savary, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Emma Bou Hanna, Florian Bressand, et al. 2024. Mixtral of experts. arXiv preprint arXiv:2401.04088
2024 arXiv
-
[26]
Chao Jin, Zili Zhang, Xuanlin Jiang, Fangyue Liu, Xin Liu, Xuanzhe Liu, and Xin Jin. 2024 a . Ragcache: Efficient knowledge caching for retrieval-augmented generation. arXiv preprint arXiv:2404.12457
2024 arXiv
-
[27]
Lifeng Jin, Baolin Peng, Linfeng Song, Haitao Mi, Ye Tian, and Dong Yu. 2024 b . Collaborative decoding of critical tokens for boosting factuality of large language models. arXiv preprint arXiv:2402.17982
2024 arXiv
-
[28]
Katie Kang, Eric Wallace, Claire Tomlin, Aviral Kumar, and Sergey Levine. 2024. Unfamiliar finetuning examples control how language models hallucinate. arXiv preprint arXiv:2403.05612
2024 arXiv
-
[29]
Waleed Kareem and Noorhan Abbas. 2023. Fighting lies with intelligence: Using large language models and chain of thoughts technique to combat fake news. In International Conference on Innovative Techniques and Applications of Artificial Intelligence, pages 253--258. Springer
2023
-
[30]
u ttler, Mike Lewis, Wen-tau Yih, Tim Rockt \
Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich K \"u ttler, Mike Lewis, Wen-tau Yih, Tim Rockt \"a schel, et al. 2020. Retrieval-augmented generation for knowledge-intensive nlp tasks. Advances in Neural Information Proc...
2020
-
[31]
Guanghua Li, Wensheng Lu, Wei Zhang, Defu Lian, Kezhong Lu, Rui Mao, Kai Shu, and Hao Liao. 2024 a . Re-search for the truth: Multi-round retrieval-augmented large language models are strong fake news detectors. arXiv preprint arXiv:2403.09747
2024 arXiv
-
[32]
Jiarui Li, Ye Yuan, and Zehua Zhang. 2024 b . Enhancing llm factual accuracy with rag to counter hallucinations: A case study on domain-specific queries in private knowledge-bases. arXiv preprint arXiv:2403.10446
2024 arXiv
-
[33]
Lincan Li, Jiaqi Li, Catherine Chen, Fred Gui, Hongjia Yang, Chenxiao Yu, Zhengguang Wang, Jianing Cai, Junlong Aaron Zhou, Bolin Shen, et al. 2024 c . Political-llm: Large language models in political science. arXiv preprint arXiv:2412.06864
2024 arXiv
-
[34]
Ming Li, Lichang Chen, Jiuhai Chen, Shwai He, Jiuxiang Gu, and Tianyi Zhou. 2024 d . Selective reflection-tuning: Student-selected data recycling for llm instruction-tuning. arXiv preprint arXiv:2402.10110
2024 arXiv
-
[35]
Minghan Li, Xilun Chen, Ari Holtzman, Beidi Chen, Jimmy Lin, Wen-tau Yih, and Xi Victoria Lin. 2024 e . Nearest neighbor speculative decoding for llm generation and attribution. arXiv preprint arXiv:2405.19325
2024 arXiv
-
[36]
Xiang Lisa Li, Ari Holtzman, Daniel Fried, Percy Liang, Jason Eisner, Tatsunori B Hashimoto, Luke Zettlemoyer, and Mike Lewis. 2023 a . Contrastive decoding: Open-ended text generation as optimization. In Proceedings of the 61st Annual Meeting of the Association for Computatio...
2023
-
[37]
Yanzhou Li, Tianlin Li, Kangjie Chen, Jian Zhang, Shangqing Liu, Wenhan Wang, Tianwei Zhang, and Yang Liu. 2024 f . Badedit: Backdooring large language models by model editing. arXiv preprint arXiv:2403.13355
2024 arXiv
-
[38]
Zhoubo Li, Ningyu Zhang, Yunzhi Yao, Mengru Wang, Xi Chen, and Huajun Chen. 2023 b . Unveiling the pitfalls of knowledge editing for large language models. arXiv preprint arXiv:2310.02129
2023 arXiv
-
[39]
Sheng-Chieh Lin, Luyu Gao, Barlas Oguz, Wenhan Xiong, Jimmy Lin, Wen-tau Yih, and Xilun Chen. 2024. Flame: Factuality-aware alignment for large language models. arXiv preprint arXiv:2405.01525
2024 arXiv
-
[40]
Stephanie Lin, Jacob Hilton, and Owain Evans. 2021. Truthfulqa: Measuring how models mimic human falsehoods. arXiv preprint arXiv:2109.07958
2021 arXiv
-
[41]
Aiwei Liu, Qiang Sheng, and Xuming Hu. 2024 a . https://doi.org/10.1145/3626772.3661377 Preventing and detecting misinformation generated by large language models . In Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retriev...
2024
-
[42]
Fan Liu, Zhao Xu, and Hao Liu. 2024 b . Adversarial tuning: Defending against jailbreak attacks for llms. arXiv preprint arXiv:2406.06622
2024 arXiv
-
[43]
Weize Liu, Guocong Li, Kai Zhang, Bang Du, Qiyuan Chen, Xuming Hu, Hongxia Xu, Jintai Chen, and Jian Wu. 2023. Mind's mirror: Distilling self-evaluation capability and comprehensive thinking from large language models. arXiv preprint arXiv:2311.09214
2023 arXiv
-
[44]
Junyu Luo, Cao Xiao, and Fenglong Ma. 2023. Zero-resource hallucination prevention for large language models. arXiv preprint arXiv:2309.02654
2023 arXiv
-
[45]
Kevin Meng, David Bau, Alex Andonian, and Yonatan Belinkov. 2022 a . Locating and editing factual associations in gpt. Advances in Neural Information Processing Systems, 35:17359--17372
2022
-
[46]
Kevin Meng, Arnab Sen Sharma, Alex Andonian, Yonatan Belinkov, and David Bau. 2022 b . Mass-editing memory in a transformer. arXiv preprint arXiv:2210.07229
2022 arXiv
-
[47]
Eric Mitchell, Charles Lin, Antoine Bosselut, Christopher D Manning, and Chelsea Finn. 2022. Memory-based model editing at scale. In International Conference on Machine Learning, pages 15817--15831. PMLR
2022
-
[48]
Mengjia Niu, Hao Li, Jie Shi, Hamed Haddadi, and Fan Mo. 2024. Mitigating hallucinations in large language models via self-refinement-enhanced knowledge retrieval. arXiv preprint arXiv:2405.06545
2024 arXiv
-
[49]
Vishnu S Pendyala and Christopher E Hall. 2024. Explaining misinformation detection using large language models. Electronics, 13(9):1673
2024
-
[50]
Baolin Peng, Michel Galley, Pengcheng He, Hao Cheng, Yujia Xie, Yu Hu, Qiuyuan Huang, Lars Liden, Zhou Yu, Weizhu Chen, et al. 2023. Check your facts and try again: Improving large language models with external knowledge and automated feedback. arXiv preprint arXiv:2302.12813
2023 arXiv
-
[51]
Mansi Phute, Alec Helbling, Matthew Hull, ShengYun Peng, Sebastian Szyller, Cory Cornelius, and Duen Horng Chau. 2023. Llm self defense: By self examination, llms know they are being tricked. arXiv preprint arXiv:2308.07308
2023 arXiv
-
[52]
Julien Piet, Maha Alrashed, Chawin Sitawarin, Sizhe Chen, Zeming Wei, Elizabeth Sun, Basel Alomair, and David Wagner. 2024. Jatmo: Prompt injection defense by task-specific finetuning. In European Symposium on Research in Computer Security, pages 105--124. Springer
2024
-
[53]
Aman Rangapur, Haoran Wang, and Kai Shu. 2023. Investigating online financial misinformation and its consequences: A computational perspective. arXiv preprint arXiv:2309.12363
2023 arXiv
-
[54]
Sippo Rossi, Alisia Marianne Michel, Raghava Rao Mukkamala, and Jason Bennett Thatcher. 2024. An early categorization of prompt injection attacks on large language models. arXiv preprint arXiv:2402.00898
2024 arXiv
-
[55]
Shrey Satapara, Parth Mehta, Debasis Ganguly, and Sandip Modha. 2024. Fighting fire with fire: Adversarial prompting to generate a misinformation detection dataset. arXiv preprint arXiv:2401.04481
2024 arXiv
-
[56]
Sander Schulhoff, Michael Ilie, Nishant Balepur, Konstantine Kahadze, Amanda Liu, Chenglei Si, Yinheng Li, Aayush Gupta, HyoJung Han, Sevien Schulhoff, et al. 2024. The prompt report: A systematic survey of prompting techniques. arXiv preprint arXiv:2406.06608
2024 arXiv
-
[57]
Weijia Shi, Xiaochuang Han, Mike Lewis, Yulia Tsvetkov, Luke Zettlemoyer, and Wen-tau Yih. 2024. Trusting your evidence: Hallucinate less with context-aware decoding. In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Lingu...
2024
-
[58]
Chenmien Tan, Ge Zhang, and Jie Fu. 2023. Massive editing for large language models via meta learning. arXiv preprint arXiv:2311.04661
2023 arXiv
-
[59]
Gemini Team, Petko Georgiev, Ving Ian Lei, Ryan Burnell, Libin Bai, Anmol Gulati, Garrett Tanzer, Damien Vincent, Zhufeng Pan, Shibo Wang, et al. 2024. Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context. arXiv preprint arXiv:2403.05530
2024 arXiv
-
[60]
Tu Vu, Mohit Iyyer, Xuezhi Wang, Noah Constant, Jerry Wei, Jason Wei, Chris Tar, Yun-Hsuan Sung, Denny Zhou, Quoc Le, et al. 2023. Freshllms: Refreshing large language models with search engine augmentation. arXiv preprint arXiv:2310.03214
2023 arXiv
-
[61]
David Wan, Mengwen Liu, Kathleen Mckeown, Markus Dreyer, and Mohit Bansal. 2023. Faithfulness-aware decoding strategies for abstractive summarization. In Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics, pages 2864--2880
2023
-
[62]
Yizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu, Noah A Smith, Daniel Khashabi, and Hannaneh Hajishirzi. 2022. Self-instruct: Aligning language models with self-generated instructions. arXiv preprint arXiv:2212.10560
2022 arXiv
-
[63]
Jialiang Wu, Yi Shen, Sijia Liu, Yi Tang, Sen Song, Xiaoyi Wang, and Longjun Cai. 2025. Improve decoding factuality by token-wise cross layer entropy of large language models. arXiv preprint arXiv:2502.03199
2025 arXiv
-
[64]
Jiaying Wu, Jiafeng Guo, and Bryan Hooi. 2024 a . Fake news in sheep's clothing: Robust fake news detection against llm-empowered style attacks. In Proceedings of the 30th ACM SIGKDD conference on knowledge discovery and data mining, pages 3367--3378
2024
-
[65]
Shangyu Wu, Ying Xiong, Yufei Cui, Haolun Wu, Can Chen, Ye Yuan, Lianming Huang, Xue Liu, Tei-Wei Kuo, Nan Guan, et al. 2024 b . Retrieval-augmented generation for natural language processing: A survey. arXiv preprint arXiv:2407.13193
2024 arXiv
-
[66]
Keyang Xuan, Li Yi, Fan Yang, Ruochen Wu, Yi R Fung, and Heng Ji. 2024. Lemma: Towards lvlm-enhanced multimodal misinformation detection with external knowledge augmentation. arXiv preprint arXiv:2402.11943
2024 arXiv
-
[67]
Boyang Xue, Fei Mi, Qi Zhu, Hongru Wang, Rui Wang, Sheng Wang, Erxin Yu, Xuming Hu, and Kam-Fai Wong. 2024. Ualign: Leveraging uncertainty estimations for factuality alignment on large language models. arXiv preprint arXiv:2412.11803
2024 arXiv
-
[68]
Shi-Qi Yan, Jia-Chen Gu, Yun Zhu, and Zhen-Hua Ling. 2024. Corrective retrieval augmented generation. arXiv preprint arXiv:2401.15884
2024 arXiv
-
[69]
Zhiyuan Yu, Xiaogeng Liu, Shunning Liang, Zach Cameron, Chaowei Xiao, and Ning Zhang. 2024. Don't listen to me: Understanding and exploring jailbreak prompts of large language models. arXiv preprint arXiv:2403.17336
2024 arXiv
-
[70]
Hongbang Yuan, Yubo Chen, Pengfei Cao, Zhuoran Jin, Kang Liu, and Jun Zhao. 2024. Beyond under-alignment: Atomic preference enhanced factuality tuning for large language models. arXiv preprint arXiv:2406.12416
2024 arXiv
-
[71]
Zhenrui Yue, Huimin Zeng, Yimeng Lu, Lanyu Shang, Yang Zhang, and Dong Wang. 2024. Evidence-driven retrieval augmented response generation for online misinformation. arXiv preprint arXiv:2403.14952
2024 arXiv
-
[72]
Hanning Zhang, Shizhe Diao, Yong Lin, Yi Fung, Qing Lian, Xingyao Wang, Yangyi Chen, Heng Ji, and Tong Zhang. 2024 a . R-tuning: Instructing large language models to say ‘i don’t know’. In Proceedings of the 2024 Conference of the North American Chapter of the Association for ...
2024
-
[73]
Jianyi Zhang, Da-Cheng Juan, Cyrus Rashtchian, Chun-Sung Ferng, Heinrich Jiang, and Yiran Chen. 2024 b . Sled: Self logits evolution decoding for improving factuality in large language models. arXiv preprint arXiv:2411.02433
2024 arXiv
-
[74]
Xiaoying Zhang, Baolin Peng, Ye Tian, Jingyan Zhou, Lifeng Jin, Linfeng Song, Haitao Mi, and Helen Meng. 2024 c . Self-alignment for factuality: Mitigating hallucinations in llms via self-evaluation. arXiv preprint arXiv:2402.09267
2024 arXiv
-
[75]
Xuan Zhang and Wei Gao. 2023. Towards llm-based fact verification on news claims with a hierarchical step-by-step prompting method. arXiv preprint arXiv:2310.00305
2023 arXiv
-
[76]
Yue Zhang, Leyang Cui, Wei Bi, and Shuming Shi. 2023. Alleviating hallucinations of large language models through induced hallucinations. arXiv preprint arXiv:2312.15710
2023 arXiv
-
[77]
Andy Zou, Zifan Wang, Nicholas Carlini, Milad Nasr, J Zico Kolter, and Matt Fredrikson. 2023. Universal and transferable adversarial attacks on aligned language models. arXiv preprint arXiv:2307.15043
2023 arXiv
-
[78]
online" 'onlinestring :=
ENTRY address archivePrefix author booktitle chapter edition editor eid eprint eprinttype howpublished institution journal key month note number organization pages publisher school series title type volume year doi pubmed url lastchecked label extra.label sort.label short.list...
-
[79]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 gl...
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.