REVIEW 5 cited by
What Evidence Do Language Models Find Convincing?
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Retrieval-augmented language models are being increasingly tasked with subjective, contentious, and conflicting queries such as "is aspartame linked to cancer". To resolve these ambiguous queries, one must search through a large range of websites and consider "which, if any, of this evidence do I find convincing?". In this work, we study how LLMs answer this question. In particular, we construct ConflictingQA, a dataset that pairs controversial queries with a series of real-world evidence documents that contain different facts (e.g., quantitative results), argument styles (e.g., appeals to authority), and answers (Yes or No). We use this dataset to perform sensitivity and counterfactual analyses to explore which text features most affect LLM predictions. Overall, we find that current models rely heavily on the relevance of a website to the query, while largely ignoring stylistic features that humans find important such as whether a text contains scientific references or is written with a neutral tone. Taken together, these results highlight the importance of RAG corpus quality (e.g., the need to filter misinformation), and possibly even a shift in how LLMs are trained to better align with human judgements.
Forward citations
Cited by 5 Pith papers
-
Generative Engine Optimization: How to Dominate AI Search
Across hundreds of query comparisons, AI search engines systematically favor earned media over brand-owned and social sources, and vary strongly by engine and language.
-
ImportSnare: Directed "Code Manual" Hijacking in Retrieval-Augmented Code Generation
Documentation poisoning with hidden ranking and suggestion sequences can make RAG-based code generators confidently recommend malicious dependencies, even at 0.01% poisoning ratios.
-
Do as We Do, Not as You Think: the Conformity of Large Language Models
In multi-agent question answering, all 11 tested LLMs abandon some correct answers to follow a unanimous wrong majority, with the effect strongest under the Doubt protocol.
-
Evaluating Large Language Models for Evidence-Based Clinical Question Answering
A new multi-source clinical QA benchmark shows GPT models answer structured guideline questions best, perform worse on narrative texts and systematic reviews, and improve with retrieved abstracts.
-
Lies, Damned Lies, and Distributional Language Statistics: Persuasion and Deception with Large Language Models
Current LLMs are about as persuasive as humans, can deceive strategically, but show small effect sizes; future techniques may increase this risk.
Discussion (0). Continue with ORCID to comment.