REVIEW 12 cited by
LM vs LM: Detecting Factual Errors via Cross Examination
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
A prominent weakness of modern language models (LMs) is their tendency to generate factually incorrect text, which hinders their usability. A natural question is whether such factual errors can be detected automatically. Inspired by truth-seeking mechanisms in law, we propose a factuality evaluation framework for LMs that is based on cross-examination. Our key idea is that an incorrect claim is likely to result in inconsistency with other claims that the model generates. To discover such inconsistencies, we facilitate a multi-turn interaction between the LM that generated the claim and another LM (acting as an examiner) which introduces questions to discover inconsistencies. We empirically evaluate our method on factual claims made by multiple recent LMs on four benchmarks, finding that it outperforms existing methods and baselines, often by a large gap. Our results demonstrate the potential of using interacting LMs for capturing factual errors.
Forward citations
Cited by 12 Pith papers
-
A Large-Scale Study of Relevance Assessments with Large Language Models: An Initial Look
Fully automatic LLM relevance judgments rank TREC 2024 retrieval runs as well as human judgments, and human-in-the-loop assistance provides no clear extra benefit.
-
HalluField: Detecting LLM Hallucinations via Field-Theoretic Modeling
HalluField flags LLM hallucinations using a hand-weighted temperature-perturbation of token-level 'free energy' (negative log-likelihood) and Shannon entropy.
-
Can Large Language Models Integrate Spatial Data? Empirical Insights into Reasoning Strengths and Computational Weaknesses
LLMs only become competitive at spatial data integration when given pre-computed geometric features; a review-and-refine prompt then exceeds hand-tuned heuristics.
-
Shaking to Reveal: Perturbation-Based Detection of LLM Hallucinations
SSP adds a learned, sample-specific noise prompt to an LLM input and scores hallucination by the cosine shift in intermediate representations, outperforming output-confidence baselines on QA benchmarks.
-
Detection and Mitigation of Hallucination in Large Reasoning Models: A Mechanistic Perspective
A logit-lens divergence score over late layers is used to detect hallucinated reasoning traces and to shape reinforcement learning rewards, with results on math, science, and multi-hop QA benchmarks.
-
When One LLM Drools, Multi-LLM Collaboration Rules
A position paper that introduces a four-level taxonomy of multi-LLM collaboration (API, text, logit, weight) and argues it is essential for reliability, pluralism, and democratization.
-
ChartInsighter: An Approach for Mitigating Hallucination in Time-series Chart Summary Generation with A Benchmark Dataset
A multi-agent LLM pipeline with external computation and self-consistency checking produces time-series chart summaries with fewer annotated hallucinations than GPT-4 or VL2NL on the authors' new benchmark.
-
Critical-Questions-of-Thought: Steering LLM reasoning with Argumentative Querying
CQoT, a pipeline that uses argumentation-theoretic critical questions to check LLM reasoning plans, improves MT-Bench reasoning and math scores by roughly 5% over baseline and CoT prompting.
-
Ontology-Constrained Generation of Domain-Specific Clinical Summaries
An ontology-guided constrained decoding method produces specialty-specific clinical summaries and lowers hallucination scores on MIMIC-III relative to greedy and beam search baselines.
-
Recon, Answer, Verify: Agents in Search of Truth
Removing annotator cues from fact-checking evidence lowers LLM scores substantially, and a three-agent question-answering pipeline, RAV, outperforms several published fact-checking baselines.
-
Counterfactual Probing for Hallucination Detection and Mitigation in Large Language Models
Counterfactual Probing detects LLM hallucinations by measuring how much a model's confidence changes when a claim is altered to a plausible but incorrect variant, then hedges flagged statements with template-based mit...
-
Label-Confidence-Aware Uncertainty Estimation in Natural Language Generation
A label-confidence-aware KL-divergence score between sampled-set Gibbs probability and greedy-answer probability improves uncertainty estimation AUROC on several QA datasets.
Discussion (0). Continue with ORCID to comment.