Pith. sign in

REVIEW 3 cited by

ConflictBank: A Benchmark for Evaluating the Influence of Knowledge Conflicts in LLM

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2408.12076 v1 pith:ZGGOVILZ submitted 2024-08-22 cs.CL cs.AI

classification cs.CLcs.AI
keywords conflictsknowledgeconflictllmsbenchmarkconflictbankmodelmodels
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Large language models (LLMs) have achieved impressive advancements across numerous disciplines, yet the critical issue of knowledge conflicts, a major source of hallucinations, has rarely been studied. Only a few research explored the conflicts between the inherent knowledge of LLMs and the retrieved contextual knowledge. However, a thorough assessment of knowledge conflict in LLMs is still missing. Motivated by this research gap, we present ConflictBank, the first comprehensive benchmark developed to systematically evaluate knowledge conflicts from three aspects: (i) conflicts encountered in retrieved knowledge, (ii) conflicts within the models' encoded knowledge, and (iii) the interplay between these conflict forms. Our investigation delves into four model families and twelve LLM instances, meticulously analyzing conflicts stemming from misinformation, temporal discrepancies, and semantic divergences. Based on our proposed novel construction framework, we create 7,453,853 claim-evidence pairs and 553,117 QA pairs. We present numerous findings on model scale, conflict causes, and conflict types. We hope our ConflictBank benchmark will help the community better understand model behavior in conflicts and develop more reliable LLMs.

Discussion (0). Sign in to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Where Facts Go Missing: A Layerwise Taxonomy and Per-Layer Attribution of Information Omission in Air-Gapped LLMAgent Pipelines

    cs.MA 2026-07 conditional novelty 6.0 of 10

    In a controlled 75,476-trial stress test, about 73% of omitted-fact failures in LLM agent pipelines are traced to deterministic middleware (redaction, pagination, truncation) rather than model behavior.

  2. Presentation, Not Mechanism: A Render Confound in Deprecation-Aware Memory Evaluation

    cs.LG 2026-07 accept novelty 6.0 of 10

    A render-matched control explains away almost all (+0.159 of +0.184) of a fine-grained revision-ledger's apparent advantage over a flat baseline, leaving a near-zero mechanism residual and making coarse invalidation t...

  3. REAL: Reading Out Transformer Activations for Precise Localization in Language Model Steering

    cs.CL 2025-06 conditional novelty 6.0 of 10

    A VQ-AE-based module-scoring method picks steering locations in LLMs, improving truthfulness and knowledge-selection steering over ITI and SPARE baselines.

Pith tools