A new 698-pair human-annotated benchmark and inconsistency typology for political language, with LLMs roughly matching individual annotators on coarse detection but not on fine-grained subtypes.
Prospects for inconsistency detection using large language models and sheaves
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
We demonstrate that large language models can produce reasonable numerical ratings of the logical consistency of claims. We also outline a mathematical approach based on sheaf theory for lifting such ratings to hypertexts such as laws, jurisprudence, and social media and evaluating their consistency globally. This approach is a promising avenue to increasing consistency in and of government, as well as to combating mis- and disinformation and related ills.
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Misleading through Inconsistency: A Benchmark for Political Inconsistencies Detection
A new 698-pair human-annotated benchmark and inconsistency typology for political language, with LLMs roughly matching individual annotators on coarse detection but not on fine-grained subtypes.