Self-consistent errors, where an LLM confidently repeats the same wrong answer, persist with model scale and evade all four mainstream detectors; a cross-model probe using another LLM's hidden states improves detection.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Too Consistent to Detect: A Study of Self-Consistent Errors in LLMs
Self-consistent errors, where an LLM confidently repeats the same wrong answer, persist with model scale and evade all four mainstream detectors; a cross-model probe using another LLM's hidden states improves detection.