TRAPSBench shows that across 16 vision-language models, answerability is decodable from hidden states while spontaneous abstention remains poor, pointing to an output-stage bottleneck in epistemic restraint.
International conference on learning representations , year=
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CV 1years
2026 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
TRAPSBench: Vision-Language Models Encode but Fail to Express Epistemic Restraint
TRAPSBench shows that across 16 vision-language models, answerability is decodable from hidden states while spontaneous abstention remains poor, pointing to an output-stage bottleneck in epistemic restraint.