A classifier that decides which hidden states to patch during inference improves 2-hop question answering from 18.45% to 23.63% on MuSiQue, but the evaluation uses the same prompts that trained the classifier.
Patchscopes: A unifying framework for inspecting hidden representations of language models, 2024
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2025 1verdicts
REJECT 1representative citing papers
citing papers explorer
-
Auto-Patching: Enhancing Multi-Hop Reasoning in Language Models
A classifier that decides which hidden states to patch during inference improves 2-hop question answering from 18.45% to 23.63% on MuSiQue, but the evaluation uses the same prompts that trained the classifier.