In GPT-2 multi-document QA, the layer gap between the first correct top-1 token prediction and its stable final form is larger when relevant information is in the middle of the context.
com/posts/AcKRB8wDpdaN6v6ru/interpreting-gpt-the-logit-lens
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Unraveling Token Prediction Refinement and Identifying Essential Layers in Language Models
In GPT-2 multi-document QA, the layer gap between the first correct top-1 token prediction and its stable final form is larger when relevant information is in the middle of the context.