A frozen Qwen2-1.5B teacher passes its hidden states through gated cross-attention into GPT-Neo-125M, which after 15 epochs generates more coherent arithmetic responses than the base small model.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2025 1verdicts
REJECT 1representative citing papers
citing papers explorer
-
LLM Modules: Knowledge Transfer from a Large to a Small Model using Enhanced Cross-Attention
A frozen Qwen2-1.5B teacher passes its hidden states through gated cross-attention into GPT-Neo-125M, which after 15 epochs generates more coherent arithmetic responses than the base small model.