In-context learning is reframed as implicit knowledge distillation, but the main claims either restate the known attention-equals-gradient-descent result or build the prompt-shift bound into the definition of MMD.
Rademacher and gaussian complexities: Risk bounds and structural results
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.LG 1years
2025 1verdicts
REJECT 1representative citing papers
citing papers explorer
-
Brewing Knowledge in Context: Distillation Perspectives on In-Context Learning
In-context learning is reframed as implicit knowledge distillation, but the main claims either restate the known attention-equals-gradient-descent result or build the prompt-shift bound into the definition of MMD.