Distilling a 7B LLM teacher into a BERT-base student with Margin-MSE loss on 170M teacher-labeled pairs yields a small student that matches or slightly beats the teacher on NDCG and improves Walmart's tail-query search metrics online.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.IR 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Knowledge Distillation for Enhancing Walmart E-commerce Search Relevance Using Large Language Models
Distilling a 7B LLM teacher into a BERT-base student with Margin-MSE loss on 170M teacher-labeled pairs yields a small student that matches or slightly beats the teacher on NDCG and improves Walmart's tail-query search metrics online.