A 0.1B-parameter model, trained by distillation from a 10B teacher plus SFT, DPO, and inference optimizations, matches a 10B model on query-focused webpage summarization and runs at about 50,000 queries per second.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Leveraging Generative Models for Real-Time Query-Driven Text Summarization in Large-Scale Web Search
A 0.1B-parameter model, trained by distillation from a 10B teacher plus SFT, DPO, and inference optimizations, matches a 10B model on query-focused webpage summarization and runs at about 50,000 queries per second.