An inference-aware scaling law that includes model aspect ratio ranks model shapes by loss and latency, producing a 1B model that is 1.8x faster without losing accuracy.
dmodel is the hidden size, fsize is the intermediate size,n layers is the number of layers, andn heads is the number of attention heads
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.LG 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Scaling Inference-Efficient Language Models
An inference-aware scaling law that includes model aspect ratio ranks model shapes by loss and latency, producing a 1B model that is 1.8x faster without losing accuracy.