Pith. sign in

Inference economics of language models

8 Pith papers cite this work. Polarity classification is still indexing.

8 Pith papers citing it
abstract

We develop a theoretical model that addresses the economic trade-off between cost per token versus serial token generation speed when deploying LLMs for inference at scale. Our model takes into account arithmetic, memory bandwidth, network bandwidth and latency constraints; and optimizes over different parallelism setups and batch sizes to find the ones that optimize serial inference speed at a given cost per token. We use the model to compute Pareto frontiers of serial speed versus cost per token for popular language models.

citation-role summary

background 3

citation-polarity summary

years

2026 7 2025 1

roles

background 3

polarities

background 3

representative citing papers

Think Before You Grid-Search: Floor-First Triage for LLM Serving

cs.PF · 2026-07-07 · conditional · novelty 6.0 · 2 refs

LLM serving should triage by five-resource analytical floors and wall ordering, not grid search; on 16×H20, TP16 is capacity-capped at ~70 while EP+DP attention reaches ~644 concurrent 8K requests.

citing papers explorer

Showing 8 of 8 citing papers.