Chiron's hierarchical backpressure autoscaler, which queues batch requests and adapts batch sizes dynamically, improves SLO attainment and GPU efficiency for LLM serving.
INFaaS: Automated model-less in- ference serving
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.DC 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Hierarchical Autoscaling for Large Language Model Serving with Chiron
Chiron's hierarchical backpressure autoscaler, which queues batch requests and adapts batch sizes dynamically, improves SLO attainment and GPU efficiency for LLM serving.