A VM-deployable controller combining dynamic MIG, PCIe-aware placement, and I/O guardrails reduces SLO miss rate by about 32 percent and p99 latency by about 15 percent at under 5 percent throughput cost on a 16-GPU A100 cluster.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.DC 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Predictable LLM Serving on GPU Clusters
A VM-deployable controller combining dynamic MIG, PCIe-aware placement, and I/O guardrails reduces SLO miss rate by about 32 percent and p99 latency by about 15 percent at under 5 percent throughput cost on a 16-GPU A100 cluster.