A simulator-plus-ML pipeline places LoRA adapters onto GPUs so a given workload needs fewer GPUs (60% claimed on average) without request starvation or memory errors.
Fast algorithms for bin packing
1 Pith paper cite this work, alongside 536 external citations. Polarity classification is still indexing.
1
Pith paper citing it
536
external citations · OpenAlex
fields
cs.DC 1years
2026 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Data Driven Optimization of GPU efficiency for Distributed LLM-Adapter Serving
A simulator-plus-ML pipeline places LoRA adapters onto GPUs so a given workload needs fewer GPUs (60% claimed on average) without request starvation or memory errors.