FluidML combines graph splitting, dynamic programming, and greedy memory allocation to optimize ML inference memory layout, but its reported improvements are inconsistent across models and its headline numbers contradict its own tables.
GPT-NeoX: Large Scale Autoregressive Language Modeling in PyTorch , 9 2023
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.LG 1years
2024 1verdicts
REJECT 1representative citing papers
citing papers explorer
-
FluidML: Fast and Memory Efficient Inference Optimization
FluidML combines graph splitting, dynamic programming, and greedy memory allocation to optimize ML inference memory layout, but its reported improvements are inconsistent across models and its headline numbers contradict its own tables.