SlideFormer uses layer-sliding async offloading, pre-allocated heterogeneous memory, and fused Triton kernels to fine-tune 123B+ models on one RTX 4090 with 1.4–6.3× higher throughput and roughly half the memory of prior offload systems.
Title resolution pending
1 Pith paper cite this work, alongside 33 external citations. Polarity classification is still indexing.
1
Pith paper citing it
33
external citations · OpenAlex
fields
cs.DC 1years
2026 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
An Efficient Heterogeneous Co-Design for Fine-Tuning on a Single GPU
SlideFormer uses layer-sliding async offloading, pre-allocated heterogeneous memory, and fused Triton kernels to fine-tune 123B+ models on one RTX 4090 with 1.4–6.3× higher throughput and roughly half the memory of prior offload systems.