D2FT dynamically schedules attention subnets to full, forward-only, or skipped operations during fine-tuning, cutting compute 40% and communication 50% with 1-2% top-1 accuracy loss on three vision datasets, plus a LoRA variant.
Train big, then compress: Rethinking model size for efficient training and inference of transformers,
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.LG 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
You Don't Need All Attentions: Distributed Dynamic Fine-Tuning for Foundation Models
D2FT dynamically schedules attention subnets to full, forward-only, or skipped operations during fine-tuning, cutting compute 40% and communication 50% with 1-2% top-1 accuracy loss on three vision datasets, plus a LoRA variant.