TokenTune reduces the memory needed for fine-tuning transformers by computing gradients only for a random subset of tokens, cutting activation cache size with little accuracy loss.
In Findings of the As- sociation for Computational Linguistics, ACL 2024, Bangkok, Thailand and virtual meeting, August 11- 16, 2024, pages 467–484
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Memory-Efficient Fine-Tuning of Transformers via Token Selection
TokenTune reduces the memory needed for fine-tuning transformers by computing gradients only for a random subset of tokens, cutting activation cache size with little accuracy loss.