TokenTune reduces the memory needed for fine-tuning transformers by computing gradients only for a random subset of tokens, cutting activation cache size with little accuracy loss.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Memory-Efficient Fine-Tuning of Transformers via Token Selection
TokenTune reduces the memory needed for fine-tuning transformers by computing gradients only for a random subset of tokens, cutting activation cache size with little accuracy loss.