Differentially private pretraining of linear attention heads for in-context linear regression has excess risk that decays like 1/(N L^3) in low dimensions and D^2/(N^2 L^2) in high dimensions, up to log factors and privacy parameters.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
stat.ML 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
How Private is Your Attention? Bridging Privacy with In-Context Learning
Differentially private pretraining of linear attention heads for in-context linear regression has excess risk that decays like 1/(N L^3) in low dimensions and D^2/(N^2 L^2) in high dimensions, up to log factors and privacy parameters.