Merging two GPT-2 variants with B-spline-blended hidden states plus autoencoders yields a single model that keeps both English and French perplexity closer to the best expert than linear interpolation.
Outrageously large neural net- works: The sparsely-gated mixture-of-experts layer
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
background 1
citation-polarity summary
fields
cs.CL 1years
2024 1verdicts
CONDITIONAL 1roles
background 1polarities
unclear 1representative citing papers
citing papers explorer
-
Superposition in Transformers: A Novel Way of Building Mixture of Experts
Merging two GPT-2 variants with B-spline-blended hidden states plus autoencoders yields a single model that keeps both English and French perplexity closer to the best expert than linear interpolation.