A targeted KL penalty on the second-ranked expert plus a jointly trained dual-model blend improves adversarial robustness of mixture-of-experts classifiers with little clean-accuracy loss.
Supplementary Material In this supplementary material, we provide the proofs of Theorems 5.4 and 5.5 in Section A.1
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.LG 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Optimizing Robustness and Accuracy in Mixture of Experts: A Dual-Model Approach
A targeted KL penalty on the second-ranked expert plus a jointly trained dual-model blend improves adversarial robustness of mixture-of-experts classifiers with little clean-accuracy loss.