Pith. sign in

REVIEW 3 cited by

Sparse Mixture-of-Experts are Domain Generalizable Learners

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2206.04046 v6 pith:KO52G54Y submitted 2022-06-08 cs.CV cs.AIcs.LG

classification cs.CVcs.AIcs.LG
keywords gmoemodelstrainedalgorithmsarchitecturedesigndomainempirical
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Human visual perception can easily generalize to out-of-distributed visual data, which is far beyond the capability of modern machine learning models. Domain generalization (DG) aims to close this gap, with existing DG methods mainly focusing on the loss function design. In this paper, we propose to explore an orthogonal direction, i.e., the design of the backbone architecture. It is motivated by an empirical finding that transformer-based models trained with empirical risk minimization (ERM) outperform CNN-based models employing state-of-the-art (SOTA) DG algorithms on multiple DG datasets. We develop a formal framework to characterize a network's robustness to distribution shifts by studying its architecture's alignment with the correlations in the dataset. This analysis guides us to propose a novel DG model built upon vision transformers, namely Generalizable Mixture-of-Experts (GMoE). Extensive experiments on DomainBed demonstrate that GMoE trained with ERM outperforms SOTA DG baselines by a large margin. Moreover, GMoE is complementary to existing DG methods and its performance is substantially improved when trained with DG algorithms.

Discussion (0). Sign in to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. DEER: Disentangled Mixture of Experts with Instance-Adaptive Routing for Generalizable Machine-Generated Text Detection

    cs.CL 2025-11 conditional novelty 6.0 of 10

    DEER, a disentangled mixture-of-experts detector with RL-based instance routing, reports F1 gains of about 1.4 in-domain and 5.3 points out-of-domain over prior MGT detectors.

  2. Generalizable Multispectral Land Cover Classification via Frequency-Aware Mixture of Low-Rank Token Experts

    cs.CV 2025-05 conditional novelty 6.0 of 10

    A parameter-efficient adapter with low-rank token experts and frequency-aware filtering improves domain generalization for multispectral land cover classification with frozen vision foundation models.

  3. Can Experts Adapt Without Training? On Test-Time Modality Generalization in MVLMs

    cs.CV 2026-07 conditional novelty 5.0 of 10

    A training-free test-time adaptation method (MoBE) routes between modality experts by entropy and adapts their prototypes/priors online, improving medical VLM accuracy by 4.3–7.2 points across benchmarks.

Pith tools