REVIEW 5 cited by
Residual Mixture of Experts
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Mixture of Experts (MoE) is able to scale up vision transformers effectively. However, it requires prohibiting computation resources to train a large MoE transformer. In this paper, we propose Residual Mixture of Experts (RMoE), an efficient training pipeline for MoE vision transformers on downstream tasks, such as segmentation and detection. RMoE achieves comparable results with the upper-bound MoE training, while only introducing minor additional training cost than the lower-bound non-MoE training pipelines. The efficiency is supported by our key observation: the weights of an MoE transformer can be factored into an input-independent core and an input-dependent residual. Compared with the weight core, the weight residual can be efficiently trained with much less computation resource, e.g., finetuning on the downstream data. We show that, compared with the current MoE training pipeline, we get comparable results while saving over 30% training cost. When compared with state-of-the-art non- MoE transformers, such as Swin-T / CvT-13 / Swin-L, we get +1.1 / 0.9 / 1.0 mIoU gain on ADE20K segmentation and +1.4 / 1.6 / 0.6 AP gain on MS-COCO object detection task with less than 3% additional training cost.
Forward citations
Cited by 5 Pith papers
-
LiMoE: Mixture of LiDAR Representation Learners from Automotive Scenes
LiMoE fuses range, voxel, and point representations via MoE gating for LiDAR pretraining, improving label-efficient semantic segmentation on 11 datasets.
-
UniRestorer: Universal Image Restoration via Adaptively Estimating Image Degradation at Proper Granularity
A multi-granularity mixture-of-experts image restoration model that routes each degraded image to an expert using both degradation and granularity estimates, outperforming all-in-one baselines.
-
Mr. DETR++: Instructive Multi-Route Training for Detection Transformers with Mixture-of-Experts
Multi-route training with instructive self-attention tokens and a route-aware mixture-of-experts raises detection mAP by 2 to 4 points across several DETR baselines at no inference cost.
-
SkySense V2: A Unified Foundation Model for Multi-modal Remote Sensing
A unified multi-modal remote sensing foundation model with adaptive patch merging, modality prompt tokens, mixture of experts, and query-based semantic aggregation contrastive learning outperforms SkySense by 1.8 poin...
-
YOLO-Master: MOE-Accelerated with Specialized Transformers for Enhanced Real-time Detection
YOLO-Master inserts a sparse Mixture-of-Experts block into a YOLO backbone, reporting 42.4% COCO AP at 1.62 ms, +0.8 AP and 18% faster than YOLOv13-N.
Discussion (0). Continue with ORCID to comment.