Pith. sign in

REVIEW 4 cited by

Exploring Training on Heterogeneous Data with Mixture of Low-rank Adapters

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.09679 v1 pith:333KXOOJ submitted 2024-06-14 cs.CV

classification cs.CV
keywords adapterslow-rankmolatrainingconflictsdataheterogeneousmitigate
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Training a unified model to take multiple targets into account is a trend towards artificial general intelligence. However, how to efficiently mitigate the training conflicts among heterogeneous data collected from different domains or tasks remains under-explored. In this study, we explore to leverage Mixture of Low-rank Adapters (MoLA) to mitigate conflicts in heterogeneous data training, which requires to jointly train the multiple low-rank adapters and their shared backbone. Specifically, we introduce two variants of MoLA, namely, MoLA-Grad and MoLA-Router, to respectively handle the target-aware and target-agnostic scenarios during inference. The former uses task identifiers to assign personalized low-rank adapters to each task, disentangling task-specific knowledge towards their adapters, thereby mitigating heterogeneity conflicts. The latter uses a novel Task-wise Decorrelation (TwD) loss to intervene the router to learn oriented weight combinations of adapters to homogeneous tasks, achieving similar effects. We conduct comprehensive experiments to verify the superiority of MoLA over previous state-of-the-art methods and present in-depth analysis on its working mechanism. Source code is available at: https://github.com/MediaBrain-SJTU/MoLA

Discussion (0). Sign in to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. GeMoE: Gating Entropy is All You Need for Uncertainty-aware Adaptive Routing in MoE-based Large Vision-Language Models

    cs.CV 2026-06 unverdicted novelty 7.0 of 10

    GeMoE adaptively sets the number of experts per token via gating entropy, retaining 99.5% of static-routing performance while raising average sparsity by 36.5%.

  2. Dementia Etiology Diagnosis via Collaborative Meta Knowledge Enhancement

    cs.LG 2026-07 conditional novelty 6.0 of 10

    A Transformer that injects site, scan-sequence, and modality-availability embeddings, regularized by a reference-model trust region, improves multi-center dementia subtype classification (mean macro AUC 85.62%).

  3. Parametric Memory Decoding for Zero-Shot Routing in LoRA-Based External Parametric Memory

    cs.LG 2026-07 conditional novelty 6.0 of 10

    PMDRouter selects LoRAs zero-shot by decoding scale-normalized linear response energy from one adapter-free backbone prefill, and leads most internal-signal baselines on a new multi-granularity EPM bench.

  4. Towards Unified Multi-task EEG Analysis with Low-Rank Adaptation

    cs.LG 2026-04 unverdicted novelty 4.0 of 10

    MTEEG uses task-specific LoRA modules to jointly adapt a pre-trained EEG model across multiple tasks, outperforming single-task baselines on most metrics in evaluations on six downstream tasks.

Pith tools