Pith. sign in

REVIEW 4 major objections 5 minor 2 cited by

Ensembles of Low-Rank Expert Adapters

T0 review · 4 major / 5 minor · reviewed 2026-08-09 · deepseek-v4-flash

Pith's one-line read Clustering fine-tuning instructions by gradient direction and training one LoRA expert per cluster beats a single full-data LoRA adapter without extra data or task-specific validation.

desk verdict The ensemble trick works, but the paper's own ablations contradict the claim that gradient clustering is what makes it work. read the letter →

arxiv 2502.00089 v1 pith:LND3R6LC submitted 2025-01-31 cs.CL cs.AIcs.LG

classification cs.CLcs.AIcs.LG
keywords low-rankadaptationgradient-basedclusteringmixtureofexpertsdeepensemblesinstructiontuningparameter-efficientfine-tuningdataselectionlanguagemodel
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

ELREA is a fine-tuning scheme that replaces one full-data LoRA adapter with a small team of expert adapters, each trained on a cluster of training instructions whose gradients point in similar directions. The paper's central claim is that this arrangement reduces conflicting gradient updates during fine-tuning and, at inference, lets the model route each new instruction to the adapters whose gradient profiles match it. On the MATH-Combined benchmark, the paper reports relative accuracy gains of 9.67% (rank 8) and 3.56% (rank 64) over the full-data LoRA baseline, with comparable gains on BBH and MMLU. The method is designed to be task-agnostic: no task-specific validation set or data features are needed, and the overhead is roughly 2.3x training time and 2.4x inference memory. If the claim is right, ELREA is a drop-in LoRA-based way to turn a diverse fine-tuning corpus into better downstream accuracy.

What carries the argument

The central object is the gradient feature $\delta(x)$ of an instruction: the flattened, epoch-averaged, normalized gradient of the next-token loss with respect to a LoRA adapter's parameters, projected with a random $\{-1,1\}$ matrix to 8,192 dimensions. This single representation does double duty: BIRCH clustering on it defines the expert training clusters, and cosine similarity to the resulting centroids, standardized and softmaxed, sets the expert weights at inference. The second load-bearing piece is the ensemble decoder of Equation 10, which sums the logits of the base adapter and all cluster experts and takes the argmax, so routing happens once per prompt and is reused across generated tokens.

What would settle it

Run ELREA on MATH-Combined at rank 8 with the cluster routing weights randomly permuted across clusters; if accuracy stays near ELREA's 20.41% rather than falling toward the Uniform Weights baseline at 20.16% or the full-data baseline at 18.61%, the gradient-routing signal is not carrying the reported gain.

Watch

Extended reading notes

Core claim

On its own terms, the paper establishes a recipe: fine-tune a base LoRA adapter on the full dataset; compute a per-instruction gradient feature from the instruction tokens only, averaged over epochs, normalized, and randomly projected to 8,192 dimensions; cluster those features with BIRCH into 8-10 groups; train one LoRA expert per cluster starting from the base adapter; and at inference weight the experts by the cosine similarity between the test instruction's gradient feature and each cluster centroid, standardized and passed through softmax, with the base adapter carrying a residual weight when the test point is far from every cluster. Across MATH-Combined and the BBH/MMLU general-reasoning tasks, the paper reports that ELREA outperforms the full-data baseline, dataset-specific adapters, MoE routing and merging, a trainable MoLE router, self-consistency, instruction-embedding clustering, and uniform-weight and random-cluster ablations, while noting that the advantage over random clustering is real but not always large.

Load-bearing premise

The load-bearing premise is that the cosine similarity between a test instruction's gradient (computed with respect to the base adapter and projected to 8,192 dimensions) and a training-cluster centroid predicts which expert adapter will generate better tokens; the paper provides no validation-free justification or ablation against alternative routing signals such as held-out perplexity.

Editorial extensions

If this is right

  • The paper's results imply that a diverse fine-tuning corpus can be exploited as a set of specialized experts rather than blended into one adapter, with the benefit concentrated where gradient conflicts are strongest.
  • The same gradient features drive both cluster formation and inference routing, so the pipeline needs no task-specific validation set or external data, a direct corollary of the task-agnostic design.
  • On stronger backbones the benefit narrows (Gemma2-9b gains 0.49 and 0.21 points vs. 1.80 and 0.94 on Gemma-2b at ranks 8 and 64 on MATH-Combined), so ELREA is most valuable for weaker base models or more diverse training data.
  • Compared with training several independent LoRA adapters, ELREA reaches similar or better accuracy with roughly half the training cost, because expert training starts from the shared base checkpoint and cluster sizes sum to the full dataset.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The paper leaves implicit that its routing signal is only weakly tested: the closest ablation is Random Cluster, not a permutation of the routing weights, so a reader should not conclude that cosine similarity to cluster centroids is the decisive component of the gain.
  • A cheap testable extension would replace gradient routing with held-out perplexity per expert or with gradients computed with respect to each expert; if accuracy does not fall, the gain is coming from ensembling rather than from gradient-based routing.
  • Because cluster assignment is a one-time cost on the training set, the same clustering could be reused across multiple target tasks or updated incrementally as new instructions arrive, which the paper does not explore.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The paper proposes ELREA, a fine-tuning framework that combines low-rank adapters with gradient-based clustering and routing. A base LoRA adapter is first trained on the full dataset; per-instance gradient features of the instruction tokens are then computed with an Adam-adjusted formula, randomly projected to 8,192 dimensions, normalized, and clustered with BIRCH. A LoRA expert is trained on each cluster, starting from the base adapter. At inference, the test instruction's gradient is compared with the cluster centroids by cosine similarity, converted to weights via a standardized softmax, and the experts' logits are combined with the base adapter's logits. Experiments on Gemma-2b and Gemma2-9b over MATH-Combined and MMLU/BBH report that ELREA outperforms full-data LoRA and several MoE/ensemble baselines, with roughly 2.3x training time and 2.4x inference memory cost.

Significance. If the reported gains are robust, ELREA would be an attractive drop-in PEFT scheme: it requires no task-specific validation data, uses only instruction-token gradients, and improves accuracy over full-data LoRA on math and general reasoning tasks at modest cost. The paper is clearly written, includes an efficiency analysis, ablates projection dimensionality and clustering algorithm, and includes multiple backbone sizes and ranks. However, the distinctive mechanism—gradient-based clustering and routing—is not yet supported by the evidence: the margins over the Random Cluster control are small and unreplicated, and the paper's own Uniform Weights ablation makes the clustering-only comparison go the wrong way. The central claim is defensible only if the authors add variance estimates and a clean 2x2 ablation separating clustering from routing.

major comments (4)
  1. [Section 5, Tables 1 and 2] The paper's central mechanistic claim—that gradient-based clustering and routing drive the gains—is not supported by the reported ablations. ELREA beats Random Cluster by only 0.11 percentage points on MATH-Combined at r=8 (20.41 vs 20.30), 0.46 points at r=64 (27.33 vs 26.87), and 0.30 points on the BBH/MMLU macro (31.44 vs 31.14), all in single runs without error bars or significance tests. Moreover, the clean comparison of clustering alone goes the wrong way: Uniform Weights (gradient clusters, uniform routing) scores 20.16, 26.48, and 30.83, below Random Cluster (20.30, 26.87, 31.14). Therefore the statement in the Ablation Studies paragraph that "gradient-based clustering method consistently outperforms the random approach" conflates clustering with routing. The authors should report multiple seeds with variance and run a full 2x2 ablation (gradient vs random clustering crossed with gradient-similarity vs uniform routing) to separate the two components.
  2. [Section 4 and Figure 3a] The task-agnostic claim is weakened by the selection of d_proj on the test benchmarks. The text states that "Following Xia et al. (2024), we set the gradient projection dimensionality for clustering d_proj to 8,192, which we show leads to the best model performance," and Figure 3a reports MMLU, BBH, and MATH-Combined accuracy as a function of d_proj. Since these are the same held-out benchmarks used in the main tables, the claim of "no task-specific validation data" does not hold for this hyperparameter. Please either set d_proj with a training-only criterion or present the main results across a range of d_proj values and show that the conclusion is unchanged.
  3. [Section 3.4, Eqs. (8)-(10)] The routing signal—cosine similarity between the test instruction gradient and cluster centroids—is introduced without evidence that it predicts expert quality, and the only direct ablation (Uniform Weights) shows the routing component adds only 0.25-0.85 percentage points in single unreplicated runs. The paper should compare against alternative routing signals (for example, gradients evaluated at the expert adapters, held-out perplexity, or text embeddings) and report whether the routing weights are stable across seeds. Without this, the ensemble may be performing well simply because of the LoRA ensemble effect already captured by the Random Cluster baseline.
  4. [Section 5, Table 1 and Limitations] The headline claim of "performance gains of 9.67% and 3.56% over M + Qbase" rests on single-run differences of 1.80 and 0.94 absolute percentage points. The paper reports no standard errors, no number of seeds, and no significance tests anywhere in the main results. Given that the Gemma2-9b gains shrink to 0.49 and 0.21 points, and the Limitations concede that gains diminish for stronger backbones, the abstract and Section 5 should either provide variance estimates or temper the strength of the claim.
minor comments (5)
  1. [Section 5, paragraph after Table 2] The sentence "This . These findings align with our discussion in § 2.3" appears to be missing words between "This" and the period; please correct the typo.
  2. [Figure 3a] The x-axis labels are malformed ("Based_proj = 512d_proj = 8192") and the two settings are hard to distinguish; please fix the labels and add a legend or clearer tick marks.
  3. [Appendix E, Random Cluster baseline] The Random Cluster description says the adapters are "uniformly weighted during inference" and sets w_rand,base = w_rand,1 = ... = 1, but it does not state whether the logits are averaged or summed. Since the weights in Eq. (10) are not required to sum to 1, please clarify the exact combination rule.
  4. [Appendix H, Figure 4] The cluster-source analysis in Figure 4 is qualitative; reporting a quantitative cluster-purity or correlation metric would make the claim that clusters correspond to data sources more convincing.
  5. [Section 3.2, Eqs. (5)-(7)] The paper correctly notes in the text that ELREA uses instruction-only gradients, unlike Xia et al. (2024), but the equations do not make this distinction explicit. Adding a subscript or sentence to Eqs. (5)-(7) clarifying that g is computed on x_instr only would improve clarity.

Circularity Check

1 steps flagged · score 2.0 of 10

Mild test-set selection of d_proj; no derivation-level circularity or load-bearing self-citation.

  1. fitted input called prediction [Section 4 (Model and Fine-Tuning); also Limitations]
    "Following Xia et al. (2024), we set the gradient projection dimensionality for clustering dproj to 8,192, which we show leads to the best model performance. ... Our hyperparameter tuning for both ELREA and the baseline models is preliminary."

    The paper selects dproj = 8,192 by comparing model performance, and the performance measure used for that selection is the same test-set accuracy later reported in Tables 1–2 as ELREA's headline gains. Thus the reported configuration is a selected maximum over the tested projection dimensionalities on the evaluation data, not an independent out-of-sample prediction. This is a mild selection-on-test-set circularity; it does not make the central clustering/routing derivation tautological.

full rationale

ELREA's derivation chain is self-contained: Eq. (5)–(7) define gradient features from the base adapter, Eq. (8)–(9) produce routing weights from cosine similarity to cluster centroids, and Eq. (10) defines the final prediction as a weighted sum of logits. None of these equations are fit to benchmark labels, and the headline numbers in Tables 1–2 are external MMLU/BBH/MATH-Combined accuracies, not identities derived from the method's inputs. The self-citations (e.g., Li et al. 2024b for deep ensembles) are not load-bearing: they are accompanied by many external citations and do not supply a uniqueness theorem or forbid alternatives. The main experimental concern is not circularity: Random Cluster is defined with uniform weights while ELREA uses gradient routing, so the claim that 'gradient-based clustering consistently outperforms the random approach' conflates clustering with routing; this is a confounded ablation, not a derivation-level circle. The one circular element is mild: Section 4 states dproj = 8192 'leads to the best model performance,' and that performance is measured on the same test sets later used for the reported gains, so the configuration is selected on the evaluation data. The Limitations section also calls the hyperparameter tuning 'preliminary.' Score 2 reflects this selection bias; the central routing/clustering claim retains independent empirical content.

Assumptions & free parameters 5 free parameters · 5 assumptions · 0 invented entities

The method rests on several unproven modeling choices: that gradient directions of instructions are a sufficient and stable signature of task expertise, that random projection preserves these directions after normalization, and that BIRCH clustering on a 5,000-point sample transfers to the full corpus. The paper's hyperparameters, especially the projection dimension, were selected using the same benchmarks that are later reported, which weakens the claim of a task-agnostic setup.

free parameters (5)
  • d_proj = 8192
    Projection dimensionality for gradient features; selected as best on MMLU/BBH/MATH benchmarks (Figure 3a), i.e., tuned to the reported test sets.
  • Initial cluster count and BIRCH iterations = 5 initial clusters, up to 3 splits, final C=8-10
    Hand-chosen; Section 3.3 says initial 5 clusters and up to three iterations, targeting 8 to 10 clusters.
  • Learning rates = 5e-5 base, 2e-5 experts
    Set from prior experience (Appendix D), not from a validation search.
  • Epochs = 2 for base and for each expert
    Chosen based on loss curves and preliminary testing (Appendix D).
  • BIRCH fitting sample size = 5000
    Random subsample used to fit clustering model (Section 3.3); size chosen for computational efficiency.
assumptions (5)
  • standard math Johnson-Lindenstrauss lemma guarantees that random Rademacher projection approximately preserves pairwise distances among gradient features.
    Invoked in Section 3.2 after Eq. (6) to justify reducing |theta_Q| to d_proj.
  • domain assumption The Adam-adapted gradient feature from LESS (Eq. 5) is a faithful influence estimate for instruction-tuning data.
    Borrowed from Xia et al. (2024) without re-justification in this paper.
  • ad hoc to paper Gradient directions of instruction tokens alone, excluding system responses, are sufficient to characterize task expertise and to route test inputs.
    Stated in Section 3.2: 'we only consider the gradient of the instruction tokens...'; this is the key design choice behind both clustering and routing.
  • ad hoc to paper Cluster centroids of training-instruction gradients generalize to test instructions from unseen distributions.
    Assumed in Section 3.4 when weights are computed via cosine similarity to centroids; no out-of-distribution guarantee is provided.
  • domain assumption BIRCH clustering on a random 5,000-point subset is stable and representative of the full dataset.
    Section 3.3 asserts robustness with preliminary experiments but only shows seed stability visualizations in Figure 5 for selected cases.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Ensembles of Low-Rank Expert Adapters." pith.science (2026). https://pith.science/paper/LND3R6LC

@misc{pith2026250200089,
  author       = {Pith},
  title        = {Pith review of: Ensembles of Low-Rank Expert Adapters},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/LND3R6LC}},
  note         = {Machine review of arXiv:2502.00089}
}
read the original abstract

The training and fine-tuning of large language models (LLMs) often involve diverse textual data from multiple sources, which poses challenges due to conflicting gradient directions, hindering optimization and specialization. These challenges can undermine model generalization across tasks, resulting in reduced downstream performance. Recent research suggests that fine-tuning LLMs on carefully selected, task-specific subsets of data can match or even surpass the performance of using the entire dataset. Building on these insights, we propose the Ensembles of Low-Rank Expert Adapters (ELREA) framework to improve the model's capability to handle diverse tasks. ELREA clusters the training instructions based on their gradient directions, representing different areas of expertise and thereby reducing conflicts during optimization. Expert adapters are then trained on these clusters, utilizing the low-rank adaptation (LoRA) technique to ensure training efficiency and model scalability. During inference, ELREA combines predictions from the most relevant expert adapters based on the input data's gradient similarity to the training clusters, ensuring optimal adapter selection for each task. Experiments show that our method outperforms baseline LoRA adapters trained on the full dataset and other ensemble approaches with similar training and inference complexity across a range of domain-specific tasks.

Figures

Figures reproduced from arXiv: 2502.00089 by the authors.

Figure 1
Figure 1. The pipeline of ELREA for fine-tuning and inference. The data points (solid and hollow [PITH_FULL_IMAGE:figures/full_fig_p004_1.png] view at source ↗
Figure 2
Figure 2. Average weight distribution across clusters for different datasets and LoRA ranks. Only [PITH_FULL_IMAGE:figures/full_fig_p009_2.png] view at source ↗
Figure 3
Figure 3. Effects of gradient projection dimensionality and selection of top- [PITH_FULL_IMAGE:figures/full_fig_p010_3.png] view at source ↗
Figures from the paper (2 more)
Figure 4
Figure 4. Figure 4: Distribution of data sources and categories within each cluster for the MATH-Combined [PITH_FULL_IMAGE:figures/full_fig_p028_4.png]
Figure 5
Figure 5. Figure 5: Examples of data clusters from MATH-Combined, generated using different random seeds [PITH_FULL_IMAGE:figures/full_fig_p029_5.png]

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Exploring the Rashomon Set for Concept-Based Models

    cs.LG 2025-11 conditional novelty 6.0 of 10

    A shared frozen backbone plus per-model LoRA adapters and a concept-diversity loss trains a set of accurate CBMs that reason through different concepts.

  2. ScatterPrism: convergence for generative simulation and inverse problems in particle and nuclear physics

    cs.LG 2026-04 unverdicted novelty 4.0 of 10

    Standard Conditional Flow Matching loss is a misleading early plateau; physics-informed metrics keep improving, so ScatterPrism and multi-metric diagnostics are needed for kinematic fidelity.

Reference graph

Works this paper leans on

97 extracted references · 26 canonical work pages · cited by 2 Pith papers

  1. [1]

    write newline

    " write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION format.date year duplicate empty "emp...

  2. [2]

    Ahmed, Rafael Rafailov, Stepan Sharkov, Xuechen Li, and Sanmi Koyejo

    Ahmed M. Ahmed, Rafael Rafailov, Stepan Sharkov, Xuechen Li, and Sanmi Koyejo. Scalable ensembling for mitigating reward overoptimisation. CoRR, abs/2406.01013, 2024. doi:10.48550/ARXIV.2406.01013. URL https://doi.org/10.48550/arXiv.2406.01013

  3. [3]

    Mathqa: Towards interpretable math word problem solving with operation-based formalisms

    Aida Amini, Saadia Gabriel, Shanchuan Lin, Rik Koncel - Kedziorski, Yejin Choi, and Hannaneh Hajishirzi. Mathqa: Towards interpretable math word problem solving with operation-based formalisms. In Jill Burstein, Christy Doran, and Thamar Solorio (eds.), Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational L...

  4. [4]

    Introducing the next generation of claude, 2024

    Anthropic. Introducing the next generation of claude, 2024. URL https://www.anthropic.com/news/claude-3-family

  5. [5]

    Beyond the imitation game: Quantifying and extrapolating the capabilities of language models

    BIG bench authors. Beyond the imitation game: Quantifying and extrapolating the capabilities of language models. Transactions on Machine Learning Research, 2023. ISSN 2835-8856. URL https://openreview.net/forum?id=uyTL5Bvosj

  6. [6]

    A survey on mixture of experts

    Weilin Cai, Juyong Jiang, Fan Wang, Jing Tang, Sunghun Kim, and Jiayi Huang. A survey on mixture of experts. CoRR, abs/2407.06204, 2024. doi:10.48550/ARXIV.2407.06204. URL https://doi.org/10.48550/arXiv.2407.06204

  7. [7]

    SWAD: domain generalization by seeking flat minima

    Junbum Cha, Sanghyuk Chun, Kyungjae Lee, Han - Cheol Cho, Seunghyun Park, Yunsung Lee, and Sungrae Park. SWAD: domain generalization by seeking flat minima. In Marc'Aurelio Ranzato, Alina Beygelzimer, Yann N. Dauphin, Percy Liang, and Jennifer Wortman Vaughan (eds.), Advances in Neural Information Processing Systems 34: Annual Conference on Neural Informa...

  8. [8]

    Frugalgpt: How to use large language models while reducing cost and improving performance

    Lingjiao Chen, Matei Zaharia, and James Zou. Frugalgpt: How to use large language models while reducing cost and improving performance. CoRR, abs/2305.05176, 2023. doi:10.48550/ARXIV.2305.05176. URL https://doi.org/10.48550/arXiv.2305.05176

Show all 97 references
  1. [9]

    Llava-mole: Sparse mixture of lora experts for mitigating data conflicts in instruction finetuning mllms

    Shaoxiang Chen, Zequn Jie, and Lin Ma. Llava-mole: Sparse mixture of lora experts for mitigating data conflicts in instruction finetuning mllms. CoRR, abs/2401.16160, 2024. doi:10.48550/ARXIV.2401.16160. URL https://doi.org/10.48550/arXiv.2401.16160

  2. [10]

    Peters, Alexander Fraser, and Jesse Dodge

    Alexandra Chronopoulou, Matthew E. Peters, Alexander Fraser, and Jesse Dodge. Adaptersoup: Weight averaging to improve generalization of pretrained language models. In Andreas Vlachos and Isabelle Augenstein (eds.), Findings of the Association for Computational Linguistics: EA...

  3. [11]

    Training verifiers to solve math word problems

    Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, Christopher Hesse, and John Schulman. Training verifiers to solve math word problems. CoRR, abs/2110.14168, 2021. URL https://...

  4. [12]

    Free dolly: Introducing the world's first truly open instruction-tuned llm, 2023

    Mike Conover, Matt Hayes, Ankit Mathur, Jianwei Xie, Jun Wan, Sam Shah, Ali Ghodsi, Patrick Wendell, Matei Zaharia, and Reynold Xin. Free dolly: Introducing the world's first truly open instruction-tuned llm, 2023. URL https://www.databricks.com/blog/2023/04/12/dolly-first-ope...

  5. [13]

    Reward model ensembles help mitigate overoptimization

    Thomas Coste, Usman Anwar, Robert Kirk, and David Krueger. Reward model ensembles help mitigate overoptimization. In The Twelfth International Conference on Learning Representations, ICLR 2024, Vienna, Austria, May 7-11, 2024 . OpenReview.net, 2024. URL https://openreview.net/...

  6. [14]

    Damai Dai, Chengqi Deng, Chenggang Zhao, R. X. Xu, Huazuo Gao, Deli Chen, Jiashi Li, Wangding Zeng, Xingkai Yu, Y. Wu, Zhenda Xie, Y. K. Li, Panpan Huang, Fuli Luo, Chong Ruan, Zhifang Sui, and Wenfeng Liang. Deepseekmoe: Towards ultimate expert specialization in mixture-of-ex...

  7. [15]

    Qlora: Efficient finetuning of quantized llms

    Tim Dettmers, Artidoro Pagnoni, Ari Holtzman, and Luke Zettlemoyer. Qlora: Efficient finetuning of quantized llms. In Alice Oh, Tristan Naumann, Amir Globerson, Kate Saenko, Moritz Hardt, and Sergey Levine (eds.), Advances in Neural Information Processing Systems 36: Annual Co...

  8. [16]

    Mixture-of-domain-adapters: Decoupling and injecting domain knowledge to pre-trained language models' memories

    Shizhe Diao, Tianyang Xu, Ruijia Xu, Jiawei Wang, and Tong Zhang. Mixture-of-domain-adapters: Decoupling and injecting domain knowledge to pre-trained language models' memories. In Anna Rogers, Jordan L. Boyd - Graber, and Naoaki Okazaki (eds.), Proceedings of the 61st Annual ...

  9. [17]

    Loramoe: Revolutionizing mixture of experts for maintaining world knowledge in language model alignment

    Shihan Dou, Enyu Zhou, Yan Liu, Songyang Gao, Jun Zhao, Wei Shen, Yuhao Zhou, Zhiheng Xi, Xiao Wang, Xiaoran Fan, Shiliang Pu, Jiang Zhu, Rui Zheng, Tao Gui, Qi Zhang, and Xuanjing Huang. Loramoe: Revolutionizing mixture of experts for maintaining world knowledge in language m...

  10. [18]

    Revisiting deep ensemble for out-of-distribution detection: A loss landscape perspective

    Kun Fang, Qinghua Tao, Xiaolin Huang, and Jie Yang. Revisiting deep ensemble for out-of-distribution detection: A loss landscape perspective. CoRR, abs/2310.14227, 2023. doi:10.48550/ARXIV.2310.14227. URL https://doi.org/10.48550/arXiv.2310.14227

  11. [19]

    Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity

    William Fedus, Barret Zoph, and Noam Shazeer. Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity. J. Mach. Learn. Res., 23: 0 120:1--120:39, 2022. URL https://jmlr.org/papers/v23/21-0998.html

  12. [20]

    Deep ensembles: A loss landscape perspective

    Stanislav Fort, Huiyi Hu, and Balaji Lakshminarayanan. Deep ensembles: A loss landscape perspective. CoRR, abs/1912.02757, 2019. URL http://arxiv.org/abs/1912.02757

  13. [21]

    Chongyang Gao, Kezhen Chen, Jinmeng Rao, Baochen Sun, Ruibo Liu, Daiyi Peng, Yawen Zhang, Xiaoyuan Guo, Jie Yang, and V. S. Subrahmanian. Higher layers need more lora experts. CoRR, abs/2402.08562, 2024. doi:10.48550/ARXIV.2402.08562. URL https://doi.org/10.48550/arXiv.2402.08562

  14. [22]

    The pile: An 800gb dataset of diverse text for language modeling

    Leo Gao, Stella Biderman, Sid Black, Laurence Golding, Travis Hoppe, Charles Foster, Jason Phang, Horace He, Anish Thite, Noa Nabeshima, Shawn Presser, and Connor Leahy. The pile: An 800gb dataset of diverse text for language modeling. CoRR, abs/2101.00027, 2021. URL https://a...

  15. [23]

    Vetrov, and Andrew Gordon Wilson

    Timur Garipov, Pavel Izmailov, Dmitrii Podoprikhin, Dmitry P. Vetrov, and Andrew Gordon Wilson. Loss surfaces, mode connectivity, and fast ensembling of dnns. In Samy Bengio, Hanna M. Wallach, Hugo Larochelle, Kristen Grauman, Nicol \` o Cesa - Bianchi, and Roman Garnett (eds....

  16. [24]

    Gemma 2: Improving open language models at a practical size

    Gemma Team . Gemma 2: Improving open language models at a practical size. CoRR, abs/2408.00118, 2024 a . doi:10.48550/ARXIV.2403.08295. URL https://doi.org/10.48550/arXiv.2408.00118

  17. [25]

    Gemma: Open models based on gemini research and technology

    Gemma Team . Gemma: Open models based on gemini research and technology. CoRR, abs/2403.08295, 2024 b . doi:10.48550/ARXIV.2403.08295. URL https://doi.org/10.48550/arXiv.2403.08295

  18. [26]

    Uncertainty estimation for language reward models

    Adam Gleave and Geoffrey Irving. Uncertainty estimation for language reward models. CoRR, abs/2203.07472, 2022. doi:10.48550/ARXIV.2203.07472. URL https://doi.org/10.48550/arXiv.2203.07472

  19. [27]

    Kwok, and Yu Zhang

    Yunhao Gou, Zhili Liu, Kai Chen, Lanqing Hong, Hang Xu, Aoxue Li, Dit - Yan Yeung, James T. Kwok, and Yu Zhang. Mixture of cluster-conditional lora experts for vision-language instruction tuning. CoRR, abs/2312.12379, 2023. doi:10.48550/ARXIV.2312.12379. URL https://doi.org/10...

  20. [28]

    Training independent subnetworks for robust prediction

    Marton Havasi, Rodolphe Jenatton, Stanislav Fort, Jeremiah Zhe Liu, Jasper Snoek, Balaji Lakshminarayanan, Andrew Mingbo Dai, and Dustin Tran. Training independent subnetworks for robust prediction. In 9th International Conference on Learning Representations, ICLR 2021, Virtua...

  21. [29]

    Towards a unified view of parameter-efficient transfer learning

    Junxian He, Chunting Zhou, Xuezhe Ma, Taylor Berg - Kirkpatrick, and Graham Neubig. Towards a unified view of parameter-efficient transfer learning. In The Tenth International Conference on Learning Representations, ICLR 2022, Virtual Event, April 25-29, 2022 . OpenReview.net,...

  22. [30]

    Measuring massive multitask language understanding

    Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. Measuring massive multitask language understanding. In 9th International Conference on Learning Representations, ICLR 2021, Virtual Event, Austria, May 3-7, 2021 . OpenReview...

  23. [31]

    Measuring mathematical problem solving with the MATH dataset

    Dan Hendrycks, Collin Burns, Saurav Kadavath, Akul Arora, Steven Basart, Eric Tang, Dawn Song, and Jacob Steinhardt. Measuring mathematical problem solving with the MATH dataset. In Joaquin Vanschoren and Sai - Kit Yeung (eds.), Proceedings of the Neural Information Processing...

  24. [32]

    Parameter-efficient transfer learning for NLP

    Neil Houlsby, Andrei Giurgiu, Stanislaw Jastrzebski, Bruna Morrone, Quentin de Laroussilhe, Andrea Gesmundo, Mona Attariyan, and Sylvain Gelly. Parameter-efficient transfer learning for NLP . In Kamalika Chaudhuri and Ruslan Salakhutdinov (eds.), Proceedings of the 36th Intern...

  25. [33]

    Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen - Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen

    Edward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen - Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. Lora: Low-rank adaptation of large language models. In The Tenth International Conference on Learning Representations, ICLR 2022, Virtual Event, April 25-29, 2022 ....

  26. [34]

    Lorahub: Efficient cross-task generalization via dynamic lo RA composition

    Chengsong Huang, Qian Liu, Bill Yuchen Lin, Tianyu Pang, Chao Du, and Min Lin. Lorahub: Efficient cross-task generalization via dynamic lo RA composition. In First Conference on Language Modeling, 2024. URL https://openreview.net/forum?id=TrloAXEJ2B

  27. [35]

    Albert Q. Jiang, Alexandre Sablayrolles, Antoine Roux, Arthur Mensch, Blanche Savary, Chris Bamford, Devendra Singh Chaplot, Diego de Las Casas, Emma Bou Hanna, Florian Bressand, Gianna Lengyel, Guillaume Bour, Guillaume Lample, L \' e lio Renard Lavaud, Lucile Saulnier, Marie...

  28. [36]

    Llm-blender: Ensembling large language models with pairwise ranking and generative fusion

    Dongfu Jiang, Xiang Ren, and Bill Yuchen Lin. Llm-blender: Ensembling large language models with pairwise ranking and generative fusion. In Anna Rogers, Jordan L. Boyd - Graber, and Naoaki Okazaki (eds.), Proceedings of the 61st Annual Meeting of the Association for Computatio...

  29. [37]

    Johnson and Joram Lindenstrauss

    William B. Johnson and Joram Lindenstrauss. Extensions of lipschitz mappings into hilbert space. Contemporary mathematics, 26: 0 189--206, 1984. URL https://api.semanticscholar.org/CorpusID:117819162

  30. [38]

    Jordan and Robert A

    Michael I. Jordan and Robert A. Jacobs. Hierarchical mixtures of experts and the EM algorithm. Neural Comput., 6 0 (2): 0 181--214, 1994. doi:10.1162/NECO.1994.6.2.181. URL https://doi.org/10.1162/neco.1994.6.2.181

  31. [39]

    Random indexing of text samples for latent semantic analysis

    Pentii Kanerva, Jan Kristoferson, and Anders Holst. Random indexing of text samples for latent semantic analysis. In Proceedings of the Annual Meeting of the Cognitive Science Society, volume 22, 2000

  32. [40]

    Kingma and Jimmy Ba

    Diederik P. Kingma and Jimmy Ba. Adam: A method for stochastic optimization. In Yoshua Bengio and Yann LeCun (eds.), 3rd International Conference on Learning Representations, ICLR 2015, San Diego, CA, USA, May 7-9, 2015, Conference Track Proceedings , 2015. URL http://arxiv.or...

  33. [41]

    o pf, Yannic Kilcher, Dimitri von R \

    Andreas K \" o pf, Yannic Kilcher, Dimitri von R \" u tte, Sotiris Anagnostidis, Zhi Rui Tam, Keith Stevens, Abdullah Barhoum, Duc Nguyen, Oliver Stanley, Rich \' a rd Nagyfi, Shahul ES, Sameer Suri, David Glushkov, Arnav Dantuluri, Andrew Maguire, Christoph Schuhmann, Huu Ngu...

  34. [42]

    Simple and scalable predictive uncertainty estimation using deep ensembles

    Balaji Lakshminarayanan, Alexander Pritzel, and Charles Blundell. Simple and scalable predictive uncertainty estimation using deep ensembles. In Isabelle Guyon, Ulrike von Luxburg, Samy Bengio, Hanna M. Wallach, Rob Fergus, S. V. N. Vishwanathan, and Roman Garnett (eds.), Adva...

  35. [43]

    Mixlora: Enhancing large language models fine-tuning with lora based mixture of experts

    Dengchun Li, Yingzi Ma, Naizheng Wang, Zhiyuan Cheng, Lei Duan, Jie Zuo, Cal Yang, and Mingjie Tang. Mixlora: Enhancing large language models fine-tuning with lora based mixture of experts. CoRR, abs/2404.15159, 2024 a . doi:10.48550/ARXIV.2404.15159. URL https://doi.org/10.48...

  36. [44]

    Prefix-tuning: Optimizing continuous prompts for generation

    Xiang Lisa Li and Percy Liang. Prefix-tuning: Optimizing continuous prompts for generation. In Chengqing Zong, Fei Xia, Wenjie Li, and Roberto Navigli (eds.), Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joi...

  37. [45]

    MUB en: Benchmarking the uncertainty of molecular representation models

    Yinghao Li, Lingkai Kong, Yuanqi Du, Yue Yu, Yuchen Zhuang, Wenhao Mu, and Chao Zhang. MUB en: Benchmarking the uncertainty of molecular representation models. Transactions on Machine Learning Research, 2024 b . ISSN 2835-8856. URL https://openreview.net/forum?id=qYceFeHgm4

  38. [46]

    When MOE meets llms: Parameter efficient fine-tuning for multi-task medical applications

    Qidong Liu, Xian Wu, Xiangyu Zhao, Yuanshao Zhu, Derong Xu, Feng Tian, and Yefeng Zheng. When MOE meets llms: Parameter efficient fine-tuning for multi-task medical applications. In Grace Hui Yang, Hongning Wang, Sam Han, Claudia Hauff, Guido Zuccon, and Yi Zhang (eds.), Proce...

  39. [47]

    Deep ensembling with no overhead for either training or testing: The all-round blessings of dynamic sparsity

    Shiwei Liu, Tianlong Chen, Zahra Atashgahi, Xiaohan Chen, Ghada Sokar, Elena Mocanu, Mykola Pechenizkiy, Zhangyang Wang, and Decebal Constantin Mocanu. Deep ensembling with no overhead for either training or testing: The all-round blessings of dynamic sparsity. In The Tenth In...

  40. [48]

    Intuition-aware mixture-of-rank-1-experts for parameter efficient finetuning

    Yijiang Liu, Rongyu Zhang, Huanrui Yang, Kurt Keutzer, Yuan Du, Li Du, and Shanghang Zhang. Intuition-aware mixture-of-rank-1-experts for parameter efficient finetuning. CoRR, abs/2404.08985, 2024 b . doi:10.48550/ARXIV.2404.08985. URL https://doi.org/10.48550/arXiv.2404.08985

  41. [49]

    Take the essence and discard the dross: A rethinking on data selection for fine-tuning large language models

    Ziche Liu, Rui Ke, Feng Jiang, and Haizhou Li. Take the essence and discard the dross: A rethinking on data selection for fine-tuning large language models. CoRR, abs/2406.14115, 2024 c . doi:10.48550/ARXIV.2406.14115. URL https://doi.org/10.48550/arXiv.2406.14115

  42. [50]

    Le, Barret Zoph, Jason Wei, and Adam Roberts

    Shayne Longpre, Le Hou, Tu Vu, Albert Webson, Hyung Won Chung, Yi Tay, Denny Zhou, Quoc V. Le, Barret Zoph, Jason Wei, and Adam Roberts. The flan collection: Designing data and methods for effective instruction tuning. In Andreas Krause, Emma Brunskill, Kyunghyun Cho, Barbara ...

  43. [51]

    Merge, ensemble, and cooperate! A survey on collaborative strategies in the era of large language models

    Jinliang Lu, Ziliang Pang, Min Xiao, Yaochen Zhu, Rui Xia, and Jiajun Zhang. Merge, ensemble, and cooperate! A survey on collaborative strategies in the era of large language models. CoRR, abs/2407.06089, 2024. doi:10.48550/ARXIV.2407.06089. URL https://doi.org/10.48550/arXiv....

  44. [52]

    Moelora: Contrastive learning guided mixture of experts on parameter-efficient fine-tuning for large language models

    Tongxu Luo, Jiahe Lei, Fangyu Lei, Weihao Liu, Shizhu He, Jun Zhao, and Kang Liu. Moelora: Contrastive learning guided mixture of experts on parameter-efficient fine-tuning for large language models. CoRR, abs/2402.12851, 2024. doi:10.48550/ARXIV.2402.12851. URL https://doi.or...

  45. [53]

    Learning to route among specialized experts for zero-shot generalization

    Mohammed Muqeeth, Haokun Liu, Yufan Liu, and Colin Raffel. Learning to route among specialized experts for zero-shot generalization. In Forty-first International Conference on Machine Learning, ICML 2024, Vienna, Austria, July 21-27, 2024 . OpenReview.net, 2024. URL https://op...

  46. [54]

    Introducing ChatGPT , 2022

    OpenAI. Introducing ChatGPT , 2022. URL https://openai.com/blog/chatgpt. (Accessed on Jun 18, 2023)

  47. [55]

    GPT-4 technical report

    OpenAI. GPT-4 technical report. CoRR, abs/2303.08774, 2023. doi:10.48550/ARXIV.2303.08774. URL https://doi.org/10.48550/arXiv.2303.08774

  48. [56]

    Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul F. Christiano, Jan Leik...

  49. [57]

    G-DIG: towards gradient-based diverse and high-quality instruction data selection for machine translation

    Xingyuan Pan, Luyang Huang, Liyan Kang, Zhicheng Liu, Yu Lu, and Shanbo Cheng. G-DIG: towards gradient-based diverse and high-quality instruction data selection for machine translation. CoRR, abs/2405.12915, 2024. doi:10.48550/ARXIV.2405.12915. URL https://doi.org/10.48550/arX...

  50. [58]

    Arkil Patel, Satwik Bhattamishra, and Navin Goyal. Are NLP models really able to solve simple math word problems? In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pp.\ 2080--2094,...

  51. [59]

    Zhang, Andrew Wang, and Jimmy Ba

    Silviu Pitis, Michael R. Zhang, Andrew Wang, and Jimmy Ba. Boosted prompt ensembles for large language models. CoRR, abs/2304.05970, 2023. doi:10.48550/ARXIV.2304.05970. URL https://doi.org/10.48550/arXiv.2304.05970

  52. [60]

    Estimating training data influence by tracing gradient descent

    Garima Pruthi, Frederick Liu, Satyen Kale, and Mukund Sundararajan. Estimating training data influence by tracing gradient descent. In Hugo Larochelle, Marc'Aurelio Ranzato, Raia Hadsell, Maria - Florina Balcan, and Hsuan - Tien Lin (eds.), Advances in Neural Information Proce...

  53. [61]

    Improving language understanding by generative pre-training

    Alec Radford, Karthik Narasimhan, Tim Salimans, Ilya Sutskever, et al. Improving language understanding by generative pre-training. 2018. URL https://cdn.openai.com/research-covers/language-unsupervised/language_understanding_paper.pdf

  54. [62]

    Sentence-bert: Sentence embeddings using siamese bert-networks

    Nils Reimers and Iryna Gurevych. Sentence-bert: Sentence embeddings using siamese bert-networks. In Kentaro Inui, Jing Jiang, Vincent Ng, and Xiaojun Wan (eds.), Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Jo...

  55. [63]

    Mome: Mixture of multimodal experts for generalist multimodal large language models

    Leyang Shen, Gongwei Chen, Rui Shao, Weili Guan, and Liqiang Nie. Mome: Mixture of multimodal experts for generalist multimodal large language models. CoRR, abs/2407.12709, 2024. doi:10.48550/ARXIV.2407.12709. URL https://doi.org/10.48550/arXiv.2407.12709

  56. [64]

    Zhao, Hongkun Yu, Kurt Keutzer, Trevor Darrell, and Denny Zhou

    Sheng Shen, Le Hou, Yanqi Zhou, Nan Du, Shayne Longpre, Jason Wei, Hyung Won Chung, Barret Zoph, William Fedus, Xinyun Chen, Tu Vu, Yuexin Wu, Wuyang Chen, Albert Webson, Yunxuan Li, Vincent Y. Zhao, Hongkun Yu, Kurt Keutzer, Trevor Darrell, and Denny Zhou. Flan-moe: Scaling i...

  57. [65]

    Sara Mahdavi, Joelle K

    Karan Singhal, Tao Tu, Juraj Gottweis, Rory Sayres, Ellery Wulczyn, Le Hou, Kevin Clark, Stephen Pfohl, Heather Cole - Lewis, Darlene Neal, Mike Schaekermann, Amy Wang, Mohamed Amin, Sami Lachgar, Philip Andrew Mansfield, Sushant Prakash, Bradley Green, Ewa Dominowska, Blaise ...

  58. [66]

    Le, Ed H

    Mirac Suzgun, Nathan Scales, Nathanael Sch \" a rli, Sebastian Gehrmann, Yi Tay, Hyung Won Chung, Aakanksha Chowdhery, Quoc V. Le, Ed H. Chi, Denny Zhou, and Jason Wei. Challenging big-bench tasks and whether chain-of-thought can solve them. In Anna Rogers, Jordan L. Boyd - Gr...

  59. [67]

    Szymanski and Michael D

    Peter T. Szymanski and Michael D. Lemmon. Adaptive mixtures of local experts are source coding solutions. In Proceedings of International Conference on Neural Networks (ICNN'88), San Francisco, CA, USA, March 28 - April 1, 1993, pp.\ 1391--1396. IEEE , 1993. doi:10.1109/ICNN.1...

  60. [68]

    Hashimoto

    Rohan Taori, Ishaan Gulrajani, Tianyi Zhang, Yann Dubois, Xuechen Li, Carlos Guestrin, Percy Liang, and Tatsunori B. Hashimoto. Stanford alpaca: An instruction-following llama model. https://github.com/tatsu-lab/stanford_alpaca, 2023

  61. [69]

    Hydralo RA : An asymmetric lo RA architecture for efficient fine-tuning

    Chunlin Tian, Zhan Shi, Zhijiang Guo, Li Li, and Cheng zhong Xu. Hydralo RA : An asymmetric lo RA architecture for efficient fine-tuning. In The Thirty-eighth Annual Conference on Neural Information Processing Systems, 2024. URL https://openreview.net/forum?id=qEpi8uWX3N

  62. [70]

    Llama: Open and efficient foundation language models

    Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie - Anne Lachaux, Timoth \' e e Lacroix, Baptiste Rozi \` e re, Naman Goyal, Eric Hambro, Faisal Azhar, Aur \' e lien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lample. Llama: Open and efficient fo...

  63. [71]

    Liu, Michael W

    Dustin Tran, Jeremiah Z. Liu, Michael W. Dusenberry, Du Phan, Mark Collier, Jie Ren, Kehang Han, Zi Wang, Zelda Mariet, Huiyi Hu, Neil Band, Tim G. J. Rudner, Karan Singhal, Zachary Nado, Joost van Amersfoort, Andreas Kirsch, Rodolphe Jenatton, Nithum Thain, Honglin Yuan, Kell...

  64. [72]

    Gomez, Lukasz Kaiser, and Illia Polosukhin

    Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. Attention is all you need. In Isabelle Guyon, Ulrike von Luxburg, Samy Bengio, Hanna M. Wallach, Rob Fergus, S. V. N. Vishwanathan, and Roman Garnett (...

  65. [73]

    Xing, and Mikhail Yurochkin

    Hongyi Wang, Felipe Maia Polo, Yuekai Sun, Souvik Kundu, Eric P. Xing, and Mikhail Yurochkin. Fusing models with complementary expertise. In The Twelfth International Conference on Learning Representations, ICLR 2024, Vienna, Austria, May 7-11, 2024 . OpenReview.net, 2024. URL...

  66. [74]

    Lora ensembles for large language model fine-tuning

    Xi Wang, Laurence Aitchison, and Maja Rudolph. Lora ensembles for large language model fine-tuning. CoRR, abs/2310.00035, 2023 a . doi:10.48550/ARXIV.2310.00035. URL https://doi.org/10.48550/arXiv.2310.00035

  67. [75]

    Le, Ed H

    Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V. Le, Ed H. Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou. Self-consistency improves chain of thought reasoning in language models. In The Eleventh International Conference on Learning Representations, ICLR 2023, Kigali,...

  68. [76]

    Multilora: Democratizing lora for better multi-task learning

    Yiming Wang, Yu Lin, Xiaodong Zeng, and Guannan Zhang. Multilora: Democratizing lora for better multi-task learning. CoRR, abs/2311.11501, 2023 c . doi:10.48550/ARXIV.2311.11501. URL https://doi.org/10.48550/arXiv.2311.11501

  69. [77]

    Smith, Iz Beltagy, and Hannaneh Hajishirzi

    Yizhong Wang, Hamish Ivison, Pradeep Dasigi, Jack Hessel, Tushar Khot, Khyathi Raghavi Chandu, David Wadden, Kelsey MacMillan, Noah A. Smith, Iz Beltagy, and Hannaneh Hajishirzi. How far can camels go? exploring the state of instruction tuning on open resources. In Alice Oh, T...

  70. [78]

    Gradient vaccine: Investigating and improving multi-task optimization in massively multilingual models

    Zirui Wang, Yulia Tsvetkov, Orhan Firat, and Yuan Cao. Gradient vaccine: Investigating and improving multi-task optimization in massively multilingual models. In 9th International Conference on Learning Representations, ICLR 2021, Virtual Event, Austria, May 3-7, 2021 . OpenRe...

  71. [79]

    Chi, Quoc V

    Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed H. Chi, Quoc V. Le, and Denny Zhou. Chain-of-thought prompting elicits reasoning in large language models. In Sanmi Koyejo, S. Mohamed, A. Agarwal, Danielle Belgrave, K. Cho, and A. Oh (eds.), Ad...

  72. [80]

    Batched low-rank adaptation of foundation models

    Yeming Wen and Swarat Chaudhuri. Batched low-rank adaptation of foundation models. In The Twelfth International Conference on Learning Representations, ICLR 2024, Vienna, Austria, May 7-11, 2024 . OpenReview.net, 2024. URL https://openreview.net/forum?id=w4abltTZ2f

  73. [81]

    Mixture of lora experts

    Xun Wu, Shaohan Huang, and Furu Wei. Mixture of lora experts. CoRR, abs/2404.13628, 2024. doi:10.48550/ARXIV.2404.13628. URL https://doi.org/10.48550/arXiv.2404.13628

  74. [82]

    LESS: selecting influential data for targeted instruction tuning

    Mengzhou Xia, Sadhika Malladi, Suchin Gururangan, Sanjeev Arora, and Danqi Chen. LESS: selecting influential data for targeted instruction tuning. CoRR, abs/2402.04333, 2024. doi:10.48550/ARXIV.2402.04333. URL https://doi.org/10.48550/arXiv.2402.04333

  75. [83]

    Data selection for language models via importance resampling

    Sang Michael Xie, Shibani Santurkar, Tengyu Ma, and Percy Liang. Data selection for language models via importance resampling. In Alice Oh, Tristan Naumann, Amir Globerson, Kate Saenko, Moritz Hardt, and Sergey Levine (eds.), Advances in Neural Information Processing Systems 3...

  76. [84]

    Openmoe: An early effort on open mixture-of-experts language models

    Fuzhao Xue, Zian Zheng, Yao Fu, Jinjie Ni, Zangwei Zheng, Wangchunshu Zhou, and Yang You. Openmoe: An early effort on open mixture-of-experts language models. CoRR, abs/2402.01739, 2024. doi:10.48550/ARXIV.2402.01739. URL https://doi.org/10.48550/arXiv.2402.01739

  77. [85]

    Fingpt: Open-source financial large language models

    Hongyang Yang, Xiao - Yang Liu, and Christina Dan Wang. Fingpt: Open-source financial large language models. CoRR, abs/2306.06031, 2023. doi:10.48550/ARXIV.2306.06031. URL https://doi.org/10.48550/arXiv.2306.06031

  78. [86]

    Solving token gradient conflict in mixture-of-experts for large vision-language model

    Longrong Yang, Dong Sheng, Chaoxiang Cai, Fan Yang, Size Li, Di Zhang, and Xi Li. Solving token gradient conflict in mixture-of-experts for large vision-language model. CoRR, abs/2406.19905, 2024. doi:10.48550/ARXIV.2406.19905. URL https://doi.org/10.48550/arXiv.2406.19905

  79. [87]

    Pushing mixture of experts to the limit: Extremely parameter efficient moe for instruction tuning

    Ted Zadouri, Ahmet \" U st \" u n, Arash Ahmadian, Beyza Ermis, Acyr Locatelli, and Sara Hooker. Pushing mixture of experts to the limit: Extremely parameter efficient moe for instruction tuning. In The Twelfth International Conference on Learning Representations, ICLR 2024, V...

  80. [88]

    Improving reinforcement learning from human feedback with efficient reward model ensemble

    Shun Zhang, Zhenfang Chen, Sunli Chen, Yikang Shen, Zhiqing Sun, and Chuang Gan. Improving reinforcement learning from human feedback with efficient reward model ensemble. CoRR, abs/2401.16635, 2024 a . doi:10.48550/ARXIV.2401.16635. URL https://doi.org/10.48550/arXiv.2401.16635

  81. [89]

    Birch: an efficient data clustering method for very large databases

    Tian Zhang, Raghu Ramakrishnan, and Miron Livny. Birch: an efficient data clustering method for very large databases. In Proceedings of the 1996 ACM SIGMOD International Conference on Management of Data, SIGMOD '96, pp.\ 103--114, New York, NY, USA, 1996. Association for Compu...

  82. [90]

    A comprehensive survey of scientific large language models and their applications in scientific discovery

    Yu Zhang, Xiusi Chen, Bowen Jin, Sheng Wang, Shuiwang Ji, Wei Wang, and Jiawei Han. A comprehensive survey of scientific large language models and their applications in scientific discovery. CoRR, abs/2406.10833, 2024 b . doi:10.48550/ARXIV.2406.10833. URL https://doi.org/10.4...

  83. [91]

    Lory: Fully differentiable mixture-of-experts for autoregressive language model pre-training

    Zexuan Zhong, Mengzhou Xia, Danqi Chen, and Mike Lewis. Lory: Fully differentiable mixture-of-experts for autoregressive language model pre-training. In First Conference on Language Modeling, 2024. URL https://openreview.net/forum?id=LKEJPySnlt

  84. [92]

    Exploring training on heterogeneous data with mixture of low-rank adapters

    Yuhang Zhou, Zihua Zhao, Siyuan Du, Haolin Li, Jiangchao Yao, Ya Zhang, and Yanfeng Wang. Exploring training on heterogeneous data with mixture of low-rank adapters. In Forty-first International Conference on Machine Learning, ICML 2024, Vienna, Austria, July 21-27, 2024 . Ope...

  85. [93]

    Llama-moe: Building mixture-of-experts from llama with continual pre-training

    Tong Zhu, Xiaoye Qu, Daize Dong, Jiacheng Ruan, Jingqi Tong, Conghui He, and Yu Cheng. Llama-moe: Building mixture-of-experts from llama with continual pre-training. CoRR, abs/2406.16554, 2024. doi:10.48550/ARXIV.2406.16554. URL https://doi.org/10.48550/arXiv.2406.16554

  86. [94]

    Sira: Sparse mixture of low rank adaptation

    Yun Zhu, Nevan Wichers, Chu - Cheng Lin, Xinyi Wang, Tianlong Chen, Lei Shu, Han Lu, Canoee Liu, Liangchen Luo, Jindong Chen, and Lei Meng. Sira: Sparse mixture of low rank adaptation. CoRR, abs/2311.09179, 2023. doi:10.48550/ARXIV.2311.09179. URL https://doi.org/10.48550/arXi...

  87. [95]

    @esa (Ref

    \@ifxundefined[1] #1\@undefined \@firstoftwo \@secondoftwo \@ifnum[1] #1 \@firstoftwo \@secondoftwo \@ifx[1] #1 \@firstoftwo \@secondoftwo [2] @ #1 \@temptokena #2 #1 @ \@temptokena \@ifclassloaded agu2001 natbib The agu2001 class already includes natbib coding, so you should ...

  88. [96]

    \@lbibitem[] @bibitem@first@sw\@secondoftwo \@lbibitem[#1]#2 \@extra@b@citeb \@ifundefined br@#2\@extra@b@citeb \@namedef br@#2 \@nameuse br@#2\@extra@b@citeb \@ifundefined b@#2\@extra@b@citeb @num @parse #2 @tmp #1 NAT@b@open@#2 NAT@b@shut@#2 \@ifnum @merge>\@ne @bibitem@firs...

  89. [97]

    The data points (solid and hollow circles) do not necessarily have a geometric correspondence to their gradient directions (arrows)

    @open @close @open @close and [1] URL: #1 \@ifundefined chapter * \@mkboth \@ifxundefined @sectionbib * \@mkboth * \@mkboth\@gobbletwo \@ifclassloaded amsart * \@ifclassloaded amsbook * \@ifxundefined @heading @heading NAT@ctr thebibliography [1] @ \@biblabel @NAT@ctr \@bibset...

Pith tools

Reviewed August 9, 2026 · model on record in the stance chip above.