REVIEW 3 major objections 4 minor 6 cited by
Doing More with Less: A Survey on Routing Strategies for Resource Optimisation in Large Language Model-Based Systems
T0 review · 3 major / 4 minor · reviewed 2026-08-09 · deepseek-v4-flash
Pith's one-line read This survey argues that LLM routing is a performance–cost optimisation problem, and that all existing strategies can be organised by timing (pre- vs post-generation) and by implementation family (similarity, supervised, reinforcement…
desk verdict A useful, well-organized survey of LLM routing that will help practitioners navigate the field; cross-paper performance comparisons are the main soft spot, but the paper itself flags the comparability problem. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the router function of Equation (1), which selects the candidate that maximises a scoring function $s(q, M)$ under a cost or budget constraint $C_M(q) \le B$. Around it, the survey builds a two-axis taxonomy: pre-generation routing (predict candidate performance before any output is generated, based on domain or complexity) versus post-generation routing (cascade routing, where each generated answer is evaluated and the query escalates to a larger model if the answer seems inadequate); and four implementation families, namely similarity-based routing (nearest neighbours, clustering, preference similarity), supervised routing (recommendation, domain classification, query complexity inference, answer confidence inference, knowledge graphs), reinforcement-learning-based routing (stateless, state-based, reward inference), and generative routing (prompts, sequence or token probabilities, repeated calls, LLM fine-tuning, code execution). This machinery organises the literature and lets the authors compare strategies on cost, generalisation, and resource requirements.
What would settle it
Run representatives of each routing family (similarity, supervised, reinforcement learning, and generative) on one shared benchmark with a fixed model pool, a single cost model, and random routing plus the best stand-alone model as baselines; if no lightweight strategy reliably beats those baselines on that common setup, the survey's comparative conclusions would not hold.
Extended reading notes
Core claim
The central claim is that routing should be treated as a performance–cost optimisation problem, written as $R_M(q) = \arg\max_{M \in \mathcal{M}} s(q, M)$ subject to $C_M(q) \le B$, and that the literature can be understood through two axes: timing (before or after generation) and implementation family. The survey finds that routing does not require expensive machinery: similarity and supervised methods can match or approach large-model quality while cutting calls to the expensive model substantially, and even the most resource-intensive family (generative routing via LLM fine-tuning) can beat much larger generalists on specific tasks. It also reports that post-generation cascade routing is generally more resource-intensive than pre-generation routing, that routers generalise poorly to new model candidates unless they project options into a shared semantic space, and that the field's evaluations are not standardised, making direct comparisons between strategies unreliable. The authors conclude that the router can be effective and lightweight, and that its gains depend on the complementarity of the candidate models.
Load-bearing premise
The survey's rankings and contrasts across strategies assume the performance and cost numbers reported in different primary studies are accurate and comparable with one another, even though the paper itself concedes in Section 5.2.2 that the field lacks a standardised evaluation framework.
Editorial extensions
If this is right
- Pre-generation routing is the cheaper default when candidate performance can be predicted from the query; post-generation (cascade) routing buys reliability at the price of extra generations per query.
- Lightweight similarity and supervised routers can approach large-model quality while cutting expensive-model calls by 50 to 80 percent in the studies reviewed, making them viable for industrial deployment.
- Routing applies beyond model choice to retrieval strategies, guardrails, prompts, and embedding selection, so whole conversational pipelines can be made adaptive.
- Routers that represent candidates in a shared semantic space, such as per-cluster performance vectors or graph embeddings, can accept new models without full retraining.
- Because the field lacks standardised evaluation, the relative rankings of strategies are provisional; adopting random, oracle, and best-standalone-model baselines is needed before the comparisons become reliable.
Reading between the lines
- A direct consequence the authors leave implicit is that routing benchmarks should start treating energy and environmental cost as first-class constraints; otherwise a router that appears cheap in API dollars may simply shift the burden onto compute and carbon.
- The taxonomy suggests a concrete modular design: a router could be updated for a new model by computing one small performance vector or cluster assignment, making the router a plug-and-play component of evolving LLM systems.
- If the field adopted the baselines the survey recommends, a routing method that only beats random routing when the candidate pool is skewed would be exposed; the survey's own comparisons already hint that some reported gains vanish under such controls.
- The optimisation framing points toward end-to-end co-adaptation: the router and the downstream pipeline could be optimised together, so routing decisions adapt to changing user behaviour and model availability rather than being fixed at training time.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This survey reviews routing strategies for LLM-based systems. It formalizes routing in Eq. (1) as a constrained performance-cost optimization problem and organizes the literature along two axes: pre-generation versus post-generation routing (Section 3) and four implementation families—similarity-based, supervised, reinforcement learning-based, and generative routing (Section 4). It also discusses industrial practice and open challenges, most notably the lack of standardized evaluation. The paper's central contribution is the taxonomy and the comparative synthesis, not new empirical results.
Significance. If the taxonomy is accepted, the survey provides a useful structured map of a fast-growing area and a concrete checklist of open problems. It is careful in several places to flag missing baselines and undocumented choices in primary studies (for example, Section 4.2.3 on Shnitzer et al. and Section 4.4.5 on AutoMix). The paper does not ship code or machine-checked artifacts, which is normal for a survey; its value lies in the synthesis. The main risk is that comparative statements inherit the heterogeneous setups of the primary papers; the paper itself acknowledges this in Section 5.2.2. The scope exclusion of answer-selection and ensemble methods is clearly stated in Section 1 and is reasonable.
major comments (3)
- [Section 4.2.3] The sentence 'It outperformed several methods proposed in the survey, including RouterBench [37], Zooter multi-perceptron [65], Sakota et al. [81]’s supervised strategy...' treats RouterBench as a routing method, but RouterBench [37] is a benchmark, not a method. The comparison should name the actual baseline methods used inside RouterBench (e.g., its MLP or k-NN baselines) or be rephrased to say that MixLLM outperformed methods evaluated on RouterBench. As written, this is an inaccurate comparative statement in a section whose purpose is to rank methods.
- [Section 4.4.5 and Section 4.2.3] Section 4.4.5 reports that AutoMix outperforms FrugalGPT and HybridLLM on QASPER and COQA, and immediately notes that Wang et al. [98] found it did not outperform random routing on RouterBench; Section 4.2.3 also lists AutoMix among the methods MixLLM outperformed. These reports are not necessarily contradictory because they come from different benchmarks, model pools, and cost budgets, but the survey does not state this explicitly and continues to use such cross-paper comparisons to characterize families throughout Sections 4.1-4.4. I recommend adding an explicit comparability caveat before each cross-paper ranking or downgrading these conclusions to 'as reported in the primary study.' This is important because Section 5.2.2 itself concedes that the field lacks a standardized evaluation framework.
- [Section 4.4.3 and Appendix A] Ning et al. [74] is described in the main text as a prompt-based routing method (Section 4.4.3) and appears in Table 1 under both query complexity inference and prompt-based routing, but Appendix A labels both rows for [74] as 'No routing.' This inconsistency makes the taxonomy hard to apply to this entry and should be resolved, either by correcting the appendix or by clarifying that the routing variant is a re-analysis of the original work.
minor comments (4)
- [Section 1, Eq. (1)] Equation (1) is under-specified: the scoring function s(q, M) and the cost C_M(q) are not tied to the generated response, and post-generation/cascade routing does not fit the argmax-over-M form because the decision depends on the output of a previously selected model. A short clarifying remark would strengthen the claimed formalization.
- [Section 4.1.1] The statement that query-similarity methods 'often fail to capture complex query-response relationships and often perform worse than random baselines' is too strong for the cited evidence; it should be attributed to specific studies and their evaluation conditions, since the comparison setups differ.
- [Throughout] There are several typos and infelicities: 'several possible scoring function' (Section 2.1), 'Leveraging multiple nearest neighbours address this limitation' (Section 4.1.1), and a missing period before 'Sikeridis et al. [87]' in Section 4.3.1. A copy-editing pass would improve readability.
- [Section 4.2.3] The discussion of Shnitzer et al. [85] correctly notes that the labeling procedure for the training set is not described, but this caveat should also be reflected in the comparative summary for that entry rather than only in the descriptive paragraph.
Circularity Check
No significant circularity: the survey formalizes and taxonomizes external results without deriving predictions from its own fitted inputs.
full rationale
This is a literature survey, not a derivation, and I found no step in which a claimed prediction reduces to an input by construction. Equation (1) is a definitional formalization of routing as constrained optimization; it is presented as a framing device, not derived from data, so it cannot be circular. The taxonomy in Sections 3 and 4 is descriptive: strategies are grouped by routing timing and implementation family based on the surveyed papers, and the comparative statements rest on the primary studies' reported numbers rather than on any parameter fitted within this survey. Section 5.2.2 explicitly concedes that 'the field lacks a standardised framework for evaluating routing strategies,' which is a validity and comparability caveat about the underlying literature, not a circularity in the survey's own reasoning. The only self-citation is Bouvard et al. [5] in Section 1, used for the claim that RAG 'was shown to reduce hallucinations' compared to fine-tuning; that claim is unrelated to the routing taxonomy or to Eq. (1), and it is not load-bearing for the survey's organizational contribution. No fitted quantity is relabeled as a prediction, no uniqueness theorem is imported from the authors' prior work, and no ansatz is smuggled in via citation. The mischaracterization of RouterBench as a method in Section 4.2.3 is an accuracy issue, not a circular one. I therefore assign a score of 0.
Assumptions & free parameters
assumptions (4)
- domain assumption The router problem is adequately captured by Eq. (1): choose a candidate maximizing a score subject to a budget constraint.
- domain assumption Primary studies' performance and cost figures are comparable enough to support the survey's qualitative comparisons.
- ad hoc to paper The proposed taxonomy, pre-generation versus post-generation and four implementation families, is exhaustive and usable.
- ad hoc to paper Answer-selection and ensemble methods are outside the scope of routing because they prioritize performance over cost efficiency.
Cite this review
Pith. "Pith review of Doing More with Less: A Survey on Routing Strategies for Resource Optimisation in Large Language Model-Based Systems." pith.science (2026). https://pith.science/paper/XVHB66OT
@misc{pith2026250200409,
author = {Pith},
title = {Pith review of: Doing More with Less: A Survey on Routing Strategies for Resource Optimisation in Large Language Model-Based Systems},
year = {2026},
howpublished = {\url{https://pith.science/paper/XVHB66OT}},
note = {Machine review of arXiv:2502.00409}
}
read the original abstract
Large Language Model (LLM)-based systems, i.e. interconnected elements that include an LLM as a central component, such as conversational agents, are usually designed with monolithic, static architectures that rely on a single, general-purpose LLM to handle all user queries. However, these systems may be inefficient as different queries may require different levels of reasoning, domain knowledge or pre-processing. While generalist LLMs (e.g. GPT-4o, Claude-Sonnet) perform well across a wide range of tasks, they may incur significant financial, energy and computational costs. These costs may be disproportionate for simpler queries, resulting in unnecessary resource utilisation. A routing mechanism can therefore be employed to route queries to more appropriate components, such as smaller or specialised models, thereby improving efficiency and optimising resource consumption. This survey aims to provide a comprehensive overview of routing strategies in LLM-based systems. Specifically, it reviews when, why, and how routing should be integrated into LLM pipelines to improve efficiency, scalability, and performance. We define the objectives to optimise, such as cost minimisation and performance maximisation, and discuss the timing of routing within the LLM workflow, whether it occurs before or after generation. We also detail the various implementation strategies, including similarity-based, supervised, reinforcement learning-based, and generative methods. Practical considerations such as industrial applications and current limitations are also examined, like standardising routing experiments, accounting for non-financial costs, and designing adaptive strategies. By formalising routing as a performance-cost optimisation problem, this survey provides tools and directions to guide future research and development of adaptive low-cost LLM-based systems.
Figures
Forward citations
Cited by 6 Pith papers
-
RedactOR: An LLM-Powered Framework for Automatic Clinical Data De-Identification
RedactOR combines schema rules, multi-pass LLM entity extraction, retrieval-based relexicalization, and ASR plus VAD audio redaction, achieving F1 0.9646 on a 100-note i2b2 2014 subsample.
-
SHERPA: A Model-Driven Framework for Large Language Model Execution
A framework that executes LLM tasks through hierarchical state machines improves output quality in 12 of 15 comparisons, but the evaluation lacks error bars and includes test-set-informed design choices.
-
Conformal Arbitrage: Risk-Controlled Balancing of Competing Objectives in Language Models
Conformal Arbitrage calibrates a score-gap threshold with conformal risk control so that a primary model can act when confident and defer to a guardian otherwise, with the expected guardrail loss bounded by a user-cho...
-
Adaptive Minds: Empowering Agents with LoRA-as-Tools
Adaptive Minds makes a base LLM select LoRA adapters as tools per query; the 5-adapter demo gets 100% routing on 25 queries, while the abstract's 30-adapter/nine-family numbers are unsupported.
-
CoE-Ops: Collaboration of LLM-based Experts for AIOps Question-Answering
CoE-Ops routes DevOps questions to specialized LLM experts using an LLM classifier plus retrieval, reporting gains on DevOps-Eval that are compromised by possible test-set leakage.
-
Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques
A survey of LLM routing and hierarchical inference techniques that proposes an unvalidated unified evaluation metric called the Inference Efficiency Score.
Reference graph
Works this paper leans on
-
[37]
‘RouterBench: A Benchmark for Multi-LLM Routing System’
Qitian Jason Hu, Jacob Bieker, Xiuyu Li, Nan Jiang, Benjamin Keigwin, Gaurav Rangan- ath, Kurt Keutzer and Shriyash Kaustubh Upadhyay. ‘RouterBench: A Benchmark for Multi-LLM Routing System’. In:Agentic Markets Workshop at ICML 2024. 2024. url: https://openreview.net/forum?id=IVXmV8Uxwh
2024
-
[65]
Jing Liu, Ruihao Gong, Mingyang Zhang, Yefei He, Jianfei Cai and Bohan Zhuang.ME- Switch: A Memory-Efficient Expert Switching Framework for Large Language Models
-
[81]
Josef Pichlmeier, Philipp Ross and Andre Luckow.Performance Characterization of Ex- pert Router for Scalable LLM Inference. 2024. arXiv:2404.15153 [cs.CL]
arXiv 2024
-
[98]
Xinyuan Wang, Yanchi Liu, Wei Cheng, Xujiang Zhao, Zhengzhang Chen, Wenchao Yu, Yanjie Fu and Haifeng Chen.MixLLM: Dynamic Routing in Mixed Large Language Mod- els. 2025. arXiv:2502.18482 [cs.CL]
arXiv 2025
-
[74]
XuefeiNing,ZinanLin,ZixuanZhou,ZifuWang,HuazhongYangandYuWang.‘Skeleton- of-Thought:PromptingLLMsforEfficientParallelGeneration’.In: The 12th International Conference on Learning Representations. 2024. url: https://openreview.net/forum? id=mqVgBbNCm9
2024
-
[1]
2024.url: https://openreview.net/forum? id=e6WrwIvgzX
Pranjal Aggarwal, Aman Madaan, Ankit Anand, Srividya Pranavi Potharaju, Swaroop Mishra, Pei Zhou, Aditya Gupta, Dheeraj Rajagopal, Karthik Kappaganthu, Yiming Yang etal.‘AutoMix:AutomaticallyMixingLanguageModels’.In: The 38th Annual Conference on Neural Information Processing Systems. 2024.url: https://openreview.net/forum? id=e6WrwIvgzX
2024
-
[2]
Jinze Bai, Shuai Bai, Yunfei Chu, Zeyu Cui, Kai Dang, Xiaodong Deng, Yang Fan, Wenbin Ge, Yu Han, Fei Huang et al.Qwen Technical Report. 2023. arXiv:2309.16609 [cs.CL]
arXiv 2023
-
[3]
Adarsh Prasad Behera, Jaya Prakash Champati, Roberto Morabito, Sasu Tarkoma and James Gross. Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques. 2025. arXiv:2506.06579 [cs.LG]
arXiv 2025
Show all 117 references
-
[4]
‘Generalization in NLI: Ways (Not) To Go Beyond Simple Heuristics’
Prajjwal Bhargava, Aleksandr Drozd and Anna Rogers. ‘Generalization in NLI: Ways (Not) To Go Beyond Simple Heuristics’. In:Proceedings of the Second Workshop on In- sights from Negative Results in NLP. Ed. by João Sedoc, Anna Rogers, Anna Rumshisky and Shabnam Tafreshi. Online...
2021 doi
-
[5]
‘Derby LLM : Évaluation comparative des approches RAG et fine-tuning’
Christophe Bouvard, Mathieu Ciancone, Antoine Gourru and Marion Schaeffer. ‘Derby LLM : Évaluation comparative des approches RAG et fine-tuning’. In:Actes de la 10 ème Conférence Nationale sur les Applications Pratiques de l’Intelligence Artificielle (APIA). Ed.byCatherineRous...
-
[6]
Ralph Allan Bradley and Milton E. Terry. ‘Rank Analysis of Incomplete Block Designs: I. The Method of Paired Comparisons.’ In:Biometrika 39 (1952), pp. 324–345.doi: https: //doi.org/10.2307/2334029. url: https://www.jstor.org/stable/2334029
1952
-
[7]
‘A Survey on Mixture of Experts in Large Language Models’
Weilin Cai, Juyong Jiang, Fan Wang, Jing Tang, Sunghun Kim and Jiayi Huang. ‘A Survey on Mixture of Experts in Large Language Models’. In:IEEE Transactions on Knowledge and Data Engineering(2025), pp. 1–20.issn: 2326-3865.doi: 10.1109/tkde. 2025.3554028. url: http://dx.doi.org...
2025
-
[8]
‘A Literature Survey of Recent Advances in Chatbots’
Guendalina Caldarini, Sardar Jaf and Kenneth McGarry. ‘A Literature Survey of Recent Advances in Chatbots’. In:Information 13.1 (2022). doi: 10.3390/info13010041. url: https://www.mdpi.com/2078-2489/13/1/41
2022 doi
-
[9]
‘An Expert is Worth One Token: Synergiz- ing Multiple Expert LLMs as Generalist via Expert Token Routing’
ZiweiChai,GuoyinWang,JingSu,TianjieZhang,XuanwenHuang,XuwuWang,Jingjing Xu, Jianbo Yuan, Hongxia Yang, Fei Wu et al. ‘An Expert is Worth One Token: Synergiz- ing Multiple Expert LLMs as Generalist via Expert Token Routing’. In:Proceedings of the 62nd Annual Meeting of the Asso...
2024 doi
-
[10]
‘A Survey on Evaluation of Large Language Models’
Yupeng Chang, Xu Wang, Jindong Wang, Yuan Wu, Linyi Yang, Kaijie Zhu, Hao Chen, Xiaoyuan Yi, Cunxiang Wang, Yidong Wang et al. ‘A Survey on Evaluation of Large Language Models’. In: ACM Trans. Intell. Syst. Technol.15.3 (2024). doi: 10 . 1145 / 3641289. url: https://doi.org/10...
2024 doi
-
[11]
Lingjiao Chen, Matei Zaharia and James Zou.FrugalGPT: How to Use Large Language Models While Reducing Cost and Improving Performance.2023.arXiv: 2305.05176 [cs.LG]
2023 arXiv
-
[12]
‘RouterDC: Query- Based Router by Dual Contrastive Learning for Assembling Large Language Models’
Shuhao Chen, Weisen Jiang, Baijiong Lin, James Kwok and Yu Zhang. ‘RouterDC: Query- Based Router by Dual Contrastive Learning for Assembling Large Language Models’. In: The 38th Annual Conference on Neural Information Processing Systems. 2024. url: https://openreview.net/forum...
2024
-
[13]
Yi Chen, JiaHao Zhao and HaoHao Han.A Survey on Collaborative Mechanisms Between Large and Small Language Models. 2025. arXiv:2505.07460 [cs.AI]
2025 arXiv
-
[14]
Yu.Harnessing Multiple Large Language Models: A Survey on LLM Ensemble
Zhijun Chen, Jingzheng Li, Pengpeng Chen, Zhuoran Li, Kai Sun, Yuankai Luo, Qianren Mao, Dingqi Yang, Hailong Sun and Philip S. Yu.Harnessing Multiple Large Language Models: A Survey on LLM Ensemble. 2025. arXiv:2502.18036 [cs.CL]
2025 arXiv
-
[16]
Yu-Neng Chuang, Helen Zhou, Prathusha Kameswara Sarma, Parikshit Gopalan, John Boccio, Sara Bolouki and Xia Hu.Learning to Route with Confidence Tokens. 2024. arXiv: 2410.13284 [cs.CL]
2024 arXiv
-
[17]
In: Journal of Machine Learning Research25.70 (2024), pp
Hyung Won Chung, Le Hou, Shayne Longpre, Barret Zoph, Yi Tai, William Fedus, YunxuanLi,XuezhiWang,MostafaDehghani,SiddharthaBrahmaetal.‘ScalingInstruction- Finetuned Language Models’. In: Journal of Machine Learning Research25.70 (2024), pp. 1–53. url: http://jmlr.org/papers/v...
2024
-
[18]
A Complete Survey on LLM-based AI Chatbots
Sumit Kumar Dam, Choong Seon Hong, Yu Qiao and Chaoning Zhang. A Complete Survey on LLM-based AI Chatbots. 2024. arXiv:2406.16937 [cs.CL]
2024 arXiv
-
[19]
Jasper Dekoninck, Maximilian Baader and Martin Vechev.A Unified Approach to Routing and Cascading for LLMs. 2025. arXiv:2410.10347 [cs.CL]
2025 arXiv
-
[20]
‘Trends in AI inference energy consumption: Beyond the performance-vs-parameter laws of deep learning’
Radosvet Desislavov, Fernando Martínez-Plumed and José Hernández-Orallo. ‘Trends in AI inference energy consumption: Beyond the performance-vs-parameter laws of deep learning’. In: Sustainable Computing: Informatics and Systems 38 (2023), p. 100857. doi: https : / / doi . org ...
2023
-
[21]
‘BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding’
Jacob Devlin, Ming-Wei Chang, Kenton Lee and Kristina Toutanova. ‘BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding’. In:Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics.2019,pp.417...
2019 doi
-
[22]
Dujian Ding, Ankur Mallick, Chi Wang, Robert Sim, Subhabrata Mukherjee, Victor Rühle, Laks V. S. Lakshmanan and Ahmed Hassan Awadallah. ‘Hybrid LLM: Cost- Efficient and Quality-Aware Query Routing’. In:The 12th International Conference on Learning Representations. 2024. url: h...
2024
-
[23]
Arpad E. Elo. Ratings of Chess Players Past and Present. Arco Pub, 1978
1978
-
[24]
In:Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics
ShahulEs,JithinJames,LuisEspinosaAnkeandStevenSchockaert.‘RAGAS:Automated Evaluation of Retrieval Augmented Generation’. In:Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics. 2024, pp. 150–
2024
-
[25]
‘Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity’
William Fedus, Barret Zoph and Noam Shazeer. ‘Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity’. In:Journal of Machine Learning Research 23.120 (2022), pp. 1–39.url: http://jmlr.org/papers/v23/21-0998.html
2022
-
[26]
In: The Thirteenth International Conference on Learning Representations
TaoFeng,YanzhenShenandJiaxuanYou.‘GraphRouter:AGraph-basedRouterforLLM Selections’. In: The Thirteenth International Conference on Learning Representations
-
[27]
Johannes Fürnkranz and Eyke Hüllermeier.Preference Learning. Ed. by Claude Sammut and Geoffrey I. Webb. Springer US, 2012, pp. 1–7.doi: 10.1007/978- 1- 4899- 7502- 7_667-1. url: https://doi.org/10.1007/978-1-4899-7502-7_667-1
2012 doi
-
[28]
Yunfan Gao, Yun Xiong, Xinyu Gao, Kangxiang Jia, Jinliu Pan, Yuxi Bi, Yi Dai, Jiawei Sun, Qianyu Guo, Meng Wang et al.Retrieval-Augmented Generation for Large Language Models: A Survey. 2024. arXiv:2312.10997 [cs.CL]
2024 arXiv
-
[29]
Modular RAG: Transform- ing RAG Systems into LEGO-like Reconfigurable Frameworks
Yunfan Gao, Yun Xiong, Meng Wang and Haofen Wang. Modular RAG: Transform- ing RAG Systems into LEGO-like Reconfigurable Frameworks. 2024. arXiv: 2407.21059 [cs.CL]
2024 arXiv
-
[30]
The Llama 3 Herd of Models
Aaron Grattafiori, Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Ka- dian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Alex Vaughan et al. The Llama 3 Herd of Models. 2024. arXiv:2407.21783 [cs.AI]
2024 arXiv
-
[31]
Khare and Christopher Re
Neel Guha, Mayee F Chen, Trevor Chow, Ishan S. Khare and Christopher Re. ‘Smoothie: Label Free Language Model Routing’. In:The 38th Annual Conference on Neural Inform- ation Processing Systems. 2024. url: https://openreview.net/forum?id=pPSWHsgqRp
2024
-
[32]
Surya Narayanan Hari and Matt Thomson.Tryage: Real-time, intelligent Routing of User Prompts to Large Language Models. 2023. arXiv:2308.11601 [cs.LG]
2023 arXiv
-
[33]
Pengcheng He, Jianfeng Gao and Weizhu Chen.DeBERTaV3: Improving DeBERTa us- ing ELECTRA-Style Pre-Training with Gradient-Disentangled Embedding Sharing. 2023. arXiv: 2111.09543 [cs.CL]
2023 arXiv
-
[34]
‘DeBERTaV3: Improving DeBERTa us- ing ELECTRA-Style Pre-Training with Gradient-Disentangled Embedding Sharing’
Pengcheng He, Jianfeng Gao and Weizhu Chen. ‘DeBERTaV3: Improving DeBERTa us- ing ELECTRA-Style Pre-Training with Gradient-Disentangled Embedding Sharing’. In: The 11th International Conference on Learning Representations. 2023. url: https:// openreview.net/forum?id=sE7-XhLxHA. 25
2023
-
[35]
‘Measuring Massive Multitask Language Understanding’
Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song and Jacob Steinhardt. ‘Measuring Massive Multitask Language Understanding’. In:In- ternational Conference on Learning Representations. 2021. url: https://openreview. net/forum?id=d7KBjmI3GmQ
2021
-
[36]
Jinwu Hu, Yufeng Wang, Shuhai Zhang, Kai Zhou, Guohao Chen, Yu Hu, Bin Xiao and Mingkui Tan.Dynamic Ensemble Reasoning for LLM Experts. 2024. arXiv:2412.07448 [cs.AI]
2024
-
[38]
Zijian Hu, Jipeng Zhang, Rui Pan, Zhaozhuo Xu, Shanshan Han, Han Jin, Alay Dilipbhai Shah, Dimitris Stripelis, Yuhang Yao, Salman Avestimehr et al.Fox-1 Technical Report
-
[39]
Irugalbandara, A
C. Irugalbandara, A. Mahendra, R. Daynauth, T. Arachchige, J. Dantanarayana, K. Flautner, L. Tang, Y. Kang and J. Mars. ‘Scaling Down to Scale Up: A Cost-Benefit Ana- lysis of Replacing OpenAI’s LLM with Open Source SLMs in Production’. In:2024 Inter- national Symposium on Per...
2024
-
[40]
Jacobs, Michael I
Robert A. Jacobs, Michael I. Jordan, Steven J. Nowlan and Geoffrey E. Hinton. ‘Adaptive Mixtures of Local Experts’. In:Neural Computation3.1 (1991), pp. 79–87.doi: 10.1162/ neco.1991.3.1.79
1991
-
[41]
arXiv: 2411.05281 [cs.CL]
-
[42]
‘Exploring the benefits of training expert language models over instruction tuning’
JoelJang,SeungoneKim,SeonghyeonYe,DoyoungKim,LajanugenLogeswaran,Moontae Lee, Kyungjae Lee and Minjoon Seo. ‘Exploring the benefits of training expert language models over instruction tuning’. In:Proceedings of the 40th International Conference on Machine Learning. JMLR.org, 2...
2023
-
[43]
‘Adaptive- RAG: Learning to Adapt Retrieval-Augmented Large Language Models through Question Complexity’
Soyeong Jeong, Jinheon Baek, Sukmin Cho, Sung Ju Hwang and Jong Park. ‘Adaptive- RAG: Learning to Adapt Retrieval-Augmented Large Language Models through Question Complexity’. In:Proceedings of the 2024 Conference of the North American Chapter of the Association for Computatio...
2024 doi
-
[44]
Swayambhoo Jain, Ravi Raju, Bo Li, Zoltan Csaki, Jonathan Li, Kaizhao Liang, Guoyao Feng, Urmish Thakkar, Anand Sampat, Raghu Prabhakar et al.Composition of Experts: A Modular Compound AI System Leveraging Large Language Models. 2024. arXiv:2412. 01868 [cs.LG]
2024
-
[45]
‘LLM-Blender: Ensembling Large Lan- guage Models with Pairwise Ranking and Generative Fusion’
Dongfu Jiang, Xiang Ren and Bill Yuchen Lin. ‘LLM-Blender: Ensembling Large Lan- guage Models with Pairwise Ranking and Generative Fusion’. In:Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics. Ed. by Anna Rogers, Jordan Boyd-Graber and Na...
2023 doi
-
[46]
Ruili Jiang, Kehai Chen, Xuefeng Bai, Zhixuan He, Juntao Li, Muyun Yang, Tiejun Zhao, Liqiang Nie and Min Zhang.A Survey on Human Preference Learning for Large Language Models. 2024. arXiv:2406.11191 [cs.CL]
2024 arXiv
-
[47]
Albert Q. Jiang, Alexandre Sablayrolles, Antoine Roux, Arthur Mensch, Blanche Savary, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Emma Bou Hanna, Florian Bressand et al.Mixtral of Experts. 2024. arXiv:2401.04088 [cs.LG]. 26
2024 arXiv
-
[48]
Kaack, Priya L
Lynn H. Kaack, Priya L. Donti, Emma Strubell, George Kamiya, Felix Creutzig and David Rolnick. ‘Aligning artificial intelligence with climate change mitigation’. In:Nature Climate Change 12.6 (2022), pp. 518–527. doi: 10.1038/s41558- 022- 01377- 7 . url: https://doi.org/10.103...
2022 doi
-
[49]
Littman and Andrew W
Leslie Pack Kaelbling, Michael L. Littman and Andrew W. Moore. ‘Reinforcement learn- ing: a survey’. In:Journal of Artificial Intelligence Research4.1 (1996), pp. 237–285
1996
-
[50]
Universal Model Routing for Efficient LLM Inference
WittawatJitkrittum,HarikrishnaNarasimhan,AnkitSinghRawat,JeeveshJuneja,Zifeng Wang, Chen-Yu Lee, Pradeep Shenoy, Rina Panigrahy, Aditya Krishna Menon and Sanjiv Kumar. Universal Model Routing for Efficient LLM Inference. 2025. arXiv:2502.08773 [cs.CL]
2025 arXiv
-
[51]
‘Contrastive Representation Learning: A Framework and Review’
Phuc Le-Khac, Graham Healy and Alan Smeaton. ‘Contrastive Representation Learning: A Framework and Review’. In: IEEE Access 8 (Jan. 2020), pp. 193907–193934. doi: 10.1109/ACCESS.2020.3031549
2020
-
[52]
Chankyu Lee, Rajarshi Roy, Mengyao Xu, Jonathan Raiman, Mohammad Shoeybi, Bryan Catanzaro and Wei Ping.NV-Embed: Improved Techniques for Training LLMs as Gener- alist Embedding Models. 2025. arXiv:2405.17428 [cs.CL]
2025 arXiv
-
[54]
SafeRoute: Adaptive Model Selection for Efficient and Accurate Safety Guardrails in Large Language Models
Seanie Lee, Dong Bok Lee, Dominik Wagner, Minki Kang, Haebin Seong, Tobias Bocklet, Juho Lee and Sung Ju Hwang. SafeRoute: Adaptive Model Selection for Efficient and Accurate Safety Guardrails in Large Language Models. 2025. arXiv:2502.12464 [cs.CL]. 27
2025 arXiv
-
[55]
‘BART: Denoising Sequence-to- Sequence Pre-training for Natural Language Generation, Translation, and Comprehen- sion’
Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Veselin Stoyanov and Luke Zettlemoyer. ‘BART: Denoising Sequence-to- Sequence Pre-training for Natural Language Generation, Translation, and Comprehen- sion’. In: Proceedings of the 58th...
2020
-
[56]
‘OrchestraLLM: Efficient Orchestration of Language Models for Dialogue State Tracking’
Chia-Hsuan Lee, Hao Cheng and Mari Ostendorf. ‘OrchestraLLM: Efficient Orchestration of Language Models for Dialogue State Tracking’. In:Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologi...
2024 doi
-
[57]
‘Making Text Embedders Few-Shot Learners’
Chaofan Li, Minghao Qin, Shitao Xiao, Jianlyu Chen, Kun Luo, Defu Lian, Yingxia Shao and Zheng Liu. ‘Making Text Embedders Few-Shot Learners’. In:The Thirteenth International Conference on Learning Representations. 2025.url: https://openreview. net/forum?id=wfLuiDjQ0u
2025
-
[58]
LLM Bandit: Cost-Efficient LLM Generation via Preference-Conditioned Dy- namic Routing
Yang Li. LLM Bandit: Cost-Efficient LLM Generation via Preference-Conditioned Dy- namic Routing. 2025. arXiv:2502.02743 [cs.LG]
2025 arXiv
-
[59]
‘Retrieval- augmented generation for knowledge-intensive NLP tasks’
Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Na- man Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel et al. ‘Retrieval- augmented generation for knowledge-intensive NLP tasks’. In:Proceedings of the 34th In- ternational Co...
2020
-
[60]
Chien-Chang Lin, Anna Y. Q. Huang and Stephen J. H. Yang. ‘A Review of AI-Driven Conversational Chatbots Implementation Methodologies and Challenges (1999–2022)’. In: Sustainability 15.5 (2023), pp. 1–13.doi: 10.3390/su15054012. url: https://www.mdpi. com/2071-1050/15/5/4012
2023 doi
-
[61]
‘ROUGE: A Package for Automatic Evaluation of Summaries’
Chin-Yew Lin. ‘ROUGE: A Package for Automatic Evaluation of Summaries’. In:Text Summarization Branches Out. Barcelona, Spain: Association for Computational Linguist- ics, 2004, pp. 74–81.url: https://aclanthology.org/W04-1013/
2004
-
[62]
‘Retrieval Augmented Generation or Long-Context LLMs? A Comprehensive Study and Hybrid Approach’
Zhuowan Li, Cheng Li, Mingyang Zhang, Qiaozhu Mei and Michael Bendersky. ‘Retrieval Augmented Generation or Long-Context LLMs? A Comprehensive Study and Hybrid Approach’. In:Proceedings of the 2024 Conference on Empirical Methods in Natural Lan- guage Processing: Industry Trac...
2024 doi
-
[63]
Yueyue Liu, Hongyu Zhang, Yuantian Miao, Van-Hoang Le and Zhiqiang Li.OptLLM: Optimal Assignment of Queries to Large Language Models. 2024. arXiv: 2405 . 15130 [cs.SE]. 28
2024
-
[66]
Sasha Luccioni, Bruna Trevelin and Margaret Mitchell.The Environmental Impacts of AI – Policy Primer. 2024. doi: 10.57967/hf/3004. url: https://doi.org/10.57967/hf/ 3004
2024 doi
-
[67]
‘Fine-Tuning LLaMA for Multi-Stage Text Retrieval’
Xueguang Ma, Liang Wang, Nan Yang, Furu Wei and Jimmy Lin. ‘Fine-Tuning LLaMA for Multi-Stage Text Retrieval’. In:SIGIR ’24: Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval. Association for Computing Machinery, 20...
2024
-
[68]
‘To- wards Optimizing SQL Generation via LLM Routing’
Mohammadhossein Malekpour, Nour Shaheen, Foutse Khomh and Amine Mhedhbi. ‘To- wards Optimizing SQL Generation via LLM Routing’. In: NeurIPS 2024 Third Table Representation Learning Workshop. 2024. url: https://openreview.net/forum?id= VYvYR7U7s3
2024
-
[69]
‘Routing to the Expert: Efficient Reward-guided Ensemble of Large Language Models’
Keming Lu, Hongyi Yuan, Runji Lin, Junyang Lin, Zheng Yuan, Chang Zhou and Jin- gren Zhou. ‘Routing to the Expert: Efficient Reward-guided Ensemble of Large Language Models’. In:Proceedings of the 2024 Conference of the North American Chapter of the As- sociation for Computati...
2024 doi
-
[70]
Alireza Mohammadshahi, Ali Shaikh and Majid Yazdani.Routoo: Learning to Route to Large Language Models Effectively. 2024. arXiv:2401.13979 [cs.CL]
2024 arXiv
-
[71]
‘Generative Representational Instruction Tuning’
Niklas Muennighoff, Hongjin SU, Liang Wang, Nan Yang, Furu Wei, Tao Yu, Aman- preet Singh and Douwe Kiela. ‘Generative Representational Instruction Tuning’. In:The Thirteenth International Conference on Learning Representations. 2025. url: https:// openreview.net/forum?id=BC4lIvfSzv
2025
-
[72]
Narendra and M
Kumpati S. Narendra and M. A. L. Thathachar. ‘Learning Automata - A Survey’. In: IEEE Transactions on Systems, Man, and Cybernetics SMC-4.4 (1974), pp. 323–334. doi: 10.1109/TSMC.1974.5408453
1974
-
[73]
Dimitrios Michael Manias, Ali Chouman and Abdallah Shami.Semantic Routing for En- hanced Performance of LLM-Assisted Intent-Based 5G Core Network Management and Orchestration. 2024. arXiv:2404.15869 [cs.NI]
2024 arXiv
-
[75]
Gonzalez, M Waleed Kadous and Ion Stoica
Isaac Ong, Amjad Almahairi, Vincent Wu, Wei-Lin Chiang, Tianhao Wu, Joseph E. Gonzalez, M Waleed Kadous and Ion Stoica. ‘RouteLLM: Learning to Route LLMs from Preference Data’. In:The Thirteenth International Conference on Learning Representa- tions. 2025. url: https://openrev...
2025
-
[76]
Gonzalez
Shishir G Patil, Tianjun Zhang, Xin Wang and Joseph E. Gonzalez. ‘Gorilla: Large Lan- guage Model Connected with Massive APIs’. In:The 38th Annual Conference on Neural Information Processing Systems. 2024. url: https : / / openreview . net / forum ? id = tBRNC6YemY
2024
-
[77]
Nguyen, Duy C
Quang H. Nguyen, Duy C. Hoang, Juliette Decugis, Saurav Manchanda, Nitesh V. Chawla and Khoa D. Doan.MetaLLM: A High-performant and Cost-efficient Dynamic Framework for Wrapping LLMs. 2024. arXiv:2407.10834 [cs.LG]. 29
2024 arXiv
-
[78]
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever et al. Language models are unsupervised multitask learners. 2019. url: https://cdn.openai. com / better - language - models / language _ models _ are _ unsupervised _ multitask _ learners.pdf
2019
-
[79]
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li and Peter J. Liu. ‘Exploring the limits of transfer learn- ing with a unified text-to-text transformer’. In:Journal of Machine Learning Research 21.1 (2020), pp. 1–67.url...
2020
-
[80]
‘Optimising Calls to Large Language Models with Uncertainty-Based Two-Tier Selection’
Guillem Ramírez, Alexandra Birch and Ivan Titov. ‘Optimising Calls to Large Language Models with Uncertainty-Based Two-Tier Selection’. In:First Conference on Language Modeling. 2024. url: https://openreview.net/forum?id=T9cOYH0wGF
2024
-
[82]
‘DistilBERT, a dis- tilledversionofBERT:smaller,faster,cheaperandlighter’.In: CoRR cs.CL/1910.01108v4 (2020)
Victor Sanh, Lysandre Debut, Julien Chaumond and Thomas Wolf. ‘DistilBERT, a dis- tilledversionofBERT:smaller,faster,cheaperandlighter’.In: CoRR cs.CL/1910.01108v4 (2020). url: https://arxiv.org/abs/1910.01108
2020 arXiv
-
[83]
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford and Oleg Klimov.Proximal Policy Optimization Algorithms. 2017. arXiv:1707.06347 [cs.LG]
2017 arXiv
-
[84]
‘HuggingGPT: Solving AI Tasks with ChatGPT and its Friends in Hugging Face’
Yongliang Shen, Kaitao Song, Xu Tan, Dongsheng Li, Weiming Lu and Yueting Zhuang. ‘HuggingGPT: Solving AI Tasks with ChatGPT and its Friends in Hugging Face’. In:37th Conference on Neural Information Processing Systems. 2023.url: https://openreview. net/forum?id=yHdTscY6Ci
2023
-
[85]
‘Fly-Swat or Cannon? Cost-Effective Language Model Choice via Meta-Modeling’
Marija Sakota, Maxime Peyrard and Robert West. ‘Fly-Swat or Cannon? Cost-Effective Language Model Choice via Meta-Modeling’. In:Proceedings of the 17th ACM Interna- tional Conference on Web Search and Data Mining. 2024, pp. 606–615. doi: 10.1145/ 3616855.3635825. url: http://d...
2024
-
[86]
‘Getting MoRE out of Mixture of Language Model Reasoning Experts’
Chenglei Si, Weijia Shi, Chen Zhao, Luke Zettlemoyer and Jordan Boyd-Graber. ‘Getting MoRE out of Mixture of Language Model Reasoning Experts’. In:Findings of the Asso- ciation for Computational Linguistics: EMNLP 2023. Ed. by Houda Bouamor, Juan Pino and Kalika Bali. Singapor...
2023
-
[87]
Dimitrios Sikeridis, Dennis Ramdass and Pranay Pareek.PickLLM: Context-Aware RL- Assisted Large Language Model Routing. 2024. arXiv:2412.12170 [cs.LG]
2024 arXiv
-
[88]
‘MoDEM: Mixture of Domain Ex- pert Models’
Toby Simonds, Kemal Kurniawan and Jey Han Lau. ‘MoDEM: Mixture of Domain Ex- pert Models’. In:Proceedings of the 22nd Annual Workshop of the Australasian Language Technology Association. Ed. by Tim Baldwin, Sergio José Rodríguez Méndez and Nicholas Kuo. 2024, pp. 75–88.url: ht...
2024
-
[89]
Tal Shnitzer, Anthony Ou, Mírian Silva, Kate Soule, Yuekai Sun, Justin Solomon, Neil Thompson and Mikhail Yurochkin.Large Language Model Routing with Benchmark Data- sets. 2024. arXiv:2309.15789 [cs.CL]. 30
2024 arXiv
-
[90]
‘Harnessing the Power of Multiple Minds: Lessons Learned from LLM Routing’
Kv Aditya Srivatsa, Kaushal Maurya and Ekaterina Kochmar. ‘Harnessing the Power of Multiple Minds: Lessons Learned from LLM Routing’. In: Proceedings of the Fifth Workshop on Insights from Negative Results in NLP. Ed. by Shabnam Tafreshi, Arjun Akula, João Sedoc, Aleksandr Dro...
2024
-
[91]
‘TensorOpera Router: A Multi-Model Router for Efficient LLM Inference’
Dimitris Stripelis, Zhaozhuo Xu, Zijian Hu, Alay Dilipbhai Shah, Han Jin, Yuhang Yao, Jipeng Zhang, Tong Zhang, Salman Avestimehr and Chaoyang He. ‘TensorOpera Router: A Multi-Model Router for Efficient LLM Inference’. In:Proceedings of the 2024 Confer- ence on Empirical Metho...
2024 doi
-
[92]
Sutton and Andrew G
Richard S. Sutton and Andrew G. Barto.Temporal-Difference Learning. The MIT Press,
-
[93]
Llama 2: Open Foundation and Fine-Tuned Chat Models
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei,NikolayBashlykov,SoumyaBatra,PrajjwalBhargava,ShrutiBhosaleetal. Llama 2: Open Foundation and Fine-Tuned Chat Models. 2023. arXiv:2307.09288 [cs.CL]
2023 arXiv
-
[94]
SeamusSomerstep,FelipeMaiaPolo,AllyssonFlavioMelodeOliveira,PrattyushMangal, Mírian Silva, Onkar Bhardwaj, Mikhail Yurochkin and Subha Maity.CARROT: A Cost Aware Rate Optimal Router. 2025. arXiv:2502.03261 [stat.ML]
2025 arXiv
-
[95]
Can Wang, Bolin Zhang, Dianbo Sui, Zhiying Tu, Xiaoyu Liu and Jiabao Kang.A sur- vey on effective invocation methods of massive LLM services. 2024. arXiv: 2402.03408 [cs.SE]. 31
2024 arXiv
-
[96]
‘Fusing Models with Complementary Expertise’
Hongyi Wang, Felipe Maia Polo, Yuekai Sun, Souvik Kundu, Eric Xing and Mikhail Yurochkin. ‘Fusing Models with Complementary Expertise’. In:The 12th International Conference on Learning Representations. 2024. url: https://openreview.net/forum? id=PhMrGCMIRL
2024
-
[97]
‘Improving Text Embeddings with Large Language Models’
Liang Wang, Nan Yang, Xiaolong Huang, Linjun Yang, Rangan Majumder and Furu Wei. ‘Improving Text Embeddings with Large Language Models’. In:Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Ed. by Lun-Wei Ku, Andre...
2024
-
[99]
Yuanshuai Wang, Xingjian Zhang, Jinkun Zhao, Siwei Wen, Peilin Feng, Shuhao Liao, Lei Huang and Wenjun Wu.Bench-CoE: a Framework for Collaboration of Experts from Benchmark. 2024. arXiv:2412.04167 [cs.AI]
2024 arXiv
-
[100]
Iulia Turc, Ming-Wei Chang, Kenton Lee and Kristina Toutanova.Well-Read Students Learn Better: The Impact of Student Initialization on Knowledge Distillation. 2019. arXiv: 1908.08962 [cs.CL]
2019 arXiv
-
[101]
‘Emergent Abilities of Large Language Models’
Jason Wei, Yi Tay, Rishi Bommasani, Colin Raffel, Barret Zoph, Sebastian Borgeaud, Dani Yogatama, Maarten Bosma, Denny Zhou, Donald Metzler et al. ‘Emergent Abilities of Large Language Models’. In:Transactions on Machine Learning Research(2022). url: https://openreview.net/for...
2022
-
[102]
‘CanLLMsExpressTheirUncertainty?AnEmpiricalEvaluationofConfidenceElicitation in LLMs’
Miao Xiong, Zhiyuan Hu, Xinyang Lu, YIFEI LI, Jie Fu, Junxian He and Bryan Hooi. ‘CanLLMsExpressTheirUncertainty?AnEmpiricalEvaluationofConfidenceElicitation in LLMs’. In:The 12th International Conference on Learning Representations. 2024.url: https://openreview.net/forum?id=g...
2024
-
[103]
An Yang, Baosong Yang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Zhou, Chengpeng Li, Chengyuan Li, Dayiheng Liu, Fei Huang et al.Qwen2 Technical Report. 2024. arXiv: 2407.10671 [cs.CL]
2024 arXiv
-
[104]
‘FLASK: Fine-grained Language Model Evaluation based on Alignment Skill Sets’
Seonghyeon Ye, Doyoung Kim, Sungdong Kim, Hyeonbin Hwang, Seungone Kim, Yongrae Jo, James Thorne, Juho Kim and Minjoon Seo. ‘FLASK: Fine-grained Language Model Evaluation based on Alignment Skill Sets’. In:The Twelfth International Conference on Learning Representations. 2024....
2024
-
[105]
‘BARTScore: Evaluating Generated Text as Text Generation’
Weizhe Yuan, Graham Neubig and Pengfei Liu. ‘BARTScore: Evaluating Generated Text as Text Generation’. In:Advances in Neural Information Processing Systems. 2021. url: https://openreview.net/forum?id=5Ya8PbvpZ9
2021
-
[106]
Hui Wei, Shenghua He, Tian Xia, Andy Wong, Jingyang Lin and Mei Han.Systematic Evaluation of LLM-as-a-Judge in LLM Alignment Tasks: Explainable Metrics and Diverse Prompt Templates. 2024. arXiv:2408.13006 [cs.CL]
2024 arXiv
-
[107]
Awadallah and Chi Wang.EcoAssistant: Using LLM Assistant More Affordably and Accurately
Jieyu Zhang, Ranjay Krishna, Ahmed H. Awadallah and Chi Wang.EcoAssistant: Using LLM Assistant More Affordably and Accurately. 2023. arXiv:2310.03046 [cs.SE]
2023 arXiv
-
[108]
Enhancing Large Language Model Performance To Answer Questions and Extract Information More Accurately
Liang Zhang, Katherine Jijo, Spurthi Setty, Eden Chung, Fatima Javid, Natan Vidra and Tommy Clifford. Enhancing Large Language Model Performance To Answer Questions and Extract Information More Accurately. 2024. arXiv:2402.01722 [cs.CL]
2024 arXiv
-
[109]
Tuo Zhang, Asal Mehradfar, Dimitrios Dimitriadis and Salman Avestimehr.Leveraging Uncertainty Estimation for Efficient LLM Routing. 2025. arXiv:2502.11021 [cs.NI]
2025 arXiv
-
[110]
An Introduction to Matrix factorization and Factorization Machines in Recommendation System, and Beyond
Yuefeng Zhang. An Introduction to Matrix factorization and Factorization Machines in Recommendation System, and Beyond. 2022. arXiv:2203.11026 [cs.IR]
2022 arXiv
-
[111]
Morley Mao.Eagle: Efficient Training-Free Router for Multi-LLM Inference
Zesen Zhao, Shuowei Jin and Z. Morley Mao.Eagle: Efficient Training-Free Router for Multi-LLM Inference. 2024. arXiv:2409.15518 [cs.LG]
2024 arXiv
-
[112]
‘Large Language Model Cas- cadeswithMixtureofThoughtRepresentationsforCost-EfficientReasoning’.In: The 12th International Conference on Learning Representations
Murong Yue, Jie Zhao, Min Zhang, Liang Du and Ziyu Yao. ‘Large Language Model Cas- cadeswithMixtureofThoughtRepresentationsforCost-EfficientReasoning’.In: The 12th International Conference on Learning Representations. 2024.url: https://openreview. net/forum?id=6okaSfANzh. 32
2024
-
[113]
‘A Robustly Optimized BERT Pre-training Approach with Post-training’
Liu Zhuang, Lin Wayne, Shi Ya and Zhao Jun. ‘A Robustly Optimized BERT Pre-training Approach with Post-training’. In:Proceedings of the 20th Chinese National Conference on Computational Linguistics. Ed. by Sheng Li, Maosong Sun, Yang Liu, Hua Wu, Kang Liu, Wanxiang Che, Shizhu...
2021
-
[114]
‘EmbedLLM: Learning Compact Representations of Large Language Models’
RichardZhuang,TianhaoWu,ZhaojinWen,AndrewLi,JiantaoJiaoandKannanRamchandran. ‘EmbedLLM: Learning Compact Representations of Large Language Models’. In:The Thirteenth International Conference on Learning Representations. 2025. url: https : //openreview.net/forum?id=Fs9EabmQrJ. ...
2025
-
[118]
‘LoraRetriever: Input-Aware LoRA Retrieval and Composition for Mixed Tasks in the Wild’
Ziyu Zhao, Leilei Gan, Guoyin Wang, Wangchunshu Zhou, Hongxia Yang, Kun Kuang and Fei Wu. ‘LoraRetriever: Input-Aware LoRA Retrieval and Composition for Mixed Tasks in the Wild’. In:ACL (Findings). 2024, pp. 4447–4462.url: https://doi.org/ 10.18653/v1/2024.findings-acl.263
2024 doi
-
[158]
url: https://aclanthology.org/2024.eacl-demo.16
2024
-
[2014]
Chap. 6, pp. 143–166. url: https : / / web . stanford . edu / class / psych209 / Readings/SuttonBartoIPRLBook2ndEd.pdf
-
[2024]
arXiv: 2406.09041 [cs.CL]
-
[2025]
url: https://openreview.net/forum?id=eU39PDsZtT
-
[8249]
url: https://aclanthology
doi: 10.18653/v1/2023.findings- emnlp.552 . url: https://aclanthology. org/2023.findings-emnlp.552/
2023 doi
Reviewed August 9, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.