REVIEW 3 major objections 4 minor 80 references
Delta Activations: A Representation for Finetuned Large Language Models
T0 review · 3 major / 4 minor · reviewed 2026-08-15 · deepseek-v4-flash
Pith's one-line read Finetuned LLMs become clusterable vectors when their internal activation shifts from the base model are averaged over generic prompts.
desk verdict A useful, cheap model-embedding method that is overclaimed as 'domain clustering' when the experiments only show dataset clustering; worth a serious referee, but the central claim needs a cross-dataset-within-domain test. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the $\Delta$ Activations vector, $\Delta_f(x) = h_f(x) - h_{\mathrm{base}}(x)$, averaged over a fixed probe dataset $D_{\mathrm{probe}}$ of five generic instruction templates. The probe prompts are deliberately task-free, so any consistent activation shift they evoke is attributed to the finetuning itself rather than to a particular task, and averaging over paraphrased templates dampens prompt-specific noise. The default extraction point is the last token of the final layer, though the paper finds 2/3 depth and weighted token averaging slightly better, and the same differencing operation applied to logits or inverse-perplexity meaning vectors defines the $\Delta$-X family. The model-agnostic $\Delta$ Meaning variant is what allows clustering across different base architectures.
What would settle it
Take one of the paper's model pools, compute $\Delta$ Activations with the published five-prompt probe set, then recompute with a second set of five paraphrased generic prompts of the same length; if the two embeddings disagree about domain membership for more than one model per backbone, or if average silhouette across both probe sets falls below, say, 0.3, the claim that the probe set is a universal lens fails.
Extended reading notes
Core claim
The central claim is that the mean difference in last-token, final-layer hidden states between a finetuned model and its base model, computed on a small fixed set of generic prompt templates, is a compact behavioral fingerprint of the finetuned model. Concretely, for $v_f = \frac{1}{N}\sum_{i=1}^N \left(h_f(x_i) - h_{\mathrm{base}}(x_i)\right)$, the paper argues that $v_f$ clusters models by finetuning domain across three open backbones, outperforming flattened LoRA weights, salient masks, and output sentence embeddings. The same vector is shown to stay informative under varied training settings and under preference optimization, and to satisfy an approximate additive property: the delta vector of a model finetuned on $D_1 \cup D_2$ is closer to the sum of the separately trained deltas than to either component alone. A five-prompt probe set of paraphrased Alpaca-style templates with no task content is sufficient, and replacing activations with a model-agnostic meaning vector ($\Delta$ Meaning) extends the representation to models finetuned from different base architectures.
Load-bearing premise
The load-bearing premise is that five short, task-free prompts reliably switch on a finetuned model's specialization in its internal activations no matter the domain, task, or backbone; if some specializations stay silent under generic prompts, the resulting clusters and additive sums will not generalize.
Editorial extensions
If this is right
- A hub that stores delta vectors can embed a newly uploaded finetuned model in one forward pass and place it in the same space as all existing models, with no retraining and no metadata.
- Domain clusters that survive learning-rate, epoch, and data-size variation make nearest-neighbor search in delta space a viable substitute for parsing model names or trusting repository labels.
- Because finetuning on 20 examples yields a task vector that lands near the cluster of fully finetuned same-domain models, task similarity can be estimated before committing to a large training run.
- The additive property implies that the delta vector of a mixed-dataset model can be approximated as the sum of component deltas, which links model composition in data space to vector arithmetic in embedding space.
- Cross-architecture clustering via Delta Meaning means model pools do not have to be restricted to a single backbone, broadening the scope of model discovery.
Reading between the lines
- A testable extension is to mine probe prompts per base model family and check whether cross-backbone domain clusters survive; if the generic lens is itself architecture-sensitive, the method's universality claim needs qualification.
- The additive property is shown with equal-proportion data mixing, so an obvious next test is whether mixture weights map to weighted vector sums; if so, delta arithmetic could predict the embedding of arbitrary data blends before training.
- The model-selection experiment found that nearest-neighbor selection hurt merging, so the practical route may be to use delta vectors to choose diverse-but-relevant subsets rather than the single most similar model.
- Because delta vectors separate models trained on the same data under different settings, they may serve as a reproducibility diagnostic that flags finetuning runs whose internal behavior drifted despite identical data.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper proposes Delta Activations, defined in Section 2.4 as the average over a fixed probe set of the difference between a finetuned model's last-token, last-layer hidden state and that of its base model. The authors evaluate this embedding by constructing pools of LoRA-finetuned LLaMA-3.1-8B, Gemma-2-9B, and Qwen-2.5-7B models across five domains, reporting silhouette scores around 0.61 and favorable comparisons to flattened weights, salient masks, and output sentence embeddings. They also report robustness to training hyperparameters, an additive property for mixed datasets, few-shot task embeddings, cross-base transfer with Delta Meaning, and a proof-of-concept BBH model-selection improvement.
Significance. The proposed representation is simple, inexpensive (one forward pass per model), and does not require metadata or joint fitting, which makes it genuinely attractive for model-hub applications. The paper contains useful ablations of probe-prompt number, length, and content, token and layer position, and checks across three backbones and several training settings, and the code release supports reproducibility. If the domain-and-task clustering claim survives the dataset-identity confound identified below, this would be a practical contribution to model discovery and reuse. At present, however, the headline claim is not yet established at the level claimed.
major comments (3)
- [Section 3.1, Table 2] The 'domain' clusters are built from three disjoint splits of a single dataset per domain (LegalBench, GSM8K, PubMedQA, HellaSwag, OPC-SFT). Consequently, the silhouette score of ~0.61 demonstrates only that models finetuned on different splits of the same dataset are close in Delta-Activation space; it does not demonstrate that models finetuned on different datasets within the same domain cluster together. Dataset-specific cues such as answer formats, vocabulary, and the training prompt templates in Appendix A.2 could drive the separation. This confound propagates to the additive property (Tables 4 and 13), the few-shot task embeddings (Section 3.3), and the BBH experiment, all of which use the same five datasets. The intra-domain experiment in Table 14 does not resolve the issue because it clusters by dataset/sub-expertise rather than showing invariance across datasets within a domain. Please add a cross-dataset-within-domain experiment (e.g., finetune on a second legal reasoning dataset and a second math dataset and test whether the two legal models land in one cluster) or soften the 'domain' claim to 'dataset identity'.
- [Section 3.2, Tables 5-7] The probe prompts and extraction location are selected via ablations on the same pools used to report the main result. The reported silhouette score is therefore conditional on those choices, and all scores are single numbers with no variance estimate or significance test. Since the pools contain only 15 models each (five clusters of three), the difference between 0.61 and, for example, the 0.51 of Delta Logits may not be meaningful. Please report bootstrap confidence intervals or cluster-level score distributions. In the robustness table, the Qwen row with different learning rates reaches only 0.23, which weakens the statement that varied training settings 'generally did not break domain-specific clustering.'
- [Section 3.2, Table 4 and Appendix B.2] The additive property is central to the paper's claims, but the similarity metric used for 'Mixed vs. D1', 'Mixed vs. D2', and 'Mixed vs. Sum' is never defined. Please specify whether cosine similarity is used, report per-pair values with a measure of dispersion, and compare the additive match to a baseline such as similarity to a random vector or to a model finetuned on an unrelated dataset. This would make clear whether the additive effect is specific to the two constituent datasets rather than a generic property of any combination.
minor comments (4)
- [Table 2] The dimension for flattened adapter weights is reported as ~2e7 while the salient mask is ~8e9; please clarify whether these are computed over different parameter sets and explain the discrepancy.
- [Section 3.2, Table 6] The caption states that the 2/3-depth layer performs best, yet the default remains the final layer. Since the difference is small (0.64 vs. 0.61), please provide a justification or a significance analysis for this choice.
- [Throughout] Please proofread for typographical errors: for example, 'LL AMA' in Section 2.3 and 'layerees' in the Table 6 caption. The captions of Figures 4 and 5 are missing closing periods.
- [Section 2.3 and Table 1] The motivating observation is supported only by a few selected output examples. Since the method itself uses activations rather than outputs, this is not a fatal issue, but it would be helpful to state explicitly that this is an exploratory observation rather than evidence for the effectiveness of the final method.
Circularity Check
No circularity found: Delta Activations are measured directly from fixed probe activations; clustering and additivity are empirical evaluations, not fitted predictions.
full rationale
No load-bearing step reduces to its own inputs. The embedding is defined as v_f = (1/N) sum_i (h_f(x_i) - h_base(x_i)) over a fixed probe dataset D_probe. No parameter is fitted to domain labels, and the silhouette scores are computed from these unoptimized vectors. The probe prompts, token position, and layer are selected through ablations that optimize clustering quality, but the reported values are measurements under the resulting protocol rather than predictions derived from a fitted quantity; choosing a representation by validation is standard method selection and does not make the evaluated clustering tautological. The additive property is an empirical comparison of v(mixed) with v(D1) + v(D2), not an identity imposed by construction. Few-shot task embeddings are produced by finetuning on held-out examples and then retrieving clusters, so the retrieval is an out-of-sample generalization check. The same-author citations (e.g., [58] and [68]) are background or motivation, not evidence that the method works. The paper's main validity concern—Section 3.1 uses one dataset per domain, so 'domain clusters' may partly reflect dataset identity rather than domain specialization—is a correctness or confound issue, not a circularity issue, because the embedding is not constructed from the domain labels or from any fitted parameter.
Assumptions & free parameters
free parameters (3)
- Probe prompt set
- Activation extraction location
- Probe dataset size N =
5
assumptions (2)
- domain assumption A generic instruction template elicits specialization-related activation shifts in finetuned LLMs.
- domain assumption Last-token hidden states of decoder-only LLMs are comparable across models sharing the same base and tokenizer.
Cite this review
Pith. "Pith review of Delta Activations: A Representation for Finetuned Large Language Models." pith.science (2026). https://pith.science/paper/W4SIPIR7
@misc{pith2026250904442,
author = {Pith},
title = {Pith review of: Delta Activations: A Representation for Finetuned Large Language Models},
year = {2026},
howpublished = {\url{https://pith.science/paper/W4SIPIR7}},
note = {Machine review of arXiv:2509.04442}
}
read the original abstract
The success of powerful open source Large Language Models (LLMs) has enabled the community to create a vast collection of post-trained models adapted to specific tasks and domains. However, navigating and understanding these models remains challenging due to inconsistent metadata and unstructured repositories. We introduce Delta Activations, a method to represent finetuned models as vector embeddings by measuring shifts in their internal activations relative to a base model. This representation allows for effective clustering by domain and task, revealing structure in the model landscape. Delta Activations also demonstrate desirable properties: it is robust across finetuning settings and exhibits an additive property when finetuning datasets are mixed. In addition, we show that Delta Activations can embed tasks via few-shot finetuning, and further explore its use for model selection and merging. We hope Delta Activations can facilitate the practice of reusing publicly available models. Code is available at https://github.com/OscarXZQ/delta_activations.
Figures
Figures from the paper (5 more)
Reference graph
Works this paper leans on
-
[1]
Training a helpful and harmless assistant with reinforcement learning from human feedback
Yuntao Bai, Andy Jones, Kamal Ndousse, Amanda Askell, Anna Chen, Nova DasSarma, Dawn Drain, Stanislav Fort, Deep Ganguli, Tom Henighan, et al. Training a helpful and harmless assistant with reinforcement learning from human feedback. arXiv preprint arXiv:2204.05862, 2022
arXiv 2022
-
[2]
Perception encoder: The best visual embeddings are not at the output of the network
Daniel Bolya, Po-Yao Huang, Peize Sun, Jang Hyun Cho, Andrea Madotto, Chen Wei, Tengyu Ma, Jiale Zhi, Jathushan Rajasegaran, Hanoona Rasheed, Junke Wang, Marco Monteiro, Hu Xu, Shiyu Dong, Nikhila Ravi, Daniel Li, Piotr Dollár, and Christoph Feichtenhofer. Perception encoder: The best visual embeddings are not at the output of the network. arXiv:2504.13181, 2025
arXiv 2025
-
[3]
Enhancing human-like responses in large language models
Ethem Ya˘gız Çalık and Talha Rüzgar Akku¸ s. Enhancing human-like responses in large language models. arXiv preprint arXiv:2501.05032, 2025
arXiv 2025
-
[4]
Generative pretraining from pixels
Mark Chen, Alec Radford, Rewon Child, Jeffrey Wu, Heewoo Jun, David Luan, and Ilya Sutskever. Generative pretraining from pixels. In ICML, 2020
2020
-
[5]
Adaptersoup: Weight averaging to improve generalization of pretrained language models
Alexandra Chronopoulou, Matthew E Peters, Alexander Fraser, and Jesse Dodge. Adaptersoup: Weight averaging to improve generalization of pretrained language models. In ACL, 2023
work page 2023
-
[6]
Scaling instruction-finetuned language models
Hyung Won Chung, Le Hou, Shayne Longpre, Barret Zoph, Yi Tay, William Fedus, Yunxuan Li, Xuezhi Wang, Mostafa Dehghani, Siddhartha Brahma, et al. Scaling instruction-finetuned language models. Journal of Machine Learning Research, 2024
2024
-
[7]
Training verifiers to solve math word problems
Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, Christopher Hesse, and John Schulman. Training verifiers to solve math word problems. arXiv preprint arXiv:2110.14168, 2021
- [8]
Show all 80 references
-
[9]
Ultrafeedback: Boosting language models with scaled ai feedback
Ganqu Cui, Lifan Yuan, Ning Ding, Guanming Yao, Bingxiang He, Wei Zhu, Yuan Ni, Guotong Xie, Ruobing Xie, Yankai Lin, et al. Ultrafeedback: Boosting language models with scaled ai feedback. In ICML, 2024
2024
-
[10]
Tim Dettmers, Mike Lewis, Younes Belkada, and Luke Zettlemoyer. Gpt3. int8 (): 8-bit matrix multiplication for transformers at scale. Advances in neural information processing systems, pp. 30318–30332, 2022
2022
-
[11]
Model swarms: Col- laborative search to adapt llm experts via swarm intelligence
Shangbin Feng, Zifeng Wang, Yike Wang, Sayna Ebrahimi, Hamid Palangi, Lesly Miculicich, Achin Kulshrestha, Nathalie Rauschmayr, Yejin Choi, Yulia Tsvetkov, et al. Model swarms: Col- laborative search to adapt llm experts via swarm intelligence. arXiv preprint arXiv:2410.11163, 2024
-
[12]
Sparsegpt: Massive language models can be accurately pruned in one-shot
Elias Frantar and Dan Alistarh. Sparsegpt: Massive language models can be accurately pruned in one-shot. In International Conference on Machine Learning, pp. 10323–10337, 2023
2023
-
[13]
Gptq: Accurate post-training quantization for generative pre-trained transformers
Elias Frantar, Saleh Ashkboos, Torsten Hoefler, and Dan Alistarh. Gptq: Accurate post-training quantization for generative pre-trained transformers. arXiv preprint arXiv:2210.17323, 2022
2022 arXiv
-
[14]
Disease database
FreedomIntelligence. Disease database. https://huggingface.co/datasets/ FreedomIntelligence/Disease_Database, 2024. Hugging Face dataset
2024
-
[15]
Ho, Christopher Ré, Adam Chilton, Aditya Narayana, Alex Chohlas-Wood, Austin Peters, Brandon Waldon, Daniel N
Neel Guha, Julian Nyarko, Daniel E. Ho, Christopher Ré, Adam Chilton, Aditya Narayana, Alex Chohlas-Wood, Austin Peters, Brandon Waldon, Daniel N. Rockmore, Diego Zambrano, Dmitry Talisman, Enam Hoque, Faiz Surani, Frank Fagan, Galit Sarfaty, Gregory M. Dickinson, Haggai Porat...
2023
-
[16]
Learning both weights and connections for efficient neural network
Song Han, Jeff Pool, John Tran, and William Dally. Learning both weights and connections for efficient neural network. In NeurIPS, 2015
2015
-
[17]
Localize-and-stitch: Efficient model merging via sparse task arithmetic
Yifei He, Yuzheng Hu, Yong Lin, Tong Zhang, and Han Zhao. Localize-and-stitch: Efficient model merging via sparse task arithmetic. Transactions on Machine Learning Research, 2024,
2024
-
[18]
Distilling the knowledge in a neural network
Geoffrey Hinton, Oriol Vinyals, and Jeff Dean. Distilling the knowledge in a neural network. arXiv preprint arXiv:1503.02531, 2015
2015 arXiv
-
[19]
Unsupervised model tree heritage recovery
Eliahu Horwitz, Asaf Shul, and Yedid Hoshen. Unsupervised model tree heritage recovery. In ICLR, 2025
2025
-
[20]
LoRA: Low-rank adaptation of large language models
Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. LoRA: Low-rank adaptation of large language models. In ICLR, 2022
2022
-
[21]
Lorahub: Efficient cross-task generalization via dynamic lora composition
Chengsong Huang, Qian Liu, Bill Yuchen Lin, Tianyu Pang, Chao Du, and Min Lin. Lorahub: Efficient cross-task generalization via dynamic lora composition. In COLM, 2024
2024
-
[22]
Opencoder: The open cookbook for top-tier code large language models
Siming Huang, Tianhao Cheng, Jason Klein Liu, Jiaran Hao, Liuyihan Song, Yang Xu, J Yang, JH Liu, Chenchen Zhang, Linzheng Chai, et al. Opencoder: The open cookbook for top-tier code large language models. arXiv preprint arXiv:2411.04905, 2024
2024 arXiv
-
[23]
Editing models with task arithmetic
Gabriel Ilharco, Marco Tulio Ribeiro, Mitchell Wortsman, Suchin Gururangan, Ludwig Schmidt, Hannaneh Hajishirzi, and Ali Farhadi. Editing models with task arithmetic. In ICLR, 2023
2023
-
[24]
Smith, Iz Beltagy, and Hannaneh Hajishirzi
Hamish Ivison, Yizhong Wang, Valentina Pyatkin, Nathan Lambert, Matthew Peters, Pradeep Dasigi, Joel Jang, David Wadden, Noah A. Smith, Iz Beltagy, and Hannaneh Hajishirzi. Camels in a changing climate: Enhancing lm adaptation with tulu 2, 2023
2023
-
[25]
What disease does this patient have? a large-scale open domain question answering dataset from medical exams
Di Jin, Eileen Pan, Nassim Oufattole, Wei-Hung Weng, Hanyi Fang, and Peter Szolovits. What disease does this patient have? a large-scale open domain question answering dataset from medical exams. arXiv preprint arXiv:2009.13081, 2020
2009 arXiv
-
[26]
Retrieval instead of fine-tuning: A retrieval-based parameter ensemble for zero-shot learning
Pengfei Jin, Peng Shu, Sekeun Kim, Qing Xiao, Sifan Song, Cheng Chen, Tianming Liu, Xiang Li, and Quanzheng Li. Retrieval instead of fine-tuning: A retrieval-based parameter ensemble for zero-shot learning. arXiv preprint arXiv:2410.09908, 2024
-
[27]
Pubmedqa: A dataset for biomedical research question answering
Qiao Jin, Bhuwan Dhingra, Zhengping Liu, William Cohen, and Xinghua Lu. Pubmedqa: A dataset for biomedical research question answering. In EMNLP-IJCNLP, 2019
2019
-
[28]
Can this model also recognize dogs? zero-shot model search from weights
Jonathan Kahana, Or Nathan, Eliahu Horwitz, and Yedid Hoshen. Can this model also recognize dogs? zero-shot model search from weights. arXiv preprint arXiv:2502.09619, 2025
2025 arXiv
-
[29]
Matrix factorization techniques for recom- mender systems
Yehuda Koren, Robert Bell, and Chris V olinsky. Matrix factorization techniques for recom- mender systems. Computer, 2009
2009
-
[30]
Imagenet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton. Imagenet classification with deep convolutional neural networks. In NIPS, 2012
2012
-
[31]
Optimal brain damage
Yann LeCun, John Denker, and Sara Solla. Optimal brain damage. In NeurIPS, 1989
1989
-
[32]
Emergent world representations: Exploring a sequence model trained on a synthetic task
Kenneth Li, Aspen K Hopkins, David Bau, Fernanda Viégas, Hanspeter Pfister, and Martin Wattenberg. Emergent world representations: Exploring a sequence model trained on a synthetic task. ICLR, 2023
2023
-
[33]
Drag-and-drop llms: Zero-shot prompt-to-weights
Zhiyuan Liang, Dongwen Tang, Yuhao Zhou, Xuanlei Zhao, Mingjia Shi, Wangbo Zhao, Zekai Li, Peihao Wang, Konstantin Schürholt, Damian Borth, et al. Drag-and-drop llms: Zero-shot prompt-to-weights. arXiv preprint arXiv:2506.16406, 2025
2025 arXiv
-
[34]
Awq: Activation- aware weight quantization for llm compression and acceleration
Ji Lin, Jiaming Tang, Haotian Tang, Shang Yang, Xingyu Dang, and Song Han. Awq: Activation- aware weight quantization for llm compression and acceleration. MlSys, 2023. 12
2023
-
[35]
Deepseek-v3 technical report
Aixin Liu, Bei Feng, Bing Xue, Bingxuan Wang, Bochao Wu, Chengda Lu, Chenggang Zhao, Chengqi Deng, Chenyu Zhang, Chong Ruan, et al. Deepseek-v3 technical report. arXiv preprint arXiv:2412.19437, 2024
2024 arXiv
-
[36]
In-context vectors: Making in context learning more effective and controllable through latent space steering
Sheng Liu, Haotian Ye, Lei Xing, and James Zou. In-context vectors: Making in context learning more effective and controllable through latent space steering. arXiv preprint arXiv:2311.06668, 2023
2023 arXiv
-
[37]
Meaning representations from trajectories in autoregressive models
Tian Yu Liu, Matthew Trager, Alessandro Achille, Pramuditha Perera, Luca Zancato, and Stefano Soatto. Meaning representations from trajectories in autoregressive models. In ICLR, 2024
2024
-
[38]
Routing to the expert: Efficient reward-guided ensemble of large language models
Keming Lu, Hongyi Yuan, Runji Lin, Junyang Lin, Zheng Yuan, Chang Zhou, and Jingren Zhou. Routing to the expert: Efficient reward-guided ensemble of large language models. arXiv preprint arXiv:2311.08692, 2023
2023 arXiv
-
[39]
Pace: Parsimonious concept engineering for large language models
Jinqi Luo, Tianjiao Ding, Kwan Ho Ryan Chan, Darshan Thaker, Aditya Chattopadhyay, Chris Callison-Burch, and René Vidal. Pace: Parsimonious concept engineering for large language models. In NeurIPS, 2024
2024
-
[40]
Stylus: Automatic adapter selection for diffusion models
Michael Luo, Justin Wong, Brandon Trabucco, Yanping Huang, Joseph E Gonzalez, Ruslan Salakhutdinov, Ion Stoica, et al. Stylus: Automatic adapter selection for diffusion models. In NeurIPS, 2024
2024
-
[41]
Wizardcoder: Empowering code large language models with evol-instruct
Ziyang Luo, Can Xu, Pu Zhao, Qingfeng Sun, Xiubo Geng, Wenxiang Hu, Chongyang Tao, Jing Ma, Qingwei Lin, and Daxin Jiang. Wizardcoder: Empowering code large language models with evol-instruct. ICLR, 2024
2024
-
[42]
Simpo: Simple preference optimization with a reference-free reward
Yu Meng, Mengzhou Xia, and Danqi Chen. Simpo: Simple preference optimization with a reference-free reward. In NeurIPS, 2024
2024
-
[43]
Efficient estimation of word representations in vector space
Tomas Mikolov, Kai Chen, Greg Corrado, and Jeffrey Dean. Efficient estimation of word representations in vector space. In ICLR Workshop, 2013
2013
-
[44]
Disease-symptoms dataset
Mohamed-Ahmed161. Disease-symptoms dataset. https://huggingface.co/datasets/ Mohamed-Ahmed161/Disease-Symptoms , 2024. Hugging Face dataset
2024
-
[45]
Sgpt: Gpt sentence embeddings for semantic search
Niklas Muennighoff. Sgpt: Gpt sentence embeddings for semantic search. arXiv preprint arXiv:2202.08904, 2022
2022 arXiv
-
[46]
Routellm: Learning to route llms from preference data
Isaac Ong, Amjad Almahairi, Vincent Wu, Wei-Lin Chiang, Tianhao Wu, Joseph E Gonzalez, M Waleed Kadous, and Ion Stoica. Routellm: Learning to route llms from preference data. In ICLR, 2024
2024
-
[47]
Task arithmetic in the tangent space: Improved editing of pre-trained models
Guillermo Ortiz-Jimenez, Alessandro Favero, and Pascal Frossard. Task arithmetic in the tangent space: Improved editing of pre-trained models. In A. Oh, T. Naumann, A. Globerson, K. Saenko, M. Hardt, and S. Levine (eds.), NeurIPS, 2023
2023
-
[48]
Towards modular llms by building and reusing a library of loras
Oleksiy Ostapenko, Zhan Su, Edoardo Maria Ponti, Laurent Charlin, Nicolas Le Roux, Matheus Pereira, Lucas Caccia, and Alessandro Sordoni. Towards modular llms by building and reusing a library of loras. arXiv preprint arXiv:2405.11157, 2024
2024 arXiv
-
[49]
Medmcqa: A large-scale multi-subject multi-choice dataset for medical domain question answering
Ankit Pal, Logesh Kumar Umapathi, and Malaikannan Sankarasubbu. Medmcqa: A large-scale multi-subject multi-choice dataset for medical domain question answering. In Proceedings of the Conference on Health, Inference, and Learning, 2022
2022
-
[50]
Smaug: Fixing failure modes of preference optimisation with dpo-positive
Arka Pal, Deep Karkhanis, Samuel Dooley, Manley Roberts, Siddartha Naidu, and Colin White. Smaug: Fixing failure modes of preference optimisation with dpo-positive. arXiv preprint arXiv:2402.13228, 2024
2024 arXiv
-
[51]
Codeelo: Benchmarking competition-level code generation of llms with human-comparable elo ratings, 2025
Shanghaoran Quan, Jiaxi Yang, Bowen Yu, Bo Zheng, Dayiheng Liu, An Yang, Xuancheng Ren, Bofei Gao, Yibo Miao, Yunlong Feng, Zekun Wang, Jian Yang, Zeyu Cui, Yang Fan, Yichang Zhang, Binyuan Hui, and Junyang Lin. Codeelo: Benchmarking competition-level code generation of llms w...
2025
-
[52]
Direct preference optimization: Your language model is secretly a reward model
Rafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D Manning, Stefano Ermon, and Chelsea Finn. Direct preference optimization: Your language model is secretly a reward model. In NeurIPS, 2023
2023
-
[53]
Learning dynamics of llm finetuning
Yi Ren and Danica J Sutherland. Learning dynamics of llm finetuning. In ICLR, 2025
2025
-
[54]
Rousseeuw
Peter J. Rousseeuw. Silhouettes: A graphical aid to the interpretation and validation of cluster analysis. Journal of Computational and Applied Mathematics, 1987
1987
-
[55]
Large language model routing with benchmark datasets
Tal Shnitzer, Anthony Ou, Mírian Silva, Kate Soule, Yuekai Sun, Justin Solomon, Neil Thomp- son, and Mikhail Yurochkin. Large language model routing with benchmark datasets. arXiv preprint arXiv:2309.15789, 2023
2023 arXiv
-
[56]
A simple and effective pruning approach for large language models
Mingjie Sun, Zhuang Liu, Anna Bair, and J Zico Kolter. A simple and effective pruning approach for large language models. arXiv preprint arXiv:2306.11695, 2023
2023 arXiv
-
[57]
Zico Kolter, and Zhuang Liu
Mingjie Sun, Xinlei Chen, J. Zico Kolter, and Zhuang Liu. Massive activations in large language models. COLM, 2024
2024
-
[58]
Zico Kolter, and Zhuang Liu
Mingjie Sun, Yida Yin, Zhiqiu Xu, J. Zico Kolter, and Zhuang Liu. Idiosyncrasies in large language models. In ICML, 2025
2025
-
[59]
Challenging big- bench tasks and whether chain-of-thought can solve them
Mirac Suzgun, Nathan Scales, Nathanael Schärli, Sebastian Gehrmann, Yi Tay, Hyung Won Chung, Aakanksha Chowdhery, Quoc V Le, Ed H Chi, Denny Zhou, et al. Challenging big- bench tasks and whether chain-of-thought can solve them. arXiv preprint arXiv:2210.09261, 2022
-
[60]
Learnware of language models: Specialized small language models can do big
Zhi-Hao Tan, Zi-Chen Zhao, Hao-Yu Shi, Xin-Yu Zhang, Peng Tan, Yang Yu, and Zhi-Hua Zhou. Learnware of language models: Specialized small language models can do big. arXiv preprint arXiv:2505.13425, 2025
2025 arXiv
-
[61]
Alpaca: A strong, replicable instruction-following model
Rohan Taori, Ishaan Gulrajani, Tianyi Zhang, Yann Dubois, Xuechen Li, Carlos Guestrin, Percy Liang, and Tatsunori B Hashimoto. Alpaca: A strong, replicable instruction-following model. Stanford Center for Research on Foundation Models, 2023
2023
-
[62]
Gemma 2: Improving open language models at a practical size
Gemma Team, Morgane Riviere, Shreya Pathak, Pier Giuseppe Sessa, Cassidy Hardin, Surya Bhupatiraju, Léonard Hussenot, Thomas Mesnard, Bobak Shahriari, Alexandre Ramé, et al. Gemma 2: Improving open language models at a practical size. arXiv preprint arXiv:2408.00118, 2024
2024 arXiv
-
[63]
Contrastive representation distillation
Yonglong Tian, Dilip Krishnan, and Phillip Isola. Contrastive representation distillation. arXiv preprint arXiv:1910.10699, 2019
1910 arXiv
-
[64]
Llama: Open and efficient foundation language models
Hugo Touvron et al. Llama: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971, 2023
2023 arXiv
-
[65]
Neural network diffusion.arXiv preprint arXiv:2402.13144, 2024
Kai Wang, Dongwen Tang, Boya Zeng, Yida Yin, Zhaopan Xu, Yukun Zhou, Zelin Zang, Trevor Darrell, Zhuang Liu, and Yang You. Neural network diffusion.arXiv preprint arXiv:2402.13144, 2024
2024 arXiv
-
[66]
MiniLMv2: Multi-head self-attention relation distillation for compressing pretrained transformers
Wenhui Wang, Hangbo Bao, Shaohan Huang, Li Dong, and Furu Wei. MiniLMv2: Multi-head self-attention relation distillation for compressing pretrained transformers. In ACL, 2021
2021
-
[67]
Sheared LLaMA: Accelerating language model pre-training via structured pruning
Mengzhou Xia, Tianyu Gao, Zhiyuan Zeng, and Danqi Chen. Sheared LLaMA: Accelerating language model pre-training via structured pruning. In ICLR, 2024
2024
-
[68]
Initializing models with larger ones
Zhiqiu Xu, Yanjie Chen, Kirill Vishniakov, Yida Yin, Zhiqiang Shen, Trevor Darrell, Lingjie Liu, and Zhuang Liu. Initializing models with larger ones. In ICLR, 2024
2024
-
[69]
Qwen2.5 technical report.arXiv preprint arXiv:2412.15115, 2024
An Yang, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chengyuan Li, Dayiheng Liu, Fei Huang, Haoran Wei, Huan Lin, Jian Yang, Jianhong Tu, Jianwei Zhang, Jianxin Yang, Jiaxi Yang, Jingren Zhou, Junyang Lin, Kai Dang, Keming Lu, Keqin Bao, Kexin Yang, Le Yu, Me...
2024 arXiv
-
[70]
Model merging in llms, mllms, and beyond: Methods, theories, applications and opportu- nities
Enneng Yang, Li Shen, Guibing Guo, Xingwei Wang, Xiaochun Cao, Jie Zhang, and Dacheng Tao. Model merging in llms, mllms, and beyond: Methods, theories, applications and opportu- nities. arXiv preprint arXiv:2408.07666, 2024
2024 arXiv
-
[71]
Outlier weighed layerwise sparsity (owl): A missing secret sauce for pruning llms to high sparsity
Lu Yin, You Wu, Zhenyu Zhang, Cheng-Yu Hsieh, Yaqing Wang, Yiling Jia, Mykola Pech- enizkiy, Yi Liang, Zhangyang Wang, and Shiwei Liu. Outlier weighed layerwise sparsity (owl): A missing secret sauce for pruning llms to high sparsity. arXiv preprint arXiv:2310.05175, 2023
-
[72]
Metamath: Bootstrap your own mathematical questions for large language models
Longhui Yu, Weisen Jiang, Han Shi, Jincheng Yu, Zhengying Liu, Yu Zhang, James T Kwok, Zhenguo Li, Adrian Weller, and Weiyang Liu. Metamath: Bootstrap your own mathematical questions for large language models. arXiv preprint arXiv:2309.12284, 2023
2023 arXiv
-
[73]
Neural phylogeny: Fine-tuning relationship detection among neural networks
Runpeng Yu and Xinchao Wang. Neural phylogeny: Fine-tuning relationship detection among neural networks. In ICLR, 2025
2025
-
[74]
Hellaswag: Can a machine really finish your sentence? In ACL, 2019
Rowan Zellers, Ari Holtzman, Yonatan Bisk, Ali Farhadi, and Yejin Choi. Hellaswag: Can a machine really finish your sentence? In ACL, 2019
2019
-
[75]
Generative modeling of weights: General- ization or memorization? arXiv preprint arXiv:2506.07998, 2025
Boya Zeng, Yida Yin, Zhiqiu Xu, and Zhuang Liu. Generative modeling of weights: General- ization or memorization? arXiv preprint arXiv:2506.07998, 2025
2025
-
[76]
Dynamic sparse no training: Training-free fine-tuning for sparse llms
Yuxin Zhang, Lirui Zhao, Mingbao Lin, Yunyun Sun, Yiwu Yao, Xingjia Han, Jared Tanner, Shiwei Liu, and Rongrong Ji. Dynamic sparse no training: Training-free fine-tuning for sparse llms. arXiv preprint arXiv:2310.08915, 2023
2023 arXiv
-
[77]
LoraRetriever: Input-aware LoRA retrieval and composition for mixed tasks in the wild
Ziyu Zhao, Leilei Gan, Guoyin Wang, Wangchunshu Zhou, Hongxia Yang, Kun Kuang, and Fei Wu. LoraRetriever: Input-aware LoRA retrieval and composition for mixed tasks in the wild. In ACL Findings, 2024
2024
-
[78]
A comprehensive survey on transfer learning
Fuzhen Zhuang, Zhiyuan Qi, Keyu Duan, Dongbo Xi, Yongchun Zhu, Hengshu Zhu, Hui Xiong, and Qing He. A comprehensive survey on transfer learning. Proceedings of the IEEE, 2020
2020
-
[79]
meaning vector
Richard Zhuang, Tianhao Wu, Zhaojin Wen, Andrew Li, Jiantao Jiao, and Kannan Ramchandran. EmbedLLM: Learning compact representations of large language models. In ICLR, 2025. 15 A Training settings For LoRA, we set rankr = 8,α = 16, targeting query, key, value, and MLP projecti...
2025
-
[2024]
Preprint available at arXiv:2408.13656
Reviewed August 15, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.