REVIEW 5 major objections 6 minor 45 references
ProDS: Preference-oriented Data Selection for Instruction Tuning
T0 review · 5 major / 6 minor · reviewed 2026-08-15 · deepseek-v4-flash
Pith's one-line read Scoring instruction data by alignment with human-preference directions, as ProDS does, matches or exceeds full-data fine-tuning at 5–20% of the data.
desk verdict A plausible incremental twist on LESS—DPO preference gradients for instruction data selection—but the core gradient-similarity step is not justified as written, one equation is wrong, and the annealing is just a threshold, so it needs major revision before the mechanism can be trusted. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the bidirectional preference gradient. DPO warm-up turns human preference into a vector direction: for each validation pair, the DPO loss gradient points from a dispreferred response toward a preferred one, and the reversed pair points the other way. Candidate training samples are represented by randomly projected supervised-fine-tuning gradients, so that cosine similarity between a training gradient and a preference gradient measures how much a sample pushes the model in the preferred direction. The scoring formula combines these similarities through an instance-level weight optimized by simulated annealing, which is the mechanism that synthesizes positive and negative preference signals into a single rank.
What would settle it
Rank the same training pool with the ProDS score and with a deliberately reversed score that swaps positive and negative preferences, train on the top subsets of each, and compare on the target benchmark; if the reversed ranking performs as well, the bidirectional preference signal is not what drives the gains. A cleaner test would train each candidate example alone and correlate its measured effect on target-preference accuracy with its ProDS score; a near-zero correlation would falsify the claim that the score ranks samples by usefulness.
Extended reading notes
Core claim
Instruction data selection benefits from treating 'preferred response' as a direction in model-parameter space rather than as a fixed target answer. The paper's central claim is that training examples whose fine-tuning gradient aligns with the DPO gradient of positive preference pairs, and does not align with the gradient of negative pairs, are the examples that transfer preference-aligned behavior to the target task. The score is formed per sample as a weighted difference of cosine similarities to the two preference directions, with the instance-level weight tuned by simulated annealing. On benchmarks with open-ended or reasoning-rich responses, the selected subsets are reported to beat full-data fine-tuning; on multiple-choice tasks where preferences carry little information beyond the correct option, the advantage is smaller.
Load-bearing premise
The method's usefulness rests on one premise: the direction in which a training example would push the model during supervised fine-tuning is a trustworthy guide to how much that example improves preference-aligned behavior on the target task, so examples whose directions match the preference direction are the ones worth keeping.
Editorial extensions
If this is right
- A data pool can be scored once offline, and the same scores reused across different target tasks and target models, since the expensive gradient computation is a one-time pass.
- Small or architecturally different selection models can choose data for larger target models; the paper reports consistent selection-model transfer across several model families, which lowers the cost of running the method.
- Separating positive from negative preference directions is load-bearing: collapsing them into a single preference score degraded performance in the paper's ablation.
- The selection rule should be most valuable where responses have rich preference structure, such as dialogue and chain-of-thought reasoning, and least valuable on tasks where the correct answer is a single option.
Reading between the lines
- This suggests the preference signal could be obtained without a full DPO warm-up, for instance from an existing reward model's gradients or from a small set of human preference judgments, which would cut the method's extra training cost.
- A natural extension is to use the bidirectional score actively, re-scoring newly generated candidate responses as preferences accumulate, rather than only ranking a fixed pool.
- The paper's response-length analysis shows the method is not merely selecting short answers; isolating which semantic preferences drive selection would be a direct next experiment.
- If the gradient-overlap proxy holds across domains, the same scoring could apply to other alignment signals, such as safety or style preferences, by swapping the objective that defines the preference direction.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. ProDS (Preference-oriented Data Selection) is a method for instruction-tuning data selection. The paper proposes scoring training samples by the cosine similarity between their projected SFT-loss gradients and projected DPO-loss gradients computed on validation preference pairs, thereby incorporating response-preference signals into selection. The method has three stages: (1) SFT and DPO warm-up on small subsets; (2) construction of positive and negative validation preference sets (with GPT-4 as judge) and computation of projected gradients; (3) scoring of training samples via cosine similarity, with positive and negative preference scores combined by an instance-level weight optimized by simulated annealing. The top-k samples are then used to fine-tune the target LLM. Experiments on Alpaca (evaluated on Vicuna, Koala, WizardLM, Self-Instruct, and LIMA via GPT-4 pairwise comparison) and on FLAN/COT/DOLLY/OpenAssistant (evaluated on MMLU, TYDIQA, and BBH with external metrics) compare against BM25, DSIR, RDS, LESS, and IFD.
Significance. If the proposed scoring mechanism were principled, ProDS would advance instruction-data selection by introducing preference information and would be supported by a broad set of experiments and ablations, including selection-model transfer and preference-direction analysis. However, the central mechanism is undermined by (a) an internal inconsistency in Equation (5), (b) the lack of any justification for comparing gradients of different losses at different checkpoints, and (c) the degenerate linear annealing objective that reduces to a per-sample threshold rule. The Alpaca evaluation is further confounded by using the same judge (GPT-4) for both preference construction and outcome evaluation, while the benchmark results show small margins over LESS and no variance estimates. The contribution is thus currently not at the level of a journal publication, though the underlying idea is worth pursuing.
major comments (5)
- [Section 3.2, Eq. (5)] Equation (5) defines G_val = [G_app_val, G_app_val], duplicating the positive-preference block instead of concatenating G_app_val with G_awy_val. As written, the negative preference direction is never used in the scoring, which directly contradicts the stated bidirectional contribution. This must be corrected and the subsequent formulas verified.
- [Section 3.2-3.3, Eqs. (3)-(6)] The core score compares projected SFT-loss gradients at θ_S with projected DPO-loss gradients at θ_D_p, i.e., gradients of two different loss functions evaluated at two different parameter checkpoints. LESS, which is cited as the basis, compares gradients of the same loss at the same checkpoint, for which influence-theoretic interpretations apply. The paper provides no argument or experiment showing that the cosine similarity between a training SFT gradient and a validation DPO gradient estimates how much a training sample improves target-aligned behavior. Without such justification, the selection signal is not a principled influence estimate; an empirical validation (e.g., correlation with actual fine-tuning gains or leave-one-out influence) is required.
- [Section 3.3, Eq. (7) and Algorithm 1] The energy function E(Λ) is linear in each Λ_i, so its global optimum over Λ_i ∈ [0,1] is the boundary point Λ_i = 1 if Γapp_i + Γawy_i > 0, otherwise Λ_i = 0. Simulated annealing therefore implements a per-sample sign threshold and provides no additional modeling beyond that rule. This overstates the contribution of the 'annealing-based integrating method.' Moreover, Algorithm 1 does not specify bounds on Λ, leaving the objective potentially unbounded; the comparison against the 'fixed' baseline in Table 3 is not an ablation of the synthesis mechanism against the true boundary optimum.
- [Section 4.1.2-4.1.3, Fig. 4/5, Tables 3-4] For the Alpaca-related test sets, the same LLM judge (GPT-4) is used both to construct the validation preference pairs (by scoring responses from Mbase and Mcmp) and to compute the final pairwise win scores. Consequently, the reported improvements may reflect alignment with the judge's particular preferences rather than a generalizable improvement in instruction following. The MMLU, TYDIQA, and BBH results are not subject to this circularity, but their margins over LESS are small.
- [Section 4.2, Table 1 and Fig. 5] No repeated runs, error bars, or significance tests are reported. The differences over LESS are small (e.g., +0.6 MMLU, +0.4 TYDIQA, +1.2 or +1.6 BBH, depending on the baseline), and the claimed consistent advantage in Fig. 5 could be within run-to-run variance. The paper should report means and standard deviations over at least three seeds, and ideally a significance test, for the main comparisons.
minor comments (6)
- [Section 3.2, Eq. (3)] The word 'drowned' should be 'drawn' in the sentence describing the projection matrix.
- [Section 3.1] The statement 'Since M_D_r is fixed during the DPO warm-up process, θ_D_p is the same as θ_S' is incorrect; it is the reference model parameters θ_D_r that remain equal to θ_S after initialization, not the policy parameters θ_D_p.
- [Table 1] The '∆' column does not specify the baseline for the comparison, e.g., whether it is Ours(7B) minus LESS(7B) or another difference; please clarify.
- [Section 4.2, Fig. 4] The caption of Figure 4 is ambiguous about what the three numbers in each row represent and for which comparison direction the wins/ties/losses are counted; please clarify.
- [Section 3.3, Eq. (6)] The text says Γapp and Γawy are computed via a 'weighted summation using the L2 norm' but does not give an explicit formula; define these quantities precisely.
- [References] The reference to Har-Peled and Kushal (2005) is a coresets paper, not a standard K-Means clustering reference; cite an appropriate clustering source.
Circularity Check
No significant circularity: ProDS's gradient-similarity score is a constructed proxy, and the benchmark evaluations are not functions of the selection score.
full rationale
ProDS's derivation chain is self-contained rather than circular. The training-sample signal G_train (Eq. 3) is the projected SFT-loss gradient at θ_S, and the preference signal G_val (Eq. 5) is the projected DPO-loss gradient at θ_D_p; neither is defined in terms of the final evaluation scores, so the cosine-similarity score Γ is a constructed proxy, not a fitted prediction of the reported metrics. The annealing objective E(Λ) (Eq. 7) is optimized over the same similarity scores rather than over test-set outcomes; even if the linear form makes the optimum a threshold rule, that only reduces the annealing procedure to a sign test, not the experimental claim to its inputs. The use of GPT-4 both to create validation preference pairs and to judge Alpaca-style test outputs is a legitimate evaluation-validity concern because the judge is shared between selection and assessment, but the final pairwise evaluation compares responses of separately fine-tuned models on test instructions and is not equal by construction to the selection score; moreover, MMLU, TYDIQA, and BBH results use external accuracy, F1, and exact-match metrics and are independent of the GPT-4 preference labels. Cited prior work such as LESS and TagCOS supplies gradient-projection machinery, but the paper does not rely on a self-citation or a uniqueness theorem, and no parameter is fitted to the metric it then claims to predict. Accordingly, no circular step meeting the quoted-equation standard is present.
Assumptions & free parameters
free parameters (5)
- Lambda per-instance weights =
boundary values (0 or 1) selected by annealing
- Projection dimension d =
8192
- Warm-up data fraction =
5% SFT, 5% DPO
- Annealing schedule =
T0=1.0, alpha=0.95, Tend=0.01, sigma unspecified
- Validation set size =
10% or 2/3 shots per sub-task
assumptions (4)
- domain assumption Gradient inner products approximate training influence
- domain assumption DPO loss gradients encode target human preferences
- domain assumption GPT-4 scores approximate human preferences
- standard math Random Rademacher projection preserves cosine similarity
Cite this review
Pith. "Pith review of ProDS: Preference-oriented Data Selection for Instruction Tuning." pith.science (2026). https://pith.science/paper/WGK73HJA
@misc{pith2026250512754,
author = {Pith},
title = {Pith review of: ProDS: Preference-oriented Data Selection for Instruction Tuning},
year = {2026},
howpublished = {\url{https://pith.science/paper/WGK73HJA}},
note = {Machine review of arXiv:2505.12754}
}
read the original abstract
Instruction data selection aims to identify a high-quality subset from the training set that matches or exceeds the performance of the full dataset on target tasks. Existing methods focus on the instruction-to-response mapping, but neglect the human preference for diverse responses. In this paper, we propose Preference-oriented Data Selection method (ProDS) that scores training samples based on their alignment with preferences observed in the target set. Our key innovation lies in shifting the data selection criteria from merely estimating features for accurate response generation to explicitly aligning training samples with human preferences in target tasks. Specifically, direct preference optimization (DPO) is employed to estimate human preferences across diverse responses. Besides, a bidirectional preference synthesis strategy is designed to score training samples according to both positive preferences and negative preferences. Extensive experimental results demonstrate our superiority to existing task-agnostic and targeted methods.
Figures
Figures from the paper (4 more)
Reference graph
Works this paper leans on
-
[1]
Jinze Bai, Shuai Bai, Yunfei Chu, Zeyu Cui, Kai Dang, Xiaodong Deng, Yang Fan, Wenbin Ge, Yu Han, Fei Huang, Binyuan Hui, Luo Ji, Mei Li, Junyang Lin, Runji Lin, Dayiheng Liu, Gao Liu, Chengqiang Lu, Keming Lu, Jianxin Ma, Rui Men, Xingzhang Ren, Xuancheng Ren, Chuanqi Tan, Sinan Tan, Jianhong Tu, Peng Wang, Shijie Wang, Wei Wang, Shengguang Wu, Benfeng X...
arXiv 2023
-
[2]
Xiao Bi, Deli Chen, Guanting Chen, Shanhuang Chen, Damai Dai, Chengqi Deng, Honghui Ding, Kai Dong, Qiushi Du, Zhe Fu, et al. 2024. Deepseek llm: Scaling open-source language models with longtermism. arXiv preprint arXiv:2401.02954
arXiv 2024
-
[3]
Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert - Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litw...
2020
-
[4]
Mayee Chen, Nicholas Roberts, Kush Bhatia, Jue Wang, Ce Zhang, Frederic Sala, and Christopher R \'e . 2024. Skill-it! a data-driven skills framework for understanding and training language models. Advances in Neural Information Processing Systems, 36
2024
-
[5]
Wei-Lin Chiang, Zhuohan Li, Zi Lin, Ying Sheng, Zhanghao Wu, Hao Zhang, Lianmin Zheng, Siyuan Zhuang, Yonghao Zhuang, Joseph E Gonzalez, et al. 2023. Vicuna: An open-source chatbot impressing gpt-4 with 90\ See https://vicuna. lmsys. org (accessed 14 April 2023), 2(3):6
2023
-
[6]
Jonathan H Clark, Eunsol Choi, Michael Collins, Dan Garrette, Tom Kwiatkowski, Vitaly Nikolaev, and Jennimaria Palomaki. 2020. Tydi qa: A benchmark for information-seeking question answering in ty pologically di verse languages. Transactions of the Association for Computational Linguistics, 8:454--470
2020
-
[7]
Mike Conover, Matt Hayes, Ankit Mathur, Jianwei Xie, Jun Wan, Sam Shah, Ali Ghodsi, Patrick Wendell, Matei Zaharia, and Reynold Xin. 2023 a . Free dolly: Introducing the world’s first truly open instruction-tuned llm. Company Blog of Databricks
work page 2023
-
[8]
Mike Conover, Matt Hayes, Ankit Mathur, Jianwei Xie, Jun Wan, Sam Shah, Ali Ghodsi, Patrick Wendell, Matei Zaharia, and Reynold Xin. 2023 b . Free dolly: Introducing the world’s first truly open instruction-tuned llm. Company Blog of Databricks
work page 2023
Show all 45 references
-
[9]
Jacob Devlin, Ming - Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Hum...
2019
- [10]
-
[11]
Yuan Ge, Yilun Liu, Chi Hu, Weibin Meng, Shimin Tao, Xiaofeng Zhao, Mahong Xia, Zhang Li, Boxing Chen, Hao Yang, Bei Li, Tong Xiao, and JingBo Zhu. 2024. Clustering and ranking: Diversity-preserved instruction selection through expert-aligned quality estimation. In Proceedings...
2024
-
[12]
Aaron Grattafiori, Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Alex Vaughan, Amy Yang, Angela Fan, Anirudh Goyal, Anthony Hartshorn, Aobo Yang, Archi Mitra, and etc. 2024. https://arxiv.org/abs/2...
2024 arXiv
-
[13]
Xiaochuang Han, Daniel Simig, Todor Mihaylov, Yulia Tsvetkov, Asli Celikyilmaz, and Tianlu Wang. 2023. Understanding in-context learning via supportive pretraining data. In The 61st Annual Meeting Of The Association For Computational Linguistics
2023
-
[14]
Kazuaki Hanawa, Sho Yokoi, Satoshi Hara, and Kentaro Inui. 2020. Evaluation of similarity-based explanations. arXiv preprint arXiv:2006.04528
2020 arXiv
-
[15]
Sariel Har-Peled and Akash Kushal. 2005. Smaller coresets for k-median and k-means clustering. In Proceedings of the twenty-first annual symposium on Computational geometry, pages 126--134
2005
-
[16]
Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. 2020. Measuring massive multitask language understanding. arXiv preprint arXiv:2009.03300
2020 arXiv
-
[17]
Albert Q Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, et al. 2023. Mistral 7b. arXiv preprint arXiv:2310.06825
2023 arXiv
-
[18]
Scott Kirkpatrick, C Daniel Gelatt Jr, and Mario P Vecchi. 1983. Optimization by simulated annealing. science, 220(4598):671--680
1983
-
[19]
o pf, Yannic Kilcher, Dimitri von R \
Andreas K \"o pf, Yannic Kilcher, Dimitri von R \"u tte, Sotiris Anagnostidis, Zhi Rui Tam, Keith Stevens, Abdullah Barhoum, Duc Nguyen, Oliver Stanley, Rich \'a rd Nagyfi, et al. 2024. Openassistant conversations-democratizing large language model alignment. Advances in Neura...
2024
-
[20]
Ming Li, Yong Zhang, Shwai He, Zhitao Li, Hongyu Zhao, Jianzong Wang, Ning Cheng, and Tianyi Zhou. 2024 a . Superfiltering: Weak-to-strong data filtering for fast instruction-tuning. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Vo...
2024
-
[21]
Ming Li, Yong Zhang, Zhitao Li, Jiuhai Chen, Lichang Chen, Ning Cheng, Jianzong Wang, Tianyi Zhou, and Jing Xiao. 2024 b . From quantity to quality: Boosting LLM performance with self-guided data selection for instruction tuning. In Proceedings of the 2024 Conference of the No...
2024
-
[22]
Yunshui Li, Binyuan Hui, Xiaobo Xia, Jiaxi Yang, Min Yang, Lei Zhang, Shuzheng Si, Ling - Hao Chen, Junhao Liu, Tongliang Liu, Fei Huang, and Yongbin Li. 2024 c . One-shot learning as instruction data prospector for large language models. In Proceedings of the 62nd Annual Meet...
2024
-
[23]
Wong, Dongfang Li, Ziyi Wang, Baotian Hu, and Min Zhang
Liangxin Liu, Xuebo Liu, Derek F. Wong, Dongfang Li, Ziyi Wang, Baotian Hu, and Min Zhang. 2024 a . https://doi.org/10.48550/ARXIV.2402.16705 Selectit: Selective instruction tuning for large language models via uncertainty-aware self-reflection . CoRR, abs/2402.16705
-
[24]
Yilun Liu, Shimin Tao, Xiaofeng Zhao, Ming Zhu, Wenbing Ma, Junhao Zhu, Chang Su, Yutai Hou, Miao Zhang, Min Zhang, et al. 2024 b . Coachlm: Automatic instruction revisions improve the data quality in llm instruction tuning. In 2024 IEEE 40th International Conference on Data E...
2024
-
[25]
Shayne Longpre, Le Hou, Tu Vu, Albert Webson, Hyung Won Chung, Yi Tay, Denny Zhou, Quoc V Le, Barret Zoph, Jason Wei, et al. 2023. The flan collection: Designing data and methods for effective instruction tuning. In International Conference on Machine Learning, pages 22631--22...
2023
- [26]
-
[27]
Sung Min Park, Kristian Georgiev, Andrew Ilyas, Guillaume Leclerc, and Aleksander Madry. 2023. Trak: Attributing model behavior at scale. arXiv preprint arXiv:2303.14186
2023 arXiv
-
[28]
Manning, Stefano Ermon, and Chelsea Finn
Rafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning, Stefano Ermon, and Chelsea Finn. 2023. Direct preference optimization: Your language model is secretly a reward model. In Advances in Neural Information Processing Systems 36: Annual Conference on Neural In...
2023
-
[29]
Stephen Robertson, Hugo Zaragoza, et al. 2009. The probabilistic relevance framework: Bm25 and beyond. Foundations and Trends in Information Retrieval , 3(4):333--389
2009
-
[30]
Mirac Suzgun, Nathan Scales, Nathanael Sch \"a rli, Sebastian Gehrmann, Yi Tay, Hyung Won Chung, Aakanksha Chowdhery, Quoc Le, Ed Chi, Denny Zhou, et al. 2023. Challenging big-bench tasks and whether chain-of-thought can solve them. In Findings of the Association for Computati...
2023
-
[31]
Rohan Taori, Ishaan Gulrajani, Tianyi Zhang, Yann Dubois, Xuechen Li, Carlos Guestrin, Percy Liang, and Tatsunori B Hashimoto. 2023. Stanford alpaca: an instruction-following llama model (2023). URL https://github. com/tatsu-lab/stanford\_alpaca, 1(9)
2023
-
[32]
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timoth \'e e Lacroix, Baptiste Rozi \`e re, Naman Goyal, Eric Hambro, Faisal Azhar, et al. 2023 a . Llama: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971
2023 arXiv
-
[33]
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. 2023 b . Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288
2023 arXiv
-
[34]
Thuy-Trang Vu, Xuanli He, Gholamreza Haffari, and Ehsan Shareghi. 2023. Koala: An index for quantifying overlaps with pre-training corpora. arXiv preprint arXiv:2303.14770
2023 arXiv
-
[35]
Yizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu, Noah A Smith, Daniel Khashabi, and Hannaneh Hajishirzi. 2023. Self-instruct: Aligning language models with self-generated instructions. In The 61st Annual Meeting Of The Association For Computational Linguistics
2023
-
[36]
Yizhong Wang, Swaroop Mishra, Pegah Alipoormolabashi, Yeganeh Kordi, Amirreza Mirzaei, Atharva Naik, Arjun Ashok, Arut Selvan Dhanasekaran, Anjana Arunkumar, David Stap, Eshaan Pathak, Giannis Karamanolakis, Haizhi Lai, Ishan Purohit, Ishani Mondal, Jacob Anderson, Kirby Kuzni...
2022
-
[37]
Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, Andrew M
Jason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, Andrew M. Dai, and Quoc V. Le. 2022 a . Finetuned language models are zero-shot learners. In The Tenth International Conference on Learning Representations, ICLR 2022, Virtual Event, Apr...
2022
-
[38]
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al. 2022 b . Chain-of-thought prompting elicits reasoning in large language models. Advances in neural information processing systems, 35:24824--24837
2022
-
[39]
Mengzhou Xia, Sadhika Malladi, Suchin Gururangan, Sanjeev Arora, and Danqi Chen. 2024. LESS: selecting influential data for targeted instruction tuning. In Forty-first International Conference on Machine Learning, ICML 2024, Vienna, Austria, July 21-27, 2024 . OpenReview.net
2024
-
[40]
Sang Michael Xie, Shibani Santurkar, Tengyu Ma, and Percy S Liang. 2023. Data selection for language models via importance resampling. Advances in Neural Information Processing Systems, 36:34201--34227
2023
-
[41]
Can Xu, Qingfeng Sun, Kai Zheng, Xiubo Geng, Pu Zhao, Jiazhan Feng, Chongyang Tao, and Daxin Jiang. 2023. Wizardlm: Empowering large language models to follow complex instructions. arXiv preprint arXiv:2304.12244
2023 arXiv
-
[42]
Jipeng Zhang, Yaxuan Qin, Renjie Pi, Weizhong Zhang, Rui Pan, and Tong Zhang. 2024. Tagcos: Task-agnostic gradient clustered coreset selection for instruction tuning data. arXiv preprint arXiv:2407.15235
2024 arXiv
-
[43]
Chunting Zhou, Pengfei Liu, Puxin Xu, Srinivasan Iyer, Jiao Sun, Yuning Mao, Xuezhe Ma, Avia Efrat, Ping Yu, Lili Yu, et al. 2024. Lima: Less is more for alignment. Advances in Neural Information Processing Systems, 36
2024
-
[44]
online" 'onlinestring :=
ENTRY address archivePrefix author booktitle chapter edition editor eid eprint eprinttype howpublished institution journal key month note number organization pages publisher school series title type volume year doi pubmed url lastchecked label extra.label sort.label short.list...
-
[45]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 gl...
Reviewed August 15, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.