REVIEW 4 major objections 5 minor 90 references
RICo: Refined In-Context Contribution for Automatic Instruction-Tuning Data Selection
T0 review · 4 major / 5 minor · reviewed 2026-08-15 · deepseek-v4-flash
Pith's one-line read RICo ranks instruction samples by in-context contribution, and 15% of selected Alpaca data outperforms the full dataset on LLaMA3.1-8B.
desk verdict Solid, genuinely new data-selection method with a plausible scoring mechanism, but the headline gains are likely optimistic because the 15% ratio is chosen post hoc on the same benchmarks and there are no variance estimates. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the global-RICo score: the average, over a diverse assessment set $\mathcal{D}_a$, of task-level scores that measure how much a candidate training sample $T_j$ reduces the model's perplexity on assessment sample $S_i$ when $T_j$ is inserted into the context. Each task-level score is $\frac{\mathrm{PPL}_\theta(S_i \mid T_j^{\mathrm{rand}}) - \mathrm{PPL}_\theta(S_i \mid T_j)}{\mathrm{PPL}_\theta(S_i) + \epsilon}$, where $T_j^{\mathrm{rand}}$ is a randomly generated sequence of the same length as $T_j$; subtracting this baseline is the fairness adjustment that removes length bias. The global score assigns equal weight to every assessment task, so a sample ranks high only if it helps broadly rather than specializing in one task type. A second component, a LoRA-trained classifier trained on top-scoring samples, reduces the cost of selecting from a large pool from $O(nm)$ to $O(m)$ inference calls.
What would settle it
Take a fixed set of candidate samples and measure, for each, the actual change in assessment-set perplexity after finetuning on that sample alone. If those measured changes do not rank-order samples the way RICo scores do, or if top-RICo subsets fail to beat full-data training on new benchmarks, the in-context proxy is not carrying the central claim.
Extended reading notes
Core claim
RICo's central claim is that a fine-grained, length-corrected perplexity difference measured in-context is a reliable estimate of a sample's contribution to gradient-based instruction tuning. For a training candidate $T_j$, the task-level score on assessment sample $S_i$ is $\frac{\mathrm{PPL}_\theta(S_i \mid T_j^{\mathrm{rand}}) - \mathrm{PPL}_\theta(S_i \mid T_j)}{\mathrm{PPL}_\theta(S_i) + \epsilon}$, and the global score averages this over a diverse assessment set. The paper argues that a higher score means greater contribution and a negative score means harm, and that selecting the top fraction by this score finds samples that are diverse and moderately difficult rather than merely hardest. The empirical assertion is that models trained on such subsets match or beat full-data models: on LLaMA3.1-8B, 15% of RICo-selected Alpaca data gives 43.37 average benchmark score versus 37.95 for full-data training, and exceeds the best prior selection method by 2.06 points.
Load-bearing premise
The central assumption is that a training sample's effect on the model when placed in the prompt at inference time predicts its effect when the model is actually trained on that sample with gradient updates.
Editorial extensions
If this is right
- Training on 15% of RICo-selected Alpaca data outperforms full-data training on LLaMA3.1-8B by 5.42 points and on Qwen2.5-3B by 1.24 points, while using roughly one-fifth to one-hundredth of the FLOPs.
- RICo outperforms established selection methods, including random, low and high perplexity, Alpagasus, Deita, Superfilter, and Nuggets, in both benchmark averages and GPT-4 pairwise comparisons.
- The selection paradigm transfers across datasets: a selector trained on Alpaca scores directly filters WizardLM and improves average scores by up to 4.54 points at 5% selected data.
- High-contribution samples tend to be diverse and of moderate difficulty, with extreme-difficulty samples avoided, suggesting an optimal data scale around 15% in this setting.
Reading between the lines
- If the RICo score truly approximates finetuning influence, it could serve as an attribution tool: the same assessment set could localize which training samples drive failures on specific benchmarks, not just rank a corpus.
- The length-bias correction is a separable component; a controlled ablation that removes only the random-sequence baseline would show how much of the gain over prior perplexity-based methods comes from fairness adjustment rather than from fine-grained task averaging.
- Because the selection model is trained on Alpaca and transfers to WizardLM, the approach may scale to much larger candidate pools, but the LoRA classifier's calibration on out-of-distribution instructions would need to be checked before trusting rankings in production-scale corpora.
- The reported optimal scale at 15% suggests that comparisons among selection methods should be made at each method's best scale; a fixed 15% comparison may favor methods whose optimum happens to fall there.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes RICo, a gradient-free method for instruction-tuning data selection. RICo scores each candidate training sample by the change in perplexity it induces when inserted into the context of an assessment set, normalized by a length-controlled random baseline and task difficulty (Eqs. 1-5). A lightweight selection model is then trained with LoRA on the RICo labels so that the deployed selection requires one forward pass per candidate. Experiments on LLaMA3.1-8B, Qwen2.5-3B, and LLaMA2-7B, using Alpaca and WizardLM training data and 12 benchmarks plus 5 pairwise evaluation sets, support the headline claims: models trained on 15% of RICo-selected Alpaca data improve LLaMA3.1-8B's average benchmark score by 5.42 points over full-data training and by 2.06 points over the best prior selection method.
Significance. If the results hold, RICo is a practically valuable contribution: it avoids gradients, provides a fairness adjustment for length bias, reports consistent gains across three model sizes, includes a cross-dataset transfer experiment, and reduces training FLOPs by an order of magnitude. The design is sensible, and the evaluation is broader than in many prior data-selection papers. The main quantitative claims, however, are weakened by two protocol issues: the 15% ratio is chosen after observing the peak of the accuracy-versus-scale curve on the same evaluation benchmarks, and all comparisons rest on single runs without variance estimates. These issues affect the magnitude and reliability of the headline numbers rather than the basic plausibility of the method.
major comments (4)
- [Section 6; Tables 1, 2, 6] The 15% selection ratio used for the headline comparison is chosen after inspecting the same evaluation benchmarks on which the gains are reported. Section 6 states that the average score 'generally rises and then declines, peaking at 15%' and that 'based on this observation, we select the model trained on 15% of RICo-selected data as the representative model for comparisons.' The +5.42-point gain over FULL (Table 1) and the +2.06-point gain over the best baseline (Table 2) are therefore the maximum of the scale curve in Table 6, and all baselines are evaluated only at this empirically chosen optimum. This makes the headline magnitude an optimistic upper bound and biases the comparison against baselines whose optimal scale may differ. Please add a validation split (e.g., hold out several benchmarks for choosing the ratio), pre-register the ratio, or report the full scale curve for all baselines; at minimum, the paper should state explicitly that the reported gain is the best of ten tested scales.
- [Section 5.1 and Table 1] All reported numbers are single runs with no variance estimates. The central comparison (RICo 15% vs FULL on LLaMA3.1-8B, 43.37 vs 37.95) and the 2.06-point advantage over the best baseline are differences between individual runs. For Qwen2.5-3B the gain is 1.24 points, and for LLaMA2-7B the 15% run is 0.09 points below FULL while the 1% run is +0.65 points, so results are unstable across models and scales. Please report standard deviations over at least three seeds or bootstrap confidence intervals over benchmarks, and discuss whether the headline differences are within noise.
- [Section 3.3] The abstract and Section 3.3 describe 'strictly linear inference complexity,' but this applies only to applying the already-trained selection model. Computing the global-RICo labels requires O(n x m) forward passes over n assessment samples and m candidates, and the paper does not state how many samples are used to train the selection model, what top-K fraction is used for labeling, or how the classifier's training-set distribution relates to the full candidate pool. Please report these details and clarify that the linear-complexity claim covers only the deployment phase; otherwise the scalability claim is incomplete.
- [Limitations and Eq. (4)] The Limitations section concedes that in-context learning 'does not replicate the exact dynamics of full batch training with gradient updates,' yet Eq. (4) assumes the perplexity delta caused by inserting a candidate into the context is a measure of that candidate's contribution to instruction tuning. This proxy is load-bearing for the claim that RICo 'accurately identifies high-contribution data.' I am not asking for a theoretical proof, but the paper should provide a direct empirical check on at least one model/dataset: correlate the RICo score ranking with influence estimates from leave-one-out retraining or a gradient-based method on a small candidate pool, or show that top-RICo subsets outperform random subsets across several independent assessment sets. Without such a check, the mechanism claim is supported only indirectly by the end-task results.
minor comments (5)
- [Introduction] The word 'gredient-free' should be 'gradient-free', and the method name is typeset as 'RIC O' throughout; please unify the notation, e.g., 'RICo'.
- [Section 4.3.2] The sentence 'These sets contain 218, 252, 80 and 300 human-curated instruction data' gives only four numbers for five test sets; please map each number to its named test set (Vicuna, WizardLM, LIMA, SInstruct, Koala).
- [Section 5.2] The text preceding Table 2 says the table reports 13 evaluation benchmarks, but Table 2 (and Table 5) contain 12 benchmarks; please correct the count.
- [Section 6 and Figure 3] The description 'generally improves and then declines' is not supported by Table 6, where 1% (42.91) is higher than 5% (42.09) and 10% (42.08); the curve is non-monotonic before peaking at 15%, so please describe it more precisely.
- [References] The t-test citation to Kendall (1937) is inaccurate for a Student t-test; please cite the appropriate statistical source or remove the citation.
Circularity Check
No circular derivation: RICo scores are computed from a separate assessment set and validated on held-out benchmarks; the headline 15% ratio is a post-hoc evaluation choice, not a construction-level circularity.
full rationale
The derivation chain is self-contained. The RICo score (Eq. 4) is defined as (PPL(S_i|T_rand_j) - PPL(S_i|T_j))/(PPL(S_i)+epsilon) on an assessment set D_a that is disjoint from training data and downstream benchmarks (Section 3.1), and the global score (Eq. 5) is a plain average over D_a. The selection model (Section 3.3) is trained with LoRA on these RICo labels, not on the 12 benchmark scores. The main comparisons (Tables 1, 2, 5) therefore evaluate the method on external targets rather than fitting to them; no equation reduces benchmark performance to the RICo score by construction. The Limitations section candidly states that ICL 'does not replicate the exact dynamics of full batch training with gradient updates', which is a modeling assumption about transfer, not a circular step. The only evaluative concern is that Section 6 chooses the reporting scale after seeing the benchmark curve ('Based on this observation, we select the model trained on 15%...'), so the headline 5.42-point and 2.06-point margins are post-hoc maxima rather than pre-specified predictions; this is a selection/evaluation bias, not a circular derivation. Minor self-citations for the ICL premise (Dong et al. 2022; Dai et al. 2023; Zheng et al. 2023) are motivational and not load-bearing to the score computation, and no uniqueness claim or fitted parameter is used to force the outcome.
Assumptions & free parameters
free parameters (3)
- selection ratio K for headline result =
15% for LLaMA3.1-8B and Qwen2.5-3B; 1% for LLaMA2-7B
- epsilon in task-RICo denominator =
unspecified small constant
- assessment set size =
1,020
assumptions (5)
- domain assumption In-context conditioning approximates the contribution of the same sample under gradient finetuning.
- domain assumption Perplexity is a valid proxy for instruction-following and downstream benchmark performance.
- domain assumption The 1,020-sample assessment set is representative of overall capability and disjoint from training and evaluation data.
- domain assumption A same-length random token sequence is a neutral baseline that removes length bias without other distortions.
- domain assumption A classifier trained on RICo labels from a sampled subset can rank the entire candidate pool.
Cite this review
Pith. "Pith review of RICo: Refined In-Context Contribution for Automatic Instruction-Tuning Data Selection." pith.science (2026). https://pith.science/paper/I5FD5HZ2
@misc{pith2026250505327,
author = {Pith},
title = {Pith review of: RICo: Refined In-Context Contribution for Automatic Instruction-Tuning Data Selection},
year = {2026},
howpublished = {\url{https://pith.science/paper/I5FD5HZ2}},
note = {Machine review of arXiv:2505.05327}
}
read the original abstract
Data selection for instruction tuning is crucial for improving the performance of large language models (LLMs) while reducing training costs. In this paper, we propose Refined Contribution Measurement with In-Context Learning (RICo), a novel gradient-free method that quantifies the fine-grained contribution of individual samples to both task-level and global-level model performance. RICo enables more accurate identification of high-contribution data, leading to better instruction tuning. We further introduce a lightweight selection paradigm trained on RICo scores, enabling scalable data selection with a strictly linear inference complexity. Extensive experiments on three LLMs across 12 benchmarks and 5 pairwise evaluation sets demonstrate the effectiveness of RICo. Remarkably, on LLaMA3.1-8B, models trained on 15% of RICo-selected data outperform full datasets by 5.42% points and exceed the best performance of widely used selection methods by 2.06% points. We further analyze high-contribution samples selected by RICo, which show both diverse tasks and appropriate difficulty levels, rather than just the hardest ones.
Figures
Figures from the paper (2 more)
Reference graph
Works this paper leans on
-
[1]
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al. 2023. Gpt-4 technical report. arXiv preprint arXiv:2303.08774
arXiv 2023
-
[2]
Alon Albalak, Yanai Elazar, Sang Michael Xie, Shayne Longpre, Nathan Lambert, Xinyi Wang, Niklas Muennighoff, Bairu Hou, Liangming Pan, Haewon Jeong, et al. 2024. A survey on data selection for language models. arXiv preprint arXiv:2402.16827
arXiv 2024
-
[3]
Alexander Bukharin and Tuo Zhao. 2023. https://api.semanticscholar.org/CorpusID:265456564 Data diversity matters for robust instruction tuning . ArXiv, abs/2311.14736
arXiv 2023
-
[4]
Yihan Cao, Yanbin Kang, and Lichao Sun. 2023. https://api.semanticscholar.org/CorpusID:259837472 Instruction mining: High-quality instruction data selection for large language models . ArXiv, abs/2307.06290
arXiv 2023
-
[5]
Ernie Chang, Pin-Jie Lin, Yang Li, Changsheng Zhao, Daeil Kim, Rastislav Rabatin, Zechun Liu, Yangyang Shi, and Vikas Chandra. 2024. https://api.semanticscholar.org/CorpusID:272828139 Target-aware language modeling via granular data sampling . In Conference on Empirical Methods in Natural Language Processing
2024
-
[6]
Lichang Chen, SHIYANG LI, Jun Yan, Hai Wang, Kalpa Gunaratna, Vikas Yadav, Zheng Tang, Vijay Srinivasan, Tianyi Zhou, Heng Huang, and Hongxia Jin. 2023. https://api.semanticscholar.org/CorpusID:259937133 Alpagasus: Training a better alpaca with fewer data . ArXiv, abs/2307.08701
arXiv 2023
-
[7]
Wei-Lin Chiang, Zhuohan Li, Ziqing Lin, Ying Sheng, Zhanghao Wu, Hao Zhang, Lianmin Zheng, Siyuan Zhuang, Yonghao Zhuang, Joseph E Gonzalez, et al. 2023. Vicuna: An open-source chatbot impressing gpt-4 with 90\ See https://vicuna. lmsys. org (accessed 14 April 2023), 2(3):6
2023
-
[8]
Hyung Won Chung, Le Hou, Shayne Longpre, Barret Zoph, Yi Tay, William Fedus, Yunxuan Li, Xuezhi Wang, Mostafa Dehghani, Siddhartha Brahma, et al. 2024. Scaling instruction-finetuned language models. Journal of Machine Learning Research, 25(70):1--53
2024
Show all 90 references
-
[9]
Peter Clark, Isaac Cowhey, Oren Etzioni, Tushar Khot, Ashish Sabharwal, Carissa Schoenick, and Oyvind Tafjord. 2018. Think you have solved question answering? try arc, the ai2 reasoning challenge. arXiv preprint arXiv:1803.05457
2018 arXiv
-
[10]
Ganqu Cui, Lifan Yuan, Ning Ding, Guanming Yao, Wei Zhu, Yuan Ni, Guotong Xie, Zhiyuan Liu, and Maosong Sun. 2023. Ultrafeedback: Boosting language models with high-quality feedback
2023
-
[11]
Damai Dai, Yutao Sun, Li Dong, Yaru Hao, Shuming Ma, Zhifang Sui, and Furu Wei. 2023. Why can gpt learn in-context? language models secretly perform gradient descent as meta-optimizers. In Findings of the Association for Computational Linguistics: ACL 2023, pages 4005--4019
2023
-
[12]
Qingxiu Dong, Li Dong, Ke Xu, Guangyan Zhou, Yaru Hao, Zhifang Sui, and Furu Wei. 2023. Large language model for science: a study on p vs. np. arXiv preprint arXiv:2309.05689
2023 arXiv
-
[13]
Qingxiu Dong, Lei Li, Damai Dai, Ce Zheng, Jingyuan Ma, Rui Li, Heming Xia, Jingjing Xu, Zhiyong Wu, Tianyu Liu, et al. 2022. A survey on in-context learning. arXiv preprint arXiv:2301.00234
2022 arXiv
-
[14]
Qianlong Du, Chengqing Zong, and Jiajun Zhang. 2023. https://api.semanticscholar.org/CorpusID:265457248 Mods: Model-oriented data selection for instruction tuning . ArXiv, abs/2311.15653
2023 arXiv
-
[15]
Yann Dubois, Chen Xuechen Li, Rohan Taori, Tianyi Zhang, Ishaan Gulrajani, Jimmy Ba, Carlos Guestrin, Percy S Liang, and Tatsunori B Hashimoto. 2023. Alpacafarm: A simulation framework for methods that learn from human feedback. Advances in Neural Information Processing System...
2023
-
[16]
Yuan Ge, Yilun Liu, Chi Hu, Weibin Meng, Shimin Tao, Xiaofeng Zhao, Hongxia Ma, Li Zhang, Boxing Chen, Hao Yang, et al. 2024. Clustering and ranking: Diversity-preserved instruction selection through expert-aligned quality estimation. arXiv preprint arXiv:2402.18191
2024 arXiv
-
[17]
David Grangier and Dan Iter. 2022. The trade-offs of domain adaptation for neural language models. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 3802--3813
2022
-
[18]
Aaron Grattafiori, Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Alex Vaughan, et al. 2024. The llama 3 herd of models. arXiv preprint arXiv:2407.21783
2024 arXiv
-
[19]
Daya Guo, Dejian Yang, Haowei Zhang, Junxiao Song, Ruoyu Zhang, Runxin Xu, Qihao Zhu, Shirong Ma, Peiyi Wang, Xiao Bi, et al. 2025. Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning. arXiv preprint arXiv:2501.12948
2025 arXiv
-
[20]
Jindong Han, Hao Liu, Jun Fang, Naiqiang Tan, and Hui Xiong. 2025. Automatic instruction data selection for large language models via uncertainty-aware influence maximization. In Proceedings of the ACM on Web Conference 2025, pages 4969--4979
2025
-
[21]
Yexiao He, Ziyao Wang, Zheyu Shen, Guoheng Sun, Yucong Dai, Yongkai Wu, Hongyi Wang, and Ang Li. 2024. Shed: Shapley-based automated dataset refinement for instruction fine-tuning. arXiv preprint arXiv:2405.00705
2024 arXiv
-
[22]
Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. 2020. Measuring massive multitask language understanding. arXiv preprint arXiv:2009.03300
2020 arXiv
-
[23]
Edward Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, and Weizhu Chen
J. Edward Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, and Weizhu Chen. 2021. https://api.semanticscholar.org/CorpusID:235458009 Lora: Low-rank adaptation of large language models . ArXiv, abs/2106.09685
2021 arXiv
-
[24]
Aaron Hurst, Adam Lerer, Adam P Goucher, Adam Perelman, Aditya Ramesh, Aidan Clark, AJ Ostrow, Akila Welihinda, Alan Hayes, Alec Radford, et al. 2024. Gpt-4o system card. arXiv preprint arXiv:2410.21276
2024 arXiv
-
[25]
Hamish Ivison, Noah A Smith, Hannaneh Hajishirzi, and Pradeep Dasigi. 2023. Data-efficient finetuning using cross-task nearest neighbors. In Findings of the Association for Computational Linguistics: ACL 2023, pages 9036--9061
2023
-
[26]
Zhengbao Jiang, Zhiqing Sun, Weijia Shi, Pedro Rodriguez, Chunting Zhou, Graham Neubig, Xi Lin, Wen-tau Yih, and Srini Iyer. 2024. https://doi.org/10.18653/v1/2024.acl-long.296 Instruction-tuned language models are better knowledge learners . In Proceedings of the 62nd Annual ...
2024 doi
-
[27]
Cathy Jiao, Weizhen Gao, Aditi Raghunathan, and Chenyan Xiong. 2025. On the feasibility of in-context probing for data attribution. In Findings of the Association for Computational Linguistics: NAACL 2025, pages 5140--5155
2025
-
[28]
M. G. Kendall. 1937. https://api.semanticscholar.org/CorpusID:4018463 Statistical methods for research workers . Nature, 139:737--737
1937
-
[29]
Diederik P Kingma and Jimmy Ba. 2014. Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980
2014 arXiv
-
[30]
o pf, Yannic Kilcher, Dimitri Von R \
Andreas K \"o pf, Yannic Kilcher, Dimitri Von R \"u tte, Sotiris Anagnostidis, Zhi Rui Tam, Keith Stevens, Abdullah Barhoum, Duc Nguyen, Oliver Stanley, Rich \'a rd Nagyfi, et al. 2023. Openassistant conversations-democratizing large language model alignment. Advances in Neura...
2023
-
[31]
Po-Nien Kung, Fan Yin, Di Wu, Kai-Wei Chang, and Nanyun Peng. 2023. https://api.semanticscholar.org/CorpusID:264832712 Active instruction tuning: Improving cross-task generalization by training on prompt sensitive tasks . In Conference on Empirical Methods in Natural Language ...
2023
-
[32]
Ming Li, Lichang Chen, Jiuhai Chen, Shwai He, Jiuxiang Gu, and Tianyi Zhou. 2024 a . https://api.semanticscholar.org/CorpusID:267682220 Selective reflection-tuning: Student-selected data recycling for llm instruction-tuning . ArXiv, abs/2402.10110
2024 arXiv
-
[33]
Ming Li, Yong Zhang, Shwai He, Zhitao Li, Hongyu Zhao, Jianzong Wang, Ning Cheng, and Tianyi Zhou. 2024 b . https://api.semanticscholar.org/CorpusID:267365346 Superfiltering: Weak-to-strong data filtering for fast instruction-tuning . In Annual Meeting of the Association for C...
2024
-
[34]
Ming Li, Yong Zhang, Zhitao Li, Jiuhai Chen, Lichang Chen, Ning Cheng, Jianzong Wang, Tianyi Zhou, and Jing Xiao. 2023 a . https://api.semanticscholar.org/CorpusID:261076515 From quantity to quality: Boosting llm performance with self-guided data selection for instruction tuni...
2023 arXiv
-
[35]
Hashimoto
Xuechen Li, Tianyi Zhang, Yann Dubois, Rohan Taori, Ishaan Gulrajani, Carlos Guestrin, Percy Liang, and Tatsunori B. Hashimoto. 2023 b . Alpacaeval: An automatic evaluator of instruction-following models. https://github.com/tatsu-lab/alpaca_eval
2023
-
[36]
Yunshui Li, Binyuan Hui, Xiaobo Xia, Jiaxi Yang, Min Yang, Lei Zhang, Shuzheng Si, Ling-Hao Chen, Junhao Liu, Tongliang Liu, et al. 2023 c . One-shot learning as instruction data prospector for large language models. arXiv preprint arXiv:2312.10302
2023 arXiv
-
[37]
Wing Lian, Bleys Goodson, Eugene Pentland, Austin Cook, Chanvichet Vong, and “Teknium”. 2023. Openorca: An open dataset of gpt augmented flan reasoning traces
2023
-
[38]
Stephanie Lin, Jacob Hilton, and Owain Evans. 2021. Truthfulqa: Measuring how models mimic human falsehoods. arXiv preprint arXiv:2109.07958
2021 arXiv
-
[40]
Fuxiao Liu, Xiaoyang Wang, Wenlin Yao, Jianshu Chen, Kaiqiang Song, Sangwoo Cho, Yaser Yacoob, and Dong Yu. 2023 b . https://api.semanticscholar.org/CorpusID:265294419 Mmc: Advancing multimodal chart understanding with large-scale instruction tuning . ArXiv, abs/2311.10774
2023 arXiv
-
[41]
Jian Liu, Leyang Cui, Hanmeng Liu, Dandan Huang, Yile Wang, and Yue Zhang. 2020. Logiqa: A challenge dataset for machine reading comprehension with logical reasoning. arXiv preprint arXiv:2007.08124
2020 arXiv
-
[42]
Wong, Dongfang Li, Ziyi Wang, Baotian Hu, and Min Zhang
Liangxin Liu, Xuebo Liu, Derek F. Wong, Dongfang Li, Ziyi Wang, Baotian Hu, and Min Zhang. 2024. https://api.semanticscholar.org/CorpusID:268032188 Selectit: Selective instruction tuning for llms via uncertainty-aware self-reflection . In Neural Information Processing Systems
2024
-
[43]
Wei Liu, Weihao Zeng, Keqing He, Yong Jiang, and Junxian He. 2023 c . https://api.semanticscholar.org/CorpusID:266551413 What makes good data for alignment? a comprehensive study of automatic data selection in instruction tuning . ArXiv, abs/2312.15685
2023 arXiv
-
[44]
Shayne Longpre, Gregory Yauney, Emily Reif, Katherine Lee, Adam Roberts, Barret Zoph, Denny Zhou, Jason Wei, Kevin Robinson, David Mimno, et al. 2024. A pretrainer’s guide to training data: Measuring the effects of data age, domain coverage, quality, & toxicity. In Proceedings...
2024
-
[45]
Yao Lu, Max Bartolo, Alastair Moore, Sebastian Riedel, and Pontus Stenetorp. 2021. Fantastically ordered prompts and where to find them: Overcoming few-shot prompt order sensitivity. arXiv preprint arXiv:2104.08786
2021 arXiv
-
[46]
Man Luo, Xin Xu, Yue Liu, Panupong Pasupat, and Mehran Kazemi. 2024. In-context learning with retrieved demonstrations for language models: A survey. arXiv preprint arXiv:2401.11624
2024 arXiv
-
[47]
Kaifeng Lyu, Haoyu Zhao, Xinran Gu, Dingli Yu, Anirudh Goyal, and Sanjeev Arora. 2024. Keeping llms aligned after fine-tuning: The crucial role of prompt templates. arXiv preprint arXiv:2402.18540
2024 arXiv
-
[48]
Conover Mike, Hayes Matt, Mathur Ankit, Xie Jianwei, Wan Jun, Shah Sam, Ghodsi Ali, Wendell Patrick, Zaharia Matei, and Xin Reynold. 2023. Free dolly: Introducing the world’s first truly open instruction-tuned llm
2023
-
[49]
Wu Minghao, Thuy-Trang Vu, Lizhen Qu, and Gholamreza Haffari. 2024. The best of both worlds: Bridging quality and diversity in data selection with bipartite graph. arXiv preprint arXiv:2410.12458
2024 arXiv
-
[50]
Swaroop Mishra, Daniel Khashabi, Chitta Baral, and Hannaneh Hajishirzi. 2021. Natural instructions: Benchmarking generalization to new tasks from natural language instructions. arXiv preprint arXiv:2104.08773, pages 839--849
2021 arXiv
-
[51]
Kevin P Murphy. 2012. Machine learning: a probabilistic perspective. MIT press
2012
-
[52]
Chatterji, Faisal Ladhak, and Tatsunori Hashimoto
Yonatan Oren, Nicole Meister, Niladri S. Chatterji, Faisal Ladhak, and Tatsunori Hashimoto. 2024. https://openreview.net/forum?id=KS8mIvetg2 Proving test set contamination in black-box language models . In The Twelfth International Conference on Learning Representations
2024
-
[53]
Aliz \'e e Pace, Jonathan Mallinson, Eric Malmi, Sebastian Krause, and Aliaksei Severyn. 2024. West-of-n: Synthetic preference generation for improved reward modeling. arXiv preprint arXiv:2401.12086
2024 arXiv
-
[54]
Xingyuan Pan, Luyang Huang, Liyan Kang, Zhicheng Liu, Yu Lu, and Shanbo Cheng. 2024. https://api.semanticscholar.org/CorpusID:269930173 G-dig: Towards gradient-based diverse and high-quality instruction data selection for machine translation . ArXiv, abs/2405.12915
2024 arXiv
-
[55]
Guilherme Penedo, Quentin Malartic, Daniel Hesslow, Ruxandra Cojocaru, Hamza Alobeidli, Alessandro Cappelli, Baptiste Pannier, Ebtesam Almazrouei, and Julien Launay. 2023. The refinedweb dataset for falcon llm: Outperforming curated corpora with web data only. Advances in Neur...
2023
-
[56]
David Rein, Betty Li Hou, Asa Cooper Stickland, Jackson Petty, Richard Yuanzhe Pang, Julien Dirani, Julian Michael, and Samuel R Bowman. 2024. Gpqa: A graduate-level google-proof q&a benchmark. In First Conference on Language Modeling
2024
-
[57]
Keisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, and Yejin Choi. 2021. Winogrande: An adversarial winograd schema challenge at scale. Communications of the ACM, 64(9):99--106
2021
-
[58]
Victor Sanh, Albert Webson, Colin Raffel, Stephen H Bach, Lintang Sutawika, Zaid Alyafeai, Antoine Chaffin, Arnaud Stiegler, Teven Le Scao, Arun Raja, et al. 2022. Multitask prompted training enables zero-shot task generalization. In ICLR 2022-Tenth International Conference on...
2022
-
[59]
Luca Soldaini, Rodney Kinney, Akshita Bhagia, Dustin Schwenk, David Atkinson, Russell Authur, Ben Bogin, Khyathi Chandu, Jennifer Dumas, Yanai Elazar, et al. 2024. Dolma: An open corpus of three trillion tokens for language model pretraining research. arXiv preprint arXiv:2402.00159
2024 arXiv
-
[60]
Ben Sorscher, Robert Geirhos, Shashank Shekhar, Surya Ganguli, and Ari Morcos. 2022. Beyond neural scaling laws: beating power law scaling via data pruning. Advances in Neural Information Processing Systems, 35:19523--19536
2022
-
[61]
Zayne Sprague, Xi Ye, Kaj Bostrom, Swarat Chaudhuri, and Greg Durrett. 2023. Musr: Testing the limits of chain-of-thought with multistep soft reasoning. arXiv preprint arXiv:2310.16049
2023 arXiv
-
[62]
Pedro Javier Ortiz Su \'a rez, Beno \^ t Sagot, and Laurent Romary. 2019. Asynchronous pipeline for processing huge corpora on medium to low resource infrastructures. In 7th Workshop on the Challenges in the Management of Large Corpora (CMLC-7). Leibniz-Institut f \"u r Deutsc...
2019
-
[63]
Mirac Suzgun, Nathan Scales, Nathanael Sch \"a rli, Sebastian Gehrmann, Yi Tay, Hyung Won Chung, Aakanksha Chowdhery, Quoc V Le, Ed H Chi, Denny Zhou, et al. 2022. Challenging big-bench tasks and whether chain-of-thought can solve them. arXiv preprint arXiv:2210.09261
2022 arXiv
-
[64]
Rohan Taori, Ishaan Gulrajani, Tianyi Zhang, Yann Dubois, Xuechen Li, Carlos Guestrin, Percy Liang, and Tatsunori B Hashimoto. 2023. Stanford alpaca: an instruction-following llama model (2023)
2023
-
[65]
Gemma Team, Thomas Mesnard, Cassidy Hardin, Robert Dadashi, Surya Bhupatiraju, Shreya Pathak, Laurent Sifre, Morgane Rivi \`e re, Mihir Sanjay Kale, Juliette Love, et al. 2024. Gemma: Open models based on gemini research and technology. arXiv preprint arXiv:2403.08295
2024 arXiv
-
[66]
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. 2023. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288
2023 arXiv
-
[67]
Thuy-Trang Vu, Xuanli He, Gholamreza Haffari, and Ehsan Shareghi. 2023. Koala: An index for quantifying overlaps with pre-training corpora. arXiv preprint arXiv:2303.14770
2023 arXiv
-
[68]
Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R Bowman. 2018. Glue: A multi-task benchmark and analysis platform for natural language understanding. arXiv preprint arXiv:1804.07461
2018 arXiv
-
[69]
Peiyi Wang, Lei Li, Liang Chen, Zefan Cai, Dawei Zhu, Binghuai Lin, Yunbo Cao, Qi Liu, Tianyu Liu, and Zhifang Sui. 2023 a . Large language models are not fair evaluators. arXiv preprint arXiv:2305.17926
2023 arXiv
-
[70]
Tianlu Wang, Ping Yu, Xiaoqing Ellen Tan, Sean O'Brien, Ramakanth Pasunuru, Jane Dwivedi-Yu, Olga Golovneva, Luke Zettlemoyer, Maryam Fazel-Zarandi, and Asli Celikyilmaz. 2023 b . Shepherd: A critic for language model generation. arXiv preprint arXiv:2308.04592
2023 arXiv
-
[71]
Yejie Wang, Keqing He, Dayuan Fu, Zhuoma Gongque, Heyang Xu, Yanxu Chen, Zhexu Wang, Yujia Fu, Guanting Dong, Muxi Diao, Jingang Wang, Mengdi Zhang, Xunliang Cai, and Weiran Xu. 2024. https://api.semanticscholar.org/CorpusID:272463773 How do your code llms perform? empowering ...
2024 arXiv
-
[72]
Yequan Wang, Jiawen Deng, Aixin Sun, and Xuying Meng. 2022 a . Perplexity from plm is unreliable for evaluating text quality. arXiv preprint arXiv:2210.05892
2022 arXiv
-
[73]
Yizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu, Noah A Smith, Daniel Khashabi, and Hannaneh Hajishirzi. 2022 b . Self-instruct: Aligning language models with self-generated instructions. arXiv preprint arXiv:2212.10560
2022 arXiv
-
[74]
Dai, and Quoc V
Jason Wei, Maarten Bosma, Vincent Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, Andrew M. Dai, and Quoc V. Le. 2021. https://api.semanticscholar.org/CorpusID:237416585 Finetuned language models are zero-shot learners . ArXiv, abs/2109.01652
2021 arXiv
-
[75]
Alexander Wettig, Aatmik Gupta, Saumya Malik, and Danqi Chen. 2024. Qurating: Selecting high-quality data for training language models. arXiv preprint arXiv:2402.09739
2024 arXiv
-
[76]
Mengzhou Xia, Sadhika Malladi, Suchin Gururangan, Sanjeev Arora, and Danqi Chen. 2024. https://api.semanticscholar.org/CorpusID:267522839 Less: Selecting influential data for targeted instruction tuning . ArXiv, abs/2402.04333
2024 arXiv
-
[77]
Can Xu, Qingfeng Sun, Kai Zheng, Xiubo Geng, Pu Zhao, Jiazhan Feng, Chongyang Tao, and Daxin Jiang. 2023. Wizardlm: Empowering large language models to follow complex instructions. arXiv preprint arXiv:2304.12244
2023 arXiv
-
[78]
Shangqing Xu and Chao Zhang. 2024. Misconfidence-based demonstration selection for llm in-context learning. arXiv preprint arXiv:2401.06301
2024 arXiv
-
[79]
Zhangchen Xu, Fengqing Jiang, Luyao Niu, Yuntian Deng, Radha Poovendran, Yejin Choi, and Bill Yuchen Lin. 2024. https://api.semanticscholar.org/CorpusID:270391432 Magpie: Alignment data synthesis from scratch by prompting aligned llms with nothing . ArXiv, abs/2406.08464
2024 arXiv
-
[80]
An Yang, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chengyuan Li, Dayiheng Liu, Fei Huang, Haoran Wei, Huan Lin, Jian Yang, Jianhong Tu, Jianwei Zhang, Jianxin Yang, Jiaxi Yang, Jingren Zhou, Junyang Lin, Kai Dang, Keming Lu, Keqin Bao, Kexin Yang, Le Yu, Me...
2024 arXiv
-
[81]
Qingyu Yin, Xuzheng He, Chak Tou Leong, Fan Wang, Yanzhao Yan, Xiaoyu Shen, and Qiang Zhang. 2024. https://doi.org/10.18653/v1/2024.findings-emnlp.239 Deeper insights without updates: The power of in-context learning over fine-tuning . In Findings of the Association for Comput...
2024 doi
-
[82]
Rowan Zellers, Ari Holtzman, Yonatan Bisk, Ali Farhadi, and Yejin Choi. 2019. Hellaswag: Can a machine really finish your sentence? arXiv preprint arXiv:1905.07830
2019 arXiv
-
[83]
Dylan Zhang, Qirun Dai, and Hao Peng. 2025. https://api.semanticscholar.org/CorpusID:276161183 The best instruction-tuning data are those that fit
2025
-
[84]
Jipeng Zhang, Yaxuan Qin, Renjie Pi, Weizhong Zhang, Rui Pan, and Tong Zhang. 2024 a . https://api.semanticscholar.org/CorpusID:271328390 Tagcos: Task-agnostic gradient clustered coreset selection for instruction tuning data . ArXiv, abs/2407.15235
2024 arXiv
-
[85]
Qi Zhang, Yiming Zhang, Haobo Wang, and Junbo Zhao. 2024 b . Recost: External knowledge guided data-efficient instruction tuning. arXiv preprint arXiv:2402.17355
2024 arXiv
-
[86]
Ce Zheng, Lei Li, Qingxiu Dong, Yuxuan Fan, Zhiyong Wu, Jingjing Xu, and Baobao Chang. 2023. Can we edit factual knowledge by in-context learning? arXiv preprint arXiv:2305.12740
2023 arXiv
-
[87]
Yu, Susan Zhang, Gargi Ghosh, Mike Lewis, Luke Zettlemoyer, and Omer Levy
Chunting Zhou, Pengfei Liu, Puxin Xu, Srini Iyer, Jiao Sun, Yuning Mao, Xuezhe Ma, Avia Efrat, Ping Yu, L. Yu, Susan Zhang, Gargi Ghosh, Mike Lewis, Luke Zettlemoyer, and Omer Levy. 2023 a . https://api.semanticscholar.org/CorpusID:258822910 Lima: Less is more for alignment . ...
2023 arXiv
-
[88]
Chunting Zhou, Pengfei Liu, Puxin Xu, Srinivasan Iyer, Jiao Sun, Yuning Mao, Xuezhe Ma, Avia Efrat, Ping Yu, Lili Yu, et al. 2023 b . Lima: Less is more for alignment. Advances in Neural Information Processing Systems, 36:55006--55021
2023
-
[89]
Jeffrey Zhou, Tianjian Lu, Swaroop Mishra, Siddhartha Brahma, Sujoy Basu, Yi Luan, Denny Zhou, and Le Hou. 2023 c . Instruction-following evaluation for large language models. arXiv preprint arXiv:2311.07911
2023 arXiv
-
[90]
online" 'onlinestring :=
ENTRY address archivePrefix author booktitle chapter edition editor eid eprint eprinttype howpublished institution journal key month note number organization pages publisher school series title type volume year doi pubmed url lastchecked label extra.label sort.label short.list...
-
[91]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 gl...
Reviewed August 15, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.