REVIEW 4 major objections 6 minor 90 references
Honey, I Shrunk the Language Model: Impact of Knowledge Distillation Methods on Performance and Explainability
T0 review · 4 major / 6 minor · reviewed 2026-08-16 · deepseek-v4-flash
Pith's one-line read Distillation can make small language models explain their answers better, not just answer correctly.
desk verdict A workmanlike distillation comparison with a solid performance finding and a genuinely confounded explainability result that needs a matched-question design before it can be the headline. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The argument rests on two mechanisms: critique-revision prompting and a synthesized training objective. Critique-revision prompting is a three-step data-generation loop in which the teacher model first writes an explanation, then critiques its own explanation, then rewrites it in light of the critique. The combined training method is the unweighted sum of a multitask loss (answer the multiple-choice question and explain the answer) and a counterfactual loss (answer correctly from the question alone, and answer incorrectly when given a question plus an intentionally wrong explanation). The counterfactual component teaches the student to be sensitive to misleading reasoning, while the multitask component teaches it to produce faithful explanations. Together with revised explanations as training data, these mechanisms produce the student that humans rate as best at completeness and contrastiveness.
What would settle it
A reader could rerun the human study but split each model's rated explanations by question difficulty (for instance, by the fraction of models that answered each question correctly) and check whether the MT+CF:Revised advantage on completeness and contrastiveness persists within difficulty-stratified subsets. If the advantage disappears once question difficulty is controlled, the paper's central explainability claim would collapse.
Extended reading notes
Core claim
The paper's central claim is that the quality of the explanations a distilled student model gives to humans can be improved substantially by improving the training data through critique-revision prompting, and that this improvement shows up specifically in two explainability dimensions: completeness and contrastiveness. On the performance side, multitask training alone yields the strongest student in terms of accuracy, while counterfactual training alone performs worst. The combined multitask-plus-counterfactual student trained on revised explanations (MT+CF:Revised) significantly outperforms the other students on those two explanation-quality dimensions in the human-grounded study, even though its accuracy is not top. The authors therefore argue that the choice of distillation method involves a trade-off between raw task performance and the perceived quality of explanations, and that for applications where explainability matters, the combined method on revised data is the better choice.
Load-bearing premise
The explainability comparison assumes that judging only explanations for questions the student answered correctly does not distort the comparison, even though the student models have different accuracy levels and therefore effectively evaluate different, potentially easier or harder, subsets of questions.
Editorial extensions
If this is right
- If the central result holds, practitioners who care about explainability can keep a small student model and a modest teacher (here a 13B model) rather than needing a much larger teacher to get high-quality human-facing explanations.
- The performance-versus-explainability trade-off suggests that evaluation of distillation methods should include human-grounded explanation ratings, not just accuracy benchmarks.
- The finding that counterfactual training alone hurts accuracy and does not improve explainability calls into question the automatic adoption of counterfactual objectives for reasoning tasks.
- The authors' observation that a 13B teacher can produce students comparable to those from a 540B teacher challenges the assumption that bigger teachers are always better for distillation.
- The framework of standardizing teacher and training settings makes different distillation methods directly comparable, which is a precondition for building a reliable body of knowledge about what works.
Reading between the lines
- The critique-revision loop could be generalized beyond data generation: the same teacher could critique and revise the student's own explanations during inference, potentially improving explanation quality without retraining.
- Because the combined method adds a counterfactual loss, the observed improvement in contrastiveness may partly reflect the student learning to reason about why other answer choices are wrong, a skill that could transfer to other multiple-choice reasoning tasks.
- A natural testable extension is to vary the temperature or number of critique-revision iterations and measure whether explanation quality continues to improve or plateaus, and whether a larger student can exploit longer revised explanations better than a smaller one.
- The limited effect sizes (VDA around 0.4) suggest that the practical benefit of the combined method may depend on the stakes of the application; for high-stakes uses, even small gains in completeness and contrastiveness could matter.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper presents a controlled comparison of knowledge-distillation variants for T5 student models (220M and 770M parameters) trained on CommonsenseQA data generated by LLaMA-2-13B. The four conditions are counterfactual training (CF:Unrevised), multitask training (MT:Unrevised), combined multitask and counterfactual training (MT+CF:Unrevised), and combined training with critique-revision-prompted revised explanations (MT+CF:Revised). Performance is measured by accuracy on the CommonsenseQA test set; explainability is measured through a within-subject human study with 117 participants who rated explanations on five dimensions. The paper reports that multitask training yields the strongest accuracy, while MT+CF:Revised shows the largest perceived completeness, contrastiveness, and overall quality, leading to the conclusion that critique-revision prompting benefits explainability even though it does not improve accuracy.
Significance. This is a useful empirical contribution: it provides a standardized comparison of data-generation and training choices for distillation, and it evaluates explainability with a human-grounded study rather than a proxy metric. The authors' public repository, fixed seeds, attention-checked participant pool, randomization of statement and task order, and use of ordinal-appropriate nonparametric tests in part of the analysis are strengths. If the main claims hold, the finding that critique-revision prompting improves perceived explanation quality without improving accuracy, and that multitask training dominates counterfactual training, is actionable for practitioners building small deployed models. The novelty is incremental rather than conceptual, but the paper fills a real gap in comparative evaluation of distillation methods.
major comments (4)
- [Section IV-B, Section V-B] Section IV-B states that the study 'included only correctly answered explanations to focus on explanation quality rather than answer correctness,' but this creates a selection confound because accuracy differs substantially across models (Table I: at 220M, MT:Unrevised is +12.66 over CF:Unrevised; at 770M, MT+CF:Revised is -1.90 versus MT:Unrevised). Each model's rated set is therefore conditioned on that model's correctness, so lower-accuracy models may be rated only on easier or systematically different questions, and Table II's n=1114 (out of a possible 1404 ratings) confirms that explanations were dropped unevenly. The Kruskal-Wallis, Dunn, and regression analyses in Section V-B do not include question identity or difficulty as a covariate or blocking factor. The reported advantage of MT+CF:Revised on completeness and contrastiveness could thus be an artifact of question-subset selection rather than explanation quality. Please re-analyze on the common set of questions answered correctly by all models, or add question fixed effects or difficulty covariates, and report per-model n and set overlap.
- [Section V-A] Section V-A excludes outliers defined as points beyond 1.5 times the IQR, but does not report how many training runs were performed per condition, how many outliers were removed, or the degrees of freedom in the ANOVA and Tukey tests. Since Section IV-A mentions that 'subsequent models with different random seeds were trained,' the number of seeds per model is essential for interpreting Table I and the associated p-values. Please report n per cell, the seed count, the excluded outlier counts, and the results of the pairwise comparisons with and without outlier removal.
- [Section V-B] The claim that MT+CF:Revised significantly outperforms the other student models on completeness and contrastiveness is stronger than the reported evidence. The Dunn test supports a significant contrastiveness advantage over all three other models, but for completeness Section V-B states only that MT+CF:Revised 'surpasses CF:Unrevised significantly'; significant pairwise differences against MT:Unrevised and MT+CF:Unrevised are not reported. The abstract and the first contribution bullet should be revised to state the dimension-specific pairwise results, or the full pairwise comparison matrix should be provided.
- [Section V-B] The human study is within-subject (each of 117 participants rated 12 explanations, three per student model), but the Kruskal-Wallis and Dunn tests and the OLS regressions in Section V-B appear to treat every rating as independent, without participant random effects or cluster-robust standard errors. If ratings from the same participant are correlated, the reported p-values (e.g., p=0.0004 for contrastiveness) may be anti-conservative. A Friedman test or a mixed-effects model with participant and possibly question random intercepts would be more appropriate; if a repeated-measures method was actually used, please describe it explicitly.
minor comments (6)
- [Section IV-A] Section IV-A states 'T5-base with 220M million (220M) parameters'; 'million' is redundant and should be removed.
- [Figure 2] Figure 2 contains the typo 'Few-short Prompting'; this should be 'Few-shot Prompting' in both the figure and its legend table.
- [Section III-C] Section III-C, Eq. (6), writes 'Lconterf actual'; this should be 'L_counterfactual'.
- [Section V-A] Section V-A contains the typo 'ANOV A approach'; this should be 'ANOVA approach'.
- [Section V-A] The comparison with the PaLM 540B teacher from [17] is not a controlled comparison because the teacher model, dataset, and training settings differ; this should be framed as a suggestive cross-paper observation rather than a firm conclusion.
- [Section IV-A] The sentence 'Initial validation with corrupted samples from the CQA dataset showed that revised explanations add context-relevant information' is vague; please specify the corruption procedure and the validation metric used.
Circularity Check
No circularity: the paper's claims are empirical comparisons with external human ratings and test-set accuracy, not derivations from fitted inputs or self-citation chains.
full rationale
This is an empirical comparison study rather than a derivation, so the circularity patterns do not apply. The student models are trained with explicit loss functions (L_multitask, L_counterfactual, L_combined in Eqs. 3, 6, 7), which are defined from the training data and prior literature, not from the evaluated outcomes. Performance is measured as accuracy on the held-out CQA test set, and explainability is measured through human ratings of generated explanations, which are external judgments rather than quantities encoded in the loss functions or data-generation procedure. The central claim that the MT+CF:Revised model improves completeness and contrastiveness is supported by Kruskal-Wallis, Dunn, and regression analyses of human ratings; nothing in those claims reduces by definition to the training objective or to the way the models were constructed. The human study's decision to include only correctly answered explanations (Section IV-B) is a potential validity confound because different models may be correct on different question subsets, but this is a question of experimental control, not circularity: the rated explanations are not the fitted inputs of the claimed prediction. Self-citations [29], [30], [33] are used only to support established evaluation methodology and are not load-bearing for the paper's new results. No step was found in which a prediction is equivalent to its input by construction or in which a fitted parameter is renamed as a predicted outcome.
Assumptions & free parameters
free parameters (2)
- Outlier exclusion rule for performance ANOVA =
1.5 × IQR
- Outlier exclusion for regression on quality =
17 outliers removed
assumptions (5)
- domain assumption Human ratings on the five Likert dimensions (plausibility, understandability, completeness, satisfaction, contrastiveness) validly measure the explainability of student models.
- domain assumption Restricting the explainability evaluation to explanations for correctly answered questions does not bias comparisons across models with different accuracies.
- domain assumption CommonsenseQA test accuracy is a valid measure of student model performance for the studied distillation methods.
- ad hoc to paper The unweighted sum of multitask and counterfactual losses is a reasonable combined training objective.
- standard math The assumptions of the parametric and nonparametric tests (normality after outlier removal, variance homogeneity) hold for the data.
Cite this review
Pith. "Pith review of Honey, I Shrunk the Language Model: Impact of Knowledge Distillation Methods on Performance and Explainability." pith.science (2026). https://pith.science/paper/YI4UWH57
@misc{pith2026250416056,
author = {Pith},
title = {Pith review of: Honey, I Shrunk the Language Model: Impact of Knowledge Distillation Methods on Performance and Explainability},
year = {2026},
howpublished = {\url{https://pith.science/paper/YI4UWH57}},
note = {Machine review of arXiv:2504.16056}
}
read the original abstract
Artificial Intelligence (AI) has increasingly influenced modern society, recently in particular through significant advancements in Large Language Models (LLMs). However, high computational and storage demands of LLMs still limit their deployment in resource-constrained environments. Knowledge distillation addresses this challenge by training a small student model from a larger teacher model. Previous research has introduced several distillation methods for both generating training data and for training the student model. Despite their relevance, the effects of state-of-the-art distillation methods on model performance and explainability have not been thoroughly investigated and compared. In this work, we enlarge the set of available methods by applying critique-revision prompting to distillation for data generation and by synthesizing existing methods for training. For these methods, we provide a systematic comparison based on the widely used Commonsense Question-Answering (CQA) dataset. While we measure performance via student model accuracy, we employ a human-grounded study to evaluate explainability. We contribute new distillation methods and their comparison in terms of both performance and explainability. This should further advance the distillation of small language models and, thus, contribute to broader applicability and faster diffusion of LLM technology.
Figures
Reference graph
Works this paper leans on
-
[1]
Language models are few-shot learners,
T. B. Brown, B. Mann, N. Ryder, M. Subbiah, J. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, S. Agarwal, A. Herbert- V oss, G. Krueger, T. Henighan, R. Child, A. Ramesh, D. M. Ziegler, J. Wu, C. Winter, C. Hesse, M. Chen, E. Sigler, M. Litwin, S. Gray, B. Chess, J. Clark, C. Berner, S. McCandlish, A. Radford, I. Sutskever, and D. Am...
2020
-
[2]
Language models are unsupervised multitask learners,
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, I. Sutskever et al., “Language models are unsupervised multitask learners,” OpenAI blog , vol. 1, no. 8, p. 9, 2019
2019
-
[3]
How many data points is a prompt worth?
T. Le Scao and A. Rush, “How many data points is a prompt worth?” in Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. Online: Association for Computational Linguistics, 2021, pp. 2627–2636. [Online]. Available: https://aclanthology.org/202 1.naacl-main.208
2021
-
[4]
W. Jin, Y . Cheng, Y . Shen, W. Chen, and X. Ren, “A good prompt is worth millions of parameters? low-resource prompt-based learning for vision-language models,” CoRR, vol. abs/2110.08484, 2021. [Online]. Available: https://arxiv.org/abs/2110.08484
arXiv 2021
-
[5]
Orca 2: Teaching small language models how to reason,
A. Mitra, L. D. Corro, S. Mahajan, A. Codas, C. Simoes, S. Agar- wal, X. Chen, A. Razdaibiedina, E. Jones, K. Aggarwal, H. Palangi, G. Zheng, C. Rosset, H. Khanpour, and A. Awadallah, “Orca 2: Teaching small language models how to reason,” 2023
2023
-
[6]
Chain-of-thought prompting elicits reasoning in large language models,
J. Wei, X. Wang, D. Schuurmans, M. Bosma, B. Ichter, F. Xia, E. H. Chi, Q. V . Le, and D. Zhou, “Chain-of-thought prompting elicits reasoning in large language models,” in Proceedings of the 36th International Conference on Neural Information Processing Systems , ser. NIPS ’22. Red Hook, NY , USA: Curran Associates Inc., 2022
2022
-
[7]
Tree of thoughts: deliberate problem solving with large language models,
S. Yao, D. Yu, J. Zhao, I. Shafran, T. L. Griffiths, Y . Cao, and K. Narasimhan, “Tree of thoughts: deliberate problem solving with large language models,” in Proceedings of the 37th International Conference on Neural Information Processing Systems , ser. NIPS ’23. Red Hook, NY , USA: Curran Associates Inc., 2023
2023
-
[8]
Evaluating Consistency and Reasoning Capabilities of Large Language Models,
Y . Saxena, S. Chopra, and A. M. Tripathi, “Evaluating Consistency and Reasoning Capabilities of Large Language Models,” Apr. 2024, arXiv:2404.16478 [cs]. [Online]. Available: http://arxiv.org/abs/2404.1 6478
arXiv 2024
Show all 90 references
-
[9]
LogicBench: Towards Systematic Evaluation of Logical Reasoning Ability of Large Language Models,
M. Parmar, N. Patel, N. Varshney, M. Nakamura, M. Luo, S. Mashetty, A. Mitra, and C. Baral, “LogicBench: Towards Systematic Evaluation of Logical Reasoning Ability of Large Language Models,” in Proceedings of the 62nd Annual Meeting of the Association for Computational Linguis...
2024
-
[10]
Towards a Mechanistic Interpretation of Multi-Step Reasoning Capabilities of Language Models,
Y . Hou, J. Li, Y . Fei, A. Stolfo, W. Zhou, G. Zeng, A. Bosselut, and M. Sachan, “Towards a Mechanistic Interpretation of Multi-Step Reasoning Capabilities of Language Models,” in Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, H. Bouam...
2023
-
[11]
Large Language Models: A Survey,
S. Minaee, T. Mikolov, N. Nikzad, M. Chenaghlu, R. Socher, X. Amatriain, and J. Gao, “Large Language Models: A Survey,” Feb. 2024, arXiv:2402.06196 [cs]. [Online]. Available: http://arxiv.org/abs/ 2402.06196
2024 arXiv
-
[12]
PaLM: Scaling Language Modeling with Pathways,
A. Chowdhery, S. Narang, J. Devlin, M. Bosma, G. Mishra, A. Roberts, P. Barham, H. W. Chung, C. Sutton, S. Gehrmann, P. Schuh, K. Shi, S. Tsvyashchenko, J. Maynez, A. Rao, P. Barnes, Y . Tay, N. Shazeer, V . Prabhakaran, E. Reif, N. Du, B. Hutchinson, R. Pope, J. Bradbury, J. ...
2022 arXiv
-
[13]
Distilling Task-Specific Knowledge from BERT into Simple Neural Networks,
R. Tang, Y . Lu, L. Liu, L. Mou, O. Vechtomova, and J. Lin, “Distilling Task-Specific Knowledge from BERT into Simple Neural Networks,” Mar. 2019, arXiv:1903.12136 [cs]. [Online]. Available: http://arxiv.org/abs/1903.12136
2019 arXiv
-
[14]
TinyBERT: Distilling BERT for Natural Language Understanding,
X. Jiao, Y . Yin, L. Shang, X. Jiang, X. Chen, L. Li, F. Wang, and Q. Liu, “TinyBERT: Distilling BERT for Natural Language Understanding,” in Findings of the Association for Computational Linguistics: EMNLP 2020 , T. Cohn, Y . He, and Y . Liu, Eds. Online: Association for Comp...
2020
-
[15]
A survey on model compression for large language models,
X. Zhu, J. Li, Y . Liu, C. Ma, and W. Wang, “A survey on model compression for large language models,” Transactions of the Association for Computational Linguistics , vol. 12, pp. 1556–1577,
-
[16]
Robustness-Reinforced Knowledge Distillation With Correlation Distance and Network Pruning,
S. Kim, G. Ham, Y . Cho, and D. Kim, “Robustness-Reinforced Knowledge Distillation With Correlation Distance and Network Pruning,” IEEE Transactions on Knowledge and Data Engineering , vol. 36, no. 12, pp. 9163–9175, Dec. 2024, conference Name: IEEE Transactions on Knowledge a...
2024
-
[17]
Distilling step-by-step! outperforming larger language models with less training data and smaller model sizes,
C.-Y . Hsieh, C.-L. Li, C.-K. Yeh, H. Nakhost, Y . Fujii, A. J. Ratner, R. Krishna, C.-Y . Lee, and T. Pfister, “Distilling step-by-step! outperforming larger language models with less training data and smaller model sizes,” ArXiv, vol. abs/2305.02301, 2023. [Online]. Availabl...
2023 arXiv
-
[18]
PanDa: Prompt Transfer Meets Knowledge Distillation for Efficient Model Adaptation,
Q. Zhong, L. Ding, J. Liu, B. Du, and D. Tao, “PanDa: Prompt Transfer Meets Knowledge Distillation for Efficient Model Adaptation,” IEEE Transactions on Knowledge and Data Engineering , vol. 36, no. 9, pp. 4835–4848, Sep. 2024, conference Name: IEEE Transactions on Knowledge a...
2024
-
[19]
Efficient Knowledge Distillation: Empowering Small Language Models with Teacher Model Insights,
M. Ballout, U. Krumnack, G. Heidemann, and K.-U. K ¨uhnberger, “Efficient Knowledge Distillation: Empowering Small Language Models with Teacher Model Insights,” Sep. 2024, arXiv:2409.12586 [cs]. [Online]. Available: http://arxiv.org/abs/2409.12586
2024 arXiv
-
[20]
Improve Student‘s Reasoning Generalizability through Cascading Decomposed CoTs Distillation,
C. Dai, K. Li, W. Zhou, and S. Hu, “Improve Student‘s Reasoning Generalizability through Cascading Decomposed CoTs Distillation,” in Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing , Y . Al-Onaizan, M. Bansal, and Y .-N. Chen, Eds. Miami,...
2024
-
[21]
Beyond imitation: Learning key reasoning steps from dual chain-of-thoughts in reasoning distillation,
——, “Beyond imitation: Learning key reasoning steps from dual chain-of-thoughts in reasoning distillation,” CoRR, vol. abs/2405.19737,
-
[22]
One teacher is enough? pre-trained language model distillation from multiple teachers,
C. Wu, F. Wu, and Y . Huang, “One teacher is enough? pre-trained language model distillation from multiple teachers,” in Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021 , C. Zong, F. Xia, W. Li, and R. Navigli, Eds. Online: Association for Computatio...
2021
- [23]
-
[24]
We’re getting a better idea of AI’s true carbon footprint,
M. Heikkil ¨a, “We’re getting a better idea of AI’s true carbon footprint,” Nov. 2022. [Online]. Available: https://www.technologyreview.com/202 2/11/14/1063192/were-getting-a-better-idea-of-ais-true-carbon-footpri nt/
2022
-
[25]
Distilling Large Vision-Language Model with Out-of-Distribution Generalizability ,
X. Li, Y . Fang, M. Liu, Z. Ling, Z. Tu, and H. Su, “ Distilling Large Vision-Language Model with Out-of-Distribution Generalizability ,” in 2023 IEEE/CVF International Conference on Computer Vision (ICCV) . Los Alamitos, CA, USA: IEEE Computer Society, Oct. 2023, pp. 2492–250...
2023
-
[26]
How Far Can Camels Go? Exploring the State of Instruction Tuning on Open Resources,
Y . Wang, H. Ivison, P. Dasigi, J. Hessel, T. Khot, K. R. Chandu, D. Wadden, K. MacMillan, N. A. Smith, I. Beltagy, and H. Hajishirzi, “How Far Can Camels Go? Exploring the State of Instruction Tuning on Open Resources,” Oct. 2023, arXiv:2306.04751 [cs]. [Online]. Available: h...
2023 arXiv
-
[27]
Energy and policy considerations for modern deep learning research,
E. Strubell, A. Ganesh, and A. McCallum, “Energy and policy considerations for modern deep learning research,” Proceedings of the AAAI Conference on Artificial Intelligence , vol. 34, no. 09, pp. 13 693–13 696, Apr. 2020. [Online]. Available: https://ojs.aaai.org/ind ex.php/AA...
2020
-
[28]
Towards a rigorous science of interpretable machine learning,
F. Doshi-Velez and B. Kim, “Towards a rigorous science of interpretable machine learning,” arXiv: Machine Learning, 2017. [Online]. Available: https://api.semanticscholar.org/CorpusID:11319376
2017
-
[29]
Code Llama: Open Foundation Models for Code,
B. Rozi `ere, J. Gehring, F. Gloeckle, S. Sootla, I. Gat, X. E. Tan, Y . Adi, J. Liu, R. Sauvestre, T. Remez, J. Rapin, A. Kozhevnikov, I. Evtimov, J. Bitton, M. Bhatt, C. C. Ferrer, A. Grattafiori, W. Xiong, JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUST 2021 11 A. D ´e...
2021 arXiv
-
[30]
On the Effect of Contextual Information on Human Delegation Behavior in Human-AI collaboration,
P. Spitzer, J. Holstein, P. Hemmer, M. V ¨ossing, N. K ¨uhl, D. Martin, and G. Satzger, “On the Effect of Contextual Information on Human Delegation Behavior in Human-AI collaboration,” Jan. 2024, arXiv:2401.04729 [cs]. [Online]. Available: http://arxiv.org/abs/2401.0 4729
2024 arXiv
-
[31]
Explainability in AI Based Applications: A Framework for Comparing Different Techniques,
A. Grobrugge, N. Mishra, J. Jakubik, and G. Satzger, “Explainability in AI Based Applications: A Framework for Comparing Different Techniques,” Oct. 2024, arXiv:2410.20873 [cs]. [Online]. Available: http://arxiv.org/abs/2410.20873
2024 arXiv
-
[32]
Multimodal explanations: Justifying decisions and pointing to the evidence,
D. H. Park, L. A. Hendricks, Z. Akata, A. Rohrbach, B. Schiele, T. Darrell, and M. Rohrbach, “Multimodal explanations: Justifying decisions and pointing to the evidence,” 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition , pp. 8779–8788, 2018. [Online]. Avail...
2018
-
[33]
Textual explanations for self-driving vehicles,
J. Kim, A. Rohrbach, T. Darrell, J. Canny, and Z. Akata, “Textual explanations for self-driving vehicles,” in Computer Vision – ECCV 2018, V . Ferrari, M. Hebert, C. Sminchisescu, and Y . Weiss, Eds. Cham: Springer International Publishing, 2018, pp. 577–593
2018
-
[34]
SCOTT: Self-consistent chain-of-thought distillation,
P. Wang, Z. Wang, Z. Li, Y . Gao, B. Yin, and X. Ren, “SCOTT: Self-consistent chain-of-thought distillation,” in Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , A. Rogers, J. Boyd-Graber, and N. Okazaki, Eds. To...
2023
-
[35]
The Impact of Imperfect XAI on Human-AI Decision-Making,
K. Morrison, P. Spitzer, V . Turri, M. Feng, N. K ¨uhl, and A. Perer, “The Impact of Imperfect XAI on Human-AI Decision-Making,” Proceedings of the ACM on Human-Computer Interaction , vol. 8, no. CSCW1, pp. 1–39, Apr. 2024, arXiv:2307.13566 [cs]. [Online]. Available: http://ar...
2024 arXiv
-
[36]
GPT-NeoX-20B: An open-source autoregressive language model,
S. Black, S. Biderman, E. Hallahan, Q. Anthony, L. Gao, L. Golding, H. He, C. Leahy, K. McDonell, J. Phang, M. Pieler, U. S. Prashanth, S. Purohit, L. Reynolds, J. Tow, B. Wang, and S. Weinbach, “GPT-NeoX-20B: An open-source autoregressive language model,” in Proceedings of Bi...
2022
-
[37]
Constitutional AI: Harmlessness from AI Feedback,
Y . Bai, S. Kadavath, S. Kundu, A. Askell, J. Kernion, A. Jones, A. Chen, A. Goldie, A. Mirhoseini, C. McKinnon, C. Chen, C. Olsson, C. Olah, D. Hernandez, D. Drain, D. Ganguli, D. Li, E. Tran-Johnson, E. Perez, J. Kerr, J. Mueller, J. Ladish, J. Landau, K. Ndousse, K. Lukosui...
2022 arXiv
-
[38]
MiniLLM: Knowledge distillation of large language models,
Y . Gu, L. Dong, F. Wei, and M. Huang, “MiniLLM: Knowledge distillation of large language models,” in The Twelfth International Conference on Learning Representations , 2024. [Online]. Available: https://openreview.net/forum?id=5h0qf7IBZZ
2024
-
[39]
Distilling the Knowledge in a Neural Network,
G. Hinton, O. Vinyals, and J. Dean, “Distilling the Knowledge in a Neural Network,” Mar. 2015, arXiv:1503.02531 [cs, stat]. [Online]. Available: http://arxiv.org/abs/1503.02531
2015 arXiv
-
[40]
Collaborative Distillation for Ultra-Resolution Universal Style Transfer,
H. Wang, Y . Li, Y . Wang, H. Hu, and M.-H. Yang, “Collaborative Distillation for Ultra-Resolution Universal Style Transfer,” Mar. 2020, arXiv:2003.08436 [cs, eess]. [Online]. Available: http://arxiv.org/abs/20 03.08436
2020 arXiv
-
[41]
ERDL: Efficient Retrieval Framework Based on Distillation from Large Language Models,
H. Yu, R. Li, Z. Zhang, S. Ye, Q. Liu, Z. Huang, and E. Chen, “ERDL: Efficient Retrieval Framework Based on Distillation from Large Language Models,” in 2024 International Conference on Computational Linguistics and Natural Language Processing (CLNLP) , Jul. 2024, pp. 83–87. [...
2024
-
[42]
Learning Efficient Object Detection Models with Knowledge Distillation,
G. Chen, W. Choi, X. Yu, T. Han, and M. Chandraker, “Learning Efficient Object Detection Models with Knowledge Distillation,” in Advances in Neural Information Processing Systems , vol. 30. Curran Associates, Inc., 2017. [Online]. Available: https://papers.nips.cc/paper files/...
2017
-
[43]
Triplet Distillation for Deep Face Recognition,
Y . Feng, H. Wang, D. T. Yi, and R. Hu, “Triplet Distillation for Deep Face Recognition,” May 2019, arXiv:1905.04457 [cs]. [Online]. Available: http://arxiv.org/abs/1905.04457
2019 arXiv
-
[44]
Training Compact Models for Low Resource Entity Tagging using Pre-trained Language Models,
P. Izsak, S. Guskin, and M. Wasserblat, “Training Compact Models for Low Resource Entity Tagging using Pre-trained Language Models,” in 2019 Fifth Workshop on Energy Efficient Machine Learning and Cognitive Computing - NeurIPS Edition (EMC2-NIPS) , Dec. 2019, pp. 44–47. [Onlin...
2019
-
[45]
A Survey on Knowledge Distillation of Large Language Models,
X. Xu, M. Li, C. Tao, T. Shen, R. Cheng, J. Li, C. Xu, D. Tao, and T. Zhou, “A Survey on Knowledge Distillation of Large Language Models,” Oct. 2024, arXiv:2402.13116 [cs]. [Online]. Available: http://arxiv.org/abs/2402.13116
2024 arXiv
-
[46]
Lora: Low-rank adaptation of large language models,
E. J. Hu, Y . Shen, P. Wallis, Z. Allen-Zhu, Y . Li, S. Wang, L. Wang, and W. Chen, “Lora: Low-rank adaptation of large language models,”
-
[47]
On the effectiveness of adapter-based tuning for pretrained language model adaptation,
R. He, L. Liu, H. Ye, Q. Tan, B. Ding, L. Cheng, J. Low, L. Bing, and L. Si, “On the effectiveness of adapter-based tuning for pretrained language model adaptation,” in Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th Interna...
2021 doi
-
[48]
The Power of Scale for Parameter-Efficient Prompt Tuning,
B. Lester, R. Al-Rfou, and N. Constant, “The Power of Scale for Parameter-Efficient Prompt Tuning,” in Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, M.-F. Moens, X. Huang, L. Specia, and S. W.-t. Yih, Eds. Online and Punta Cana, Domini...
2021
-
[49]
Efficiency Optimization of Large-Scale Language Models Based on Deep Learning in Natural Language Processing Tasks,
T. Mei, Y . Zi, X. Cheng, Z. Gao, Q. Wang, and H. Yang, “Efficiency Optimization of Large-Scale Language Models Based on Deep Learning in Natural Language Processing Tasks,” in 2024 IEEE 2nd International Conference on Sensors, Electronics and Computer Engineering (ICSECE), Au...
2024
-
[50]
LoPT: Low-Rank Prompt Tuning for Parameter Efficient Language Models,
S. Guo, S. Damani, and K.-h. Chang, “LoPT: Low-Rank Prompt Tuning for Parameter Efficient Language Models,” 2024, version Number: 1. [Online]. Available: https://arxiv.org/abs/2406.19486
2024 arXiv
-
[51]
A survey on the explainability of su- pervised machine learning,
N. Burkart and M. F. Huber, “A survey on the explainability of su- pervised machine learning,” Journal of Artificial Intelligence Research , vol. 70, pp. 245–317, 2021
2021
-
[52]
Baize: An Open- Source Chat Model with Parameter-Efficient Tuning on Self-Chat Data,
C. Xu, D. Guo, N. Duan, and J. McAuley, “Baize: An Open- Source Chat Model with Parameter-Efficient Tuning on Self-Chat Data,” in Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing , H. Bouamor, J. Pino, and K. Bali, Eds. Singapore: Associat...
2023
-
[53]
Causability and explainability of artificial intelligence in medicine,
A. Holzinger, G. Langs, H. Denk, K. Zatloukal, and H. M ¨uller, “Causability and explainability of artificial intelligence in medicine,” Wiley Interdisciplinary Reviews. Data Mining and Knowledge Discovery, vol. 9, 2019. [Online]. Available: https://api.semanticscholar.org/Cor...
2019
-
[54]
Large Language Models Are Reasoning Teachers,
N. Ho, L. Schmid, and S.-Y . Yun, “Large Language Models Are Reasoning Teachers,” Jun. 2023, arXiv:2212.10071 [cs]. [Online]. Available: http://arxiv.org/abs/2212.10071
2023 arXiv
-
[55]
Specializing Smaller Language Models towards Multi-Step Reasoning,
Y . Fu, H. Peng, L. Ou, A. Sabharwal, and T. Khot, “Specializing Smaller Language Models towards Multi-Step Reasoning,” Jan. 2023, arXiv:2301.12726 [cs]. [Online]. Available: http://arxiv.org/abs/2301.1 2726
2023 arXiv
-
[56]
Teaching small language models to reason,
L. C. Magister, J. Mallinson, J. Adamek, E. Malmi, and A. Severyn, “Teaching small language models to reason,” in Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers) , A. Rogers, J. Boyd-Graber, and N. Okazaki, Eds. ...
2023
-
[57]
Orca: Progressive Learning from Complex Explanation Traces of GPT-4,
S. Mukherjee, A. Mitra, G. Jawahar, S. Agarwal, H. Palangi, and A. Awadallah, “Orca: Progressive Learning from Complex Explanation Traces of GPT-4,” Jun. 2023, arXiv:2306.02707 [cs]. [Online]. Available: http://arxiv.org/abs/2306.02707
2023 arXiv
-
[58]
Explain Yourself! Leveraging Language Models for Commonsense Reasoning,
N. F. Rajani, B. McCann, C. Xiong, and R. Socher, “Explain Yourself! Leveraging Language Models for Commonsense Reasoning,” in Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics . Florence, Italy: Association for Computational Linguistics, ...
2019
-
[59]
Sci-CoT: Leveraging Large Language Models for Enhanced Knowledge Distillation in Small Models for Scientific QA,
Y . Ma, C. Fan, and H. Jiang, “Sci-CoT: Leveraging Large Language Models for Enhanced Knowledge Distillation in Small Models for Scientific QA,” in 2023 9th International Conference on Computer and Communications (ICCC), Dec. 2023, pp. 2394–2398, iSSN: 2837-7109. [Online]. Ava...
2023
-
[60]
Learning the Difference that Makes a Difference with Counterfactually-Augmented Data,
D. Kaushik, E. Hovy, and Z. C. Lipton, “Learning the Difference that Makes a Difference with Counterfactually-Augmented Data,” Feb. 2020, arXiv:1909.12434 [cs, stat]. [Online]. Available: http: //arxiv.org/abs/1909.12434
2020 arXiv
-
[61]
Measuring Association Between Labels and Free-Text Rationales,
S. Wiegreffe, A. Marasovi ´c, and N. A. Smith, “Measuring Association Between Labels and Free-Text Rationales,” in Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing , M.-F. Moens, X. Huang, L. Specia, and S. W.-t. Yih, Eds. Online and Punta...
2021
-
[62]
Self-consistency improves chain of thought reasoning in language models,
X. Wang, J. Wei, D. Schuurmans, Q. V . Le, E. H. Chi, S. Narang, A. Chowdhery, and D. Zhou, “Self-consistency improves chain of thought reasoning in language models,” in The Eleventh International Conference on Learning Representations , 2023. [Online]. Available: https://open...
2023
-
[63]
Llama 2: Open Foundation and Fine-Tuned Chat Models,
H. Touvron, L. Martin, K. Stone, P. Albert, A. Almahairi, Y . Babaei, N. Bashlykov, S. Batra, P. Bhargava, S. Bhosale, D. Bikel, L. Blecher, C. Canton Ferrer, M. Chen, G. Cucurull, D. Esiobu, J. Fernandes, J. Fu, W. Fu, B. Fuller, C. Gao, V . Goswami, N. Goyal, A. Hartshorn, S...
2023 arXiv
-
[64]
RLCD: Reinforcement learning from contrastive distillation for LM alignment,
K. Yang, D. Klein, A. Celikyilmaz, N. Peng, and Y . Tian, “RLCD: Reinforcement learning from contrastive distillation for LM alignment,” in The Twelfth International Conference on Learning Representations ,
-
[65]
CommonsenseQA: A Question Answering Challenge Targeting Commonsense Knowledge,
A. Talmor, J. Herzig, N. Lourie, and J. Berant, “CommonsenseQA: A Question Answering Challenge Targeting Commonsense Knowledge,” in Proceedings of the 2019 Conference of the North . Minneapolis, Minnesota: Association for Computational Linguistics, 2019, pp. 4149–4158. [Online...
2019
-
[66]
Transformers: State- of-the-art natural language processing,
T. Wolf, L. Debut, V . Sanh, J. Chaumond, C. Delangue, A. Moi, P. Cistac, T. Rault, R. Louf, M. Funtowicz, J. Davison, S. Shleifer, P. von Platen, C. Ma, Y . Jernite, J. Plu, C. Xu, T. Le Scao, S. Gugger, M. Drame, Q. Lhoest, and A. Rush, “Transformers: State- of-the-art natur...
2020
-
[67]
Generating visual explanations,
L. A. Hendricks, Z. Akata, M. Rohrbach, J. Donahue, B. Schiele, and T. Darrell, “Generating visual explanations,” inComputer Vision – ECCV 2016, B. Leibe, J. Matas, N. Sebe, and M. Welling, Eds. Cham: Springer International Publishing, 2016, pp. 3–19
2016
-
[68]
Available: https://openreview.net/forum?id=v3XXtxWK i6
[Online]. Available: https://openreview.net/forum?id=v3XXtxWK i6
-
[69]
Exploring the limits of transfer learning with a unified text-to-text transformer,
C. Raffel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y . Zhou, W. Li, and P. J. Liu, “Exploring the limits of transfer learning with a unified text-to-text transformer,” J. Mach. Learn. Res. , vol. 21, no. 1, Jan. 2020
2020
-
[70]
Exploring Evaluation Methods for Interpretable Machine Learning: A Survey,
N. Alangari, M. El Bachir Menai, H. Mathkour, and I. Almosallam, “Exploring Evaluation Methods for Interpretable Machine Learning: A Survey,” Information, vol. 14, no. 8, p. 469, Aug. 2023. [Online]. Available: https://www.mdpi.com/2078-2489/14/8/469
2023
-
[71]
Evaluating the Quality of Machine Learning Explanations: A Survey on Methods and Metrics,
J. Zhou, A. H. Gandomi, F. Chen, and A. Holzinger, “Evaluating the Quality of Machine Learning Explanations: A Survey on Methods and Metrics,” Electronics, vol. 10, no. 5, p. 593, Mar. 2021. [Online]. Available: https://www.mdpi.com/2079-9292/10/5/593
2021
-
[72]
Prolific · Quickly find research participants you can trust
Prolific, “Prolific · Quickly find research participants you can trust.”
-
[73]
Oxford Learner’s Dictionaries | Find definitions, trans- lations, and grammar explanations at Oxford Learner’s Dictionaries,
O. Dictionary, “Oxford Learner’s Dictionaries | Find definitions, trans- lations, and grammar explanations at Oxford Learner’s Dictionaries,”
-
[74]
Interpretation quality score for measuring the quality of interpretability methods,
Y . Xie, S. V osoughi, and S. Hassanpour, “Interpretation quality score for measuring the quality of interpretability methods,” ArXiv, vol. abs/2205.12254, 2022. [Online]. Available: https://api.semanticscholar. org/CorpusID:249018098
2022 arXiv
-
[75]
Machine Learning Interpretability: A Survey on Methods and Metrics,
D. V . Carvalho, E. M. Pereira, and J. S. Cardoso, “Machine Learning Interpretability: A Survey on Methods and Metrics,” Electronics, vol. 8, no. 8, p. 832, Jul. 2019. [Online]. Available: https://www.mdpi.com/2079-9292/8/8/832
2019
-
[76]
Leakage-adjusted simulatability: Can models generate non-trivial explanations of their behavior in natural language?
P. Hase, S. Zhang, H. Xie, and M. Bansal, “Leakage-adjusted simulatability: Can models generate non-trivial explanations of their behavior in natural language?” in Findings of the Association for Computational Linguistics: EMNLP 2020 , T. Cohn, Y . He, and Y . Liu, Eds. Online...
2020
-
[77]
REV: Information-theoretic evaluation of free-text rationales,
H. Chen, F. Brahman, X. Ren, Y . Ji, Y . Choi, and S. Swayamdipta, “REV: Information-theoretic evaluation of free-text rationales,” in Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , A. Rogers, J. Boyd-Graber, a...
2023
-
[78]
A collection of principles for guiding and evaluating large language models,
K. Hebenstreit, R. Praas, and M. Samwald, “A collection of principles for guiding and evaluating large language models,” 2023. [Online]. Available: https://arxiv.org/abs/2312.10059
2023 arXiv
-
[79]
Available: https://www.oxfordlearnersdictionaries.com/
[Online]. Available: https://www.oxfordlearnersdictionaries.com/
-
[80]
Measures for explainable AI: Explanation goodness, user satisfaction, mental models, curiosity, trust, and human-AI performance,
R. R. Hoffman, S. T. Mueller, G. Klein, and J. Litman, “Measures for explainable AI: Explanation goodness, user satisfaction, mental models, curiosity, trust, and human-AI performance,” Frontiers in Computer Science , vol. 5, p. 1096257, Feb. 2023. [Online]. Available: https:/...
2023
-
[81]
Exploratory Data Analysis,
A. P ´aez and G. Boisjoly, “Exploratory Data Analysis,” in Discrete Choice Analysis with R . Cham: Springer International Publishing, 2022, pp. 25–64, series Title: Use R! [Online]. Available: https: //link.springer.com/10.1007/978-3-031-20719-8 2
2022 doi
-
[82]
Likert scales, levels of measurement and the “laws
G. Norman, “Likert scales, levels of measurement and the “laws” of statistics,” Advances in Health Sciences Education , vol. 15, no. 5, pp. 625–632, Dec. 2010. [Online]. Available: http://link.springer.com/10.1 007/s10459-010-9222-y
2010
-
[83]
Measuring the Quality of Explanations: The System Causability Scale (SCS): Comparing Human and Machine Explanations,
A. Holzinger, A. Carrington, and H. M ¨uller, “Measuring the Quality of Explanations: The System Causability Scale (SCS): Comparing Human and Machine Explanations,” KI - K ¨unstliche Intelligenz , vol. 34, no. 2, pp. 193–198, Jun. 2020. [Online]. Available: http://link.springe...
2020 doi
-
[84]
A Critique and Improvement of the
A. Vargha, H. D. Delaney, and A. Vargha, “A Critique and Improvement of the ”CL” Common Language Effect Size Statistics of McGraw and Wong,” Journal of Educational and Behavioral Statistics, vol. 25, no. 2, p. 101, 2000. [Online]. Available: http://links.jstor.org/sici?sici=10...
2000
-
[85]
H. O. Mayer, Interview und schriftliche Befragung: Entwicklung, Durchf¨uhrung und Auswertung , 4th ed., ser. 150 Jahre Wissen f ¨ur die Zukunft. M ¨unchen Wien: Oldenbourg, 2008
2008
-
[86]
Bhattacherjee, Social Science Research: Principles, Methods and Practices
A. Bhattacherjee, Social Science Research: Principles, Methods and Practices. Open Textbook Library, 01 2012
2012
-
[89]
Nonparametric Pairwise Multiple Comparisons in Independent Groups using Dunn’s Test,
A. Dinno, “Nonparametric Pairwise Multiple Comparisons in Independent Groups using Dunn’s Test,” The Stata Journal: Promoting communications on statistics and Stata , vol. 15, no. 1, pp. 292–300, Apr. 2015. [Online]. Available: http://journals.sagepub.com/doi/10.1177/1536867X1...
2015 doi
-
[2021]
Available: https://arxiv.org/abs/2106.09685
[Online]. Available: https://arxiv.org/abs/2106.09685
-
[2023]
Available: https://www.prolific.com/
[Online]. Available: https://www.prolific.com/
-
[2024]
Available: https://aclanthology.org/2024.tacl-1.85/
[Online]. Available: https://aclanthology.org/2024.tacl-1.85/
2024
Reviewed August 16, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.