REVIEW 5 major objections 5 minor 94 references
Resource-Efficient Automatic Software Vulnerability Assessment via Knowledge Distillation and Particle Swarm Optimization
T0 review · 5 major / 5 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read A particle-swarm-searched student model can score vulnerability severity at 0.6% of the teacher's size while keeping 89.3% of its accuracy.
desk verdict A workmanlike engineering paper on compressing CodeBERT for CVSS scoring with KD and PSO, but the PSO fitness proxy is unvalidated and the 'state-of-the-art' claim rests on a single BiLSTM baseline. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is a fitness function that lets particle swarm optimization search without training: it combines GFLOPs, a hardware-independent proxy for computational capacity and therefore expected performance, with the size difference between teacher and student, a proxy for the capacity gap, to score each candidate architecture. PSO then updates particle positions with an inertia weight and acceleration coefficients toward individual and global bests; the winning configuration is trained by knowledge distillation with loss $L_{\mathrm{KD}} = \alpha L_{\mathrm{CE}} + (1-\alpha)T^2 L_{\mathrm{KL}}$, so the student learns from soft teacher logits at temperature $T$ while also using ground-truth labels. This two-stage design is what turns a $476$ MB teacher into a $3$ MB student without evaluating any intermediate architecture by training.
What would settle it
Take a random sample of architectures from the same 13-dimensional search space, train and evaluate each after distillation, and rank them by real accuracy; if the PSO-selected architecture is not near the top of that ranking, then the fitness proxy does not actually find the optimal student and the reported gains are not a systematic result.
Extended reading notes
Core claim
The central claim is that a compact student model for automated vulnerability assessment need not be hand-designed: a 13-dimensional hyperparameter space (tokenizer type, vocabulary size, layer count, hidden size, and similar settings) can be searched by particle swarm optimization, and the best configuration can then be distilled from CodeBERT to reach 3 MB. On the authors' version of MegaVul, PSO-KDVA retains 89.3% of the teacher's accuracy, cuts model size to 0.6% of the original, reduces training time to 27.9%, and beats the BiLSTM distillation baseline by 1.7% accuracy at 40% of that baseline's size. The paper also reports that compression hurts Matthews correlation coefficient more than accuracy, especially on the minority "Low" severity class, so the retained accuracy is not uniform across classes.
Load-bearing premise
The whole method assumes that a student's GFLOPs and its size difference from the teacher predict which architecture will be most accurate after distillation, even though no candidate is ever trained during the search.
Editorial extensions
If this is right
- A 3 MB severity-scoring model can replace a 476 MB CodeBERT teacher, so CVSS scoring could run inside IDEs, CI pipelines, or embedded scanners.
- PSO-KDVA uses 60% fewer parameters than the BiLSTM distillation baseline while scoring 1.7% higher accuracy, so the particle swarm search buys real efficiency over a fixed hand-chosen student.
- Training the compressed model takes 27.9% of the teacher's training time, and architecture search is 34.88% faster with PSO than with a genetic algorithm.
- Compression is uneven: MCC falls 37.99% at 3 MB even though accuracy falls only 10.73%, so deployment should account for worse performance on rare severity classes.
Reading between the lines
- The paper leaves implicit that, if the GFLOPs and size-difference proxies are reliable, the same PSO-then-distill pipeline should transfer to other large code encoders beyond CodeBERT, producing similarly compact students without re-deriving a search space.
- Because the fitness function never trains a candidate, swapping GFLOPs for a cheap training-free accuracy estimator would turn the search from a heuristic into an empirically verifiable optimization.
- The large MCC drop points to a natural extension the paper names but does not implement: a class-weighted or focal distillation loss to protect minority-severity classes.
- The enhanced MegaVul subset contains C++ only, so testing PSO-KDVA on other programming languages would reveal whether the 13-dimensional search space and the fitness balance generalize beyond C++.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes PSO-KDVA, a two-stage framework for automatic severity scoring of software vulnerabilities. Stage 1 uses particle swarm optimization to search over a 13-dimensional student architecture space, with a fitness function built from GFLOPs and teacher-student model-size difference rather than from trained performance. Stage 2 distills a CodeBERT teacher into the selected student using task-specific knowledge distillation. On an enhanced MegaVul dataset of 12,071 C/C++ vulnerabilities with four CVSS v3 severity classes, the authors report compressing a 476 MB CodeBERT teacher to a 3 MB student with 89.3% of teacher accuracy, 1.7% higher accuracy than a BiLSTM baseline with 60% fewer parameters, and reductions of 72.1% in training time and 34.88% in search time relative to a genetic algorithm.
Significance. If the empirical claims hold, the paper offers a practically relevant result: a 3 MB model that retains most of CodeBERT's CVSS scoring accuracy would lower deployment barriers for vulnerability triage. The paper has several strengths that should be acknowledged: it targets a real application where model compression has been little studied; it reports MCC and minority-class accuracy in addition to overall accuracy, showing awareness of class imbalance; it includes sensitivity analyses for learning rate and distillation temperature; and it compares PSO search time against a genetic algorithm. The central limitations are that the PSO fitness function is an unvalidated proxy, the reported accuracy differences lack statistical support, and the baseline comparison is very narrow. These issues are fixable: adding proxy validation, error bars, and additional baselines would substantially strengthen the contribution. The paper does not contain internally inconsistent derivations; the contribution is empirical, and the main concerns are evidential rather than logical.
major comments (5)
- [Section 3.1.2, Fitness Function] The fitness function that drives the architecture search is never written down. The text says it should account for GFLOPs and model-size difference, and Algorithm 1 repeatedly calls Fitness(...), but no equation, weight, or normalization is given. In addition, the proxy is unvalidated: no analysis shows that higher GFLOPs or larger size difference correlates with post-distillation accuracy or MCC on this task. Since Algorithm 1 returns the proxy-optimal particle, the reported 3 MB student is a proxy-optimal architecture, not necessarily one that is optimal or even good after distillation. Please specify the exact fitness formula and report a correlation analysis between the proxy and actual student performance, or compare the PSO-selected architecture against random search and against training the top-k proxy candidates.
- [Sections 5.2 and 5.3, Experimental Results] The experimental section reports single-run point estimates without error bars, confidence intervals, or significance tests. The headline claims of 89.3% accuracy retention and 1.7% accuracy improvement over BiLSTM are differences of a few percentage points on a four-class, imbalanced problem; with no repeated seeds, it is impossible to tell whether these differences are meaningful. Please report mean and standard deviation over at least 5-10 runs with the same architecture, and state whether differences are statistically significant (e.g., paired bootstrap or a McNemar-style test).
- [Section 5.3, MCC Analysis] The accuracy-based headline understates the compression cost. The same table shows that the 3 MB student's MCC drops 37.99% relative to CodeBERT (35.53 to 22.03), while its accuracy drops only 10.73%. Given the highly imbalanced class distribution (e.g., Low has 238 training examples versus High's 4,454), accuracy is dominated by majority classes. The central claim should be rephrased to report MCC retention (about 62.0% of teacher MCC) alongside accuracy retention, and the minority-class results should be included in the abstract-level claims.
- [Section 5.2, Baseline Comparison] The comparison that supports the claim of outperforming 'state-of-the-art baselines' consists of a single baseline, BiLSTM (Tang et al., 2019). No comparisons are made to other knowledge-distillation compression methods for code models (e.g., the 3 MB code-model compression of Shi et al., TinyBERT, or DistilBERT) or to other SVA models (e.g., DeepCVA or graph-based methods). With only one baseline, the superiority claim is not established. Please add at least one additional modern baseline and temper the wording accordingly.
- [Data Availability and Reproducibility] The manuscript states that data 'will be made available on request' and does not provide a repository for the PSO-KDVA implementation, the fitness function, or the exact hyperparameters of the final student architecture. This is a reproducibility concern for an empirical systems paper. Please release the dataset and code, or a detailed configuration with all search and training settings, as part of the revision.
minor comments (5)
- [Throughout] The manuscript contains encoding artifacts such as 'BiLSTM ����', inconsistent 'PSO-KDV A' spacing, and garbled symbols in equations and tables; these should be fixed in the camera-ready version.
- [Abstract and Section 5.2] The phrase 'state-of-the-art baselines' overstates the single BiLSTM comparison; consider replacing it with 'a BiLSTM baseline' or adding the missing baselines.
- [Section 3.1.2, Fitness Function] The text says GFLOPs values range from 1 to 3 and model-size differences from 0 to 1, but no scaling or weighting of the two terms is reported; specify how they are combined numerically.
- [Introduction and Abstract] The contribution summary is repeated nearly verbatim in the abstract and the introduction; please consolidate to avoid redundancy.
- [Section 7, Related Work] The discussion of recent knowledge-distillation papers in image and multimodal tasks is tangential and not connected to the proposed SVA method; consider trimming it to keep the paper focused.
Circularity Check
No significant circularity: the accuracy and MCC results are measured after actual training on a held-out test split, not derived from the PSO fitness function.
full rationale
The paper's central claims are empirical, not derivationally circular. The PSO search in Section 3.1.2 maximizes a proxy fitness combining GFLOPs and teacher-student size difference without training or evaluating candidate student architectures; after the search, the chosen architecture is actually trained via knowledge distillation and assessed on a held-out test set. The reported 89.3% accuracy retention, the MCC values, and the 1.7% accuracy advantage over the BiLSTM baseline are measured from the trained student, so they do not reduce to the search inputs. The 99.4% model-size reduction is a direct consequence of optimizing a size-difference term in the fitness function, meaning it is more a statement of the optimization objective than an independent prediction, but the paper does not present it as a derived prediction and it is not used to manufacture the accuracy result. Hyperparameter choices such as learning rate and temperature are explored in sensitivity analyses and selected before the final comparison, which is a mild tuning practice rather than a circular fit of the headline result. The GFLOPs-based assumption that higher GFLOPs generally imply better performance is unsupported and the state-of-the-art comparison rests on a single BiLSTM baseline, but these are validity and completeness concerns, not instances of self-definition, fitted-input-as-prediction, or load-bearing self-citation. No uniqueness theorem is imported from the authors' prior work, and no cited earlier result is doing the work of the empirical evaluation.
Assumptions & free parameters
free parameters (5)
- KD temperature T =
not explicitly stated, sensitivity analysis over T=2 to 20
- learning rate =
5e-4
- KD loss weight alpha =
unspecified in text
- fitness function weights =
not shown in extracted text
- PSO hyperparameters (inertia, c1, c2, swarm size, iterations) =
unspecified ranges
assumptions (3)
- domain assumption GFLOPs is a valid monotonic proxy for student model post-distillation accuracy
- domain assumption CVSS v3 severity labels in the enhanced MegaVul dataset are accurate and representative
- domain assumption Knowledge distillation from CodeBERT logits transfers sufficient task knowledge to a small student
Cite this review
Pith. "Pith review of Resource-Efficient Automatic Software Vulnerability Assessment via Knowledge Distillation and Particle Swarm Optimization." pith.science (2026). https://pith.science/paper/3R7FHPGJ
@misc{pith2026250802840,
author = {Pith},
title = {Pith review of: Resource-Efficient Automatic Software Vulnerability Assessment via Knowledge Distillation and Particle Swarm Optimization},
year = {2026},
howpublished = {\url{https://pith.science/paper/3R7FHPGJ}},
note = {Machine review of arXiv:2508.02840}
}
read the original abstract
The increasing complexity of software systems has led to a surge in cybersecurity vulnerabilities, necessitating efficient and scalable solutions for vulnerability assessment. However, the deployment of large pre-trained models in real-world scenarios is hindered by their substantial computational and storage demands. To address this challenge, we propose a novel resource-efficient framework that integrates knowledge distillation and particle swarm optimization to enable automated vulnerability assessment. Our framework employs a two-stage approach: First, particle swarm optimization is utilized to optimize the architecture of a compact student model, balancing computational efficiency and model capacity. Second, knowledge distillation is applied to transfer critical vulnerability assessment knowledge from a large teacher model to the optimized student model. This process significantly reduces the model size while maintaining high performance. Experimental results on an enhanced MegaVul dataset, comprising 12,071 CVSS (Common Vulnerability Scoring System) v3 annotated vulnerabilities, demonstrate the effectiveness of our approach. Our approach achieves a 99.4% reduction in model size while retaining 89.3% of the original model's accuracy. Furthermore, it outperforms state-of-the-art baselines by 1.7% in accuracy with 60% fewer parameters. The framework also reduces training time by 72.1% and architecture search time by 34.88% compared to traditional genetic algorithms.
Reference graph
Works this paper leans on
-
[1]
S. Shah, B. M. Mehtre, An overview of vulnerability assessment and pen- etration testing techniques, Journal of Computer Virology and Hacking Techniques 11 (2015) 27–49
2015
-
[2]
T. H. Le, H. Chen, M. A. Babar, A survey on data-driven software vul- nerability assessment and prioritization, ACM Computing Surveys 55 (5) (2022) 1–39
2022
-
[3]
T. H. M. Le, B. Sabir, M. A. Babar, Automated software vulnerability assessment with concept drift, in: 2019 IEEE/ACM 16th International Conference on Mining Software Repositories (MSR), IEEE, 2019, pp. 371– 382
2019
-
[4]
G. Yang, Y. Zhou, X. Chen, X. Zhang, T. Han, T. Chen, Exploitgen: Template-augmented exploit code generation based on codebert, Journal of Systems and Software 197 (2023) 111577
2023
-
[5]
S. Zhou, U. Alon, S. Agarwal, G. Neubig, Codebertscore: Evaluating code generation with pretrained models of code, arXiv preprint arXiv:2302.05527 (2023)
arXiv 2023
-
[6]
Hanif, S
H. Hanif, S. Maffeis, Vulberta: Simplified source code pre-training for vul- nerability detection, in: 2022 International joint conference on neural net- works (IJCNN), IEEE, 2022, pp. 1–8
2022
-
[7]
H. Wu, Z. Zhang, S. Wang, Y. Lei, B. Lin, Y. Qin, H. Zhang, X. Mao, Peculiar: Smart contract vulnerability detection based on crucial data flow 46 graph and pre-training techniques, in: 2021 IEEE 32nd International Sym- posium on Software Reliability Engineering (ISSRE), IEEE, 2021, pp. 378– 389
2021
-
[8]
Z. Feng, D. Guo, D. Tang, N. Duan, X. Feng, M. Gong, L. Shou, B. Qin, T. Liu, D. Jiang, et al., Codebert: A pre-trained model for programming and natural languages, arXiv preprint arXiv:2002.08155 (2020)
arXiv 2020
Show all 94 references
-
[9]
G. A. Aye, G. E. Kaiser, Sequence model design for code completion in the modern ide, arXiv preprint arXiv:2004.05249 (2020)
2020 arXiv
-
[10]
Svyatkovskiy, S
A. Svyatkovskiy, S. Lee, A. Hadjitofi, M. Riechert, J. V. Franco, M. Al- lamanis, Fast and memory-efficient neural code completion, in: 2021 IEEE/ACM 18th International Conference on Mining Software Reposito- ries (MSR), IEEE, 2021, pp. 329–340
2021
-
[11]
X. Wei, S. K. Gonugondla, S. Wang, W. Ahmad, B. Ray, H. Qian, X. Li, V. Kumar, Z. Wang, Y. Tian, et al., Towards greener yet powerful code generation via quantization: An empirical study, in: Proceedings of the 31st ACM Joint European Software Engineering Conference and Sympos...
2023
-
[12]
J. Shi, Z. Yang, D. Lo, Efficient and green large language models for software engineering: Vision and the road ahead, arXiv preprint arXiv:2404.04566 (2024)
2024 arXiv
-
[13]
M. Jin, S. Shahriar, M. Tufano, X. Shi, S. Lu, N. Sundaresan, A. Svy- atkovskiy, Inferfix: End-to-end program repair with llms, in: Proceedings of the 31st ACM Joint European Software Engineering Conference and Symposium on the Foundations of Software Engineering, 2023, pp. 1646– 1656
2023
-
[14]
Y. Peng, S. Gao, C. Gao, Y. Huo, M. Lyu, Domain knowledge matters: Improving prompts with fix templates for repairing python type errors, in: 47 Proceedings of the 46th IEEE/ACM International Conference on Software Engineering, 2024, pp. 1–13
2024
-
[15]
Sch¨ afer, S
M. Sch¨ afer, S. Nadi, A. Eghbali, F. Tip, An empirical evaluation of using large language models for automated unit test generation, IEEE Transac- tions on Software Engineering (2023)
2023
-
[16]
Z. Yuan, Y. Lou, M. Liu, S. Ding, K. Wang, Y. Chen, X. Peng, No more manual tests? evaluating and improving chatgpt for unit test generation, arXiv preprint arXiv:2305.04207 (2023)
2023 arXiv
-
[17]
Clark, Electra: Pre-training text encoders as discriminators rather than generators, arXiv preprint arXiv:2003.10555 (2020)
K. Clark, Electra: Pre-training text encoders as discriminators rather than generators, arXiv preprint arXiv:2003.10555 (2020)
2020 arXiv
-
[18]
Choudhary, V
T. Choudhary, V. Mishra, A. Goswami, J. Sarangapani, A comprehensive survey on model compression and acceleration, Artificial Intelligence Re- view 53 (2020) 5113–5155
2020
-
[19]
Cheng, D
Y. Cheng, D. Wang, P. Zhou, T. Zhang, Model compression and acceler- ation for deep neural networks: The principles, progress, and challenges, IEEE Signal Processing Magazine 35 (1) (2018) 126–136
2018
-
[20]
L. Deng, G. Li, S. Han, L. Shi, Y. Xie, Model compression and hardware acceleration for neural networks: A comprehensive survey, Proceedings of the IEEE 108 (4) (2020) 485–532
2020
-
[21]
Z. Liu, M. Sun, T. Zhou, G. Huang, T. Darrell, Rethinking the value of network pruning, arXiv preprint arXiv:1810.05270 (2018)
2018 arXiv
-
[22]
T. Lin, S. U. Stich, L. Barba, D. Dmitriev, M. Jaggi, Dynamic model pruning with feedback, arXiv preprint arXiv:2006.07253 (2020)
2020 arXiv
-
[23]
Jiang, S
Y. Jiang, S. Wang, V. Valls, B. J. Ko, W.-H. Lee, K. K. Leung, L. Tassiulas, Model pruning enables efficient federated learning on edge devices, IEEE Transactions on Neural Networks and Learning Systems 34 (12) (2022) 10374–10386. 48
2022
-
[24]
J. Shi, Z. Yang, B. Xu, H. J. Kang, D. Lo, Compressing pre-trained models of code into 3 mb (ase’22). association for computing machinery, new york, ny, usa, article 24, 12 pages (2023)
2023
-
[25]
Polino, R
A. Polino, R. Pascanu, D. Alistarh, Model compression via distillation and quantization, arXiv preprint arXiv:1802.05668 (2018)
2018 arXiv
-
[26]
A. Fan, P. Stock, B. Graham, E. Grave, R. Gribonval, H. Jegou, A. Joulin, Training with quantization noise for extreme model compression, arXiv preprint arXiv:2004.07320 (2020)
2020 arXiv
-
[27]
Ganesh, Y
P. Ganesh, Y. Chen, X. Lou, M. A. Khan, Y. Yang, H. Sajjad, P. Nakov, D. Chen, M. Winslett, Compressing large-scale transformer-based models: A case study on bert, Transactions of the Association for Computational Linguistics 9 (2021) 1061–1080
2021
-
[28]
W. Park, D. Kim, Y. Lu, M. Cho, Relational knowledge distillation, in: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2019, pp. 3967–3976
2019
-
[29]
J. Gou, B. Yu, S. J. Maybank, D. Tao, Knowledge distillation: A survey, International Journal of Computer Vision 129 (6) (2021) 1789–1819
2021
-
[30]
B. Zhao, Q. Cui, R. Song, Y. Qiu, J. Liang, Decoupled knowledge distil- lation, in: Proceedings of the IEEE/CVF Conference on computer vision and pattern recognition, 2022, pp. 11953–11962
2022
-
[31]
J. Shi, Z. Yang, H. J. Kang, B. Xu, J. He, D. Lo, Greening large language models of code, in: Proceedings of the 46th International Conference on Software Engineering: Software Engineering in Society, 2024, pp. 142–153
2024
-
[32]
M. Gao, Y. Wang, L. Wan, Residual error based knowledge distillation, Neurocomputing 433 (2021) 154–161
2021
-
[33]
C. Ni, L. Shen, X. Yang, Y. Zhu, S. Wang, Megavul: Ac/c++ vulnerability dataset with comprehensive code representations, in: 2024 IEEE/ACM 21st 49 International Conference on Mining Software Repositories (MSR), IEEE, 2024, pp. 738–742
2024
-
[34]
R. Tang, Y. Lu, L. Liu, L. Mou, O. Vechtomova, J. Lin, Distilling task- specific knowledge from bert into simple neural networks, arXiv preprint arXiv:1903.12136 (2019)
2019 arXiv
-
[35]
N. Jain, U. Nangia, J. Jain, A review of particle swarm optimization, Jour- nal of The Institution of Engineers (India): Series B 99 (2018) 407–411
2018
-
[36]
T. M. Shami, A. A. El-Saleh, M. Alswaitti, Q. Al-Tashi, M. A. Summakieh, S. Mirjalili, Particle swarm optimization: A comprehensive survey, Ieee Access 10 (2022) 10031–10061
2022
-
[37]
Prajapati, J
A. Prajapati, J. K. Chhabra, A particle swarm optimization-based heuristic for software module clustering problem, Arabian Journal for Science and Engineering 43 (12) (2018) 7083–7094
2018
-
[38]
Lambora, K
A. Lambora, K. Gupta, K. Chopra, Genetic algorithm-a literature review, in: 2019 international conference on machine learning, big data, cloud and parallel computing (COMITCon), IEEE, 2019, pp. 380–384
2019
-
[39]
Mirjalili, S
S. Mirjalili, S. Mirjalili, Genetic algorithm, Evolutionary algorithms and neural networks: theory and applications (2019) 43–55
2019
-
[40]
X. Du, B. Chen, Y. Li, J. Guo, Y. Zhou, Y. Liu, Y. Jiang, Leopard: Iden- tifying vulnerable code for vulnerability assessment through program met- rics, in: 2019 IEEE/ACM 41st International Conference on Software Engi- neering (ICSE), IEEE, 2019, pp. 60–71
2019
-
[41]
Humayun, N
M. Humayun, N. Jhanjhi, M. F. Almufareh, M. I. Khalil, Security threat and vulnerability assessment and measurement in secure software develop- ment, Comput. Mater. Contin 71 (2022) 5039–5059
2022
-
[42]
Common Vulnerability Scoring System
2024. Common Vulnerability Scoring System. ���������������������� ����� . 50
2024
-
[43]
H. Holm, K. K. Afridi, An expert-based investigation of the common vul- nerability scoring system, Computers & Security 53 (2015) 18–30
2015
-
[44]
A. Beck, S. Rass, Using neural networks to aid cvss risk aggregation—an empirically validated approach, Journal of Innovation in Digital Ecosystems 3 (2) (2016) 148–154
2016
-
[45]
K. Liu, Y. Zhou, Q. Wang, X. Zhu, Vulnerability severity prediction with deep neural network, in: 2019 5th international conference on big data and information analytics (BigDIA), IEEE, 2019, pp. 114–119
2019
-
[46]
Babalau, D
I. Babalau, D. Corlatescu, O. Grigorescu, C. Sandescu, M. Dascalu, Sever- ity prediction of software vulnerabilities based on their text description, in: 2021 23rd International Symposium on Symbolic and Numeric Algorithms for Scientific Computing (SYNASC), IEEE, 2021, pp. 171–177
2021
-
[47]
Ganesh, T
S. Ganesh, T. Ohlsson, F. Palma, Predicting security vulnerabilities using source code metrics, in: 2021 Swedish workshop on data science (SweDS), IEEE, 2021, pp. 1–7
2021
-
[48]
M. D. Purba, A. Ghosh, B. J. Radford, B. Chu, Software vulnerability de- tection using large language models, in: 2023 IEEE 34th International Sym- posium on Software Reliability Engineering Workshops (ISSREW), IEEE, 2023, pp. 112–119
2023
-
[49]
Khanfir, M
A. Khanfir, M. Jimenez, M. Papadakis, Y. Le Traon, Codebert-nt: code naturalness via codebert, in: 2022 IEEE 22nd International Conference on Software Quality, Reliability and Security (QRS), IEEE, 2022, pp. 936–947
2022
-
[50]
Sonnekalb, B
T. Sonnekalb, B. Gruner, C.-A. Brust, P. M¨ ader, Generalizability of code clone detection on codebert, in: Proceedings of the 37th IEEE/ACM Inter- national Conference on Automated Software Engineering, 2022, pp. 1–3
2022
-
[51]
Khajezade, J
M. Khajezade, J. J. Wu, F. H. Fard, G. Rodr ´ ıguez-P´ erez, M. S. Shehata, Investigating the efficacy of large language models for code clone detec- 51 tion, in: Proceedings of the 32nd IEEE/ACM International Conference on Program Comprehension, 2024, pp. 161–165
2024
-
[52]
Nguyen, D
V.-A. Nguyen, D. Q. Nguyen, V. Nguyen, T. Le, Q. H. Tran, D. Phung, Regvd: Revisiting graph neural networks for vulnerability detection, in: Proceedings of the ACM/IEEE 44th International Conference on Software Engineering: Companion Proceedings, 2022, pp. 178–182
2022
-
[53]
X. Zhou, T. Zhang, D. Lo, Large language model for vulnerability detec- tion: Emerging results and future directions, in: Proceedings of the 2024 ACM/IEEE 44th International Conference on Software Engineering: New Ideas and Emerging Results, 2024, pp. 47–51
2024
-
[54]
Mukherjee, A
P. Mukherjee, A. Das, A. K. Bhunia, P. P. Roy, Cogni-net: Cognitive fea- ture learning through deep visual perception, in: 2019 IEEE International Conference on Image Processing (ICIP), IEEE, 2019, pp. 4539–4543
2019
-
[55]
Z. Peng, Z. Li, J. Zhang, Y. Li, G.-J. Qi, J. Tang, Few-shot image recogni- tion with knowledge transfer, in: Proceedings of the IEEE/CVF interna- tional conference on computer vision, 2019, pp. 441–449
2019
-
[56]
Y. Pan, F. He, H. Yu, A novel enhanced collaborative autoencoder with knowledge distillation for top-n recommender systems, Neurocomputing 332 (2019) 137–148
2019
-
[57]
X. Chen, Y. Zhang, H. Xu, Z. Qin, H. Zha, Adversarial distillation for efficient recommendation with external knowledge, ACM Transactions on Information Systems (TOIS) 37 (1) (2018) 1–28
2018
-
[58]
Y. Bai, J. Yi, J. Tao, Z. Tian, Z. Wen, Learn spelling from teachers: Trans- ferring knowledge from language models to sequence-to-sequence speech recognition, arXiv preprint arXiv:1907.06017 (2019)
2019 arXiv
-
[59]
R. W. Ng, X. Liu, P. Swietojanski, Teacher-student training for text- independent speaker recognition, in: 2018 IEEE Spoken Language Tech- nology Workshop (SLT), IEEE, 2018, pp. 1044–1051. 52
2018
-
[60]
W. Wang, F. Wei, L. Dong, H. Bao, N. Yang, M. Zhou, Minilm: Deep self- attention distillation for task-agnostic compression of pre-trained trans- formers, Advances in Neural Information Processing Systems 33 (2020) 5776–5788
2020
-
[61]
Liang, H
C. Liang, H. Jiang, Z. Li, X. Tang, B. Yin, T. Zhao, Homodistil: Homo- topic task-agnostic distillation of pre-trained transformers, arXiv preprint arXiv:2302.09632 (2023)
2023 arXiv
-
[62]
W.-H. Li, H. Bilen, Knowledge distillation for multi-task learning, in: Com- puter Vision–ECCV 2020 Workshops: Glasgow, UK, August 23–28, 2020, Proceedings, Part VI 16, Springer, 2020, pp. 163–176
2020
-
[63]
J. H. Cho, B. Hariharan, On the efficacy of knowledge distillation, in: Proceedings of the IEEE/CVF international conference on computer vision, 2019, pp. 4794–4802
2019
-
[64]
Kennedy, R
J. Kennedy, R. Eberhart, Particle swarm optimization, in: Proceedings of ICNN’95-international conference on neural networks, Vol. 4, ieee, 1995, pp. 1942–1948
1995
-
[65]
C. Wu, M. Wang, X. Chu, K. Wang, L. He, Low-precision floating-point arithmetic for high-performance fpga-based cnn acceleration, ACM Trans- actions on Reconfigurable Technology and Systems (TRETS) 15 (1) (2021) 1–21
2021
-
[66]
H. Zhou, L. Song, J. Chen, Y. Zhou, G. Wang, J. Yuan, Q. Zhang, Rethink- ing soft labels for knowledge distillation: A bias-variance tradeoff perspec- tive, arXiv preprint arXiv:2102.00650 (2021)
2021 arXiv
-
[67]
Aguilar, Y
G. Aguilar, Y. Ling, Y. Zhang, B. Yao, X. Fan, C. Guo, Knowledge distilla- tion from internal representations, in: Proceedings of the AAAI conference on artificial intelligence, Vol. 34, 2020, pp. 7350–7357
2020
-
[68]
B. Zi, S. Zhao, X. Ma, Y.-G. Jiang, Revisiting adversarial robustness distillation: Robust soft labels make student better, in: Proceedings of 53 the IEEE/CVF International Conference on Computer Vision, 2021, pp. 16443–16452
2021
-
[69]
W. Liu, P. Zhou, Z. Zhao, Z. Wang, H. Deng, Q. Ju, Fast- bert: a self-distilling bert with adaptive inference time, arXiv preprint arXiv:2004.02178 (2020)
2020 arXiv
-
[70]
Vilar, M
D. Vilar, M. Federico, A statistical extension of byte-pair encoding, in: Pro- ceedings of the 18th International Conference on Spoken Language Trans- lation (IWSLT 2021), 2021, pp. 263–275
2021
-
[71]
C. Pan, M. Lu, B. Xu, An empirical study on software defect prediction using codebert model, Applied Sciences 11 (11) (2021) 4793
2021
-
[72]
Kudo, Subword regularization: Improving neural network translation models with multiple subword candidates, arXiv preprint arXiv:1804.10959 (2018)
T. Kudo, Subword regularization: Improving neural network translation models with multiple subword candidates, arXiv preprint arXiv:1804.10959 (2018)
2018 arXiv
-
[73]
Karampatsis, H
R.-M. Karampatsis, H. Babii, R. Robbes, C. Sutton, A. Janes, Big code!= big vocabulary: Open-vocabulary models for source code, in: Proceedings of the ACM/IEEE 42nd International Conference on Software Engineering, 2020, pp. 1073–1085
2020
-
[74]
Jiang, Z
X. Jiang, Z. Zheng, C. Lyu, L. Li, L. Lyu, Treebert: A tree-based pre- trained model for programming language, in: Uncertainty in Artificial In- telligence, PMLR, 2021, pp. 54–63
2021
-
[75]
A. A. Ishtiaq, M. Hasan, M. M. A. Haque, K. S. Mehrab, T. Muttaqueen, T. Hasan, A. Iqbal, R. Shahriyar, Bert2code: Can pretrained language models be leveraged for code search?, arXiv preprint arXiv:2104.08017 (2021)
2021 arXiv
-
[76]
Elfwing, E
S. Elfwing, E. Uchibe, K. Doya, Sigmoid-weighted linear units for neural network function approximation in reinforcement learning, Neural networks 107 (2018) 3–11. 54
2018
-
[77]
Mashhadi, H
E. Mashhadi, H. Hemmati, Applying codebert for automated program re- pair of java simple bugs, in: 2021 IEEE/ACM 18th International Confer- ence on Mining Software Repositories (MSR), IEEE, 2021, pp. 505–509
2021
-
[78]
X. Zhou, D. Han, D. Lo, Assessing generalizability of codebert, in: 2021 IEEE International Conference on Software Maintenance and Evolution (ICSME), IEEE, 2021, pp. 425–436
2021
-
[79]
Sharma, F
R. Sharma, F. Chen, F. Fard, D. Lo, An exploratory study on code atten- tion in bert, in: Proceedings of the 30th IEEE/ACM International Confer- ence on Program Comprehension, 2022, pp. 437–448
2022
-
[80]
L. Yang, Z. Li, D. Wang, H. Miao, Z. Wang, Software defects prediction based on hybrid particle swarm optimization and sparrow search algorithm, Ieee Access 9 (2021) 60865–60879
2021
-
[81]
Hinton, Distilling the knowledge in a neural network, arXiv preprint arXiv:1503.02531 (2015)
G. Hinton, Distilling the knowledge in a neural network, arXiv preprint arXiv:1503.02531 (2015)
2015 arXiv
-
[82]
Radosavovic, P
I. Radosavovic, P. Doll´ ar, R. Girshick, G. Gkioxari, K. He, Data distillation: Towards omni-supervised learning, in: Proceedings of the IEEE conference on computer vision and pattern recognition, 2018, pp. 4119–4128
2018
-
[83]
X. Jiao, Y. Yin, L. Shang, X. Jiang, X. Chen, L. Li, F. Wang, Q. Liu, Tiny- bert: Distilling bert for natural language understanding, arXiv preprint arXiv:1909.10351 (2019)
2019 arXiv
-
[84]
T. H. M. Le, D. Hin, R. Croft, M. A. Babar, Deepcva: Automated commit- level vulnerability assessment with deep multi-task learning, in: 2021 36th IEEE/ACM International Conference on Automated Software Engineering (ASE), IEEE, 2021, pp. 717–729
2021
-
[85]
Rajbhandari, J
S. Rajbhandari, J. Rasley, O. Ruwase, Y. He, Zero: Memory optimizations toward training trillion parameter models, in: SC20: International Confer- ence for High Performance Computing, Networking, Storage and Analysis, IEEE, 2020, pp. 1–16. 55
2020
-
[86]
Gorodkin, Comparing two k-category assignments by a k-category cor- relation coefficient, Computational biology and chemistry 28 (5-6) (2004) 367–374
J. Gorodkin, Comparing two k-category assignments by a k-category cor- relation coefficient, Computational biology and chemistry 28 (5-6) (2004) 367–374
2004
-
[87]
Z. Han, X. Li, Z. Xing, H. Liu, Z. Feng, Learning to predict severity of software vulnerability using only vulnerability description, in: 2017 IEEE International conference on software maintenance and evolution (ICSME), IEEE, 2017, pp. 125–136
2017
-
[88]
Spanos, L
G. Spanos, L. Angelis, A multi-target approach to estimate software vulner- ability characteristics and severity scores, Journal of Systems and Software 146 (2018) 152–166
2018
-
[89]
T. H. M. Le, M. A. Babar, On the use of fine-grained vulnerable code statements for software vulnerability assessment models, in: Proceedings of the 19th International Conference on Mining Software Repositories, 2022, pp. 621–633
2022
-
[90]
J. Hao, S. Luo, L. Pan, A novel vulnerability severity assessment method for source code based on a graph neural network, Information and Software Technology 161 (2023) 107247
2023
-
[91]
W. Zhou, F. Sun, Q. Jiang, R. Cong, J.-N. Hwang, Wavenet: Wavelet network with knowledge distillation for rgb-t salient object detection, IEEE Transactions on Image Processing 32 (2023) 3027–3039
2023
-
[92]
W. Zhou, X. Yang, W. Yan, Q. Jiang, Hybrid knowledge distillation for rgb- t crowd density estimation in smart surveillance systems, IEEE Internet of Things Journal (2024)
2024
-
[93]
W. Zhou, Y. Wang, X. Qian, Knowledge distillation and contrastive learn- ing for detecting visible-infrared transmission lines using separated stagger registration network, IEEE Transactions on Circuits and Systems I: Regu- lar Papers (2025). 56
2025
-
[94]
W. Zhou, B. Jian, Y. Liu, Feature contrast difference and enhanced network for rgb-d indoor scene classification in internet of things, IEEE Internet of Things Journal (2025). Chaoyang Gao is currently pursuing the Master degree at the School of Ar- tificial Intelligence and C...
2025
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.