REVIEW 4 major objections 5 minor 33 references
A weighted scoring model for penetration-testing tools ranks BeEF plus Metasploit as the most suitable combination for system penetration testing.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
A survey-style paper that proposes a weighted tool-scoring rubric, then demonstrates routine host and web penetration tests on vulnerable virtual machines.
T0 review reviewed 2026-08-04 challenge →
load-bearing objection A competent practical survey that overclaims its tool-selection model: the BeEF+Metasploit recommendation is a tautology of the author-assigned capability matrix, not a validated result. the 4 major comments →
A Comprehensive Evaluation and Practice of System Penetration Testing
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
Core claim
The paper's central claim is that tool selection in penetration testing can be made objective by scoring each tool with a weighted sum over nine functional capabilities: host scanning, password cracking, web scanning, social engineering, vulnerability discovery, exploit, session control, report generation, and visualization. The authors define O = VωV + EωE + WωW + HωH + PωP + CωC + RωR + SωS + IωI and propose three weight schemes for routine testing, enterprise internal-network assessment, and red-team exercises. Applying the formula to a capability matrix for mainstream tools, they find the combination of BeEF and Metasploit yields the highest score and recommend it. The claim is accompani
What carries the argument
The weighted-suitability formula O = VωV + EωE + WωW + HωH + PωP + CωC + RωR + SωS + IωI, together with a nine-function capability matrix (Table 10) and the three scenario-specific weight tables (Tables 7–9). The formula converts each tool's binary function presence into a single suitability score; the capability matrix and weights are the inputs that make the ranking, so the recommendation (BeEF + Metasploit) follows arithmetically.
Load-bearing premise
The load-bearing premise is that the nine weights in the three schemes and the binary capability scores in Table 10 reflect how much each function actually matters in practice; if those numbers are arbitrary, the ranking and the BeEF + Metasploit recommendation are arbitrary.
What would settle it
Recompute O scores with an independently derived weight set or with the capability matrix filled from official tool documentation; if BeEF + Metasploit is not highest, the claim collapses. Or run BeEF + Metasploit against the next two ranked combinations on identical target ranges and measure success rate, time, and coverage; a lower success rate would contradict the suitability claim.
If this is right
- Practitioners can use the three weight schemes to score and compare tools for routine testing, enterprise intranet assessments, and red-team exercises.
- Combining BeEF with Metasploit is presented as the highest-scoring tool pairing, giving testers a concrete default combination.
- The six-phase process (preparation and information gathering, vulnerability detection, post-testing, reporting, and retesting) gives teams a standardized workflow with a closed-loop retest.
- The experimental reproductions show that the process and tools can be applied to real target machines for Windows and Linux host attacks and web vulnerabilities like SQL injection and file upload.
Where Pith is reading between the lines
- The ranking's objectivity is inherited from the weights; a sensitivity analysis varying Tables 7–9 would show whether BeEF + Metasploit's top position is stable or a byproduct of one scheme.
- The same weighted-sum scaffolding could serve as a scoring function for automated tool orchestration, letting a planner choose per-phase tools by maximizing O under scenario constraints.
- Comparing predicted O values with actual metrics from the paper's own experiments (time to compromise, number of shells, coverage) would test whether 'suitability' tracks performance.
- The capability matrix is binary (has/does not have); grading capabilities (e.g., depth of SQLi support) would likely change rankings and is a natural extension.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper presents a broad survey and practice-oriented account of system penetration testing. It proposes a six-phase penetration testing process, describes a set of mainstream tools (Nmap, Masscan, Nessus, OpenVAS, Metasploit, Burp Suite, SQLMap, etc.), and introduces a quantitative tool-selection model in which a weighted sum of nine binary capability indicators, O = VωV + EωE + WωW + HωH + PωP + CωC + RωR + SωS + IωI, is used as a suitability score. Three weighting schemes are proposed for different testing scenarios. The paper then reports host and web penetration experiments in a virtualized lab and discusses several real-world security incidents. The headline conclusion is that combining BeEF and Metasploit yields the highest O value and is therefore a good tool combination for penetration testing.
Significance. If the quantitative tool-selection model were sound, it would provide a useful reference for practitioners choosing and combining penetration testing tools. The paper also contains useful descriptive material: the process outline, the tool summaries, the reproduced attacks against Windows and Linux targets, and the case-study lessons are generally accurate and readable. However, the central quantitative contribution is not currently supported. The weighting schemes are asserted without external justification, the capability matrix in Table 10 is undocumented, and the top ranking of BeEF+Metasploit follows by construction from that matrix under any positive weights. The experimental sections do not test BeEF or the recommended combination, so the model is neither calibrated nor validated. These issues are load-bearing for the paper's stated contribution of providing 'objective reference criteria' for tool selection, and they require substantive revision rather than copy-editing.
major comments (4)
- [§4.2, Tables 7–9] The three weighting schemes are presented as reflecting 'industry standards, common security threats, and their impact,' but no citation, survey, expert elicitation, or empirical calibration is provided. Since the O-score and the resulting tool ranking are direct arithmetic consequences of these weights, the claimed objectivity of the model rests on an unsupported premise. Please justify the weights from a documented source, or reframe the schemes as illustrative examples and remove the objectivity claim.
- [§4.2, Table 10 and Figure 5] The capability matrix appears to be author-constructed, but no methodology, source, or audit trail is given for the binary entries. More importantly, the combined row 'BeEF & Metasploit' is the only entry with all nine capabilities. Since every weight in each of the three schemes is positive and the weights sum to 100%, that row necessarily receives O = 100%, the maximum possible score, while every other row is missing at least one capability and scores strictly lower. Therefore the conclusion that 'the weighted sum for combining BeEF and Metasploit is the highest' is independent of the three weighting schemes and is a tautology of the matrix entries. The paper needs to document the matrix's provenance, provide a sensitivity analysis, and ideally benchmark the matrix against independent tool-capability data before using it to support a recommendation.
- [§4.2–§4.4] The central claim that 'a higher O value indicates a greater suitability for system penetration testing' is asserted, not demonstrated, and the experimental sections do not test the model's predictive validity. The host and web experiments use Nmap, Nessus, Metasploit, Hydra, Burp Suite, and DVWA, but no experiment involves BeEF or the BeEF+Metasploit combination. Consequently, the statement that 'these two tools have also delivered solid results in practical applications' is unsupported by the evidence in this paper. Please either validate the model against actual penetration outcomes or substantially weaken the claimed conclusion.
- [§6; §4.2 linear-additivity assumption] The suitability model assumes that a tool's overall value is a linearly additive weighted sum of binary indicators. This ignores interactions between capabilities, diminishing returns, and the possibility that a tool may have a capability only partially or in a context-dependent manner. The limitations paragraph in §6 acknowledges restricted target machines and DVWA difficulty levels, but it does not address this modeling assumption. The paper should discuss or test the linear-additivity assumption, or explicitly frame the O-score as a crude screening heuristic rather than an objective reference criterion.
minor comments (5)
- [Throughout] Typographical and formatting issues: 'OpenV AS' should be 'OpenVAS'; 'Creenbone' should be 'Greenbone'; 'W AF' should be 'WAF' in §4.1.3; 'permission allocation' in §4.2 should be 'weight allocation.'
- [Tables 6 and 13] The SMB port is consistently written as '455' rather than '445'. Please correct this in both the Nmap/Masscan comparison and the port-scan table.
- [Figure 5] The bar chart for weighted sums would benefit from axis labels, a legend identifying the three schemes, and a caption that explains the data source for the plotted values.
- [Introduction, reference [18]] Reference [18] is cited for State Grid's machine-learning-based automated penetration testing, but the cited work is 'Standard Penetration Test State-of-the-Art Report' by Ivan K. Nixon, which appears to be a geotechnical engineering source. Please verify and replace this citation with a relevant reference.
- [§5] The case studies are informative but are not connected to the proposed tool-selection model or the six-phase process. A sentence in each case study linking the lessons learned to the proposed framework would strengthen the paper's coherence.
Circularity Check
No significant circularity: the weighted-sum ranking is a transparent definitional rubric; unsupported weights are a validity issue, not circularity.
full rationale
The central quantitative claim is a scoring model, not a fitted prediction. The paper defines O = V·ω_V + E·ω_E + ... in §4.2 and states that a higher O indicates greater suitability. The conclusion that the BeEF & Metasploit combination has the highest weighted sum follows arithmetically from the author-assigned capability matrix (Table 10) and the positive weights in Tables 7–9: that row contains all nine capabilities, so under any positive weights it receives the maximum possible score. This is a direct consequence of the model's definitions, but it is not circular in the sense required here: no capability value or weight is derived from the conclusion, and the conclusion is not used to define the score. The model is transparently a rubric; the real weakness is that the weights are asserted to be 'based on industry standards, common security threats, and their impact' with no citation or empirical calibration, and the capability assignments in Table 10 are unaudited. Those are validity/correctness limitations, not circularity. The paper also contains many self-citations, but they support background material and threat taxonomies and are not load-bearing for the tool-ranking formula; there is no self-citation chain that forces the BeEF+Metasploit result, and no uniqueness theorem is imported from the authors' prior work. The experimental shortcomings acknowledged in §6 — restricted target machines and DVWA levels — weaken external validation but do not create an equation-level feedback loop. Because no claim reduces by construction to its own inputs or to a self-citation, the appropriate finding is no significant circularity.
Axiom & Free-Parameter Ledger
free parameters (4)
- Scheme 1 weight vector (ω_V=20%, ω_E=18%, ω_W=15%, ω_H=12%, ω_P=10%, ω_C=8%, ω_R=7%, ω_S=6%, ω_I=4%) =
sums to 100%
- Scheme 2 weight vector (ω_V=25%, ω_W=20%, ω_R=15%, ω_H=15%, ω_P=10%, ω_E=8%, other=7%) =
sums to 100%
- Scheme 3 weight vector (ω_E=25%, ω_S=20%, ω_C=15%, ω_V=15%, ω_P=10%, ω_I=5%, other=10%) =
sums to 100%
- Tool capability matrix (Table 10 binary entries) =
0/1 assignments for each tool across 9 functions
axioms (3)
- domain assumption The nine-category function decomposition (Host Scanning, Password Cracking, Web Scanning, Social Engineering, Vulnerability Discovery, Exploit, Session Control, Report Generation, Visualization Interface) is a complete and meaningful representation of penetration-testing tool capability.
- domain assumption The weight allocations reflect 'industry standards, common security threats, and their impact'.
- domain assumption The six-phase penetration testing process is appropriate and standard.
Cite this review
Pith. "Pith review of A Comprehensive Evaluation and Practice of System Penetration Testing." pith.science (2026). https://pith.science/paper/26RZXOIV
@misc{pith2026251026555,
author = {Pith},
title = {Pith review of: A Comprehensive Evaluation and Practice of System Penetration Testing},
year = {2026},
howpublished = {\url{https://pith.science/paper/26RZXOIV}},
note = {Machine review of arXiv:2510.26555}
}
read the original abstract
With the rapid advancement of information technology, the complexity of applications continues to increase, and the cybersecurity challenges we face are also escalating. This paper aims to investigate the methods and practices of system security penetration testing, exploring how to enhance system security through systematic penetration testing processes and technical approaches. It also examines existing penetration tools, analyzing their strengths, weaknesses, and applicable domains to guide penetration testers in tool selection. Furthermore, based on the penetration testing process outlined in this paper, appropriate tools are selected to replicate attack processes using target ranges and target machines. Finally, through practical case analysis, lessons learned from successful attacks are summarized to inform future research.
Figures
Reference graph
Works this paper leans on
-
[1]
A systematic literature review on penetration testing in networks: future research directions.Applied Sciences, 13(12):6986, 2023
Mariam Alhamed and MM Hafizur Rahman. A systematic literature review on penetration testing in networks: future research directions.Applied Sciences, 13(12):6986, 2023
2023
-
[2]
Rifqi Azis and Setiadi Yazid. Pengujian kerentanan website wordpress dengan menggunakan penetration testing untuk menghasilkan website yang aman.Jurnal Restikom: Riset Teknik Informatika Dan Komputer, 3(3):93–105, 2021
2021
-
[3]
An overview of penetration testing.International Journal of Network Security & Its Applications, 3(6):19, 2011
Aileen G Bacudio, Xiaohong Yuan, Bei-Tseng Bill Chu, and Monique Jones. An overview of penetration testing.International Journal of Network Security & Its Applications, 3(6):19, 2011
2011
-
[4]
In 33rd USENIX Security Symposium (USENIX Security 24), pages 847–864, 2024
Gelei Deng, Yi Liu, V ´ ıctor Mayoral-Vilches, Peng Liu, Yuekang Li, Yuan Xu, Tianwei Zhang, Yang Liu, Martin Pinzger, and Stefan Rass.{PentestGPT}: Eval- uating and harnessing large language models for automated penetration testing. In 33rd USENIX Security Symposium (USENIX Security 24), pages 847–864, 2024
2024
-
[5]
Vulnerability assessment and penetration testing
Dr Jason Edwards. Vulnerability assessment and penetration testing. InMastering cybersecurity: Strategies, technologies, and best practices, pages 371–412. Springer, 2024
2024
-
[6]
Penetration testing operating systems: Exploiting vulnerabilities
Evan Gardner, Gurmeet Singh, and Weihao Qu. Penetration testing operating systems: Exploiting vulnerabilities. In2024 International Conference on Commu- nications, Computing, Cybersecurity, and Informatics (CCCI), pages 1–9. IEEE, 2024
2024
-
[7]
Information security based on llm approaches: A review.arXiv preprint arXiv:2507.18215, 2025
Chang Gong, Zhongwen Li, and Xiaoqi Li. Information security based on llm approaches: A review.arXiv preprint arXiv:2507.18215, 2025
arXiv 2025
-
[8]
Getting pwn’d by ai: Penetration testing with large language models
Andreas Happe and J¨ urgen Cito. Getting pwn’d by ai: Penetration testing with large language models. InProceedings of the 31st ACM joint european software 32 engineering conference and symposium on the foundations of software engineering, pages 2082–2086, 2023
2082
-
[9]
Comparative analysis of blockchain systems.arXiv preprint arXiv:2505.08652, 2025
Jiaqi Huang, Yuanzheng Niu, Xiaoqi Li, and Zongwei Li. Comparative analysis of blockchain systems.arXiv preprint arXiv:2505.08652, 2025
Pith/arXiv arXiv 2025
-
[10]
Research on penetration testing pro- cedures based on kali system
Zhan Jiayan, Ma Haifei, and Chen Gengjie. Research on penetration testing pro- cedures based on kali system. In2023 4th International Conference on Computers and Artificial Intelligence Technology (CAIT), pages 271–276. IEEE, 2023
2023
-
[11]
Dechao Kong, Xiaoqi Li, and Wenkai Li. Uechecker: Detecting unchecked external call vulnerabilities in dapps via graph analysis.arXiv preprint arXiv:2508.01343, 2025
arXiv 2025
-
[12]
Research on evaluation index system for software vulnerability analysis methods
Jin Li, Min-Huan Huang, Shuai-Bing Lu, Hu Li, and Jin-Fu Chen. Research on evaluation index system for software vulnerability analysis methods. In2019 IEEE Fourth International Conference on Data Science in Cyberspace (DSC), pages 522–
-
[13]
Interaction-aware vulner- ability detection in smart contract bytecodes.IEEE Transactions on Dependable and Secure Computing, 2025
Wenkai Li, Xiaoqi Li, Yingjie Mao, and Yuqing Zhang. Interaction-aware vulner- ability detection in smart contract bytecodes.IEEE Transactions on Dependable and Secure Computing, 2025
2025
-
[14]
Haiyang Liu, Yingjie Mao, and Xiaoqi Li. An empirical analysis of eos blockchain: Architecture, contract, and security.arXiv preprint arXiv:2505.15051, 2025
Pith/arXiv arXiv 2025
-
[15]
Yuhe Luo, Zhongwen Li, and Xiaoqi Li. Movescanner: Analysis of security risks of move smart contracts.arXiv preprint arXiv:2508.17964, 2025
arXiv 2025
-
[16]
A trusted os penetration testing scheme based on metasploit and beef
Hengji Miao, Lei Shang, Weihong Gan, Chenle Zhang, Zhuo Guan, and Zhihong Ge. A trusted os penetration testing scheme based on metasploit and beef. In2024 4th International Conference on Blockchain Technology and Information Security (ICBCTIS), pages 278–282. IEEE, 2024
2024
-
[17]
Natlm: Detecting defects in nft smart contracts leveraging llm.arXiv preprint arXiv:2508.01351, 2025
Yuanzheng Niu, Xiaoqi Li, and Wenkai Li. Natlm: Detecting defects in nft smart contracts leveraging llm.arXiv preprint arXiv:2508.01351, 2025
arXiv 2025
-
[18]
Standard penetration test state-of-the-art report
Ivan K Nixon. Standard penetration test state-of-the-art report. InPenetration Testing, volume 1, pages 3–22. Routledge, 2021
2021
-
[19]
Vulnerability assess- ment and penetration testing: a portable solution implementation
Rajiv Pandey, Vutukuru Jyothindar, and Umesh K Chopra. Vulnerability assess- ment and penetration testing: a portable solution implementation. In2020 12th International Conference on Computational Intelligence and Communication Net- works (CICN), pages 398–402. IEEE, 2020
2020
-
[20]
Performance analysis of vul- nerability detection tools and techniques
Kailash Kumar Pareek and Gaurav Kumar Ameta. Performance analysis of vul- nerability detection tools and techniques. In2024 Parul International Conference on Engineering and Technology (PICET), pages 1–5. IEEE, 2024
2024
-
[21]
Mining characteristics of vulnerable smart contracts across lifecycle stages.IET Blockchain, 5(1):e70016, 2025
Hongli Peng, Wenkai Li, and Xiaoqi Li. Mining characteristics of vulnerable smart contracts across lifecycle stages.IET Blockchain, 5(1):e70016, 2025. 33
2025
-
[22]
Hongli Peng, Xiaoqi Li, and Wenkai Li. Multicfv: Detecting control flow vulner- abilities in smart contracts leveraging multimodal deep learning.arXiv preprint arXiv:2508.01346, 2025
arXiv 2025
-
[23]
Chengxin Shen, Zhongwen Li, Xiaoqi Li, and Zongwei Li. When blockchain meets crawlers: Real-time market analytics in solana nft markets.arXiv preprint arXiv:2506.02892, 2025
arXiv 2025
-
[24]
Analysis of vulnera- bility characteristics for automated penetration testing
Yaroslav Stefinko, Andrian Piskozub, and Anatolii Obshta. Analysis of vulnera- bility characteristics for automated penetration testing. In2024 IEEE 17th Inter- national Conference on Advanced Trends in Radioelectronics, Telecommunications and Computer Engineering (TCSET), pages 449–453. IEEE, 2024
2024
-
[25]
Ai-based vulnerability analysis of nft smart contracts
Xin Wang and Xiaoqi Li. Ai-based vulnerability analysis of nft smart contracts. arXiv preprint arXiv:2504.16113, 2025
arXiv 2025
-
[26]
Elsevier, 2025
Thomas Wilhelm.Professional penetration testing: Creating and learning in a hacking lab. Elsevier, 2025
2025
-
[27]
Security analysis of chatgpt: Threats and privacy risks.arXiv preprint arXiv:2508.09426, 2025
Yushan Xiang, Zhongwen Li, and Xiaoqi Li. Security analysis of chatgpt: Threats and privacy risks.arXiv preprint arXiv:2508.09426, 2025
arXiv 2025
-
[28]
Wei Zhang, Ju Xing, and Xiaoqi Li. Penetration testing for system security: Meth- ods and practical approaches.arXiv preprint arXiv:2505.19174, 2025
arXiv 2025
-
[29]
Risk assessment and security analysis of large language models.arXiv preprint arXiv:2508.17329, 2025
Xiaoyan Zhang, Dongyang Lyu, and Xiaoqi Li. Risk assessment and security analysis of large language models.arXiv preprint arXiv:2508.17329, 2025
Pith/arXiv arXiv 2025
-
[30]
Research on the speed and accuracy of full port scanning
Jinxiong Zhao, Lan Yang, Chi Zhang, and Jinpeng Zhang. Research on the speed and accuracy of full port scanning. In2023 IEEE 6th Information Technology, Networking, Electronic and Automation Control Conference (ITNEC), volume 6, pages 1159–1162. IEEE, 2023
2023
-
[31]
Nig-ap: A new method for automated penetration testing.Frontiers of Information Technology & Electronic Engineering, 20(9):1277–1288, 2019
Tian-yang Zhou, Yi-chao Zang, Jun-hu Zhu, and Qing-xian Wang. Nig-ap: A new method for automated penetration testing.Frontiers of Information Technology & Electronic Engineering, 20(9):1277–1288, 2019
2019
-
[32]
Blockchain security based on cryp- tography: a review.arXiv preprint arXiv:2508.01280, 2025
Wenwen Zhou, Dongyang Lyu, and Xiaoqi Li. Blockchain security based on cryp- tography: a review.arXiv preprint arXiv:2508.01280, 2025
Pith/arXiv arXiv 2025
-
[33]
Huanhuan Zou, Zongwei Li, and Xiaoqi Li. Malicious code detection in smart contracts via opcode vectorization.arXiv preprint arXiv:2504.12720, 2025. 34
arXiv 2025
This paper was first reviewed by deepseek-v4-flash on August 4, 2026.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.