REVIEW 2 major objections 24 references
A Tsetlin Machine detects PDF malware at 98 percent accuracy while explaining each decision through readable logical rules.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-13 04:06 UTC pith:I6PYJDAE
load-bearing objection Solid first application of Tsetlin Machines to PDF malware: competitive accuracy, real interpretability demos, and a new public dataset, tempered by single-corpus evaluation after heavy cleaning. the 2 major comments →
Leveraging Interpretable Tsetlin Machine for PDF Malware Detection
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
A Tsetlin Machine trained on 42 static numerical features extracted from PDF files without execution classifies them as benign or malicious at 98.02 percent accuracy (macro precision 96.03 percent, recall 95.95 percent, F1 95.99 percent) on the held-out RIT-PDFMal-2026 test set, matching or closely trailing strong black-box classifiers while supplying intrinsic explanations through clause activations, class-vote totals, and signed feature contributions.
What carries the argument
The Tsetlin Machine: a rule-based learner that builds conjunctive clauses from binary literals and their negations, then aggregates signed clause votes (bounded by threshold T) to decide the class. Those clauses are the human-readable decision rules that make classification transparent.
Load-bearing premise
The 42 static numerical features, after discarding more than a third of samples as duplicates and undersampling the benign class, still capture the patterns that matter for real-world PDF malware.
What would settle it
Re-run the identical feature extraction, preprocessing, and TM training pipeline on a fresh, independently collected PDF corpus that keeps natural class balance and includes malware families arriving after the original collection window; a large drop below the reported 98 percent accuracy would falsify the generalization claim.
If this is right
- Security operators can inspect activated clauses and top feature contributions to audit why any given PDF was flagged or cleared.
- Static analysis plus TM inference runs in roughly three microseconds per sample, supporting high-volume scanning.
- Competitive accuracy is obtained without post-hoc explanation tools such as SHAP or LIME.
- The same clause representation surfaces which structural PDF elements (JavaScript, Encrypt, Launch, XFA, etc.) dominate malicious patterns.
Where Pith is reading between the lines
- Because the model relies only on static structural counts, it can serve as a cheap first-stage filter before slower sandboxes or dynamic analysis.
- The higher false-negative rate relative to false positives suggests some malware families remain under-covered by the learned clauses; expanding the clause budget or feature set is a direct next experiment.
- Readable rules open a path for human-in-the-loop editing in which analysts disable or refine clauses that fire on known false positives.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes a Tsetlin Machine (TM) framework for static PDF malware detection. It extracts 42 numerical structural features without executing files, preprocesses the RIT-PDFMal-2026 corpus (duplicate removal, train-only undersampling, min-max scaling, KBinsDiscretizer binarization), and trains a TM with 250 clauses to classify benign vs. malicious PDFs. On a held-out 20% test split the model reports 98.02% accuracy and 95.99% macro F1, competitive with Random Forest (98.28%) and superior to XGBoost/LightGBM, while offering lower inference latency than RF and intrinsic interpretability via class-vote scores, clause-activation heatmaps, and feature-contribution plots. The authors position the combination of competitive accuracy, efficiency, and built-in explainability as the main contribution for practical PDF malware detection.
Significance. If the reported accuracy and interpretability claims hold under realistic distributions, the work supplies a concrete, low-latency alternative to black-box ensembles for a high-volume attack vector. The TM’s propositional clauses and vote/feature visualizations are a genuine methodological advantage over post-hoc SHAP/LIME explanations commonly applied to PDF detectors. The experimental pipeline is transparent (stratified split, train-only undersampling, five-fold CV, class-wise metrics, confusion matrix, seven baselines). The main limitation is that all evidence rests on a single, heavily cleaned public corpus; external validation or multi-dataset results would be required before the practical-deployment claim can be considered established. Within those bounds the contribution is solid and of clear interest to the malware-detection community.
major comments (2)
- Section V-C1 and Tables II/IV: After discarding 8,966 duplicates (36.84% of the original 24,337 samples) only 2,222 malicious files remain. Random undersampling is then applied solely to the training split, so the absolute number of malicious test examples is small. No ablation restores the original class prior, re-inserts near-duplicates, or evaluates an external hold-out corpus. Consequently the headline 98.02% accuracy / 95.99% macro-F1 figures (and the learned clauses shown in Figs. 7–15) may be optimistic relative to live PDF streams; at minimum the paper should quantify sensitivity to these preprocessing choices or report results on an independent corpus.
- Section V-D / Table VI: The state-of-the-art comparison juxtaposes methods evaluated on Contagio, Evasive-2022 and RIT-PDFMal-2026. Because the datasets differ in collection period, feature sets and class balance, the claim of “competitive performance … with existing methods” is only weakly supported. Either re-evaluate the baselines on the same RIT-PDFMal-2026 split or clearly qualify the comparison as non-head-to-head.
Circularity Check
No circularity: purely empirical supervised classification with held-out test metrics; self-citations are unrelated and non-load-bearing.
full rationale
The paper's central claims (98.02% accuracy / 95.99% macro F1 on RIT-PDFMal-2026, competitive with RF, plus intrinsic TM interpretability via clauses/votes/feature contributions) are obtained by standard supervised training and evaluation: static feature extraction, duplicate removal, stratified train/test split, random undersampling only on train, min-max + KBinsDiscretizer binarization, TM training with fixed hyperparameters (clauses=250, T=15, s=5, 50 epochs), and measurement of accuracy/precision/recall/F1/confusion matrix/inference time on the untouched test set (Tables III–V, Figs. 5–6). There is no mathematical derivation, uniqueness theorem, or fitted constant that is later re-presented as a prediction. Self-citations ([7]–[9], [23], [24]) concern the author's prior radio-map transfer-learning and speech-quality work and do not underwrite any malware-detection equation, feature set, or performance number; the TM itself is cited to Granmo and the dataset to Alani. Preprocessing choices affect absolute numbers but do not create a by-construction identity between inputs and reported outputs. The evaluation is therefore self-contained against the stated benchmark.
Axiom & Free-Parameter Ledger
free parameters (5)
- number of clauses =
250
- voting threshold T =
15
- specificity s =
5
- KBinsDiscretizer n_bins =
15
- training epochs =
50
axioms (3)
- domain assumption Static numerical features extracted without execution are sufficient to discriminate malicious from benign PDFs at the reported accuracy.
- domain assumption Random undersampling of the majority class on the training split does not destroy the decision boundary needed for generalization to the original imbalanced test distribution.
- domain assumption Propositional clauses learned by the Tsetlin Machine constitute faithful, human-interpretable explanations of the model’s decisions.
read the original abstract
In the digital era, Portable Document Format (PDF) is one of the most widely used file formats for storing and exchanging digital documents due to its platform independence and rich functionality. However, these same capabilities have also made PDF files an attractive attack vector for cyberattackers, who embed malicious code within seemingly legitimate documents to compromise target systems. This paper presents a novel interpretable Tsetlin Machine (TM)-based framework for PDF malware detection. The proposed framework extracts salient features from PDF documents through static analysis without executing the files and employs rule-based learning to accurately classify benign and malicious PDF documents. Numerical evaluation on the RIT-PDFMal-2026 dataset demonstrates that the proposed framework achieves competitive performance, attaining an accuracy of 98.02% compared with several ML classifiers and existing methods. Moreover, the proposed framework provides intrinsic interpretability by transparently explaining its classification decisions. The combination of competitive detection performance, computational efficiency, and intrinsic interpretability makes the proposed framework a promising solution for practical PDF malware detection.
Figures
Reference graph
Works this paper leans on
-
[1]
Digital files presence in the world, 2025,
CloudFiles, “Digital files presence in the world, 2025,” Accessed on July 05, 2026. [Online]. Available: https://www.cloudfiles.io/blog/how -many-files-are-there-in-the-world
2025
-
[2]
Malware Detection in PDF and Office Documents: A Survey,
P. Singh, S. Tapaswi, and S. Gupta, “Malware Detection in PDF and Office Documents: A Survey,”Information Security Journal: A Global Perspective, vol. 29, no. 3, pp. 134–153, 2020
2020
-
[3]
Malicious PDF Attacks on Microsoft Windows, 2026,
R. Informatica, “Malicious PDF Attacks on Microsoft Windows, 2026,” Accessed on July 05, 2026. [Online]. Available: https: //reisinformatica.com/malicious-pdf-attacks-on-microsoft-windows-202 6-protection-guide-for-canadian-businesses/
2026
-
[4]
V APD: An Anomaly Detection Model for PDF Malware Forensics with Adversarial Robust- ness,
S. Liu, J. Ming, Y . Zhou, J. Fu, and G. Peng, “V APD: An Anomaly Detection Model for PDF Malware Forensics with Adversarial Robust- ness,” in34th USENIX Security Symposium, 2025, pp. 4759–4778
2025
-
[5]
PDF Malware Detection Based on Optimizable Decision Trees,
Q. Abu Al-Haija and H. Qattous, “PDF Malware Detection Based on Optimizable Decision Trees,”Electronics, vol. 11, no. 19, p. 3142, 2022
2022
-
[6]
Leveraging Machine Learning-Based PDF Malware Detection in Snort,
F. Chbib and R. Khatoun, “Leveraging Machine Learning-Based PDF Malware Detection in Snort,” inInt. Conference on Electrical, Computer, Communications and Mechatronics Engineering. IEEE, 2024, pp. 1–6
2024
-
[7]
Leveraging Transfer learning for Radio Map Estimation via Mixture of Experts,
R. K. Jaiswal, M. Elnourani, S. Deshmukh, and B. Beferull-Lozano, “Leveraging Transfer learning for Radio Map Estimation via Mixture of Experts,”IEEE TCCN, vol. 12, pp. 846–863, 2025
2025
-
[8]
Location-free Indoor Radio Map Estimation using Transfer learning,
R. Jaiswal, M. Elnourani, S. Deshmukh, and B. Beferull-Lozano, “Location-free Indoor Radio Map Estimation using Transfer learning,” in97th Vehicular Technology Conference. IEEE, 2023, pp. 1–7
2023
-
[9]
A Data-driven Transfer Learning Method for Indoor Radio Map Estima- tion,
R. K. Jaiswal, M. Elnourani, S. Deshmukh, and B. Beferull-Lozano, “A Data-driven Transfer Learning Method for Indoor Radio Map Estima- tion,”IEEE TVT, vol. 75, no. 3, pp. 4261–4277, 2026
2026
-
[10]
TransNet: Unseen Malware Variants Detection using Deep Transfer Learning,
C. Rong, G. Gou, M. Cui, Z. Li, and L. Guo, “TransNet: Unseen Malware Variants Detection using Deep Transfer Learning,” inInt. Conf. on Security and Privacy in Communication Systems, 2020, pp. 84–101
2020
-
[11]
A Unified Approach to Interpreting Model Predictions,
S. M. Lundberg and S.-I. Lee, “A Unified Approach to Interpreting Model Predictions,” inNIPS, 2017, pp. 1–10
2017
-
[12]
A Comprehensive Review of Tsetlin Machines: Concepts, Applications, Analysis, and the Future,
S. Kundu, S. M. Mishra, G. Trivedi, and F. Merchant, “A Comprehensive Review of Tsetlin Machines: Concepts, Applications, Analysis, and the Future,”IEEE IoT Journal, vol. 13, no. 10, pp. 20 105–20 127, 2026
2026
-
[13]
O.-C. Granmo, “The Tsetlin Machine–A Game Theoretic Bandit Driven Approach to Optimal Pattern Recognition with Propositional Logic,” arXiv preprint arXiv:1804.01508, pp. 1–42, 2018
Pith/arXiv arXiv 2018
-
[14]
Breiman, J
L. Breiman, J. Friedman, R. A. Olshen, and C. J. Stone,Classification and Regression Trees. Chapman and Hall/CRC, 2017
2017
-
[15]
Alpaydin,Introduction to Machine Learning
E. Alpaydin,Introduction to Machine Learning. MIT press, 2020
2020
-
[16]
D. W. Hosmer Jr, S. Lemeshow, and R. X. Sturdivant,Applied Logistic Regression. John Wiley & Sons, 2013
2013
-
[17]
Xgboost: A Scalable Tree Boosting System,
T. Chen and C. Guestrin, “Xgboost: A Scalable Tree Boosting System,” in22nd ACM SIGKDD International Conference on Knowledge Discov- ery and Data Mining, 2016, pp. 785–794
2016
-
[18]
Lightgbm: A Highly Efficient Gradient Boosting Decision Tree,
G. Ke, Q. Meng, and T. Finley, “Lightgbm: A Highly Efficient Gradient Boosting Decision Tree,” inNIPS, 2017, pp. 1–9
2017
-
[19]
Handling Imbalanced Data: SMOTE vs Random Undersam- pling,
S. Mishra, “Handling Imbalanced Data: SMOTE vs Random Undersam- pling,”International Research Journal of Engineering and Technology, vol. 4, no. 8, pp. 317–320, 2017
2017
-
[20]
Bhagwat, M
R. Bhagwat, M. Abdolahnejad, and M. Moocarme,Applied Deep Learn- ing with Keras: Solve Complex Real-life Problems with the Simplicity of Keras. Packt Publishing Ltd, 2019
2019
-
[21]
RIT-PDFMal-2026: A Comprehensive Benchmark Dataset for PDF Malware Detection,
M. M. Alani and E. Damiani, “RIT-PDFMal-2026: A Comprehensive Benchmark Dataset for PDF Malware Detection,”IEEE Access, vol. 14, pp. 97 841–97 855, 2026
2026
-
[22]
Malicious Sample,
VirusTotal, “Malicious Sample,” Accessed on July 05, 2026. [Online]. Available: https://www.virustotal.com/gui/home/upload
2026
-
[23]
Performance Analysis of V oice Activity Detector in Pres- ence of Non-stationary Noise,
R. Jaiswal, “Performance Analysis of V oice Activity Detector in Pres- ence of Non-stationary Noise,” in11th International Conf. on Robotics, Vision, Signal Processing and Power Applications, 2022, pp. 59–65
2022
-
[24]
Non-intrusive Speech Quality Assess- ment using Context-aware Neural Networks,
R. K. Jaiswal and R. K. Dubey, “Non-intrusive Speech Quality Assess- ment using Context-aware Neural Networks,”International Journal of Speech Technology, vol. 25, no. 4, pp. 947–965, 2022
2022
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.