REVIEW 3 major objections 4 minor 58 references
Living off the Analyst: Harvesting Features from Yara Rules for Malware Detection
T0 review · 3 major / 4 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read Analyst-written Yara signatures can be repurposed as static malware-detection features, giving a modest but real boost beyond conventional features.
desk verdict A genuinely new feature-mining idea that currently rests on an undated Yara-rule corpus, so the headline 1.8% gain should not be trusted until the temporal leakage is ruled out. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central mechanism is the sub-signature: each individual string literal inside a Yara rule, extracted as a separate binary feature that records whether that string occurs in a file, with the original rule's Boolean condition discarded. Because whole rules fire too rarely to be useful, the sub-signatures restore occurrence frequency. A Lasso logistic-regression model performs joint feature selection, and a gradient-boosted tree model then fits interactions among the selected sub-signatures, effectively reconstructing rule-like logic from data. The paper also evaluates three ways to bring in side information during selection: independent, conditional on the base model's predictions, and stacked with the original features.
What would settle it
Run the same pipeline with the 300 selected sub-signatures on a separate collection of Windows PE binaries gathered after EMBER's 2018 cutoff; if the low-false-positive AUC is not above the EMBER-only baseline, the reported 1.8% improvement is an artifact of the benchmark rather than a portable gain.
Extended reading notes
Core claim
The paper's central claim is that Yara signatures written by human analysts in the course of normal malware work can be re-purposed as machine-learning features for Windows PE malware detection. The extraction is deliberately crude: every string literal in a Yara rule becomes a candidate binary feature, and the rule's condition logic is dropped. After sparse linear selection on the EMBER training set, a few hundred sub-signatures, combined with the original EMBER features and trained into gradient-boosted trees, outperform the EMBER-only baseline, surpassing its 96.6% accuracy and achieving a 1.8% relative improvement at a false-positive rate of 0.01%. The selected features are not all rare family-specific markers or all generic strings; they form a power-law mixture, and roughly 30% of the information they carry is not linearly reconstructable from the original EMBER features.
Load-bearing premise
The load-bearing premise is that whether a small string literal from a Yara rule appears in a file is itself a stable, meaningful signal, even though the original rule's Boolean conditions are discarded and the rules were not filtered for quality or relevance.
Editorial extensions
If this is right
- Static malware detectors can be improved by mining the Yara rules analysts already produce, with no additional analyst labor.
- A relatively small set of a few hundred sub-signatures is enough to reach the performance plateau, making the approach computationally practical.
- Because tree-based models can reconstruct interactions among sub-signatures, the discarded Yara condition logic is not wasted information.
- The selected sub-signatures carry information distinct from conventional hand-engineered features, so they remain useful when added to an already strong baseline.
- The approach offers a possible low-cost way to refresh detector features as analysts write new signatures for emerging threats.
Reading between the lines
- A natural extension, not explored in the paper, would be a continuous pipeline: as new public Yara rules are published, fresh sub-signatures could be harvested automatically, giving detectors a cheap defense against concept drift.
- Because Yara rules are known to be subvertible, the selected sub-signatures should not be treated as adversarial-robust features; combining them with monotonic or non-negative models would be a separate test.
- The paper's maximal-correlation analysis is linear; a direct test of the 'new information' claim would ablate the Yara features entirely and compare low-false-positive AUC on a time-separated corpus, which would also reveal whether the 1.8% gain transfers beyond EMBER.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes harvesting sub-signatures from publicly available YARA rules and using their occurrence in Windows PE files as additional features for malware detection. The authors break each YARA rule into individual string literals, discard the boolean conditions, and build a binary feature vector indicating substring presence in each file. Using L1-regularized logistic regression to select a small number of such features, they combine them with the original EMBER 2018 features and train an XGBoost classifier. They report a relative improvement of 1.8% at a low false-positive rate of 0.01% over an EMBER-only baseline, along with manual analysis of the selected sub-signatures showing power-law behavior and dual-purpose strings. The evaluation uses EMBER's time-based train/test split.
Significance. If the central claim survives scrutiny, the contribution is a practical and low-cost way to augment static malware detectors with analyst knowledge, and the manual feature analysis provides useful insight into what kinds of sub-signatures are informative. The use of EMBER's held-out time split and the comparison against an EMBER-only baseline are appropriate controls. The paper does not ship code or artifacts, and it gives no uncertainty estimates, but the core idea is clearly presented and the empirical setup is straightforward. The main risk is temporal leakage from undated YARA rules, which could invalidate the headline improvement.
major comments (3)
- [Section III-A, Section IV-B] Section III-A states 'we did not place any restrictions on how or by whom the rules are written. Nor did we attempt to filter the rules in each repository.' This is load-bearing because EMBER 2018 is split by time (Section III-A), with test malware first seen in 2018, while the YARA rules were harvested from current GitHub repositories (Appendix A) at research time. Any rule written after a test sample's first-seen date can turn its sub-signatures into retrospective label indicators. The example in Section IV-D of a selected sub-signature identifying 2015 SPHINX MOTH does not establish that the rule predates the 2018 test samples; the rule could have been written much later. The claimed 1.8% relative improvement at FPR<0.01 (Figure 5(b)) could therefore be a temporal leakage artifact rather than evidence that analyst-written signatures add prospective value. The authors should rerun the experiment using only rules with commit dates or archive snapshots before the beginning of the EMBER 2018 test period, and report the gain under that restriction.
- [Figure 5(b), Section IV-B] The abstract and Section IV-B report a 1.8% relative improvement at a low false-positive rate of 0.01%, but the paper gives no measure of variance for this figure. XGBoost training is stochastic (the Optuna tuning in Section III-F uses random search), and Lasso feature selection depends on the training sample. A difference this small could easily be within run-to-run noise. Please provide repeated-seed experiments or bootstrap confidence intervals for AUC at FPR<0.01 and test accuracy for both the EMBER-only baseline and the best EMBER+YARA model, so the reader can judge whether the claimed improvement is statistically meaningful.
- [Algorithm 1, line 9; Section III-A] The paper represents each sub-signature occurrence as 1{r_j ∈ sample_i} (Algorithm 1, line 9) and says the binary vectors are computed over the corpus, but it never specifies the matching procedure. Are sub-signatures matched case-sensitively as ASCII byte strings anywhere in the file? Are YARA string modifiers such as 'wide' or 'nocase' preserved when the rules are decomposed? Without this detail, the reader cannot assess the validity of the feature extraction or reproduce it. Please specify the exact search procedure, including byte encoding, case handling, and whether the scan is performed on the raw PE file bytes or on extracted strings.
minor comments (4)
- [Section IV-D] The text refers to 'SPING MOTH' but the malware family is SPHINX MOTH; please fix this typo.
- [Section III-F] The description of tree-based methods contains the typo 'Random Rorest'; it should be 'Random Forest'.
- [Figure 8 caption] The caption says 'Darker values are higher correlation' but the displayed matrix has many dark regions; please add a colorbar with a consistent scale and label the axes for the EMBER and YARA feature blocks.
- [Section III-B and III-F] Equation (1) uses λ for the L1 penalty, while Section III-F says the regularization strength is the parameter C from scikit-learn; clarify that C is the inverse of the penalty strength.
Circularity Check
No significant circularity: the Yara-feature pipeline is an ordinary supervised-learning evaluation on a held-out EMBER test split; self-citations are peripheral.
full rationale
The derivation chain is empirical rather than definitional. Yara rules are decomposed into string-literal sub-signatures whose occurrence in a file is a binary feature; this feature definition depends only on analyst-written rules and file content, not on labels. Lasso (Eq. 1) selects sub-signatures using training labels, and XGBoost is then fit on the selected features plus EMBER features, but every performance claim is evaluated on the held-out EMBER 2018 test split. Thus the 1.8% relative improvement at FPR<0.01 is a genuine out-of-sample result, not a fitted input renamed as a prediction. The maximal-correlation analysis in Section IV-E is descriptive and is not presented as an independent prediction. The paper's self-citations, notably [49] and [45,47,46], support peripheral observations about analyst behavior and n-gram-like power-law occurrence; they are not load-bearing for the central claim. The one substantive validity concern is temporal leakage: Section III-A states the authors 'did not place any restrictions on how or by whom the rules are written. Nor did we attempt to filter the rules in each repository,' so some harvested rules may postdate EMBER 2018 test samples. That is a data-leakage / experimental-validity risk, not a circularity risk, because the feature definition does not by construction encode test labels and the question of whether rule dates cause leakage is empirical. No equation in the paper reduces to its own input, and no parameter is fitted to a subset and then reported as a prediction of a closely related quantity. Therefore the paper warrants a circularity score of 0.
Assumptions & free parameters
free parameters (3)
- Lasso regularization strength (lambda or C) =
chosen via validation; exact value not reported
- XGBoost hyperparameters =
tuned with Optuna, 100 iterations; final values in Appendix not shown in text
- Number of selected Yara sub-signatures (k) =
approximately 300
assumptions (4)
- domain assumption EMBER 2018 labels and the train/test time split are free of information leakage and reliable enough for evaluating malware detection.
- domain assumption The 22 public Yara rule repositories are representative of analyst-written signatures and require no filtering.
- domain assumption Treating each Yara string literal as an independent binary feature, ignoring the boolean conditions that combine them, preserves discriminative signal.
- standard math Lasso logistic regression and XGBoost behave as standard and converge to reliable models on this data distribution.
Cite this review
Pith. "Pith review of Living off the Analyst: Harvesting Features from Yara Rules for Malware Detection." pith.science (2026). https://pith.science/paper/AJ3447UY
@misc{pith2026241118516,
author = {Pith},
title = {Pith review of: Living off the Analyst: Harvesting Features from Yara Rules for Malware Detection},
year = {2026},
howpublished = {\url{https://pith.science/paper/AJ3447UY}},
note = {Machine review of arXiv:2411.18516}
}
read the original abstract
A strategy used by malicious actors is to "live off the land," where benign systems and tools already available on a victim's systems are used and repurposed for the malicious actor's intent. In this work, we ask if there is a way for anti-virus developers to similarly re-purpose existing work to improve their malware detection capability. We show that this is plausible via YARA rules, which use human-written signatures to detect specific malware families, functionalities, or other markers of interest. By extracting sub-signatures from publicly available YARA rules, we assembled a set of features that can more effectively discriminate malicious samples from benign ones. Our experiments demonstrate that these features add value beyond traditional features on the EMBER 2018 dataset. Manual analysis of the added sub-signatures shows a power-law behavior in a combination of features that are specific and unique, as well as features that occur often. A prior expectation may be that the features would be limited in being overly specific to unique malware families. This behavior is observed, and is apparently useful in practice. In addition, we also find sub-signatures that are dual-purpose (e.g., detecting virtual machine environments) or broadly generic (e.g., DLL imports).
Figures
Figures from the paper (6 more)
Reference graph
Works this paper leans on
-
[1]
Hojjat Aghakhani, Fabio Gritti, Francesco Mecca, Mar- tina Lindorfer, Stefano Ortolani, Davide Balzarotti, Gio- vanni Vigna, and Christopher Kruegel. “When Malware is Packin’ Heat; Limits of Machine Learning Classifiers Based on Static Analysis Features”. en. In: Proceedings 2020 Network and Distributed System Security Sympo- sium. San Diego, CA: Internet...
work page 2020
-
[2]
Optuna: A Next- generation Hyperparameter Optimization Framework
Takuya Akiba, Shotaro Sano, Toshihiko Yanase, Takeru Ohta, and Masanori Koyama. “Optuna: A Next- generation Hyperparameter Optimization Framework”. In: Proceedings of the 25th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining . Series Title: KDD ’19. New York, NY , USA: ACM, 2019
work page 2019
-
[3]
Yara: The pattern matching swiss knife for malware researchers (and everyone else)
Victor Alvarez. Yara: The pattern matching swiss knife for malware researchers (and everyone else). 2013
work page 2013
-
[4]
Ember: an open dataset for training static pe malware machine learning models
Hyrum S Anderson and Phil Roth. “Ember: an open dataset for training static pe malware machine learning models”. In: arXiv preprint arXiv:1804.04637 (2018)
arXiv 2018
-
[5]
Humans vs. Machines in Malware Classification
Simone Aonzo, Yufei Han, Alessandro Mantovani, and Davide Balzarotti. “Humans vs. Machines in Malware Classification”. en. In: 2023
work page 2023
-
[6]
Automatisierte Signaturgener- ierung f ¨ur Malware-St ¨amme
Christian Blichmann. “Automatisierte Signaturgener- ierung f ¨ur Malware-St ¨amme”. PhD thesis. Technical University of Dortmund, 2008
work page 2008
-
[7]
Challenges and pitfalls in malware research
Marcus Botacin, Fabricio Ceschin, Ruimin Sun, Daniela Oliveira, and Andr ´e Gr ´egio. “Challenges and pitfalls in malware research”. In: Computers & Security 106 (2021)
work page 2021
-
[8]
An- tiViruses under the Microscope: A Hands-On Perspec- tive
Marcus Botacin, Felipe Duarte Domingues, Fabr ´ıcio Ceschin, Raphael Machnicki, Marco Antonio Zanata Alves, Paulo L ´ıcio de Geus, and Andr ´e Gr ´egio. “An- tiViruses under the Microscope: A Hands-On Perspec- tive”. In: Computers & Security i (Oct. 2021)
work page 2021
Show all 58 references
-
[9]
The other guys: automated anal- ysis of marginalized malware
Marcus Felipe Botacin, Paulo L ´ıcio de Geus, and Andr´e Ricardo Abed Gr´egio. “The other guys: automated anal- ysis of marginalized malware”. In: Journal of Computer Virology and Hacking Techniques (2017)
2017
-
[10]
Random forests
Leo Breiman. “Random forests”. In: Machine learning 45 (2001)
2001
-
[11]
The MalSource Dataset: Quantifying Complexity and Code Reuse in Malware Development
Alejandro Calleja, Juan Tapiador, and Juan Caballero. “The MalSource Dataset: Quantifying Complexity and Code Reuse in Malware Development”. In: IEEE Trans- actions on Information Forensics and Security 14.12 (Dec. 2019). arXiv: 1811.06888
2019 arXiv
-
[12]
About the Robustness and Looseness of Yara Rules
Gerardo Canfora, Mimmo Carapella, Andrea Del Vec- chio, Laura Nardi, Antonio Pirozzi, and Corrado Aaron Visaggio. “About the Robustness and Looseness of Yara Rules”. In: Testing Software and Systems: 32nd IFIP WG 6.1 International Conference, ICTSS 2020, Naples, Italy, Decembe...
2020
-
[13]
Xgboost: A scalable tree boosting system
Tianqi Chen and Carlos Guestrin. “Xgboost: A scalable tree boosting system”. In: Proceedings of the 22nd acm sigkdd international conference on knowledge discovery and data mining . 2016
2016
-
[14]
Y AMME: a Y Ara-byte-signatures Metamorphic Mutation Engine
Antonio Coscia, Vincenzo Dentamaro, Stefano Galan- tucci, Antonio Maci, and Giuseppe Pirlo. “Y AMME: a Y Ara-byte-signatures Metamorphic Mutation Engine”. In: IEEE Transactions on Information Forensics and Se- curity 18 (2023). Conference Name: IEEE Transactions on Information...
2023
-
[15]
Adversarial Robustness with Non-uniform Perturbations
Ecenaz Erdemir, Jeffrey Bickford, Luca Melis, and Ser- gul Aydore. “Adversarial Robustness with Non-uniform Perturbations”. In: Advances in Neural Information Pro- cessing Systems. Ed. by M. Ranzato, A. Beygelzimer, Y . Dauphin, P.S. Liang, and J. Wortman Vaughan. V ol. 34. Cu...
2021
-
[16]
LIBLINEAR: A library for large linear classification
Rong-En Fan, Kai-Wei Chang, Cho-Jui Hsieh, Xiang- Rui Wang, and Chih-Jen Lin. “LIBLINEAR: A library for large linear classification”. In: the Journal of ma- chine Learning research 9 (2008)
2008
-
[17]
Non-Negative Net- works Against Adversarial Attacks
William Fleshman, Edward Raff, Jared Sylvester, Steven Forsyth, and Mark McLean. “Non-Negative Net- works Against Adversarial Attacks”. In: AAAI-2019 Workshop on Artificial Intelligence for Cyber Security (2019). arXiv: 1806.06108
2019 arXiv
-
[18]
Automatic Generation of String Signatures for Malware Detection
Kent Griffin, Scott Schneider, Xin Hu, and Tzi-cker Chiueh. “Automatic Generation of String Signatures for Malware Detection”. In: Recent Advances in Intrusion Detection (RAID). 2009
2009
-
[19]
RS-Del: Edit Distance Robustness Certificates for Sequence Classifiers via Randomized Deletion
Zhuoqun Huang, Neil G Marchant, Keane Lucas, Lujo Bauer, Olga Ohrimenko, and Benjamin Rubinstein. “RS-Del: Edit Distance Robustness Certificates for Sequence Classifiers via Randomized Deletion”. In: Advances in Neural Information Processing Systems . Ed. by A. Oh, T. Naumann,...
2023
-
[20]
Adversarially Robust Malware Detec- tion Using Monotonic Classification
Inigo Incer, Michael Theodorides, Sadia Afroz, and David Wagner. “Adversarially Robust Malware Detec- tion Using Monotonic Classification”. In: Proceedings of the Fourth ACM International Workshop on Security and Privacy Analytics . Series Title: IWSPA ’18. New York, NY , USA:...
2018
-
[21]
Transcend: Detecting Concept Drift in Mal- ware Classification Models
Roberto Jordaney, Kumar Sharad, Santanu K Dash, Zhi Wang, Davide Papini, Ilia Nouretdinov, and Lorenzo Cavallaro. “Transcend: Detecting Concept Drift in Mal- ware Classification Models”. In: 26th USENIX Secu- rity Symposium (USENIX Security 17) . Vancouver, BC: {USENIX} Associ...
2017
-
[22]
A Framework for Cluster and Classifier Evaluation in the Absence of Reference Labels
Robert J Joyce, Edward Raff, and Charles Nicholas. “A Framework for Cluster and Classifier Evaluation in the Absence of Reference Labels”. In: Proceedings of the 14th ACM Workshop on Artificial Intelligence and Security (AISec ’21). arXiv: 2109.11126v1. Association for Computi...
2021 arXiv
-
[23]
Rank-1 Similarity Matrix Decomposition For Model- ing Changes in Antivirus Consensus Through Time
Robert J Joyce, Edward Raff, and Charles Nicholas. “Rank-1 Similarity Matrix Decomposition For Model- ing Changes in Antivirus Consensus Through Time”. In: Proceedings of the Conference on Applied Ma- chine Learning for Information Security . arXiv: 2201.00757v1. 2021
2021 arXiv
-
[24]
Lightgbm: A highly efficient gradient boosting deci- sion tree
Guolin Ke, Qi Meng, Thomas Finley, Taifeng Wang, Wei Chen, Weidong Ma, Qiwei Ye, and Tie-Yan Liu. “Lightgbm: A highly efficient gradient boosting deci- sion tree”. In: Advances in neural information process- ing systems 30 (2017)
2017
-
[25]
Autograph: toward automated, distributed worm signature detection
Hyang-Ah Kim and Brad Karp. “Autograph: toward automated, distributed worm signature detection”. In: Proceedings of the 13th conference on USENIX Security Symposium - Volume 13 . 2004
2004
-
[26]
The Dropper Effect: Insights into Malware Distribution with Downloader Graph Analytics
Bum Jun Kwon, Jayanta Mondal, Jiyong Jang, Leyla Bilge, and Tudor Dumitras , . “The Dropper Effect: Insights into Malware Distribution with Downloader Graph Analytics”. In: Proceedings of the 22Nd ACM SIGSAC Conference on Computer and Communications Security. Series Title: CCS...
2015
-
[27]
Learning from Context : Exploiting and Interpreting File Path Informa- tion for Better Malware Detection
Adarsh Kyadige and Ethan M Rudd. “Learning from Context : Exploiting and Interpreting File Path Informa- tion for Better Malware Detection”. In: ArXiv e-prints (2019). arXiv: 1905.06987v1
2019 arXiv
-
[28]
Stuxnet: Dissecting a Cyberwarfare Weapon
Ralph Langner. “Stuxnet: Dissecting a Cyberwarfare Weapon”. In: IEEE Security & Privacy Magazine 9.3 (May 2011). ISBN: 1540-7993
2011
-
[29]
Large-Scale Identification of Malicious Singleton Files
Bo Li, Kevin Roundy, Chris Gates, and Yevgeniy V orobeychik. “Large-Scale Identification of Malicious Singleton Files”. In: 7TH ACM Conference on Data and Application Security and Privacy . 2017
2017
-
[30]
Automatically Generate Mal- ware Detection Rules By Extracting Risk Information
Haocong Li and Jie Li. “Automatically Generate Mal- ware Detection Rules By Extracting Risk Information”. In: Proceedings of the 2024 5th International Confer- ence on Computing, Networks and Internet of Things . New York, NY , USA: Association for Computing Ma- chinery, July 2024
2024
-
[31]
PackGenome: Automatically Generating Robust Y ARA Rules for Accurate Malware Packer Detection
Shijia Li, Jiang Ming, Pengda Qiu, Qiyuan Chen, Lanqing Liu, Huaifeng Bao, Qiang Wang, and Chunfu Jia. “PackGenome: Automatically Generating Robust Y ARA Rules for Accurate Malware Packer Detection”. In: Proceedings of the 2023 ACM SIGSAC Conference on Computer and Communicati...
2023
-
[32]
High-Dimensional Distributed Sparse Classification with Scalable Communication- Efficient Global Updates
Fred Lu, Ryan R. Curtin, Edward Raff, Francis Fer- raro, and James Holt. “High-Dimensional Distributed Sparse Classification with Scalable Communication- Efficient Global Updates”. In: Proceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining . Barce...
2024
-
[33]
Curtin, Edward Raff, Francis Ferraro, and James Holt
Fred Lu, Ryan R. Curtin, Edward Raff, Francis Ferraro, and James Holt. Optimizing the Optimal Weighted Av- erage: Efficient Distributed Sparse Classification. 2024
2024
-
[34]
Para- graph: Thwarting Signature Learning by Training Ma- liciously
James Newsome, Brad Karp, and Dawn Song. “Para- graph: Thwarting Signature Learning by Training Ma- liciously”. In: Proceedings of the 9th International Conference on Recent Advances in Intrusion Detection . Series Title: RAID’06. Berlin, Heidelberg: Springer- Verlag, 2006
2006
-
[35]
Poly- graph: Automatically Generating Signatures for Poly- morphic Worms
James Newsome, Brad Karp, and Dawn Song. “Poly- graph: Automatically Generating Signatures for Poly- morphic Worms”. In: Proceedings of the 2005 IEEE Symposium on Security and Privacy . Series Title: SP ’05. Washington, DC, USA: IEEE Computer Society, 2005
2005
-
[36]
Out of Distribution Data Detection Using Dropout Bayesian Neural Networks
Andre T Nguyen, Fred Lu, Gary Lopez Munoz, Ed- ward Raff, Charles Nicholas, and James Holt. “Out of Distribution Data Detection Using Dropout Bayesian Neural Networks”. In: Proceedings of the 36th AAAI Conference on Artificial Intelligence . 2022
2022
-
[37]
Leveraging Uncertainty for Improved Static Malware Detection Under Extreme False Positive Constraints
Andre T. Nguyen, Edward Raff, Charles Nicholas, and James Holt. “Leveraging Uncertainty for Improved Static Malware Detection Under Extreme False Positive Constraints”. In: IJCAI-21 1st International Workshop on Adaptive Cyber Defense . arXiv: 2108.04081. 2021
2021 arXiv
-
[38]
Would a File by Any Other Name Seem as Malicious?
Andre T Nguyen, Edward Raff, and Aaron Sant-Miller. “Would a File by Any Other Name Seem as Malicious?” In: 2019 IEEE International Conference on Big Data (Big Data). IEEE, Dec. 2019
2019
-
[39]
Small Effect Sizes in Malware Detection? Make Harder Train/Test Splits!
Tirth Patel, Fred Lu, Edward Raff, Charles Nicholas, Cynthia Matuszek, and James Holt. “Small Effect Sizes in Malware Detection? Make Harder Train/Test Splits!” In: Proceedings of the Conference on Applied Machine Learning in Information Security (2023)
2023
-
[40]
Scikit-learn: Machine learning in Python
Fabian Pedregosa, Ga ¨el Varoquaux, Alexandre Gram- fort, Vincent Michel, Bertrand Thirion, Olivier Grisel, Mathieu Blondel, Peter Prettenhofer, Ron Weiss, Vin- cent Dubourg, et al. “Scikit-learn: Machine learning in Python”. In: the Journal of machine Learning research 12 (2011)
2011
-
[41]
TESSER- ACT: Eliminating Experimental Bias in Malware Clas- sification across Space and Time
Feargus Pendlebury, Fabio Pierazzi, Roberto Jordaney, Johannes Kinder, and Lorenzo Cavallaro. “TESSER- ACT: Eliminating Experimental Bias in Malware Clas- sification across Space and Time”. In: 28th USENIX Security Symposium (USENIX Security 19). Santa Clara, CA: USENIX Associ...
2019
-
[42]
Misleading worm signature generators using deliberate noise injection
R. Perdisci, D. Dagon, Wenke Lee, P. Fogla, and M. Sharif. “Misleading worm signature generators using deliberate noise injection”. In: 2006 IEEE Symposium on Security and Privacy (S&P’06) . IEEE, 2006
2006
-
[43]
Asymmetric Certified Robust- ness via Feature-Convex Neural Networks
Samuel Pfrommer, Brendon G. Anderson, Julien Piet, and Somayeh Sojoudi. “Asymmetric Certified Robust- ness via Feature-Convex Neural Networks”. In: Thirty- seventh Conference on Neural Information Processing Systems. 2023
2023
-
[44]
Getting Passive Aggressive About False Positives: Patching Deployed Malware Detectors
Edward Raff, Bobby Filar, and James Holt. “Getting Passive Aggressive About False Positives: Patching Deployed Malware Detectors”. In: 2020 International Conference on Data Mining Workshops (ICDMW) . IEEE, Nov. 2020
2020
-
[45]
KiloGrams: Very Large N-Grams for Malware Classification
Edward Raff, William Fleming, Richard Zak, Hyrum Anderson, Bill Finlayson, Charles Nicholas, and Mark McLean. KiloGrams: Very Large N-Grams for Malware Classification. arXiv: 1908.00200. 2019
1908 arXiv
-
[46]
Hash-Grams On Many-Cores and Skewed Distributions
Edward Raff and Mark McLean. “Hash-Grams On Many-Cores and Skewed Distributions”. In: 2018 IEEE International Conference on Big Data (Big Data) . IEEE, Dec. 2018
2018
-
[47]
Hash-Grams: Faster N-Gram Features for Classification and Malware Detection
Edward Raff and Charles Nicholas. “Hash-Grams: Faster N-Gram Features for Classification and Malware Detection”. In: Proceedings of the ACM Symposium on Document Engineering 2018 . Halifax, NS, Canada: ACM, 2018
2018
-
[48]
An investigation of byte n-gram features for malware classification
Edward Raff, Richard Zak, Russell Cox, Jared Sylvester, Paul Yacci, Rebecca Ward, Anna Tracy, Mark McLean, and Charles Nicholas. “An investigation of byte n-gram features for malware classification”. en. In: Journal of Computer Virology and Hacking Techniques 14.1 (Feb. 2018)
2018
-
[49]
Automatic Yara Rule Gener- ation Using Biclustering
Edward Raff, Richard Zak, Gary Lopez Munoz, William Fleming, Hyrum S. Anderson, Bobby Filar, Charles Nicholas, and James Holt. “Automatic Yara Rule Gener- ation Using Biclustering”. In: 13th ACM Workshop on Artificial Intelligence and Security (AISec’20) . arXiv: 2009.03779. 2020
2009 arXiv
-
[50]
Florian Roth. yarGen. 2013
2013
-
[51]
Is Function Similarity Over-Engineered? Building a Benchmark
Rebecca Saul, Chang Liu, Noah Fleischmann, Richard J Zak, Kristopher Micinski, Edward Raff, and James Holt. “Is Function Similarity Over-Engineered? Building a Benchmark”. In: The Thirty-eight Conference on Neural Information Processing Systems Datasets and Bench- marks Track. 2024
2024
-
[52]
Tracking concept drift in malware families
Anshuman Singh, Andrew Walenstein, and Arun Lakhotia. “Tracking concept drift in malware families”. In: Proceedings of the 5th ACM workshop on Security and artificial intelligence - AISec ’12 (2012). ISBN: 9781450316644
2012
-
[53]
Guilt by Association: Large Scale Malware Detection by Mining File-relation Graphs
Acar Tamersoy, Kevin Roundy, and Duen Horng Chau. “Guilt by Association: Large Scale Malware Detection by Mining File-relation Graphs”. In: Proceedings of the 20th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining . Series Title: KDD ’14. New York, NY ...
2014
-
[54]
A Close Look at a Daily Dataset of Malware Samples
Xabier Ugarte-Pedrero, Mariano Graziano, and Da- vide Balzarotti. “A Close Look at a Daily Dataset of Malware Samples”. In: ACM Trans. Priv. Secur. 22.1 (2019)
2019
-
[55]
An Ob- servational Investigation of Reverse Engineers ’ Pro- cesses
Daniel V otipka, Seth M. Rabin, Kristopher Micinski, Jeffrey S. Foster, and Michelle M. Mazurek. “An Ob- servational Investigation of Reverse Engineers ’ Pro- cesses”. In: USENIX Security Symposium . 2019
2019
-
[56]
Stabilizing Linear Passive-Aggressive Online Learning with Weighted Reservoir Sampling
Skyler Wu, Fred Lu, Edward Raff, and James Holt. “Stabilizing Linear Passive-Aggressive Online Learning with Weighted Reservoir Sampling”. In: The Thirty- eighth Annual Conference on Neural Information Pro- cessing Systems. 2024
2024
-
[57]
Combining File Content and File Relations for Cloud Based Malware Detection
Yanfang Ye, Tao Li, Shenghuo Zhu, Weiwei Zhuang, Egemen Tas, Umesh Gupta, and Melih Abdulhayoglu. “Combining File Content and File Relations for Cloud Based Malware Detection”. In: Proceedings of the 17th ACM SIGKDD International Conference on Knowledge Discovery and Data Mini...
2011
-
[58]
An Inside Look into the Practice of Malware Analysis
Miuyin Yong Wong, Matthew Landen, Manos Anton- akakis, Douglas M Blough, Elissa M Redmiles, and Mustaque Ahamad. “An Inside Look into the Practice of Malware Analysis”. In: Proceedings of the 2021 ACM SIGSAC Conference on Computer and Communications Security. Series Title: CCS...
2021
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.