Pith. sign in

REVIEW 3 major objections 4 minor 58 references

Living off the Analyst: Harvesting Features from Yara Rules for Malware Detection

T0 review · 3 major / 4 minor · reviewed 2026-08-12 · deepseek-v4-flash

Pith's one-line read Analyst-written Yara signatures can be repurposed as static malware-detection features, giving a modest but real boost beyond conventional features.

desk verdict A genuinely new feature-mining idea that currently rests on an undated Yara-rule corpus, so the headline 1.8% gain should not be trusted until the temporal leakage is ruled out. read the letter →

arxiv 2411.18516 v1 pith:AJ3447UY submitted 2024-11-27 cs.CR cs.LG

classification cs.CRcs.LG
keywords YararulesmalwaredetectionstaticfeaturesfeatureengineeringLassoselectionWindowsPEfilesEMBERdatasetsub-signatures
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper asks whether antivirus developers can 'live off the analyst': repurpose the YARA signatures analysts already write in their normal work as new features for static malware detection. It breaks each rule into individual string literals, called sub-signatures, discards the Boolean conditions that combine them, and uses sparse linear selection to pick a few hundred useful ones. On the EMBER 2018 corpus, adding these selected features to the standard features improves detection, with a relative gain of 1.8% at a 0.01% false-positive rate. The contribution matters because static detectors need regularly updated features and analyst time is expensive; if the approach holds, existing signature-writing effort can double as feature engineering.

What carries the argument

The central mechanism is the sub-signature: each individual string literal inside a Yara rule, extracted as a separate binary feature that records whether that string occurs in a file, with the original rule's Boolean condition discarded. Because whole rules fire too rarely to be useful, the sub-signatures restore occurrence frequency. A Lasso logistic-regression model performs joint feature selection, and a gradient-boosted tree model then fits interactions among the selected sub-signatures, effectively reconstructing rule-like logic from data. The paper also evaluates three ways to bring in side information during selection: independent, conditional on the base model's predictions, and stacked with the original features.

What would settle it

Run the same pipeline with the 300 selected sub-signatures on a separate collection of Windows PE binaries gathered after EMBER's 2018 cutoff; if the low-false-positive AUC is not above the EMBER-only baseline, the reported 1.8% improvement is an artifact of the benchmark rather than a portable gain.

Watch

Extended reading notes

Core claim

The paper's central claim is that Yara signatures written by human analysts in the course of normal malware work can be re-purposed as machine-learning features for Windows PE malware detection. The extraction is deliberately crude: every string literal in a Yara rule becomes a candidate binary feature, and the rule's condition logic is dropped. After sparse linear selection on the EMBER training set, a few hundred sub-signatures, combined with the original EMBER features and trained into gradient-boosted trees, outperform the EMBER-only baseline, surpassing its 96.6% accuracy and achieving a 1.8% relative improvement at a false-positive rate of 0.01%. The selected features are not all rare family-specific markers or all generic strings; they form a power-law mixture, and roughly 30% of the information they carry is not linearly reconstructable from the original EMBER features.

Load-bearing premise

The load-bearing premise is that whether a small string literal from a Yara rule appears in a file is itself a stable, meaningful signal, even though the original rule's Boolean conditions are discarded and the rules were not filtered for quality or relevance.

Editorial extensions

If this is right

  • Static malware detectors can be improved by mining the Yara rules analysts already produce, with no additional analyst labor.
  • A relatively small set of a few hundred sub-signatures is enough to reach the performance plateau, making the approach computationally practical.
  • Because tree-based models can reconstruct interactions among sub-signatures, the discarded Yara condition logic is not wasted information.
  • The selected sub-signatures carry information distinct from conventional hand-engineered features, so they remain useful when added to an already strong baseline.
  • The approach offers a possible low-cost way to refresh detector features as analysts write new signatures for emerging threats.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A natural extension, not explored in the paper, would be a continuous pipeline: as new public Yara rules are published, fresh sub-signatures could be harvested automatically, giving detectors a cheap defense against concept drift.
  • Because Yara rules are known to be subvertible, the selected sub-signatures should not be treated as adversarial-robust features; combining them with monotonic or non-negative models would be a separate test.
  • The paper's maximal-correlation analysis is linear; a direct test of the 'new information' claim would ablate the Yara features entirely and compare low-false-positive AUC on a time-separated corpus, which would also reveal whether the 1.8% gain transfers beyond EMBER.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. The paper proposes harvesting sub-signatures from publicly available YARA rules and using their occurrence in Windows PE files as additional features for malware detection. The authors break each YARA rule into individual string literals, discard the boolean conditions, and build a binary feature vector indicating substring presence in each file. Using L1-regularized logistic regression to select a small number of such features, they combine them with the original EMBER 2018 features and train an XGBoost classifier. They report a relative improvement of 1.8% at a low false-positive rate of 0.01% over an EMBER-only baseline, along with manual analysis of the selected sub-signatures showing power-law behavior and dual-purpose strings. The evaluation uses EMBER's time-based train/test split.

Significance. If the central claim survives scrutiny, the contribution is a practical and low-cost way to augment static malware detectors with analyst knowledge, and the manual feature analysis provides useful insight into what kinds of sub-signatures are informative. The use of EMBER's held-out time split and the comparison against an EMBER-only baseline are appropriate controls. The paper does not ship code or artifacts, and it gives no uncertainty estimates, but the core idea is clearly presented and the empirical setup is straightforward. The main risk is temporal leakage from undated YARA rules, which could invalidate the headline improvement.

major comments (3)
  1. [Section III-A, Section IV-B] Section III-A states 'we did not place any restrictions on how or by whom the rules are written. Nor did we attempt to filter the rules in each repository.' This is load-bearing because EMBER 2018 is split by time (Section III-A), with test malware first seen in 2018, while the YARA rules were harvested from current GitHub repositories (Appendix A) at research time. Any rule written after a test sample's first-seen date can turn its sub-signatures into retrospective label indicators. The example in Section IV-D of a selected sub-signature identifying 2015 SPHINX MOTH does not establish that the rule predates the 2018 test samples; the rule could have been written much later. The claimed 1.8% relative improvement at FPR<0.01 (Figure 5(b)) could therefore be a temporal leakage artifact rather than evidence that analyst-written signatures add prospective value. The authors should rerun the experiment using only rules with commit dates or archive snapshots before the beginning of the EMBER 2018 test period, and report the gain under that restriction.
  2. [Figure 5(b), Section IV-B] The abstract and Section IV-B report a 1.8% relative improvement at a low false-positive rate of 0.01%, but the paper gives no measure of variance for this figure. XGBoost training is stochastic (the Optuna tuning in Section III-F uses random search), and Lasso feature selection depends on the training sample. A difference this small could easily be within run-to-run noise. Please provide repeated-seed experiments or bootstrap confidence intervals for AUC at FPR<0.01 and test accuracy for both the EMBER-only baseline and the best EMBER+YARA model, so the reader can judge whether the claimed improvement is statistically meaningful.
  3. [Algorithm 1, line 9; Section III-A] The paper represents each sub-signature occurrence as 1{r_j ∈ sample_i} (Algorithm 1, line 9) and says the binary vectors are computed over the corpus, but it never specifies the matching procedure. Are sub-signatures matched case-sensitively as ASCII byte strings anywhere in the file? Are YARA string modifiers such as 'wide' or 'nocase' preserved when the rules are decomposed? Without this detail, the reader cannot assess the validity of the feature extraction or reproduce it. Please specify the exact search procedure, including byte encoding, case handling, and whether the scan is performed on the raw PE file bytes or on extracted strings.
minor comments (4)
  1. [Section IV-D] The text refers to 'SPING MOTH' but the malware family is SPHINX MOTH; please fix this typo.
  2. [Section III-F] The description of tree-based methods contains the typo 'Random Rorest'; it should be 'Random Forest'.
  3. [Figure 8 caption] The caption says 'Darker values are higher correlation' but the displayed matrix has many dark regions; please add a colorbar with a consistent scale and label the axes for the EMBER and YARA feature blocks.
  4. [Section III-B and III-F] Equation (1) uses λ for the L1 penalty, while Section III-F says the regularization strength is the parameter C from scikit-learn; clarify that C is the inverse of the penalty strength.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the Yara-feature pipeline is an ordinary supervised-learning evaluation on a held-out EMBER test split; self-citations are peripheral.

full rationale

The derivation chain is empirical rather than definitional. Yara rules are decomposed into string-literal sub-signatures whose occurrence in a file is a binary feature; this feature definition depends only on analyst-written rules and file content, not on labels. Lasso (Eq. 1) selects sub-signatures using training labels, and XGBoost is then fit on the selected features plus EMBER features, but every performance claim is evaluated on the held-out EMBER 2018 test split. Thus the 1.8% relative improvement at FPR<0.01 is a genuine out-of-sample result, not a fitted input renamed as a prediction. The maximal-correlation analysis in Section IV-E is descriptive and is not presented as an independent prediction. The paper's self-citations, notably [49] and [45,47,46], support peripheral observations about analyst behavior and n-gram-like power-law occurrence; they are not load-bearing for the central claim. The one substantive validity concern is temporal leakage: Section III-A states the authors 'did not place any restrictions on how or by whom the rules are written. Nor did we attempt to filter the rules in each repository,' so some harvested rules may postdate EMBER 2018 test samples. That is a data-leakage / experimental-validity risk, not a circularity risk, because the feature definition does not by construction encode test labels and the question of whether rule dates cause leakage is empirical. No equation in the paper reduces to its own input, and no parameter is fitted to a subset and then reported as a prediction of a closely related quantity. Therefore the paper warrants a circularity score of 0.

Assumptions & free parameters 3 free parameters · 4 assumptions · 0 invented entities

The central claim rests on several fitted parameters: the Lasso regularization strength, the XGBoost hyperparameters, and the chosen number of Yara features. The method also depends on domain assumptions about EMBER label reliability, the representativeness of public Yara repositories, and the validity of treating sub-signature occurrence as a feature. No new entities such as particles, forces, or mediators are introduced.

free parameters (3)
  • Lasso regularization strength (lambda or C) = chosen via validation; exact value not reported
    Controls the number of selected Yara sub-signatures, varied across Figures 4 and 5 to produce curves.
  • XGBoost hyperparameters = tuned with Optuna, 100 iterations; final values in Appendix not shown in text
    Model performance depends on these tuned values, and the main text does not report them.
  • Number of selected Yara sub-signatures (k) = approximately 300
    Chosen from the plateau in accuracy and low-FPR AUC curves; this selection is empirical.
assumptions (4)
  • domain assumption EMBER 2018 labels and the train/test time split are free of information leakage and reliable enough for evaluating malware detection.
    All performance claims rely on this dataset and split, as stated in Section III-A and used throughout Section IV.
  • domain assumption The 22 public Yara rule repositories are representative of analyst-written signatures and require no filtering.
    Section III-A explicitly declines to filter repositories, so any mismatch between public rules and production signatures transfers directly to the features.
  • domain assumption Treating each Yara string literal as an independent binary feature, ignoring the boolean conditions that combine them, preserves discriminative signal.
    This is the central design choice in Section III-A and Figure 1; if false, the reconstructed tree rules may not correspond to actual malware indicators.
  • standard math Lasso logistic regression and XGBoost behave as standard and converge to reliable models on this data distribution.
    Used as the feature selector and final classifier; these are standard algorithms, but no formal guarantees are given for this data.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Living off the Analyst: Harvesting Features from Yara Rules for Malware Detection." pith.science (2026). https://pith.science/paper/AJ3447UY

@misc{pith2026241118516,
  author       = {Pith},
  title        = {Pith review of: Living off the Analyst: Harvesting Features from Yara Rules for Malware Detection},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/AJ3447UY}},
  note         = {Machine review of arXiv:2411.18516}
}
read the original abstract

A strategy used by malicious actors is to "live off the land," where benign systems and tools already available on a victim's systems are used and repurposed for the malicious actor's intent. In this work, we ask if there is a way for anti-virus developers to similarly re-purpose existing work to improve their malware detection capability. We show that this is plausible via YARA rules, which use human-written signatures to detect specific malware families, functionalities, or other markers of interest. By extracting sub-signatures from publicly available YARA rules, we assembled a set of features that can more effectively discriminate malicious samples from benign ones. Our experiments demonstrate that these features add value beyond traditional features on the EMBER 2018 dataset. Manual analysis of the added sub-signatures shows a power-law behavior in a combination of features that are specific and unique, as well as features that occur often. A prior expectation may be that the features would be limited in being overly specific to unique malware families. This behavior is observed, and is apparently useful in practice. In addition, we also find sub-signatures that are dual-purpose (e.g., detecting virtual machine environments) or broadly generic (e.g., DLL imports).

Figures

Figures reproduced from arXiv: 2411.18516 by the authors.

Figure 1
Figure 1. Example of a Yara signature from https://github.com/tjnel/yara [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Distributional statistics of Yara sub-signatures over the Ember dataset [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗
Figure 3
Figure 3. Using only Yara sub-signatures as features, we see that only a limited [PITH_FULL_IMAGE:figures/full_fig_p004_3.png] view at source ↗
Figures from the paper (6 more)
Figure 4
Figure 4. Figure 4: In this experiment, linear classifiers are trained on a subset of Yara sub [PITH_FULL_IMAGE:figures/full_fig_p005_4.png]
Figure 6
Figure 6. Figure 6: Comparing various machine learning algorithms as YARA rules are [PITH_FULL_IMAGE:figures/full_fig_p006_6.png]
Figure 5
Figure 5. Figure 5: Performance of our full model as presented in Algo. 1, showing [PITH_FULL_IMAGE:figures/full_fig_p006_5.png]
Figure 7
Figure 7. Figure 7: The Empirical Cumulative Distribution Function (ECDF) of selected [PITH_FULL_IMAGE:figures/full_fig_p007_7.png]
Figure 8
Figure 8. Figure 8: Inter-correlation matrix among top 50 Ember and Yara features as [PITH_FULL_IMAGE:figures/full_fig_p008_8.png]
Figure 9
Figure 9. Figure 9: Maximal correlation between selected Yara sub-signatures with the [PITH_FULL_IMAGE:figures/full_fig_p008_9.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

58 extracted references · 51 canonical work pages

  1. [1]

    When Malware is Packin’ Heat; Limits of Machine Learning Classifiers Based on Static Analysis Features

    Hojjat Aghakhani, Fabio Gritti, Francesco Mecca, Mar- tina Lindorfer, Stefano Ortolani, Davide Balzarotti, Gio- vanni Vigna, and Christopher Kruegel. “When Malware is Packin’ Heat; Limits of Machine Learning Classifiers Based on Static Analysis Features”. en. In: Proceedings 2020 Network and Distributed System Security Sympo- sium. San Diego, CA: Internet...

  2. [2]

    Optuna: A Next- generation Hyperparameter Optimization Framework

    Takuya Akiba, Shotaro Sano, Toshihiko Yanase, Takeru Ohta, and Masanori Koyama. “Optuna: A Next- generation Hyperparameter Optimization Framework”. In: Proceedings of the 25th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining . Series Title: KDD ’19. New York, NY , USA: ACM, 2019

  3. [3]

    Yara: The pattern matching swiss knife for malware researchers (and everyone else)

    Victor Alvarez. Yara: The pattern matching swiss knife for malware researchers (and everyone else). 2013

  4. [4]

    Ember: an open dataset for training static pe malware machine learning models

    Hyrum S Anderson and Phil Roth. “Ember: an open dataset for training static pe malware machine learning models”. In: arXiv preprint arXiv:1804.04637 (2018)

  5. [5]

    Humans vs. Machines in Malware Classification

    Simone Aonzo, Yufei Han, Alessandro Mantovani, and Davide Balzarotti. “Humans vs. Machines in Malware Classification”. en. In: 2023

  6. [6]

    Automatisierte Signaturgener- ierung f ¨ur Malware-St ¨amme

    Christian Blichmann. “Automatisierte Signaturgener- ierung f ¨ur Malware-St ¨amme”. PhD thesis. Technical University of Dortmund, 2008

  7. [7]

    Challenges and pitfalls in malware research

    Marcus Botacin, Fabricio Ceschin, Ruimin Sun, Daniela Oliveira, and Andr ´e Gr ´egio. “Challenges and pitfalls in malware research”. In: Computers & Security 106 (2021)

  8. [8]

    An- tiViruses under the Microscope: A Hands-On Perspec- tive

    Marcus Botacin, Felipe Duarte Domingues, Fabr ´ıcio Ceschin, Raphael Machnicki, Marco Antonio Zanata Alves, Paulo L ´ıcio de Geus, and Andr ´e Gr ´egio. “An- tiViruses under the Microscope: A Hands-On Perspec- tive”. In: Computers & Security i (Oct. 2021)

Show all 58 references
  1. [9]

    The other guys: automated anal- ysis of marginalized malware

    Marcus Felipe Botacin, Paulo L ´ıcio de Geus, and Andr´e Ricardo Abed Gr´egio. “The other guys: automated anal- ysis of marginalized malware”. In: Journal of Computer Virology and Hacking Techniques (2017)

  2. [10]

    Random forests

    Leo Breiman. “Random forests”. In: Machine learning 45 (2001)

  3. [11]

    The MalSource Dataset: Quantifying Complexity and Code Reuse in Malware Development

    Alejandro Calleja, Juan Tapiador, and Juan Caballero. “The MalSource Dataset: Quantifying Complexity and Code Reuse in Malware Development”. In: IEEE Trans- actions on Information Forensics and Security 14.12 (Dec. 2019). arXiv: 1811.06888

  4. [12]

    About the Robustness and Looseness of Yara Rules

    Gerardo Canfora, Mimmo Carapella, Andrea Del Vec- chio, Laura Nardi, Antonio Pirozzi, and Corrado Aaron Visaggio. “About the Robustness and Looseness of Yara Rules”. In: Testing Software and Systems: 32nd IFIP WG 6.1 International Conference, ICTSS 2020, Naples, Italy, Decembe...

  5. [13]

    Xgboost: A scalable tree boosting system

    Tianqi Chen and Carlos Guestrin. “Xgboost: A scalable tree boosting system”. In: Proceedings of the 22nd acm sigkdd international conference on knowledge discovery and data mining . 2016

  6. [14]

    Y AMME: a Y Ara-byte-signatures Metamorphic Mutation Engine

    Antonio Coscia, Vincenzo Dentamaro, Stefano Galan- tucci, Antonio Maci, and Giuseppe Pirlo. “Y AMME: a Y Ara-byte-signatures Metamorphic Mutation Engine”. In: IEEE Transactions on Information Forensics and Se- curity 18 (2023). Conference Name: IEEE Transactions on Information...

  7. [15]

    Adversarial Robustness with Non-uniform Perturbations

    Ecenaz Erdemir, Jeffrey Bickford, Luca Melis, and Ser- gul Aydore. “Adversarial Robustness with Non-uniform Perturbations”. In: Advances in Neural Information Pro- cessing Systems. Ed. by M. Ranzato, A. Beygelzimer, Y . Dauphin, P.S. Liang, and J. Wortman Vaughan. V ol. 34. Cu...

  8. [16]

    LIBLINEAR: A library for large linear classification

    Rong-En Fan, Kai-Wei Chang, Cho-Jui Hsieh, Xiang- Rui Wang, and Chih-Jen Lin. “LIBLINEAR: A library for large linear classification”. In: the Journal of ma- chine Learning research 9 (2008)

  9. [17]

    Non-Negative Net- works Against Adversarial Attacks

    William Fleshman, Edward Raff, Jared Sylvester, Steven Forsyth, and Mark McLean. “Non-Negative Net- works Against Adversarial Attacks”. In: AAAI-2019 Workshop on Artificial Intelligence for Cyber Security (2019). arXiv: 1806.06108

  10. [18]

    Automatic Generation of String Signatures for Malware Detection

    Kent Griffin, Scott Schneider, Xin Hu, and Tzi-cker Chiueh. “Automatic Generation of String Signatures for Malware Detection”. In: Recent Advances in Intrusion Detection (RAID). 2009

  11. [19]

    RS-Del: Edit Distance Robustness Certificates for Sequence Classifiers via Randomized Deletion

    Zhuoqun Huang, Neil G Marchant, Keane Lucas, Lujo Bauer, Olga Ohrimenko, and Benjamin Rubinstein. “RS-Del: Edit Distance Robustness Certificates for Sequence Classifiers via Randomized Deletion”. In: Advances in Neural Information Processing Systems . Ed. by A. Oh, T. Naumann,...

  12. [20]

    Adversarially Robust Malware Detec- tion Using Monotonic Classification

    Inigo Incer, Michael Theodorides, Sadia Afroz, and David Wagner. “Adversarially Robust Malware Detec- tion Using Monotonic Classification”. In: Proceedings of the Fourth ACM International Workshop on Security and Privacy Analytics . Series Title: IWSPA ’18. New York, NY , USA:...

  13. [21]

    Transcend: Detecting Concept Drift in Mal- ware Classification Models

    Roberto Jordaney, Kumar Sharad, Santanu K Dash, Zhi Wang, Davide Papini, Ilia Nouretdinov, and Lorenzo Cavallaro. “Transcend: Detecting Concept Drift in Mal- ware Classification Models”. In: 26th USENIX Secu- rity Symposium (USENIX Security 17) . Vancouver, BC: {USENIX} Associ...

  14. [22]

    A Framework for Cluster and Classifier Evaluation in the Absence of Reference Labels

    Robert J Joyce, Edward Raff, and Charles Nicholas. “A Framework for Cluster and Classifier Evaluation in the Absence of Reference Labels”. In: Proceedings of the 14th ACM Workshop on Artificial Intelligence and Security (AISec ’21). arXiv: 2109.11126v1. Association for Computi...

  15. [23]

    Rank-1 Similarity Matrix Decomposition For Model- ing Changes in Antivirus Consensus Through Time

    Robert J Joyce, Edward Raff, and Charles Nicholas. “Rank-1 Similarity Matrix Decomposition For Model- ing Changes in Antivirus Consensus Through Time”. In: Proceedings of the Conference on Applied Ma- chine Learning for Information Security . arXiv: 2201.00757v1. 2021

  16. [24]

    Lightgbm: A highly efficient gradient boosting deci- sion tree

    Guolin Ke, Qi Meng, Thomas Finley, Taifeng Wang, Wei Chen, Weidong Ma, Qiwei Ye, and Tie-Yan Liu. “Lightgbm: A highly efficient gradient boosting deci- sion tree”. In: Advances in neural information process- ing systems 30 (2017)

  17. [25]

    Autograph: toward automated, distributed worm signature detection

    Hyang-Ah Kim and Brad Karp. “Autograph: toward automated, distributed worm signature detection”. In: Proceedings of the 13th conference on USENIX Security Symposium - Volume 13 . 2004

  18. [26]

    The Dropper Effect: Insights into Malware Distribution with Downloader Graph Analytics

    Bum Jun Kwon, Jayanta Mondal, Jiyong Jang, Leyla Bilge, and Tudor Dumitras , . “The Dropper Effect: Insights into Malware Distribution with Downloader Graph Analytics”. In: Proceedings of the 22Nd ACM SIGSAC Conference on Computer and Communications Security. Series Title: CCS...

  19. [27]

    Learning from Context : Exploiting and Interpreting File Path Informa- tion for Better Malware Detection

    Adarsh Kyadige and Ethan M Rudd. “Learning from Context : Exploiting and Interpreting File Path Informa- tion for Better Malware Detection”. In: ArXiv e-prints (2019). arXiv: 1905.06987v1

  20. [28]

    Stuxnet: Dissecting a Cyberwarfare Weapon

    Ralph Langner. “Stuxnet: Dissecting a Cyberwarfare Weapon”. In: IEEE Security & Privacy Magazine 9.3 (May 2011). ISBN: 1540-7993

  21. [29]

    Large-Scale Identification of Malicious Singleton Files

    Bo Li, Kevin Roundy, Chris Gates, and Yevgeniy V orobeychik. “Large-Scale Identification of Malicious Singleton Files”. In: 7TH ACM Conference on Data and Application Security and Privacy . 2017

  22. [30]

    Automatically Generate Mal- ware Detection Rules By Extracting Risk Information

    Haocong Li and Jie Li. “Automatically Generate Mal- ware Detection Rules By Extracting Risk Information”. In: Proceedings of the 2024 5th International Confer- ence on Computing, Networks and Internet of Things . New York, NY , USA: Association for Computing Ma- chinery, July 2024

  23. [31]

    PackGenome: Automatically Generating Robust Y ARA Rules for Accurate Malware Packer Detection

    Shijia Li, Jiang Ming, Pengda Qiu, Qiyuan Chen, Lanqing Liu, Huaifeng Bao, Qiang Wang, and Chunfu Jia. “PackGenome: Automatically Generating Robust Y ARA Rules for Accurate Malware Packer Detection”. In: Proceedings of the 2023 ACM SIGSAC Conference on Computer and Communicati...

  24. [32]

    High-Dimensional Distributed Sparse Classification with Scalable Communication- Efficient Global Updates

    Fred Lu, Ryan R. Curtin, Edward Raff, Francis Fer- raro, and James Holt. “High-Dimensional Distributed Sparse Classification with Scalable Communication- Efficient Global Updates”. In: Proceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining . Barce...

  25. [33]

    Curtin, Edward Raff, Francis Ferraro, and James Holt

    Fred Lu, Ryan R. Curtin, Edward Raff, Francis Ferraro, and James Holt. Optimizing the Optimal Weighted Av- erage: Efficient Distributed Sparse Classification. 2024

  26. [34]

    Para- graph: Thwarting Signature Learning by Training Ma- liciously

    James Newsome, Brad Karp, and Dawn Song. “Para- graph: Thwarting Signature Learning by Training Ma- liciously”. In: Proceedings of the 9th International Conference on Recent Advances in Intrusion Detection . Series Title: RAID’06. Berlin, Heidelberg: Springer- Verlag, 2006

  27. [35]

    Poly- graph: Automatically Generating Signatures for Poly- morphic Worms

    James Newsome, Brad Karp, and Dawn Song. “Poly- graph: Automatically Generating Signatures for Poly- morphic Worms”. In: Proceedings of the 2005 IEEE Symposium on Security and Privacy . Series Title: SP ’05. Washington, DC, USA: IEEE Computer Society, 2005

  28. [36]

    Out of Distribution Data Detection Using Dropout Bayesian Neural Networks

    Andre T Nguyen, Fred Lu, Gary Lopez Munoz, Ed- ward Raff, Charles Nicholas, and James Holt. “Out of Distribution Data Detection Using Dropout Bayesian Neural Networks”. In: Proceedings of the 36th AAAI Conference on Artificial Intelligence . 2022

  29. [37]

    Leveraging Uncertainty for Improved Static Malware Detection Under Extreme False Positive Constraints

    Andre T. Nguyen, Edward Raff, Charles Nicholas, and James Holt. “Leveraging Uncertainty for Improved Static Malware Detection Under Extreme False Positive Constraints”. In: IJCAI-21 1st International Workshop on Adaptive Cyber Defense . arXiv: 2108.04081. 2021

  30. [38]

    Would a File by Any Other Name Seem as Malicious?

    Andre T Nguyen, Edward Raff, and Aaron Sant-Miller. “Would a File by Any Other Name Seem as Malicious?” In: 2019 IEEE International Conference on Big Data (Big Data). IEEE, Dec. 2019

  31. [39]

    Small Effect Sizes in Malware Detection? Make Harder Train/Test Splits!

    Tirth Patel, Fred Lu, Edward Raff, Charles Nicholas, Cynthia Matuszek, and James Holt. “Small Effect Sizes in Malware Detection? Make Harder Train/Test Splits!” In: Proceedings of the Conference on Applied Machine Learning in Information Security (2023)

  32. [40]

    Scikit-learn: Machine learning in Python

    Fabian Pedregosa, Ga ¨el Varoquaux, Alexandre Gram- fort, Vincent Michel, Bertrand Thirion, Olivier Grisel, Mathieu Blondel, Peter Prettenhofer, Ron Weiss, Vin- cent Dubourg, et al. “Scikit-learn: Machine learning in Python”. In: the Journal of machine Learning research 12 (2011)

  33. [41]

    TESSER- ACT: Eliminating Experimental Bias in Malware Clas- sification across Space and Time

    Feargus Pendlebury, Fabio Pierazzi, Roberto Jordaney, Johannes Kinder, and Lorenzo Cavallaro. “TESSER- ACT: Eliminating Experimental Bias in Malware Clas- sification across Space and Time”. In: 28th USENIX Security Symposium (USENIX Security 19). Santa Clara, CA: USENIX Associ...

  34. [42]

    Misleading worm signature generators using deliberate noise injection

    R. Perdisci, D. Dagon, Wenke Lee, P. Fogla, and M. Sharif. “Misleading worm signature generators using deliberate noise injection”. In: 2006 IEEE Symposium on Security and Privacy (S&P’06) . IEEE, 2006

  35. [43]

    Asymmetric Certified Robust- ness via Feature-Convex Neural Networks

    Samuel Pfrommer, Brendon G. Anderson, Julien Piet, and Somayeh Sojoudi. “Asymmetric Certified Robust- ness via Feature-Convex Neural Networks”. In: Thirty- seventh Conference on Neural Information Processing Systems. 2023

  36. [44]

    Getting Passive Aggressive About False Positives: Patching Deployed Malware Detectors

    Edward Raff, Bobby Filar, and James Holt. “Getting Passive Aggressive About False Positives: Patching Deployed Malware Detectors”. In: 2020 International Conference on Data Mining Workshops (ICDMW) . IEEE, Nov. 2020

  37. [45]

    KiloGrams: Very Large N-Grams for Malware Classification

    Edward Raff, William Fleming, Richard Zak, Hyrum Anderson, Bill Finlayson, Charles Nicholas, and Mark McLean. KiloGrams: Very Large N-Grams for Malware Classification. arXiv: 1908.00200. 2019

  38. [46]

    Hash-Grams On Many-Cores and Skewed Distributions

    Edward Raff and Mark McLean. “Hash-Grams On Many-Cores and Skewed Distributions”. In: 2018 IEEE International Conference on Big Data (Big Data) . IEEE, Dec. 2018

  39. [47]

    Hash-Grams: Faster N-Gram Features for Classification and Malware Detection

    Edward Raff and Charles Nicholas. “Hash-Grams: Faster N-Gram Features for Classification and Malware Detection”. In: Proceedings of the ACM Symposium on Document Engineering 2018 . Halifax, NS, Canada: ACM, 2018

  40. [48]

    An investigation of byte n-gram features for malware classification

    Edward Raff, Richard Zak, Russell Cox, Jared Sylvester, Paul Yacci, Rebecca Ward, Anna Tracy, Mark McLean, and Charles Nicholas. “An investigation of byte n-gram features for malware classification”. en. In: Journal of Computer Virology and Hacking Techniques 14.1 (Feb. 2018)

  41. [49]

    Automatic Yara Rule Gener- ation Using Biclustering

    Edward Raff, Richard Zak, Gary Lopez Munoz, William Fleming, Hyrum S. Anderson, Bobby Filar, Charles Nicholas, and James Holt. “Automatic Yara Rule Gener- ation Using Biclustering”. In: 13th ACM Workshop on Artificial Intelligence and Security (AISec’20) . arXiv: 2009.03779. 2020

  42. [50]

    Florian Roth. yarGen. 2013

  43. [51]

    Is Function Similarity Over-Engineered? Building a Benchmark

    Rebecca Saul, Chang Liu, Noah Fleischmann, Richard J Zak, Kristopher Micinski, Edward Raff, and James Holt. “Is Function Similarity Over-Engineered? Building a Benchmark”. In: The Thirty-eight Conference on Neural Information Processing Systems Datasets and Bench- marks Track. 2024

  44. [52]

    Tracking concept drift in malware families

    Anshuman Singh, Andrew Walenstein, and Arun Lakhotia. “Tracking concept drift in malware families”. In: Proceedings of the 5th ACM workshop on Security and artificial intelligence - AISec ’12 (2012). ISBN: 9781450316644

  45. [53]

    Guilt by Association: Large Scale Malware Detection by Mining File-relation Graphs

    Acar Tamersoy, Kevin Roundy, and Duen Horng Chau. “Guilt by Association: Large Scale Malware Detection by Mining File-relation Graphs”. In: Proceedings of the 20th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining . Series Title: KDD ’14. New York, NY ...

  46. [54]

    A Close Look at a Daily Dataset of Malware Samples

    Xabier Ugarte-Pedrero, Mariano Graziano, and Da- vide Balzarotti. “A Close Look at a Daily Dataset of Malware Samples”. In: ACM Trans. Priv. Secur. 22.1 (2019)

  47. [55]

    An Ob- servational Investigation of Reverse Engineers ’ Pro- cesses

    Daniel V otipka, Seth M. Rabin, Kristopher Micinski, Jeffrey S. Foster, and Michelle M. Mazurek. “An Ob- servational Investigation of Reverse Engineers ’ Pro- cesses”. In: USENIX Security Symposium . 2019

  48. [56]

    Stabilizing Linear Passive-Aggressive Online Learning with Weighted Reservoir Sampling

    Skyler Wu, Fred Lu, Edward Raff, and James Holt. “Stabilizing Linear Passive-Aggressive Online Learning with Weighted Reservoir Sampling”. In: The Thirty- eighth Annual Conference on Neural Information Pro- cessing Systems. 2024

  49. [57]

    Combining File Content and File Relations for Cloud Based Malware Detection

    Yanfang Ye, Tao Li, Shenghuo Zhu, Weiwei Zhuang, Egemen Tas, Umesh Gupta, and Melih Abdulhayoglu. “Combining File Content and File Relations for Cloud Based Malware Detection”. In: Proceedings of the 17th ACM SIGKDD International Conference on Knowledge Discovery and Data Mini...

  50. [58]

    An Inside Look into the Practice of Malware Analysis

    Miuyin Yong Wong, Matthew Landen, Manos Anton- akakis, Douglas M Blough, Elissa M Redmiles, and Mustaque Ahamad. “An Inside Look into the Practice of Malware Analysis”. In: Proceedings of the 2021 ACM SIGSAC Conference on Computer and Communications Security. Series Title: CCS...

Pith tools

Reviewed August 12, 2026 · model on record in the stance chip above.