REVIEW 2 cited by
A Survey of Machine Learning Methods and Challenges for Windows Malware Classification
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Malware classification is a difficult problem, to which machine learning methods have been applied for decades. Yet progress has often been slow, in part due to a number of unique difficulties with the task that occur through all stages of the developing a machine learning system: data collection, labeling, feature creation and selection, model selection, and evaluation. In this survey we will review a number of the current methods and challenges related to malware classification, including data collection, feature extraction, and model construction, and evaluation. Our discussion will include thoughts on the constraints that must be considered for machine learning based solutions in this domain, and yet to be tackled problems for which machine learning could also provide a solution. This survey aims to be useful both to cybersecurity practitioners who wish to learn more about how machine learning can be applied to the malware problem, and to give data scientists the necessary background into the challenges in this uniquely complicated space.
Forward citations
Cited by 2 Pith papers
-
Adaptive Malware Detection using Sequential Feature Selection: A Dueling Double Deep Q-Network (D3QN) Framework for Intelligent Classification
A D3QN agent that jointly selects features and classifies malware reaches about 99% accuracy on two benchmarks, but the claimed efficiency gain fails because the full feature vector is always in the network input.
-
CodableLLM: Automating Decompiled and Source Code Mapping for LLM Dataset Generation
CodableLLM automatically aligns decompiled binary functions with source code functions to generate training datasets for code LLMs.
Discussion (0). Sign in to comment.