Pith. sign in

REVIEW 3 cited by

A Survey of Machine Learning Methods and Challenges for Windows Malware Classification

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2006.09271 v2 pith:CR3LI26C submitted 2020-06-15 cs.CR cs.LGstat.APstat.ML

classification cs.CRcs.LGstat.APstat.ML
keywords learningmachinemalwarechallengesclassificationdatamethodssurvey
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Malware classification is a difficult problem, to which machine learning methods have been applied for decades. Yet progress has often been slow, in part due to a number of unique difficulties with the task that occur through all stages of the developing a machine learning system: data collection, labeling, feature creation and selection, model selection, and evaluation. In this survey we will review a number of the current methods and challenges related to malware classification, including data collection, feature extraction, and model construction, and evaluation. Our discussion will include thoughts on the constraints that must be considered for machine learning based solutions in this domain, and yet to be tackled problems for which machine learning could also provide a solution. This survey aims to be useful both to cybersecurity practitioners who wish to learn more about how machine learning can be applied to the malware problem, and to give data scientists the necessary background into the challenges in this uniquely complicated space.

Discussion (0). Sign in to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. EMBER2024 -- A Benchmark Dataset for Holistic Evaluation of Malware Classifiers

    cs.CR 2025-06 conditional novelty 7.0 of 10

    EMBER2024 provides a 3.2-million-file, six-format, seven-task malware benchmark with a dedicated challenge set of antivirus-evading samples.

  2. Adaptive Malware Detection using Sequential Feature Selection: A Dueling Double Deep Q-Network (D3QN) Framework for Intelligent Classification

    cs.LG 2025-07 reject novelty 5.0 of 10

    A D3QN agent that jointly selects features and classifies malware reaches about 99% accuracy on two benchmarks, but the claimed efficiency gain fails because the full feature vector is always in the network input.

  3. CodableLLM: Automating Decompiled and Source Code Mapping for LLM Dataset Generation

    cs.SE 2025-07 conditional novelty 5.0 of 10

    CodableLLM automatically aligns decompiled binary functions with source code functions to generate training datasets for code LLMs.

Pith tools