Pith. sign in

REVIEW 1 cited by

Malware Detection with LSTM using Opcode Language

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1906.04593 v1 pith:MSH4A77K submitted 2019-06-10 cs.CR cs.SE

classification cs.CRcs.SE
keywords malwaredetectionopcodeperformlanguagelstmproposeapproach
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Nowadays, with the booming development of Internet and software industry, more and more malware variants are designed to perform various malicious activities. Traditional signature-based detection methods can not detect variants of malware. In addition, most behavior-based methods require a secure and isolated environment to perform malware detection, which is vulnerable to be contaminated. In this paper, similar to natural language processing, we propose a novel and efficient approach to perform static malware analysis, which can automatically learn the opcode sequence patterns of malware. We propose modeling malware as a language and assess the feasibility of this approach. First, We use the disassembly tool IDA Pro to obtain opcode sequence of malware. Then the word embedding technique is used to learn the feature vector representation of opcode. Finally, we propose a two-stage LSTM model for malware detection, which use two LSTM layers and one mean-pooling layer to obtain the feature representations of opcode sequences of malwares. We perform experiments on the dataset that includes 969 malware and 123 benign files. In terms of malware detection and malware classification, the evaluation results show our proposed method can achieve average AUC of 0.99 and average AUC of 0.987 in best case, respectively.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Malware Classification using a Hybrid Hidden Markov Model-Convolutional Neural Network

    cs.LG 2024-12 conditional novelty 3.0 of 10

    Converting HMM hidden-state sequences into images and classifying them with a CNN yields 0.9781 accuracy on a 7-family Malicia subset, a 0.0023 gain over the authors' HMM-RF baseline.

Pith tools