Pith. sign in

REVIEW 2 cited by

SOREL-20M: A Large Scale Benchmark Dataset for Malicious PE Detection

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2012.07634 v1 pith:IVTAI3E6 submitted 2020-12-14 cs.CR

classification cs.CR
keywords featuresdatasetmalwaremillionsamplesadditionalcodedetection
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In this paper we describe the SOREL-20M (Sophos/ReversingLabs-20 Million) dataset: a large-scale dataset consisting of nearly 20 million files with pre-extracted features and metadata, high-quality labels derived from multiple sources, information about vendor detections of the malware samples at the time of collection, and additional ``tags'' related to each malware sample to serve as additional targets. In addition to features and metadata, we also provide approximately 10 million ``disarmed'' malware samples -- samples with both the optional\_headers.subsystem and file\_header.machine flags set to zero -- that may be used for further exploration of features and detection strategies. We also provide Python code to interact with the data and features, as well as baseline neural network and gradient boosted decision tree models and their results, with full training and evaluation code, to serve as a starting point for further experimentation.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. EMBER2024 -- A Benchmark Dataset for Holistic Evaluation of Malware Classifiers

    cs.CR 2025-06 conditional novelty 7.0 of 10

    EMBER2024 provides a 3.2-million-file, six-format, seven-task malware benchmark with a dedicated challenge set of antivirus-evading samples.

  2. Adaptive Malware Detection using Sequential Feature Selection: A Dueling Double Deep Q-Network (D3QN) Framework for Intelligent Classification

    cs.LG 2025-07 reject novelty 5.0 of 10

    A D3QN agent that jointly selects features and classifies malware reaches about 99% accuracy on two benchmarks, but the claimed efficiency gain fails because the full feature vector is always in the network input.

Pith tools