REVIEW 5 major objections 5 minor 24 references
MalVol-25: A Diverse, Labelled and Detailed Volatile Memory Dataset for Malware Detection and Response Testing and Validation
T0 review · 5 major / 5 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read MalVol-25 pairs clean and infected RAM dumps to give malware-detection models a system-state view.
desk verdict A small new memory-dump dataset with a DOI, but the manuscript's validation and RL-readiness claims are unsupported by the evidence presented. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is a paired memory snapshot: a clean RAM dump and an infected RAM dump of the same virtual machine, taken under standardised timing after controlled malware execution. The dataset is built by automated malware execution in an isolated, regularly reset virtualised environment, with snapshots later examined using established memory-forensics tooling and manual cross-checking. This pairing is what enables state-transition modelling: the clean snapshot defines the baseline state, the infected snapshot defines the post-infection state, and their difference supplies the behavioural and environmental features for machine-learning and reinforcement-learning models.
What would settle it
For each of the 15 named malware families, inspect the corresponding infected memory dump for that family's distinctive artifacts (for example, WannaCry's mutex or ransom-note remnants, Cerber's process names, or GandCrab's encryption traces); if a substantial share of claimed-infected dumps contains no trace of the named malware while the clean baseline does, the central claim of correctly labelled, diverse infection states collapses.
Extended reading notes
Core claim
The central claim is that MalVol-25 provides a diverse, labelled, and detailed volatile-memory resource for training and testing malware detection and response systems, where current datasets are limited in diversity, labelling, or detail. The authors construct it by pairing each malware sample with a clean baseline memory snapshot and a post-infection snapshot taken after the malware has had time to execute and manifest behaviour, across Windows 7, 8.1, 10, and 11. They argue that the resulting 30 dumps encode the state change caused by infection, and that automated forensic analysis and manual inspection confirm visible differences in processes, network patterns, and injected code between clean and infected states. The paper's stated significance is that system-state and transition modelling of this kind is what reinforcement-learning and agentic AI approaches need to make detection and response decisions.
Load-bearing premise
The dataset's central premise is that each infected snapshot actually contains active, correctly labelled malware behaviour; the paper says malware was given time to execute and that validation was performed, but it provides no per-sample evidence such as malware hashes, process lists, or infection logs to verify that the right malware ran in each dump.
Editorial extensions
If this is right
- Researchers can train and benchmark memory-based malware detectors on a common, labelled set of clean and infected dumps instead of ad hoc collections.
- The paired clean/infected structure supports reinforcement-learning formulations where system states and transitions are explicit, enabling detection and response strategies beyond static classification.
- Multi-OS and multi-family coverage allows evaluation of whether detection generalises across Windows versions and malware categories such as ransomware, trojans, and worms.
- The documented, integrity-checked release supports reproducible comparison of forensic and machine-learning pipelines, a step toward standard benchmarks in memory-level malware analysis.
Reading between the lines
- The dataset's practical value will depend on per-sample verification artifacts (hashes, process lists, timestamps, infection logs) being published alongside the dumps; the paper currently asserts validation without showing them.
- At 30 dumps and 15 families, the resource is more likely a testbed or benchmark seed than a training-scale corpus, and pairing it with existing larger datasets may be needed for deep-learning experiments.
- If the infection-timing protocol is standardised and released as part of the toolkit, the dataset could be extended to continuous time-series snapshots, which would materially strengthen the claimed reinforcement-learning use case.
- An independent audit using different memory-forensics tooling than the authors used would be a cheap and direct way to test label reliability and strengthen trust in the clean/infected pairing.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper introduces MalVol-25, a volatile-memory dataset constructed by running 15 malware samples (mostly ransomware, with a worm, a trojan, and injection/crypto malware) in isolated Windows 7/8.1/10/11 virtual machines, capturing a clean baseline memory dump and an infected memory dump for each pair. The authors claim the resulting 30 dumps are validated via Volatility Framework artifact extraction and manual inspection, documented with cryptographic checksums, and available at IEEE Dataport. They further claim the dataset supports machine learning, agentic AI, reinforcement-learning state-transition modelling, and multimodal analysis. The paper contains no per-snapshot artifact evidence, no label schema, no checksum manifest, and no experimental evaluation; several broad claims in the abstract and Sections IV-VII exceed what the methodology and reported scale can support.
Significance. If the dataset actually contains correctly labelled, verified infected snapshots, it would be a modest but useful addition to the small set of public memory-dump datasets, particularly because it provides paired clean/infected captures and reports ethical-approval and containment procedures. The strengths are a controlled experimental setup, use of a public repository (IEEE Dataport), a clear pairing design, and a stated intent to combine automated and manual validation. There are no equations or fitted parameters, so the usual circularity concerns do not apply; the heavy reliance on the authors' own reference [3] in the motivation does not by itself invalidate the dataset. However, the manuscript currently does not provide the evidence needed to assess those strengths, and the stated scale and missing modalities fall short of the abstract's and Section VII's promises.
major comments (5)
- [Section III-C and Table I] The described dataset contains only 30 memory dumps (one clean and one infected per malware/OS pair), but the abstract and Section VII claim the dataset enables modelling system states and transitions and RL-based malware detection. With no repeated trials, no timestamps, and no temporal sequence, one snapshot per state cannot support state-transition or reinforcement-learning claims. Either add the missing captures or substantially weaken these claims.
- [Section IV] The validation claim is unsupported. Section IV asserts that the Volatility Framework and manual inspection revealed anomalous processes and network patterns, but it provides no plugin names, no process names, no network connection lists, no injected-code indicators, no artifact counts, and no per-sample results. Since the dataset's central value depends on correct labels and verified infection, include a validation appendix or companion manifest with per-snapshot evidence and a concrete label schema.
- [Section V and Data Availability] The paper states that cryptographic checksums and standardised naming were used to preserve data integrity, but no checksums, file manifest, naming convention, or directory structure are shown. Without a manifest listing each file and its hash, the integrity and replicability claims cannot be checked by readers. Add this information either in the paper or as a linked file in the dataset repository.
- [Section VII] Section VII claims that the dataset includes 'time-series memory snapshots synchronised with multimodal data such as network traffic, system call traces, and logs.' Section III describes only RAM snapshot acquisition; no such synchronised multimodal collection is described in the methodology, and no corresponding data files are identified. This claim should be removed or the described data must be added to the dataset.
- [Abstract and Section III-B] The abstract and Section III-B emphasise 'multiple operating systems' and 'diverse platforms,' but Table I covers only Windows 7, 8.1, 10, and 11—four releases of the same OS family. The OS-diversity claim is overstated and should be reworded to 'multiple Windows versions' unless non-Windows systems are added to the dataset.
minor comments (5)
- [Figure 1] Figure 1 appears as a caption with no visible diagram in the manuscript; please include the actual figure or remove the reference.
- [General formatting] There are typographical artifacts such as 'V olatility' and 'MALV ADA' with irregular spacing, and missing spaces in 'SectionII' and 'SectionIII'; a careful proofreading pass is needed.
- [Section VI and Section VIII] Section VI concedes virtualization constraints and narrow coverage, while Section VIII calls the dataset 'comprehensive' and 'well-validated'; these statements should be reconciled to avoid overclaiming.
- [Data Availability] The Data Availability statement gives a DOI but no inventory of file names, sizes, or formats; a brief table of contents for the dataset would help readers understand what is being released.
- [References] Reference [5] contains a mixed DOI prefix that does not match the journal article, and reference [3] is the authors' own prior work cited repeatedly in the motivation; the former should be corrected and the latter should not be used as the sole basis for the RL-readiness claims.
Circularity Check
No significant circularity: the dataset is an empirical artifact and no claimed result reduces to its inputs; self-citations are motivational, not load-bearing.
full rationale
The paper's central claim is the production of a labelled memory-snapshot dataset (30 clean and infected dumps) plus documentation and validation. This is an empirical artifact, not a derived result: Section III-C describes a data-collection protocol (clean baseline, controlled infection, snapshot acquisition) with no equations, fitted parameters, or predictions that could be equivalent to the inputs by construction. Section IV's validation via Volatility and manual inspection is an external check of snapshot contents; regardless of whether the evidence is sufficient, it does not define infected snapshots in terms of the dataset's own claimed utility. The only author-overlapping citations are [3], [21], and [22], used in the introduction and related work to motivate the need for diverse datasets and to point to prior forensic and RL work. These are not invoked as uniqueness theorems, ansatz justifications, or fitted inputs, and the dataset's value does not rest on them. Sections VI and VII contain scope overclaims (e.g., time-series and multimodal synchronization) but those are support gaps, not circular reductions. Accordingly, no circularity step meets the standard of a quotable reduction to the paper's own inputs.
Assumptions & free parameters
assumptions (3)
- domain assumption Virtual machine memory snapshots faithfully represent real-world malware behaviour for detection training.
- domain assumption Each malware sample executed successfully and produced observable infection states in memory.
- domain assumption Clean baseline snapshots are truly clean and infected snapshots are correctly labelled.
Cite this review
Pith. "Pith review of MalVol-25: A Diverse, Labelled and Detailed Volatile Memory Dataset for Malware Detection and Response Testing and Validation." pith.science (2026). https://pith.science/paper/QDJOM4JF
@misc{pith2026250703993,
author = {Pith},
title = {Pith review of: MalVol-25: A Diverse, Labelled and Detailed Volatile Memory Dataset for Malware Detection and Response Testing and Validation},
year = {2026},
howpublished = {\url{https://pith.science/paper/QDJOM4JF}},
note = {Machine review of arXiv:2507.03993}
}
read the original abstract
This paper addresses the critical need for high-quality malware datasets that support advanced analysis techniques, particularly machine learning and agentic AI frameworks. Existing datasets often lack diversity, comprehensive labelling, and the complexity necessary for effective machine learning and agent-based AI training. To fill this gap, we developed a systematic approach for generating a dataset that combines automated malware execution in controlled virtual environments with dynamic monitoring tools. The resulting dataset comprises clean and infected memory snapshots across multiple malware families and operating systems, capturing detailed behavioural and environmental features. Key design decisions include applying ethical and legal compliance, thorough validation using both automated and manual methods, and comprehensive documentation to ensure replicability and integrity. The dataset's distinctive features enable modelling system states and transitions, facilitating RL-based malware detection and response strategies. This resource is significant for advancing adaptive cybersecurity defences and digital forensic research. Its scope supports diverse malware scenarios and offers potential for broader applications in incident response and automated threat mitigation.
Figures
Reference graph
Works this paper leans on
-
[3]
Dunsin, D., Ghanem, M.C., Ouazzane, K. and Vassilev, V ., 2025. Re- inforcement learning for an efficient and effective malware investigation during cyber Incident response. High-Confidence Computing, p.100299. https://doi.org/10.1016/j.hcc.2025.100299
arXiv 2025
-
[1]
Malik, M.I., Ibrahim, A., Hannay, P. and Sikos, L.F., 2023. Developing resilient cyber-physical systems: a review of state-of-the-art malware detection approaches, gaps, and future directions. Computers, 12(4), p.79. https://www.mdpi.com/2073-431X/12/4/79
work page 2023
-
[2]
Or-Meir, O., Nissim, N., Elovici, Y . and Rokach, L., 2019. Dy- namic malware analysis in the modern era—A state of the art survey. ACM Computing Surveys (CSUR), 52(5), pp.1-48. https://dl.acm.org/doi/abs/10.1145/3329786
doi:10.1145/3329786 2019
- [4]
-
[5]
Creation of a Dataset Modeling the Behavior of Malware in IoT Devices
Huertas Celdrán, A., Pérez Fernández, D., Ferrández-Pastor, F.J. and García Clemente, F.J. (2022) “Creation of a Dataset Modeling the Behavior of Malware in IoT Devices”, Sensors, 22(23), p. 9177. https: //doi.org/10.1007/978-3-030-96737-6_11
-
[6]
Nguyen, P.S., Huy, T.N., Tuan, T.A., Trung, P.D. and Long, H.V . (2025) ‘Hybrid feature extraction and integrated deep learning for cloud-based malware detection’, Computers & Security , 150, p. 104233. https://www.sciencedirect.com/science/article/abs/pii/S016740482400539X
work page 2025
-
[7]
Raducu, R., Villagrasa-Labrador, A., Rodríguez, R.J. and Álvarez, P. (2025) ‘MALV ADA: A framework for generating datasets of malware execution traces’, SoftwareX, 30, p. 102082. https://doi.org/10.1016/j.softx.2025.102082
-
[8]
Damasevicius, R., Venckauskas, A., Grigaliunas, S., Toldinas, J., Morke- vicius, N., Aleliunas, T. and Smuikys, P., 2020. LITNET-2020: An annotated real-world network flow dataset for network intrusion detection. Electronics, 9(5), p.800. https://www.mdpi.com/2079-9292/9/5/800
work page 2020
Show all 24 references
-
[9]
and Aloul, F., 2021
Abdalgawad, N., Sajun, A., Kaddoura, Y ., Zualkernan, I.A. and Aloul, F., 2021. Generative deep learning to detect cyberattacks for the IoT-23 dataset. IEEE Access, 10, pp.6430-6441. https://ieeexplore.ieee.org/abstract/document/9667357
2021
-
[10]
and Alavi, S.E., 2021, May
Abbasi, F., Naderan, M. and Alavi, S.E., 2021, May. Anomaly detection in Internet of Things using feature selection and classification based on Logistic Regression and Artificial Neural Network on N-BaIoT dataset. In 2021 5th International Conference on Internet of Things and ...
2021
-
[11]
and Venter, H., 2023
Singh, A., Ikuesan, R.A. and Venter, H., 2023. MalFe—Malware Feature Engineering Generation Platform. Computers, 12(10), p.201. Available at: https://doi.org/10.3390/computers12100201
2023 doi
-
[12]
and Goranin, N., 2019
Ceponis, D. and Goranin, N., 2019. Evaluation of deep learning methods efficiency for malicious and benign system calls classification on the AWSCTD. Security and Communication Networks, 2019(1), p.2317976. https://doi.org/10.1155/2019/2317976
2019 doi
-
[13]
and Barner, K.E
Makkawy, S.J., De Lucia, M.J. and Barner, K.E. (2025) ‘MalVis: A Large-Scale Image-Based Framework and Dataset for Advancing Android Malware Classification’, https://arxiv.org/abs/2505.12106
2025 arXiv
-
[14]
and Chau, D.H., 2022, October
Freitas, S., Duggal, R. and Chau, D.H., 2022, October. MalNet: A large-scale image database of malicious software. In Proceedings of the 31st ACM International Conference on Information & Knowledge Management (pp. 3948–3952). https://dl.acm.org/doi/abs/10.1145/3511808.3557533
2022
-
[15]
and Noever, S.E
Noever, D. and Noever, S.E. (2021) ‘Virus-MNIST: A benchmark malware dataset’. https://arxiv.org/abs/2103.00602
2021 arXiv
-
[16]
and Youn, J.M., 2025
Musaev, A., Anorboev, A. and Youn, J.M., 2025. Optimized Epoch Selection Ensemble: Integrating Custom CNN and Fine- Tuned MobileNetV2 for Malimg Dataset Classification. IEEE Access. https://ieeexplore.ieee.org/abstract/document/10909518
2025
-
[17]
and Le Traon, Y ., 2016, May
Allix, K., Bissyandé, T.F., Klein, J. and Le Traon, Y ., 2016, May. Androzoo: Collecting millions of android apps for the research community. In Proceedings of the 13th international conference on mining software repositories (pp. 468-471). https://dl.acm.org/doi/abs/10.1145/2...
2016
-
[18]
and Kalita, J.K., 2020, December
Borah, P., Bhattacharyya, D.K. and Kalita, J.K., 2020, December. Malware dataset generation and evaluation. In 2020 IEEE 4th Conference on Information & Communication Technology (CICT) (pp. 1–6). IEEE. https://doi.org/10.1109/CICT51604.2020.9312053
2020
-
[19]
and Binder, A.,
Sadek, I., Chong, P., Rehman, S.U., Elovici, Y . and Binder, A.,
-
[20]
Research on the Construction of Malware Variant Datasets and Their Detection Method
Lu, F., Cai, Z., Lin, Z., Bao, Y ., & Tang, M., 2022. Research on the Construction of Malware Variant Datasets and Their Detection Method. Applied Sciences , 12(15), 7546. https://doi.org/10.3390/app12157546
2022 doi
-
[21]
and Dunsin, D., 2025
Ghanem, M.C., Almeida Palmieri, E., Sowinski-Mydlarz, V ., Al-Sudani, S. and Dunsin, D., 2025. Weaponized IoT: a comprehensive comparative forensic analysis of Hacker Raspberry Pi and PC Kali Linux machine. IoT, 6(18), pp.1-23. IoT 2025, 6(1), 18; https://doi.org/10.3390/iot6010018
2025 doi
-
[22]
Animesh Singh Basnet, Mohamed Chahine Ghanem, Dipo Dunsin, Hamza Kheddar, and Wiktor Sowinski-Mydlarz. 2025. Advanced Per- sistent Threats (APT) Attribution Using Deep Reinforcement Learning. ACM Digital Threats, May 2025. https://doi.org/10.1145/3736654
2025 doi
- [2018]
-
[2019]
Data in brief, 26, p.104437
Memory snapshot dataset of a compromised host with malware using obfuscation evasion techniques. Data in brief, 26, p.104437. https://doi.org/10.1016/j.dib.2019.104437
2019
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.