Pith. sign in

REVIEW 3 cited by

TESSERACT: Eliminating Experimental Bias in Malware Classification across Space and Time (Extended Version)

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.01359 v2 pith:R5DPSUGQ submitted 2024-02-02 cs.LG cs.CRcs.PF

classification cs.LGcs.CRcs.PF
keywords performancebiasclassifierdatastudiestimeadditionallybiases
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Machine learning (ML) plays a pivotal role in detecting malicious software. Despite the high F1-scores reported in numerous studies reaching upwards of 0.99, the issue is not completely solved. Malware detectors often experience performance decay due to constantly evolving operating systems and attack methods, which can render previously learned knowledge insufficient for accurate decision-making on new inputs. This paper argues that commonly reported results are inflated due to two pervasive sources of experimental bias in the detection task: spatial bias caused by data distributions that are not representative of a real-world deployment; and temporal bias caused by incorrect time splits of data, leading to unrealistic configurations. To address these biases, we introduce a set of constraints for fair experiment design, and propose a new metric, AUT, for classifier robustness in real-world settings. We additionally propose an algorithm designed to tune training data to enhance classifier performance. Finally, we present TESSERACT, an open-source framework for realistic classifier comparison. Our evaluation encompasses both traditional ML and deep learning methods, examining published works on an extensive Android dataset with 259,230 samples over a five-year span. Additionally, we conduct case studies in the Windows PE and PDF domains. Our findings identify the existence of biases in previous studies and reveal that significant performance enhancements are possible through appropriate, periodic tuning. We explore how mitigation strategies may support in achieving a more stable and better performance over time by employing multiple strategies to delay performance decay.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. How Benchmarks and Evaluation Protocols Shape Conclusions in Provenance-Based Intrusion Detection

    cs.CR 2026-08 conditional novelty 6.0 of 10

    A simple lexical allowlist matches or beats learned provenance-based detectors on three of four audited E3 datasets, showing that measured PIDS gains often reflect benchmark shortcuts rather than richer modeling.

  2. Malaika: Understanding Malware through Tri-Grounded Agentic Reasoning

    cs.CR 2026-07 conditional novelty 6.0 of 10

    A tri-grounded multi-agent harness improves precision and auditability of Android malware behavior reports over fixed LLM pipelines and frontier coding agents.

  3. SoK: The Pitfalls of Deep Reinforcement Learning for Cybersecurity

    cs.LG 2026-02 accept novelty 6.0 of 10

    Across 66 DRL-for-cybersecurity papers, the authors identify 11 recurring methodological pitfalls—averaging 5.8 per paper—and demonstrate their impact in four environments.

Pith tools