REVIEW 54 references
A compression based framework for the detection of anomalies in heterogeneous data sources
T0 review · reviewed 2026-08-14 · deepseek-v4-flash
Pith's one-line read A single NCD-plus-SVM pipeline reaches 0.77 to 0.95 accuracy on five text classification tasks, but the parameter-free claim is undercut by per-dataset tuning and training-set evaluation.
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
The authors apply this single pipeline to five problems: malicious HTTP queries, SMS spam, algorithmically generated domains, Twitter sentiment, and movie review sentiment. The reported accuracy ranges from about 0.77 for Twitter to 0.95 for HTTP. Those numbers are close to, but generally below, specialized state-of-the-art systems. The benefit claimed is convenience: no manual feature design and no task-specific preprocessing.
Two caveats matter. The paper calls the method parameter-free, but the number of generators and the SVM settings are different in every experiment and seem to be chosen by trying several values. Also, the main accuracy numbers are measured on the data used to train the classifier, with no separate test set, so real-world performance may be lower. No code is released, which makes exact reproduction harder.
Extended reading notes
Core claim
The abstract states: 'we propose a parameter-free methodology to detect security incidents from structured text regardless its nature. We use the Normalized Compression Distance to obtain a set of features that can be used by a Support Vector Machine to classify events from a heterogeneous cybersecurity environment.' If true, the same no-preprocessing pipeline would produce usable binary classifiers, with reported accuracies between 0.77 and 0.95, across HTTP requests, SMS spam, DGA domains, Twitter, and movie reviews. The 'parameter-free' part is not true as written, because k, C and gamma are tuned per dataset.
Load-bearing premise
The quantitative claims assume that accuracy measured on the I set, the same data used to train the SVM, estimates generalization to new text. Section 3 says the classifier is trained on the attribute vectors from I and quality is measured, with results averaged over random G/I partitions; no independent test set is used for the headline tables. Because k, C and gamma are also selected on these data (Tables 3, 5, 7, 9, 11), the reported numbers are at risk of optimism. If this assumption fails, the specific accuracies and AUCs in the paper are unsupported, although the qualitative flexibility claim might survive.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Assumptions & free parameters
free parameters (3)
- k (number of attribute generators) =
80 (best for HTTP and DGA), 160 (best for spam and movies), swept over 8 to 160
- SVM regularization parameter C =
0.01 to 100 depending on dataset and k
- SVM kernel parameter gamma =
0.1 to 100 for RBF experiments; linear kernel for Twitter
assumptions (4)
- domain assumption Gzip compressed length is an adequate proxy for Kolmogorov complexity in these text domains.
- domain assumption Balanced random splits G and I are representative of each class distribution.
- ad hoc to paper Performance measured on the SVM training set (I) estimates generalization.
- standard math An SVM with RBF or linear kernel can separate the NCD feature vectors.
Cite this review
Pith. "Pith review of A compression based framework for the detection of anomalies in heterogeneous data sources." pith.science (2026). https://pith.science/paper/Q4PIEICV
@misc{pith2026190800417,
author = {Pith},
title = {Pith review of: A compression based framework for the detection of anomalies in heterogeneous data sources},
year = {2026},
howpublished = {\url{https://pith.science/paper/Q4PIEICV}},
note = {Machine review of arXiv:1908.00417}
}
read the original abstract
Nowadays, information and communications technology systems are fundamental assets of our social and economical model, and thus they should be properly protected against the malicious activity of cybercriminals. Defence mechanisms are generally articulated around tools that trace and store information in several ways, the simplest one being the generation of plain text files coined as security logs. This log files are usually inspected, in a semi-automatic way, by security analysts to detect events that may affect system integrity. On this basis, we propose a parameter-free methodology to detect security incidents from structured text regardless its nature. We use the Normalized Compression Distance to obtain a set of features that can be used by a Support Vector Machine to classify events from a heterogeneous cybersecurity environment. In specific, we explore and validate the application of our methodology in four different cybersecurity domains: HTTP anomaly identification, spam detection, Domain Generation Algorithms tracking and sentiment analysis. The results obtained show the validity and flexibility of our approach in different security scenarios with a low configuration burden.
Figures
Reference graph
Works this paper leans on
-
[1]
OECD: The Economic Impact of ICT. (2004)
work page 2004
-
[2]
Technical report, ENISA (2019)
Sfakianakis, A., Douligeris, C., Marinos, L., Lourenço, M., Raghimi, O.: ENISA Threat Landscape Report 2018. Technical report, ENISA (2019)
work page 2019
- [3]
-
[4]
In: 24th USENIX Security Symposium (USENIX Security 15)
Sabottke, C., Suciu, O., Dumitras,, T.: Vulnerability disclosure in the age of social media: exploiting twitter for predicting real-world exploits. In: 24th USENIX Security Symposium (USENIX Security 15). (2015) 1041–1056
work page 2015
-
[5]
Curry, S., Kirda, E., Schwartz, E., Stewart, W., Yoran, A.: Big data fuels intelligence-driven security. RSA Security Brief (2013)
work page 2013
-
[6]
Keogh, E., Lonardi, S., Ratanamahatana, C.A.: Towards parameter-free data mining. In: Proceedings of the tenth ACM SIGKDD international conference on Knowledge discovery and data mining. (2004) 206–215
work page 2004
-
[7]
Ferragina, P., Giancarlo, R., Greco, V., Manzini, G., Valiente, G.: Compression-based classifi- cation of biological sequences and structures via the universal similarity metric: experimental assessment. BMC bioinformatics (2007)
work page 2007
-
[8]
In: 2006 IEEE International Symposium on Information Theory, IEEE (2006) 2309–2313
Cilibrasi, R., Vitanyi, P.: Automatic Extraction of Meaning from the Web. In: 2006 IEEE International Symposium on Information Theory, IEEE (2006) 2309–2313
work page 2006
Show all 54 references
-
[9]
IEEE Transactions on Information Theory (4) (2005) 1523–1545
Cilibrasi, R., Vitányi, P.M.B.: Clustering by compression. IEEE Transactions on Information Theory (4) (2005) 1523–1545
2005
-
[10]
PhD thesis, Tel-Aviv (2008)
Yahalom, S.: URI Anomaly Detection using Similarity Metrics. PhD thesis, Tel-Aviv (2008)
2008
-
[11]
In: International Joint Conference SOCO’17-CISIS’17-ICEUTE’17 León, Spain, September 6–8, 2017, Proceeding, Springer, Cham (2017) 661–671 10
de la Torre-Abaitua, G., Lago-Fernández, L.F., Arroyo, D.: A parameter-free method for the detection of web attacks. In: International Joint Conference SOCO’17-CISIS’17-ICEUTE’17 León, Spain, September 6–8, 2017, Proceeding, Springer, Cham (2017) 661–671 10
2017
-
[12]
In: International Conference on Human and Social Analytics (HUSO 2015)
Hee, C.V., Lefever, E., Verhoeven, B., Mennes, J., Desmet, B., Pauw, G.D., Daelemans, W., Hoste, V.: Automatic detection and prevention of cyberbullying. In: International Conference on Human and Social Analytics (HUSO 2015). (2015)
2015
-
[13]
Text Analytics for Cybersecurity and Online Safety (TA-COS) (2016)
Killam, R., Cook, P., Stakhanova, N.: Android malware classification through analysis of string literals. Text Analytics for Cybersecurity and Online Safety (TA-COS) (2016)
2016
-
[14]
In: Sensors
Hernandez-Suarez, A., Sanchez-Perez, G., Toscano-Medina, K., Martinez-Hernandez, V., Meana, H.M.P., Olivares-Mercado, J., Sanchez, V.: Social sentiment sensor in twitter for predicting cyber-attacks using𝓁1 regularization. In: Sensors. (2018)
2018
-
[15]
Computers & Security (1-2) (2009) 18–28
García-Teodoro, P., Díaz-Verdejo, J., Maciá-Fernández, G., Vázquez, E.: Anomaly-based network intrusion detection: Techniques, systems and challenges. Computers & Security (1-2) (2009) 18–28
2009
-
[16]
International Journal of current Engineering and Scientific Research (IJCESR) (9) (2016) 107–112
Chaurasia, M.A.: Comparative study of data mining techniques in intrusion dectection. International Journal of current Engineering and Scientific Research (IJCESR) (9) (2016) 107–112
2016
-
[17]
Communications Surveys & Tutorials, IEEE (1) (2014) 303–336
Bhuyan, M.H., Bhattacharyya, D.K., Kalita, J.K.: Network Anomaly Detection: Methods, Systems and Tools. Communications Surveys & Tutorials, IEEE (1) (2014) 303–336
2014
-
[18]
ArXiv e-prints (2017) 1–43
Hodo, E., Bellekens, X., Hamilton, A., Tachtatzis, C., Robert, A.: Shallow and Deep Networks Intrusion Detection System : A Taxonomy and Survey. ArXiv e-prints (2017) 1–43
2017
-
[19]
In: Proceedings of the 10th ACM Conference on Computer and Communications Security, New York, NY, USA, ACM (2003) 251–261
Kruegel, C., Vigna, G.: Anomaly Detection of Web-based Attacks. In: Proceedings of the 10th ACM Conference on Computer and Communications Security, New York, NY, USA, ACM (2003) 251–261
2003
-
[20]
ArXiv e-prints (2017)
Dong, Y., Zhang, Y.: Adaptively Detecting Malicious Queries in Web Attacks. ArXiv e-prints (2017)
2017
-
[21]
In: Proceedings of the Anti-phishing Working Groups 2Nd Annual eCrime Researchers Summit
Abu-Nimeh, S., Nappa, D., Wang, X., Nair, S.: A comparison of machine learning techniques for phishing detection. In: Proceedings of the Anti-phishing Working Groups 2Nd Annual eCrime Researchers Summit. eCrime ’07, New York, NY, USA, ACM (2007) 60–69
2007
-
[22]
Prabhakar, D.: A novel method of spam mail detection using text based clustering approach
Mallikarjunappa, B., R. Prabhakar, D.: A novel method of spam mail detection using text based clustering approach. International Journal of Computer Applications (08 2010)
2010
-
[23]
Jabatan Sistem dan Teknologi Komputer, Fakulti Sains Komputer dan Teknologi Maklumat, Universiti Malaya (2010)
Tee, H., dan Teknologi Komputer, U.M.J.S.: FPGA Unsolicited Commercial Email Inline Filter Design Using Levenshtein Distance Algorithm and Longest Common Subsequence Al- gorithm. Jabatan Sistem dan Teknologi Komputer, Fakulti Sains Komputer dan Teknologi Maklumat, Universiti M...
2010
-
[24]
In Weber, R.O., Richter, M.M., eds.: Case-Based Reasoning Research and Development, Berlin, Heidelberg, Springer Berlin Heidelberg (2007) 314–328
Delany, S.J., Bridge, D.: Catching the drift: Using feature-free case-based reasoning for spam filtering. In Weber, R.O., Richter, M.M., eds.: Case-Based Reasoning Research and Development, Berlin, Heidelberg, Springer Berlin Heidelberg (2007) 314–328
2007
-
[25]
Cybernetics and Systems (6-7) (2013) 533–549
Prilepok, M., Berek, P., Platos, J., Snasel, V.: spam detection using data compression and signatures. Cybernetics and Systems (6-7) (2013) 533–549
2013
-
[26]
Bratko, A., Filipič, B., Cormack, G.V., Lynam, T.R., Zupan, B.: Spam filtering using statis- tical data compression models. J. Mach. Learn. Res. (December 2006) 2673–2698
2006
-
[27]
In: Proceedings of the 23rd International Conference on World Wide Web
Thomas, M., Mohaisen, A.: Kindred domains: Detecting and clustering botnet domains using dns traffic. In: Proceedings of the 23rd International Conference on World Wide Web. WWW ’14 Companion, ACM (2014) 707–712
2014
-
[28]
In: Proceedings of the 21st USENIX Conference on Security Symposium
Antonakakis, M., Perdisci, R., Nadji, Y., Vasiloglou, N., Abu-Nimeh, S., Lee, W., Dagon, D.: From throw-away traffic to bots: Detecting the rise of dga-based malware. In: Proceedings of the 21st USENIX Conference on Security Symposium. Security’12, Berkeley, CA, USA, USENIX Asso...
2012
-
[29]
CoRR (2016) 11
Woodbridge, J., Anderson, H.S., Ahuja, A., Grant, D.: Predicting domain generation algo- rithms with long short-term memory networks. CoRR (2016) 11
2016
-
[30]
In Traore, I., Woungang, I., Awad, A., eds.: Intelligent, Secure, and Dependable Systems in Distributed and Cloud Environments, Cham, Springer International Publishing (2017) 19–34
Ahluwalia, A., Traore, I., Ganame, K., Agarwal, N.: Detecting broad length algorithmically generated domains. In Traore, I., Woungang, I., Awad, A., eds.: Intelligent, Secure, and Dependable Systems in Distributed and Cloud Environments, Cham, Springer International Publishing...
2017
-
[31]
Expert Systems with Applications124 (2019) 156 – 163
Selvi, J., Rodríguez, R.J., Soria-Olivas, E.: Detection of algorithmically generated malicious domain names using masked n-grams. Expert Systems with Applications124 (2019) 156 – 163
2019
-
[32]
In: Proceedings of the Seventh Symposium on Information and Communication Technology
Tong, V., Nguyen, G.: A method for detecting dga botnet based on semantic and cluster analysis. In: Proceedings of the Seventh Symposium on Information and Communication Technology. SoICT ’16, New York, NY, USA, ACM (2016) 272–277
2016
-
[33]
Al-Rowaily, K., Abulaish, M., Al-Hasan Haldar, N., Al-Rubaian, M.: Bisal - a bilingual sentiment analysis lexicon to analyze dark web forums for cyber security. Digit. Investig. 14(C) (2015) 53–62
2015
-
[34]
In: 2014 IEEE Joint Intelligence and Security Informatics Confer- ence
Weifeng, L., Hsinchun, C.: Identifyingtopsellersinundergroundeconomyusingdeeplearning- based sentiment analysis. In: 2014 IEEE Joint Intelligence and Security Informatics Confer- ence. (2014)
2014
-
[35]
CoRR abs/1804.05276 (2018)
Deb, A., Lerman, K., Ferrara, E.: Predicting cyber events by leveraging hacker sentiment. CoRR abs/1804.05276 (2018)
2018 arXiv
-
[36]
In: ASONAM
Mittal, S., Das, P.K., Mulwad, V., Joshi, A., Finin, T.: Cybertwitter: Using twitter to generate alerts for cybersecurity threats and vulnerabilities. In: ASONAM. (2016)
2016
-
[37]
Ain Shams Engineering Journal (4) (2014) 1093–1113
Medhat, W., Hassan, A., Korashy, H.: Sentiment analysis algorithms and applications: A survey. Ain Shams Engineering Journal (4) (2014) 1093–1113
2014
-
[38]
In: Mining text data
Liu, B., Zhang, L.: A survey of opinion mining and sentiment analysis. In: Mining text data. Springer (2012) 415–463
2012
-
[39]
In: Proceedings of the 38th International ACM SIGIR Conference on Research and Development in Information Retrieval, ACM (2015) 959–962
Severyn, A., Moschitti, A.: Twitter sentiment analysis with deep convolutional neural net- works. In: Proceedings of the 38th International ACM SIGIR Conference on Research and Development in Information Retrieval, ACM (2015) 959–962
2015
-
[40]
In: Proceedings of COLING 2014, the 25th International Conference on Computational Linguistics: Technical Papers
dos Santos, C., Gatti, M.: Deep convolutional neural networks for sentiment analysis of short texts. In: Proceedings of COLING 2014, the 25th International Conference on Computational Linguistics: Technical Papers. (2014) 69–78
2014
-
[41]
In: Proceedings of the 8th international workshop on semantic evaluation (SemEval 2014)
Tang, D., Wei, F., Qin, B., Liu, T., Zhou, M.: Coooolll: A deep learning system for twit- ter sentiment classification. In: Proceedings of the 8th international workshop on semantic evaluation (SemEval 2014). (2014) 208–212
2014
-
[42]
Bzip2 vs
Tukaani-project: A Quick Benchmark: Gzip vs. Bzip2 vs. LZMA. [Online]. Available from: http://tukaani.org/lzma/benchmarks.html [Accessed 2017-04-06]
2017
-
[43]
[Online]
scikit learn: scikit-learn: machine learning in Python - scikit-learn 0.18.1 documentation. [Online]. Available from:http://scikit-learn.org/stable/ [Accessed 2017-03-29]
2017
-
[44]
Security and Communi- cation Networks 8(16) (2015) 2750–2767
Torrano-Gimenez, C., Nguyen, H.T., Alvarez, G., Franke, K.: Combining expert knowledge with automatic feature extraction for reliable web attack detection. Security and Communi- cation Networks 8(16) (2015) 2750–2767
2015
-
[45]
[Online]
CSIC-dataset: HTTP DATASET CSIC 2010. [Online]. Available from: http://www.isi. csic.es/dataset/ [Accessed 2017-03-29]
2010
-
[46]
In: Computational Intelligence in Security for Information Systems
Nguyen, H.T., Torrano-Gimenez, C., Alvarez, G., Petrović, S., Franke, K.: Application of the generic feature selection measure in detection of web attacks. In: Computational Intelligence in Security for Information Systems. Springer (2011) 25–32
2011
-
[47]
In: Proceedings of 52nd annual meeting of the association for computational linguistics: system demonstrations
Manning, C., Surdeanu, M., Bauer, J., Finkel, J., Bethard, S., McClosky, D.: The stanford corenlp natural language processing toolkit. In: Proceedings of 52nd annual meeting of the association for computational linguistics: system demonstrations. (2014) 55–60 12
2014
-
[48]
In: Proceedings of the 11th ACM Symposium on Document Engineering
Almeida, T.A., Hidalgo, J.M.G., Yamakami, A.: Contributions to the study of sms spam fil- tering: New collection and results. In: Proceedings of the 11th ACM Symposium on Document Engineering. DocEng ’11, New York, NY, USA, ACM (2011) 259–262
2011
-
[49]
CoRR (2017)
Lison, P., Mavroeidis, V.: Automatic detection of malware-generated domains with recurrent neural models. CoRR (2017)
2017
-
[50]
Technical report, Stanford University (2009)
Go, A., Bhayani, R., Huang, L.: Twitter Sentiment Classification using Distant Supervision. Technical report, Stanford University (2009)
2009
-
[51]
In: Proceedings of the 49th Annual Meeting of the Association for Computational Linguistics: Human Language Technologies, Association for Computational Linguistics (2011) 142–150
Maas, A.L., Daly, R.E., Pham, P.T., Huang, D., Ng, A.Y., Potts, C.: Learning word vectors for sentiment analysis. In: Proceedings of the 49th Annual Meeting of the Association for Computational Linguistics: Human Language Technologies, Association for Computational Linguistics...
2011
-
[52]
arXiv preprint arXiv:1604.03850 (2016)
Lillis, D., Becker, B., O’Sullivan, T., Scanlon, M.: Current challenges and future research areas for digital forensic investigation. arXiv preprint arXiv:1604.03850 (2016)
2016 arXiv
-
[53]
Logic Journal of the IGPL (2019) In Press
de la Torre-Abaitua, G., Lago-Fernández, L., Arroyo, D.: On the application of compression based metrics to identifying anomalous behaviour in web traffic. Logic Journal of the IGPL (2019) In Press
2019
-
[54]
IEEE transactions on evolutionary computation (1997) 67–82 13
Wolpert, D.H., Macready, W.G.: No free lunch theorems for optimization. IEEE transactions on evolutionary computation (1997) 67–82 13
1997
Reviewed August 14, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.