Pith. sign in

REVIEW 54 references

A compression based framework for the detection of anomalies in heterogeneous data sources

T0 review · reviewed 2026-08-14 · deepseek-v4-flash

Pith's one-line read A single NCD-plus-SVM pipeline reaches 0.77 to 0.95 accuracy on five text classification tasks, but the parameter-free claim is undercut by per-dataset tuning and training-set evaluation.

arxiv 1908.00417 v1 pith:Q4PIEICV submitted 2019-08-01 cs.CR

classification cs.CR
keywords securitycompressioncybersecuritydetectdetectiondifferenteventsfiles
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Most security tools use different features for different jobs: URLs, SMS messages, domain names, and tweets each need their own rules. This paper asks whether one text-based trick can handle all of them. For every text, the method estimates how similar it is to a set of reference text files, called attribute generators, by comparing compressed sizes. Text that compresses well together with a file is considered close to that file. This produces a small vector of numbers for each text, and a support vector machine learns to separate the two classes from those vectors.

The authors apply this single pipeline to five problems: malicious HTTP queries, SMS spam, algorithmically generated domains, Twitter sentiment, and movie review sentiment. The reported accuracy ranges from about 0.77 for Twitter to 0.95 for HTTP. Those numbers are close to, but generally below, specialized state-of-the-art systems. The benefit claimed is convenience: no manual feature design and no task-specific preprocessing.

Two caveats matter. The paper calls the method parameter-free, but the number of generators and the SVM settings are different in every experiment and seem to be chosen by trying several values. Also, the main accuracy numbers are measured on the data used to train the classifier, with no separate test set, so real-world performance may be lower. No code is released, which makes exact reproduction harder.

Extended reading notes

Core claim

The abstract states: 'we propose a parameter-free methodology to detect security incidents from structured text regardless its nature. We use the Normalized Compression Distance to obtain a set of features that can be used by a Support Vector Machine to classify events from a heterogeneous cybersecurity environment.' If true, the same no-preprocessing pipeline would produce usable binary classifiers, with reported accuracies between 0.77 and 0.95, across HTTP requests, SMS spam, DGA domains, Twitter, and movie reviews. The 'parameter-free' part is not true as written, because k, C and gamma are tuned per dataset.

Load-bearing premise

The quantitative claims assume that accuracy measured on the I set, the same data used to train the SVM, estimates generalization to new text. Section 3 says the classifier is trained on the attribute vectors from I and quality is measured, with results averaged over random G/I partitions; no independent test set is used for the headline tables. Because k, C and gamma are also selected on these data (Tables 3, 5, 7, 9, 11), the reported numbers are at risk of optimism. If this assumption fails, the specific accuracies and AUCs in the paper are unsupported, although the qualitative flexibility claim might survive.

Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Assumptions & free parameters 3 free parameters · 4 assumptions · 0 invented entities

The central claim rests on tuned parameters (k, C, gamma) despite the 'parameter-free' label, plus the assumption that gzip-based NCD features and training-set accuracy generalize. No new entities are introduced.

free parameters (3)
  • k (number of attribute generators) = 80 (best for HTTP and DGA), 160 (best for spam and movies), swept over 8 to 160
    The abstract calls the method parameter-free, but k is swept and the best value is reported in Tables 3, 5, 7, 9, 11.
  • SVM regularization parameter C = 0.01 to 100 depending on dataset and k
    C is listed per experiment and tuned; for Twitter it is set by 10-fold cross-validation.
  • SVM kernel parameter gamma = 0.1 to 100 for RBF experiments; linear kernel for Twitter
    Gamma is chosen per experiment; no selection procedure is described for the RBF experiments.
assumptions (4)
  • domain assumption Gzip compressed length is an adequate proxy for Kolmogorov complexity in these text domains.
    Eq. (3) defines NCD using C(x), the gzip compressed size; no evidence is given that gzip preserves the discriminative information for all five tasks.
  • domain assumption Balanced random splits G and I are representative of each class distribution.
    Section 3 constructs generators from G and trains on I; results are averaged over random splits, but no stratification details beyond class balance are given.
  • ad hoc to paper Performance measured on the SVM training set (I) estimates generalization.
    Section 3 describes training and measuring quality on the same I set; no independent test set is used for the headline numbers, an unstated methodological assumption.
  • standard math An SVM with RBF or linear kernel can separate the NCD feature vectors.
    This is a standard machine learning assumption; kernel choice varies by task.

how reviews work

0 comments
Cite this review

Pith. "Pith review of A compression based framework for the detection of anomalies in heterogeneous data sources." pith.science (2026). https://pith.science/paper/Q4PIEICV

@misc{pith2026190800417,
  author       = {Pith},
  title        = {Pith review of: A compression based framework for the detection of anomalies in heterogeneous data sources},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/Q4PIEICV}},
  note         = {Machine review of arXiv:1908.00417}
}
read the original abstract

Nowadays, information and communications technology systems are fundamental assets of our social and economical model, and thus they should be properly protected against the malicious activity of cybercriminals. Defence mechanisms are generally articulated around tools that trace and store information in several ways, the simplest one being the generation of plain text files coined as security logs. This log files are usually inspected, in a semi-automatic way, by security analysts to detect events that may affect system integrity. On this basis, we propose a parameter-free methodology to detect security incidents from structured text regardless its nature. We use the Normalized Compression Distance to obtain a set of features that can be used by a Support Vector Machine to classify events from a heterogeneous cybersecurity environment. In specific, we explore and validate the application of our methodology in four different cybersecurity domains: HTTP anomaly identification, spam detection, Domain Generation Algorithms tracking and sentiment analysis. The results obtained show the validity and flexibility of our approach in different security scenarios with a low configuration burden.

Figures

Figures reproduced from arXiv: 1908.00417 by the authors.

Figure 1
Figure 1. End to end representation of the proposed methodology. The set of texts composing the dataset are first [PITH_FULL_IMAGE:figures/full_fig_p004_1.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

54 extracted references · 54 canonical work pages

  1. [1]

    OECD: The Economic Impact of ICT. (2004)

  2. [2]

    Technical report, ENISA (2019)

    Sfakianakis, A., Douligeris, C., Marinos, L., Lourenço, M., Raghimi, O.: ENISA Threat Landscape Report 2018. Technical report, ENISA (2019)

  3. [3]

    Syngress, Boston (2013)

    : Logging and log management. Syngress, Boston (2013)

  4. [4]

    In: 24th USENIX Security Symposium (USENIX Security 15)

    Sabottke, C., Suciu, O., Dumitras,, T.: Vulnerability disclosure in the age of social media: exploiting twitter for predicting real-world exploits. In: 24th USENIX Security Symposium (USENIX Security 15). (2015) 1041–1056

  5. [5]

    RSA Security Brief (2013)

    Curry, S., Kirda, E., Schwartz, E., Stewart, W., Yoran, A.: Big data fuels intelligence-driven security. RSA Security Brief (2013)

  6. [6]

    In: Proceedings of the tenth ACM SIGKDD international conference on Knowledge discovery and data mining

    Keogh, E., Lonardi, S., Ratanamahatana, C.A.: Towards parameter-free data mining. In: Proceedings of the tenth ACM SIGKDD international conference on Knowledge discovery and data mining. (2004) 206–215

  7. [7]

    BMC bioinformatics (2007)

    Ferragina, P., Giancarlo, R., Greco, V., Manzini, G., Valiente, G.: Compression-based classifi- cation of biological sequences and structures via the universal similarity metric: experimental assessment. BMC bioinformatics (2007)

  8. [8]

    In: 2006 IEEE International Symposium on Information Theory, IEEE (2006) 2309–2313

    Cilibrasi, R., Vitanyi, P.: Automatic Extraction of Meaning from the Web. In: 2006 IEEE International Symposium on Information Theory, IEEE (2006) 2309–2313

Show all 54 references
  1. [9]

    IEEE Transactions on Information Theory (4) (2005) 1523–1545

    Cilibrasi, R., Vitányi, P.M.B.: Clustering by compression. IEEE Transactions on Information Theory (4) (2005) 1523–1545

  2. [10]

    PhD thesis, Tel-Aviv (2008)

    Yahalom, S.: URI Anomaly Detection using Similarity Metrics. PhD thesis, Tel-Aviv (2008)

  3. [11]

    In: International Joint Conference SOCO’17-CISIS’17-ICEUTE’17 León, Spain, September 6–8, 2017, Proceeding, Springer, Cham (2017) 661–671 10

    de la Torre-Abaitua, G., Lago-Fernández, L.F., Arroyo, D.: A parameter-free method for the detection of web attacks. In: International Joint Conference SOCO’17-CISIS’17-ICEUTE’17 León, Spain, September 6–8, 2017, Proceeding, Springer, Cham (2017) 661–671 10

  4. [12]

    In: International Conference on Human and Social Analytics (HUSO 2015)

    Hee, C.V., Lefever, E., Verhoeven, B., Mennes, J., Desmet, B., Pauw, G.D., Daelemans, W., Hoste, V.: Automatic detection and prevention of cyberbullying. In: International Conference on Human and Social Analytics (HUSO 2015). (2015)

  5. [13]

    Text Analytics for Cybersecurity and Online Safety (TA-COS) (2016)

    Killam, R., Cook, P., Stakhanova, N.: Android malware classification through analysis of string literals. Text Analytics for Cybersecurity and Online Safety (TA-COS) (2016)

  6. [14]

    In: Sensors

    Hernandez-Suarez, A., Sanchez-Perez, G., Toscano-Medina, K., Martinez-Hernandez, V., Meana, H.M.P., Olivares-Mercado, J., Sanchez, V.: Social sentiment sensor in twitter for predicting cyber-attacks using𝓁1 regularization. In: Sensors. (2018)

  7. [15]

    Computers & Security (1-2) (2009) 18–28

    García-Teodoro, P., Díaz-Verdejo, J., Maciá-Fernández, G., Vázquez, E.: Anomaly-based network intrusion detection: Techniques, systems and challenges. Computers & Security (1-2) (2009) 18–28

  8. [16]

    International Journal of current Engineering and Scientific Research (IJCESR) (9) (2016) 107–112

    Chaurasia, M.A.: Comparative study of data mining techniques in intrusion dectection. International Journal of current Engineering and Scientific Research (IJCESR) (9) (2016) 107–112

  9. [17]

    Communications Surveys & Tutorials, IEEE (1) (2014) 303–336

    Bhuyan, M.H., Bhattacharyya, D.K., Kalita, J.K.: Network Anomaly Detection: Methods, Systems and Tools. Communications Surveys & Tutorials, IEEE (1) (2014) 303–336

  10. [18]

    ArXiv e-prints (2017) 1–43

    Hodo, E., Bellekens, X., Hamilton, A., Tachtatzis, C., Robert, A.: Shallow and Deep Networks Intrusion Detection System : A Taxonomy and Survey. ArXiv e-prints (2017) 1–43

  11. [19]

    In: Proceedings of the 10th ACM Conference on Computer and Communications Security, New York, NY, USA, ACM (2003) 251–261

    Kruegel, C., Vigna, G.: Anomaly Detection of Web-based Attacks. In: Proceedings of the 10th ACM Conference on Computer and Communications Security, New York, NY, USA, ACM (2003) 251–261

  12. [20]

    ArXiv e-prints (2017)

    Dong, Y., Zhang, Y.: Adaptively Detecting Malicious Queries in Web Attacks. ArXiv e-prints (2017)

  13. [21]

    In: Proceedings of the Anti-phishing Working Groups 2Nd Annual eCrime Researchers Summit

    Abu-Nimeh, S., Nappa, D., Wang, X., Nair, S.: A comparison of machine learning techniques for phishing detection. In: Proceedings of the Anti-phishing Working Groups 2Nd Annual eCrime Researchers Summit. eCrime ’07, New York, NY, USA, ACM (2007) 60–69

  14. [22]

    Prabhakar, D.: A novel method of spam mail detection using text based clustering approach

    Mallikarjunappa, B., R. Prabhakar, D.: A novel method of spam mail detection using text based clustering approach. International Journal of Computer Applications (08 2010)

  15. [23]

    Jabatan Sistem dan Teknologi Komputer, Fakulti Sains Komputer dan Teknologi Maklumat, Universiti Malaya (2010)

    Tee, H., dan Teknologi Komputer, U.M.J.S.: FPGA Unsolicited Commercial Email Inline Filter Design Using Levenshtein Distance Algorithm and Longest Common Subsequence Al- gorithm. Jabatan Sistem dan Teknologi Komputer, Fakulti Sains Komputer dan Teknologi Maklumat, Universiti M...

  16. [24]

    In Weber, R.O., Richter, M.M., eds.: Case-Based Reasoning Research and Development, Berlin, Heidelberg, Springer Berlin Heidelberg (2007) 314–328

    Delany, S.J., Bridge, D.: Catching the drift: Using feature-free case-based reasoning for spam filtering. In Weber, R.O., Richter, M.M., eds.: Case-Based Reasoning Research and Development, Berlin, Heidelberg, Springer Berlin Heidelberg (2007) 314–328

  17. [25]

    Cybernetics and Systems (6-7) (2013) 533–549

    Prilepok, M., Berek, P., Platos, J., Snasel, V.: spam detection using data compression and signatures. Cybernetics and Systems (6-7) (2013) 533–549

  18. [26]

    Bratko, A., Filipič, B., Cormack, G.V., Lynam, T.R., Zupan, B.: Spam filtering using statis- tical data compression models. J. Mach. Learn. Res. (December 2006) 2673–2698

  19. [27]

    In: Proceedings of the 23rd International Conference on World Wide Web

    Thomas, M., Mohaisen, A.: Kindred domains: Detecting and clustering botnet domains using dns traffic. In: Proceedings of the 23rd International Conference on World Wide Web. WWW ’14 Companion, ACM (2014) 707–712

  20. [28]

    In: Proceedings of the 21st USENIX Conference on Security Symposium

    Antonakakis, M., Perdisci, R., Nadji, Y., Vasiloglou, N., Abu-Nimeh, S., Lee, W., Dagon, D.: From throw-away traffic to bots: Detecting the rise of dga-based malware. In: Proceedings of the 21st USENIX Conference on Security Symposium. Security’12, Berkeley, CA, USA, USENIX Asso...

  21. [29]

    CoRR (2016) 11

    Woodbridge, J., Anderson, H.S., Ahuja, A., Grant, D.: Predicting domain generation algo- rithms with long short-term memory networks. CoRR (2016) 11

  22. [30]

    In Traore, I., Woungang, I., Awad, A., eds.: Intelligent, Secure, and Dependable Systems in Distributed and Cloud Environments, Cham, Springer International Publishing (2017) 19–34

    Ahluwalia, A., Traore, I., Ganame, K., Agarwal, N.: Detecting broad length algorithmically generated domains. In Traore, I., Woungang, I., Awad, A., eds.: Intelligent, Secure, and Dependable Systems in Distributed and Cloud Environments, Cham, Springer International Publishing...

  23. [31]

    Expert Systems with Applications124 (2019) 156 – 163

    Selvi, J., Rodríguez, R.J., Soria-Olivas, E.: Detection of algorithmically generated malicious domain names using masked n-grams. Expert Systems with Applications124 (2019) 156 – 163

  24. [32]

    In: Proceedings of the Seventh Symposium on Information and Communication Technology

    Tong, V., Nguyen, G.: A method for detecting dga botnet based on semantic and cluster analysis. In: Proceedings of the Seventh Symposium on Information and Communication Technology. SoICT ’16, New York, NY, USA, ACM (2016) 272–277

  25. [33]

    Al-Rowaily, K., Abulaish, M., Al-Hasan Haldar, N., Al-Rubaian, M.: Bisal - a bilingual sentiment analysis lexicon to analyze dark web forums for cyber security. Digit. Investig. 14(C) (2015) 53–62

  26. [34]

    In: 2014 IEEE Joint Intelligence and Security Informatics Confer- ence

    Weifeng, L., Hsinchun, C.: Identifyingtopsellersinundergroundeconomyusingdeeplearning- based sentiment analysis. In: 2014 IEEE Joint Intelligence and Security Informatics Confer- ence. (2014)

  27. [35]

    CoRR abs/1804.05276 (2018)

    Deb, A., Lerman, K., Ferrara, E.: Predicting cyber events by leveraging hacker sentiment. CoRR abs/1804.05276 (2018)

  28. [36]

    In: ASONAM

    Mittal, S., Das, P.K., Mulwad, V., Joshi, A., Finin, T.: Cybertwitter: Using twitter to generate alerts for cybersecurity threats and vulnerabilities. In: ASONAM. (2016)

  29. [37]

    Ain Shams Engineering Journal (4) (2014) 1093–1113

    Medhat, W., Hassan, A., Korashy, H.: Sentiment analysis algorithms and applications: A survey. Ain Shams Engineering Journal (4) (2014) 1093–1113

  30. [38]

    In: Mining text data

    Liu, B., Zhang, L.: A survey of opinion mining and sentiment analysis. In: Mining text data. Springer (2012) 415–463

  31. [39]

    In: Proceedings of the 38th International ACM SIGIR Conference on Research and Development in Information Retrieval, ACM (2015) 959–962

    Severyn, A., Moschitti, A.: Twitter sentiment analysis with deep convolutional neural net- works. In: Proceedings of the 38th International ACM SIGIR Conference on Research and Development in Information Retrieval, ACM (2015) 959–962

  32. [40]

    In: Proceedings of COLING 2014, the 25th International Conference on Computational Linguistics: Technical Papers

    dos Santos, C., Gatti, M.: Deep convolutional neural networks for sentiment analysis of short texts. In: Proceedings of COLING 2014, the 25th International Conference on Computational Linguistics: Technical Papers. (2014) 69–78

  33. [41]

    In: Proceedings of the 8th international workshop on semantic evaluation (SemEval 2014)

    Tang, D., Wei, F., Qin, B., Liu, T., Zhou, M.: Coooolll: A deep learning system for twit- ter sentiment classification. In: Proceedings of the 8th international workshop on semantic evaluation (SemEval 2014). (2014) 208–212

  34. [42]

    Bzip2 vs

    Tukaani-project: A Quick Benchmark: Gzip vs. Bzip2 vs. LZMA. [Online]. Available from: http://tukaani.org/lzma/benchmarks.html [Accessed 2017-04-06]

  35. [43]

    [Online]

    scikit learn: scikit-learn: machine learning in Python - scikit-learn 0.18.1 documentation. [Online]. Available from:http://scikit-learn.org/stable/ [Accessed 2017-03-29]

  36. [44]

    Security and Communi- cation Networks 8(16) (2015) 2750–2767

    Torrano-Gimenez, C., Nguyen, H.T., Alvarez, G., Franke, K.: Combining expert knowledge with automatic feature extraction for reliable web attack detection. Security and Communi- cation Networks 8(16) (2015) 2750–2767

  37. [45]

    [Online]

    CSIC-dataset: HTTP DATASET CSIC 2010. [Online]. Available from: http://www.isi. csic.es/dataset/ [Accessed 2017-03-29]

  38. [46]

    In: Computational Intelligence in Security for Information Systems

    Nguyen, H.T., Torrano-Gimenez, C., Alvarez, G., Petrović, S., Franke, K.: Application of the generic feature selection measure in detection of web attacks. In: Computational Intelligence in Security for Information Systems. Springer (2011) 25–32

  39. [47]

    In: Proceedings of 52nd annual meeting of the association for computational linguistics: system demonstrations

    Manning, C., Surdeanu, M., Bauer, J., Finkel, J., Bethard, S., McClosky, D.: The stanford corenlp natural language processing toolkit. In: Proceedings of 52nd annual meeting of the association for computational linguistics: system demonstrations. (2014) 55–60 12

  40. [48]

    In: Proceedings of the 11th ACM Symposium on Document Engineering

    Almeida, T.A., Hidalgo, J.M.G., Yamakami, A.: Contributions to the study of sms spam fil- tering: New collection and results. In: Proceedings of the 11th ACM Symposium on Document Engineering. DocEng ’11, New York, NY, USA, ACM (2011) 259–262

  41. [49]

    CoRR (2017)

    Lison, P., Mavroeidis, V.: Automatic detection of malware-generated domains with recurrent neural models. CoRR (2017)

  42. [50]

    Technical report, Stanford University (2009)

    Go, A., Bhayani, R., Huang, L.: Twitter Sentiment Classification using Distant Supervision. Technical report, Stanford University (2009)

  43. [51]

    In: Proceedings of the 49th Annual Meeting of the Association for Computational Linguistics: Human Language Technologies, Association for Computational Linguistics (2011) 142–150

    Maas, A.L., Daly, R.E., Pham, P.T., Huang, D., Ng, A.Y., Potts, C.: Learning word vectors for sentiment analysis. In: Proceedings of the 49th Annual Meeting of the Association for Computational Linguistics: Human Language Technologies, Association for Computational Linguistics...

  44. [52]

    arXiv preprint arXiv:1604.03850 (2016)

    Lillis, D., Becker, B., O’Sullivan, T., Scanlon, M.: Current challenges and future research areas for digital forensic investigation. arXiv preprint arXiv:1604.03850 (2016)

  45. [53]

    Logic Journal of the IGPL (2019) In Press

    de la Torre-Abaitua, G., Lago-Fernández, L., Arroyo, D.: On the application of compression based metrics to identifying anomalous behaviour in web traffic. Logic Journal of the IGPL (2019) In Press

  46. [54]

    IEEE transactions on evolutionary computation (1997) 67–82 13

    Wolpert, D.H., Macready, W.G.: No free lunch theorems for optimization. IEEE transactions on evolutionary computation (1997) 67–82 13

Pith tools

Reviewed August 14, 2026 · model on record in the stance chip above.