Pith. sign in

REVIEW 2 cited by

Malicious URL Detection using Machine Learning: A Survey

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1701.07179 v3 pith:43QTNR7F submitted 2017-01-25 cs.LG cs.CR

classification cs.LGcs.CR
keywords maliciouslearningmachinedetectionresearchsurveyarticleblacklists
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Malicious URL, a.k.a. malicious website, is a common and serious threat to cybersecurity. Malicious URLs host unsolicited content (spam, phishing, drive-by exploits, etc.) and lure unsuspecting users to become victims of scams (monetary loss, theft of private information, and malware installation), and cause losses of billions of dollars every year. It is imperative to detect and act on such threats in a timely manner. Traditionally, this detection is done mostly through the usage of blacklists. However, blacklists cannot be exhaustive, and lack the ability to detect newly generated malicious URLs. To improve the generality of malicious URL detectors, machine learning techniques have been explored with increasing attention in recent years. This article aims to provide a comprehensive survey and a structural understanding of Malicious URL Detection techniques using machine learning. We present the formal formulation of Malicious URL Detection as a machine learning task, and categorize and review the contributions of literature studies that addresses different dimensions of this problem (feature representation, algorithm design, etc.). Further, this article provides a timely and comprehensive survey for a range of different audiences, not only for machine learning researchers and engineers in academia, but also for professionals and practitioners in cybersecurity industry, to help them understand the state of the art and facilitate their own research and practical applications. We also discuss practical issues in system design, open research challenges, and point out some important directions for future research.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. URL2Graph++: Unified Semantic-Structural-Character Learning for Malicious URL Detection

    cs.CR 2025-09 conditional novelty 5.0 of 10

    URL2Graph++ fuses BERT semantics, character CNN features, and dual word/character co-occurrence graphs to report state-of-the-art malicious URL detection on three public datasets.

  2. Phishing Webpage Detection: Unveiling the Threat Landscape and Investigating Detection Techniques

    cs.CR 2025-09 conditional novelty 2.0 of 10

    A survey categorizing phishing webpage detection into URL, content, and visual approaches, with an analysis of research gaps and suggested directions.

Pith tools