Pith. sign in

REVIEW 4 major objections 6 minor 62 references

An Analysis of Architectural and Operational Dynamics of Phishkits in the Wild

T0 review · 4 major / 6 minor · reviewed 2026-08-10 · deepseek-v4-flash

Pith's one-line read This paper analyzes 1,300 phishing kits and argues that heavy code reuse and known evasion tricks make most phishkit-built sites predictable enough to detect at scale.

desk verdict A valuable descriptive census of phishkit exfiltration channels and evasion artifacts, whose reuse-rate figure needs methodological transparency before the broad predictability claim can be taken at face value. read the letter →

arxiv 2608.07451 v1 pith:OJATJJKV submitted 2026-08-07 cs.CR cs.NIcs.SI

classification cs.CRcs.NIcs.SI
keywords phishkitsphishingattacksevasionandcloakingcodereuseTelegrambotexfiltrationtrafficattributionstructuralsimilaritydetection
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper analyzes 1,300 phishing kits observed between 2020 and 2023 to understand what makes phishing pages cheap to build and hard to catch. It reports that the kits' core modules—page templates, traffic filtering, and stolen-data exfiltration—are heavily reused, with 276 kits sharing 100% structurally identical PHP layouts, 126 kits using identical IP-blocking code, and 284 kits using no evasion at all. It also finds that messaging-bot APIs have become the main exfiltration channel: 292 bot tokens embedded in the live-feed kits produced 8,635 leaked messages with credentials, card numbers, and SMS codes when polled over 101 days. The paper argues that because most phishkits are this predictable, structural-similarity and known-evasion-pattern detection can catch a large share of phishkit-built websites at scale.

What carries the argument

The load-bearing object is the three-module structure common to modern phishkits: an app template loader that clones a legitimate login page, a backdoor that sends stolen data to the operator over email or messaging-bot APIs, and a traffic-attribution layer that inspects IP address, ISP, user agent, hostname, or geolocation before serving the fake page. The argument is carried by a structural-similarity pipeline that parses each kit's PHP source into control-flow graphs, embeds those graphs as vectors using the Node2Vec method, and clusters them with hierarchical agglomerative clustering, so near-identical layouts become visible as clusters rather than as superficial template differences. That pipeline, together with shared blocklists and repeated bot tokens, is what transforms "kits look different" into "the ecosystem is homogeneous and predictable."

What would settle it

Run the same structural clustering on an independent, larger collection of phishkits gathered from multiple commercial feeds and underground forums over a longer window; if the proportion of kits with 100% structurally identical PHP layouts falls well below the 21% observed here, or if most new kits carry bespoke evasion code, the predictability claim is contradicted.

Watch

Extended reading notes

Core claim

The paper's central claim is that the phishkit ecosystem is far less sophisticated than many accounts suggest: the components that keep a phishing page alive are nearly identical across kits and have not changed critically over the 2020–2023 observation window. The evidence is a set of large overlapping measurements: 600 of 1,300 kits (46.1%) filter incoming traffic by IP address, 312 (24%) block by ISP, 498 call a geolocation service, and 126 kits ship the same evasion file set. Structural analysis of the PHP control-flow graphs puts 276 kits (21%) in clusters with 100% identical structure, with the largest cluster of 17 banking kits sharing 82–100% similarity. On the operational side, 292 Telegram bot tokens appeared in the 732 live-feed kits, and controlled polling of those bots over 101 days captured 8,635 messages, including 370 credit-card numbers and 509 username/password pairs. From these observations the paper concludes that most phishkits manifest similar, predictable behavior and that unsupervised or semi-supervised structure-aware detection can exploit that predictability at scale.

Load-bearing premise

The conclusion depends on the 1,300 collected kits—732 from a single commercial feed covering about three months and 589 from one public repository—being representative of the broader phishkit ecosystem; if more diverse or sophisticated kits escape collection, the true ecosystem is less predictable than the paper claims.

Editorial extensions

If this is right

  • A detector that compares the control-flow structure of a live phishing page against known kit clusters could flag a substantial share of sites without retraining on each new campaign.
  • Because 284 kits (21.8%) apply no evasion at all, ordinary crawler-based blocklist scanning should already catch them, and a structural detector would catch the other kits that only hide behind IP and user-agent filtering.
  • The widespread, near-identical use of messaging-bot exfiltration means that tracking and revoking the embedded bot tokens and email addresses can disrupt data theft across many campaigns at once.
  • The stability of core components over the study window means behavioral signatures built from current kits are likely to retain value for some time, reducing the need for constant retraining.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • One consequence the paper leaves implicit is that defenders can treat a single discovered kit as a starting point for a whole family: because kit code is reused heavily, one cluster's shared artifacts can be used to hunt for its other deployments, effectively turning each new sample into multiple detections.
  • A testable extension would be to measure how quickly the predictability advantage decays: after this analysis is public, adversarial kit authors may diversify their evasion code, and the share of structurally identical kits in a fresh post-2023 sample would show whether the ecosystem responds to published measurement.
  • Another unstated implication is that the lower-bound caveat cuts both ways: the same homogeneity that makes detection easy also means a single published countermeasure, such as blocking the observed messaging-bot infrastructure, could force kit authors to rework a large fraction of their operational chain at once.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 6 minor

Summary. The paper presents an observational study of 1,300 phishkits collected from 2020 to 2023, drawing on a commercial daily feed (732 kits) and a public repository (589 kits). It analyzes backdoor mechanisms (email and Telegram), traffic attribution and cloaking, structural code reuse over time, and characterizes the ecosystem as highly predictable and therefore easier to defend against at scale. The authors also report responsible handling of leaked victim data, including IRB approval and disclosure to a credit-card issuer.

Significance. If the quantitative claims hold, the paper provides useful longitudinal evidence on the commoditization of phishing infrastructure. The study's strengths include a real-world dataset spanning two collection sources, live interaction with Telegram bots, a clear limitations section, and internally consistent arithmetic. The claim that heavy code reuse makes phishkit-based phishing predictable is relevant for detection research. However, the headline code-reuse measurements depend on unspecified clustering hyperparameters and an unstratified dataset, so the quantitative support for the main claim needs substantial clarification before the paper is publication-ready.

major comments (4)
  1. [Section VI (Analysis of Phishkits Over the Years)] The central quantitative claim that 276 kits (21%) have 'structurally identical layouts and 100% structural similarity' is not auditable as reported. The pipeline uses Node2Vec embeddings and agglomerative clustering, but the embedding dimension, random-walk parameters, distance threshold, and the definition of the '100%' cutoff are not specified. Since cluster membership is used to assert '100% structural similarity,' the metric is partly circular. Please report all hyperparameters and the exact similarity threshold, or replace this with a direct pairwise code-similarity measure (e.g., normalized AST/CFG edit distance) computed with a fixed threshold. Also state whether the 21% figure is stable under reasonable parameter variation.
  2. [Section III (Methodology) and Abstract] The dataset is dominated by 732 kits from a single commercial feed collected over a three-month window (Dec 21, 2022–Mar 20, 2023) and 589 kits from one public repository. The abstract and conclusion generalize to 'a large number of cases' and 'detection ... at scale' without the Section VIII lower-bound caveat. If the commercial feed preferentially surfaces kits that match known phishkit families, the code-reuse estimates are inflated. Please report the similarity statistics separately for the two sources and by collection period, and show that the largest clusters are not an artifact of the short feed window. At a minimum, the abstract should carry the same 'lower-bound estimate' qualification as Section VIII.
  3. [Section VIII (Limitations) vs. Section VI] The first paragraph of Section VIII states that 'our approach for creating clusters was based on the network address and the company names,' with judgments made from reverse DNS lookups. This conflicts with the clustering methodology in Section VI, which builds clusters from Node2Vec embeddings of PHP control-flow graphs. The sentence appears to describe the sector classification in Table III, not the code-similarity clusters, but as written it undermines the methodological narrative. Please rewrite this passage to disambiguate the cluster-construction process for code similarity from the classification of target sectors.
  4. [Section IV (Backdoor Mechanisms) and Ethics] The paper states in the Ethics subsection that 'we did not perform any analysis of the leaked data,' yet Section IV reports 370 leaked credit card instances, 425 SMS verification codes, and 509 username/password pairs, and the abstract summarizes these counts. These statements are inconsistent unless the authors distinguish 'categorization for responsible disclosure' from 'analysis for research.' If the counts were needed for disclosure, please state that explicitly and describe how the counting was performed without using the PII for research. If the counts came from a different workflow, clarify that workflow.
minor comments (6)
  1. [References] References [21] and [22] are duplicates of the same paper (Bijmans et al., USENIX Security 2021) and should be merged.
  2. [Abstract and Section V] The abstract reports 284 phishkits (21.8%) with no evasion mechanism, while Section V reports 600 kits with IP blocking and 439 with user-agent attribution. Please clarify the overlap between these categories so the reader can follow how the 284 figure is derived with respect to the other 600 and 439 counts.
  3. [Section VI] The largest cluster is described as having similarity scores 'ranging from 82%–100%' after the paper says all 276 kits in similar clusters had '100% structural similarity.' Please clarify whether the 82–100% range refers to pairwise similarities within the cluster and whether the 100% figure applies only to a subset or to a particular centroid-based definition.
  4. [Section VI] The selection of Node2Vec over spectral embedding and GNNs is justified qualitatively, but no evaluation of the chosen set of features (call graph and CFG) is reported. A short comparison of clustering quality across the three approaches would make the choice reproducible.
  5. [Section III] The phrase 'After removing duplicate and irrelevant samples, and accounting for 21 kits shared between the two sources' should specify how 'irrelevant' was determined and whether 'duplicate' was based on hashes, filenames, or a stricter comparison.
  6. [Section I] The introduction's contribution list is brief and does not mention the longitudinal analysis or the code-reuse clustering, which are the paper's most distinctive and risky parts. Consider expanding the contribution list to state these explicitly.

Circularity Check

1 steps flagged · score 2.0 of 10

Clustering similarity statistic is mildly self-referential, but the main findings are direct observations and not circular.

  1. self definitional [Section VI, 'Analysis of Phishkits Over the Years', clustering paragraph]
    "We computed pairwise cosine distance on the embeddings and applied agglomerative clustering [34] to generate clusters of similar phishkits. ... the generated clusters show that 276 kits (21%) had structurally identical layouts and 100% structural similarity."

    The clusters are the output of a similarity-threshold procedure on the same Node2Vec embeddings, so reporting that the clusters contain structurally identical kits is partly a restatement of the clustering criterion rather than an independent estimate of code reuse. The 21% figure is not a prediction or an external validation; it describes the structure imposed by agglomerative clustering of cosine distances. This is self-definitional in the limited sense that the 'similar clusters' are defined by the pairwise similarity that is then cited as evidence of similarity. The circularity is mild because the paper also reports direct file-level evidence of reuse, such as 126 phishkits using identical evasion files, so the central conclusion does not stand or fall on this statistic alone.

full rationale

This is an observational measurement study rather than a derivation. The main findings, including IP blocklisting behavior, Telegram bot use, email-based exfiltration, and the presence or absence of evasion mechanisms, are direct code-level observations that do not depend on any fitted model or predictive parameter. No parameter is fit to a subset and then renamed as a prediction, and no uniqueness theorem or ansatz is imported from the authors' prior work; self-citations appear only in related-work context and are not load-bearing. The single mildly circular element is the Section VI clustering statistic, where the 'structurally identical' result is produced by the same Node2Vec-cosine-agglomerative procedure used to define the clusters, making the 21% figure partly a restatement of the similarity criterion. However, the paper's broader code-reuse conclusion is corroborated by independent direct observations, including 126 phishkits sharing identical evasion files, identical IP blocklists, and repeated Telegram bot tokens. The collection-bias concern raised about the commercial feed is a validity limitation rather than circularity, and Section VIII explicitly frames the conclusions as a lower-bound estimate of the full phishkit ecosystem. Overall, the paper's central claim is not forced by its own definitions or by self-citation, so the circularity score is low.

Assumptions & free parameters 4 free parameters · 5 assumptions · 0 invented entities

The central claims rest on a manually assembled dataset, hand-chosen analysis parameters, and several background assumptions about the completeness of static analysis and the representativeness of the sample. These are not unreasonable for an observational security study, but they limit independent verification.

free parameters (4)
  • Node2Vec embedding dimension and walk parameters = not reported
    The structural similarity clusters in Section VI depend on these hyperparameters; without them the clustering result is not reproducible.
  • Agglomerative clustering distance threshold = not reported
    The number of clusters (171) and the 'structurally identical' claim (276 kits) depend on the chosen cutoff, which is not stated.
  • Regex patterns for evasion detection = not reported
    The classification of 600 IP-blocking kits and 284 no-evasion kits (Section V) relies on manually curated patterns that are not fully enumerated.
  • Similarity cutoff for '100% structural similarity' = not reported
    The claim that 276 kits have identical layouts requires a similarity threshold that the paper does not define.
assumptions (5)
  • domain assumption PHP static analysis via AST/CFG captures the functionally relevant behavior of phishkits.
    Section VI focuses on PHP source code; if evasion or exfiltration logic lives in JavaScript or other files, the measurements would be incomplete.
  • domain assumption The collected dataset is representative enough as a lower-bound estimate of the wild phishkit ecosystem.
    Section VIII states this is a lower-bound estimate and that undetected kits may be more sophisticated; the conclusion generalizes from this sample.
  • standard math Node2Vec embeddings and agglomerative clustering yield meaningful structural similarity groups.
    Section VI applies these unsupervised methods without reporting validation or hyperparameter choices.
  • domain assumption Polling Telegram bots with getUpdates does not materially alter the bots' behavior or the exfiltration sample.
    Section IV describes polling every 3 hours for 101 days; this interaction could affect message delivery or trigger bot shutdown.
  • domain assumption Regex-based detection of evasion mechanisms is complete enough to classify kits as having 'no evasion mechanism'.
    Section V identifies 284 kits with no evasion based on the analyzed patterns; unknown patterns could miss more subtle evasion.

how reviews work

0 comments
Cite this review

Pith. "Pith review of An Analysis of Architectural and Operational Dynamics of Phishkits in the Wild." pith.science (2026). https://pith.science/paper/OJATJJKV

@misc{pith2026260807451,
  author       = {Pith},
  title        = {Pith review of: An Analysis of Architectural and Operational Dynamics of Phishkits in the Wild},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/OJATJJKV}},
  note         = {Machine review of arXiv:2608.07451}
}
read the original abstract

Phishing attacks have always been a favored vector for adversaries to defraud users, bypass modern defense mechanisms, and penetrate critical systems. Among all the elements contributing to the creation and deployment of successful phishing attacks, phishkits stand out as a crucial parameter. Phishkits often facilitate creating and deploying compelling phishing pages, implement evasion strategies, and establish and maintain backdoors with remote adversaries for exchanging leaked data. In this work, we performed an analysis of 1,300 modern phishkits collected from 2020 to 2023. We analyzed the architecture, source code, communication channels, and the nature of leaked data shared with adversaries. We identified mechanisms for dynamic redirection and attributing incoming web traffic as part of the evasion and cloaking mechanism. We also observed heavy reliance on current messaging services for exchanging stolen data with phishers. That said, our analysis shows that the number of phishkits with advanced functionalities is quite small. We identified 284 (21.8%) phishkits that did not use any form of evasion mechanism. We also observed that while there were differences in the implementation details of phishkits, the major components that keep phishing pages functional were very similar or even identical across kits. The level of code reuse and heavy reliance on known tricks to build pre-packaged phishing pages make a large number of cases predictable, which can potentially make the detection of these adversarial operations even easier at scale.

Figures

Figures reproduced from arXiv: 2608.07451 by the authors.

Figure 1
Figure 1. Fraudulent website created using a phishkit. These fake websites are clones of the legitimate entity and contain a login form. If users input their [PITH_FULL_IMAGE:figures/full_fig_p003_1.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

62 extracted references · 53 canonical work pages

  1. [23]

    What happens after you leak your password: Understanding credential sharing on phishing sites,

    P. Peng, C. Xu, L. Quinn, H. Hu, B. Viswanath, and G. Wang, “What happens after you leak your password: Understanding credential sharing on phishing sites,” inProceedings of the 2019 ACM Asia Conference on Computer and Communications Security, ser. Asia CCS ’19. New York, NY , USA: Association for Computing Machinery, 2019, p. 181–192. [Online]. Available...

  2. [21]

    Catching phishers by their bait: Investigating the dutch phishing landscape through phishing kit detection,

    H. Bijmans, T. Booij, A. Schwedersky, A. Nedgabat, and R. van Wegberg, “Catching phishers by their bait: Investigating the dutch phishing landscape through phishing kit detection,” in30th USENIX Security Symposium (USENIX Security 21). USENIX Association, Aug. 2021, pp. 3757–3774. [Online]. Available: https://www.usenix.org/conference/usenixsecurity21/pre...

  3. [22]

    Catching phishers by their bait: Investigating the dutch phishing landscape through phishing kit detection,

    ——, “Catching phishers by their bait: Investigating the dutch phishing landscape through phishing kit detection,” in30th USENIX Security Symposium (USENIX Security 21), 2021, pp. 3757–3774

  4. [1]

    Top Phishing Statistics for 2023: Latest Figures and Trends,

    StationX - Gary Smith, “Top Phishing Statistics for 2023: Latest Figures and Trends,” https://www.stationx.net/phishing-statistics/, Accessed: 10- 12-2023

  5. [2]

    2023 Phishing Report Reveals 47.2% Surge in Phishing At- tacks Last Year,

    ZScaler, “2023 Phishing Report Reveals 47.2% Surge in Phishing At- tacks Last Year,” https://www.zscaler.com/blogs/security-research/2023- phishing-report-reveals-47-2-surge-phishing-attacks-last-year, Accessed: 10-12-2023

  6. [3]

    New Phishing-as-a-Service Platform Lets Cybercriminals Generate Convincing Phishing Pages,

    The Hacker News, “New Phishing-as-a-Service Platform Lets Cybercriminals Generate Convincing Phishing Pages,” https://thehackernews.com/2023/05/new-phishing-as-service-platform- lets.html, Accessed: 10-12-2023

  7. [4]

    Researchers Uncover Thriving Phishing Kit Market on Telegram Channels,

    ——, “Researchers Uncover Thriving Phishing Kit Market on Telegram Channels,” https://thehackernews.com/2023/04/researchers- uncover-thriving-phishing.html, Accessed: 10-12-2023

  8. [5]

    Phishing as a service continues to plague business users,

    SiliconAngle, “Phishing as a service continues to plague business users,” https://siliconangle.com/2023/08/31/phishing-service-continues- plague-business-users/, Accessed: 10-12-2023

Show all 62 references
  1. [6]

    A machine-learning based unbiased phishing detection approach,

    H. Shirazi, L. Zweigle, and I. Ray, “A machine-learning based unbiased phishing detection approach,” inProceedings of the 17th International Joint Conference on e-Business and Telecommunications (ICETE 2020)- SECRYPT, 2020

  2. [7]

    Survey and taxonomy of adversarial reconnaissance techniques,

    S. Roy, N. Sharmin, J. C. Acosta, C. Kiekintveld, and A. Laszka, “Survey and taxonomy of adversarial reconnaissance techniques,”ACM Computing Surveys, vol. 55, no. 6, pp. 1–38, 2022

  3. [8]

    Adversarial autoencoder data synthesis for enhancing machine learning- based phishing detection algorithms,

    H. Shirazi, S. R. Muramudalige, I. Ray, A. P. Jayasumana, and H. Wang, “Adversarial autoencoder data synthesis for enhancing machine learning- based phishing detection algorithms,”IEEE Transactions on Services Computing, 2023

  4. [9]

    Knowledge ex- pansion and counterfactual interaction for{Reference-Based}phishing detection,

    R. Liu, Y . Lin, Y . Zhang, P. H. Lee, and J. S. Dong, “Knowledge ex- pansion and counterfactual interaction for{Reference-Based}phishing detection,” in32nd USENIX Security Symposium (USENIX Security 23), 2023, pp. 4139–4156

  5. [10]

    Spacephish: The evasion- space of adversarial attacks against phishing website detectors using machine learning,

    G. Apruzzese, M. Conti, and Y . Yuan, “Spacephish: The evasion- space of adversarial attacks against phishing website detectors using machine learning,” inProceedings of the 38th Annual Computer Security Applications Conference, 2022, pp. 171–185

  6. [11]

    Uncovering the cloak: A systematic review of techniques used to conceal phishing websites,

    W. Li, S. Manickam, S. U. A. Laghari, and Y .-W. Chong, “Uncovering the cloak: A systematic review of techniques used to conceal phishing websites,”IEEE Access, 2023. 8

  7. [12]

    Crawlphish: Large-scale analysis of client-side cloaking techniques in phishing,

    P. Zhang, A. Oest, H. Cho, Z. Sun, R. Johnson, B. Wardman, S. Sarker, A. Kapravelos, T. Bao, R. Wanget al., “Crawlphish: Large-scale analysis of client-side cloaking techniques in phishing,” in2021 IEEE Symposium on Security and Privacy (SP). IEEE, 2021, pp. 1109–1124

  8. [13]

    Stargazer: Long-term and multiregional measurement of timing/geolocation-based cloaking,

    S. Fujii, T. Sato, S. Aoki, Y . Tsuda, N. Kawaguchi, T. Shigemoto, and M. Terada, “Stargazer: Long-term and multiregional measurement of timing/geolocation-based cloaking,”IEEE Access, 2023

  9. [14]

    I’m spartacus, no, i’m spartacus: Proactively protecting users from phishing by inten- tionally triggering cloaking behavior,

    P. Zhang, Z. Sun, S. Kyung, H. W. Behrens, Z. L. Basque, H. Cho, A. Oest, R. Wang, T. Bao, Y . Shoshitaishviliet al., “I’m spartacus, no, i’m spartacus: Proactively protecting users from phishing by inten- tionally triggering cloaking behavior,” inProceedings of the 2022 ACM S...

  10. [15]

    OpenPhish.com,

    OpenPhish, “OpenPhish.com,” https://openphish.com/, Accessed: 10-12- 2023

  11. [16]

    Phishhunt.io,

    Phishhunt, “Phishhunt.io,” https://phishunt.io/, Accessed: 10-12-2023

  12. [17]

    Phishing kits collected from phishunt,

    Daniel Lopez, “Phishing kits collected from phishunt,” https://github.com/0xDanielLopez/phishing kits, Accessed: 10-12- 2023

  13. [18]

    Google Chrome Privacy Whitepaper,

    Google, “Google Chrome Privacy Whitepaper,” https://www.google.com/chrome/privacy/whitepaper.html#malware, Accessed: 10-12-2023

  14. [19]

    There is no free phish: An analysis of

    M. Cova, C. Kruegel, and G. Vigna, “There is no free phish: An analysis of” free” and live phishing kits.”WOOT, vol. 8, pp. 1–8, 2008

  15. [20]

    Phisheye: Live monitoring of sandboxed phishing kits,

    X. Han, N. Kheir, and D. Balzarotti, “Phisheye: Live monitoring of sandboxed phishing kits,” inProceedings of the 2016 ACM SIGSAC Conference on Computer and Communications Security, 2016, pp. 1402–1413

  16. [24]

    Telegram Bot Platform,

    Telegram, “Telegram Bot Platform,” https://telegram.org/blog/bot- revolution, Accessed: 10-12-2023

  17. [25]

    Ev- eryone is Different: Client-side Diversification for Defending Against Extension Fingerprinting,

    E. Trickel, O. Starov, A. Kapravelos, N. Nikiforakis, and A. Doupe, “Ev- eryone is Different: Client-side Diversification for Defending Against Extension Fingerprinting,” inProceedings of the USENIX Security Symposium, 2019

  18. [26]

    Uncovering cloaking web pages with hybrid detection approaches,

    J. Deng, H. Chen, and J. Sun, “Uncovering cloaking web pages with hybrid detection approaches,” in2013 International Symposium on Computational and Business Intelligence. IEEE, 2013, pp. 291–296

  19. [27]

    Cloaker catcher: a client-based cloaking detection system,

    R. Duan, W. Wang, and W. Lee, “Cloaker catcher: a client-based cloaking detection system,”arXiv preprint arXiv:1710.01387, 2017

  20. [28]

    Surveylance: Automatically detecting online survey scams,

    A. Kharraz, W. Robertson, and E. Kirda, “Surveylance: Automatically detecting online survey scams,” in2018 IEEE Symposium on Security and Privacy (SP). IEEE, 2018, pp. 70–86

  21. [29]

    Dial one for scam: Analyzing and detecting technical support scams,

    N. Miramirkhani, O. Starov, and N. Nikiforakis, “Dial one for scam: Analyzing and detecting technical support scams,”CoRR, vol. abs/1607.06891, 2016. [Online]. Available: http://arxiv.org/abs/1607.06891

  22. [30]

    Google Safe Browsing,

    Google, “Google Safe Browsing,” https://safebrowsing.google.com/, Ac- cessed: 10-12-2023

  23. [31]

    Spectral embedding of graph networks,

    S. Deutsch and S. Soatto, “Spectral embedding of graph networks,”CoRR, vol. abs/2009.14441, 2020. [Online]. Available: https://arxiv.org/abs/2009.14441

  24. [32]

    node2vec: Scalable feature learning for networks,

    A. Grover and J. Leskovec, “node2vec: Scalable feature learning for networks,”CoRR, vol. abs/1607.00653, 2016. [Online]. Available: http://arxiv.org/abs/1607.00653

  25. [33]

    Graph neural networks: A review of methods and applications,

    J. Zhou, G. Cui, Z. Zhang, C. Yang, Z. Liu, and M. Sun, “Graph neural networks: A review of methods and applications,”CoRR, vol. abs/1812.08434, 2018. [Online]. Available: http://arxiv.org/abs/1812.08434

  26. [34]

    M. L. Zepeda-Mendoza and O. Resendis-Antonio,Hierarchical Agglomerative Clustering. New York, NY: Springer New York, 2013, pp. 886–887. [Online]. Available: https://doi.org/10.1007/978-1-4419- 9863-7 1371

  27. [35]

    A lustrum of malware network communication: Evolution and insights,

    C. Lever, P. Kotzias, D. Balzarotti, J. Caballero, and M. Antonakakis, “A lustrum of malware network communication: Evolution and insights,” in 2017 IEEE Symposium on Security and Privacy (SP). IEEE, 2017, pp. 788–804

  28. [36]

    SimilarWeb,

    SimilarWeb, “SimilarWeb,” https://www.similarweb.com/, Accessed: 10- 12-2023

  29. [37]

    Social engineering attacks: A survey,

    F. Salahdine and N. Kaabouch, “Social engineering attacks: A survey,” Future Internet, vol. 11, no. 4, p. 89, 2019

  30. [38]

    An overview of social engineering malware: Trends, tactics, and implications,

    S. Abraham and I. Chengalur-Smith, “An overview of social engineering malware: Trends, tactics, and implications,”Technology in Society, vol. 32, no. 3, pp. 183–196, 2010. [Online]. Available: https://www.sciencedirect.com/science/article/pii/S0160791X10000497

  31. [39]

    Measuring Pay-per-Install: The commoditization of malware distribution,

    J. Caballero, C. Grier, C. Kreibich, and V . Paxson, “Measuring Pay-per-Install: The commoditization of malware distribution,” in20th USENIX Security Symposium (USENIX Security 11). San Francisco, CA: USENIX Association, Aug. 2011. [Online]. Available: https://www.usenix.org/c...

  32. [40]

    The dropper effect: Insights into malware distribution with downloader graph analytics,

    B. J. Kwon, J. Mondal, J. Jang, L. Bilge, and T. Dumitras ¸, “The dropper effect: Insights into malware distribution with downloader graph analytics,” inProceedings of the 22nd ACM SIGSAC Conference on Computer and Communications Security, ser. CCS ’15. New York, NY , USA: Ass...

  33. [41]

    Real-time detection of malware downloads via large-scale url->file->machine graph mining,

    B. Rahbarinia, M. Balduzzi, and R. Perdisci, “Real-time detection of malware downloads via large-scale url->file->machine graph mining,” inProceedings of the 11th ACM on Asia Conference on Computer and Communications Security, ser. ASIA CCS ’16. New York, NY , USA: Assoc...

  34. [42]

    An End-to-End Analysis of Covid-Themed Scams in the Wild,

    B. Ousat, M. A. Tofighi, and A. Kharraz, “An End-to-End Analysis of Covid-Themed Scams in the Wild,” inACM ASIA Conference on Computer and Communications Security (ASIACCS ’23), Melbourne, Australia, July 2023

  35. [43]

    Constructs of deceit: exploring nuances in modern social engineering attacks,

    M. A. Tofighi, B. Ousat, J. Zandi, E. Schafir, and A. Kharraz, “Constructs of deceit: exploring nuances in modern social engineering attacks,” inInternational conference on detection of intrusions and malware, and vulnerability assessment. Springer, 2024, pp. 107–127

  36. [44]

    Cloak of visibility: Detecting when machines browse a different web,

    L. Invernizzi, K. Thomas, A. Kapravelos, O. Comanescu, J.-M. Picod, and E. Bursztein, “Cloak of visibility: Detecting when machines browse a different web,” inProceedings of the 37th IEEE Symposium on Security and Privacy, 2016

  37. [45]

    Inside a phisher’s mind: Understanding the anti-phishing ecosystem through phishing kit analysis,

    A. Oest, Y . Safei, A. Doup ´e, G.-J. Ahn, B. Wardman, and G. Warner, “Inside a phisher’s mind: Understanding the anti-phishing ecosystem through phishing kit analysis,” in2018 APWG Symposium on Electronic Crime Research (eCrime), 2018, pp. 1–12

  38. [46]

    Phishfarm: A scalable framework for measuring the effectiveness of evasion techniques against browser phishing blacklists,

    A. Oest, Y . Safaei, A. Doup ´e, G.-J. Ahn, B. Wardman, and K. Tyers, “Phishfarm: A scalable framework for measuring the effectiveness of evasion techniques against browser phishing blacklists,” in2019 IEEE Symposium on Security and Privacy (SP), 2019, pp. 1344–1361

  39. [47]

    A good fishman knows all the angles: A critical evaluation of google’s phishing page classifier,

    C. Miao, J. Feng, W. You, W. Shi, J. Huang, and B. Liang, “A good fishman knows all the angles: A critical evaluation of google’s phishing page classifier,” inProceedings of the 2023 ACM SIGSAC Conference on Computer and Communications Security, ser. CCS ’23. New York, NY , US...

  40. [48]

    Cookieless Monster: Exploring the Ecosystem of Web-Based Device Fingerprinting,

    N. Nikiforakis, A. Kapravelos, W. Joosen, C. Kruegel, F. Piessens, and G. Vigna, “Cookieless Monster: Exploring the Ecosystem of Web-Based Device Fingerprinting,” in2013 IEEE Symposium on Security and Privacy, SP 2013, Berkeley, CA, USA, May 19-22, 2013. IEEE Computer Society,...

  41. [49]

    ”How Unique Is Your Web Browser?

    P. Eckersley, “”How Unique Is Your Web Browser?”,” inPrivacy Enhancing Technologies, M. J. Atallah and N. J. Hopper, Eds. Berlin, Heidelberg: Springer Berlin Heidelberg, 2010, pp. 1–18

  42. [50]

    Device fingerprinting for augmenting web authentication: Classification and analysis of methods,

    F. Alaca and P. C. van Oorschot, “Device fingerprinting for augmenting web authentication: Classification and analysis of methods,” inProceed- ings of the 32nd Annual Conference on Computer Security Applications, ser. ACSAC ’16. New York, NY , USA: Association for Computing Ma...

  43. [51]

    The web never forgets: Persistent tracking mechanisms in the wild,

    G. Acar, C. Eubank, S. Englehardt, M. Juarez, A. Narayanan, and C. Diaz, “The web never forgets: Persistent tracking mechanisms in the wild,” inProceedings of the 2014 ACM SIGSAC Conference on Computer and Communications Security, ser. CCS ’14. Association for Computing Machin...

  44. [52]

    Cloak of visibility: Detecting when machines browse a different web,

    L. Invernizzi, K. Thomas, A. Kapravelos, O. Comanescu, J.-M. Picod, and E. Bursztein, “Cloak of visibility: Detecting when machines browse a different web,” in2016 IEEE Symposium on Security and Privacy (SP). IEEE, 2016, pp. 743–758

  45. [53]

    Understanding, measuring, and detecting modern technical support scams,

    J. Liu, P. Pun, P. Vadrevu, and R. Perdisci, “Understanding, measuring, and detecting modern technical support scams,” in2023 IEEE 8th European Symposium on Security and Privacy (EuroS&P). IEEE, 2023, pp. 18–38

  46. [54]

    A comprehensive study cyber attacks and countermeasures,

    S. Nirwan and B. K. Dhaliwal, “A comprehensive study cyber attacks and countermeasures,” in2023 International Conference on Inventive Computation Technologies (ICICT). IEEE, 2023, pp. 1103–1107

  47. [55]

    Phish- ing attacks in social engineering: A review

    K. S. Adu-Manu, R. K. Ahiable, J. K. Appati, and E. E. Mensah, “Phish- ing attacks in social engineering: A review.”Journal of Cybersecurity (2579-0072), vol. 4, no. 4, 2022

  48. [56]

    Large-scale analysis of pop-up scam on typosquatting urls,

    T. Dam, L. D. Klausner, D. Buhov, and S. Schrittwieser, “Large-scale analysis of pop-up scam on typosquatting urls,” inProceedings of the 14th International Conference on Availability, Reliability and Security, 2019, pp. 1–9

  49. [57]

    Beyond phish: Toward detecting fraudulent e-commerce websites at scale,

    M. Bitaab, H. Cho, A. Oest, Z. Lyu, W. Wang, J. Abraham, R. Wang, T. Bao, Y . Shoshitaishvili, and A. Doup ´e, “Beyond phish: Toward detecting fraudulent e-commerce websites at scale,” in2023 IEEE Symposium on Security and Privacy (SP). IEEE Computer Society, 2023, pp. 2566–2583

  50. [58]

    Understanding the domain registration behavior of spammers,

    S. Hao, M. Thomas, V . Paxson, N. Feamster, C. Kreibich, C. Grier, and S. Hollenbeck, “Understanding the domain registration behavior of spammers,” inProceedings of the 2013 Internet Measurement Conference, IMC 2013, Barcelona, Spain, October 23-25, 2013, K. Papagiannaki, P. K...

  51. [59]

    A framework for detection and measurement of phishing attacks,

    S. Garera, N. Provos, M. Chew, and A. D. Rubin, “A framework for detection and measurement of phishing attacks,” inProceedings of the 2007 ACM workshop on Recurring malcode, 2007, pp. 1–8

  52. [60]

    Techniques and solutions for addressing ransomware attacks,

    A. Kharraz, “Techniques and solutions for addressing ransomware attacks,” Ph.D. dissertation, Northeastern University, 2017

  53. [61]

    Devphish: Exploring social engineer- ing in software supply chain attacks on developers,

    H. Siadati, S. Jafarikhah, E. Sahin, T. Hernandez, E. Tripp, D. Khryashchev, and A. Kharraz, “Devphish: Exploring social engineer- ing in software supply chain attacks on developers,” in2024 IEEE 15th Annual Ubiquitous Computing, Electronics & Mobile Communication Conference (...

  54. [62]

    The shift and surge of phishing during the covid-19 pandemic: An analysis of over 1100 targeted top-level domains used by dutch firms,

    R. Hoheisel, G. van Capelleveen, D. Sarmah, and M. Junger, “The shift and surge of phishing during the covid-19 pandemic: An analysis of over 1100 targeted top-level domains used by dutch firms,”Available at SSRN 4215129, 2022. 10

Pith tools

Reviewed August 10, 2026 · model on record in the stance chip above.