{"id":"54664a78-239b-4461-bb53-b96a867e6090","arxiv_id":"1908.02449","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"low","formal_verification":"none","parameter_count":0,"one_line_summary":"The paper's main contribution is two freely available, manually reviewed datasets of phishing and Tor hidden service website screenshots with labels and hashes.","lead":"This paper releases two labeled datasets of website screenshots: about 460 phishing images and about 37,500 Tor hidden service images. The datasets are meant to help security analysts build tools that automatically classify, search, and correlate website screenshots.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"AIL dataset's label completeness and reliability are unquantified; the paper itself says the classification is partial.","rationale":"The reader's weakest assumption correctly identifies label quality as the load-bearing concern. The paper itself provides direct evidence for this concern: Section 2.2 explicitly admits the AIL classification is partial, and Section 2.3 gives no inter-annotator agreement or other quality metrics. The label files are described as outputs of a manual process with no validation, and the reported label-frequency figures are based on samples rather than the full dataset. Because the primary contribution is the datasets themselves, incomplete or unreliable labels undermine the stated value for training classifiers. The availability of the datasets is a separate, more easily checked prerequisite; even if the URLs are live, the scientific utility of the AIL dataset remains uncertain. The missing citation in Section 3.2.1 is a minor editorial issue, not load-bearing. The reader's CONDITIONAL verdict already accounts for these gaps, so no change is needed. The concrete test of counting labeled files and measuring agreement would resolve the concern empirically.","tokens_in":183,"tokens_out":4351,"duration_ms":54720,"concrete_test":"Download circl-ail-dataset-01 from the provided URL. Count the number of image files (expect about 37,500) and parse the DataTurks label file to compute the proportion of images with at least one label. Independently, have two annotators label a random sample of 50 images using the same MISP dark-web taxonomy and any added labels, then compute Cohen's kappa between each annotator and the provided labels. If the labeled proportion is below 50% or kappa is below 0.6, the 'human-classified' claim is not supported for the full dataset.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim is the availability of two human-verified and human-classified datasets. For the AIL dataset, this claim is insecure. Section 2.2 states the classification 'is partial to date and will be improved and updated as soon as classification operations had been achieved,' but no fraction of labeled versus unlabeled images is reported. The label-frequency figures (Figures 2b and 3a) show only 800 and 9,500 sampled pictures out of roughly 37,500, implying that the provided label file may cover only a small subset. Section 2.3 describes a manual review by a single reviewer ('We manually reviewed datasets') with no inter-annotator agreement, no quality metrics, and no external validation. Because the stated contribution is human-classified data, the utility of the AIL dataset for training classifiers rests on label completeness and reliability that the paper never quantifies. If most images lack labels or if labels are inconsistent, the dataset's value is materially reduced and the central claim is overstated.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper announces two open datasets of website screenshots for security research: circl-phishing-dataset-01 (about 460 phishing screenshots) and circl-ail-dataset-01 (about 37,500 Tor hidden-service screenshots). It describes the data sources (MISP, URLAbuse, AIL), the file formats (JSON hash mappings, DataTurks label exports, VisJS cluster and graph exports), the naming convention based on content hashes, and the manual labeling process using DataTurks and VisJS-Classiificator. A short usage example illustrates fuzzy-hash based image matching on the phishing dataset, and future work mentions extending the datasets and improving the AIL classification. The authors state explicitly that the main contribution is the availability of the two human-verified and human-classified datasets, not the research results shown in the usage examples.","tokens_in":7461,"tokens_out":1305,"duration_ms":15264,"significance":"If the two datasets are actually downloadable, correctly described by the provided hash files, and accompanied by usable label files, the paper serves a useful community function: it addresses a genuine gap, since few open datasets of website screenshots exist specifically for phishing and Tor hidden services. The hash-to-filename reference JSON is a concrete reproducibility aid, and the MISP dark-web taxonomy-based labels give the AIL dataset a standard vocabulary. The paper is also honest in scope: it does not claim algorithmic novelty and explicitly labels the matching experiments as preliminary. The central weakness is that the papers stated contribution is the availability of human-classified data, yet the completeness and quality of that classification are never measured, especially for the AIL dataset.","major_comments":[{"comment":"The central claim is the availability of two human-classified datasets, but for circl-ail-dataset-01 the paper reports only that the classification 'is partial to date' and provides no fraction of labeled versus unlabeled images. The label-frequency figures are based on 800 and 9,500 sampled pictures out of roughly 37,500, which suggests the label file may cover only a minority of the images. The authors should state explicitly how many of the 37,500 pictures have at least one label, how many are unlabeled, and whether the label file contains entries for all pictures or only for the labeled subset.","section":"Section 2.2, Figures 2b and 3"},{"comment":"The manual classification is described as done by 'we' with no indication of the number of annotators, no inter-annotator agreement, and no quality metric. Since the paper claims human classification as a contribution, at minimum the authors should report the number of reviewers, the level of consistency (e.g., a random re-labeling sample or a second pass), and the criteria used to resolve ambiguous cases. Without such information, the reliability of the labels as ground truth for training classifiers is unquantified.","section":"Section 2.3"},{"comment":"The usage example is presented as evidence that the dataset supports image-matching research, but it contains no quantitative evaluation: there are no numbers of true/false matches, no comparison of algorithms, and no measure of retrieval accuracy. This would be acceptable if the section were explicitly framed as an illustration only, but the current text claims that structures 'can be easily detected' and content 'can be matched with some confidence.' Please either add quantitative results (e.g., precision/recall on a labeled subset) or rephrase the claims as purely anecdotal.","section":"Section 4 and Figures 5a/5b"}],"minor_comments":[{"comment":"The phrase 'human-verified and human-classified' is used as the main characterization of both datasets, but the AIL dataset is only partially classified. Please qualify the claim in the abstract, e.g., 'human-verified, and human-classified for the phishing dataset and partially human-classified for the AIL dataset.'","section":"Abstract and Section 1.1"},{"comment":"The collision-handling description is unclear: it says bytes are 'temporary adding bytes to each colliding file,' which suggests modifying the file content to resolve a hash-name collision. If the file bytes are changed, the hash-to-file mapping would change as well. Please clarify whether this operation modifies the picture bytes or only the filename input to the codename generator.","section":"Section 2.2"},{"comment":"The two added labels are written differently in the text ('error page' and 'other') and in the annex ('error_page' and 'other'). Please unify the notation so users can parse the label file unambiguously.","section":"Section 3.2.1 and Annex 7.1.5"},{"comment":"The citation for the MISP dark-web taxonomy is missing a key or full reference; the text shows 'taxonomy9[?]'. Please add the complete reference for the MISP taxonomies repository.","section":"References"},{"comment":"Several minor typos and formatting issues occur, including 'screen-captures' inconsistency in the title, '10000' without thousands separator, 'classiﬁcation' with unusual ligatures, and the mention of 'Douglas-Quaid' and 'Carl-Hauser' as footnotes rather than as proper references. These do not affect the technical content.","section":"Throughout"}],"recommendation":"major_revision","confidential_remarks":"The paper is a dataset-announcement resource paper. Its acceptance hinges on the datasets actually being accessible and the label files being complete enough for the stated purpose. I checked that the URLs are given and the formats are specified, but I could not verify downloadability from the manuscript text alone. The classification-completeness issue for the AIL dataset is the main risk; if the authors can report the labeled/unlabeled counts, the paper would be much stronger. I would not reject the paper for the lack of inter-annotator metrics, but the authors should add at least a basic quality statement because the contribution is explicitly the human classification."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Bottom line: the two public datasets are the contribution, and they appear to be real and usable. The paper is a resource announcement, not an experimental evaluation, and judged on that basis it mostly holds up. The soft spot is the AIL label file: the abstract calls the data “human-verified and human-classified,” but the paper later says the classification is partial and gives no fraction of images labeled. The label-frequency figures are computed on 800 and 9,500 sampled pictures out of roughly 37,500, so I would want the authors to state clearly how many of the archive's images actually have labels in the provided JSON. Section 2.3 says “we manually reviewed datasets” with a single reviewer; no inter-annotator agreement or quality metric is given. That matters because the value of a labeled dataset for training classifiers depends on label reliability. None of this invalidates the resource, but the abstract overstates the current state of the AIL labels.\n\nThe phishing dataset is small (about 460 images) but fine for what it is; the label lists are transparent, and the hashes and JSON formats are a nice touch. The usage examples are preliminary and appropriately labeled as such. Citation pattern is fine; self-citations to CIRCL tools are expected since the data came from those tools. One minor knock: references [2] and [4] are missing bibliographic details (venue/journal), easily fixed. I did not re-download the archives, so I cannot independently confirm the pictures are all there, but the description and hashes are concrete enough.\n\nWho is this for? People building visual classifiers or correlators for phishing and onion sites; for them the dataset is a useful starting point. The AIL corpus is large and could enable work that previously had no public benchmark. A serious referee should ask for the missing label-completeness numbers and a label-quality estimate, but the paper deserves that review rather than a desk reject.","headline":"The datasets are a genuine and useful public resource, but the abstract overstates the completeness and reliability of the AIL labels.","tokens_in":7904,"tokens_out":1957,"would_cite":false,"duration_ms":19941,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper's main contribution is not an algorithm but the release of two human-verified, human-classified datasets of phishing and Tor hidden-service screenshots, each image keyed by cryptographic hashes and supplied with labels for…","keywords":["phishing","Tor hidden services","onion domains","screenshot datasets","visual image matching","image classification","open data","threat intelligence"],"falsifier":"Re-annotate a random sample of a few hundred images from each dataset with independent labelers and compute agreement (for example, Cohen's kappa) against the published labels; if agreement is low on common classes such as 'other' and 'error page', the datasets' usefulness as ground truth collapses.","tokens_in":7114,"feed_emoji":"🎣","tokens_out":7232,"duration_ms":73143,"temperature":0.7,"pith_summary":"The paper's central claim is that it provides ready-made, human-verified datasets for visual correlation and automatic classification of phishing and Tor hidden-service screenshots. One set holds roughly 460 screenshots of phishing pages; the other holds roughly 37,500 screenshots of onion domains scraped from the Tor network. Each image has a stable hash triplet and a human-readable name, and the phishing set comes with two independent label schemes plus a graph of clusters. The author argues that these labeled images are the missing raw material for training automatic tools that would let security analysts correlate screenshots instead of doing it by hand.","feed_headline":"38,000 labeled screenshots of phishing and dark-web sites go public","feed_subtitle":"Human-verified labels and hashes let analysts train classifiers to correlate phishing and onion pages.","key_machinery":"The load-bearing object is the dataset archive itself: screenshots named by hashing each file's bytes into a human-readable adjective-triple, shipped with a JSON reference mapping names to MD5, SHA1, and SHA256, plus label files that map filenames to classes. That naming scheme makes picture mentions stable across tools, while the label files provide the ground truth an automatic classifier would train on; the hash reference makes the corpus auditable and de-duplicable.","core_discovery":"On the paper's own terms, the discovery is the availability, not a classification algorithm: two free and open datasets that have been reviewed picture by picture, stripped of personal or harmful content, labelled by hand, and shipped with per-file MD5, SHA1, and SHA256 hashes. The phishing dataset (about 460 images) is labelled twice, once with a collaborative tagging tool and once with a graph-based classifier that also yields cluster membership; the dark-web dataset (about 37,500 images) carries partial labels drawn from a public dark-web taxonomy. The paper demonstrates on a small sample that fuzzy-hash matching against these screenshots can group pages by visual structure or theme, but explicitly frames that as a glimpse, not the contribution.","pith_inferences":["Extension: because each screenshot ships with a hash triplet, independent researchers could relabel a random subset and quantify agreement, giving the corpus an inter-annotator reliability measure the paper leaves out.","Extension: the partially labeled dark-web set could be exploited with label propagation from the labeled images, letting researchers train coarse topic classifiers without full manual annotation.","Extension: the same hash-readable naming and labeling conventions transfer naturally to screenshots from malware sandboxes or email attachments, broadening visual correlation beyond phishing and onion pages.","Extension: if the label noise is low, simple visual features alone should separate broad classes such as login forms, marketplaces, and error pages; this is a direct test readers can run against the published labels."],"forward_implications":["Analysts can download both archives and use the label files as training ground truth for screenshot classifiers without collecting or labeling their own corpus.","The phishing set's brand and login-form labels support clustering that can track a phishing brand across domain changes.","The dark-web taxonomy labels allow automatic triage of onion sites by topic, motivation, and structure, including scam, marketplace, and login-form categories.","The per-file hashes let a future tool de-duplicate or verify screenshots and let different teams compare results on identical images.","Even partial automation of screenshot classification would reduce the manual correlation burden on security teams."],"supporting_citations":[{"why":"Supplies the open-source analysis framework whose daily scraping generated the roughly 37,500 onion-domain screenshots in the dark-web dataset.","marker":"[1]"},{"why":"Establishes the prior result that screenshots can be classified with convolutional networks, the line of work the datasets are meant to feed.","marker":"[2]"},{"why":"Demonstrates classification-mining techniques for phishing-website prediction, one of the only prior phishing-specific benchmarks this dataset complements.","marker":"[3]"},{"why":"Provides the keypoint-feature phishing-detection result that the visual-matching usage example extends and benchmarks against.","marker":"[4]"},{"why":"Supplies the collaborative threat-intelligence platform whose public phishing events were a source of screenshots, and whose taxonomy the dark-web labels follow.","marker":"[5]"}],"fun_headline_variants":["Free open dataset: 38k screenshots of phishing and dark-web sites","Hand-labeled screenshots of phishing and onion sites go open","37,500 dark-web and 460 phishing screenshots released openly","Open dataset: verified screenshots for phishing and Tor research","38,000 human-checked screenshots of phishing and .onion sites"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the hand-made labels, made by one reviewer without documented agreement checks, are accurate enough to serve as training ground truth; if the labels are noisy or inconsistent, classifiers trained on them inherit that noise.","fun_headline_variants_meta":{"raw":{"variants":["Free open dataset: 38k screenshots of phishing and dark-web sites","Hand-labeled screenshots of phishing and onion sites go open","37,500 dark-web and 460 phishing screenshots released openly","Open dataset: verified screenshots for phishing and Tor research","38,000 human-checked screenshots of phishing and .onion sites"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000559,"raw_usage":{"total_tokens":2576,"prompt_tokens":781,"completion_tokens":1795,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":397,"completion_tokens_details":{"reasoning_tokens":1704}},"tokens_in":397,"tokens_out":1795,"duration_ms":13525,"temperature":1.0,"reasoning_tokens":1704,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T14:42:51.961169+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-annotate a random sample of a few hundred images from each dataset with independent labelers and compute agreement (for example, Cohen's kappa) against the published labels; if agreement is low on common classes such as 'other' and 'error page', the datasets' usefulness as ground truth collapses.","supporting_citations":[],"review_version":1}