{"id":"bb5927d1-5f80-436b-9d01-5a3e11903817","arxiv_id":"2506.01848","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"Using CVE mentions mapped to CAPEC attack patterns, the authors identify forum communities and report that about 4% of actors are high-skill, high-commitment experts.","lead":"This paper maps cybercrime forum posts that mention known vulnerabilities onto attack-pattern categories, then clusters the posters by skill, commitment, and activity. It reports that communities of actors share attack-pattern interests and that roughly 4% of the analyzed population qualify as expert actors.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The key claim rests on an unvalidated proxy: CVE mentions plus MITRE skill ratings are treated as evidence of expertise, and the 4% expert figure is additionally sensitive to the discretionary exclusion of a high-skill, high-commitment cluster.","rationale":"The paper is internally coherent and reports a plausible set of descriptive outputs: modularity 0.473 for the Leiden partition, silhouette 0.569 for k-means with k=8, and clear community labels after content analysis. I do not dispute the numerical pipeline or claim that the authors were careless; the limitations section is unusually candid. The problem is construct validity: the central claim that CVE/CAPEC posts identify technically expert actors requires that a CVE mention plus a MITRE skill-level label be a genuine signal of expertise, and the paper never tests this against any independent ground truth. Because the authors themselves acknowledge the proxy, black-box skill metric, and potential overestimation, the reader's CONDITIONAL verdict is the right verdict. I keep it unchanged: the analysis may well be correct, but it should be released with a validation check before the 4% and amateur-majority findings are used operationally. The additional sensitivity around cluster 1 reinforces the same conclusion: the exact percentage depends on a discretionary cluster-naming decision, so an external validation step is not optional if the headline is to be trusted.","tokens_in":974,"tokens_out":1027,"duration_ms":109176,"concrete_test":"Extract a stratified random sample of posts from the 14 cluster-2 (Professional) actors, the 21 cluster-1 actors, and a matched control sample from low-skill clusters such as cluster 4. Have two cybercrime analysts, blind to cluster labels and hypotheses, independently rate each post on a 3-point scale of demonstrated technical proficiency (0 = bare CVE mention or copy-paste, 1 = contextual discussion, 2 = original exploit code or technical analysis). Pre-specify that the expert label is supported only if the mean rating for cluster 2 is significantly higher than the control sample after accounting for inter-rater agreement. In the same pass, compare cluster 1 with cluster 2; if they are indistinguishable on demonstrated proficiency, recompute the expert proportion with both clusters counted as experts, which changes the headline from about 4% to about 9.7%.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The load-bearing inference is made in Sections III.C and III.F.1: an actor is linked to CAPECs whenever any fetched post mentions a CVE, and the actor's skill level is then taken from MITRE's CAPEC 'Skill Level Required' metric (70th percentile of the actor's assigned values). Nothing checks whether the post demonstrates ability: a one-line CVE mention, copied news, a question, or an undercover researcher's post all produce the same edge and the same skill assignment. Sections III.F.2-III.F.5 then derive commitment from the fraction of CVE posts falling in the actor's Leiden community and cluster actors into 'Professionals' (cluster 2, N=14, 3.9% of 359). The authors concede in Section VII that CVE-CAPEC mapping loses information, that MITRE's metric is a 'black box', and that the expertise measure is proxy-based and static; these concessions highlight exactly where the central claim is weakest. The 4% headline is additionally fragile: cluster 1 (N=21, centroid [2.81; 97.62; 5.14]) has high skill and high commitment by the paper's own variables, yet is renamed Pro-Amateur on the grounds of a short active window; if that discretionary reclassification is removed, 'Professionals' become about 9.7% of the 359-actor sample. The claim that a CVE/CAPEC pipeline isolates a scarce set of genuinely expert actors is therefore not established by the reported analyses.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a CVE/CAPEC-based pipeline to identify technically expert actors in cybercrime forums. It builds a bimodal actor-CAPEC network from posts mentioning CVEs, applies Leiden community detection to find communities of interest, measures each actor's skill level from MITRE's CAPEC 'Skill Level Required' metric, computes commitment as the share of in-community CVE/CAPEC posts, and then applies k-means clustering on skill, commitment, and activity rate. The clusters are interpreted with Bouchard and Nguyen's professional-criminal framework, yielding the headline results that 'professionals' (key expert actors) are about 4% of the 359-actor sample and that amateurs are about 54.87% of the sample. The paper frames this as a method to reduce the population of interest for cyber threat intelligence resource allocation.","tokens_in":16424,"tokens_out":4842,"duration_ms":47465,"significance":"If the method and its central numbers held, the contribution would be practically valuable: it operationalizes technical expertise in forum data at scale, extends key-hacker identification beyond centrality and reputation measures, and offers a concrete scarcity estimate for monitoring. Strengths include the use of a public standardized vulnerability/attack-pattern taxonomy, transparent reporting of many filtering thresholds and cluster centroids, and a candid limitations section. The main quantitative claims are nevertheless built on an unvalidated expertise proxy and on cluster-labeling decisions that are not fully determined by the reported criteria, so the contribution is currently more methodological than empirical. The paper is a reasonable candidate after substantial revision: the core pipeline is coherent, but the evidence supporting the 4% and 54.87% headline figures needs to be materially strengthened.","major_comments":[{"comment":"The central expertise measure is unvalidated. Any post mentioning a CVE creates an actor-CAPEC edge, and the actor's skill level is then read off MITRE's 'Skill Level Required' for the mapped CAPECs; no part of the pipeline checks whether the post demonstrates ability, intent, or understanding. A one-line CVE mention, a question, copied news, or an undercover researcher's post all receive the same edge and the same skill assignment. The limitations section (VII) explicitly concedes CVE-CAPEC information loss, the 'black box' nature of MITRE's metric, and the proxy-based, static character of the expertise measure. Since the 4% and 54.87% headline figures are direct outputs of this proxy, the paper needs at least a manual content-validation subsample and a comparison against a content-based or reputation-based skill signal before those numbers can be treated as estimates of actual expertise.","section":"III.C, III.F.1, VII"},{"comment":"The labeling of cluster 1 as Pro-Amateur is not consistent with the paper's stated definition of professionals. Cluster 1 has centroid [2.81; 97.62; 5.14], i.e., high skill and 97.62% commitment, and differs from cluster 2 (Professionals) essentially in activity rate and in the short active window. The paper's own framework defines professionals as high skill and high commitment, with activity rate added only as a descriptive third variable. Reclassifying cluster 1 as professionals changes the key-expert share from 14/359 = 3.90% to 35/359 = 9.75%, which is a material change to the scarcity claim. The paper needs an explicit, pre-specified decision rule for when short-lived activity overrides the two defining dimensions, or a sensitivity analysis reporting both variants.","section":"V.C, Table IX"},{"comment":"The headline proportions are not shown to be robust to the distribution-driven thresholds. The final 359-actor sample is produced by removing CAPECs with in-degree above 500, dropping actors with fewer than four specialized posts, taking the 70th percentile of each actor's skill list, requiring at least 50% of a post's CAPECs to fall in the actor's community, and selecting eight k-means clusters; several of these choices are justified by the same data ('elbow', 'easier to work with', silhouette score). No sensitivity analysis reports how the 4% professional share or the 54.87% amateur share changes under neighboring thresholds or different values of k. Because the paper's central contribution is a scarcity estimate, a threshold-stability table is needed before the 'tiny proportion' claim is supportable.","section":"III.C.1, III.F.2, V.C"},{"comment":"Assigning each CAPEC its highest Skill Level Required scenario, and then taking the 70th percentile of the resulting list, systematically inflates skill scores. The paper acknowledges the overestimation in Section VII but does not quantify its effect on the final cluster assignments. Since the professional/amateur categories are defined by cutoffs on this inflated scale, an actor's category can change when the scenario rule or the percentile choice is varied; a sensitivity analysis over these choices is necessary to establish that the 4% figure is not an artifact of the skill-scoring rule.","section":"III.F.1"}],"minor_comments":[{"comment":"The text on the Recon community says members 'sit at the top for average number of 76 CAPECs: they have a link with 61 CAPECs,' which is internally inconsistent; Table VIII reports a mean out-degree of 61, so the text should be corrected.","section":"V.B"},{"comment":"The phrase 'Using Slash in Using Slashes' with CAPEC IDs 79, 64, 78, 76 is garbled; please list the CAPEC names and IDs cleanly.","section":"III.C.1"},{"comment":"The column header 'Nb CAPECs' is ambiguous: the table appears to report counts of skill-level values across actors' lists, not just the number of CAPECs per skill level; please clarify the denominator in the caption or the column header.","section":"III.F.1, Table III"},{"comment":"Table I lacks a visible table number/caption in the extracted text, and the citation to Bouchard and Nguyen [5] does not include a year or publisher in the reference list; please complete the bibliographic entry.","section":"II.C, Table I"},{"comment":"Reference [8] appears incomplete ('Chi chi. Two approaches to the study of experts' characteristics'); please update it.","section":"References"},{"comment":"There are several spacing and typographical errors, including 'known asthe key hacker identification problem' in the Introduction; a careful proofreading pass is needed.","section":"Throughout"}],"recommendation":"major_revision","confidential_remarks":"The paper is a lightly edited version of a peer-reviewed IEEE eCrime 2024 paper; the arXiv version should state what, if anything, has changed relative to the published version. The main technical risk is the unvalidated expertise proxy, and I would ask for content-validation and sensitivity analyses before the quantitative claims are accepted. If the original forum data cannot be shared for privacy or legal reasons, a pseudo-data validation or a benchmark against a labeled subset would still substantially improve the support for the headline figures."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"This paper applies standard network and clustering methods to a new measurement problem: identifying technically expert actors in cybercrime forums via CVE mentions mapped to CAPEC attack patterns. That combination—CVE/CAPEC mapping, Leiden community detection, and Bouchard and Nguyen's professional/amateur framework—is genuinely new in the key-hacker identification literature, and the paper is a clear, honest description of a plausible pipeline.\n\nWhat it does well: the methodology is transparent, the thresholds are stated explicitly, and the limitations section is unusually candid. The authors acknowledge CVE-CAPEC information loss, MITRE's skill metric being a black box, the proxy-based and static nature of their expertise measure, and the real possibility that undercover investigators or security analysts get classified alongside malicious actors. The community detection results (modularity 0.473, eight interpretable communities) pass the plausibility test.\n\nThe soft spots are real, though. The load-bearing assumption is that a post mentioning a CVE, and the CAPEC mapped through CWE identifiers, reflects the author's technical skill at that attack pattern. A one-line CVE reference, a copied news item, or a question all produce the same edge and the same skill assignment. Nothing in the paper checks whether the post demonstrates ability, so 'skill level' is more accurately 'interest level.' The skill values themselves come from MITRE's black-box 'Skill Level Required' metric, with the highest scenario assigned and parent/child imputation, which likely overestimates skill.\n\nThe 4% expert figure is also fragile. Cluster 1 (N=21, skill 2.81, commitment 97.62%) has high skill and high commitment by the paper's own variables, but is reclassified as pro-amateur because its members were active for only one day. If that discretionary call is reversed, the professional share jumps from 3.9% to about 9.7%. The authors give a reason (activity rate), but it is not a strong enough reason to carry the headline.\n\nThe paper also releases no data or code, so no external validation is possible. The descriptive claim that amateurs dominate (about 55%) is probably robust to threshold choices, but the scarcity claim that makes the paper interesting is not.\n\nThis is a useful proof-of-concept for threat-intelligence teams, not a validated measurement. It deserves a serious referee, but a reviewer should demand sensitivity analyses and ideally validation against ground truth (known experts, forum roles, qualitative coding of posts).\n\nRecommendation: engage with it, but treat the 4% figure and the expertise proxy as hypotheses, not findings.","headline":"A cleanly written, exploratory pipeline for finding technically specialized forum actors; plausible but the expertise proxy is unvalidated and the 4% headline is fragile to one discretionary cluster reclassification.","tokens_in":16883,"tokens_out":1573,"would_cite":false,"duration_ms":19009,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that matching the CVE vulnerabilities mentioned in cybercrime forum posts to the CAPEC attack patterns they map to reveals coherent communities of actors with shared technical interests, and that the high-skill…","keywords":["cybercrime forums","key hacker identification","technical expertise","CAPEC","CVE mentions","bimodal network","community detection","k-means clustering"],"falsifier":"The central claim would be falsified if qualitative review of the identified expert actors showed that they mostly repost CVE announcements or are security researchers and law enforcement rather than technically engaged attackers, or if the community structure and the 4% expert share disappeared when an alternative CVE-to-CAPEC mapping or a different community-detection method was used.","tokens_in":15799,"feed_emoji":"🎯","tokens_out":5174,"duration_ms":48630,"temperature":0.7,"pith_summary":"This paper tries to establish that technical expertise in cybercrime forums can be measured from the vulnerabilities actors mention in their posts, and that this measurement splits a large forum population into a small group of expert actors worth monitoring. By mapping each mentioned CVE to the corresponding attack pattern in the CAPEC taxonomy, the authors build a two-mode network connecting actors to attack patterns and run community detection. The communities that emerge group actors interested in similar attack techniques, such as privilege escalation, cross-site scripting, reconnaissance, and impersonation. Within those communities, actors are scored on skill, commitment, and activity, and clustered into the categories of a criminological framework. The central result is that high-skill, high-commitment experts make up about 4% of the studied population, while about half are amateurs.","feed_headline":"Forum posts narrow 4,400 hackers to a 4% expert core","feed_subtitle":"Matching CVE mentions to attack patterns reveals small expert communities worth watching in cybercrime forums.","key_machinery":"The central object is the bimodal actor-CAPEC network, in which each forum actor is linked to the attack patterns (CAPECs) corresponding to the CVEs they mention, with the CVE-to-CAPEC mapping performed through shared CWE weakness identifiers. Community detection on this network reveals groups of actors interested in similar attack patterns. Expertise is then operationalized with two facets from a criminological framework: skill level, taken from the 70th percentile of the CAPEC skill-level values associated with an actor, and commitment, the share of an actor's posts that reference their community's attack patterns; a third variable, activity rate, measures posting frequency over the actor's active period. K-means clustering on these three variables partitions the sample into eight clusters interpreted through the professional, pro-amateur, average career criminal, and amateur categories.","core_discovery":"The author's central claim is that the actor-CAPEC bimodal network displays a genuine community structure that groups actors by shared interest in attack patterns, and that key expert actors—those with high skill and high commitment in their community—represent about 4% of the study population. This is established by linking 2,321 actors to 263 attack patterns through CVE mentions, detecting eight communities of interest, and then clustering the 359 actors with enough posts on skill level, commitment, and activity rate. The result is that the four categories of the criminological framework are present, with professionals (the experts) at 3.90%, pro-amateurs at 31.20%, average career criminals at 10.02%, and amateurs at 54.87%. The paper therefore claims that a CVE/CAPEC-based measurement can reduce a large forum population to a small set of technically specialized actors for cyber threat intelligence.","pith_inferences":["The method could be applied to other online communities or marketplaces where CVE mentions appear, with the threshold parameters likely needing re-estimation for each new sample.","The skill measurement is static and could be extended into a longitudinal study of how actors move between the amateur, pro-amateur, and professional categories over time.","Because the paper explicitly notes that cybersecurity analysts and law enforcement may be classified alongside malicious actors, the 4% figure may overstate the number of genuinely malicious experts.","A natural testable extension is to validate the expert labels against independent indicators, such as the sale of exploit code or detailed technical tutorials, rather than only CVE mentions."],"forward_implications":["If correct, threat intelligence teams can reduce a forum population of thousands to a small set of technically specialized actors, about 4%, worth monitoring.","Forums contain coherent groups of actors focused on particular attack patterns, such as privilege escalation or XSS, so intelligence can be organized by attack technique rather than by forum.","About half of the studied population shows little technical expertise, suggesting most forum members are not the primary threat.","The pro-amateur group, roughly 31%, has high skill but low commitment and may be the talent pool from which future experts emerge."],"supporting_citations":[{"why":"Supplies the two-facet skill/commitment framework and the professional, pro-amateur, average career criminal, and amateur classification used to label the clusters.","marker":"[5]"},{"why":"Provides the community-detection algorithm used to partition the actor-CAPEC network into communities of interest.","marker":"[47]"},{"why":"Defines the key hacker identification problem that this study addresses.","marker":"[30]"},{"why":"Provides the modularity measure used to assess the strength of the detected community structure.","marker":"[34]"},{"why":"Supplies the k-means clustering algorithm used to identify key expert actors within the sample.","marker":"[26]"},{"why":"Supports the expectation that key actors form a small fraction of the forum population.","marker":"[29]"},{"why":"Provides the definition of expertise on which the study's conceptualization of technical expertise is based.","marker":"[13]"}],"fun_headline_variants":["CVE-based method spots 4% expert core in cybercrime forums","Community detection finds skilled hackers in forums: only 4% experts","Bimodal network trims 2,321 forum users to 4% technical elites","Technical expertise mapping reveals small expert circle in dark forums"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that an actor mentioning a CVE, and the CAPEC attack pattern mapped to that CVE through CWE identifiers, is a valid proxy for the actor's interest in and technical skill at that attack pattern; if this proxy fails, the communities and expert labels are artifacts of keyword matching.","fun_headline_variants_meta":{"raw":{"variants":["CVE-based method spots 4% expert core in cybercrime forums","Community detection finds skilled hackers in forums: only 4% experts","Bimodal network trims 2,321 forum users to 4% technical elites","Technical expertise mapping reveals small expert circle in dark forums"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000216,"raw_usage":{"total_tokens":1433,"prompt_tokens":947,"completion_tokens":486,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":563,"completion_tokens_details":{"reasoning_tokens":408}},"tokens_in":563,"tokens_out":486,"duration_ms":5381,"temperature":1.0,"reasoning_tokens":408,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T11:32:10.599869+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"The central claim would be falsified if qualitative review of the identified expert actors showed that they mostly repost CVE announcements or are security researchers and law enforcement rather than technically engaged attackers, or if the community structure and the 4% expert share disappeared when an alternative CVE-to-CAPEC mapping or a different community-detection method was used.","supporting_citations":[{"cited_title":"Professionals or amateurs? revisiting the notion of professional crime in the context of cannabis cultivation","cited_arxiv_id":null,"evidence_quote":"Supplies the two-facet skill/commitment framework and the professional, pro-amateur, average career criminal, and amateur classification used to label the clusters."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the community-detection algorithm used to partition the actor-CAPEC network into communities of interest."},{"cited_title":"Mining key- hackers on darkweb forums","cited_arxiv_id":null,"evidence_quote":"Defines the key hacker identification problem that this study addresses."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the modularity measure used to assess the strength of the detected community structure."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the k-means clustering algorithm used to identify key expert actors within the sample."},{"cited_title":"Community finding of malware and exploit vendors on darkweb marketplaces","cited_arxiv_id":null,"evidence_quote":"Supports the expectation that key actors form a small fraction of the forum population."},{"cited_title":"Anders Ericsson","cited_arxiv_id":null,"evidence_quote":"Provides the definition of expertise on which the study's conceptualization of technical expertise is based."}],"review_version":1}