{"id":"02894ff5-885e-4377-a0de-fab15757468e","arxiv_id":"1908.05897","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Across 367 phishing papers in the ACM Digital Library, only 51 include user studies, and most of those omit participant demographics, weakening their external validity.","lead":"Researchers reviewed phishing papers in the ACM Digital Library and found that only about 14 percent of them actually tested or interviewed real users. They also found that most of those user studies failed to report basic details about participants, such as age, gender, or race, which makes it hard to trust or compare the results.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 13.9% prevalence uses the raw 367-paper search count as denominator; after the paper's own exclusion criteria leave 253 relevant papers, the figure should be 51/253 ≈ 20.2%.","rationale":"The most load-bearing problem is the denominator mismatch. The paper's central quantitative claim is the 13.9% figure, and it is computed in a way that contradicts §3.1. The reader's weakest assumption about ACM DL representativeness is a legitimate scope concern, but the authors explicitly define their corpus as ACM DL, so that concern is about how far conclusions generalize rather than whether the within-corpus arithmetic is right. The denominator issue is more immediate: even granting the ACM-only scope, the reported percentage is not what the method produces. This is not a matter of external consensus; it is an internal inconsistency between §3.1 and the abstract/conclusion. The corrected 20.2% still supports the qualitative narrative (user studies are a minority), so the paper need not be rejected, but the headline statistic must be fixed. This reinforces the reader's CONDITIONAL verdict: the paper is acceptable only with a corrected denominator and a statement of what population the 13.9% refers to. I therefore leave the verdict as UNCHANGED, and I partially agree with the reader because they noted the denominator inconsistency in their rationale even though they selected ACM corpus as the weakest assumption.","tokens_in":10151,"tokens_out":6668,"duration_ms":60419,"concrete_test":"Recompute the prevalence from §3.1: take the 367 search results, apply the stated exclusion criteria to verify the 253 count, and divide 51 by the verified number. If 51/253 ≈ 20.2%, the abstract's 13.9% is an artifact. As a second check, confirm that none of the 51 user-focused papers came from the 114 excluded records; if any did, the numerator and denominator are misaligned. This is a purely arithmetic/screening reproducibility check requiring no new data.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 3.1 reports that an ACM DL search for 'phishing' returned 367 papers, and after excluding extended abstracts, non-English works, and papers not primarily about phishing, 253 relevant papers remained. The 51 user-focused papers were identified from that screened set (Section 3.1). Yet the abstract and Section 6 compute the prevalence as 51/367 = 13.9%. This mixes the screened numerator with an unscreened denominator, so the headline statistic is not the share of user-focused work among relevant phishing papers in the corpus. Using the paper's own inclusion criteria, the rate is 51/253 ≈ 20.2%. The gap is not trivial: the claim 'only 13.9%' is the paper's main quantitative contribution, and choosing the raw search count over the screened count makes the field look less user-focused than the review's own screening says. The qualitative findings about underreporting and demographic gaps are less affected, but the central prevalence number, as stated in the abstract, is not derivable from the described method. A corrected statement should either use 253 as the denominator or explicitly justify why the broad search count is the intended population.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper reports a systematic literature review of phishing research in the ACM Digital Library. The authors searched ACM DL for the term \"phishing,\" obtained 367 records, screened these to 253 relevant papers after excluding extended abstracts, non-English works, and papers not primarily about phishing, and then identified 51 papers containing user studies. Using thematic coding by three researchers, they analyze the technical and individual attributes studied, the reported demographics of participants, and the methodologies used. The central claim is that only 13.9% of phishing papers in the corpus are user-focused and that even these papers frequently omit participant age, gender, and race/ethnicity information. The paper concludes with recommendations for improved demographic reporting and more representative recruitment in future phishing research.","tokens_in":10474,"tokens_out":8126,"duration_ms":78480,"significance":"The review addresses an important gap: the extent to which phishing research validates tools and findings with human participants. The three-coder thematic coding process with an inter-coder reliability of 87.9% is a clear strength, and the detailed counts of demographic reporting gaps (e.g., only 3 of 51 papers report race/ethnicity) are useful for the community. However, the headline prevalence of 13.9% is not derivable from the described method: the 51 user-focused papers were selected from the 253 screened papers, so the relevant rate is about 20.2%. This is a central, fixable error. The absence of the included-study list and codebook further limits verification. If the prevalence is corrected and the corpus scope is carefully qualified, the paper would be a useful resource.","major_comments":[{"comment":"The headline prevalence statistic is inconsistent with the screening procedure. Section 3.1 states that the ACM DL search returned 367 papers, that exclusions left 253 relevant papers, and that the 51 user-focused papers were identified from those 253. The abstract's figure of 13.9% equals 51/367, which mixes a screened numerator with an unscreened denominator; the rate among relevant papers is 51/253 ≈ 20.2%. Section 6's phrase \"13.9% of relevant published papers\" is directly contradicted by this arithmetic. Please recompute the prevalence, or explicitly justify why the unscreened 367 is the intended population, and adjust all prevalence statements in the abstract, findings, and conclusion accordingly.","section":"Abstract; §3.1; §6"},{"comment":"The reporting-quality findings are internally inconsistent. Section 4.4 states that 37 of the 51 papers \"did not mention the age range, gender distribution, or racial/ethnic backgrounds of the participants,\" but Section 4.4.1 then says that 14 of those 37 papers included some kind of age range. If 14 papers reported age range, then age range was missing in a different set of 37 papers. Please clarify which count refers to which demographic attribute and provide a table of reporting counts for age, gender, race/ethnicity, and participant number.","section":"§4.4 and §4.4.1"},{"comment":"The review is not reproducible in its current form: there is no list of the 51 included papers, no codebook for the thematic codes, and no flow diagram showing the number of papers excluded at each stage. This matters because the central counts (51, 253, and the demographic subcounts) cannot be checked without this material. Please add an appendix or supplementary file containing the included-study list, the coding scheme, and the exclusion flow.","section":"§3.1–§3.2"},{"comment":"The corpus is drawn exclusively from one publisher's database, but the title and several sentences in Sections 5 and 6 describe the article as reviewing \"phishing research\" or \"user-centered phishing research\" without the qualifier \"in the ACM Digital Library.\" Because the prevalence figure can vary with the choice of venues (human-factors venues versus security-engineering venues), the paper should either restrict all unqualified prevalence and trend claims to ACM DL-indexed papers or include a comparison with other venues and an explicit limitations paragraph discussing corpus-selection bias.","section":"§3.1; §5; §6"}],"minor_comments":[{"comment":"The inter-coder reliability metric is not defined; please state whether the reported 23.7%, 56.7%, and 87.9% are percentage agreement, Cohen's kappa, or another coefficient, and describe when in the reconciliation process the final ICR was measured.","section":"§3.2"},{"comment":"The search protocol is underspecified: only the single term \"phishing\" is mentioned, and no date range or database-specific filters are reported. Please state whether variant terms (e.g., \"spear phishing,\" \"phish\") were considered and how the 2004–2018 window was defined.","section":"§3.1"},{"comment":"For clarity, replace the fractions \"about ⅓\" and \"a little over ⅔\" with their exact values (15/51 ≈ 29% and 37/51 ≈ 73%).","section":"§4.4"},{"comment":"Figure 2 would be easier to read with labeled axes and the exact number of publications per year displayed above each bar.","section":"Figure 2"},{"comment":"Several references are incomplete or lack publication venues, for example Das et al. (2017) has no venue and Colarik and Janczewski (2007) is a URL-only citation; the reference list should be completed according to the journal's style.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The denominator error is the main obstacle and appears correctable; the demographic-reporting findings are likely to survive the correction. I would encourage the editor to require the supplementary dataset, including the list of included studies and the codebook, as a condition of revision."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper is a systematic review of user-focused phishing research in the ACM Digital Library, and its real value is the coding taxonomy and the detailed accounting of what the 51 user studies actually did. The finding that two-thirds of those studies omit basic demographics, and that only a handful report race/ethnicity, is concrete and worth citing. The inter-coder reliability process is described honestly, including the low first-pass ICR and the iterative discussion that brought it to 87.9%. That is more transparency than most reviews give you.\n\nThe soft spot is the headline number. Section 3.1 says the ACM DL search returned 367 papers, exclusions left 253 relevant papers, and the 51 user-focused papers were identified from that screened set. The abstract and conclusion then report 51/367 = 13.9%. If the denominator is the relevant corpus, the rate is 51/253 ≈ 20.2%. The gap is not cosmetic: the 13.9% figure is the paper's central quantitative claim, and it makes the field look less user-focused than the authors' own screening implies. The conclusion even says \"13.9% of relevant published papers,\" which is not what the arithmetic shows. This needs a correction, either by using 253 as the denominator or by explicitly justifying why the raw search count is the intended population.\n\nTwo other soft spots, both minor relative to the fixable denominator problem. The corpus is single-publisher, so any generalization to \"phishing research\" as a whole is not supported; the authors should either soften the language or add a comparison to IEEE, USENIX, and SOUPS proceedings they already cite. And the coded dataset is not provided, which makes the thematic counts hard to verify. Neither undermines the qualitative findings about demographic underreporting, but both limit the paper's evidentiary weight.\n\nWho gets value from this: people working on usable security or phishing interventions who want a quick map of what user studies exist in ACM venues and what they tend to overlook. The taxonomy of technical vs. individual attributes and the recruitment-bias observations are the genuinely useful parts. The self-citations are background and not a problem.\n\nI would send this to peer review. The denominator mistake is load-bearing but fixable, and the review itself is careful enough that a revision with a corrected statistic and a narrower scope statement would be a solid contribution. If you are the editor, ask for the corrected denominator, a public dataset or at least a full coding table, and a disclaimer about the ACM-only scope.","headline":"A useful systematic review of user-focused phishing research in ACM DL, but the headline 13.9% prevalence figure is wrong because it divides the screened numerator by the unscreened search count; the corrected rate is about 20.2%.","tokens_in":10873,"tokens_out":1229,"would_cite":true,"duration_ms":13988,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Only 13.9% of phishing research studies the people being phished, a systematic review finds.","keywords":["phishing","user studies","systematic literature review","usable security","participant demographics","social engineering","human factors","recruitment bias"],"falsifier":"Repeat the same search, exclusion, and coding protocol on a different publisher's digital library covering the same years; if the user-study share is substantially higher or demographic reporting is routine there, the paper's characterization of phishing research does not generalize.","tokens_in":9942,"feed_emoji":"🎣","tokens_out":5413,"duration_ms":50101,"temperature":0.7,"pith_summary":"This paper asks how much phishing research actually studies the people being phished. Reviewing papers in a major computing publisher's digital library, it finds that most phishing research does not: only 13.9% of 367 papers used user-study methods such as interviews, surveys, or in-lab tests. Even among those 51 papers, about a third omit participant counts and more than two-thirds omit age, gender, or race/ethnicity. If the review is right, the field's defenses and training tools are being built and tested without strong evidence about the humans they are meant to protect.","feed_headline":"Only 13.9% of phishing research studies real users","feed_subtitle":"A systematic review finds most phishing papers skip human participants and many omit demographics.","key_machinery":"The central mechanism is a systematic literature review protocol: a keyword search, exclusion criteria for non-relevant or non-English items, and independent thematic coding by three researchers, with inter-coder reliability raised to 87.9% after discussion rounds. The codebook groups phishing-attack attributes into technical attributes, individual attributes, benefits, and threats; this coding determines which papers count as user-focused and what those user studies actually measure.","core_discovery":"The paper claims that, in the phishing literature indexed by the digital library it searched, only 51 of 367 papers (13.9%) center users through interviews, surveys, or in-lab studies, and that even these 51 frequently omit participant counts, age, gender, and race/ethnicity. The authors read this as evidence that user-focused phishing research remains a small and underreported fraction of the field, concentrated in human-computer interaction venues and leaning toward usability testing of tools rather than understanding users' mental models and behaviors. The review therefore calls for routine demographic reporting and for recruiting participant pools that mirror the populations phishers actually target.","pith_inferences":["If the pattern extends beyond the searched library, the 13.9% figure could understate user studies published in venues that emphasize human factors, so the field's true user focus might be higher than reported.","The demographic reporting gaps suggest that reviewers and venues should adopt reporting checklists requiring participant counts, age, gender, and ethnicity before acceptance.","The concentration of user studies in human-computer interaction venues implies that technical security venues may rarely see user evidence, which could explain why tools are evaluated without users.","A testable extension would be to run the same review protocol on a different publisher's digital library to see whether the underreporting is a venue-specific norm or a field-wide one."],"forward_implications":["Future phishing research should report participant demographics as a matter of course.","Researchers should recruit participant pools that mirror the populations targeted by phishers, with balanced gender, wider age ranges, and varied racial, ethnic, and cultural backgrounds.","More work should combine surveys and usability tests with interviews and qualitative analysis to understand user mental models.","Training and warning tools, however well-engineered, lack a demonstrated basis in user evidence until tested with representative users.","The field should treat user studies as a core component of phishing defense, not an optional supplement to technical solutions."],"supporting_citations":[{"why":"Supplies the origin of the term 'phishing' and the 1996 timeline used to date the field.","marker":"Kay, 2004"},{"why":"Anchors the first user-centered phishing study in the dataset, marking the start of the trend line.","marker":"Garfinkel and Miller, 2005"},{"why":"Supplies the result that a good phishing site fools 90% of participants, evidence that visual deception defeats users.","marker":"Dhamija et al., 2006"},{"why":"Provides the age-related usability finding used to argue demographics matter in reporting.","marker":"Chadwick-Dias et al., 2003"},{"why":"Provides the gender-related website-identification finding used to argue gender distribution should be reported.","marker":"Djamasbi et al., 2007"},{"why":"A demographic analysis of phishing susceptibility, used as an example of survey-based user research.","marker":"Sheng et al., 2010"},{"why":"Example of a controlled phishing attack combined with interviews, one of the interview-based studies counted.","marker":"Kumaraguru et al., 2007"},{"why":"Survey-based study of phishing training providers, used as an example of survey methodology.","marker":"Wash and Cooper, 2018"}],"fun_headline_variants":["Phishing papers skip users: 86% omit human studies","Only 1 in 7 phishing studies involve real users","Systematic review: phishing research lacks user focus","13.9% of phishing papers study users, rest don't","Phishing literature rarely tests humans, review finds"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The paper assumes the single digital library it searched is a fair sample of phishing research as a whole; if that library overrepresents tool-building venues, the 13.9% user-study share is an artifact of the corpus.","fun_headline_variants_meta":{"raw":{"variants":["Phishing papers skip users: 86% omit human studies","Only 1 in 7 phishing studies involve real users","Systematic review: phishing research lacks user focus","13.9% of phishing papers study users, rest don't","Phishing literature rarely tests humans, review finds"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000137,"raw_usage":{"total_tokens":1102,"prompt_tokens":848,"completion_tokens":254,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":464,"completion_tokens_details":{"reasoning_tokens":174}},"tokens_in":464,"tokens_out":254,"duration_ms":3209,"temperature":1.0,"reasoning_tokens":174,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T13:01:00.634735+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Repeat the same search, exclusion, and coding protocol on a different publisher's digital library covering the same years; if the user-study share is substantially higher or demographic reporting is routine there, the paper's characterization of phishing research does not generalize.","supporting_citations":[],"review_version":1}