Pith. sign in

REVIEW 4 major objections 5 minor 1 cited by

All About Phishing: Exploring User Research through a Systematic Literature Review

T0 review · 4 major / 5 minor · reviewed 2026-08-14 · deepseek-v4-flash

Pith's one-line read Only 13.9% of phishing research studies the people being phished, a systematic review finds.

desk verdict A useful systematic review of user-focused phishing research in ACM DL, but the headline 13.9% prevalence figure is wrong because it divides the screened numerator by the unscreened search count; the corrected rate is about 20.2%. read the letter →

arxiv 1908.05897 v1 pith:B2I6LJLY submitted 2019-08-16 cs.CR cs.CYcs.HC

classification cs.CRcs.CYcs.HC
keywords phishinguserstudiessystematicliteraturereviewusablesecurityparticipantdemographicssocialengineeringhumanfactorsrecruitmentbias
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper asks how much phishing research actually studies the people being phished. Reviewing papers in a major computing publisher's digital library, it finds that most phishing research does not: only 13.9% of 367 papers used user-study methods such as interviews, surveys, or in-lab tests. Even among those 51 papers, about a third omit participant counts and more than two-thirds omit age, gender, or race/ethnicity. If the review is right, the field's defenses and training tools are being built and tested without strong evidence about the humans they are meant to protect.

What carries the argument

The central mechanism is a systematic literature review protocol: a keyword search, exclusion criteria for non-relevant or non-English items, and independent thematic coding by three researchers, with inter-coder reliability raised to 87.9% after discussion rounds. The codebook groups phishing-attack attributes into technical attributes, individual attributes, benefits, and threats; this coding determines which papers count as user-focused and what those user studies actually measure.

What would settle it

Repeat the same search, exclusion, and coding protocol on a different publisher's digital library covering the same years; if the user-study share is substantially higher or demographic reporting is routine there, the paper's characterization of phishing research does not generalize.

Watch

Extended reading notes

Core claim

The paper claims that, in the phishing literature indexed by the digital library it searched, only 51 of 367 papers (13.9%) center users through interviews, surveys, or in-lab studies, and that even these 51 frequently omit participant counts, age, gender, and race/ethnicity. The authors read this as evidence that user-focused phishing research remains a small and underreported fraction of the field, concentrated in human-computer interaction venues and leaning toward usability testing of tools rather than understanding users' mental models and behaviors. The review therefore calls for routine demographic reporting and for recruiting participant pools that mirror the populations phishers actually target.

Load-bearing premise

The paper assumes the single digital library it searched is a fair sample of phishing research as a whole; if that library overrepresents tool-building venues, the 13.9% user-study share is an artifact of the corpus.

Editorial extensions

If this is right

  • Future phishing research should report participant demographics as a matter of course.
  • Researchers should recruit participant pools that mirror the populations targeted by phishers, with balanced gender, wider age ranges, and varied racial, ethnic, and cultural backgrounds.
  • More work should combine surveys and usability tests with interviews and qualitative analysis to understand user mental models.
  • Training and warning tools, however well-engineered, lack a demonstrated basis in user evidence until tested with representative users.
  • The field should treat user studies as a core component of phishing defense, not an optional supplement to technical solutions.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the pattern extends beyond the searched library, the 13.9% figure could understate user studies published in venues that emphasize human factors, so the field's true user focus might be higher than reported.
  • The demographic reporting gaps suggest that reviewers and venues should adopt reporting checklists requiring participant counts, age, gender, and ethnicity before acceptance.
  • The concentration of user studies in human-computer interaction venues implies that technical security venues may rarely see user evidence, which could explain why tools are evaluated without users.
  • A testable extension would be to run the same review protocol on a different publisher's digital library to see whether the underreporting is a venue-specific norm or a field-wide one.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. This paper reports a systematic literature review of phishing research in the ACM Digital Library. The authors searched ACM DL for the term "phishing," obtained 367 records, screened these to 253 relevant papers after excluding extended abstracts, non-English works, and papers not primarily about phishing, and then identified 51 papers containing user studies. Using thematic coding by three researchers, they analyze the technical and individual attributes studied, the reported demographics of participants, and the methodologies used. The central claim is that only 13.9% of phishing papers in the corpus are user-focused and that even these papers frequently omit participant age, gender, and race/ethnicity information. The paper concludes with recommendations for improved demographic reporting and more representative recruitment in future phishing research.

Significance. The review addresses an important gap: the extent to which phishing research validates tools and findings with human participants. The three-coder thematic coding process with an inter-coder reliability of 87.9% is a clear strength, and the detailed counts of demographic reporting gaps (e.g., only 3 of 51 papers report race/ethnicity) are useful for the community. However, the headline prevalence of 13.9% is not derivable from the described method: the 51 user-focused papers were selected from the 253 screened papers, so the relevant rate is about 20.2%. This is a central, fixable error. The absence of the included-study list and codebook further limits verification. If the prevalence is corrected and the corpus scope is carefully qualified, the paper would be a useful resource.

major comments (4)
  1. [Abstract; §3.1; §6] The headline prevalence statistic is inconsistent with the screening procedure. Section 3.1 states that the ACM DL search returned 367 papers, that exclusions left 253 relevant papers, and that the 51 user-focused papers were identified from those 253. The abstract's figure of 13.9% equals 51/367, which mixes a screened numerator with an unscreened denominator; the rate among relevant papers is 51/253 ≈ 20.2%. Section 6's phrase "13.9% of relevant published papers" is directly contradicted by this arithmetic. Please recompute the prevalence, or explicitly justify why the unscreened 367 is the intended population, and adjust all prevalence statements in the abstract, findings, and conclusion accordingly.
  2. [§4.4 and §4.4.1] The reporting-quality findings are internally inconsistent. Section 4.4 states that 37 of the 51 papers "did not mention the age range, gender distribution, or racial/ethnic backgrounds of the participants," but Section 4.4.1 then says that 14 of those 37 papers included some kind of age range. If 14 papers reported age range, then age range was missing in a different set of 37 papers. Please clarify which count refers to which demographic attribute and provide a table of reporting counts for age, gender, race/ethnicity, and participant number.
  3. [§3.1–§3.2] The review is not reproducible in its current form: there is no list of the 51 included papers, no codebook for the thematic codes, and no flow diagram showing the number of papers excluded at each stage. This matters because the central counts (51, 253, and the demographic subcounts) cannot be checked without this material. Please add an appendix or supplementary file containing the included-study list, the coding scheme, and the exclusion flow.
  4. [§3.1; §5; §6] The corpus is drawn exclusively from one publisher's database, but the title and several sentences in Sections 5 and 6 describe the article as reviewing "phishing research" or "user-centered phishing research" without the qualifier "in the ACM Digital Library." Because the prevalence figure can vary with the choice of venues (human-factors venues versus security-engineering venues), the paper should either restrict all unqualified prevalence and trend claims to ACM DL-indexed papers or include a comparison with other venues and an explicit limitations paragraph discussing corpus-selection bias.
minor comments (5)
  1. [§3.2] The inter-coder reliability metric is not defined; please state whether the reported 23.7%, 56.7%, and 87.9% are percentage agreement, Cohen's kappa, or another coefficient, and describe when in the reconciliation process the final ICR was measured.
  2. [§3.1] The search protocol is underspecified: only the single term "phishing" is mentioned, and no date range or database-specific filters are reported. Please state whether variant terms (e.g., "spear phishing," "phish") were considered and how the 2004–2018 window was defined.
  3. [§4.4] For clarity, replace the fractions "about ⅓" and "a little over ⅔" with their exact values (15/51 ≈ 29% and 37/51 ≈ 73%).
  4. [Figure 2] Figure 2 would be easier to read with labeled axes and the exact number of publications per year displayed above each bar.
  5. [References] Several references are incomplete or lack publication venues, for example Das et al. (2017) has no venue and Colarik and Janczewski (2007) is a URL-only citation; the reference list should be completed according to the journal's style.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: the 13.9% prevalence is a self-contained corpus count; the abstract's denominator inconsistency is a statistical correctness issue, not a circular reduction.

full rationale

The paper's central quantitative claim is an empirical prevalence count from a systematic search of the ACM Digital Library. The count is produced by the paper's own screening and coding process (Section 3.1: 367 raw search results, 253 after exclusion, 51 user-focused studies from those 253), and it is not derived from a fitted parameter, a prediction, or the authors' prior results. There are no equations or models whose output is equivalent to an input. The three self-citations (Das et al. 2017, 2018, 2019) appear only as background support in the introduction, discussion, and race/ethnicity section; none carries the central 13.9% claim, so per the rules they do not constitute circularity. The thematic coding categories are described as emergent from a random sample of the corpus and were applied with inter-coder reliability, so the qualitative findings are not a renaming of a pre-existing result. I flag one non-circular correctness issue: the abstract and conclusion report 51/367 = 13.9% while also calling the 253-paper post-exclusion set 'relevant'; using the paper's own screened denominator would give 51/253 ≈ 20.2%. This is a denominator consistency problem, not a circular reduction, because the numerator and denominator are both observable counts rather than one being constructed from the other. The limitation that only ACM Digital Library was searched affects external validity but does not make any claim equivalent to its own input. No significant circularity found.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

The review's central figures depend on the corpus choice, the search and screening rules, and the coding reliability. There are no free parameters and no invented entities; the main burden is the assumption that the ACM DL subset and the authors' coding are representative and reliable.

assumptions (3)
  • domain assumption The ACM Digital Library alone is a representative corpus for characterizing phishing user research.
    Section 3.1 gathers data only from ACM DL, while the conclusion generalizes to phishing research; no comparison with IEEE, USENIX, or other security venues is made.
  • domain assumption A keyword search for 'phishing' combined with the stated exclusion criteria captures all relevant user-focused phishing papers in the corpus.
    Section 3.1 relies on this search to define both the denominator and the candidate set for user-study screening, so any misspecification changes the prevalence count.
  • domain assumption The three coders' thematic categories and the reported inter-coder reliability yield measurements that are stable enough for the reported percentages.
    Section 3.2 reports ICR improved from 23.7% to 87.9% after discussion; this is an internal reliability measure rather than an external validation of the coding scheme.

how reviews work

0 comments
Cite this review

Pith. "Pith review of All About Phishing: Exploring User Research through a Systematic Literature Review." pith.science (2026). https://pith.science/paper/B2I6LJLY

@misc{pith2026190805897,
  author       = {Pith},
  title        = {Pith review of: All About Phishing: Exploring User Research through a Systematic Literature Review},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/B2I6LJLY}},
  note         = {Machine review of arXiv:1908.05897}
}
read the original abstract

Phishing is a well-known cybersecurity attack that has rapidly increased in recent years. It poses legitimate risks to businesses, government agencies, and all users due to sensitive data breaches, subsequent financial and productivity losses, and social and personal inconvenience. Often, these attacks use social engineering techniques to deceive end-users, indicating the importance of user-focused studies to help prevent future attacks. We provide a detailed overview of phishing research that has focused on users by conducting a systematic literature review of peer-reviewed academic papers published in ACM Digital Library. Although published work on phishing appears in this data set as early as 2004, we found that of the total number of papers on phishing (N = 367) only 13.9% (n = 51) focus on users by employing user study methodologies such as interviews, surveys, and in-lab studies. Even within this small subset of papers, we note a striking lack of attention to reporting important information about methods and participants (e.g., the number and nature of participants), along with crucial recruitment biases in some of the research.

Figures

Figures reproduced from arXiv: 1908.05897 by the authors.

Figure 1
Figure 1. Word cloud depicting relative representation of conference publication venues in our data set of 51 papers. 3. Study Methodology Our systematic literature review focused on published research on phishing. We collected our data by starting with all the research publications on phishing that are included in the ACM Digital Library. We performed the data extraction using ACM’s export feature and then implemented a qual… view at source ↗
Figure 2
Figure 2. Number of publications in our data set (n = 51) by year of publication Throughout the course of our systematic literature review, we found specific trends in phishing research. Although the term “phishing" was coined in 1996 (Kay, 2004), academic researchers did not begin publishing about phishing until 2004 (Dunham, 2004). The first user-centered study we see in our data set was from 2005 (Garfinkel and Miller, 200… view at source ↗

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. "It's Like Not Being Able to Read and Write": Narrowing the Digital Divide for Older Adults and Leveraging the Role of Digital Educators in Ireland

    cs.CR 2025-02 conditional novelty 4.0 of 10

    Interviews with 34 Irish digital educators show learner-led, step-by-step teaching works best, while fear, outdated devices, rural connectivity, funding, and transport gaps block older adults' digital inclusion.

Reference graph

Works this paper leans on

7 extracted references · 7 canonical work pages · cited by 1 Pith paper

  1. [1]

    Phishing

    Introduction Phishing is one of the most effective and well-known cyber threats, leading to millions of compromised credentials and contributing to 90% of data breaches (Retruster.com, 2019). Phishing scams are becoming increasingly more deceptive with sophisticated attacks that are able to manipulate end-users through, for instance, spoofed websites, tar...

  2. [2]

    socially engineer

    Background Literature During a phishing attack, attackers use digital deception to get their victims to reveal confidential information about themselves. The success of the deception depends on how well the attackers mimic legitimate services and contacts and how well the user can distinguish between and/or act appropriately toward what is fake and what i...

  3. [3]

    phishing

    Study Methodology Our systematic literature review focused on published research on phishing. We collected our data by starting with all the research publications on phishing that are included in the ACM Digital Library. We performed the data extraction using ACM’s export feature and then implemented a qualitative assessment protocol that utilized exclusi...

  4. [4]

    phishing

    Findings Figure 2: Number of publications in our data set (n = 51) by year of publication Throughout the course of our systematic literature review, we found specific trends in phishing research. Although the term “phishing" was coined in 1996 (Kay, 2004), academic researchers did not begin publishing about phishing until 2004 (Dunham, 2004). The first us...

  5. [5]

    Discussion and Implications Throughout our systematic literature review of user studies in published ACM papers on phishing, we found that there are a number of identifiable trends. The breadth of the research (40 out of 51 papers) concentrates primarily on the technical attributes of phishing attacks, such as the content and appearance of spoofed website...

  6. [6]

    In 2018, the Federal Bureau of Investigation estimated that companies around the world lost $12 billion because of business email compromises (Digitalinformationworld.com, 2019)

    Conclusion Phishing attacks are one of the oldest known cyber -attacks, resulting in the loss of billions of dollars every year (Moore and Clayton, 2011; Tian et al., 2018). In 2018, the Federal Bureau of Investigation estimated that companies around the world lost $12 billion because of business email compromises (Digitalinformationworld.com, 2019). Prev...

  7. [7]

    Cyber security deception

    References Almeshekah, M.H. and Spafford, E.H., (2016), “Cyber security deception”, in Vipin, Swarup and Wang, C. (Eds) Cyber deception , Springer Publishing, Switzerland, pp 23-50, ISBN: 978-3-319-32697-9. Alsharnouby, M., Alaca, F. and Chiasson, S. (2015), “Why phishing still works: User strategies for combating phishing attacks”, International Journal ...

Pith tools

Reviewed August 14, 2026 · model on record in the stance chip above.