REVIEW 5 major objections 6 minor 1 cited by
Evaluating User Perception of Multi-Factor Authentication: A Systematic Review
T0 review · 5 major / 6 minor · reviewed 2026-08-14 · deepseek-v4-flash
Pith's one-line read A systematic review of 623 multi-factor authentication papers from 2018 finds that only 9.1% include any user evaluation, and those 57 studies portray low adoption as inevitable while documenting avoidance under mandatory use.
desk verdict Useful meta-analytic data on authentication user studies, but the MFA-specific 9.1% figure is undercut by a corpus that mixes in password/SFA papers. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The argument is carried by a two-stage review instrument. First, a systematic literature review with a six-category coding scheme (cyber threat testing, traditional authentication schemes, industry manufacturers, new authentication technologies, user-based studies, organizational implementation) classifies 623 papers harvested from four academic databases in 2018. Second, a meta-analytic codebook applied to the 57 user-focused papers extracts risk-perception themes, recruitment demographics, and methods (experiments versus surveys), and it is this second instrument that produces the findings about adoption inevitability, avoidance, and demographic bias.
What would settle it
A replication that reports exact database queries, per-database hit counts, and screening exclusions; if the resulting denominator of MFA-focused 2018 papers is not 623, or the share with user evaluation is not 9.1 percent, the review's headline numbers collapse.
Extended reading notes
Core claim
The central discovery is a quantified mismatch between technical output and user evidence in 2018 MFA research. Using a coding taxonomy applied to 623 papers, the paper categorizes 48.2 percent ($m = 300$) as new authentication technologies but only 9.1 percent ($n = 57$) as user-based studies; within that user-focused set, it finds the prevailing research narrative that lower adoption is inevitable, avoidance is pervasive under mandatory use, and risk-perception work concentrates on password memorability and usability rather than the trade-offs users actually face. It also documents reporting deficits: 91.5 percent of the user studies do not report testing technical expertise, 71.2 percent do not report compensation, 31 of 57 do not report participants' educational background, and the modal participant pool is college students.
Load-bearing premise
The derived percentages assume that the four-database, keyword-based search captured the full population of 2018 MFA research; if the sample is incomplete or contaminated by single-factor authentication papers, every percentage in the review shifts.
Editorial extensions
If this is right
- If the 9.1 percent ratio is accurate, the field's default next step should be user evaluation of existing MFA schemes, not new proposals.
- The finding that avoidance is pervasive under mandatory use implies that adoption studies should study voluntary and mandated contexts separately and design for the mandated case.
- Demographic skew toward university students means published usability results may not transfer to older, less tech-literate, or disabled populations; broadening recruitment is a direct implication.
- Reporting gaps in age, gender, education, compensation, and expertise suggest future user studies should adopt standard reporting items, and reviewers should expect them.
- The 57-study meta-analysis provides a baseline that future systematic reviews can use to measure whether the field's user-evidence share is growing.
Reading between the lines
- A field that mostly proposes new authentication schemes while rarely testing them with users cannot produce reliable adoption forecasts; the paper's numbers imply that deployment decisions in organizations are being made without strong peer-reviewed user evidence.
- The paper's finding that researchers frame low adoption as inevitable may itself contribute to low adoption: designers who expect rejection may optimize for security performance rather than for the friction points that drive avoidance.
- A testable extension is to re-run the same review on a later year with the same inclusion rules; if the user-study share has not moved away from 9.1 percent, the research culture has not absorbed the message.
- Because the paper does not report exact search strings or per-step screening counts, an independent replication with full transparency would be the cleanest way to confirm the denominator and the derived percentages.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper reports a systematic literature review of 623 papers published in 2018 that the authors describe as primarily focused on multi-factor authentication (MFA). It claims that only 9.1% of these papers (n=57) performed any user evaluation, and it meta-analyzes those 57 user-focused studies to identify risk perception categories, participant recruitment biases, and methodological/reporting problems. The paper concludes that MFA research is heavily skewed toward proposing new technologies rather than understanding users, and that existing user studies suffer from demographic and reporting biases.
Significance. If the central claims were fully supported, the paper would provide a valuable quantified snapshot of a gap in usable security research: a large majority of MFA-related work not grounded in user evidence, and the user evidence that exists concentrated in university populations. The paper's strengths include a clearly motivated research question, an explicit attempt to follow systematic-review methodology adapted from Stowell et al., and detailed coding of demographic and methodological features across the 57 user studies. However, the significance is currently limited by methodological transparency problems: the search and screening process is not reported at a level that permits replication, the corpus appears to mix single-factor password research with MFA research, and several reported statistics are internally inconsistent. The dataset and coding scheme could be a useful community resource if these issues are resolved.
major comments (5)
- [Section 3, Section 4.1, Table 3] The abstract's claim that N=623 papers 'primarily focused on MFA technologies' is not supported by the inclusion criteria and the paper's own tabulations. Inclusion criterion (5) explicitly admits papers focused on 'password, 2FA, and MFA,' and Section 4.1 reports 143 papers under 'Traditional Authentication Schemes.' More importantly, Table 3 places 51 of the 57 user-focused studies under single-factor/password categories (conventional passwords, password creation, password management, password meter, password cracking, password guessability, student-created passwords). This means the 9.1% figure is a denominator over a corpus that is not MFA-specific, and the headline 'only 9.1% of MFA papers include user evaluation' may substantially misstate the state of MFA research. The authors should either restrict the corpus to genuinely MFA-focused work or clearly reframe the claim as applying to authentication research broadly.
- [Section 3 (Methods)] The search and screening process is not described at a level that permits replication or assessment of selection bias. The text states only that four databases were searched with keywords 'multi-factor authentication,' 'two factor authentication,' and 'password' via Publish or Perish, but it does not provide exact query strings, per-database hit counts, numbers of papers excluded at title, abstract, and full-text screening stages, or reasons for exclusion. Without a PRISMA-style flow diagram and the underlying search strings, the reader cannot judge whether the 623-paper corpus is a complete or representative sample of 2018 authentication research, and the entire 9.1% denominator is therefore not auditable. This is a load-bearing omission because the paper's central quantitative claim depends on the denominator.
- [Table 4] Table 4 contains arithmetic inconsistencies that undermine confidence in the meta-analysis statistics. The compensation row reports 17 paid studies and 42 not reported, which sums to 59, not the stated n=57; the percentages 28.9% and 71.2% correspond to 17/59 and 42/59, not to proportions of 57. The gender subtable also has issues: 'Gender Based Studies 3 (5.1%)' and 'Mentions Gender For Study 4 (6.8%)' sum with 'Non-Gender Studies 52 (88.1%)' to 59, again exceeding n=57. The education subcategories (5+8+9+2+2=26) do not cover all 57 papers and lack a clear denominator. These tables need corrected denominators and explicit reporting of which categories are exclusive versus overlapping.
- [Abstract and Section 4.2] The abstract states that the meta-analysis showed 'avoidance was pervasive among mandatory use,' but this finding does not appear as a synthesized result anywhere in Section 4.2. Section 4.2 reports risk-perception categories (Table 2), password-related study types (Table 3), participant demographics (Table 4), and methods (Section 4.2.4), but it never presents a thematic synthesis on adoption inevitability or avoidance under mandatory use. Either the analysis supporting this abstract claim should be added to the body, with quotes or effect sizes from the 57 studies, or the abstract should be revised to match the reported findings.
- [Section 4.2.4] The methods subsection reports that '21 out of 57' studies performed usability testing of existing or proposed MFA, whereas earlier in the same subsection it says 25 studies were on newly proposed schemes, with 16 using usability feature testing and 9 using in-lab experiments, which sums to 25. The relationship among these numbers is unclear: are the 21 usability-testing studies a subset of the 25, or a separate category? The text should define mutually exclusive or explicitly overlapping categories and consistently report counts and percentages so that the meta-analysis is internally coherent.
minor comments (6)
- [Throughout] The manuscript contains numerous grammatical and typographical errors that should be corrected in revision, including 'For our research, began by performing' in Section 1, 'a two of them discuss' in Section 5, and the fragment 'Only eighteen Gender' preceding Table 4.
- [Section 3 (Methods)] The sentence 'Papers were included if they met the following criteria' lists six criteria, but criteria (1) and (2) are separated by a period and the list formatting is inconsistent; this should be cleaned up for readability and to avoid ambiguity about which criteria apply at which stage.
- [References] Several references are incomplete or inconsistently formatted, such as the Statistica citation in the abstract/introduction, the 'Das et al. 2019.' entry with a stray period, and the mixture of citation styles (e.g., some entries have publisher locations, others do not).
- [Figure 1] Figure 1 is referenced in Section 3 but is not described in sufficient detail in the text; the figure should be self-contained or accompanied by a narrative walkthrough of the screening funnel, including the number of papers at each stage.
- [Section 4.1] The sentence 'Majority of our collected sample set (N = 48.2%)' uses 'N' where a percentage is intended; it should read 'the majority of our collected sample set (48.2%)' to avoid confusion between sample size and percentage.
- [Section 4.2.2 and Table 3] The text says '16% of the user studies focused on understanding the password security understanding of the users,' but Table 2 lists 'Understanding Password Security' as 25 (44.0%) and Table 3 uses different subcategories; the authors should clarify which table supports which percentage and whether percentages are out of all 57 studies or out of the password-focused subset.
Circularity Check
No significant circularity: the review's percentages come from coding an external corpus; self-citations are background, not load-bearing.
full rationale
The paper's derivation chain is a systematic literature review: it constructs a corpus from databases, screens it, codes papers, and reports percentages. None of the central quantities (N=623, 9.1% user evaluation, n=57) is defined in terms of the conclusion; they are counts obtained by applying inclusion/exclusion criteria to external papers. The inclusion criterion (5) admits password-focused papers ('Papers that primarily focused on authentication technologies. Such as password, 2FA, and MFA tool and technologies'), and Table 3 shows many of the 57 user studies concern traditional passwords; this is a potential validity or corpus-construction problem, not circularity, because the 9.1% figure is still an empirical count rather than an equation equivalent to its input. Self-citations (Das et al. 2018a/b, 2019a/b) are used as supporting examples or background claims about usability challenges; they are not invoked to justify the corpus percentages or the coding outcomes. The abstract's claim that avoidance is 'pervasive among mandatory use' is not visibly synthesized in Section 4.2, but an unsupported or missing result is a reporting/correctness concern, not a circular derivation. No step reduces to its own input by definition or by self-citation, so the appropriate score is 0.
Assumptions & free parameters
assumptions (3)
- domain assumption Search of Google Scholar, ACM, Science Direct, and IEEE via Publish or Perish with keywords 'multi-factor authentication', 'two factor authentication', and 'password' identifies the population of 2018 MFA papers.
- domain assumption The authors' categorization of papers into user-focused versus non-user-focused studies is reliable.
- domain assumption Papers published in 2018 are representative of current MFA user perception trends.
Cite this review
Pith. "Pith review of Evaluating User Perception of Multi-Factor Authentication: A Systematic Review." pith.science (2026). https://pith.science/paper/LWATNYVA
@misc{pith2026190805901,
author = {Pith},
title = {Pith review of: Evaluating User Perception of Multi-Factor Authentication: A Systematic Review},
year = {2026},
howpublished = {\url{https://pith.science/paper/LWATNYVA}},
note = {Machine review of arXiv:1908.05901}
}
read the original abstract
Security vulnerabilities of traditional single factor authentication has become a major concern for security practitioners and researchers. To mitigate single point failures, new and technologically advanced Multi-Factor Authentication (MFA) tools have been developed as security solutions. However, the usability and adoption of such tools have raised concerns. An obvious solution can be viewed as conducting user studies to create more user-friendly MFA tools. To learn more, we performed a systematic literature review of recently published academic papers (N = 623) that primarily focused on MFA technologies. While majority of these papers (m = 300) proposed new MFA tools, only 9.1% of papers performed any user evaluation research. Our meta-analysis of user focused studies (n = 57) showed that researchers found lower adoption rate to be inevitable for MFAs, while avoidance was pervasive among mandatory use. Furthermore, we noted several reporting and methodological discrepancies in the user focused studies. We identified trends in participant recruitment that is indicative of demographic biases.
Figures
Forward citations
Cited by 1 Pith paper
-
The Impact of Security and Privacy Controls on Users' Emotional Engagement with Generative AI Chatbots
In a vignette study of 354 U.S. participants, deletion-based privacy controls outperformed all other controls in increasing willingness to engage with GenAI chatbots for emotional support, while technically complex co...
Reference graph
Works this paper leans on
-
[1]
reported using internet daily (Statistic 2018)
Introduction Online user presence increased considerably in the last decade (Kemp 2017), where in 2018, 89% adults in the U.S. reported using internet daily (Statistic 2018). Such exponential growth in users and data (Patil & Seshadri 2014) has warranted security practitioners to become more concerned with online data security (Al Hasib 2009) and access c...
work page 2017
-
[2]
Designing and Evaluating mHealth Interventions for Vulnerable Populations: A Systematic Review
Related Work MFA involves multi-layer authentication scheme to mitigate risks of single factor sign-ons, such as, password breaches and unauthorized access of trusted devices (Hwang et al. 2002). Previous research on MFA primarily focused on the technological improvement of authentication and access control to address existing weakness in various areas su...
work page 2002
-
[3]
Methods We adapted the study methodology for the literature review of Multi-Factor Authentication from Stowell et al.’s work (Stowell et al. 2018). Additionally, we modified the protocol to better fit our research needs. Methods utilized in our research involve the following steps: (1) Data Collection through database search, (2) Data Screening involving:...
work page 2018
-
[4]
We started our data collection by generating a large sample of papers related to a set of keywords from four major databases: ACM, IEEE Xplore, Google Scholar and Science Direct. We also performed a Quality Assessment of the papers to ensure that they met our inclusion or exclusion criteria. Papers were included if they met the following criteria: (1) The...
work page 2018
-
[5]
Below are major findings we’ve discovered during the research
Findings During our systematic literature review, we investigated the existing set of literature based around user studies in multi-factor authentication for paving the path for future studies by underlining existing gaps in research. Below are major findings we’ve discovered during the research. 4.1. Overall Analysis We conducted a thorough coding analys...
work page 2018
-
[6]
Conclusion Multi-factor authentication improves online data security by implementing multiple factors in addition to single factor sign -on. Usability of such security technologies often comes across as a challenge for security practitioners, researchers, designers, and developers. Through systematic literature review ( N = 623) we aimed at understanding ...
work page 2018
-
[7]
References Abo-Zahhad, M., Ahmed, S.M. and Abbas, S.N., 2016. A new multi -level approach to EEG based human authentication using eye blinking. Pattern Recognition Letters , 82, pp.216 -225.Al Hasib, A. (2009), ‘Threats of online social networks’, IJCSNS International Journal of Computer Science and Network Security 9(11), 288–93. Almoctar, H., Irani, P.,...
work page 2009
-
[18]
(pp. 239-253). Braz, C. and Robert, J.M., 2006, April. Security and usability: the case of the user authentication methods. In IHM (Vol. 6, pp. 199-203). Brereton, P., Kitc henham, B.A., Budgen, D., Turner, M. and Khalil, M., 2007. Lessons from applying the systematic literature review process within the software engineering domain. Journal of systems and...
work page 2006
Show all 9 references
-
[2012]
Reusable authentication experience tool . U.S. Patent 8,136,148. Chithra, P.L. and Sathva, K., 2018, February. Pristine PixC aptcha as Graphical Password for Secure eBanking Using Gaussian Elimination and Cleaves Algorithm. In 2018 International Conference on Computer, Communi...
2019 arXiv
Reviewed August 14, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.