{"id":"2352b0bf-4d98-4474-a0a8-29de7fd9dcc3","arxiv_id":"1908.05901","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Only 9.1% of 2018 multi-factor authentication papers evaluated users, and the 57 user studies that exist show recruitment bias, missing demographics, and inconsistent reporting.","lead":"An analysis of 623 papers on multi-factor authentication published in 2018 finds that only 9.1 percent include any user evaluation. A meta-analysis of the 57 user-focused studies identifies demographic bias and reporting gaps that limit what can be concluded about adoption.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 9.1% figure and the 57-study meta-analysis mix MFA with single-factor password studies; the central claim about MFA research may not be about MFA at all.","rationale":"The reader identified irreproducible search as the fragile premise. My independent read found a more directly damaging, internally evidenced flaw: the paper's own inclusion criteria and Table 3 show that the 623-paper corpus and the 57 'user-focused' studies include a substantial share of single-factor/password-only work. If those studies are not MFA evaluations, the 9.1% denominator and the meta-analytic findings cannot be attributed to MFA research. This is not simply an outside-consensus disagreement; it is an internal mismatch between the paper's stated object (MFA) and its operationalization (authentication broadly). The abstract's claim about avoidance under mandatory use also lacks a supporting synthesis in Section 4.2, so the headline is stronger than the results. Because the paper could be repaired by re-running the screen and re-coding the sample, I do not call for rejection, but the conditional verdict should require that correction.","tokens_in":9081,"tokens_out":4832,"duration_ms":43102,"concrete_test":"Request or reconstruct the screening corpus and the list of 57 user-focused studies. Independently classify each of the 57 papers as 'evaluates a multi-factor authentication mechanism' versus 'single-factor/password-only' using a strict MFA definition, and re-run the Section 3 screen with inclusion criterion (5) restricted to papers whose primary object is an MFA scheme. If the number of MFA user studies drops materially below 57, or if the ratio of user-evaluation papers in the MFA-only denominator differs from 9.1%, the paper's central claim fails. Also verify whether any of the 57 papers state that avoidance was pervasive under mandatory MFA use; if none do, the abstract overclaims.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim is that only 9.1% of MFA-focused 2018 papers performed user evaluation and that meta-analysis of 57 user studies shows MFA adoption problems. Section 3's inclusion criterion (5) explicitly admits papers 'primarily focused on authentication technologies. Such as password, 2FA, and MFA,' so the 623-paper corpus is not necessarily MFA-focused: Section 4.1 itself reports 143 papers on traditional (SFA) authentication schemes. More tellingly, Table 3 ('Traditional Single-factor Authentication') lists 51 of the 57 user-focused studies under SFA/password categories (conventional passwords 8, password creation 12, password management 16, password meter 8, password cracking 2, password guessability 2, student-created passwords 3). Even if these categories overlap with MFA work, they show that a large majority of the 'user focused studies' are not necessarily MFA evaluations. The abstract's additional assertion that these studies found avoidance 'pervasive among mandatory use' does not appear as a synthesized finding in Section 4.2; the section reports risk-perception categories, demographics, and methods instead. Consequently, the headline quantitative gap and the adoption conclusions may be artifacts of corpus construction and coding rather than properties of MFA research.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports a systematic literature review of 623 papers published in 2018 that the authors describe as primarily focused on multi-factor authentication (MFA). It claims that only 9.1% of these papers (n=57) performed any user evaluation, and it meta-analyzes those 57 user-focused studies to identify risk perception categories, participant recruitment biases, and methodological/reporting problems. The paper concludes that MFA research is heavily skewed toward proposing new technologies rather than understanding users, and that existing user studies suffer from demographic and reporting biases.","tokens_in":9346,"tokens_out":3474,"duration_ms":36365,"significance":"If the central claims were fully supported, the paper would provide a valuable quantified snapshot of a gap in usable security research: a large majority of MFA-related work not grounded in user evidence, and the user evidence that exists concentrated in university populations. The paper's strengths include a clearly motivated research question, an explicit attempt to follow systematic-review methodology adapted from Stowell et al., and detailed coding of demographic and methodological features across the 57 user studies. However, the significance is currently limited by methodological transparency problems: the search and screening process is not reported at a level that permits replication, the corpus appears to mix single-factor password research with MFA research, and several reported statistics are internally inconsistent. The dataset and coding scheme could be a useful community resource if these issues are resolved.","major_comments":[{"comment":"The abstract's claim that N=623 papers 'primarily focused on MFA technologies' is not supported by the inclusion criteria and the paper's own tabulations. Inclusion criterion (5) explicitly admits papers focused on 'password, 2FA, and MFA,' and Section 4.1 reports 143 papers under 'Traditional Authentication Schemes.' More importantly, Table 3 places 51 of the 57 user-focused studies under single-factor/password categories (conventional passwords, password creation, password management, password meter, password cracking, password guessability, student-created passwords). This means the 9.1% figure is a denominator over a corpus that is not MFA-specific, and the headline 'only 9.1% of MFA papers include user evaluation' may substantially misstate the state of MFA research. The authors should either restrict the corpus to genuinely MFA-focused work or clearly reframe the claim as applying to authentication research broadly.","section":"Section 3, Section 4.1, Table 3"},{"comment":"The search and screening process is not described at a level that permits replication or assessment of selection bias. The text states only that four databases were searched with keywords 'multi-factor authentication,' 'two factor authentication,' and 'password' via Publish or Perish, but it does not provide exact query strings, per-database hit counts, numbers of papers excluded at title, abstract, and full-text screening stages, or reasons for exclusion. Without a PRISMA-style flow diagram and the underlying search strings, the reader cannot judge whether the 623-paper corpus is a complete or representative sample of 2018 authentication research, and the entire 9.1% denominator is therefore not auditable. This is a load-bearing omission because the paper's central quantitative claim depends on the denominator.","section":"Section 3 (Methods)"},{"comment":"Table 4 contains arithmetic inconsistencies that undermine confidence in the meta-analysis statistics. The compensation row reports 17 paid studies and 42 not reported, which sums to 59, not the stated n=57; the percentages 28.9% and 71.2% correspond to 17/59 and 42/59, not to proportions of 57. The gender subtable also has issues: 'Gender Based Studies 3 (5.1%)' and 'Mentions Gender For Study 4 (6.8%)' sum with 'Non-Gender Studies 52 (88.1%)' to 59, again exceeding n=57. The education subcategories (5+8+9+2+2=26) do not cover all 57 papers and lack a clear denominator. These tables need corrected denominators and explicit reporting of which categories are exclusive versus overlapping.","section":"Table 4"},{"comment":"The abstract states that the meta-analysis showed 'avoidance was pervasive among mandatory use,' but this finding does not appear as a synthesized result anywhere in Section 4.2. Section 4.2 reports risk-perception categories (Table 2), password-related study types (Table 3), participant demographics (Table 4), and methods (Section 4.2.4), but it never presents a thematic synthesis on adoption inevitability or avoidance under mandatory use. Either the analysis supporting this abstract claim should be added to the body, with quotes or effect sizes from the 57 studies, or the abstract should be revised to match the reported findings.","section":"Abstract and Section 4.2"},{"comment":"The methods subsection reports that '21 out of 57' studies performed usability testing of existing or proposed MFA, whereas earlier in the same subsection it says 25 studies were on newly proposed schemes, with 16 using usability feature testing and 9 using in-lab experiments, which sums to 25. The relationship among these numbers is unclear: are the 21 usability-testing studies a subset of the 25, or a separate category? The text should define mutually exclusive or explicitly overlapping categories and consistently report counts and percentages so that the meta-analysis is internally coherent.","section":"Section 4.2.4"}],"minor_comments":[{"comment":"The manuscript contains numerous grammatical and typographical errors that should be corrected in revision, including 'For our research, began by performing' in Section 1, 'a two of them discuss' in Section 5, and the fragment 'Only eighteen Gender' preceding Table 4.","section":"Throughout"},{"comment":"The sentence 'Papers were included if they met the following criteria' lists six criteria, but criteria (1) and (2) are separated by a period and the list formatting is inconsistent; this should be cleaned up for readability and to avoid ambiguity about which criteria apply at which stage.","section":"Section 3 (Methods)"},{"comment":"Several references are incomplete or inconsistently formatted, such as the Statistica citation in the abstract/introduction, the 'Das et al. 2019.' entry with a stray period, and the mixture of citation styles (e.g., some entries have publisher locations, others do not).","section":"References"},{"comment":"Figure 1 is referenced in Section 3 but is not described in sufficient detail in the text; the figure should be self-contained or accompanied by a narrative walkthrough of the screening funnel, including the number of papers at each stage.","section":"Figure 1"},{"comment":"The sentence 'Majority of our collected sample set (N = 48.2%)' uses 'N' where a percentage is intended; it should read 'the majority of our collected sample set (48.2%)' to avoid confusion between sample size and percentage.","section":"Section 4.1"},{"comment":"The text says '16% of the user studies focused on understanding the password security understanding of the users,' but Table 2 lists 'Understanding Password Security' as 25 (44.0%) and Table 3 uses different subcategories; the authors should clarify which table supports which percentage and whether percentages are out of all 57 studies or out of the password-focused subset.","section":"Section 4.2.2 and Table 3"}],"recommendation":"major_revision","confidential_remarks":"The paper addresses an important topic and the authors' coding effort is evident, but the central claims currently rest on a corpus whose composition and screening are not fully described. The mixing of password/SFA papers with MFA papers is the most significant concern; if the authors can supply a complete search protocol and either re-analyze the data with a clearly defined MFA scope or reframe the claims to authentication research broadly, the contribution could be salvageable. I also note that the abstract overstates findings not present in the body; this should be corrected regardless of the direction of revision."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThis is a systematic review with a useful, new quantitative claim: among 623 authentication papers from 2018, only 57 (9.1%) did any user evaluation, and those 57 are heavily skewed toward university samples, underreport demographics, and rarely compensate participants. That last set of meta-analytic observations about reporting and recruitment bias is the genuinely useful contribution, and it is grounded in reproducible coding of the 57 papers.\n\nWhat the paper does well: the coding effort is real, the tables on risk perception and demographic reporting give a concrete picture of the literature's blind spots, and the authors are transparent about their own inclusion/exclusion criteria. The prior-work connection to Stowell et al. is appropriate.\n\nThe soft spots are real and two of them matter. First, the corpus is not specifically MFA. Inclusion criterion (5) explicitly admits papers \"primarily focused on authentication technologies. Such as password, 2FA, and MFA,\" and Table 3 lists 51 of the 57 user studies under single-factor password categories. So the headline \"only 9.1% of MFA papers did user evaluation\" is probably false; the actual number for MFA-only papers is smaller than 57, and the abstract's MFA framing is not supported by the data. This is the load-bearing flaw, and the stress-test note is correct. Second, the abstract's claim that avoidance was \"pervasive among mandatory use\" does not appear as a synthesized finding in Section 4.2. The section discusses risk perception categories and demographics, not adoption-avoidance findings. That's an abstract/body mismatch.\n\nMinor issues: arithmetic in Table 4 is off (17 paid + 42 not reported = 59, not 57, and percentages sum to >100); the search protocol lacks exact query strings, per-database counts, and exclusion counts at each screening step, so the 623 denominator is not independently reproducible; and there are typos (e.g., \"Only eighteen Gender of the 57 papers mentioned\" is mangled).\n\nNone of this kills the core observation that user evaluation is rare and demographically skewed in this literature. But the paper needs reframing as a review of authentication-user studies broadly, or it needs to re-do the corpus construction to isolate MFA. I would send it to peer review with a request for major revision: release the search protocol and article list, correct the inconsistency, and align the abstract with the body. The usable security community would benefit from a trustworthy version of this review.","headline":"Useful meta-analytic data on authentication user studies, but the MFA-specific 9.1% figure is undercut by a corpus that mixes in password/SFA papers.","tokens_in":9857,"tokens_out":2226,"would_cite":false,"duration_ms":20342,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A systematic review of 623 multi-factor authentication papers from 2018 finds that only 9.1% include any user evaluation, and those 57 studies portray low adoption as inevitable while documenting avoidance under mandatory use.","keywords":["multi-factor authentication","systematic literature review","user study","usable security","adoption","demographic bias","meta-analysis","two-factor authentication"],"falsifier":"A replication that reports exact database queries, per-database hit counts, and screening exclusions; if the resulting denominator of MFA-focused 2018 papers is not 623, or the share with user evaluation is not 9.1 percent, the review's headline numbers collapse.","tokens_in":8921,"feed_emoji":"🔐","tokens_out":6831,"duration_ms":60616,"temperature":0.7,"pith_summary":"Sifting one year of published research on multi-factor authentication (MFA), the paper finds that only 57 of 623 MFA-focused papers (9.1 percent) included any user evaluation, while 300 papers proposed new authentication technologies. Meta-analyzing those 57 user studies, it reports that researchers tended to treat low adoption as inevitable and found that users avoided MFA when it was mandatory. The user studies also show systematic demographic skews: recruitment leans on university students, gender analysis is rare, compensation and technical expertise are often unreported, and only three studies mention populations with disabilities. If these proportions hold, the security community is designing and deploying MFA on a thin and biased user-evidence base.","feed_headline":"Only 9.1% of MFA research tests real users","feed_subtitle":"A systematic review of 623 papers finds user studies rare, university-skewed, and pessimistic about adoption.","key_machinery":"The argument is carried by a two-stage review instrument. First, a systematic literature review with a six-category coding scheme (cyber threat testing, traditional authentication schemes, industry manufacturers, new authentication technologies, user-based studies, organizational implementation) classifies 623 papers harvested from four academic databases in 2018. Second, a meta-analytic codebook applied to the 57 user-focused papers extracts risk-perception themes, recruitment demographics, and methods (experiments versus surveys), and it is this second instrument that produces the findings about adoption inevitability, avoidance, and demographic bias.","core_discovery":"The central discovery is a quantified mismatch between technical output and user evidence in 2018 MFA research. Using a coding taxonomy applied to 623 papers, the paper categorizes 48.2 percent ($m = 300$) as new authentication technologies but only 9.1 percent ($n = 57$) as user-based studies; within that user-focused set, it finds the prevailing research narrative that lower adoption is inevitable, avoidance is pervasive under mandatory use, and risk-perception work concentrates on password memorability and usability rather than the trade-offs users actually face. It also documents reporting deficits: 91.5 percent of the user studies do not report testing technical expertise, 71.2 percent do not report compensation, 31 of 57 do not report participants' educational background, and the modal participant pool is college students.","pith_inferences":["A field that mostly proposes new authentication schemes while rarely testing them with users cannot produce reliable adoption forecasts; the paper's numbers imply that deployment decisions in organizations are being made without strong peer-reviewed user evidence.","The paper's finding that researchers frame low adoption as inevitable may itself contribute to low adoption: designers who expect rejection may optimize for security performance rather than for the friction points that drive avoidance.","A testable extension is to re-run the same review on a later year with the same inclusion rules; if the user-study share has not moved away from 9.1 percent, the research culture has not absorbed the message.","Because the paper does not report exact search strings or per-step screening counts, an independent replication with full transparency would be the cleanest way to confirm the denominator and the derived percentages."],"forward_implications":["If the 9.1 percent ratio is accurate, the field's default next step should be user evaluation of existing MFA schemes, not new proposals.","The finding that avoidance is pervasive under mandatory use implies that adoption studies should study voluntary and mandated contexts separately and design for the mandated case.","Demographic skew toward university students means published usability results may not transfer to older, less tech-literate, or disabled populations; broadening recruitment is a direct implication.","Reporting gaps in age, gender, education, compensation, and expertise suggest future user studies should adopt standard reporting items, and reviewers should expect them.","The 57-study meta-analysis provides a baseline that future systematic reviews can use to measure whether the field's user-evidence share is growing."],"supporting_citations":[{"why":"Provides the systematic literature review methodology that the paper adapts for MFA.","marker":"Stowell et al. 2018"},{"why":"Grounds the systematic review process and reporting lessons used to structure the search and screening.","marker":"Brereton et al. 2007"},{"why":"One of the two-phase YubiKey usability studies the paper uses to illustrate vendor-focused user research and the absence of disability-population studies.","marker":"Reynolds et al. 2018"},{"why":"A two-phase user study cited to argue that user research yields actionable improvements in MFA adoption.","marker":"Das et al. 2018b"},{"why":"Supplies the example of a university-student participant pool that substantiates the recruitment-bias finding.","marker":"Naiakshina et al. 2018"},{"why":"Provides evidence of university-based recruitment and compensation details in a password-strength field study.","marker":"Becker et al. 2018"},{"why":"Documents the low monetary crowdsourcing rewards reported in some user studies.","marker":"Kankane et al. 2018"},{"why":"Shows the college-educated participant demographics that support the educational-background bias claim.","marker":"Gratian et al. 2018"},{"why":"The only study the review identifies as potentially benefiting disabled users, anchoring the disability-population gap.","marker":"Almoctar et al. 2018"}],"fun_headline_variants":["Only 1 in 11 MFA papers evaluates real users","MFA research: user studies only 9.1% of papers","9.1% of MFA papers test users; rest assume","User studies absent from most MFA research","MFA adoption seen as inevitable in user studies"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The derived percentages assume that the four-database, keyword-based search captured the full population of 2018 MFA research; if the sample is incomplete or contaminated by single-factor authentication papers, every percentage in the review shifts.","fun_headline_variants_meta":{"raw":{"variants":["Only 1 in 11 MFA papers evaluates real users","MFA research: user studies only 9.1% of papers","9.1% of MFA papers test users; rest assume","User studies absent from most MFA research","MFA adoption seen as inevitable in user studies"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000708,"raw_usage":{"total_tokens":3157,"prompt_tokens":879,"completion_tokens":2278,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":495,"completion_tokens_details":{"reasoning_tokens":2196}},"tokens_in":495,"tokens_out":2278,"duration_ms":16205,"temperature":1.0,"reasoning_tokens":2196,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T13:00:52.861433+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A replication that reports exact database queries, per-database hit counts, and screening exclusions; if the resulting denominator of MFA-focused 2018 papers is not 623, or the share with user evaluation is not 9.1 percent, the review's headline numbers collapse.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the systematic literature review methodology that the paper adapts for MFA."}],"review_version":1}