Pith. sign in

REVIEW 4 major objections 5 minor 10 references

MFA is a Waste of Time! Understanding Negative Connotation Towards MFA Applications via User Generated Content

T0 review · 4 major / 5 minor · reviewed 2026-08-14 · deepseek-v4-flash

Pith's one-line read The paper argues that negative reviews of MFA apps cluster on concrete usability failures rather than on doubts about security.

desk verdict A useful descriptive study of MFA app reviews, undermined by an unsupported causal leap from complaints to non-adoption. read the letter →

arxiv 1908.05902 v1 pith:QLDJXLBK submitted 2019-08-16 cs.CR cs.HCcs.LG

classification cs.CRcs.HCcs.LG
keywords multi-factorauthenticationMFAusabilityuser-generatedcontentsentimentanalysisappstorereviewsadoptionbarriersuserexperience
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper claims that negative user reviews of multi-factor authentication (MFA) apps are not primarily a complaint about security inconvenience, but a record of specific usability failures: setup, backup and migration, device compatibility, integration quality, and a sense of being forced to use MFA without explanation. Analyzing 12,500 reviews of five app-based MFA tools across app stores and an organizational site, the authors find 56.9% of comments negative, while positive comments tend to be generic. The authors argue these usability defects, not lack of security awareness, drive non-adoption of MFA. If correct, the result implies that improving configuration, recovery, and device support, along with risk communication, would increase adoption more than simply telling users MFA is secure.

What carries the argument

The mechanism is the aggregate user-generated review itself, treated as a large-scale usability diary. The authors collected 12,500 comments about Duo, Google Authenticator, Microsoft Authenticator, Authy, and Okta from the Apple App Store, Google Play, Amazon, and an internal organizational site; filtered them for sentence-length content and spam; ran sentiment and keyword extraction with a commercial text-analytics service; grouped comments by application version to track changes across releases; and qualitatively coded random samples (M=300 per category). The load-bearing step is the clustering of negative keywords and themes, which turns unstructured complaints into the recurring categories of setup, backup/migration, compatibility, integration, and forced use that the recommendations target.

What would settle it

Survey a representative sample of people who were offered MFA but never enrolled, and ask whether their reasons are setup difficulty, device incompatibility, backup concerns, or a belief that MFA is unnecessary. If awareness-related reasons, such as not knowing MFA was important, dominate over the usability categories, the paper's central attribution to usability defects would be weakened.

Watch

Extended reading notes

Core claim

The central discovery is that users' negative attitudes toward MFA apps cluster on concrete engineering and policy failures rather than on skepticism about security value. Across the collected reviews, the dominant complaint categories are backup and migration (lost or non-transferable codes when changing devices), first-time setup difficulty, limited device compatibility, crashes and poor integration, and 'forced-to-use' deployments where employers or universities require MFA without explaining its benefit. The paper reports that negative comments (56.9%) outnumber positive ones and that positive reviews are mostly generic ('great app') while negative reviews name specific problems. The authors infer from this pattern that MFA non-adoption follows from avoidable usability and training gaps, and they recommend step-by-step setup instructions, risk communication that explains authentication factors, non-interrupting backup verification, NFC-based credential migration, and pre-deployment pilot testing.

Load-bearing premise

The paper relies on reviews written by people who installed or were required to use an MFA app, so its claim that negative experiences lead to non-adoption cannot be tested against people who never adopted MFA in the first place.

Editorial extensions

If this is right

  • If usability failures are the main barrier, MFA adoption in organizations should rise when setup is guided, backup and migration are seamless, and device compatibility is maintained.
  • Positive but generic reviews should not be read as endorsement; services that rely on required-use deployments may see high ratings while still harboring latent resistance.
  • Version-by-version review patterns can serve as a regression signal for MFA developers, with drops after updates indicating usability regressions.
  • Risk communication that explains what 'something you have' means could reduce the 'why do I need this?' resistance seen in forced-to-use reviews.
  • Improving recovery paths, such as backup codes and device migration, may matter as much as initial enrollment in retaining MFA users.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A direct testable extension would be to turn these complaint categories into a survey for non-adopters: people who never installed an MFA app would be asked whether setup difficulty, backup concerns, device incompatibility, or perceived benefit most explains their refusal, which would directly test the paper's causal inference.
  • The same review-mining approach could be applied to hardware tokens and biometric MFA to see whether the complaint categories shift, and to password managers to compare usability failure profiles.
  • The finding that forced deployment breeds hostility suggests that mandatory MFA programs paired with explanation and easy recovery may convert reluctant users, while mandatory deployment without communication may entrench resistance.
  • Because the data are cross-sectional, the version-trend analysis cannot separate app changes from changes in the reviewer population; a controlled deployment study with pre- and post-measures would separate those effects.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. This paper analyzes 12,500 user-generated reviews of five application-based MFA tools (Duo, Google Authenticator, Microsoft Authenticator, Authy, Okta) collected from Apple App Store, Google Play, Amazon Marketplace, and an unnamed organization's internal review site. The authors use the Microsoft Azure Text Analytics API for sentiment scoring and keyword extraction, plus a qualitative review of 300 randomly selected comments. They report that 56.9% of comments express negative sentiment and identify three main complaint themes: backup and migration difficulties, setup/compatibility/integration issues, and forced-to-use perceptions. Based on these findings, the paper recommends user training and risk communication, improved backup and migration mechanisms, and more extensive application and integration testing to improve MFA adoption. The stated central claim is that the majority of users face configuration, compatibility, backup, and risk-communication problems that lead to non-adoption of MFA.

Significance. If the findings hold, the paper provides a large-scale, naturally occurring complement to controlled usability studies of MFA, and its thematic results are broadly plausible and consistent with prior work. The paper's strength is its use of real user reviews rather than laboratory tasks, and its descriptive theme extraction is a reasonable starting point for understanding MFA usability complaints. However, the central causal claim about non-adoption is not supported by the data, and several methodological details are missing. The paper is likely to be useful to the usable-security community as descriptive evidence of user dissatisfaction, but its stronger contribution as evidence about adoption barriers is currently not justified.

major comments (4)
  1. [Abstract and Section 4] The abstract claims that the identified problems are 'leading to non-adoption of MFA.' This causal claim is not testable with the presented data. The reviews come from users who installed or were required to use MFA apps; the dataset contains no observations from individuals who never enrolled, abandoned setup, or uninstalled and refused to return. A user who complains about backup or setup may continue using the app, especially in organizational deployments. The data can support statements about dissatisfaction or negative connotation, but not a causal relationship to non-adoption. The authors should either temper the claim to 'associated with adoption barriers' or provide supplementary evidence (e.g., enrollment/uninstall logs or interviews with non-adopters).
  2. [Section 3.2 and Section 4] The sentiment analysis lacks validation and statistical support. The headline figure of 56.9% negative sentiment is presented without confidence intervals or a significance test. The per-application sentiment scores in Table 1 are given as point estimates with no uncertainty, so the claimed ordering (e.g., Authy highest, Okta lowest) cannot be distinguished from noise. The paper also does not report the accuracy or performance of the Azure Text Analytics API on this specific corpus, and the qualitative analysis of 300 comments has no inter-rater reliability measure or codebook. Without these, the descriptive results are not established beyond anecdote.
  3. [Section 3.1] The data-filtering rules may bias the sample. Specifically, comments shorter than 100 characters are discarded, yet many short negative expressions—such as 'waste of time' or 'bad app'—are exactly the kind of content the paper's title highlights. The manuscript does not report how many reviews were removed by this filter or by the anti-spam filter, nor does it give the distribution of the 12,500 comments across the five applications, the three marketplaces, and the internal organization site. This makes the representativeness of the sample impossible to assess and the analysis difficult to reproduce.
  4. [Section 4] The claim that 'positive comments are over-generic' and that many positive ratings are due to organizational mandates ('Great App' because it was instructed by their organization) is not supported by any quantitative or qualitative evidence presented in the paper. The authors do not define how this classification was made or provide examples and counts from the 300-comment subset. This claim is load-bearing for the conclusion that positive sentiment reflects compliance rather than satisfaction, so it needs explicit support or should be removed.
minor comments (5)
  1. [Section 3.1] The sentence 'we ran additional filters (N = 12500) within the collected dataset' is ambiguous: it is unclear whether 12,500 is the size before or after filtering. Please use distinct notation such as N_collected and N_final.
  2. [Figure 1 caption] The caption reads 'Word Cloud for showing the distribution of title (Fig.1) and contents (Fig.2)', but Figure 1 is a word cloud and Figure 2 appears to be a version-rating plot. This should be corrected to refer to the actual figures.
  3. [Section 1] 'We conclude by annotating crucial issues in current MFA implementation in section 4' conflicts with the paper's structure: Section 4 is Results and Analysis, not the conclusion. Please revise the wording.
  4. [Section 4] 'Duo Security and Okta had major review decent during development iterations' appears to be a typo; 'decent' should presumably be 'descent'.
  5. [References] There are small reference formatting errors: 'San Fransisco' should be 'San Francisco'; the entry for Das, Wang, Tingle & Camp (2019) has 'L.J' without a space; and the page range for Fu et al. is written as '1276– 1284' with an unwanted space.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the review analysis is self-contained and does not reduce to its inputs.

full rationale

This paper is an observational, empirical study rather than a derivation. It collects 12,500 user reviews of app-based MFA tools, applies sentiment analysis and keyword clustering, and qualitatively codes a random sample of 300 comments per category to identify usability themes such as backup/migration, setup/compatibility, and forced-to-use. The abstract's claim that these problems are 'leading to non-adoption of MFA' is a causal inference drawn from review content, but it is not a statement that is true by construction: the dataset contains only reviews from people who used or were required to use MFA apps, and it includes no observations of non-adopters. That is an external-validity or evidentiary limitation, not circular reasoning. No parameter is fitted to a subset of data and then renamed as a prediction; no target result is defined in terms of an input variable; and no uniqueness theorem or prior result is invoked to forbid alternatives. The self-citations to earlier usability studies by the same authors (e.g., Das, Dingman & Camp 2018; Das, Russo, Dingman, Dev, Kenny & Camp 2018) are used as background motivation and to support design recommendations, but the paper's central findings come from the review corpus itself and do not depend on those prior papers for their validity. Thus the derivation chain is self-contained, and the paper warrants a circularity score of 0.

Assumptions & free parameters 0 free parameters · 4 assumptions · 0 invented entities

No new parameters or entities are introduced. The analysis rests on the validity of review data, automated sentiment scores, and the filtering rules.

assumptions (4)
  • domain assumption User-generated reviews from app stores and one organizational site are representative of MFA users' experiences.
    Section 3.1 collects reviews as the sole data source; reviews are subject to self-selection and are not a random sample of users.
  • domain assumption Microsoft Azure Text Analytics sentiment scores and the 0.5 threshold correctly classify positive and negative reviews in this corpus.
    Section 3.2 and Section 4 use these scores without validating them against manual labels on this dataset.
  • domain assumption The filtering rules (complete sentence, at least 100 characters, Akismet spam filter) do not introduce systematic bias into the sentiment distribution.
    Section 3.1 describes these rules; the effect of excluding short reviews is not analyzed.
  • domain assumption The five selected applications are representative of the app-based MFA market.
    Section 3.1 lists the apps; no justification is given for excluding other popular MFA apps.

how reviews work

0 comments
Cite this review

Pith. "Pith review of MFA is a Waste of Time! Understanding Negative Connotation Towards MFA Applications via User Generated Content." pith.science (2026). https://pith.science/paper/QLDJXLBK

@misc{pith2026190805902,
  author       = {Pith},
  title        = {Pith review of: MFA is a Waste of Time! Understanding Negative Connotation Towards MFA Applications via User Generated Content},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/QLDJXLBK}},
  note         = {Machine review of arXiv:1908.05902}
}
read the original abstract

Traditional single-factor authentication possesses several critical security vulnerabilities due to single-point failure feature. Multi-factor authentication (MFA), intends to enhance security by providing additional verification steps. However, in practical deployment, users often experience dissatisfaction while using MFA, which leads to non-adoption. In order to understand the current design and usability issues with MFA, we analyze aggregated user generated comments (N = 12,500) about application-based MFA tools from major distributors, such as, Amazon, Google Play, Apple App Store, and others. While some users acknowledge the security benefits of MFA, majority of them still faced problems with initial configuration, system design understanding, limited device compatibility, and risk trade-offs leading to non-adoption of MFA. Based on these results, we provide actionable recommendations in technological design, initial training, and risk communication to improve the adoption and user experience of MFA.

Figures

Figures reproduced from arXiv: 1908.05902 by the authors.

Figure 1
Figure 1. Word Cloud for showing the distribution of title (Fig.1) and contents (Fig.2) of the user reviews are mostly targeted towards specific issues, such as device incompatibility, lack of user tool understanding, etc. Additionally, we noted that negative sentiments (56.9%) overpowered the positive yet generic connotation towards MFA. We compared overall ratings, as well as grouped user ratings, with the application versi… view at source ↗
Figure 2
Figure 2. shows that the user reviews for most applications change over a period based on their versions, however, not showing a positive linear acceptability trend. Duo Security and Okta had major review decent during development iterations, Authy had generally higher review scores across iterations, and Microsoft Authenticator and Google Authenticator had minor improvements. Combined with NLP analysis in figures 3 and 4 of … view at source ↗
Figure 3
Figure 3. Sentiment score for review titles from [0,1]. 0 represents extreme negative emotion while 1 represents the extreme positive motion. 2. Setup, Compatibility, Application and Integration Quality: As technology evolves, users who have adopted to MFA, express higher demands for MFA device compatibility, such as with smart watches. In addition, users expressed difficulties [PITH_FULL_IMAGE:figures/full_fig_p006_3.png] view at source ↗
Figures from the paper (1 more)
Figure 4
Figure 4. Figure 4: Sentiment score for the content of the user reviews valued between [0,1]. Some applications implement secure backups, but they require users to confirm their backup password routinely, since developers are concerned about users’ memorization issues. Unfortunately, this…

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

10 extracted references · 10 canonical work pages

  1. [1]

    it’s not actually that horrible

    Althobaiti, M. M. & Mayhew, P. (2014), Security and usability of authenticating process of online banking: User experience study, in ‘Security Technology (ICCST), 2014 International Carnahan Conference on’, IEEE, pp. 1–6. Amin, A., ul Haq, I. & Nazir, M. (2017), ‘Two factor authentication’, International Journal of Computer Science and Mobile Computing. A...

  2. [30]

    & Liu, B

    Jindal, N. & Liu, B. (2008), Opinion spam and analysis, in ‘Proceedings of the 2008 international conference on web search and data mining’, ACM, pp. 219 –230. Joyce, R. (2016), ‘Disrupting nation state hackers’, USENIX Enigma. San Fransisco, CA. Li, F. H., Huang, M., Yang, Y. & Zhu, X. (2011), Learning to identify review spam, in ‘Twenty-Second Internati...

  3. [39]

    & Nipprt -Eng, C

    Das, S., Kim, A., Tingl e, Z. & Nipprt -Eng, C. (2019), All about phishing exploring user research through a systematic literature review, in ‘Proceedings of the Thirteenth International Symposium on Human Aspects of Information Security & Assurance (HAISA 2019)’. Das, S., Wang, B., Tingle, Z. & Camp, L.J (2019), Evaluating User Perception of Multi -Facto...

  4. [56]

    S., Douglas, G., Richardson, T

    Weir, C. S., Douglas, G., Richardson, T. & Jack, M. (2010), ‘Usable security: User preferences for authentication methods in ebanking and the effects of experience’, Interacting with Computers 22(3), 153–

  5. [164]

    & Tygar, J

    Whitten, A. & Tygar, J. D. (1999), Why johnny can’t encrypt: A usability evaluation of pgp 5.0., in ‘USENIX Security Symposium’, Vol

  6. [200]

    Pang, B., Lee, L. et al. (2008), ‘Opinion mining and sentiment analysis’, Foundations and Trends R in Information Retrieval 2(1–2), 1–135. Ting, D. M., Hussain, O. & LaRoche, G. (2015), ‘Systems and methods for multi-factor authentication’. US Patent 9,118,656. Vasa, R., Hoon, L., Mouzakis, K. & Noguchi, A. (2012), A preliminary analysis of mobile app use...

  7. [240]

    & Sadeh, N

    Fu, B., Lin, J., Li, L., Faloutsos, C., Hong, J. & Sadeh, N. ( 2013), Why people hate your app: Making sense of user feedback in a mobile app store, in ‘Proceedings of the 19th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining’, KDD ’13, ACM, New York, NY, USA, pp. 1276–

  8. [456]

    & Camp, L

    Das, S., Dingman, A. & Camp, L. J. (2018), Why johnny doesn’t use two factor a twophase usability study of the fido u2f security key, in ‘2018 International Conference on Financial Cryptography and Data Security (FC)’. Das, S., Russo, G., Dingman, A. C., Dev, J., Kenny, O. & Camp, L. J. (2018), A qualitative study on usability and acceptability of yubico ...

Show all 10 references
  1. [1284]

    (2007), ‘A comparison of website user authentication mechanisms’, Computer Fraud & Security 2007(9), 5–9

    Furnell, S. (2007), ‘A comparison of website user authentication mechanisms’, Computer Fraud & Security 2007(9), 5–9. Hwang, M. -S. & Li, L. -H. (2000), ‘A new remote user authentication scheme using smart cards’, IEEE Transactions on consumer Electronics 46(1), 28–

  2. [2007]

    S., Das, S

    Md Noman, A. S., Das, S. & Patil, S. (2019), Rejected by techies: Understanding facebook non- adoption by experts via user generated content, in ‘Proceedings of the 2019 CHI Conference on Human Factors in Computing Systems’, ACM. Mukherjee, A., Liu, B. & Glance, N. (2012), Spo...

Pith tools

Reviewed August 14, 2026 · model on record in the stance chip above.