REVIEW 4 major objections 5 minor 10 references
MFA is a Waste of Time! Understanding Negative Connotation Towards MFA Applications via User Generated Content
T0 review · 4 major / 5 minor · reviewed 2026-08-14 · deepseek-v4-flash
Pith's one-line read The paper argues that negative reviews of MFA apps cluster on concrete usability failures rather than on doubts about security.
desk verdict A useful descriptive study of MFA app reviews, undermined by an unsupported causal leap from complaints to non-adoption. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The mechanism is the aggregate user-generated review itself, treated as a large-scale usability diary. The authors collected 12,500 comments about Duo, Google Authenticator, Microsoft Authenticator, Authy, and Okta from the Apple App Store, Google Play, Amazon, and an internal organizational site; filtered them for sentence-length content and spam; ran sentiment and keyword extraction with a commercial text-analytics service; grouped comments by application version to track changes across releases; and qualitatively coded random samples (M=300 per category). The load-bearing step is the clustering of negative keywords and themes, which turns unstructured complaints into the recurring categories of setup, backup/migration, compatibility, integration, and forced use that the recommendations target.
What would settle it
Survey a representative sample of people who were offered MFA but never enrolled, and ask whether their reasons are setup difficulty, device incompatibility, backup concerns, or a belief that MFA is unnecessary. If awareness-related reasons, such as not knowing MFA was important, dominate over the usability categories, the paper's central attribution to usability defects would be weakened.
Extended reading notes
Core claim
The central discovery is that users' negative attitudes toward MFA apps cluster on concrete engineering and policy failures rather than on skepticism about security value. Across the collected reviews, the dominant complaint categories are backup and migration (lost or non-transferable codes when changing devices), first-time setup difficulty, limited device compatibility, crashes and poor integration, and 'forced-to-use' deployments where employers or universities require MFA without explaining its benefit. The paper reports that negative comments (56.9%) outnumber positive ones and that positive reviews are mostly generic ('great app') while negative reviews name specific problems. The authors infer from this pattern that MFA non-adoption follows from avoidable usability and training gaps, and they recommend step-by-step setup instructions, risk communication that explains authentication factors, non-interrupting backup verification, NFC-based credential migration, and pre-deployment pilot testing.
Load-bearing premise
The paper relies on reviews written by people who installed or were required to use an MFA app, so its claim that negative experiences lead to non-adoption cannot be tested against people who never adopted MFA in the first place.
Editorial extensions
If this is right
- If usability failures are the main barrier, MFA adoption in organizations should rise when setup is guided, backup and migration are seamless, and device compatibility is maintained.
- Positive but generic reviews should not be read as endorsement; services that rely on required-use deployments may see high ratings while still harboring latent resistance.
- Version-by-version review patterns can serve as a regression signal for MFA developers, with drops after updates indicating usability regressions.
- Risk communication that explains what 'something you have' means could reduce the 'why do I need this?' resistance seen in forced-to-use reviews.
- Improving recovery paths, such as backup codes and device migration, may matter as much as initial enrollment in retaining MFA users.
Reading between the lines
- A direct testable extension would be to turn these complaint categories into a survey for non-adopters: people who never installed an MFA app would be asked whether setup difficulty, backup concerns, device incompatibility, or perceived benefit most explains their refusal, which would directly test the paper's causal inference.
- The same review-mining approach could be applied to hardware tokens and biometric MFA to see whether the complaint categories shift, and to password managers to compare usability failure profiles.
- The finding that forced deployment breeds hostility suggests that mandatory MFA programs paired with explanation and easy recovery may convert reluctant users, while mandatory deployment without communication may entrench resistance.
- Because the data are cross-sectional, the version-trend analysis cannot separate app changes from changes in the reviewer population; a controlled deployment study with pre- and post-measures would separate those effects.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper analyzes 12,500 user-generated reviews of five application-based MFA tools (Duo, Google Authenticator, Microsoft Authenticator, Authy, Okta) collected from Apple App Store, Google Play, Amazon Marketplace, and an unnamed organization's internal review site. The authors use the Microsoft Azure Text Analytics API for sentiment scoring and keyword extraction, plus a qualitative review of 300 randomly selected comments. They report that 56.9% of comments express negative sentiment and identify three main complaint themes: backup and migration difficulties, setup/compatibility/integration issues, and forced-to-use perceptions. Based on these findings, the paper recommends user training and risk communication, improved backup and migration mechanisms, and more extensive application and integration testing to improve MFA adoption. The stated central claim is that the majority of users face configuration, compatibility, backup, and risk-communication problems that lead to non-adoption of MFA.
Significance. If the findings hold, the paper provides a large-scale, naturally occurring complement to controlled usability studies of MFA, and its thematic results are broadly plausible and consistent with prior work. The paper's strength is its use of real user reviews rather than laboratory tasks, and its descriptive theme extraction is a reasonable starting point for understanding MFA usability complaints. However, the central causal claim about non-adoption is not supported by the data, and several methodological details are missing. The paper is likely to be useful to the usable-security community as descriptive evidence of user dissatisfaction, but its stronger contribution as evidence about adoption barriers is currently not justified.
major comments (4)
- [Abstract and Section 4] The abstract claims that the identified problems are 'leading to non-adoption of MFA.' This causal claim is not testable with the presented data. The reviews come from users who installed or were required to use MFA apps; the dataset contains no observations from individuals who never enrolled, abandoned setup, or uninstalled and refused to return. A user who complains about backup or setup may continue using the app, especially in organizational deployments. The data can support statements about dissatisfaction or negative connotation, but not a causal relationship to non-adoption. The authors should either temper the claim to 'associated with adoption barriers' or provide supplementary evidence (e.g., enrollment/uninstall logs or interviews with non-adopters).
- [Section 3.2 and Section 4] The sentiment analysis lacks validation and statistical support. The headline figure of 56.9% negative sentiment is presented without confidence intervals or a significance test. The per-application sentiment scores in Table 1 are given as point estimates with no uncertainty, so the claimed ordering (e.g., Authy highest, Okta lowest) cannot be distinguished from noise. The paper also does not report the accuracy or performance of the Azure Text Analytics API on this specific corpus, and the qualitative analysis of 300 comments has no inter-rater reliability measure or codebook. Without these, the descriptive results are not established beyond anecdote.
- [Section 3.1] The data-filtering rules may bias the sample. Specifically, comments shorter than 100 characters are discarded, yet many short negative expressions—such as 'waste of time' or 'bad app'—are exactly the kind of content the paper's title highlights. The manuscript does not report how many reviews were removed by this filter or by the anti-spam filter, nor does it give the distribution of the 12,500 comments across the five applications, the three marketplaces, and the internal organization site. This makes the representativeness of the sample impossible to assess and the analysis difficult to reproduce.
- [Section 4] The claim that 'positive comments are over-generic' and that many positive ratings are due to organizational mandates ('Great App' because it was instructed by their organization) is not supported by any quantitative or qualitative evidence presented in the paper. The authors do not define how this classification was made or provide examples and counts from the 300-comment subset. This claim is load-bearing for the conclusion that positive sentiment reflects compliance rather than satisfaction, so it needs explicit support or should be removed.
minor comments (5)
- [Section 3.1] The sentence 'we ran additional filters (N = 12500) within the collected dataset' is ambiguous: it is unclear whether 12,500 is the size before or after filtering. Please use distinct notation such as N_collected and N_final.
- [Figure 1 caption] The caption reads 'Word Cloud for showing the distribution of title (Fig.1) and contents (Fig.2)', but Figure 1 is a word cloud and Figure 2 appears to be a version-rating plot. This should be corrected to refer to the actual figures.
- [Section 1] 'We conclude by annotating crucial issues in current MFA implementation in section 4' conflicts with the paper's structure: Section 4 is Results and Analysis, not the conclusion. Please revise the wording.
- [Section 4] 'Duo Security and Okta had major review decent during development iterations' appears to be a typo; 'decent' should presumably be 'descent'.
- [References] There are small reference formatting errors: 'San Fransisco' should be 'San Francisco'; the entry for Das, Wang, Tingle & Camp (2019) has 'L.J' without a space; and the page range for Fu et al. is written as '1276– 1284' with an unwanted space.
Circularity Check
No significant circularity: the review analysis is self-contained and does not reduce to its inputs.
full rationale
This paper is an observational, empirical study rather than a derivation. It collects 12,500 user reviews of app-based MFA tools, applies sentiment analysis and keyword clustering, and qualitatively codes a random sample of 300 comments per category to identify usability themes such as backup/migration, setup/compatibility, and forced-to-use. The abstract's claim that these problems are 'leading to non-adoption of MFA' is a causal inference drawn from review content, but it is not a statement that is true by construction: the dataset contains only reviews from people who used or were required to use MFA apps, and it includes no observations of non-adopters. That is an external-validity or evidentiary limitation, not circular reasoning. No parameter is fitted to a subset of data and then renamed as a prediction; no target result is defined in terms of an input variable; and no uniqueness theorem or prior result is invoked to forbid alternatives. The self-citations to earlier usability studies by the same authors (e.g., Das, Dingman & Camp 2018; Das, Russo, Dingman, Dev, Kenny & Camp 2018) are used as background motivation and to support design recommendations, but the paper's central findings come from the review corpus itself and do not depend on those prior papers for their validity. Thus the derivation chain is self-contained, and the paper warrants a circularity score of 0.
Assumptions & free parameters
assumptions (4)
- domain assumption User-generated reviews from app stores and one organizational site are representative of MFA users' experiences.
- domain assumption Microsoft Azure Text Analytics sentiment scores and the 0.5 threshold correctly classify positive and negative reviews in this corpus.
- domain assumption The filtering rules (complete sentence, at least 100 characters, Akismet spam filter) do not introduce systematic bias into the sentiment distribution.
- domain assumption The five selected applications are representative of the app-based MFA market.
Cite this review
Pith. "Pith review of MFA is a Waste of Time! Understanding Negative Connotation Towards MFA Applications via User Generated Content." pith.science (2026). https://pith.science/paper/QLDJXLBK
@misc{pith2026190805902,
author = {Pith},
title = {Pith review of: MFA is a Waste of Time! Understanding Negative Connotation Towards MFA Applications via User Generated Content},
year = {2026},
howpublished = {\url{https://pith.science/paper/QLDJXLBK}},
note = {Machine review of arXiv:1908.05902}
}
read the original abstract
Traditional single-factor authentication possesses several critical security vulnerabilities due to single-point failure feature. Multi-factor authentication (MFA), intends to enhance security by providing additional verification steps. However, in practical deployment, users often experience dissatisfaction while using MFA, which leads to non-adoption. In order to understand the current design and usability issues with MFA, we analyze aggregated user generated comments (N = 12,500) about application-based MFA tools from major distributors, such as, Amazon, Google Play, Apple App Store, and others. While some users acknowledge the security benefits of MFA, majority of them still faced problems with initial configuration, system design understanding, limited device compatibility, and risk trade-offs leading to non-adoption of MFA. Based on these results, we provide actionable recommendations in technological design, initial training, and risk communication to improve the adoption and user experience of MFA.
Figures
Figures from the paper (1 more)
Reference graph
Works this paper leans on
-
[1]
it’s not actually that horrible
Althobaiti, M. M. & Mayhew, P. (2014), Security and usability of authenticating process of online banking: User experience study, in ‘Security Technology (ICCST), 2014 International Carnahan Conference on’, IEEE, pp. 1–6. Amin, A., ul Haq, I. & Nazir, M. (2017), ‘Two factor authentication’, International Journal of Computer Science and Mobile Computing. A...
work page 2014
-
[30]
Jindal, N. & Liu, B. (2008), Opinion spam and analysis, in ‘Proceedings of the 2008 international conference on web search and data mining’, ACM, pp. 219 –230. Joyce, R. (2016), ‘Disrupting nation state hackers’, USENIX Enigma. San Fransisco, CA. Li, F. H., Huang, M., Yang, Y. & Zhu, X. (2011), Learning to identify review spam, in ‘Twenty-Second Internati...
work page 2008
-
[39]
Das, S., Kim, A., Tingl e, Z. & Nipprt -Eng, C. (2019), All about phishing exploring user research through a systematic literature review, in ‘Proceedings of the Thirteenth International Symposium on Human Aspects of Information Security & Assurance (HAISA 2019)’. Das, S., Wang, B., Tingle, Z. & Camp, L.J (2019), Evaluating User Perception of Multi -Facto...
work page 2019
-
[56]
S., Douglas, G., Richardson, T
Weir, C. S., Douglas, G., Richardson, T. & Jack, M. (2010), ‘Usable security: User preferences for authentication methods in ebanking and the effects of experience’, Interacting with Computers 22(3), 153–
work page 2010
-
[164]
Whitten, A. & Tygar, J. D. (1999), Why johnny can’t encrypt: A usability evaluation of pgp 5.0., in ‘USENIX Security Symposium’, Vol
work page 1999
-
[200]
Pang, B., Lee, L. et al. (2008), ‘Opinion mining and sentiment analysis’, Foundations and Trends R in Information Retrieval 2(1–2), 1–135. Ting, D. M., Hussain, O. & LaRoche, G. (2015), ‘Systems and methods for multi-factor authentication’. US Patent 9,118,656. Vasa, R., Hoon, L., Mouzakis, K. & Noguchi, A. (2012), A preliminary analysis of mobile app use...
work page 2008
-
[240]
Fu, B., Lin, J., Li, L., Faloutsos, C., Hong, J. & Sadeh, N. ( 2013), Why people hate your app: Making sense of user feedback in a mobile app store, in ‘Proceedings of the 19th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining’, KDD ’13, ACM, New York, NY, USA, pp. 1276–
work page 2013
-
[456]
Das, S., Dingman, A. & Camp, L. J. (2018), Why johnny doesn’t use two factor a twophase usability study of the fido u2f security key, in ‘2018 International Conference on Financial Cryptography and Data Security (FC)’. Das, S., Russo, G., Dingman, A. C., Dev, J., Kenny, O. & Camp, L. J. (2018), A qualitative study on usability and acceptability of yubico ...
work page 2018
Show all 10 references
-
[1284]
(2007), ‘A comparison of website user authentication mechanisms’, Computer Fraud & Security 2007(9), 5–9
Furnell, S. (2007), ‘A comparison of website user authentication mechanisms’, Computer Fraud & Security 2007(9), 5–9. Hwang, M. -S. & Li, L. -H. (2000), ‘A new remote user authentication scheme using smart cards’, IEEE Transactions on consumer Electronics 46(1), 28–
2007
-
[2007]
S., Das, S
Md Noman, A. S., Das, S. & Patil, S. (2019), Rejected by techies: Understanding facebook non- adoption by experts via user generated content, in ‘Proceedings of the 2019 CHI Conference on Human Factors in Computing Systems’, ACM. Mukherjee, A., Liu, B. & Glance, N. (2012), Spo...
2019
Reviewed August 14, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.