{"id":"896ccee5-41b2-4eb9-a6d4-c27451716166","arxiv_id":"1908.05902","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"An analysis of 12,500 MFA app reviews finds negative sentiment (56.9%) driven by setup, backup, and forced-use problems, and offers design and training recommendations.","lead":"The authors analyzed 12,500 app store reviews of five multi-factor authentication apps and found that most negative comments are about setup, backup, forced adoption, and poor compatibility. The study turns user complaints into practical design and training recommendations for MFA vendors and organizations.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Review data cannot support the claim that usability complaints 'lead to non-adoption' because the dataset contains no observations of non-adopters.","rationale":"The reader identified the same load-bearing assumption: the data contain only reviews from adopters/current users, so the inference to non-adoption is untestable. I agree this is the single most load-bearing concern because the paper's title, abstract, and recommendations are framed around adoption and rejection, not just negative sentiment. The descriptive findings (negative reviews center on setup, compatibility, backup, risk communication) are plausible and supported by representative quotes, but they cannot carry the causal weight the paper assigns them. A concrete re-analysis of the dataset counting explicit refusal/uninstall statements would directly test whether any evidence of non-adoption is present. Because the paper does not share data or code, this re-analysis would require the authors to provide the dataset; if they cannot, the conditional status is warranted.","tokens_in":7116,"tokens_out":6537,"duration_ms":66073,"concrete_test":"Re-analyze the existing 12,500-review dataset to count comments that explicitly describe non-adoption behavior (e.g., 'refused to enroll', 'I deleted the app', 'I will not use this', 'stopped using MFA') and report the proportion and representative examples. If such explicit non-adoption statements are rare or absent, the abstract's 'leading to non-adoption' claim is unsupported by the data.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The abstract's central claim is that the majority of users faced configuration, compatibility, backup, and risk-communication problems 'leading to non-adoption of MFA.' The empirical basis is 12,500 reviews from app stores and one internal organizational site (Section 3.1). Those reviews are written by people who installed or were required to use an MFA app; the dataset contains no responses from individuals who never enrolled, abandoned setup, or uninstalled and refused to return. Section 4 identifies themes such as backup/migration and setup difficulty from these reviews, and Section 5 turns them into recommendations to increase adoption, but the data cannot establish a causal relation between complaints and non-adoption. A user who complains may still continue using the app (as many organizational users do). Thus the paper's core contribution, that MFA's usability defects are a principal driver of non-adoption, is not testable with the presented evidence.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper analyzes 12,500 user-generated reviews of five application-based MFA tools (Duo, Google Authenticator, Microsoft Authenticator, Authy, Okta) collected from Apple App Store, Google Play, Amazon Marketplace, and an unnamed organization's internal review site. The authors use the Microsoft Azure Text Analytics API for sentiment scoring and keyword extraction, plus a qualitative review of 300 randomly selected comments. They report that 56.9% of comments express negative sentiment and identify three main complaint themes: backup and migration difficulties, setup/compatibility/integration issues, and forced-to-use perceptions. Based on these findings, the paper recommends user training and risk communication, improved backup and migration mechanisms, and more extensive application and integration testing to improve MFA adoption. The stated central claim is that the majority of users face configuration, compatibility, backup, and risk-communication problems that lead to non-adoption of MFA.","tokens_in":7204,"tokens_out":4355,"duration_ms":42779,"significance":"If the findings hold, the paper provides a large-scale, naturally occurring complement to controlled usability studies of MFA, and its thematic results are broadly plausible and consistent with prior work. The paper's strength is its use of real user reviews rather than laboratory tasks, and its descriptive theme extraction is a reasonable starting point for understanding MFA usability complaints. However, the central causal claim about non-adoption is not supported by the data, and several methodological details are missing. The paper is likely to be useful to the usable-security community as descriptive evidence of user dissatisfaction, but its stronger contribution as evidence about adoption barriers is currently not justified.","major_comments":[{"comment":"The abstract claims that the identified problems are 'leading to non-adoption of MFA.' This causal claim is not testable with the presented data. The reviews come from users who installed or were required to use MFA apps; the dataset contains no observations from individuals who never enrolled, abandoned setup, or uninstalled and refused to return. A user who complains about backup or setup may continue using the app, especially in organizational deployments. The data can support statements about dissatisfaction or negative connotation, but not a causal relationship to non-adoption. The authors should either temper the claim to 'associated with adoption barriers' or provide supplementary evidence (e.g., enrollment/uninstall logs or interviews with non-adopters).","section":"Abstract and Section 4"},{"comment":"The sentiment analysis lacks validation and statistical support. The headline figure of 56.9% negative sentiment is presented without confidence intervals or a significance test. The per-application sentiment scores in Table 1 are given as point estimates with no uncertainty, so the claimed ordering (e.g., Authy highest, Okta lowest) cannot be distinguished from noise. The paper also does not report the accuracy or performance of the Azure Text Analytics API on this specific corpus, and the qualitative analysis of 300 comments has no inter-rater reliability measure or codebook. Without these, the descriptive results are not established beyond anecdote.","section":"Section 3.2 and Section 4"},{"comment":"The data-filtering rules may bias the sample. Specifically, comments shorter than 100 characters are discarded, yet many short negative expressions—such as 'waste of time' or 'bad app'—are exactly the kind of content the paper's title highlights. The manuscript does not report how many reviews were removed by this filter or by the anti-spam filter, nor does it give the distribution of the 12,500 comments across the five applications, the three marketplaces, and the internal organization site. This makes the representativeness of the sample impossible to assess and the analysis difficult to reproduce.","section":"Section 3.1"},{"comment":"The claim that 'positive comments are over-generic' and that many positive ratings are due to organizational mandates ('Great App' because it was instructed by their organization) is not supported by any quantitative or qualitative evidence presented in the paper. The authors do not define how this classification was made or provide examples and counts from the 300-comment subset. This claim is load-bearing for the conclusion that positive sentiment reflects compliance rather than satisfaction, so it needs explicit support or should be removed.","section":"Section 4"}],"minor_comments":[{"comment":"The sentence 'we ran additional filters (N = 12500) within the collected dataset' is ambiguous: it is unclear whether 12,500 is the size before or after filtering. Please use distinct notation such as N_collected and N_final.","section":"Section 3.1"},{"comment":"The caption reads 'Word Cloud for showing the distribution of title (Fig.1) and contents (Fig.2)', but Figure 1 is a word cloud and Figure 2 appears to be a version-rating plot. This should be corrected to refer to the actual figures.","section":"Figure 1 caption"},{"comment":"'We conclude by annotating crucial issues in current MFA implementation in section 4' conflicts with the paper's structure: Section 4 is Results and Analysis, not the conclusion. Please revise the wording.","section":"Section 1"},{"comment":"'Duo Security and Okta had major review decent during development iterations' appears to be a typo; 'decent' should presumably be 'descent'.","section":"Section 4"},{"comment":"There are small reference formatting errors: 'San Fransisco' should be 'San Francisco'; the entry for Das, Wang, Tingle & Camp (2019) has 'L.J' without a space; and the page range for Fu et al. is written as '1276– 1284' with an unwanted space.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The paper is a descriptive, exploratory study that, with appropriate claim softening, could be a valuable usable-security contribution. The present version overstates the strength of the evidence by asserting a causal link to non-adoption that the data cannot support. The methodological gaps (no uncertainty quantification, no inter-rater reliability, opaque filtering) are fixable through revision and reporting. If the authors are unwilling to temper the central claim, the paper would be difficult to defend. I would encourage the editor to request a major revision focusing on claim calibration and methodological transparency."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague—\n\nTwo things to know. This paper does something genuinely useful: it mines 12,500 app store reviews for five MFA apps and shows that negative reviews concentrate on setup, backup/migration, compatibility, and forced adoption, while positive reviews are mostly generic \"Great App\" comments. The second thing is that the abstract and conclusion overclaim that these complaints lead to non-adoption, and the dataset cannot support that. The reviews come from people who installed and used the apps; there are no observations of people who never enrolled, abandoned setup, or uninstalled and refused to return. A complaint is not evidence of abandonment.\n\nThe descriptive core is solid enough. The five apps in Table 1 show a plausible spread—Authy fares best, Google Authenticator and Okta worst. The representative quotes are well chosen and match known usability problems from earlier qualitative work on 2FA (Colnago et al. 2018, Das et al. 2018). The point that positive reviews are often generic while negative reviews are specific is a nice observation, and the recommendations on NFC-based migration and non-interrupting backup verification are reasonable. This is a legitimate new application of sentiment analysis to a new corpus, even if the resulting themes are not conceptually novel.\n\nThe soft spots are real. The methods section is under-specified: the filtering criteria say N=12500 both before and after filtering, which has to be a typo; there is no detail on how the Azure Text Analytics sentiment threshold was applied, no inter-rater reliability for the qualitative coding of the 300 comments, and no explanation of where the 56.9% negative figure comes from given that Table 1 only gives per-app average sentiment scores. There are no error bars or statistical tests anywhere. For a descriptive study that is acceptable, but the paper should say so and stop making causal claims.\n\nThe central interpretive leap—from negative reviews to non-adoption—shows up in the abstract, Section 4, and Section 5. It is load-bearing for the paper's framing. The fix is straightforward: reframe the contribution as characterizing user dissatisfaction and identifying usability barriers among current users, and note that the relationship to non-adoption is a hypothesis for future work. That is a much weaker claim but still a useful one.\n\nWho this is for: usable security researchers looking at MFA adoption, and practitioners designing authentication flows. It deserves a serious referee because the data collection effort is real and the descriptive findings are citable even after the claims are softened. I would send it to peer review with a request for major revision on methods transparency and claim framing, and ideally a release of the filtered dataset so others can reproduce the sentiment numbers.\n\nRecommendation: engage, but require the adoption claims to be cut or reframed before publication.","headline":"A useful descriptive study of MFA app reviews, undermined by an unsupported causal leap from complaints to non-adoption.","tokens_in":7755,"tokens_out":2598,"would_cite":true,"duration_ms":25808,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper argues that negative reviews of MFA apps cluster on concrete usability failures rather than on doubts about security.","keywords":["multi-factor authentication","MFA usability","user-generated content","sentiment analysis","app store reviews","adoption barriers","authentication","user experience"],"falsifier":"Survey a representative sample of people who were offered MFA but never enrolled, and ask whether their reasons are setup difficulty, device incompatibility, backup concerns, or a belief that MFA is unnecessary. If awareness-related reasons, such as not knowing MFA was important, dominate over the usability categories, the paper's central attribution to usability defects would be weakened.","tokens_in":6878,"feed_emoji":"🔐","tokens_out":8761,"duration_ms":70441,"temperature":0.7,"pith_summary":"The paper claims that negative user reviews of multi-factor authentication (MFA) apps are not primarily a complaint about security inconvenience, but a record of specific usability failures: setup, backup and migration, device compatibility, integration quality, and a sense of being forced to use MFA without explanation. Analyzing 12,500 reviews of five app-based MFA tools across app stores and an organizational site, the authors find 56.9% of comments negative, while positive comments tend to be generic. The authors argue these usability defects, not lack of security awareness, drive non-adoption of MFA. If correct, the result implies that improving configuration, recovery, and device support, along with risk communication, would increase adoption more than simply telling users MFA is secure.","feed_headline":"MFA app reviews: usability, not ignorance, drives rejection","feed_subtitle":"Fixable usability flaws, not security doubts, drive negative reviews of MFA apps.","key_machinery":"The mechanism is the aggregate user-generated review itself, treated as a large-scale usability diary. The authors collected 12,500 comments about Duo, Google Authenticator, Microsoft Authenticator, Authy, and Okta from the Apple App Store, Google Play, Amazon, and an internal organizational site; filtered them for sentence-length content and spam; ran sentiment and keyword extraction with a commercial text-analytics service; grouped comments by application version to track changes across releases; and qualitatively coded random samples (M=300 per category). The load-bearing step is the clustering of negative keywords and themes, which turns unstructured complaints into the recurring categories of setup, backup/migration, compatibility, integration, and forced use that the recommendations target.","core_discovery":"The central discovery is that users' negative attitudes toward MFA apps cluster on concrete engineering and policy failures rather than on skepticism about security value. Across the collected reviews, the dominant complaint categories are backup and migration (lost or non-transferable codes when changing devices), first-time setup difficulty, limited device compatibility, crashes and poor integration, and 'forced-to-use' deployments where employers or universities require MFA without explaining its benefit. The paper reports that negative comments (56.9%) outnumber positive ones and that positive reviews are mostly generic ('great app') while negative reviews name specific problems. The authors infer from this pattern that MFA non-adoption follows from avoidable usability and training gaps, and they recommend step-by-step setup instructions, risk communication that explains authentication factors, non-interrupting backup verification, NFC-based credential migration, and pre-deployment pilot testing.","pith_inferences":["A direct testable extension would be to turn these complaint categories into a survey for non-adopters: people who never installed an MFA app would be asked whether setup difficulty, backup concerns, device incompatibility, or perceived benefit most explains their refusal, which would directly test the paper's causal inference.","The same review-mining approach could be applied to hardware tokens and biometric MFA to see whether the complaint categories shift, and to password managers to compare usability failure profiles.","The finding that forced deployment breeds hostility suggests that mandatory MFA programs paired with explanation and easy recovery may convert reluctant users, while mandatory deployment without communication may entrench resistance.","Because the data are cross-sectional, the version-trend analysis cannot separate app changes from changes in the reviewer population; a controlled deployment study with pre- and post-measures would separate those effects."],"forward_implications":["If usability failures are the main barrier, MFA adoption in organizations should rise when setup is guided, backup and migration are seamless, and device compatibility is maintained.","Positive but generic reviews should not be read as endorsement; services that rely on required-use deployments may see high ratings while still harboring latent resistance.","Version-by-version review patterns can serve as a regression signal for MFA developers, with drops after updates indicating usability regressions.","Risk communication that explains what 'something you have' means could reduce the 'why do I need this?' resistance seen in forced-to-use reviews.","Improving recovery paths, such as backup codes and device migration, may matter as much as initial enrollment in retaining MFA users."],"supporting_citations":[{"why":"Justifies mining app-store reviews to learn what users care about in an application.","marker":"Fu et al. 2013"},{"why":"Supplies the lexicon-based opinion-mining technique behind the review scoring.","marker":"Ding, Liu & Yu 2008"},{"why":"Anchors the sentiment-analysis step used to split reviews into positive and negative.","marker":"Pang, Lee et al. 2008"},{"why":"Provides prior evidence on organization-wide MFA adoption that the review analysis extends.","marker":"Colnago et al. 2018"},{"why":"Shows enrollment and verification problems plagued FIDO U2F, the same setup barrier found here.","marker":"Das, Dingman & Camp 2018"},{"why":"Motivates the large-N collection strategy because most app reviews are short and information-poor.","marker":"Vasa et al. 2012"},{"why":"Precedent for using user-generated content to study non-adoption of a security technology.","marker":"Md Noman et al. 2019"},{"why":"Establishes the pattern that security tools fail when usability is neglected.","marker":"Whitten & Tygar 1999"}],"fun_headline_variants":["MFA backlash: setup and backup, not security doubts","Why MFA apps get panned: lost codes and forced use","Usability pain drives MFA rejection, not security skepticism","MFA reviews: users blame design, not distrust","Fix MFA setup and backup, or users will bail"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The paper relies on reviews written by people who installed or were required to use an MFA app, so its claim that negative experiences lead to non-adoption cannot be tested against people who never adopted MFA in the first place.","fun_headline_variants_meta":{"raw":{"variants":["MFA backlash: setup and backup, not security doubts","Why MFA apps get panned: lost codes and forced use","Usability pain drives MFA rejection, not security skepticism","MFA reviews: users blame design, not distrust","Fix MFA setup and backup, or users will bail"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000252,"raw_usage":{"total_tokens":1523,"prompt_tokens":867,"completion_tokens":656,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":483,"completion_tokens_details":{"reasoning_tokens":574}},"tokens_in":483,"tokens_out":656,"duration_ms":6657,"temperature":1.0,"reasoning_tokens":574,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T13:00:32.101029+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Survey a representative sample of people who were offered MFA but never enrolled, and ask whether their reasons are setup difficulty, device incompatibility, backup concerns, or a belief that MFA is unnecessary. If awareness-related reasons, such as not knowing MFA was important, dominate over the usability categories, the paper's central attribution to usability defects would be weakened.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Anchors the sentiment-analysis step used to split reviews into positive and negative."},{"cited_title":"& Camp, L","cited_arxiv_id":null,"evidence_quote":"Shows enrollment and verification problems plagued FIDO U2F, the same setup barrier found here."},{"cited_title":"S., Das, S","cited_arxiv_id":null,"evidence_quote":"Precedent for using user-generated content to study non-adoption of a security technology."},{"cited_title":"& Tygar, J","cited_arxiv_id":null,"evidence_quote":"Establishes the pattern that security tools fail when usability is neglected."}],"review_version":1}