{"id":"af77fb1c-f93a-4f6b-a270-38c28fa8af93","arxiv_id":"2505.12104","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"Quishing emails are as effective as traditional link-based phishing in driving employees to fake login pages, and OSINT-fed LLM emails lured up to 31.5 percent of openers in one company.","lead":"An international team sent over 71,000 phishing emails to employees of three companies to measure how QR-code and AI-written phishing emails perform against real workers. The results show QR-code lures work just as well as ordinary links, while low-cost AI-generated emails can fool a surprisingly large share of recipients in smaller organizations.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"RQ1's headline equivalence result is dominated by C_h, where E_Q was an uncontrolled company-run simulation (Jan 2024, all employees) compared with the authors' later E_B on a random half; the button-vs-QR variable is entangled with timing, population, and prior training exposure.","rationale":"The strongest claim combines RQ1 and RQ2, but RQ1 is the pillar that requires a controlled comparison between two delivery mechanisms. The paper only approximates that control at C_s and C_m; at C_h, E_Q is a company-run campaign from January 2024 and E_B is a later author-run campaign on a random half of employees. Because C_h supplies almost all of the aggregate data, the equivalence result and TOST are essentially measuring the similarity of two different campaigns rather than the effect of a QR code versus a button. The paper's limitation section addresses the missing credential data for C_h, but it does not address this much more serious threat to the RQ1 comparison. A reanalysis that treats company as a random effect or restricts to the controlled sites would settle whether the conclusion survives. I agree with the reader's weakest assumption. I also note RQ3 relies on only three points and should be reframed, but that is not the load-bearing issue for the paper's central takeaway, and the reader already flags it. The verdict remains conditional because the paper's field data are valuable and the concern is testable, not because the current analysis is conclusive.","tokens_in":35149,"tokens_out":5777,"duration_ms":62454,"concrete_test":"Recompute RQ1 without pooling raw counts: fit a mixed-effects logistic regression with company as a random intercept and email type (E_B vs E_Q) as the fixed effect, and separately run the chi-square/TOST on the C_s+C_m subset only. If the 95% confidence interval for the E_B−E_Q difference no longer falls within ±1%, or the company random-intercept variance is large, the equivalence conclusion is an artifact of the uncontrolled C_h comparison rather than of the interaction mechanism.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 4.2.2 states that C_h did not require E_Q because the company had recently run its own QR simulation (Jan 2024) and shared results; the authors then designed E_B to match 'that version of E_Q'. Thus the largest company does not provide a controlled button-vs-QR comparison: E_Q was sent earlier, to all employees, through C_h's own infrastructure, while E_B was sent later by the authors to a random half. The aggregate is dominated by this site (34,031 of 34,610 E_Q emails and 17,751 of 18,339 E_B emails), so the near-identical aggregate click rates and the TOST equivalence at ±1% are essentially a C_h statement. Any differential effect of timing (January vs April/May), of having recently been exposed to a QR simulation and its follow-up training, or of the unverifiable 'matching' between the two C_h emails is absorbed into the measured equivalence. The paper acknowledges the lack of control in §7.2 but only defends the missing credential data, not the confounded comparison; the C_s and C_m results are too small to rescue the equivalence claim. Since RQ1 is the paper's first central pillar, the 'same effectiveness' conclusion is not established unless the aggregate analysis can be shown to be insensitive to C_h.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper presents a field study of phishing susceptibility across three organizations (an SME, a mid-size financial company, and a large manufacturer), based on 71,309 sent emails. Three email types are compared: a traditional button-based phishing email (E_B), a nearly identical QR-code phishing email (E_Q), and an OSINT-fed LLM-generated phishing email (E_L). The authors claim that E_B and E_Q have practically the same effectiveness (RQ1), that LLM/OSINT email is cheap and highly effective especially against smaller companies (RQ2), and that a perceived-phishing-awareness score is a statistically significant predictor of phishing susceptibility (RQ3). The paper also reports a survey of 131 employees and an original demonstration that QR-code emails evade a commercial filter.","tokens_in":35372,"tokens_out":8007,"duration_ms":84550,"significance":"The paper contains rare and valuable multi-organization field data, with fine-grained per-company breakdowns, a transparent appendix, and ethical disclosures. The filter-evasion demonstration in Appendix B is a useful original contribution, and the dataset could serve as a benchmark for later work. However, the two central inferential claims need substantial qualification: the RQ1 equivalence conclusion is dominated by an uncontrolled company-run quishing campaign at C_h, and the RQ3 regression is fit to only three company-level points. The descriptive material is significant, but the paper's strong conclusions currently outrun the evidence.","major_comments":[{"comment":"The RQ1 equivalence claim is not established because the largest company does not provide a controlled button-vs-QR comparison. For C_h, E_Q was run by the company itself in January 2024 and sent to all employees, while E_B was sent by the authors in April/May 2024 to a random half of employees; the two conditions therefore differ in timing, population, sender infrastructure, and prior exposure to quishing simulations and training. The statement in §4.2.2 that these variations are harmless because E_Q and E_B 'retain consistent properties within each organization' is not justified for C_h, and the claim that E_B was designed to 'match' the company-run E_Q is unverifiable under NDA. Since C_h accounts for 34,031 of 34,610 E_Q emails and 17,751 of 18,339 E_B emails, the aggregate click rates and the TOST equivalence at ±1% are essentially an uncontrolled C_h statement. Recomputing on C_s and C_m alone gives E_B 14/321 (4.4%) versus E_Q 20/330 (6.1%), a difference of about 1.7 percentage points with a confidence interval far wider than ±1%; the equivalence conclusion cannot be reproduced on the controlled subset. The authors should either provide a sensitivity analysis that excludes C_h, or substantially weaken the RQ1 conclusion in §5.2, the abstract, and the contributions.","section":"§4.2.2, §5.2, Table 2"},{"comment":"The RQ3 claim that perceived phishing awareness is a 'predictor' of phishing susceptibility is based on a linear regression with only three company-level data points (n=3), which yields R²=1, p<.001, and Spearman's ρ=-1 by construction. Adding the aggregate as a fourth point does not solve the problem because that point is a weighted combination of the same three observations and is not independent. With three points, the model cannot provide meaningful evidence for a predictive relationship; a perfect fit is a mathematical artifact. The paper should reframe RQ3 as an exploratory observation about three companies, and should not present the fitted line, or the derived PPA=1 or PPA=4.7 predictions, as a general predictive model.","section":"§6.2, Fig. 11"}],"minor_comments":[{"comment":"The text states that for E_Q '25,172 employees opened it, and 1,970 (8.49%) visited the landing page,' but Table 2 reports 7.8% for this ratio (1,970/25,172 = 7.83%). Please reconcile the rate in the text with the table.","section":"§5.2"},{"comment":"The company-specific one-tailed chi-square p-values should be verified: for C_m, chi-square=0.514 cannot yield a one-tailed p-value of 1.0 under the stated directional hypothesis. Even if the conclusion is unchanged, the reported p-value appears incorrect and should be recomputed or the test described more precisely.","section":"§5.2"},{"comment":"The questionnaire was distributed after the phishing simulations, and the paper assumes it is 'reasonable to expect' that respondents had also taken part in the simulation. This should be stated more cautiously as an assumption, since the responses cannot be linked to individual simulation outcomes.","section":"§4.3"}],"recommendation":"major_revision","confidential_remarks":"The paper contains valuable field data and a useful filter-evasion experiment, but the two headline claims (RQ1 equivalence and RQ3 prediction) are not supported by the current analysis. A major revision should require either a sensitivity analysis excluding C_h for RQ1 and a downgrade of RQ3 to an exploratory finding, or a clear statement that these claims are not established. The descriptive dataset itself is worth publishing if the conclusions are brought in line with the evidence."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this paper has a genuinely new empirical core and should go to peer review, but the central RQ1 claim is weaker than the abstract suggests because the largest company's quishing arm was not controlled by the authors. Read it for the descriptive LLM/OSINT results and the transparency of the appendix.\n\nThe novelty is real: the authors' own systematic review of 11 top venues found no prior user study of QR-code phishing in organizations, and this three-company field study is the first to compare button-based and QR-code phishing on real employees. The LLM/OSINT extension across companies is also a new data point, and the paper is unusually forthcoming about the exact prompts, infrastructure, and questionnaire items, which makes the work reusable for replication. That transparency earns credit.\n\nThe soft spot is exactly where the stress-test note lands. For C_h, E_Q was run by the company itself in January 2024, sent to all employees, while E_B was run by the authors months later on a random half. The two arms differ in timing, population, prior training exposure, and infrastructure, so the button-vs-QR variable is not isolated. Since C_h contributes 34,031 of 34,610 E_Q emails and 17,751 of 18,339 E_B emails, the aggregate equivalence and the TOST result are essentially a statement about that uncontrolled comparison. The paper acknowledges the missing credential data in §7.2 but does not confront the confounded comparison itself. I also agree the per-company descriptive data are plausible on their own: C_h's internal numbers (8.1% vs 7.9%) happen to look similar, but that is not a substitute for a controlled experiment.\n\nThe other issues are real but smaller. RQ3 is a three-point regression with a perfect fit, and the paper already calls it a gross generalization; the abstract nevertheless sells it as a predictor. The p-value mismatch (abstract says .552, main text says .276) needs fixing. These are correctable.\n\nWhat survives scrutiny: the LLM/OSINT results are well documented, the 31.5% visit rate at C_m is a striking and concerning datapoint, and the cross-company comparison of training and reporting behavior is a useful addition to the phishing literature. The paper is not a waste of anyone's time.\n\nRecommendation: send it to peer review, but with a required major revision: reanalyze RQ1 without the uncontrolled C_h data (or explicitly label it as a separate non-controlled case study), report confidence intervals on all effectiveness ratios, correct the p-values, and reframe RQ3 as exploratory rather than predictive. With those changes, this becomes a solid empirical contribution. I would bring it to our reading group and would cite the LLM/OSINT data.","headline":"Useful first multi-org field data on quishing and LLM/OSINT phishing, but the headline RQ1 equivalence claim rides on a company-run, uncontrolled E_Q at the largest site and needs a reanalysis before it can be taken at face value.","tokens_in":35980,"tokens_out":1807,"would_cite":true,"duration_ms":21164,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"QR-code phishing emails lure employees to fake login pages just as often as traditional click-through buttons, and LLM-written emails add a cheap, high-impact vector.","keywords":["phishing","quishing","QR codes","LLM-generated phishing","OSINT","phishing simulation","perceived phishing awareness","organizational security"],"falsifier":"Run a single-organization randomized controlled trial where E_B and E_Q are sent simultaneously from the same platform to randomly split employees, with identical pretext and landing page; if the QR arm's landing-page visit rate falls outside the ±1% equivalence margin relative to the button arm, the paper's RQ1 claim fails. A cheaper check: re-analyze the multinational data using only the concurrently randomized half of employees who received E_B and the matched E_Q population, and see whether the 8.1% vs 7.9% aggregate result still holds.","tokens_in":34896,"feed_emoji":"🔳","tokens_out":5193,"duration_ms":51923,"temperature":0.7,"pith_summary":"The paper tries to establish that two emerging phishing vectors are at least as dangerous as the classic one. Across 71,309 emails sent to employees of three organizations, QR-code phishing (\"quishing\") brought users to a fake login page at the same rate as a nearly identical email with a click-through button, with an equivalence test placing the difference below one percent. The paper also tries to show that feeding publicly available company information to a free large language model produces phishing emails that are cheap to craft and highly effective, with over 30% of opened emails in the mid-sized company leading to the landing page. A survey of 131 employees suggests that higher self-reported \"perceived phishing awareness\" predicts lower measured susceptibility to these campaigns. If these claims hold, organizations cannot assume that QR codes add protective friction, and they face a new class of low-cost, personalized phishing.","feed_headline":"QR-code phishing lures users as well as button clicks","feed_subtitle":"In 71k simulated emails, quishing matched click-through rates while LLM-written emails excelled against smaller firms.","key_machinery":"The load-bearing object is a controlled pair of near-identical phishing emails: E_B embeds the URL in a click-through button; E_Q replaces that button with a QR code while keeping the design, pretext, and landing page the same. This isolates the interaction mechanism as the variable of interest. The second mechanism is an OSINT-to-LLM pipeline: data from an employer-rating site, social-network posts, and press releases is summarized by a free LLM through a five-prompt sequence to produce a persuasive company-survey invitation (E_L). The third mechanism is a 40-question survey, rooted in knowledge-attitude-behavior principles, whose aggregated score (PPA) is regressed against the phishing click-through rate (PS) per company.","core_discovery":"On the paper's own terms, the central discovery is a pair of empirical equivalences and one correlation. RQ1: employees who open a quishing email reach the credential-harvesting webpage about as often as employees who open a traditional button-based email (aggregate 8.0% vs 8.5% of opened emails, p=.276 for the one-tailed difference, with a TOST equivalence within ±1%). RQ2: an email written by a free LLM using OSINT from employer-rating sites, LinkedIn, and press releases outperformed both traditional emails in the small and mid-sized companies (66.6% and 31.5% of opened emails led to the landing page respectively, versus 22.2% and 3.9% for the button email), although it underperformed them at the multinational. RQ3: a linear regression across the three companies found perceived phishing awareness a significant negative predictor of phishing susceptibility (slope -24.2, p<.001).","pith_inferences":["The paper's aggregate quishing equivalence is dominated by the multinational's earlier QR campaign, which was not run concurrently or on the same platform as the button email; a fully matched randomized comparison could either confirm or narrow the equivalence.","If the PPA-PS relationship generalizes beyond three companies, annual survey-based PPA scoring could become a low-cost diagnostic for organizations before investing in phishing training.","The OSINT+LLM result implies attackers can automate the reconnaissance-to-email pipeline; one testable extension is measuring whether adding urgency or loss cues to E_L pushes its already high click rates higher.","The lower credential-submission rate after QR scans may be a device artifact (credentials stored in a password manager on a work device); testing with a mobile-friendly password autofill landing page would separate friction from skepticism."],"forward_implications":["Security teams cannot treat QR codes as a natural defense: quishing reached the landing page at the same rate as a button, so training and detection must treat QR codes as a first-class phishing vector.","LLM-generated emails fed with public information are cheap to produce and can outperform carefully crafted traditional phishing emails, particularly in smaller, more homogeneous organizations.","Perceived phishing awareness, measured by survey, may serve as a rough predictor of organizational phishing susceptibility, allowing pre-emptive training prioritization.","Because quishing URLs are hard for standard filters to see, client-side QR scanning with URL checks becomes a plausible defense, alongside including quishing in simulated phishing exercises.","The finding that fewer credentials were submitted after QR scans relative to button clicks suggests the device gap between scanning and typing credentials may be a real mitigation point, though the paper attributes this partly to experimental setup."],"supporting_citations":[{"why":"Supplies the motivational landscape: phishing prevalence, rising LLM usage by attackers, and training gaps.","marker":"[10]"},{"why":"Grounds the claim that QR codes evade email filters, motivating the quishing research question.","marker":"[4]"},{"why":"Provides the OSINT-plus-LLM crafting methodology that E_L adapts, plus a benchmark visit rate for comparison.","marker":"[29]"},{"why":"A prior LLM spear-phishing human study used as a comparative benchmark for E_L's effectiveness.","marker":"[45]"},{"why":"Prior user-study evidence of human susceptibility to QR-code phishing, used as a baseline for RQ1.","marker":"[89]"},{"why":"The chi-square test procedure used to test whether button and QR emails differ in click-through.","marker":"[76]"},{"why":"The two one-sided test (TOST) procedure that establishes equivalence within the ±1% margin.","marker":"[88]"},{"why":"Supplies the questionnaire indicators and framework used to compute the perceived phishing awareness score.","marker":"[30]"},{"why":"Provides the knowledge-attitude-behavior principles on which the survey design is based.","marker":"[86]"}],"fun_headline_variants":["Quishing ties click-based phishing in 71k email test","LLM-written phishing emails outperform in smaller firms","QR codes and AI pretexting slip past current phishing defenses","Phishing drills: quishing ties clicks, LLMs beat humans in small firms","Higher self-reported awareness correlates with phishing resilience"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"For RQ1, the paper assumes the button and QR emails differ only in interaction mechanism within each organization, but at the largest company the QR campaign was run earlier by the company itself on all employees, so timing, population, platform, and prior training exposure are not controlled.","fun_headline_variants_meta":{"raw":{"variants":["Quishing ties click-based phishing in 71k email test","LLM-written phishing emails outperform in smaller firms","QR codes and AI pretexting slip past current phishing defenses","Phishing drills: quishing ties clicks, LLMs beat humans in small firms","Higher self-reported awareness correlates with phishing resilience"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001271,"raw_usage":{"total_tokens":5242,"prompt_tokens":1031,"completion_tokens":4211,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":647,"completion_tokens_details":{"reasoning_tokens":4139}},"tokens_in":647,"tokens_out":4211,"duration_ms":30526,"temperature":1.0,"reasoning_tokens":4139,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T20:41:05.670057+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run a single-organization randomized controlled trial where E_B and E_Q are sent simultaneously from the same platform to randomly split employees, with identical pretext and landing page; if the QR arm's landing-page visit rate falls outside the ±1% equivalence margin relative to the button arm, the paper's RQ1 claim fails. A cheaper check: re-analyze the multinational data using only the concurrently randomized half of employees who received E_B and the matched E_Q population, and see whether the 8.1% vs 7.9% aggregate result still holds.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Prior user-study evidence of human susceptibility to QR-code phishing, used as a baseline for RQ1."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"The chi-square test procedure used to test whether button and QR emails differ in click-through."},{"cited_title":"Systems and Software(2024)","cited_arxiv_id":null,"evidence_quote":"The two one-sided test (TOST) procedure that establishes equivalence within the ±1% margin."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the knowledge-attitude-behavior principles on which the survey design is based."}],"review_version":1}