{"id":"24a6daab-1f84-4c66-a70a-08356a0a68f6","arxiv_id":"2505.09875","paper_version":2,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Unintended consequences of web-browsing GUI agents fall into input, action, output, and feedback failures, with impacts ranging from frustration and financial loss to privacy breaches and eroded trust.","lead":"This paper collects 221 Reddit posts and 14 interviews to catalog what goes wrong when AI agents browse the web for people. It groups the failures, their impacts, and user workarounds into a framework meant to guide safer agent design.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The taxonomy's construct validity is the weak link: several core categories are anticipated tradeoffs or hypothetical fears, not 'unanticipated, negative outcomes' as defined in Section 1.","rationale":"I read the paper as an exploratory qualitative study whose central contribution is an initial, user-grounded taxonomy of GUI-agent problems and their perceived consequences. The strongest part of the evidence is the triangulation of Reddit posts and interviews, the transparent interview protocol in Appendix C, and the honest limitations section. The paper does not overclaim prevalence or causal certainty, and it explicitly frames the analysis as an initial characterization. However, the load-bearing assumption is not only that the sample is representative, but that the analyzed material is actually about 'unintended consequences' under the authors' own definition. Slow response and high cost are included as core Feedback-stage phenomena, yet they are usually foreseeable tradeoffs; several security and privacy 'influences' are anticipatory anxieties rather than experienced harms. This means Sections 4.5.3, 4.5.4, and parts of 5.2 may not satisfy the paper's boundary condition for a UC. That is a construct-validity issue rather than a sampling issue, which is why my agreement with the reader is only partial. The proposed audit would settle whether the categories survive a strict application of the definition. Even if the audit shows a substantial share of non-UC items, the appropriate outcome is not rejection but a CONDITIONAL acceptance with a required reframing: either narrow the taxonomy to experienced, unanticipated outcomes, or explicitly broaden the definition and adjust the title and claims accordingly. Since the reader already assigned CONDITIONAL and that remains the correct verdict, I recommend UNCHANGED.","tokens_in":25301,"tokens_out":3812,"duration_ms":42154,"concrete_test":"Conduct a two-rater audit of every quoted evidence item (P# and I# quotes) used in Sections 4 and 5. For each quote, raters independently classify it as E (experienced, unanticipated negative outcome), A (anticipated tradeoff, e.g., known slowness or cost), or H (hypothetical risk or fear with no reported incident). Compute per-category proportions and inter-rater agreement (Cohen's kappa). If, for example, more than 30% of quotes in the Feedback stage are A, or more than 50% of security/privacy influence quotes are H, then the taxonomy conflates UCs with general dissatisfaction and perceived risk; the paper should either restrict the taxonomy to E items or explicitly redefine UCs to include anticipated tradeoffs and perceived risks, then relabel the claims accordingly.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper defines UCs as 'unanticipated, negative outcomes' (Section 1), so the central claim that failures cluster into a four-stage lifecycle taxonomy holds only if every core category is actually a UC under that definition. Two parts of the data do not clearly satisfy this. First, Slow System Response (§4.5.3) and High Operational Costs (§4.5.4) are central Feedback-stage categories, but these are typically anticipated tradeoffs: users often know before engaging that agents are slow or expensive, and quotes such as 'too slow and expensive for any real-life scenarios' (P70) read as evaluations of known limitations, not reports of surprise. A user who chooses an agent despite known latency has not experienced an 'unanticipated' outcome. Second, much of the security/privacy 'influence' evidence in §5.2 is phrased as anxiety or hypothetical risk ('users expressed anxieties,' 'fears,' 'worried,' 'could'), not as documented harms. The central claim's cascade—task failures and frustration → security/privacy harms → erosion of trust—is therefore only partially grounded in experienced UCs; for the middle level, the data consist largely of perceived risk and speculation. If a substantial share of the quotes are anticipated tradeoffs or hypothetical concerns, the resulting taxonomy is better described as a taxonomy of negative experiences, worries, and costs, and the paper's own definition of 'unintended' needs to be reconciled—either by broadening the definition or by removing or relabeling those categories. The reader flagged sampling and retrospective recall; this is a more internal concern because it threatens the very object being characterized, not just the sample's representativeness.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper presents a qualitative, exploratory study of unintended consequences (UCs) in human collaboration with LLM-based GUI agents for web browsing. The authors analyzed 221 Reddit posts (after manually screening 1,850) and conducted 14 semi-structured interviews. They induce a four-stage lifecycle taxonomy of UC phenomena (Input: comprehension/planning; Action: execution/interaction; Output: generation; Feedback: adjustment/viability) and then relate these to three levels of influence: negative impacts on task execution and user experience, security/privacy risks, and broader erosion of trust and social/ethical concerns. They also synthesize user-initiated mitigations (system-oriented and user-oriented) and derive design implications for safer, more transparent GUI agents. The paper's contribution is framed as an initial characterization of UCs around GUI agents, grounded in users' retrospective self-reports.","tokens_in":25548,"tokens_out":4355,"duration_ms":44896,"significance":"The paper addresses a timely and under-studied topic: the real-world failure modes of LLM-driven GUI agents as experienced by users, rather than on benchmarks or curated tasks. Its strength is that the taxonomy is induced from original user data (Reddit posts and interviews), with many categories supported by direct quotes, and the methods are transparent, including an appendix with the interview protocol and ethical considerations. If the taxonomy holds, it offers a useful starting map for future HCI/CSCW research on agent failures, user coping strategies, and design implications. The paper does not ship machine-checked proofs or code, but it does provide reproducible qualitative data excerpts and a clear methodological description. The main risk is construct validity: the paper defines UCs as 'unanticipated, negative outcomes,' yet several prominent categories appear to be anticipated tradeoffs or hypothetical perceived risks rather than unanticipated harms. The influence cascade ('task failures and frustration → security/privacy harms → erosion of trust') is only partially grounded because much of the middle-level evidence is phrased as anxiety or speculation.","major_comments":[{"comment":"Section 1 defines UCs as 'unanticipated, negative outcomes in the interactive processes of these GUI agents.' However, the Feedback-stage categories 'Slow System Response' (§4.5.3) and 'High Operational Costs' (§4.5.4) are typically anticipated tradeoffs: users often know before engaging that agents may be slow or expensive. Quotes such as P70's 'too slow and expensive for any real-life scenarios' read as evaluations of known limitations, not reports of surprise. If these categories are retained as core UCs, the definition must be broadened (e.g., to 'negative outcomes and costs, whether anticipated or not'), or these categories should be repositioned as 'costs/limitations' outside the UC taxonomy. As written, the central claim that the four-stage taxonomy characterizes UCs is not consistently supported by the data.","section":"§1, §4.5.3, §4.5.4"},{"comment":"Much of the evidence for the security/privacy influence level is phrased as anxiety, fear, or hypothetical risk rather than documented harm: 'Users expressed significant anxieties,' 'Fears of covert monitoring,' and numerous 'could'/'might' formulations (e.g., P5, P7, P140). The abstract and §5.3 claim a cascade from task failures through 'privacy violations and security vulnerabilities' to eroded trust, but the middle level is largely grounded in perceived risk and speculation. The paper should either re-label these sub-themes as 'perceived security and privacy risks' and adjust the cascade claim accordingly, or provide examples of actual, experienced violations. Without this, the RQ2 findings overstate the empirical grounding of the influence model.","section":"§5.2, Abstract"},{"comment":"The screening of the 1,850 Reddit posts (Section 3.1) and the subsequent categorization appear to have been performed without any reported inter-coder reliability check, and the interview sample (Section 3.2) is skewed toward male participants (12 of 14). For an initial exploratory taxonomy these choices are defensible, but the paper should explicitly acknowledge these constraints when presenting the four-stage taxonomy as a 'characterization of UCs' (Section 1) and should temper generalizability claims, especially for categories such as slow response, cost, and security concerns where self-selected complainants and retrospective recall may skew the data.","section":"§3.1, §3.2"}],"minor_comments":[{"comment":"The manuscript uses a clearly unfinished template: the author block says 'Trovato et al.' on the second page, the received/revised/accepted dates are '2007/2009', and the conference placeholder is 'Conference acronym ’XX, Woodstock, NY.' These must be corrected before any submission.","section":"Title page / template"},{"comment":"The phrase 'characterizes three UCs from three perspectives' is confusing; since the paper characterizes many UCs, the intended meaning is likely 'characterizes UCs from three perspectives: phenomena, influence, and mitigation.' Please rephrase.","section":"Abstract"},{"comment":"The phrase 'execute tasks flawedly' in the introduction should be 'execute tasks flawedly' → 'execute tasks in a flawed way' or 'flawed execution'.","section":"§1, 'flawedly'"},{"comment":"The statement that 'this paper distinguishes critical agent UCs from hallucination' is plausible but the operationalization of this distinction in the coding process is not described; please clarify how the analysts separated interaction-driven UCs from hallucination-driven factual errors.","section":"§7.1"},{"comment":"Several cells in Table 2 repeat the generic mitigation 'Enhancing agent capabilities' without specifying which capability addresses the phenomenon (e.g., 'Poor UI adaptability' vs. 'Element misidentification'); consider splitting or adding brief examples to make the mapping informative.","section":"Table 2"}],"recommendation":"major_revision","confidential_remarks":"The core qualitative contribution is timely and appropriate for a CSCW/HCI venue, and the authors have been transparent about their methods. The main substantive issues are the mismatch between the stated definition of 'unintended consequences' and several categories in the taxonomy, and the conflation of perceived security/privacy risks with experienced harms in the influence model. Both are fixable within the manuscript's scope. The lack of inter-coder reliability is not fatal for an exploratory study but should be explicitly disclosed. Also, the placeholder template (Trovato et al., 2007/2009 dates) must be cleaned before any formal submission; this is a presentation issue but would be embarrassing in review."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The one thing to know: this is a competent, honest piece of exploratory qualitative work that gives us a new four-stage lifecycle taxonomy (input, action, output, feedback) for the failures users hit with LLM-based GUI agents in web browsing. It is built from 221 Reddit posts and 14 interviews, and the synthesis is original. No prior work cited here does this for GUI agents specifically.\n\nWhat it does well: the quotes are well chosen and actually illustrate the categories; the authors are upfront about the exploratory nature and the limitations of their sample. The coding seems careful for what it is, and there is no circularity problem – the taxonomy is induced from the data, and the few self-citations are background only.\n\nThe soft spot is more internal. The paper defines UCs as 'unanticipated, negative outcomes,' but two of the four Feedback-stage categories – slow system response and high operational costs – are typically anticipated tradeoffs. Users often know before engaging that agents are slow or expensive. A user who chooses an agent despite known latency is not reporting an 'unanticipated' outcome. Similarly, much of the security and privacy 'influence' evidence in Section 5.2 is phrased as anxiety, fear, or hypothetical risk ('users expressed anxieties,' 'worried,' 'could'), not as documented harms. So the middle level of the claimed cascade – from task failures to security/privacy harms to eroded trust – is only partially grounded in experienced UCs. The taxonomy therefore reads less as a map of unintended consequences and more as a map of negative experiences, worries, and costs. That is still useful, but the authors should either broaden their definition or relabel and reposition those categories.\n\nSampling limitations are real but not fatal: self-selected Reddit posts and 14 interviewees (12 male, 2 female) with no inter-coder reliability. For an exploratory study, that is acceptable, and the authors acknowledge it.\n\nWho this is for: HCI and CSCW researchers working on agent failures, and GUI-agent developers looking for a structured list of user-reported issues with mitigation practices. It deserves serious peer review, with the expectation of a revision that reconciles the 'unintended' definition with the categories that are really about expected tradeoffs and perceived risks.","headline":"A genuinely useful exploratory taxonomy of GUI-agent failures, but the paper's own definition of 'unintended consequences' is stretched by cost and slowness categories and by speculative privacy worries.","tokens_in":26103,"tokens_out":1659,"would_cite":true,"duration_ms":17917,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Users of LLM-powered browser agents suffer failures that cluster into a four-stage lifecycle whose harms cascade from frustration to security breaches to lost trust.","keywords":["unintended consequences","GUI agents","web browsing","large language models","human-agent collaboration","qualitative study","user trust","agent failure taxonomy"],"falsifier":"Instrument real GUI-agent sessions end to end, recording the screen, the agent's action stream, and the user's interventions across a few hundred naturally occurring tasks, then compare the logged failures against the four-stage taxonomy and against the causes users attribute in follow-up interviews. If a substantial share of failures fits no lifecycle stage, or if user-attributed causes are routinely contradicted by the logs, the taxonomy's completeness and the reliability of retrospective accounts both weaken.","tokens_in":25115,"feed_emoji":"⚠️","tokens_out":9622,"duration_ms":85960,"temperature":0.7,"pith_summary":"The paper claims that the failures users meet when large language model (LLM)-based agents operate a graphical user interface (GUI) during web browsing are not scattered accidents but cluster into a four-stage lifecycle: input (misunderstanding instructions and poor planning), action (faulty clicks, misidentifying screen elements, weak adaptation to changing pages), output (inaccurate or off-target results), and feedback (no error recovery, slow responses, high token costs). Drawing on 221 screened Reddit posts and 14 interviews, it argues these failures cascade through three levels of harm: from task failure and frustration, to security and privacy violations, to eroded trust and broader ethical worries. The paper also catalogs what users already do about it, such as sandboxing the agent, rewriting prompts, and confirming sensitive steps, and turns the whole picture into design directions for future agents. If the taxonomy holds, agents going wrong stops being a vague worry and becomes a concrete, designable list of failure modes, each with user mitigations attached.","feed_headline":"Browser-agent failures follow a predictable four-stage pattern","feed_subtitle":"A taxonomy maps failures from misunderstood prompts to lost trust, giving agent designers concrete targets to fix.","key_machinery":"The load-bearing object is the four-stage lifecycle taxonomy of agent failures, covering input (comprehension and planning), action (execution and interaction), output (generation), and feedback (adjustment and viability), cross-mapped to a three-level influence cascade that runs from task and experience harms, through security and privacy harms, to trust and societal concerns. The taxonomy is the lens that turns scattered user complaints into ordered categories, and the argument is that each category is designable: every failure stage carries identifiable user mitigations and expectations, for example instruction-misinterpretation failures push users toward more precise prompting while security concerns push them toward sandboxing and permission limits. The empirical machinery underneath it is an exploratory mixed qualitative design of social media analysis followed by semi-structured interviews, analyzed with iterative reflexive thematic analysis, with the interviews serving to validate, deepen, and triangulate the post-derived categories.","core_discovery":"The central discovery, stated on the paper's own terms, is an initial characterization of unintended consequences that real users experience with LLM-based GUI agents during web browsing, built from a two-phase qualitative study of 221 Reddit posts and 14 semi-structured interviews. The characterization has three parts. Phenomenologically, failures occur across the agent's operational lifecycle in four stages: the input stage (complex setup, failed task decomposition, instruction misinterpretation, knowledge gaps), the action stage (faulty GUI actions, poor UI adaptability, element misidentification, external requirement conflicts, platform incompatibility), the output stage (inaccurate or semantically unsatisfactory results), and the feedback stage (poor error handling, difficult parameter and instruction tuning, slow responses, high operational costs). In terms of influence, these failures progress from direct task and experience harms (financial loss, frustration, increased workload, denial of service) to security and privacy harms (unauthorized access, data exposure, malicious exploitation, surveillance and social engineering) and finally to eroded trust and socio-ethical concerns. On mitigation, the paper reports that users already deploy system-oriented strategies (better error detection, permission limits, isolated environments) and user-oriented ones (precise prompting, direct oversight, confirmation for sensitive actions), and expect future agents to be more reliable, controllable, transparent, personalized, secure, and locally processed.","pith_inferences":["Editorial inference: the taxonomy doubles as a ready-made coding scheme, so a larger corpus of reviews, support tickets, or benchmark failure logs could be scored against the four stages to produce the prevalence estimates the paper explicitly does not claim.","Editorial inference: the cascade structure implies an untested ordering hypothesis, that hardening input comprehension and feedback recovery should reduce downstream privacy and trust harms more than patching outputs, which a controlled comparison of interventions could test.","Editorial inference: the paper's emphasis on user labor suggests a quantifiable design metric, an oversight burden measured as extra time, confirmations, and prompt revisions per completed task, that future agent evaluations could track.","Editorial inference: because the interviewees were mostly male and technically sophisticated (12 of 14), the relative severity of harms, especially privacy anxieties, may shift in a broader or less technical population; a replication with a more diverse sample would test how much of the emphasis is sample-driven."],"forward_implications":["Failure becomes characterizable before it is fixable: the four lifecycle stages give designers and researchers a checklist of where GUI agents break, from instruction comprehension through feedback processing.","Security and privacy harms are presented as downstream of interaction failures, not primarily of model hallucination, so interface design and control mechanisms count as legitimate security interventions.","The extensive user coping behavior the paper documents, including sandboxing, prompt rewriting, manual oversight, and confirmation gates, is read as evidence that current agents offload too much labor onto users rather than as a stable solution.","User expectations for reliability, controllability, transparency, personalization, security, and local processing become concrete requirements that future GUI-agent designs can be evaluated against.","The mitigation mapping ties specific phenomena to specific strategies, such as knowledge gaps to enhanced capabilities and external requirement conflicts to controlling the operational environment, yielding per-failure design targets instead of a general call for safer AI."],"supporting_citations":[{"why":"Together these supply the mixed-method template of social-media analysis plus semi-structured interviews that the study's design follows.","marker":"[2, 47]"},{"why":"Provide the reflexive thematic-analysis methodology used to code the integrated Reddit and interview dataset.","marker":"[9, 10]"},{"why":"Foundational study of unintended consequences arising from design and socio-technical systems that this paper extends to GUI agents.","marker":"[26]"},{"why":"The values-levers account of how overlooked design elements enable negative societal impact, cited as the precedent for studying unintended consequences.","marker":"[44]"},{"why":"Prior work on how users attribute responsibility for AI unintended consequences, the closest existing research this paper builds on and distinguishes itself from.","marker":"[16]"},{"why":"Provides the general AI-system threat taxonomy that the paper contrasts with its GUI-agent-specific failure taxonomy.","marker":"[18]"},{"why":"Survey of LLM security and privacy concerns used to show that existing risk categories do not cover GUI-agent interaction failures.","marker":"[46]"},{"why":"Real-world case study of a commercial GUI agent that motivates the focus on deployment failures beyond benchmark performance.","marker":"[35]"}],"fun_headline_variants":["Browser agents fail in four stages, from prompt to trust loss","User study maps AI browsing failures and fixes","Four-stage failure taxonomy for web-browsing AI agents","How users experience and counter AI browser agent errors","From misreads to misuse: taxonomy of browser agent harms"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The entire taxonomy rests on what users chose to report: 221 self-selected Reddit posts and 14 volunteers (12 male, 2 female) recalling past agent sessions from memory, with no direct observation, no screen logs, and no reliability check on the manual screening that kept 221 posts out of 1,850.","fun_headline_variants_meta":{"raw":{"variants":["Browser agents fail in four stages, from prompt to trust loss","User study maps AI browsing failures and fixes","Four-stage failure taxonomy for web-browsing AI agents","How users experience and counter AI browser agent errors","From misreads to misuse: taxonomy of browser agent harms"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000717,"raw_usage":{"total_tokens":3236,"prompt_tokens":974,"completion_tokens":2262,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":590,"completion_tokens_details":{"reasoning_tokens":2185}},"tokens_in":590,"tokens_out":2262,"duration_ms":15389,"temperature":1.0,"reasoning_tokens":2185,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T21:22:06.748708+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Instrument real GUI-agent sessions end to end, recording the screen, the agent's action stream, and the user's interventions across a few hundred naturally occurring tasks, then compare the logged failures against the four-stage taxonomy and against the causes users attribute in follow-up interviews. If a substantial share of failures fits no lifecycle stage, or if user-attributed causes are routinely contradicted by the logs, the taxonomy's completeness and the reliability of retrospective accounts both weaken.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"The values-levers account of how overlooked design elements enable negative societal impact, cited as the precedent for studying unintended consequences."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Survey of LLM security and privacy concerns used to show that existing risk categories do not cover GUI-agent interaction failures."}],"review_version":1}