REVIEW 3 major objections 5 minor 3 cited by
Characterizing Unintended Consequences in Human-GUI Agent Collaboration for Web Browsing
T0 review · 3 major / 5 minor · reviewed 2026-08-15 · deepseek-v4-flash
Pith's one-line read Users of LLM-powered browser agents suffer failures that cluster into a four-stage lifecycle whose harms cascade from frustration to security breaches to lost trust.
desk verdict A genuinely useful exploratory taxonomy of GUI-agent failures, but the paper's own definition of 'unintended consequences' is stretched by cost and slowness categories and by speculative privacy worries. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the four-stage lifecycle taxonomy of agent failures, covering input (comprehension and planning), action (execution and interaction), output (generation), and feedback (adjustment and viability), cross-mapped to a three-level influence cascade that runs from task and experience harms, through security and privacy harms, to trust and societal concerns. The taxonomy is the lens that turns scattered user complaints into ordered categories, and the argument is that each category is designable: every failure stage carries identifiable user mitigations and expectations, for example instruction-misinterpretation failures push users toward more precise prompting while security concerns push them toward sandboxing and permission limits. The empirical machinery underneath it is an exploratory mixed qualitative design of social media analysis followed by semi-structured interviews, analyzed with iterative reflexive thematic analysis, with the interviews serving to validate, deepen, and triangulate the post-derived categories.
What would settle it
Instrument real GUI-agent sessions end to end, recording the screen, the agent's action stream, and the user's interventions across a few hundred naturally occurring tasks, then compare the logged failures against the four-stage taxonomy and against the causes users attribute in follow-up interviews. If a substantial share of failures fits no lifecycle stage, or if user-attributed causes are routinely contradicted by the logs, the taxonomy's completeness and the reliability of retrospective accounts both weaken.
Extended reading notes
Core claim
The central discovery, stated on the paper's own terms, is an initial characterization of unintended consequences that real users experience with LLM-based GUI agents during web browsing, built from a two-phase qualitative study of 221 Reddit posts and 14 semi-structured interviews. The characterization has three parts. Phenomenologically, failures occur across the agent's operational lifecycle in four stages: the input stage (complex setup, failed task decomposition, instruction misinterpretation, knowledge gaps), the action stage (faulty GUI actions, poor UI adaptability, element misidentification, external requirement conflicts, platform incompatibility), the output stage (inaccurate or semantically unsatisfactory results), and the feedback stage (poor error handling, difficult parameter and instruction tuning, slow responses, high operational costs). In terms of influence, these failures progress from direct task and experience harms (financial loss, frustration, increased workload, denial of service) to security and privacy harms (unauthorized access, data exposure, malicious exploitation, surveillance and social engineering) and finally to eroded trust and socio-ethical concerns. On mitigation, the paper reports that users already deploy system-oriented strategies (better error detection, permission limits, isolated environments) and user-oriented ones (precise prompting, direct oversight, confirmation for sensitive actions), and expect future agents to be more reliable, controllable, transparent, personalized, secure, and locally processed.
Load-bearing premise
The entire taxonomy rests on what users chose to report: 221 self-selected Reddit posts and 14 volunteers (12 male, 2 female) recalling past agent sessions from memory, with no direct observation, no screen logs, and no reliability check on the manual screening that kept 221 posts out of 1,850.
Editorial extensions
If this is right
- Failure becomes characterizable before it is fixable: the four lifecycle stages give designers and researchers a checklist of where GUI agents break, from instruction comprehension through feedback processing.
- Security and privacy harms are presented as downstream of interaction failures, not primarily of model hallucination, so interface design and control mechanisms count as legitimate security interventions.
- The extensive user coping behavior the paper documents, including sandboxing, prompt rewriting, manual oversight, and confirmation gates, is read as evidence that current agents offload too much labor onto users rather than as a stable solution.
- User expectations for reliability, controllability, transparency, personalization, security, and local processing become concrete requirements that future GUI-agent designs can be evaluated against.
- The mitigation mapping ties specific phenomena to specific strategies, such as knowledge gaps to enhanced capabilities and external requirement conflicts to controlling the operational environment, yielding per-failure design targets instead of a general call for safer AI.
Reading between the lines
- Editorial inference: the taxonomy doubles as a ready-made coding scheme, so a larger corpus of reviews, support tickets, or benchmark failure logs could be scored against the four stages to produce the prevalence estimates the paper explicitly does not claim.
- Editorial inference: the cascade structure implies an untested ordering hypothesis, that hardening input comprehension and feedback recovery should reduce downstream privacy and trust harms more than patching outputs, which a controlled comparison of interventions could test.
- Editorial inference: the paper's emphasis on user labor suggests a quantifiable design metric, an oversight burden measured as extra time, confirmations, and prompt revisions per completed task, that future agent evaluations could track.
- Editorial inference: because the interviewees were mostly male and technically sophisticated (12 of 14), the relative severity of harms, especially privacy anxieties, may shift in a broader or less technical population; a replication with a more diverse sample would test how much of the emphasis is sample-driven.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper presents a qualitative, exploratory study of unintended consequences (UCs) in human collaboration with LLM-based GUI agents for web browsing. The authors analyzed 221 Reddit posts (after manually screening 1,850) and conducted 14 semi-structured interviews. They induce a four-stage lifecycle taxonomy of UC phenomena (Input: comprehension/planning; Action: execution/interaction; Output: generation; Feedback: adjustment/viability) and then relate these to three levels of influence: negative impacts on task execution and user experience, security/privacy risks, and broader erosion of trust and social/ethical concerns. They also synthesize user-initiated mitigations (system-oriented and user-oriented) and derive design implications for safer, more transparent GUI agents. The paper's contribution is framed as an initial characterization of UCs around GUI agents, grounded in users' retrospective self-reports.
Significance. The paper addresses a timely and under-studied topic: the real-world failure modes of LLM-driven GUI agents as experienced by users, rather than on benchmarks or curated tasks. Its strength is that the taxonomy is induced from original user data (Reddit posts and interviews), with many categories supported by direct quotes, and the methods are transparent, including an appendix with the interview protocol and ethical considerations. If the taxonomy holds, it offers a useful starting map for future HCI/CSCW research on agent failures, user coping strategies, and design implications. The paper does not ship machine-checked proofs or code, but it does provide reproducible qualitative data excerpts and a clear methodological description. The main risk is construct validity: the paper defines UCs as 'unanticipated, negative outcomes,' yet several prominent categories appear to be anticipated tradeoffs or hypothetical perceived risks rather than unanticipated harms. The influence cascade ('task failures and frustration → security/privacy harms → erosion of trust') is only partially grounded because much of the middle-level evidence is phrased as anxiety or speculation.
major comments (3)
- [§1, §4.5.3, §4.5.4] Section 1 defines UCs as 'unanticipated, negative outcomes in the interactive processes of these GUI agents.' However, the Feedback-stage categories 'Slow System Response' (§4.5.3) and 'High Operational Costs' (§4.5.4) are typically anticipated tradeoffs: users often know before engaging that agents may be slow or expensive. Quotes such as P70's 'too slow and expensive for any real-life scenarios' read as evaluations of known limitations, not reports of surprise. If these categories are retained as core UCs, the definition must be broadened (e.g., to 'negative outcomes and costs, whether anticipated or not'), or these categories should be repositioned as 'costs/limitations' outside the UC taxonomy. As written, the central claim that the four-stage taxonomy characterizes UCs is not consistently supported by the data.
- [§5.2, Abstract] Much of the evidence for the security/privacy influence level is phrased as anxiety, fear, or hypothetical risk rather than documented harm: 'Users expressed significant anxieties,' 'Fears of covert monitoring,' and numerous 'could'/'might' formulations (e.g., P5, P7, P140). The abstract and §5.3 claim a cascade from task failures through 'privacy violations and security vulnerabilities' to eroded trust, but the middle level is largely grounded in perceived risk and speculation. The paper should either re-label these sub-themes as 'perceived security and privacy risks' and adjust the cascade claim accordingly, or provide examples of actual, experienced violations. Without this, the RQ2 findings overstate the empirical grounding of the influence model.
- [§3.1, §3.2] The screening of the 1,850 Reddit posts (Section 3.1) and the subsequent categorization appear to have been performed without any reported inter-coder reliability check, and the interview sample (Section 3.2) is skewed toward male participants (12 of 14). For an initial exploratory taxonomy these choices are defensible, but the paper should explicitly acknowledge these constraints when presenting the four-stage taxonomy as a 'characterization of UCs' (Section 1) and should temper generalizability claims, especially for categories such as slow response, cost, and security concerns where self-selected complainants and retrospective recall may skew the data.
minor comments (5)
- [Title page / template] The manuscript uses a clearly unfinished template: the author block says 'Trovato et al.' on the second page, the received/revised/accepted dates are '2007/2009', and the conference placeholder is 'Conference acronym ’XX, Woodstock, NY.' These must be corrected before any submission.
- [Abstract] The phrase 'characterizes three UCs from three perspectives' is confusing; since the paper characterizes many UCs, the intended meaning is likely 'characterizes UCs from three perspectives: phenomena, influence, and mitigation.' Please rephrase.
- [§1, 'flawedly'] The phrase 'execute tasks flawedly' in the introduction should be 'execute tasks flawedly' → 'execute tasks in a flawed way' or 'flawed execution'.
- [§7.1] The statement that 'this paper distinguishes critical agent UCs from hallucination' is plausible but the operationalization of this distinction in the coding process is not described; please clarify how the analysts separated interaction-driven UCs from hallucination-driven factual errors.
- [Table 2] Several cells in Table 2 repeat the generic mitigation 'Enhancing agent capabilities' without specifying which capability addresses the phenomenon (e.g., 'Poor UI adaptability' vs. 'Element misidentification'); consider splitting or adding brief examples to make the mapping informative.
Circularity Check
No circularity: taxonomy is induced from original Reddit and interview data; the two self-citations are background only.
full rationale
The paper derives its taxonomy from original qualitative data: manual screening of 1,850 Reddit posts down to 221 posts, plus 14 semi-structured interviews, analyzed with reflexive thematic analysis (Braun and Clarke). There is no equation, fitted parameter, or benchmark through which a later claim could reduce to an earlier input. The central categories (input/action/output/feedback failures) are induced from participant quotes and incidents, so the taxonomy is not equivalent by construction to the paper's working definition of UCs as 'unanticipated, negative outcomes.' The skeptical concern that Slow System Response (§4.5.3) and High Operational Costs (§4.5.4) may be anticipated tradeoffs rather than unanticipated outcomes is a construct-validity critique of how well the data satisfy the paper's own definition; it is not a circular derivation, because those categories could in principle have failed to appear or could have been excluded, and the paper never defines 'slow' or 'expensive' as equivalent to 'unanticipated.' The only author self-citations are refs [78] and [79], cited in Related Work to support the existence of 'specific privacy protection systems'; these are not load-bearing for the empirical claims. The paper also explicitly acknowledges its methodological limits in §7.5, admitting that the study 'did not involve direct observations and controls of GUI agent use in lab settings' and was 'not designed to quantify prevalence or establish causal links.' That limitation reduces the strength of the claims but does not create circularity. Overall, the contribution is an empirical characterization rather than a derived prediction, and no load-bearing step reduces to its own input.
Assumptions & free parameters
assumptions (3)
- domain assumption Reddit posts and interview participants are treated as accurate, representative reporters of GUI agent experiences.
- domain assumption Manual screening of 1,850 Reddit posts and thematic coding are reliable without inter-coder agreement.
- domain assumption The definition of UCs as 'unanticipated, negative outcomes' is consistent with including slow operation and high cost as UCs.
Cite this review
Pith. "Pith review of Characterizing Unintended Consequences in Human-GUI Agent Collaboration for Web Browsing." pith.science (2026). https://pith.science/paper/DGEVPBXW
@misc{pith2026250509875,
author = {Pith},
title = {Pith review of: Characterizing Unintended Consequences in Human-GUI Agent Collaboration for Web Browsing},
year = {2026},
howpublished = {\url{https://pith.science/paper/DGEVPBXW}},
note = {Machine review of arXiv:2505.09875}
}
read the original abstract
The proliferation of Large Language Model (LLM)-based Graphical User Interface (GUI) agents in web browsing scenarios present complex unintended consequences (UCs). This paper characterizes three UCs from three perspectives: phenomena, influence and mitigation, drawing on social media analysis (N=221 posts) and semi-structured interviews (N=14). Key phenomenon for UCs include agents' deficiencies in comprehending instructions and planning tasks, challenges in executing accurate GUI interactions and adapting to dynamic interfaces, the generation of unreliable or misaligned outputs, and shortcomings in error handling and feedback processing. These phenomena manifest as influences from unanticipated actions and user frustration, to privacy violations and security vulnerabilities, and further to eroded trust and wider ethical concerns. Our analysis also identifies user-initiated mitigation, such as technical adjustments and manual oversight, and provides implications for designing future LLM-based GUI agents that are robust, user-centric, and transparent, fostering a crucial balance between automation and human oversight.
Figures
Forward citations
Cited by 3 Pith papers
-
Dark Patterns Meet GUI Agents: LLM Agent Susceptibility to Manipulative Interfaces and the Role of Human Oversight
GUI agents frequently fall for deceptive interface designs, often without recognizing them, and human supervision of agents improves avoidance only partially while introducing new attention and workload costs.
-
Towards Aligning Personalized Conversational Recommendation Agents with Users' Privacy Preferences
Privacy management for conversational AI agents is reframed as a dynamic alignment problem in which agents learn a user's latent privacy-utility reward function from feedback.
-
A Call for Collaborative Intelligence: Why Human-Agent Systems Should Precede AI Autonomy
A position paper arguing that LLM-based human-agent systems, not fully autonomous agents, should be the immediate goal for AI development.
Reference graph
Works this paper leans on
-
[1]
Omer Akgul, Richard Roberts, Moses Namara, Dave Levin, and Michelle L Mazurek. 2022. Investigating influencer VPN ads on YouTube. In 2022 IEEE Symposium on Security and Privacy (SP) . IEEE, 876–892
2022
-
[2]
Lize Alberts, Ulrik Lyngs, and Max Van Kleek. 2024. Computers as bad social actors: Dark patterns and anti-patterns in interfaces that act socially. Proceedings of the ACM on Human-Computer Interaction 8, CSCW1 (2024), 1–25
2024
-
[3]
James E Allen, Curry I Guinn, and Eric Horvtz. 1999. Mixed-initiative interaction. IEEE Intelligent Systems and their Applications 14, 5 (1999), 14–23
1999
-
[4]
Anthropic. 2025. Computer use (beta). https://docs.anthropic.com/en/docs/build-with-claude/computer-use/ Accessed: 2025-02-19
2025
-
[5]
Abdulgaffar O Arikewuyo, Kayode K Eluwole, and Bahire Özad. 2021. Influence of lack of trust on romantic relationship problems: The mediating role of partner cell phone snooping. Psychological Reports 124, 1 (2021), 348–365
2021
-
[6]
Eugene Bagdasarian, Ren Yi, Sahra Ghalebikesabi, Peter Kairouz, Marco Gruteser, Sewoong Oh, Borja Balle, and Daniel Ramage. 2024. AirGapAgent: Protecting privacy-conscious conversational agents. In Proceedings of the 2024 on ACM SIGSAC Conference on Computer and Communications Security. 3868–3882
2024
-
[7]
Michael Bailey, David Dittrich, Erin Kenneally, and Doug Maughan. 2012. The menlo report. IEEE Security & Privacy 10, 2 (2012), 71–75
2012
-
[8]
Tom L Beauchamp et al. 2008. The belmont report. The Oxford textbook of clinical research ethics (2008), 149–155
2008
Show all 106 references
-
[9]
Virginia Braun and Victoria Clarke. 2019. Reflecting on reflexive thematic analysis. Qualitative research in sport, exercise and health 11, 4 (2019), 589–597
2019
-
[10]
Virginia Braun and Victoria Clarke. 2024. Thematic analysis. In Encyclopedia of quality of life and well-being research . Springer, 7187–7193
2024
-
[11]
Jed R Brubaker, Casey Fiesler, Michael Madaio, John Tang, and Richmond Y Wong. 2024. Generative AI Going Awry: Enabling Designers to Proactively Avoid It in CSCW Applications. InCompanion Publication of the 2024 Conference on Computer-Supported Cooperative Work and Social Comp...
2024
-
[12]
John M Carroll. 2003. Making use: scenario-based design of human-computer interactions . MIT press
2003
-
[13]
Howard Carter. 2002. Misuse of library public access computers: Balancing privacy, accountability, and security. Journal of library administration 36, 4 (2002), 29–48. , Vol. 1, No. 1, Article . Publication date: September 2018. 24 • Trovato et al
2002
-
[14]
Chaoran Chen, Weijun Li, Wenxin Song, Yanfang Ye, Yaxing Yao, and Toby Jia-Jun Li. 2024. An empathy-based sandbox approach to bridge the privacy gap among attitudes, goals, knowledge, and behaviors. In Proceedings of the 2024 CHI Conference on Human Factors in Computing Systems. 1–28
2024
-
[15]
Chaoran Chen, Zhiping Zhang, Bingcan Guo, Shang Ma, Ibrahim Khalilov, Simret A Gebreegziabher, Yanfang Ye, Ziang Xiao, Yaxing Yao, Tianshi Li, et al. 2025. The Obvious Invisible Threat: LLM-Powered GUI Agents’ Vulnerability to Fine-Print Injections. arXiv preprint arXiv:2504.1...
2025 arXiv
-
[16]
Yu-Ting Chen, Hsin-Yi Sandy Tsai, and Chien Wen Yuan. 2024. Exploring How Users Attribute Responsibilities Across Different Stakeholders in Human-AI Interaction. In Companion Publication of the 2024 Conference on Computer-Supported Cooperative Work and Social Computing. 202–208
2024
-
[17]
Antonella De Angeli, Sheryl Brahnam, Peter Wallis, and Alan Dix. 2006. Misuse and abuse of interactive technologies. InCHI’06 Extended Abstracts on Human Factors in Computing Systems . 1647–1650
2006
-
[18]
Zehang Deng, Yongjian Guo, Changzhou Han, Wanlun Ma, Junwu Xiong, Sheng Wen, and Yang Xiang. 2025. Ai agents under threat: A survey of key security challenges and future pathways. Comput. Surveys 57, 7 (2025), 1–36
2025
-
[19]
Michael Dorin and Sergio Montenegro. 2021. Ethical lapses create complicated and problematic software. In 2021 IEEE/ACM 2nd International Workshop on Ethics in Software Engineering Research and Practice (SEthics) . IEEE, 1–4
2021
-
[20]
Salma Elsayed-Ali, Sara E Berger, Vagner Figueredo De Santana, and Juana Catalina Becerra Sandoval. 2023. Responsible & Inclusive Cards: An online card tool to promote critical reflection in technology industry work practices. In Proceedings of the 2023 CHI Conference on Human...
2023
-
[21]
Lia Emanuel, Joel Fischer, Wendy Ju, and Saiph Savage. 2016. Innovations in autonomous systems: Challenges and opportunities for human-agent collaboration. In Proceedings of the 19th ACM Conference on Computer Supported Cooperative Work and Social Computing Companion. 193–196
2016
-
[22]
Brett Eterovic-Soric, Kim-Kwang Raymond Choo, Helen Ashman, and Sameera Mubarak. 2017. Stalking the stalkers–detecting and deterring stalking behaviours using technology: A review. Computers & security 70 (2017), 278–289
2017
-
[23]
Don Fallis. 2015. What is disinformation? Library trends 63, 3 (2015), 401–426
2015
-
[24]
Matthias Fassl, Simon Anell, Sabine Houy, Martina Lindorfer, and Katharina Krombholz. 2022. Comparing User Perceptions of {Anti-Stalkerware} Apps with the Technical Reality. In Eighteenth Symposium on Usable Privacy and Security (SOUPS 2022) . 135–154
2022
-
[25]
Casey Fiesler, Jessica Pater, Janet Read, Jessica Vitak, and Michael Zimmer. 2023. Internet research ethics: A cscw community discussion. In Companion Publication of the 2023 Conference on Computer Supported Cooperative Work and Social Computing . 566–568
2023
-
[26]
Mary Flanagan. 2014. Making a Difference in and through Playful Design. In Proceedings of the 17th ACM conference on Computer supported cooperative work & social computing . 1–2
2014
-
[27]
Jessica R Frampton and Jesse Fox. 2021. Monitoring, creeping, or surveillance? A synthesis of online social information seeking concepts. Review of Communication Research 9 (2021), 1–42
2021
-
[28]
Kelsey R Fulton, Samantha Katcher, Kevin Song, Marshini Chetty, Michelle L Mazurek, Chloé Messdaghi, and Daniel Votipka. 2023. Vulnerability discovery for all: Experiences of marginalization in vulnerability discovery. In 2023 IEEE Symposium on Security and Privacy (SP). IEEE,...
2023
-
[29]
Andrea Gallardo, Hanseul Kim, Tianying Li, Lujo Bauer, and Lorrie Cranor. 2022. Detecting{iPhone} security compromise in simulated stalking scenarios: Strategies and obstacles. In Eighteenth Symposium on Usable Privacy and Security (SOUPS 2022) . 291–312
2022
-
[30]
Difei Gao, Lei Ji, Zechen Bai, Mingyu Ouyang, Peiran Li, Dongxing Mao, Qinchen Wu, Weichen Zhang, Peiyi Wang, Xiangwu Guo, et al
-
[31]
Irit Hadar, Tomer Hasson, Oshrat Ayalon, Eran Toch, Michael Birnhack, Sofia Sherman, and Arod Balissa. 2018. Privacy by designers: software developers’ privacy mindset. Empirical Software Engineering 23 (2018), 259–289
2018
-
[32]
Allyson I Hauptman, Wen Duan, and Nathan J Mcneese. 2022. The components of trust for collaborating with ai colleagues. InCompanion Publication of the 2022 Conference on Computer Supported Cooperative Work and Social Computing . 72–75
2022
-
[33]
Snooping
Skyler T Hawk, Andrik Becht, and Susan Branje. 2016. “Snooping” as a distinct parental monitoring strategy: Comparisons with overt solicitation and control. Journal of Research on Adolescence 26, 3 (2016), 443–458
2016
-
[34]
Ming-Tung Hong, Jesse Josua Benjamin, and Claudia Müller-Birn. 2018. Coordinating agents: Promoting shared situational awareness in collaborative sensemaking. In Companion of the 2018 ACM Conference on Computer Supported Cooperative Work and Social Computing . 217–220
2018
-
[35]
Siyuan Hu, Mingyu Ouyang, Difei Gao, and Mike Zheng Shou. 2024. The Dawn of GUI Agent: A Preliminary Case Study with Claude 3.5 Computer Use. arXiv preprint arXiv:2411.10323 (2024)
2024 arXiv
-
[36]
Nicolas Huaman, Sabrina Amft, Marten Oltrogge, Yasemin Acar, and Sascha Fahl. 2022. They would do better if they worked together: Interaction problems between password managers and the web. IEEE Security & Privacy 20, 2 (2022), 49–60
2022
-
[37]
Tian Huang, Chun Yu, Weinan Shi, Zijian Peng, David Yang, Weiqi Sun, and Yuanchun Shi. [n. d.]. Prompt2Task: Automating UI Tasks on Smartphones from Textual Prompts. ACM Transactions on Computer-Human Interaction ([n. d.]). , Vol. 1, No. 1, Article . Publication date: Septembe...
2018
-
[38]
Jun Li Jeung and Janet Yi-Ching Huang. 2023. Correct me if I am wrong: exploring how AI outputs affect user perception and trust. In Companion publication of the 2023 conference on computer supported cooperative work and social computing . 323–327
2023
-
[39]
Ziwei Ji, Nayeon Lee, Rita Frieske, Tiezheng Yu, Dan Su, Yan Xu, Etsuko Ishii, Ye Jin Bang, Andrea Madotto, and Pascale Fung. 2023. Survey of hallucination in natural language generation. Comput. Surveys 55, 12 (2023), 1–38
2023
-
[40]
Samuel Judson, Matthew Elacqua, Filip Cano, Timos Antonopoulos, Bettina Könighofer, Scott J Shapiro, and Ruzica Piskac. 2024. soid: A Tool for Legal Accountability for Automated Decision Making. In International Conference on Computer Aided Verification . Springer, 233–246
2024
-
[41]
Hyunggu Jung, Woosuk Seo, Seokwoo Song, and Sungmin Na. 2023. Toward value scenario generation through large language models. In Companion Publication of the 2023 Conference on Computer Supported Cooperative Work and Social Computing . 212–220
2023
-
[42]
Amy K Karlson, AJ Bernheim Brush, and Stuart Schechter. 2009. Can I borrow your phone? Understanding concerns when sharing mobile phones. In Proceedings of the SIGCHI Conference on Human Factors in Computing Systems . 1647–1650
2009
-
[43]
Jacek A Kopec and John M Esdaile. 1990. Bias in case-control studies. A review. Journal of epidemiology and community health 44, 3 (1990), 179
1990
-
[44]
Nicolas LaLone. 2014. Values levers and the unintended consequences of design. In Proceedings of the companion publication of the 17th ACM conference on Computer supported cooperative work & social computing . 189–192
2014
-
[45]
Youwei Li, Yangyang Li, and Yangzhao Yang. 2024. Test-Agent: A Multimodal App Automation Testing Framework Based on the Large Language Model. In 2024 IEEE 4th International Conference on Digital Twins and Parallel Intelligence (DTPI) . IEEE, 609–614
2024
-
[46]
Yuanchun Li, Hao Wen, Weijun Wang, Xiangyu Li, Yizhen Yuan, Guohong Liu, Jiacheng Liu, Wenxing Xu, Xiang Wang, Yi Sun, et al
-
[47]
Kai Lukoff, Ulrik Lyngs, Himanshu Zade, J Vera Liao, James Choi, Kaiyue Fan, Sean A Munson, and Alexis Hiniker. 2021. How the design of youtube influences user sense of agency. In Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems . 1–17
2021
-
[48]
arXiv preprint arXiv:2401.05459 (2024)
Personal llm agents: Insights and survey about the capability, efficiency and security. arXiv preprint arXiv:2401.05459 (2024)
2024 arXiv
-
[49]
Diogo Marques, Ildar Muslukhov, Tiago Guerreiro, Luís Carriço, and Konstantin Beznosov. 2016. Snooping on mobile phones: Prevalence and trends. In Twelfth Symposium on Usable Privacy and Security (SOUPS 2016) . 159–174
2016
-
[50]
Diogo Marques, Tiago Guerreiro, and Luis Carriço. 2014. Measuring snooping behavior with surveys: it’s how you ask it. In CHI’14 Extended Abstracts on Human Factors in Computing Systems . 2479–2484
2014
-
[51]
Phoebe Moh, Pubali Datta, Noel Warford, Adam Bates, Nathan Malkin, and Michelle L Mazurek. 2023. Characterizing everyday misuse of smart home devices. In 2023 IEEE Symposium on Security and Privacy (SP) . IEEE, 2835–2849
2023
-
[52]
Nora McDonald, Karla Badillo-Urquiola, Morgan G Ames, Nicola Dell, Elizabeth Keneski, Manya Sleeper, and Pamela J Wisniewski. 2020. Privacy and power: Acknowledging the importance of privacy research and design for vulnerable populations. In Extended Abstracts of the 2020 CHI ...
2020
-
[53]
Ildar Muslukhov, Yazan Boshmaf, Cynthia Kuo, Jonathan Lester, and Konstantin Beznosov. 2013. Know your enemy: the risk of unauthorized access in smartphones by insiders. In Proceedings of the 15th international conference on Human-computer interaction with mobile devices and s...
2013
-
[54]
Gustavo Moreira, Edyta Paulina Bogucka, Marios Constantinides, and Daniele Quercia. 2025. The Hall of AI Fears and Hopes: Comparing the Views of AI Influencers and those of Members of the US Public Through an Interactive Platform. In Proceedings of the 2025 CHI Conference on H...
2025
-
[55]
Dang Nguyen, Jian Chen, Yu Wang, Gang Wu, Namyong Park, Zhengmian Hu, Hanjia Lyu, Junda Wu, Ryan Aponte, Yu Xia, et al. 2024. Gui agents: A survey. arXiv preprint arXiv:2412.13501 (2024)
2024
-
[56]
Moses Namara, Daricia Wilkinson, Kelly Caine, and Bart P Knijnenburg. 2020. Emotional and practical considerations towards the adoption and abandonment of vpns as a privacy-enhancing technology. (2020)
2020
-
[57]
Virpi Oksman. 2010. The mobile phone-A medium in itself . VTT
2010
-
[58]
Michelle Oâ  2Reilly and Nikki Kiyimba. 2015. Advanced qualitative research: A guide to using theory. (2015)
2015
-
[59]
Michelle O’Reilly, Nikki Kiyimba, and Alison Drewett. 2021. Mixing qualitative methods versus methodologies: A critical reflection on communication and power in inpatient care. Counselling and psychotherapy research 21, 1 (2021), 66–76
2021
-
[60]
OpenAI. 2025. Introducing Operator. https://openai.com/index/introducing-operator/ Accessed: 2025-02-19
2025
-
[61]
Andreas Poller, Laura Kocksch, Sven Türpe, Felix Anand Epp, and Katharina Kinder-Kurlanda. 2017. Can security become a routine? A study of organizational change in an agile software development group. InProceedings of the 2017 ACM conference on computer supported cooperative w...
2017
-
[62]
Joon Sung Park, Joseph O’Brien, Carrie Jun Cai, Meredith Ringel Morris, Percy Liang, and Michael S Bernstein. 2023. Generative agents: Interactive simulacra of human behavior. In Proceedings of the 36th annual acm symposium on user interface software and technology . 1–22
2023
-
[63]
Harshini Sri Ramulu, Helen Schmitt, Dominik Wermke, and Yasemin Acar. 2024. Security and privacy software creators’ perspectives on unintended consequences. In 33rd USENIX Security Symposium (USENIX Security 24) . 3259–3276
2024
-
[64]
Chen Qian, Wei Liu, Hongzhang Liu, Nuo Chen, Yufan Dang, Jiahao Li, Cheng Yang, Weize Chen, Yusheng Su, Xin Cong, et al. 2024. Chatdev: Communicative agents for software development. InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Vol...
2024
-
[65]
Yangjun Ruan, Honghua Dong, Andrew Wang, Silviu Pitis, Yongchao Zhou, Jimmy Ba, Yann Dubois, Chris J Maddison, and Tatsunori Hashimoto. 2023. Identifying the risks of lm agents with an lm-emulated sandbox. arXiv preprint arXiv:2309.15817 (2023)
2023 arXiv
-
[66]
Karen Renaud, Graham Johnson, and Jacques Ophoff. 2020. Dyslexia and password usage: accessibility in authentication design. In Human Aspects of Information Security and Assurance: 14th IFIP WG 11.12 International Symposium, HAISA 2020, Mytilene, Lesbos, Greece, July 8–10, 202...
2020
-
[67]
Ben Shneiderman. 2002. Promoting universal usability with multi-layer interface design. ACM SIGCAPH computers and the physically handicapped 73-74 (2002), 1–8
2002
-
[68]
Yucheng Shi, Wenhao Yu, Wenlin Yao, Wenhu Chen, and Ninghao Liu. 2025. Towards trustworthy gui agents: A survey.arXiv preprint arXiv:2503.23434 (2025)
2025
-
[69]
Jonathan A Tran, Katie S Yang, Katie Davis, and Alexis Hiniker. 2019. Modeling the engagement-disengagement cycle of compulsive phone use. In Proceedings of the 2019 CHI conference on human factors in computing systems . 1–14
2019
-
[70]
Xiangru Tang, Qiao Jin, Kunlun Zhu, Tongxin Yuan, Yichi Zhang, Wangchunshu Zhou, Meng Qu, Yilun Zhao, Jian Tang, Zhuosheng Zhang, et al. 2024. Prioritizing safeguarding over autonomy: Risks of llm agents for science. arXiv preprint arXiv:2402.04247 (2024)
2024 arXiv
-
[71]
Shuai Wang, Weiwen Liu, Jingxuan Chen, Yuqi Zhou, Weinan Gan, Xingshan Zeng, Yuhan Che, Shuai Yu, Xinlong Hao, Kun Shao, et al
-
[72]
Wong, and Yaxing Yao
Jessica Vitak, Michael Zimmer, Anna Lenhart, Sunyup Park, Richmond Y. Wong, and Yaxing Yao. 2021. Designing for data awareness: addressing privacy and security concerns about “smart” technologies. In Companion Publication of the 2021 Conference on Computer Supported Cooperativ...
2021
-
[73]
Jingzhou Ye, Yao Li, Wenting Zou, and Xueqiang Wang. 2025. From Awareness to Action: The Effects of Experiential Learning on Educating Users about Dark Patterns. In Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems . 1–22
2025
-
[74]
arXiv preprint arXiv:2411.04890 (2024)
Gui agents with foundation models: A comprehensive survey. arXiv preprint arXiv:2411.04890 (2024)
2024 arXiv
-
[75]
David Wright and Paul De Hert. 2012. Privacy impact assessment. Vol. 6. Springer
2012
-
[76]
Mingyuan Zhang, Zhaolin Cheng, Sheung Ting Ramona Shiu, Jiacheng Liang, Cong Fang, Zhengtao Ma, Le Fang, and Stephen Jia Wang
-
[77]
Chaoyun Zhang, Shilin He, Jiaxu Qian, Bowen Li, Liqun Li, Si Qin, Yu Kang, Minghua Ma, Guyue Liu, Qingwei Lin, et al. 2024. Large language model-brained gui agents: A survey. arXiv preprint arXiv:2411.18279 (2024)
2024 arXiv
-
[78]
Chi Zhang, Zhao Yang, Jiaxuan Liu, Yanda Li, Yucheng Han, Xin Chen, Zebiao Huang, Bin Fu, and Gang Yu. 2025. Appagent: Multimodal agents as smartphone users. In Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems . 1–20
2025
-
[79]
Shuning Zhang, Xin Yi, Haobin Xing, Lyumanshan Ye, Yongquan Hu, and Hewu Li. 2024. Adanonymizer: Interactively Navigating and Balancing the Duality of Privacy and Output Performance in Human-LLM Interaction. arXiv preprint arXiv:2410.15044 (2024)
2024 arXiv
-
[80]
Zeyu Zhang, Xiaohe Bo, Chen Ma, Rui Li, Xu Chen, Quanyu Dai, Jieming Zhu, Zhenhua Dong, and Ji-Rong Wen. 2024. A survey on the memory mechanism of large language model based agents. arXiv preprint arXiv:2404.13501 (2024)
2024 arXiv
-
[81]
Muru Zhang, Ofir Press, William Merrill, Alisa Liu, and Noah A Smith. 2023. How language model hallucinations can snowball. arXiv preprint arXiv:2305.13534 (2023)
2023 arXiv
-
[82]
Ghost of the past
Shuning Zhang, Lyumanshan Ye, Xin Yi, Jingyu Tang, Bo Shui, Haobin Xing, Pengfei Liu, and Hewu Li. 2024. " Ghost of the past": identifying and resolving privacy leakage from LLM’s memory through proactive user interaction. arXiv preprint arXiv:2410.14931 (2024)
2024 arXiv
-
[85]
Kangjia Zhao, Jiahui Song, Leigang Sha, HaoZhan Shen, Zhi Chen, Tiancheng Zhao, Xiubo Liang, and Jianwei Yin. 2024. GUI Testing Arena: A Unified Benchmark for Advancing Autonomous GUI Testing Agent. arXiv preprint arXiv:2412.18426 (2024). A Ethical Considerations We acknowledg...
2024 arXiv
-
[86]
Could you describe your first experience using a GUI Agent, or your most memorable experience with a GUI Agent? What prompted you to start using GUI Agents?
-
[87]
In what situations do you typically use GUI Agents?
-
[88]
Could you describe a recent task you completed using a GUI Agent in detail?
-
[89]
Have you ever used GUI Agents for browser automation tasks? If so, could you describe one of your most recent experiences?
-
[90]
What different GUI Agents have you used? How would you compare them in terms of the following factors: functionality, usability, privacy and security concerns, and efficiency? C.2 Section 2: Unexpected Behaviors and Impacts
-
[91]
Have you ever encountered any unexpected behavior from a GUI Agent? If yes, could you describe the most memorable instance? (e.g., did the Agent perform tasks you did not request, or execute tasks in an unexpected way?)
-
[92]
Have you ever experienced a situation where a GUI Agent was unable to complete a task? If yes, could you describe the most memorable instance in detail?
-
[93]
Have you ever received incorrect or misleading information from a GUI Agent? If yes, could you provide some specific examples?
-
[94]
Have you encountered any privacy or security issues related to GUI Agents? (e.g., did an Agent access your personal information without authorization, or did its behavior result in data leakage?)
-
[95]
Which of these unexpected behaviors had the most significant impact on you? What were the consequences? (e.g., time loss, financial loss, privacy breaches, security risks, decreased user experience)
-
[96]
Did you later discover the reason for the unexpected behavior of the GUI Agent? Or do you still not know the cause? If you discovered it, what was the reason, and how did you find out? Was it due to configuration errors, limited model capabilities, insufficient context underst...
-
[97]
What measures have you taken to prevent or respond to these unexpected behaviors when using GUI Agents? (e.g., limiting Agent permissions, manually checking Agent operations, using virtual machines or containers to run Agents)
-
[98]
Do you think your coping strategies were effective? If yes, to what extent?
-
[99]
What problems do you still encounter after the strategies you adopted?
-
[100]
Do you know why these strategies were ineffective or why problems persisted? If not, do you have any guesses? , Vol. 1, No. 1, Article . Publication date: September 2018. 28 • Trovato et al. C.4 Section 4: User Expectations for Developers
2018
-
[101]
What functions or features do you think GUI Agents should offer to reduce unexpected behaviors?
-
[102]
How could they help you better control the Agent’s actions and prevent unexpected behaviors? C.5 Section 5: Future Outlook
-
[103]
What do you think is the future trend of GUI Agents?
-
[104]
How do you think GUI Agents will change our work and lifestyle?
-
[105]
What are your expectations for the future development of GUI Agents? What new functions would you like GUI Agents to achieve?
-
[106]
Is there anything else you would like to add? Received 20 February 2007; revised 12 March 2009; accepted 5 June 2009 , Vol. 1, No. 1, Article . Publication date: September 2018
2007
-
[2023]
In Companion Publication of the 2023 Conference on Computer Supported Cooperative Work and Social Computing
Towards Human-Centred AI-Co-Creation: A Three-Level Framework for Effective Collaboration between Human and AI. In Companion Publication of the 2023 Conference on Computer Supported Cooperative Work and Social Computing . 312–316
2023
-
[2024]
In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition
AssistGUI: Task-Oriented PC Graphical User Interface Automation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 13289–13298
Reviewed August 15, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.