REVIEW 4 major objections 5 minor 2 cited by
GUI agents are highly susceptible to dark patterns because they prioritize task completion over safety or privacy, and human oversight only partially mitigates the risk.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-04 17:36 UTC pith:BOWTHDPP
load-bearing objection Useful first map of GUI-agent susceptibility to dark patterns, but the human-oversight half rests on a video-review proxy and a numbers error; Phase 1 is the stronger contribution. the 4 major comments →
Dark Patterns Meet GUI Agents: LLM Agent Susceptibility to Manipulative Interfaces and the Role of Human Oversight
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The paper claims that GUI agents were highly susceptible to dark patterns primarily because they prioritize task completion over safety or privacy considerations. In a two-phase study covering 16 dark-pattern types, agents often avoided manipulation incidentally without explicitly recognizing it, and when they did recognize a manipulative design, they rarely took protective action if doing so required extra steps. Humans failed on overlapping patterns but for different reasons—cognitive shortcuts and habitual compliance—while human-agent teams improved avoidance in most tasks yet still succumbed to certain designs, and the oversight itself produced attentional tunneling, cognitive load, and
What carries the argument
The key machinery is a two-phase experimental design that separates awareness from avoidance: agents' reasoning traces (their stated justifications before acting) are used to measure whether they recognized a dark pattern, while predefined behavioral criteria measure whether they actually resisted it. This distinction lets the paper attribute failures to recognition gaps versus goal-driven prioritization. In Phase 2, a minimal 'watch mode' interface—users watched a pre-recorded agent video and signaled disagreement by pausing or skipping—served as the proxy for human oversight.
Load-bearing premise
Phase 2's oversight findings assume that pausing or skipping a pre-recorded video of an agent is a faithful proxy for supervising a live agent with real confirmation dialogs and the ability to take over.
What would settle it
Run the same 16-task study with a live agent under human supervision, with real confirmation prompts and takeover control, and compare avoidance rates and cognitive load to the video-review condition; if live oversight yields substantially different outcomes—for example, higher avoidance without the reported attention costs—the paper's oversight conclusions would not generalize.
If this is right
- GUI agent benchmarks should include safety-sensitive metrics (e.g., attack success rate, protected task completion) rather than raw task completion, because unsafe completions currently masquerade as competence.
- Training and alignment should model human caution signals—hesitation, checking fine print, deselecting defaults—and penalize unsafe shortcuts, since current objectives reward speed over safe completion.
- Oversight interfaces should integrate reasoning with actions (inline highlights, inspection traces, compact status timelines) to reduce attentional tunneling and cognitive load.
- Deployment of GUI agents in high-stakes domains such as finance, healthcare, or government is premature under current designs because errors cascade across action chains and accountability is unclear.
- Regulatory frameworks for dark patterns should extend beyond human deception to cover agent-mediated deception, with ex-ante assignment of liability.
Where Pith is reading between the lines
- The awareness-avoidance gap suggests that apparently safe agent behavior may be an artifact of simple tasks; on complex pages with multiple manipulative elements, incidental avoidance could collapse, making the 'illusion of safety' worse than it already appears.
- The video-review oversight proxy leaves open whether live oversight with real confirmation dialogs and takeover ability would be more effective or more burdensome; if the proxy fails, the reported oversight benefits and costs describe video auditing, not genuine human-agent collaboration.
- A testable prediction follows from the paper's own limitation discussion: agents trained with step-level risk rewards should show higher awareness-to-avoidance conversion on dark-pattern tasks, directly testing the claim that the gap stems from goal-driven optimization rather than capability.
- The attentional tunneling finding implies a new failure mode: oversight may make the human more manipulable via the agent's chosen path, so future work could measure whether overseers miss visually salient dark patterns outside the agent's action stream.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper presents a two-phase empirical study of how LLM-powered GUI agents, humans, and human-agent teams respond to 16 dark patterns. Phase 1 tests six agents (four Browser Use scaffolded LLMs and two end-to-end agents: Operator and Claude Computer Use) on 16 bespoke websites, coding awareness from reasoning traces and avoidance from predefined behavioral criteria. Phase 2 is a within-subjects study (N=22) comparing a human-only condition with a human-oversight condition in which participants watch pre-recorded video of Operator and pause or skip playback to signal approval or disagreement. The authors report that agents frequently avoid dark patterns without awareness, that humans and agents fail on similar patterns through different mechanisms, and that human oversight improves avoidance while introducing attentional tunneling, cognitive load, and reduced user control. The paper draws design implications for automation boundaries, safe-completion metrics, mixed-initiative handover, and regulation.
Significance. The Phase 1 descriptive findings are timely and useful: the paper operationalizes 16 dark patterns from Gray et al.'s ontology, distinguishes awareness from avoidance via reasoning traces, and documents a plausible awareness-avoidance gap and divergence between scaffolded and end-to-end agents. The task set and coding pipeline are reusable assets. However, the Phase 2 oversight condition uses pre-recorded video review rather than live supervision of an agent, and every RQ3 result depends on that proxy. In addition, the reported '14 of 16 tasks' improvement claim contradicts Table 3, and the Phase 2 comparisons lack inferential statistics. If the oversight findings were valid, they would strengthen the case that current watch modes are insufficient; as presented, the evidence is not yet sufficient to support the RQ3 conclusions. The contribution is conditional on reworking or properly delimiting the oversight study.
major comments (4)
- [Section 4.1.3 and Section 6] The human oversight condition (Section 4.1.3) uses pre-recorded Operator videos; participants pause, skip, or continue playback. This is not the watch mode of a deployed GUI agent, where confirmation requests are triggered by live actions and the user can take over or modify the trajectory. Consequently, the RQ3 results (Table 3, Figures 4-5) demonstrate video-review behavior, not human oversight of agents. The paper's central framing ('Human oversight improved avoidance...') overstates the evidence. The Limitations section (Section 6) does not acknowledge this proxy threat. This is a load-bearing validity threat: if the proxy fails, the oversight conclusions are unsupported.
- [Section 4.2.2 and Table 3] The sentence 'in 14 of 16 tasks, participants were less likely to fall for dark patterns when supervising the agent' is not supported by the paper's own data. Table 3 shows 10 tasks with higher avoidance under oversight, 4 ties (Adding Steps, Choice Overload, Urgency, Shaming), and 2 declines (Social Proof, Forced Registration). This numeric discrepancy should be corrected, and the absence of statistical testing on these rates should be addressed.
- [Section 4.2 and Section 4.1.4] Phase 2 comparisons rest on descriptive percentages without significance tests, effect sizes, or confidence intervals. With N=22 and each participant exposed to 8 of 16 patterns per condition, per-task cell sizes are roughly 11; differences such as Bad Defaults 33.3% vs 80% appear large but may not be statistically reliable. Claims about improved avoidance, attentional tunneling, and cognitive load need quantitative modeling (e.g., mixed-effects logistic regression or exact tests) before they can be taken as evidence.
- [Section 3.1.4 and Table 2] Phase 1 reports a single execution per agent-dark-pattern pair (Table 2 uses ✓/✗). There is no information about replication or run-to-run variance. The finding that agents 'often fail to recognize dark patterns' therefore describes a convenience sample of trajectories, not a stable property of agent classes. Please provide the number of runs or explicitly frame Table 2 as illustrative single-trial outcomes.
minor comments (5)
- [Abstract] Typo in the first sentence: 'Thedark patterns' should be 'Dark patterns'.
- [Section 2.3] The phrase 'a minimal watch mode' suggests a design that is then studied via video; the paper should be upfront about the video-review method from the outset, not only in Section 4.1.3.
- [Table 2] The '/ban' notation is unexplained in the caption; it should be expanded as 'agent halted for user confirmation' or similar.
- [Section 4.2.2] The claim of 'heightened cognitive load' appears to be inferred from qualitative participant statements; a standardized measure (e.g., NASA-TLX) or at least a clear operationalization would strengthen the claim.
- [References] The paper cites only arXiv v1 of [11] (fine-print injection); please check whether the published/peer-reviewed version should be cited instead.
Circularity Check
No circularity; empirical measurements are externally anchored and the central susceptibility/oversight claims rest on independent data.
full rationale
This is an empirical study, not a derivation, and the key constructs are anchored to external benchmarks. The 16 dark patterns and avoidance criteria are taken from Gray et al.'s taxonomy and translated into behavioral criteria in Table 1; agent susceptibility is measured against those external definitions, not against the paper's conclusions. Awareness is operationalized as explicit recognition in reasoning traces (Section 3.1.4), and the finding that awareness is low is a measurement result, not a tautology: the construct could have come out high, and the paper separately shows awareness sometimes occurs without avoidance. The Phase 2 oversight findings use a video-review proxy (Section 4.1.3), which is a validity threat to RQ3 but not circularity: pausing/skipping is the operational definition of disagreement in that condition, and the observed avoidance rates are empirical outcomes. Self-citations ([11], [12], [50]) appear in motivation and related work but are not load-bearing for the new measurements; no uniqueness theorem or ansatz is imported from them. The paper's own limitations (Section 6) acknowledge threats such as retrospective reasoning and static websites, but these do not reduce results to inputs. The only notable inconsistency—'14 of 16 tasks' in Section 4.2.2 vs. 10 improvements in Table 3—is a numerical/claims mismatch, not a circular step.
Axiom & Free-Parameter Ledger
free parameters (2)
- Avoidance criteria per dark pattern (16 hand-defined success thresholds)
- Awareness coding rule
axioms (5)
- domain assumption Gray et al.'s ontology is a valid basis for selecting and instantiating the 16 dark patterns
- domain assumption A single run per agent-task cell is representative of that agent's typical behavior
- domain assumption Reasoning traces (the added 'thinking' field) reflect the agent's actual decision process
- ad hoc to paper Pausing or skipping a pre-recorded video is equivalent to vetoing or approving a live agent action
- domain assumption Static single-pattern websites isolate dark pattern effects adequately for the stated conclusions
invented entities (1)
-
Unsafe success (concept)
independent evidence
read the original abstract
The dark patterns, deceptive interface designs manipulating user behaviors, have been extensively studied for their effects on human decision-making and autonomy. Yet, with the rising prominence of LLM-powered GUI agents that automate tasks from high-level intents, understanding how dark patterns affect agents is increasingly important. We present a two-phase empirical study examining how agents, human participants, and human-AI teams respond to 16 types of dark patterns across diverse scenarios. Phase 1 highlights that agents often fail to recognize dark patterns, and even when aware, prioritize task completion over protective action. Phase 2 revealed divergent failure modes: humans succumb due to cognitive shortcuts and habitual compliance, while agents falter from procedural blind spots. Human oversight improved avoidance but introduced costs such as attentional tunneling and cognitive load. Our findings show neither humans nor agents are uniformly resilient, and collaboration introduces new vulnerabilities, suggesting design needs for transparency, adjustable autonomy, and oversight.
Figures
Forward citations
Cited by 2 Pith papers
-
DPAgent-in-the-Middle: Agentic Defense and Repair Against AI-Groomed Deceptive Patterns
DPAgent is an agentic framework that detects 90.98% of AI-groomed deceptive samples and repairs 77% of deceptive interfaces while exploring 80% of pattern types with 10% of baseline page visits.
-
Comparing Human Oversight Strategies for Computer-Use Agents
Oversight strategy in computer-use agents shapes exposure to problematic actions more reliably than correction success, with plan-based approaches reducing occurrences but not uniformly improving interventions.
Reference graph
Works this paper leans on
-
[1]
http://web.archive.org/web/20220525230009/https://www.deceptive.design/types Archived version
2010.Deceptive Design - Types of Deceptive Design. http://web.archive.org/web/20220525230009/https://www.deceptive.design/types Archived version
arXiv 2010
-
[2]
Manus AI
2025. Manus AI. https://www.manusai.io/. [Accessed 08-09-2025]
2025
-
[3]
OpenAI launches Operator, an AI agent that performs tasks autonomously
2025. OpenAI launches Operator, an AI agent that performs tasks autonomously. https://techcrunch.com/2025/01/23/openai-launches-operator-an- ai-agent-that-performs-tasks-autonomously. 2025-01-23
2025
-
[4]
Jacob Aagaard, Miria Emma Clausen Knudsen, Per Bækgaard, and Kevin Doherty. 2022. A Game of Dark Patterns: Designing Healthy, Highly- Engaging Mobile Games. InExtended Abstracts of the 2022 CHI Conference on Human Factors in Computing Systems(New Orleans, LA, USA)(CHI EA ’22). Association for Computing Machinery, New York, NY, USA, Article 438, 8 pages. d...
arXiv 2022
-
[5]
2024.Computer Use (Beta)
Anthropic. 2024.Computer Use (Beta). https://docs.anthropic.com/en/docs/buildwith-claude/computer-use
2024
-
[7]
Kerstin Bongard-Blanchy, Arianna Rossi, Salvador Rivas, Sophie Doublet, Vincent Koenig, and Gabriele Lenzini. 2021. ”I am Definitely Manipulated, Even When I am Aware of it. It’s Ridiculous!” - Dark Patterns from the End-User Perspective. InProceedings of the 2021 ACM Designing Interactive Systems Conference(Virtual Event, USA)(DIS ’21). Association for C...
arXiv 2021
-
[8]
Christoph Bösch, Benjamin Erb, Frank Kargl, Henning Kopp, and Stefan Pfattheicher. 2016. Tales from the Dark Side: Privacy Dark Strategies and Privacy Dark Patterns.Proceedings on Privacy Enhancing Technologies2016 (2016), 237 – 254. doi:10.1515/popets-2016-0038
-
[9]
Evan Caragay, Katherine Xiong, Jonathan Zong, and Daniel Jackson. 2024. Beyond Dark Patterns: A Concept-Based Framework for Ethical Software Design. InProceedings of the 2024 CHI Conference on Human Factors in Computing Systems(Honolulu, HI, USA)(CHI ’24). Association for Computing Machinery, New York, NY, USA, Article 291, 16 pages. doi:10.1145/3613904.3642781
arXiv 2024
-
[10]
Chai, Lanbo She, Rui Fang, Spencer Ottarson, Cody Littley, Changsong Liu, and Kenneth Hanson
Joyce Y. Chai, Lanbo She, Rui Fang, Spencer Ottarson, Cody Littley, Changsong Liu, and Kenneth Hanson. 2014. Collaborative effort towards common ground in situated human-robot dialogue. InProceedings of the 2014 ACM/IEEE International Conference on Human-Robot Interaction (Bielefeld, Germany)(HRI ’14). Association for Computing Machinery, New York, NY, US...
arXiv 2014
-
[11]
Chaoran Chen, Zhiping Zhang, Bingcan Guo, Shang Ma, Ibrahim Khalilov, Simret Araya Gebreegziabher, Yanfang Ye, Ziang Xiao, Yaxing Yao, Tianshi Li, and Toby Jia-Jun Li. 2025. The Obvious Invisible Threat: LLM-Powered GUI Agents’ Vulnerability to Fine-Print Injections.ArXiv abs/2504.11281 (2025). doi:10.48550/arXiv.2504.11281
-
[12]
Chaoran Chen, Zhiping Zhang, Ibrahim Khalilov, Bingcan Guo, Simret A Gebreegziabher, Yanfang Ye, Ziang Xiao, Yaxing Yao, Tianshi Li, and Toby Jia-Jun Li. 2025. Toward a human-centered evaluation framework for trustworthy llm-powered gui agents.arXiv preprint arXiv:2504.17934(2025)
Pith/arXiv arXiv 2025
-
[14]
Yanda Chen, Joe Benton, Ansh Radhakrishnan, Jonathan Uesato, Carson Denison, John Schulman, Arushi Somani, Peter Hase, Misha Wagner, Fabien Roger, et al. 2025. Reasoning Models Don’t Always Say What They Think.arXiv preprint arXiv:2505.05410(2025). doi:10.48550/arXiv.2505.05410
-
[15]
Pengzhou Cheng, Zheng Wu, Zongru Wu, Aston Zhang, Zhuosheng Zhang, and Gongshen Liu. 2025. OS-Kairos: Adaptive Interaction for MLLM-Powered GUI Agents.ArXivabs/2503.16465 (2025). https://api.semanticscholar.org/CorpusID:277244134
Pith/arXiv arXiv 2025
-
[16]
Inyoung Cheong. 2025. Epistemic and Emotional Harms of Generative AI: Towards Human-Centered First Amendment. (2 Sept. 2025). https: //papers.ssrn.com/sol3/papers.cfm?abstract_id=5435335 Preprint posted on SSRN; 71 pages
2025
-
[17]
Nazli Cila. 2022. Designing Human-Agent Collaborations: Commitment, responsiveness, and support. InProceedings of the 2022 CHI Conference on Human Factors in Computing Systems(New Orleans, LA, USA)(CHI ’22). Association for Computing Machinery, New York, NY, USA, Article 420, Dark Patterns Meet GUI Agents 21 18 pages. doi:10.1145/3491102.3517500
arXiv 2022
-
[18]
Katherine M Collins, Catherine Wong, Jiahai Feng, Megan Wei, and Joshua B Tenenbaum. 2022. Structured, flexible, and robust: benchmarking and improving large language models towards more human-like behavior in out-of-distribution reasoning tasks.arXiv preprint arXiv:2205.05718(2022)
Pith/arXiv arXiv 2022
-
[20]
Linda Di Geronimo, Larissa Braz, Enrico Fregnan, Fabio Palomba, and Alberto Bacchelli. 2020. UI Dark Patterns and Where to Find Them: A Study on Mobile Applications and User Perception. InProceedings of the 2020 CHI Conference on Human Factors in Computing Systems(Honolulu, HI, USA) (CHI ’20). Association for Computing Machinery, New York, NY, USA, 1–14. ...
arXiv 2020
-
[21]
European Data Protection Board. 2022. Guidelines 3/2022 on Dark patterns in social media platform interfaces: How to recognise and avoid them. Public consultation document. https://www.edpb.europa.eu/our-work-tools/documents/public-consultations/2022/guidelines-32022-dark-patterns- social-media_en
2022
-
[22]
European Union. 2024. Artificial Intelligence Act: Article 14 - Human Oversight. https://artificialintelligenceact.eu/article/14/ Accessed: 2025-01-19
2024
-
[23]
Ivan Evtimov, Arman Zharmagambetov, Aaron Grattafiori, Chuan Guo, and Kamalika Chaudhuri. 2025. WASP: Benchmarking Web Agent Security Against Prompt Injection Attacks.ArXivabs/2504.18575 (2025). https://api.semanticscholar.org/CorpusID:278166059
Pith/arXiv arXiv 2025
-
[24]
Cedric Faas, Richard Bergs, Sarah Sterz, Markus Langer, and Anna Maria Feit. 2024. Give Me a Choice: The Consequences of Restricting Choices Through AI-Support for Perceived Autonomy, Motivational Variables, and Decision Performance. arXiv:2410.07728 [cs.HC] https: //arxiv.org/abs/2410.07728
Pith/arXiv arXiv 2024
-
[25]
K. J. Kevin Feng, David W. McDonald, and Amy X. Zhang. 2025. Levels of Autonomy for AI Agents. arXiv:2506.12469 [cs.HC] https://arxiv.org/abs/ 2506.12469
Pith/arXiv arXiv 2025
-
[26]
K. J. Kevin Feng, Kevin Pu, Matt Latzke, Tal August, Pao Siangliulue, Jonathan Bragg, Daniel S. Weld, Amy X. Zhang, and Joseph Chee Chang. 2025. Cocoa: Co-Planning and Co-Execution with AI Agents. arXiv:2412.10999 [cs.HC] https://arxiv.org/abs/2412.10999
arXiv 2025
-
[27]
Gray and Shruthi Sai Chivukula
Colin M. Gray and Shruthi Sai Chivukula. 2019. Ethical Mediation in UX Practice. InProceedings of the 2019 CHI Conference on Human Factors in Computing Systems(Glasgow, Scotland Uk)(CHI ’19). Association for Computing Machinery, New York, NY, USA, 1–11. doi:10.1145/3290605.3300408
arXiv 2019
-
[28]
Gray, Shruthi Sai Chivukula, Kassandra Melkey, and Rhea Manocha
Colin M. Gray, Shruthi Sai Chivukula, Kassandra Melkey, and Rhea Manocha. 2021. Understanding “Dark” Design Roles in Computing Education. In Proceedings of the 17th ACM Conference on International Computing Education Research(Virtual Event, USA)(ICER 2021). Association for Computing Machinery, New York, NY, USA, 225–238. doi:10.1145/3446871.3469754
arXiv 2021
-
[29]
Gray, Yubo Kou, Bryan Battles, Joseph Hoggatt, and Austin L
Colin M. Gray, Yubo Kou, Bryan Battles, Joseph Hoggatt, and Austin L. Toombs. 2018. The Dark (Patterns) Side of UX Design. InProceedings of the 2018 CHI Conference on Human Factors in Computing Systems(Montreal QC, Canada)(CHI ’18). Association for Computing Machinery, New York, NY, USA, 1–14. doi:10.1145/3173574.3174108
arXiv 2018
-
[30]
Gray, Lorena Sanchez Chamorro, Ike Obi, and Ja-Nae Duane
Colin M. Gray, Lorena Sanchez Chamorro, Ike Obi, and Ja-Nae Duane. 2023. Mapping the Landscape of Dark Patterns Scholarship: A Systematic Literature Review. InCompanion Publication of the 2023 ACM Designing Interactive Systems Conference(Pittsburgh, PA, USA)(DIS ’23 Companion). Association for Computing Machinery, New York, NY, USA, 188–193. doi:10.1145/3...
arXiv 2023
-
[31]
Gray, Cristiana Santos, Nataliia Bielova, and Thomas Mildner
Colin M. Gray, Cristiana Santos, Nataliia Bielova, and Thomas Mildner. 2023. An Ontology of Dark Patterns Knowledge: Foundations, Definitions, and a Pathway for Shared Knowledge-Building.Proceedings of the 2024 CHI Conference on Human Factors in Computing Systems(2023). https: //api.semanticscholar.org/CorpusID:262044555
2023
-
[32]
Gray, Cristiana Teixeira Santos, Nataliia Bielova, and Thomas Mildner
Colin M. Gray, Cristiana Teixeira Santos, Nataliia Bielova, and Thomas Mildner. 2024. An Ontology of Dark Patterns Knowledge: Foundations, Definitions, and a Pathway for Shared Knowledge-Building. InProceedings of the 2024 CHI Conference on Human Factors in Computing Systems (Honolulu, HI, USA)(CHI ’24). Association for Computing Machinery, New York, NY, ...
arXiv 2024
-
[33]
P. M. Groves and R. F. Thompson. 1970. Habituation: a dual-process theory.Psychological Review77 (1970), 419–450. Issue 5. doi:10.1037/h0029810
-
[34]
Chao Hao, Shuai Wang, and Kaiwen Zhou. 2025. Uncertainty-Aware GUI Agent: Adaptive Perception through Component Recommendation and Human-in-the-Loop Refinement. arXiv:2508.04025 [cs.AI] doi:10.48550/arXiv.2508.04025
-
[35]
Kasper Hornbæk, Per Ola Kristensson, and Antti Oulasvirta. 2025. 141Collaboration. InIntroduction to Human-Computer Interaction. Oxford University Press. arXiv:https://academic.oup.com/book/0/chapter/528999722/chapter-pdf/64021510/oso-9780192864543-chapter-8.pdf doi:10.1093/ oso/9780192864543.003.0008
arXiv 2025
-
[36]
Eric Horvitz. 1999. Principles of mixed-initiative user interfaces. InProceedings of the SIGCHI Conference on Human Factors in Computing Systems (Pittsburgh, Pennsylvania, USA)(CHI ’99). Association for Computing Machinery, New York, NY, USA, 159–166. doi:10.1145/302979.303030
arXiv 1999
-
[37]
Siyuan Hu, Mingyu Ouyang, Difei Gao, and Mike Zheng Shou. 2024. The Dawn of GUI Agent: A Preliminary Case Study with Claude 3.5 Computer Use. arXiv:2411.10323 [cs.AI] doi:10.48550/arXiv.2411.10323
-
[38]
Xu, Tianyue Ou, Shuyan Zhou, Jeffrey P
Faria Huq, Zora Zhiruo Wang, Frank F. Xu, Tianyue Ou, Shuyan Zhou, Jeffrey P. Bigham, and Graham Neubig. 2025. CowPilot: A Framework for Autonomous and Human-Agent Collaborative Web Navigation. InProceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (System D...
-
[39]
Jon M Jachimowicz, Shannon Duncan, Elke U Weber, and Eric J Johnson. 2019. When and why defaults influence decisions: A meta-analysis of default effects.Behavioural Public Policy3, 2 (2019), 159–186. 22 Tang et al
2019
-
[40]
Geunwoo Kim, Pierre Baldi, and Stephen McAleer. 2023. Language models can solve computer tasks.Advances in Neural Information Processing Systems36 (2023), 39648–39677
2023
-
[41]
W. J. Ladeira, W. M. Lim, F. d. O. Santini, T. Rasul, M. G. Perin, and L. Altınay. 2023. A meta-analysis on the effects of product scarcity.Psychology & Marketing40 (2023), 1267–1279. Issue 7. doi:10.1002/mar.21816
-
[42]
Vera Liao, Yunfeng Zhang, and Chenhao Tan
Vivian Lai, Samuel Carton, Rajat Bhatnagar, Q. Vera Liao, Yunfeng Zhang, and Chenhao Tan. 2022. Human-AI Collaboration via Conditional Delegation: A Case Study of Content Moderation. InProceedings of the 2022 CHI Conference on Human Factors in Computing Systems(New Orleans, LA, USA)(CHI ’22). Association for Computing Machinery, New York, NY, USA, Article...
arXiv 2022
-
[43]
Leiva, Yunfei Xue, Avya Bansal, Hamed R
Luis A. Leiva, Yunfei Xue, Avya Bansal, Hamed R. Tavakoli, Tuðçe Köroðlu, Jingzhou Du, Niraj R. Dayama, and Antti Oulasvirta. 2020. Understanding Visual Saliency in Mobile User Interfaces. In22nd International Conference on Human-Computer Interaction with Mobile Devices and Services (Oldenburg, Germany)(MobileHCI ’20). Association for Computing Machinery,...
arXiv 2020
-
[44]
Meng Li, Xiang Wang, Liming Nie, Chenglin Li, Yang Liu, Yangyang Zhao, Lei Xue, and Kabir Sulaiman Said. 2024. A Comprehensive Study on Dark Patterns.ArXivabs/2412.09147 (2024). https://api.semanticscholar.org/CorpusID:274656001
Pith/arXiv arXiv 2024
-
[45]
Zeyi Liao, Lingbo Mo, Chejian Xu, Mintong Kang, Jiawei Zhang, Chaowei Xiao, Yuan Tian, Bo Li, and Huan Sun. 2025. EIA: ENVIRONMEN- TAL INJECTION ATTACK ON GENERALIST WEB AGENTS FOR PRIVACY LEAKAGE. InThe Thirteenth International Conference on Learning Representations. https://openreview.net/forum?id=xMOLUzo2Lk
2025
-
[46]
Henry Lieberman. 1997. Autonomous interface agents. InProceedings of the ACM SIGCHI Conference on Human Factors in Computing Systems (Atlanta, Georgia, USA)(CHI ’97). Association for Computing Machinery, New York, NY, USA, 67–74. doi:10.1145/258549.258592
arXiv 1997
-
[47]
Chris Lu, Cong Lu, Robert Tjarko Lange, Jakob Foerster, Jeff Clune, and David Ha. 2024. The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery. arXiv:2408.06292 [cs.AI] https://arxiv.org/abs/2408.06292
Pith/arXiv arXiv 2024
-
[48]
Yijie Lu, Tianjie Ju, Manman Zhao, Xinbei Ma, Yuan Guo, and Zhuosheng Zhang. 2025. EVA: Red-Teaming GUI Agents via Evolving Indirect Prompt Injection.ArXivabs/2505.14289 (2025). https://arxiv.org/abs/2505.14289
Pith/arXiv arXiv 2025
-
[49]
Yadong Lu, Jianwei Yang, Yelong Shen, and Ahmed Awadallah. 2024. OmniParser for Pure Vision Based GUI Agent. arXiv:2408.00203 [cs.CV] https://arxiv.org/abs/2408.00203
Pith/arXiv arXiv 2024
-
[50]
Yuwen Lu, Chao Zhang, Yuewen Yang, Yaxing Yao, and Toby Jia-Jun Li. 2024. From Awareness to Action: Exploring End-User Empowerment Interventions for Dark Patterns in UX.Proc. ACM Hum.-Comput. Interact.8, CSCW1, Article 59 (April 2024), 41 pages. doi:10.1145/3637336
doi:10.1145/3637336 2024
-
[51]
Pattie Maes, Ben Shneiderman, and Jim Miller. 1997. Intelligent software agents vs. user-controlled direct manipulation: a debate. InCHI ’97 Extended Abstracts on Human Factors in Computing Systems(Atlanta, Georgia)(CHI EA ’97). Association for Computing Machinery, New York, NY, USA, 105–106. doi:10.1145/1120212.1120281
arXiv 1997
-
[52]
Friedman, Eli Lucherini, Jonathan Mayer, Marshini Chetty, and Arvind Narayanan
Arunesh Mathur, Gunes Acar, Michael J. Friedman, Eli Lucherini, Jonathan Mayer, Marshini Chetty, and Arvind Narayanan. 2019. Dark Patterns at Scale: Findings from a Crawl of 11K Shopping Websites.Proc. ACM Hum.-Comput. Interact.3, CSCW, Article 81 (Nov. 2019), 32 pages. doi:10.1145/3359183
doi:10.1145/3359183 2019
-
[53]
Arunesh Mathur, Mihir Kshirsagar, and Jonathan Mayer. 2021. What Makes a Dark Pattern... Dark? Design Attributes, Normative Considerations, and Measurement Methods. InProceedings of the 2021 CHI Conference on Human Factors in Computing Systems(Yokohama, Japan)(CHI ’21). Association for Computing Machinery, New York, NY, USA, Article 360, 18 pages. doi:10....
arXiv 2021
-
[54]
Arunesh Mathur, Arvind Narayanan, and Marshini Chetty. 2018. Endorsements on Social Media.Proceedings of the ACM on Human-Computer Interaction2 (2018), 1 – 26. https://api.semanticscholar.org/CorpusID:4317015
2018
-
[55]
Nora McDonald, Sarita Schoenebeck, and Andrea Forte. 2019. Reliability and Inter-rater Reliability in Qualitative Research: Norms and Guidelines for CSCW and HCI Practice.Proc. ACM Hum.-Comput. Interact.3, CSCW, Article 72 (Nov. 2019), 23 pages. doi:10.1145/3359174
doi:10.1145/3359174 2019
-
[57]
Woźniak, Rainer Malaka, and Jasmin Niess
Thomas Mildner, Daniel Fidel, Evropi Stefanidi, Paweł W. Woźniak, Rainer Malaka, and Jasmin Niess. 2025. A Comparative Study of How People With and Without ADHD Recognise and Avoid Dark Patterns on Social Media. InProceedings of the 2025 CHI Conference on Human Factors in Computing Systems (CHI ’25). Association for Computing Machinery, New York, NY, USA,...
arXiv 2025
-
[58]
Thomas Mildner, Merle Freye, Gian-Luca Savino, Philip R. Doyle, Benjamin R. Cowan, and Rainer Malaka. 2023. Defending Against the Dark Arts: Recognising Dark Patterns in Social Media. InProceedings of the 2023 ACM Designing Interactive Systems Conference(Pittsburgh, PA, USA)(DIS ’23). Association for Computing Machinery, New York, NY, USA, 2362–2374. doi:...
arXiv 2023
-
[59]
Thomas Mildner and Gian-Luca Savino. 2021. Ethical User Interfaces: Exploring the Effects of Dark Patterns on Facebook. InExtended Abstracts of the 2021 CHI Conference on Human Factors in Computing Systems(Yokohama, Japan)(CHI EA ’21). Association for Computing Machinery, New York, NY, USA, Article 464, 7 pages. doi:10.1145/3411763.3451659
arXiv 2021
-
[60]
Montgomery
Douglas C. Montgomery. 2017.Design and Analysis of Experiments(9th ed.). John Wiley & Sons, Hoboken, NJ
2017
-
[61]
Hussein Mozannar, Gagan Bansal, Cheng Tan, Adam Fourney, Victor Dibia, Jingya Chen, Jack Gerrits, Tyler Payne, Matheus Kunzler Maldaner, Madeleine Grunde-McLaughlin, Eric Zhu, Griffin Bassman, Jacob Alber, Peter Chang, Ricky Loynd, Friederike Niedtner, Ece Kamar, Maya Murad, Rafah Hosn, and Saleema Amershi. 2025. Magentic-UI: Towards Human-in-the-loop Age...
-
[62]
2024.Browser Use: Enable AI to control your browser
Magnus Müller and Gregor Žunič. 2024.Browser Use: Enable AI to control your browser. https://github.com/browser-use/browser-use
2024
-
[63]
Ahmed, Puneet Mathur, Seunghyun Yoon, Lina Yao, Branislav Kveton, Jihyung Kil, Thien Huu Nguyen, Trung Bui, Tianyi Zhou, Ryan A
Dang Nguyen, Jian Chen, Yu Wang, Gang Wu, Namyong Park, Zhengmian Hu, Hanjia Lyu, Junda Wu, Ryan Aponte, Yu Xia, Xintong Li, Jing Shi, Hongjie Chen, Viet Dac Lai, Zhouhang Xie, Sungchul Kim, Ruiyi Zhang, Tong Yu, Mehrab Tanjim, Nesreen K. Ahmed, Puneet Mathur, Seunghyun Yoon, Lina Yao, Branislav Kveton, Jihyung Kil, Thien Huu Nguyen, Trung Bui, Tianyi Zho...
2025
-
[64]
Qian Niu, Junyu Liu, Ziqian Bi, Pohsun Feng, Benji Peng, Keyu Chen, Ming Li, Lawrence KQ Yan, Yichao Zhang, Caitlyn Heqi Yin, et al. 2024. Large language models and cognitive science: A comprehensive review of similarities, differences, and challenges.arXiv preprint arXiv:2409.02387(2024)
arXiv 2024
-
[65]
Jantawan Noiwan and Anthony F. Norcio. 2006. Cultural differences on attention and perceived usability: Investigating color combinations of animated graphics.International Journal of Human-Computer Studies64, 2 (2006), 103–122. doi:10.1016/j.ijhcs.2005.06.004
-
[66]
OECD. 2022.Dark Commercial Patterns. OECD Digital Economy Papers, No. 336. OECD Publishing, Paris. doi:10.1787/44f5e846-en
-
[67]
2025.Introducing Operator-Safety and Privacy
OpenAI. 2025.Introducing Operator-Safety and Privacy. https://openai.com/index/introducing-operator/
2025
-
[68]
R. Parasuraman, T.B. Sheridan, and C.D. Wickens. 2000. A model for types and levels of human interaction with automation.IEEE Transactions on Systems, Man, and Cybernetics - Part A: Systems and Humans30, 3 (2000), 286–297. doi:10.1109/3468.844354
arXiv 2000
-
[69]
Park, Simon Goldstein, Aidan O’Gara, Michael Chen, and Dan Hendrycks
Peter S. Park, Simon Goldstein, Aidan O’Gara, Michael Chen, and Dan Hendrycks. 2024. AI deception: A survey of examples, risks, and potential solutions.Patterns5, 5 (2024), 100988. doi:10.1016/j.patter.2024.100988
arXiv 2024
-
[70]
Yi-Hao Peng, Dingzeyu Li, Jeffrey P. Bigham, and Amy Pavel. 2025. Morae: Proactively Pausing UI Agents for User Choices. arXiv:2508.21456 [cs.HC]
Pith/arXiv arXiv 2025
-
[71]
Denise F. Polit and Cheryl Tatano Beck. 2009. Qualitative research and content validity: developing best practices based on science and experience. Quality of Life Research18, 9 (2009), 1263–1278. doi:10.1007/s11136-009-9540-9
-
[72]
Kevin Pu, Daniel Lazaro, Ian Arawjo, Haijun Xia, Ziang Xiao, Tovi Grossman, and Yan Chen. 2025. Assistance or Disruption? Exploring and Evaluating the Design and Trade-offs of Proactive AI Programming Support. InProceedings of the 2025 CHI Conference on Human Factors in Computing Systems (CHI ’25). Association for Computing Machinery, New York, NY, USA, A...
arXiv 2025
-
[73]
Clayton D. Rothwell, Valerie L. Shalin, and Griffin D. Romigh. 2021. Comparison of Common Ground Models for Human–Computer Dialogue: Evidence for Audience Design.ACM Trans. Comput.-Hum. Interact.28, 2, Article 9 (April 2021), 35 pages. doi:10.1145/3410876
-
[74]
Vildan Salikutluk, Janik Schöpper, Franziska Herbert, Katrin Scheuermann, Eric Frodl, Dirk Balfanz, Frank Jäkel, and Dorothea Koert. 2024. An Evaluation of Situational Autonomy for Human-AI Collaboration in a Shared Workspace Setting. InProceedings of the 2024 CHI Conference on Human Factors in Computing Systems(Honolulu, HI, USA)(CHI ’24). Association fo...
arXiv 2024
-
[75]
Shelle Santana, Steven K. Dallas, and Vicki G. Morwitz. 2020. Consumer reactions to drip pricing.Marketing Science39, 1 (Jan 2020), 188–210. doi:10.1287/mksc.2019.1207
arXiv 2020
-
[76]
René Schäfer, Paul Miles Preuschoff, René Röpke, Sarah Sahabi, and Jan Borchers. 2024. Fighting Malicious Designs: Towards Visual Countermeasures Against Dark Patterns. InProceedings of the 2024 CHI Conference on Human Factors in Computing Systems(Honolulu, HI, USA)(CHI ’24). Association for Computing Machinery, New York, NY, USA, Article 296, 13 pages. d...
arXiv 2024
- [77]
-
[78]
Jonathan Shaki, Sarit Kraus, and Michael Wooldridge. 2023. Cognitive effects in large language models. InECAI 2023. IOS Press, 2105–2112. doi:10.3233/FAIA230505
-
[79]
Yijia Shao, Tianshi Li, Weiyan Shi, Yanchen Liu, and Diyi Yang. 2024. Privacylens: Evaluating privacy norm awareness of language models in action. Advances in Neural Information Processing Systems37 (2024), 89373–89407
2024
-
[80]
Yijia Shao, Vinay Samuel, Yucheng Jiang, John Yang, and Diyi Yang. 2025. Collaborative Gym: A Framework for Enabling and Evaluating Human-Agent Collaboration. arXiv:2412.15701 [cs.AI] https://arxiv.org/abs/2412.15701
arXiv 2025
-
[81]
Yucheng Shi, Wenhao Yu, Zaitang Li, Yonglin Wang, Hongming Zhang, Ninghao Liu, Haitao Mi, and Dong Yu. 2025. MobileGUI-RL: Advancing Mobile GUI Agent through Reinforcement Learning in Online Environment.ArXivabs/2507.05720 (2025). doi:10.48550/arXiv.2507.05720
-
[82]
Kristina Suchotzki and Matthias Gamer. 2024. Detecting deception with artificial intelligence: promises and perils.Trends in Cognitive Sciences28, 6 (2024), 481–483. doi:10.1016/j.tics.2024.04.002
-
[83]
Zhongxiang Sun, Qipeng Wang, Haoyu Wang, Xiao Zhang, and Jun Xu. 2025. Detection and Mitigation of Hallucination in Large Reasoning Models: A Mechanistic Perspective.arXiv preprint arXiv:2505.12886(2025). doi:10.48550/arXiv.2505.12886
-
[84]
Siddharth Suresh, Kushin Mukherjee, Xizheng Yu, Wei-Chun Huang, Lisa Padua, and Timothy Rogers. 2023. Conceptual structure coheres in human cognition but not in large language models. InProceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, Houda Bouamor, Juan Pino, and Kalika Bali (Eds.). Association for Computational Lin...
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.