REVIEW 3 major objections 4 minor 102 references
Plover: Steering GUI Agents through Plan-Centric Interaction
T0 review · 3 major / 4 minor · reviewed 2026-08-01 · deepseek-v4-flash
Pith's one-line read Plover shows that many GUI-agent failures become repairable when the task plan stays visible and corrections stay localized, with 23 of 26 benchmark failures improved by mixed-initiative interaction.
desk verdict Plover is a credible systems paper with an honest upper-bound recovery result; just don't let the abstract sell the 88% as proof that plan visibility is what rescues failures. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central mechanism is the persistent plan artifact, a shared state representation with the invariant that executed steps are immutable and only the pending suffix can be revised (C_{t+1}=C_t). On top of this, Plover implements Intelligent Replanning in two modes: User-Driven IR, where natural-language guidance or multimodal annotations (strokes, shapes, text overlays captured as primitives with a bounding box) generate localized plan proposals, and System-Driven IR, where a watchdog detects behavioral loop repetition and visual non-progress (via dHash Hamming distance) and injects a structured failure message that prompts the model to propose a recovery step with rationale. The versioned
What would settle it
Run the same 26 benchmark failures with naive participants who are not told the failure cause, providing only the Plover interface as the intervention channel. If the mixed-initiative recovery rate falls to near the autonomous baseline (no significant improvement over the 0% success on these originally failed tasks), the claim that failures are structurally repairable through visible plans is not supported for realistic users.
Extended reading notes
Core claim
Plover's core claim is that GUI-agent failures are structurally recoverable when plans are externalized and repairs are localized. The system keeps a versioned plan artifact, separates immutable executed steps from an editable pending suffix, and supports user-driven interventions (natural-language guidance, plan edits, and screenshot annotations) plus system-driven replanning triggered by non-progress detection. In the benchmark repair analysis, 26 autonomous failures were re-run in a mixed-initiative setting; 23 improved (17 complete successes, 6 partial), only 3 remained failures, and all 10 autonomous partial successes became complete successes. The paper also characterizes which failure
Load-bearing premise
The 88% recovery figure depends on an expert user who already knows what went wrong and what the correct target is; if ordinary users cannot detect drift or formulate correct localized corrections, the recoverability claim may not transfer to practice.
Editorial extensions
If this is right
- If recoverability holds beyond the expert setting, GUI agents can be deployed in long-horizon, high-friction workflows with a human steering loop instead of requiring near-perfect autonomy.
- The plan-invariant design means corrections preserve executed history, so each intervention is cheaper and less disruptive than re-prompting or restarting the whole task.
- System-driven non-progress detection (repeated semantic actions plus visual stability) can act as a reusable watchdog that catches drift before it propagates, independent of the specific planner or executor.
- The failure taxonomy suggests that perception errors and state misinterpretations are cheaply repairable with language or annotations, while compound failures require catching the initial planning error earlier.
- Exposing plans as versioned, diffable artifacts provides a natural audit trail for when and why an agent's behavior changed, which can support verification and post-hoc analysis.
Reading between the lines
- A natural next test, which the paper does not run, is a study with non-expert users who are not told the failure cause: if recovery rates drop to near the autonomous baseline, the 'structurally recoverable' claim would need to be re-scoped from an upper bound to a property that depends on user diagnostic skill.
- The plan artifact as a coordination protocol could generalize beyond GUI automation to other long-horizon agent domains (e.g., data-cleaning pipelines or robotics task plans) where partial progress is valuable and corrections must be localized.
- The System-Driven IR watchdog (behavioral repetition + perceptual-hash stability) is a concrete, model-agnostic component that could be extracted and benchmarked on its own to measure how many agent stalls it catches before a human would notice.
- The paper itself flags that visible plans may inflate user confidence; an empirical study measuring whether users over-accept plan proposals when the system looks confident would be a direct extension.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces Plover, a plan-centric GUI automation system that externalizes task plans as persistent, editable artifacts and supports mixed-initiative repair through natural-language guidance, multimodal annotation, plan edits, and system-driven replanning. The authors report a formative study with six participants, a benchmark repair study on 38 OSWorld-Verified tasks (26 autonomous non-successes re-run with expert interventions, yielding 23 improved, 17 complete successes, 6 partial successes, 3 failures), and a scenario-based stability analysis with trajectory-derived prompts. The central claim is that many GUI-agent failures are structurally repairable when plans remain visible and interventions are localized, and that explicit replanning improves transparency, controllability, and adaptability.
Significance. If the central claim were fully established, the paper would make a useful contribution to human-agent interaction for GUI automation: it proposes a concrete design space (persistent plans as coordination artifacts, localized repair, visible replanning) and provides a failure taxonomy that could guide future interface design. The paper is honest in labeling the benchmark repair study as an upper bound on plan-centric recoverability, and the appendices contain substantial implementation and evaluation detail, including per-task results and a formative study summary. These are real strengths. However, the main empirical evidence does not currently separate the effect of the plan-centric interface from the effect of an expert oracle user, so the causal design conclusions (DG1–DG5, 'explicit replanning helps') are not yet established by the data.
major comments (3)
- [Section 5.1, Table 1; Abstract] The load-bearing empirical claim—23/26 non-success cases improved, 88% recovery—is measured with the first author, who knows each failure cause and the correct target, supplying all interventions. There is no control condition that strips away the plan-centric affordances (plan panel, plan editing, annotation, visible replanning) while keeping the same underlying agent and the same expert. As written, the 88% figure is an upper bound on recoverability by an informed oracle, not evidence that plan visibility and localized repair cause the improvement. The paper's own limitation statement in Section 6 ('our design assumes users generally know the intended path well enough to correct the agent') does not carry through to the abstract/conclusion, which assert a causal role for visible plans. Add an ablation or control (e.g., same expert corrections issued as text prompts to the same base age
- [Section 5.2, 'Trajectory Sampling and Prompt Reconstruction' and Table 2] The scenario stability analysis synthesizes task prompts from sampled interaction trajectories and then compares the replayed plans and final states against the same trajectories. This creates a circularity: the prompt is derived from the reference trajectory, so plan alignment metrics (coverage 0.62, order 0.41, actionability 0.97) partly measure reconstruction from a trajectory-derived instruction rather than general plan quality. The browser-vs-desktop visual fidelity differences are still informative, but the plan alignment results should be presented as a property of this reverse-synthesis setup, not as evidence about Plover's planning under independently authored user instructions. A sanity check with manually authored prompts, or a clear caveat, is needed before these numbers are used to support the 'opportunity for users to inspect and correct' argument.
- [Section 5.3, 'Characterizing Repairable Failures'] The failure-mode analysis is based on the same 26 cases and the same expert interventions. Counts such as 'Execution Drift appeared in 46% (n=12/26)' and 'NL Guidance resolved 11 cases' are reported without any uncertainty or sensitivity analysis. With n=26 and intervention choices made by a single expert who already knows the failure causes, small counts can easily flip; the recovered vs. unrecovered distinction is not robust enough to support the strong claim that 'compound failures' are fundamentally harder. At minimum, report bootstrap or exact binomial confidence intervals and clarify that all recovery counts are conditional on the expert's choice of intervention.
minor comments (4)
- [Appendix C, Algorithm 1] System-Driven IR relies on hardcoded thresholds (REPEAT_SEQ_L3_R3, dHash Hamming distance > 40) with no sensitivity analysis. Since Section 5.3 attributes 9 successful recoveries to System-Driven IR, the threshold choices can materially affect the results; report how varying them changes detection and downstream recovery.
- [Section 5.1, 'no regressions'] The statement 'no regressions were observed' only covers the 26 autonomous non-success cases; it does not address whether the mixed-initiative interaction could degrade autonomous successes, since those were not re-run in the mixed-initiative condition. Please state this scope explicitly.
- [Table 1 and Table 2] Table 1's 'Improv. Rate' is not formally defined; for Multi-App, 6S+2P out of 10 corresponds to 80%, but the reader must infer the denominator. In Table 2, the 'Overall Average' row for MSE (939.57) is the mean of scenario averages, not the mean over all trials; clarify the aggregation.
- [Throughout] The paper uses 'Conference’17' in the ACM reference format and several placeholder-style citations (e.g., the DOI is 'XXXXXXX.XXXXXXX'). Please update the formatting to the final venue style and correct minor typographical issues such as the 'MI (a)' label in Figure 5.
Circularity Check
No significant circularity: the 88% recoverability result is an externally grounded upper-bound measurement with the expert-oracle caveat explicitly acknowledged; remaining concerns are validity limitations, not circular reductions.
full rationale
Plover's central empirical claim is an upper-bound measurement on the external OSWorld-Verified benchmark, not a quantity derived from a fitted parameter, an ansatz, or a load-bearing self-citation. The paper states the intervention condition explicitly: 'This setup establishes an upper bound on plan-centric recoverability rather than typical user performance' (Section 5.1), and the corresponding limitation is restated in Section 6 ('our design assumes users generally know the intended path well enough to correct the agent when needed'). Because the expert interventions are acknowledged as ideal rather than typical, the 88% recovery figure is an honest conditional result, not a hidden assumption presented as a finding. The scenario-based stability analysis (Section 5.2) synthesizes prompts from reference trajectories and then compares generated plans against those trajectories, which introduces non-independence; however the reported alignment metrics are moderate (coverage 0.62, order 0.41), so the result is not forced by construction. The only self-citation is a related-work mention ([10], multimodal interaction) and is not load-bearing. No uniqueness theorem, no ansatz smuggled via citation, and no equation reduces the conclusion to its input. The absence of a chat-only or invisible-plan control is a causal-identification limitation, but under the circularity criteria it does not constitute circularity.
Assumptions & free parameters
free parameters (4)
- Visual non-progress dHash threshold τ =
40 (Hamming distance)
- Repetition pattern REPEAT_SEQ_L3_R3 =
repeated action subsequence of length 3 seen 3 times
- Per-scenario similarity thresholds (SSIM/MSE/dHash) =
e.g., Firefox Fillable Form High: SSIM≥0.98, MSE≤100, dHash≤2; LibreOffice thresholds differ
- Selected trajectories per scenario =
5 per scenario from 100 exploration trials, after manual verification
assumptions (6)
- domain assumption Claude 4.5 Sonnet via computer-use can decompose tasks into deterministic UI instructions and ground them in the executor.
- domain assumption Image-similarity metrics (SSIM, MSE, dHash) are valid proxies for GUI execution stability and task-state equivalence.
- domain assumption The 38 OSWorld-Verified tasks that failed for Claude 4.5 Sonnet are representative of GUI-agent failures in general.
- domain assumption An expert who knows failure causes can stand in for a typical user to establish 'structural recoverability'.
- domain assumption GPA phase taxonomy and trajectory-derived phases are a meaningful ground truth for plan alignment.
- domain assumption The invariant C_{t+1}=C_t (completed steps immutable) is a sound design constraint for repair.
Cite this review
Pith. "Pith review of Plover: Steering GUI Agents through Plan-Centric Interaction." pith.science (2026). https://pith.science/paper/KG7UY24M
@misc{pith2026260715193,
author = {Pith},
title = {Pith review of: Plover: Steering GUI Agents through Plan-Centric Interaction},
year = {2026},
howpublished = {\url{https://pith.science/paper/KG7UY24M}},
note = {Machine review of arXiv:2607.15193}
}
read the original abstract
Graphical user interface (GUI) automation remains challenging in real-world environments, where dynamic layouts, unexpected dialogs, and evolving interface states can cause autonomous agents to drift from user intent. Recent vision-based multimodal agents improve flexibility by operating directly over screenshots and natural language instructions, but planning and adaptation often remain internal, limiting users' ability to inspect, supervise, or correct system behavior. We present Plover, a plan-centric vision-based GUI automation system that externalizes task plans and replanning as persistent, inspectable, and revisable artifacts. Through a planner--executor architecture, Plover supports explicit supervision of evolving execution, localized correction through editable plans, natural-language guidance, and screenshot-grounded interventions, while preserving prior progress during repair. A formative study with six participants informed the interaction design. We then evaluate Plover through benchmark failure-case repair and scenario-based workflow analyses. Our results show that many autonomous GUI-agent failures are structurally repairable when plans remain visible and interventions are localized, and that explicit replanning helps make GUI automation more transparent, controllable, and adaptable.
Figures
Figures from the paper (6 more)
Reference graph
Works this paper leans on
-
[1]
Mohamed Aghzal, Gregory J Stein, and Ziyu Yao. 2026. Why do LLM-based web agents fail? A hierarchical planning perspective. InProceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 32157–32180
2026
-
[2]
Saleema Amershi, Dan Weld, Mihaela Vorvoreanu, Adam Fourney, Besmira Nushi, Penny Collisson, Jina Suh, Shamsi Iqbal, Paul N Bennett, Kori Inkpen, et al. 2019. Guidelines for human-AI interaction. InProceedings of the 2019 CHI Conference on Human Factors in Computing Systems. 1–13
2019
-
[3]
Nishant Balepur, Matthew Shu, Yoo Yeon Sung, Seraphina Goldfarb-Tarrant, Shi Feng, Fumeng Yang, Rachel Rudinger, and Jordan Lee Boyd-Graber. 2025. A Good Plan is Hard to Find: Aligning Models with Preferences is Misaligned with What Helps Users. InProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, EMNLP 2025, Suzhou, C...
2025
-
[4]
Tanvir Bhathal and Asanshay Gupta. 2025. Websight: A vision-first architecture for robust web agents.arXiv preprint arXiv:2508.16987(2025)
arXiv 2025
-
[5]
Léo Boisvert, Megh Thakkar, Maxime Gasse, Massimo Caccia, Thibault L De Chezelles, Quentin Cappart, Nicolas Chapados, Alexandre Lacoste, and Alexandre Drouin. 2024. Workarena++: Towards compositional planning and reasoning-based common knowledge work tasks.Advances in Neural Information Processing Systems37 (2024), 5996–6051
2024
-
[6]
Sacha Brisset, Romain Rouvoy, Lionel Seinturier, and Renaud Pawlak. 2022. Er- ratum: Leveraging Flexible Tree Matching to repair broken locators in web automation scripts.Inf. Softw. Technol.144 (2022), 106754. doi:10.1016/J.INFSOF. 2021.106754
arXiv 2022
-
[7]
Meyer, Yuning Chai, Dennis Park, and Yong Jae Lee
Mu Cai, Haotian Liu, Siva Karthik Mustikovela, Gregory P. Meyer, Yuning Chai, Dennis Park, and Yong Jae Lee. 2024. ViP-LLaVA: Making Large Multimodal Models Understand Arbitrary Visual Prompts. InIEEE/CVF Conference on Com- puter Vision and Pattern Recognition, CVPR 2024, Seattle, W A, USA, June 16-22,
2024
-
[8]
Fadnis, Kartik Talamadupula, Mishal Dholakia, Biplav Srivastava, Jeffrey O
Tathagata Chakraborti, Kshitij P. Fadnis, Kartik Talamadupula, Mishal Dholakia, Biplav Srivastava, Jeffrey O. Kephart, and Rachel K. E. Bellamy. 2018. Visualiza- tions for an Explainable Planning Agent. InProceedings of the Twenty-Seventh International Joint Conference on Artificial Intelligence, IJCAI 2018, July 13-19, 2018, Stockholm, Sweden, Jérôme Lan...
2018
Show all 102 references
-
[9]
Tathagata Chakraborti, Sarath Sreedharan, Sachin Grover, and Subbarao Kamb- hampati. 2019. Plan explanations as model reconciliation–an empirical study. In 2019 14th ACM/IEEE International Conference on Human-Robot Interaction (HRI). Ieee, 258–266
2019
-
[10]
Juntong Chen, Jiang Wu, Jiajing Guo, Vikram Mohanty, Xueming Li, Jorge Pi- azentin Ono, Wenbin He, Liu Ren, and Dongyu Liu. 2025. InterChat: Enhancing generative visual analytics using multimodal interactions. InComputer Graphics Forum, Vol. 44. Wiley Online Library, e70112
2025
-
[11]
Weihao Chen, Chun Yu, Huadong Wang, Zheng Wang, Lichen Yang, Yukun Wang, Weinan Shi, and Yuanchun Shi. 2023. From gap to synergy: Enhancing contextual understanding through human-machine collaboration in personalized systems. InProceedings of the 36th Annual ACM Symposium on U...
2023
-
[12]
Ziming Cheng, Zhiyuan Huang, Junting Pan, Zhaohui Hou, and Mingjie Zhan
-
[13]
Riccardo Coppola, Luca Ardito, and Marco Torchiano. 2019. Fragility of layout-based and visual GUI test scripts: an assessment study on a hybrid mo- bile application. InProceedings of the 10th ACM SIGSOFT International Work- shop on Automating TEST Case Design, Selection, and ...
2019
-
[14]
Laradji, Manuel Del Verme, Tom Marty, David Vázquez, Nicolas Chapados, and Alexandre Lacoste
Alexandre Drouin, Maxime Gasse, Massimo Caccia, Issam H. Laradji, Manuel Del Verme, Tom Marty, David Vázquez, Nicolas Chapados, and Alexandre Lacoste
-
[15]
Anurag Dwarakanath, Neville Dubash, and Sanjay Podder. 2018. Machines that test Software like Humans.CoRRabs/1809.09455 (2018). arXiv:1809.09455 http://arxiv.org/abs/1809.09455
2018 arXiv
-
[16]
Lutfi Eren Erdogan, Nicholas Lee, Sehoon Kim, Suhong Moon, Hiroki Furuta, Gopala Anumanchipalli, Kurt Keutzer, and Amir Gholami. 2025. Plan-and-Act: Improving Planning of Agents for Long-Horizon Tasks. InProceedings of the 42nd International Conference on Machine Learning (Pro...
2025
-
[17]
WorkArena: How Capable are Web Agents at Solving Common Knowledge Work Tasks?Forty-first International Conference on Machine Learning, ICML 2024, Vienna, Austria, July 21-27, 2024235 (2024), 11642–11662
2024
-
[18]
Boyu Gou, Demi Ruohan Wang, Boyuan Zheng, Yanan Xie, Cheng Chang, Yiheng Shu, Huan Sun, and Yu Su. 2025. Navigating the digital world as humans do: Universal visual grounding for gui agents. InInternational Conference on Learning Representations, Vol. 2025. 30851–30883
2025
-
[19]
Nitesh Goyal, Minsuk Chang, and Michael Terry. 2024. Designing for Human- Agent Alignment: Understanding what humans want from their agents. InEx- tended Abstracts of the CHI Conference on Human Factors in Computing Systems. 1–6
2024
-
[20]
KJ Kevin Feng, Kevin Pu, Matt Latzke, Tal August, Pao Siangliulue, Jonathan Bragg, Daniel S Weld, Amy X Zhang, and Joseph Chee Chang. 2026. Cocoa: Co-planning and co-execution with ai agents. InProceedings of the 2026 CHI Conference on Human Factors in Computing Systems. ACM, 1–23
2026
-
[21]
Grigorev, Alexey K
Danil S. Grigorev, Alexey K. Kovalev, and Aleksandr I. Panov. 2025. VerifyLLM: LLM-Based Pre-Execution Task Plan Verification for Robots. InIEEE/RSJ Interna- tional Conference on Intelligent Robots and Systems, IROS 2025, Hangzhou, China, October 19-25, 2025. IEEE, 18489–18496
2025
-
[22]
Xiangwu Guo, Difei Gao, and Mike Zheng Shou. 2025. AUTO-Explorer: Automated Data Collection for GUI Agent.CoRRabs/2511.06417 (2025). arXiv:2511.06417 doi:10.48550/ARXIV.2511.06417
2025 doi
-
[23]
Maria Fernanda Granda, Otto Parra, and Bryan Alba-Sarango. 2021. Towards a Model-Driven Testing Framework for GUI Test Cases Generation from User Stories.. InENASE. 453–460
2021
-
[24]
Hongliang He, Wenlin Yao, Kaixin Ma, Wenhao Yu, Yong Dai, Hongming Zhang, Zhenzhong Lan, and Dong Yu. 2024. Webvoyager: Building an end-to-end web agent with large multimodal models. InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Vol...
2024
-
[25]
Theodore D Hellmann and Frank Maurer. 2011. Rule-based exploratory testing of graphical user interfaces. In2011 Agile Conference. IEEE, 107–116
2011
-
[26]
Chao Hao, Shuai Wang, and Kaiwen Zhou. 2025. Uncertainty-aware gui agent: Adaptive perception through component recommendation and human-in-the- loop refinement.arXiv preprint arXiv:2508.04025(2025)
2025 arXiv
-
[27]
Eric Horvitz. 1999. Principles of Mixed-Initiative User Interfaces. InProceeding of the CHI ’99 Conference on Human Factors in Computing Systems: The CHI is the Limit, Pittsburgh, PA, USA, May 15-20, 1999, Marian G. Williams and Mark W. Altom (Eds.). ACM, 159–166. doi:10.1145/...
1999
-
[28]
Siyuan Hu, Mingyu Ouyang, Difei Gao, and Mike Zheng Shou. 2024. The dawn of gui agent: A preliminary case study with claude 3.5 computer use.arXiv preprint arXiv:2411.10323(2024)
2024 arXiv
-
[29]
Peter Hofmann, Caroline Samp, and Nils Urbach. 2020. Robotic process automa- tion.Electronic markets30, 1 (2020), 99–106
2020
-
[30]
Zhiyuan Huang, Ziming Cheng, Junting Pan, Zhaohui Hou, and Mingjie Zhan
-
[31]
Faria Huq, Zora Zhiruo Wang, Frank F Xu, Tianyue Ou, Shuyan Zhou, Jeffrey P Bigham, and Graham Neubig. 2025. Cowpilot: a framework for autonomous and human-agent collaborative web navigation. InProceedings of the 2025 Confer- ence of the Nations of the Americas Chapter of the ...
2025
-
[32]
Wenyue Hua, Mengting Wan, Jagannath Vadrevu, Ryan Nadel, Yongfeng Zhang, and Chi Wang. 2025. Interactive speculative planning: Enhance agent efficiency through co-design of system and user interface. InInternational Conference on Learning Representations, Vol. 2025. 14256–14283
2025
-
[33]
Arushi Jain, Shubham Paliwal, Monika Sharma, Lovekesh Vig, and Gau- tam Shroff. 2024. SmartFlow: Robotic Process Automation using LLMs. arXiv:2405.12842 [cs.RO] https://arxiv.org/abs/2405.12842
2024 arXiv
-
[34]
InProceedings of the Computer Vision and Pattern Recognition Conference
Spiritsight agent: Advanced gui agent with one look. InProceedings of the Computer Vision and Pattern Recognition Conference. 29490–29500
-
[35]
Linjia Kang, Zhimin Wang, Yongkang Zhang, Duo Wu, Jinghe Wang, Ming Ma, Haopeng Yan, and Zhi Wang. 2026. Learning with Challenges: Adaptive Difficulty-Aware Data Generation for Mobile GUI Agent Training.arXiv preprint arXiv:2601.22781(2026)
2026
-
[36]
Alayt Issak, Jeba Rezwana, and Casper Harteveld. 2025. MOSAAIC: Managing Optimization towards Shared Autonomy, Authority, and Initiative in Co-creation. arXiv:2505.11481 [cs.AI] https://arxiv.org/abs/2505.11481
2025 arXiv
-
[37]
Jing Yu Koh, Robert Lo, Lawrence Jang, Vikram Duvvur, Ming Lim, Po-Yu Huang, Graham Neubig, Shuyan Zhou, Russ Salakhutdinov, and Daniel Fried. 2024. Vi- sualwebarena: Evaluating multimodal agents on realistic visual web tasks. In Proceedings of the 62nd Annual Meeting of the A...
2024
-
[38]
Mitchell, and Anupam Datta
Allison Sihan Jia, Daniel Huang, Nikhil Vytla, Nirvika Choudhury, Shayak Sen, John C. Mitchell, and Anupam Datta. 2025. What Is Your Agent’s GPA? A Frame- work for Evaluating Agent Goal-Plan-Action Alignment.CoRRabs/2510.08847 (2025)
2025
-
[39]
Kevin Qinghong Lin, Linjie Li, Difei Gao, Zhengyuan Yang, Shiwei Wu, Zechen Bai, Stan Weixian Lei, Lijuan Wang, and Mike Zheng Shou. 2025. Showui: One vision-language-action model for gui visual agent. InProceedings of the Computer Vision and Pattern Recognition Conference. 19...
2025
-
[40]
Anjali Khurana, Xiaotian Su, April Yi Wang, and Parmit K. Chilana. 2025. Do It For Me vs. Do It With Me: Investigating User Perceptions of Different Paradigms of Automation in Copilots for Feature-Rich Software. InProceedings of the 2025 CHI Conference on Human Factors in Comp...
2025
-
[41]
Yadong Lu, Jianwei Yang, Yelong Shen, and Ahmed Hassan Awadallah. 2025. OmniParser for Pure Vision Based GUI Agent. https://openreview.net/forum? id=C6hUK6Q1Pi
2025
-
[42]
Toby Jia-Jun Li, Amos Azaria, and Brad A Myers. 2017. SUGILITE: creating multimodal smartphone automation by demonstration. InProceedings of the 2017 CHI conference on human factors in computing systems. 6038–6049
2017
-
[43]
Jordan Madden, Moxanki Bhavsar, Lhamo Dorje, and Xiaohua Li. 2024. Ro- bustness of Practical Perceptual Hashing Algorithms to Hash-Evasion and Hash- Inversion Attacks. InThe Third Workshop on New Frontiers in Adversarial Machine Learning. https://openreview.net/forum?id=hraOxsleRl
2024
-
[44]
Tao Long, Xuanming Zhang, Sitong Wang, Zhou Yu, and Lydia B Chilton. 2025. DoubleAgents: Interactive Simulations for Alignment in Agentic AI.arXiv preprint arXiv:2509.12626(2025)
2025 arXiv
-
[45]
Hussein Mozannar, Gagan Bansal, Cheng Tan, Adam Fourney, Victor Dibia, Jingya Chen, Jack Gerrits, Tyler Payne, Matheus Kunzler Maldaner, Madeleine Grunde-McLaughlin, et al. 2025. Magentic-ui: Towards human-in-the-loop agen- tic systems.arXiv preprint arXiv:2507.22358(2025)
2025 arXiv
-
[46]
Shang Ma, Xusheng Xiao, and Yanfang Ye. 2025. Agent+ P: Guiding UI Agents via Symbolic Planning.arXiv preprint arXiv:2510.06042(2025)
2025
-
[47]
Rodriguez, Montek Kalsi, Nicolas Chapados, M
Shravan Nayak, Xiangru Jian, Kevin Qinghong Lin, Juan A. Rodriguez, Montek Kalsi, Nicolas Chapados, M. Tamer Özsu, Aishwarya Agrawal, David Vazquez, Christopher Pal, Perouz Taslakian, Spandana Gella, and Sai Rajeswar. 2025. UI- Vision: A Desktop-centric GUI Benchmark for Visua...
2025
-
[48]
Valérie Maquil, Dimitra Anastasiou, Hoorieh Afkari, Adrien Coppens, Johannes Hermen, and Lou Schwartz. 2023. Establishing Awareness through Pointing Gestures during Collaborative Decision-Making in a Wall-Display Environment. InExtended Abstracts of the 2023 CHI Conference on ...
2023
-
[49]
Bigham, and Amy Pavel
Yi-Hao Peng, Dingzeyu Li, Jeffrey P. Bigham, and Amy Pavel. 2025. Morae: Proactively Pausing UI Agents for User Choices. InProceedings of the 38th Annual ACM Symposium on User Interface Software and Technology, UIST 2025, Busan, Korea, 28 September 2025 - 1 October 2025, Andre...
2025
-
[50]
Arpit Narechania, Shunan Guo, Eunyee Koh, Alex Endert, and Jane Hoffswell
-
[51]
Utilizing Provenance as an Attribute for Visual Data Analysis: A Design Probe With ProvenanceLens.IEEE Trans. Vis. Comput. Graph.31, 10 (2025), 8452–8465
2025
-
[52]
Ju Qian, Zhengyu Shang, Shuoyan Yan, Yan Wang, and Lin Chen. 2020. Roscript: a visual script driven truly non-intrusive robotic testing system for touch screen applications. InProceedings of the ACM/IEEE 42nd International Conference on Software Engineering. 297–308
2020
-
[53]
Mehrab Tanjim, Nesreen K
Dang Nguyen, Jian Chen, Yu Wang, Gang Wu, Namyong Park, Zhengmian Hu, Hanjia Lyu, Junda Wu, Ryan Aponte, Yu Xia, Xintong Li, Jing Shi, Hongjie Chen, Viet Dac Lai, Zhouhang Xie, Sungchul Kim, Ruiyi Zhang, Tong Yu, Md. Mehrab Tanjim, Nesreen K. Ahmed, Puneet Mathur, Seunghyun Yo...
2025
-
[54]
Stephan Rabanser, Sayash Kapoor, Peter Kirgis, Kangheng Liu, Saiteja Utpala, and Arvind Narayanan. 2026. Towards a Science of AI Agent Reliability.arXiv preprint arXiv:2602.16666(2026)
2026 arXiv
-
[55]
Christopher Potts and Moritz Sudhof. 2026. Invisible failures in human-AI inter- actions.arXiv preprint arXiv:2603.15423(2026)
2026 arXiv
-
[56]
Petr Průcha, Michaela Matoušková, and Jan Strnad. 2025. Are LLM Agents the New RPA? A Comparative Study with RPA Across Enterprise Workflows.arXiv preprint arXiv:2509.04198(2025)
2025 arXiv
-
[57]
Minjie Shen, Yanshu Li, Lulu Chen, and Qikai Yang. 2025. From mind to ma- chine: The rise of manus ai as a fully autonomous digital agent.arXiv preprint arXiv:2505.02024(2025)
2025 arXiv
- [58]
-
[59]
When to Hand Off, When to Work Together
Kihoon Son, Hyewon Lee, DaEun Choi, Yoonsu Kim, Tae Soo Kim, Yoonjoo Lee, John Joon Young Chung, HyunJoon Jung, and Juho Kim. 2026. " When to Hand Off, When to Work Together": Expanding Human-Agent Co-Creative Collaboration through Concurrent Interaction.arXiv preprint arXiv:2...
2026 arXiv
-
[60]
Peter Shaw, Mandar Joshi, James Cohan, Jonathan Berant, Panupong Pasupat, Hexiang Hu, Urvashi Khandelwal, Kenton Lee, and Kristina N Toutanova. 2023. From pixels to ui actions: Learning to follow instructions via graphical user interfaces.Advances in Neural Information Process...
2023
-
[61]
Erfan Shayegani, Keegan Hines, Yue Dong, Nael Abu-Ghazaleh, Roman Lutz, Spencer Whitehead, Vidhisha Balachandran, Besmira Nushi, and Vibhav Vineet
-
[62]
Qiushi Sun, Kanzhi Cheng, Zichen Ding, Chuanyang Jin, Yian Wang, Fangzhi Xu, Zhenyu Wu, Chengyou Jia, Liheng Chen, Zhoumianze Liu, Ben Kao, Guohao Li, Junxian He, Yu Qiao, and Zhiyong Wu. 2025. OS-Genesis: Automating GUI Agent Trajectory Construction via Reverse Task Synthesis...
2025
-
[63]
Fei Tang, Haolei Xu, Hang Zhang, Siqi Chen, Xingyu Wu, Yongliang Shen, Wenqi Zhang, Guiyang Hou, Zeqi Tan, Yuchen Yan, et al. 2025. A survey on (m) llm-based gui agents.arXiv preprint arXiv:2504.13865(2025)
2025 arXiv
-
[64]
Aleksandar Shtedritski, Christian Rupprecht, and Andrea Vedaldi. 2023. What does CLIP know about a red circle? Visual prompt engineering for VLMs. In IEEE/CVF International Conference on Computer Vision, ICCV 2023, Paris, France, October 1-6, 2023. IEEE, 11953–11963. doi:10.11...
2023
-
[65]
Yuanrong Tang, Huiling Peng, Bingxi Zhao, Hengyang Ding, Hanchao Song, Tianhong Wang, Chen Zhong, and Jiangtao Gong. 2026. Human Tool: An MCP- Style Framework for Human-Agent Collaboration.arXiv preprint arXiv:2602.12953 (2026)
2026
-
[66]
Yunpeng Song, Yiheng Bian, Yongtao Tang, Guiyu Ma, and Zhongmin Cai. 2024. Visiontasker: Mobile task automation using vision based ui understanding and llm task planning. InProceedings of the 37th Annual ACM Symposium on User Interface Software and Technology. 1–17
2024
-
[67]
Zihe Song, S. M. Hasan Mansur, Ravishka Rathnasuriya, Yumna Fatima, Wei Yang, Kevin Moran, and Wing Lam. 2025. Can You Mimic Me? Exploring the Use of Android Record & Replay Tools in Debugging. In12th IEEE/ACM In- ternational Conference on Mobile Software Engineering and Syste...
2025
-
[68]
Vanshika Vats, Marzia Binta Nizam, Minghao Liu, Ziyuan Wang, Richard Ho, Mohnish Sai Prasad, Vincent Titterton, Sai Venkat Malreddy, Riya Aggarwal, Yan- wen Xu, et al. 2024. A Survey on Human-AI Collaboration with Large Foundation Models.arXiv preprint arXiv:2403.04931(2024)
2024 arXiv
-
[69]
Bowen Wang, Xinyuan Wang, Jiaqi Deng, Tianbao Xie, Ryan Li, Yanzhe Zhang, Junli Wang, Dunjie Lu, Zicheng Gong, Gavin Li, Toh Jing Hua, Wei-Lin Chiang, Ion Stoica, Diyi Yang, Yu Su, Yi Zhang, Zhiguo Wang, Victor Zhong, and Tao Yu
-
[70]
Jingyu Tang, Chaoran Chen, Jiawen Li, Zhiping Zhang, Bingcan Guo, Ibrahim Khalilov, Simret Araya Gebreegziabher, Bingsheng Yao, Dakuo Wang, Yanfang Ye, et al. 2026. Dark patterns meet gui agents: Llm agent susceptibility to manip- ulative interfaces and the role of human overs...
2026
-
[71]
Hui Wei, Zihao Zhang, Shenghua He, Tian Xia, Shijia Pan, and Fei Liu. 2025. PlanGenLLMs: A Modern Survey of LLM Planning Capabilities. InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2025, Vienna, Austria, ...
2025
-
[72]
Kimi Team, Tongtong Bai, Yifan Bai, Yiping Bao, SH Cai, Yuan Cao, Y Charles, HS Che, Cheng Chen, Guanduo Chen, et al . 2026. Kimi K2. 5: Visual Agentic Intelligence.arXiv preprint arXiv:2602.02276(2026)
2026 arXiv
-
[73]
Ching-Yi Tsai, Nicole Tacconi, Andrew D Wilson, and Parastoo Abtahi. 2026. Uncertain Pointer: Situated Feedforward Visualizations for Ambiguity-Aware AR Target Selection.arXiv preprint arXiv:2602.13433(2026)
2026
-
[74]
Judith Wewerka and Manfred Reichert. 2020. Robotic Process Automation - A Systematic Literature Review and Assessment Framework.CoRRabs/2012.11951 (2020). arXiv:2012.11951 https://arxiv.org/abs/2012.11951
2020 arXiv
-
[75]
Rossi, Ruiyi Zhang, Subrata Mitra, Dimitris N
Junda Wu, Zhehao Zhang, Yu Xia, Xintong Li, Zhaoyang Xia, Aaron Chang, Tong Yu, Sungchul Kim, Ryan A. Rossi, Ruiyi Zhang, Subrata Mitra, Dimitris N. Metaxas, Lina Yao, Jingbo Shang, and Julian J. McAuley. 2024. Visual Prompting in Multimodal Large Language Models: A Survey.CoR...
-
[76]
InThe Fourteenth International Conference on Learning Representations
Computer Agent Arena: Toward Human-Centric Evaluation and Analysis of Computer-Use Agents. InThe Fourteenth International Conference on Learning Representations. https://openreview.net/forum?id=3x4SDbXbgl
-
[77]
Bovik, Hamid R
Zhou Wang, Alan C. Bovik, Hamid R. Sheikh, and Eero P. Simoncelli. 2004. Image quality assessment: from error visibility to structural similarity.IEEE Trans. Image Process.13, 4 (2004), 600–612
2004
-
[78]
Tianbao Xie, Danyang Zhang, Jixuan Chen, Xiaochuan Li, Siheng Zhao, Ruisheng Cao, Toh Jing Hua, Zhoujun Cheng, Dongchan Shin, Fangyu Lei, Yitao Liu, Yiheng Xu, Shuyan Zhou, Silvio Savarese, Caiming Xiong, Victor Zhong, and Tao Yu
-
[79]
Hao Wen, Yuanchun Li, Guohong Liu, Shanhui Zhao, Tao Yu, Toby Jia-Jun Li, Shiqi Jiang, Yunhao Liu, Yaqin Zhang, and Yunxin Liu. 2024. Autodroid: Llm-powered task automation in android. InProceedings of the 30th Annual International Con- ference on Mobile Computing and Networki...
2024
- [80]
-
[81]
Yuhao Yang, Zhen Yang, Zi-Yi Dou, Anh Nguyen, Keen You, Omar Attia, Andrew Szot, Michael Feng, Ram Ramrakhya, Alexander Toshev, et al. 2025. Ultracua: A foundation model for computer use agents with hybrid action.arXiv preprint arXiv:2510.17790(2025)
2025 arXiv
-
[82]
Shunyu Yao, Howard Chen, John Yang, and Karthik Narasimhan. 2022. Webshop: Towards scalable real-world web interaction with grounded language agents. Advances in Neural Information Processing Systems35 (2022), 20744–20757
2022
-
[83]
Penghao Wu, Shengnan Ma, Bo Wang, Jiaheng Yu, Lewei Lu, and Ziwei Liu. 2026. Gui-reflection: Empowering multimodal gui models with self-reflection behavior. Advances in Neural Information Processing Systems38 (2026), 101861–101896
2026
-
[84]
Zhiyong Wu, Zhenyu Wu, Fangzhi Xu, Yian Wang, Qiushi Sun, Chengyou Jia, Kanzhi Cheng, Zichen Ding, Liheng Chen, Paul Pu Liang, and Yu Qiao. 2025. OS- ATLAS: Foundation Action Model for Generalist GUI Agents. InThe Thirteenth International Conference on Learning Representations...
2025
-
[85]
Ryan Yen, Jian Zhao, and Daniel Vogel. 2025. Code Shaping: Iterative Code Editing with Free-form AI-Interpreted Sketching. InProceedings of the 2025 CHI Conference on Human Factors in Computing Systems, CHI 2025, YokohamaJapan, 26 Conference’17, July 2017, Washington, DC, USA ...
2025
-
[86]
OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments. InAdvances in Neural Information Processing Systems 38: Annual Conference on Neural Information Processing Systems 2024, NeurIPS 2024, Vancouver, BC, Canada, December 10 - 15, 2024, Ami...
2024
-
[87]
Yuan Xu, Shaowen Xiang, Yizhi Song, Ruoting Sun, and Xin Tong. 2026. DuetUI: A Bidirectional Context Loop for Human-Agent Co-Generation of Task-Oriented Interfaces.Proceedings of the 2026 CHI Conference on Human Factors in Computing Systems, CHI 2026, Barcelona, Spain, April 1...
2026
-
[88]
John Yang, Carlos E Jimenez, Alexander Wettig, Kilian Lieret, Shunyu Yao, Karthik Narasimhan, and Ofir Press. 2024. Swe-agent: Agent-computer interfaces enable automated software engineering.Advances in Neural Information Processing Systems37 (2024), 50528–50652
2024
-
[89]
Yichun Zhang, Xiangwu Guo, Yauhong Goh, Jessica Hu, Zhiheng Chen, Xin Wang, Difei Gao, and Mike Zheng Shou. 2026. ShowUI-Aloha: Human-Taught GUI Agent.CoRRabs/2601.07181 (2026)
2026
-
[90]
Zhou Zhao, Shengyu Zhang, Liang Wang, Xiangxin Zhou, Zhaokai Wang, Kun Kuang, Fei Wu, Wangchunshu Zhou, Shuofei Qiao, Jiwei Li, Guoyin Wang, Ziyu Zhao, Hongxia Yang, Fan Wu, Jiasheng Ye, Shenzhi Wang, Ruixuan Xiao, Tieyong Zeng, Yuhuai Li, Yuchen Eleanor Jiang, Meiling Tao, Xu...
2025 arXiv
-
[91]
Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik R Narasimhan, and Yuan Cao. 2022. React: Synergizing reasoning and acting in language models. InThe eleventh international conference on learning representations
2022
-
[92]
Tom Yeh, Tsung-Hsiang Chang, and Robert C. Miller. 2009. Sikuli: using GUI screenshots for search and automation. InProceedings of the 22nd Annual ACM Symposium on User Interface Software and Technology, Victoria, BC, Canada, Octo- ber 4-7, 2009, Andrew D. Wilson and François ...
2009
-
[94]
Shengcheng Yu, Chunrong Fang, Mingzhe Du, Yuchen Ling, Zhenyu Chen, and Zhendong Su. 2024. Practical Non-Intrusive GUI Exploration Testing with Visual- based Robotic Arms. InProceedings of the 46th IEEE/ACM International Conference on Software Engineering, ICSE 2024, Lisbon, P...
2024
-
[95]
Hyeonggeun Yun and Jinkyu Jang. 2025. Interaction-Driven Browsing: A Human- in-the-Loop Conceptual Framework Informed by Human Web Browsing for Browser-Using Agents.CoRRabs/2509.12049 (2025). doi:10.48550/ARXIV.2509. 12049
2025 doi
-
[96]
Shaojie Zhang, Ruoceng Zhang, Pei Fu, Shaokang Wang, Jiahui Yang, Xin Du, Bin Qin, Ying Huang, Zhenbo Luo, and Jian Luan. 2026. Btl-ui: Blink-think-link reasoning model for gui agent.Advances in Neural Information Processing Systems 38 (2026), 56035–56056
2026
-
[99]
Shuyan Zhou, Frank F Xu, Hao Zhu, Xuhui Zhou, Robert Lo, Abishek Sridhar, Xianyi Cheng, Tianyue Ou, Yonatan Bisk, Daniel Fried, et al. 2024. Webarena: A realistic web environment for building autonomous agents. InInternational Conference on Learning Representations, Vol. 2024....
2024
-
[100]
Added” and “Changed
Henry Peng Zou, Wei-Chieh Huang, Yaozu Wu, Jizhou Guo, Yankai Chen, Chunyu Miao, Hoang H Nguyen, Yue Zhou, Weizhi Zhang, Liancheng Fang, et al. 2026. Llm-based human-agent collaboration and interaction systems: A survey.Findings of the Association for Computational Linguistics...
2026
-
[101]
SUMMARY: <one short imperative step sentence>
-
[102]
I will”, “Let’s
RATIONALE: <1–2 sentences explaining the detected failure and why the proposed next action helps> - The SUMMARY must: •start with a strong action verb •be written as a standalone executable step •not contain “I will”, “Let’s”, or future tense •not mention internal tool names -...
2017
-
[2024]
doi:10.1109/CVPR52733.2024.01227
IEEE, 12914–12923. doi:10.1109/CVPR52733.2024.01227
2024
-
[2025]
Conference’17, July 2017, Washington, DC, USA Venkatesan et al
Navi-plus: Managing Ambiguous GUI Navigation Tasks with Follow-up Questions.arXiv preprint arXiv:2503.24180(2025). Conference’17, July 2017, Washington, DC, USA Venkatesan et al
2025 arXiv
-
[2026]
In The Fourteenth International Conference on Learning Representations
Just Do It!? Computer-Use Agents Exhibit Blind Goal-Directedness. In The Fourteenth International Conference on Learning Representations. https: //openreview.net/forum?id=9W4bPRsEIT
Reviewed August 1, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.