REVIEW 3 major objections 4 minor 102 references
Plover shows that many GUI-agent failures become repairable when the task plan stays visible and corrections stay localized, with 23 of 26 benchmark failures improved by mixed-initiative interaction.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-01 23:52 UTC pith:KG7UY24M
load-bearing objection Plover is a credible systems paper with an honest upper-bound recovery result; just don't let the abstract sell the 88% as proof that plan visibility is what rescues failures. the 3 major comments →
Plover: Steering GUI Agents through Plan-Centric Interaction
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
Plover's core claim is that GUI-agent failures are structurally recoverable when plans are externalized and repairs are localized. The system keeps a versioned plan artifact, separates immutable executed steps from an editable pending suffix, and supports user-driven interventions (natural-language guidance, plan edits, and screenshot annotations) plus system-driven replanning triggered by non-progress detection. In the benchmark repair analysis, 26 autonomous failures were re-run in a mixed-initiative setting; 23 improved (17 complete successes, 6 partial), only 3 remained failures, and all 10 autonomous partial successes became complete successes. The paper also characterizes which failure
What carries the argument
The central mechanism is the persistent plan artifact, a shared state representation with the invariant that executed steps are immutable and only the pending suffix can be revised (C_{t+1}=C_t). On top of this, Plover implements Intelligent Replanning in two modes: User-Driven IR, where natural-language guidance or multimodal annotations (strokes, shapes, text overlays captured as primitives with a bounding box) generate localized plan proposals, and System-Driven IR, where a watchdog detects behavioral loop repetition and visual non-progress (via dHash Hamming distance) and injects a structured failure message that prompts the model to propose a recovery step with rationale. The versioned
Load-bearing premise
The 88% recovery figure depends on an expert user who already knows what went wrong and what the correct target is; if ordinary users cannot detect drift or formulate correct localized corrections, the recoverability claim may not transfer to practice.
What would settle it
Run the same 26 benchmark failures with naive participants who are not told the failure cause, providing only the Plover interface as the intervention channel. If the mixed-initiative recovery rate falls to near the autonomous baseline (no significant improvement over the 0% success on these originally failed tasks), the claim that failures are structurally repairable through visible plans is not supported for realistic users.
If this is right
- If recoverability holds beyond the expert setting, GUI agents can be deployed in long-horizon, high-friction workflows with a human steering loop instead of requiring near-perfect autonomy.
- The plan-invariant design means corrections preserve executed history, so each intervention is cheaper and less disruptive than re-prompting or restarting the whole task.
- System-driven non-progress detection (repeated semantic actions plus visual stability) can act as a reusable watchdog that catches drift before it propagates, independent of the specific planner or executor.
- The failure taxonomy suggests that perception errors and state misinterpretations are cheaply repairable with language or annotations, while compound failures require catching the initial planning error earlier.
- Exposing plans as versioned, diffable artifacts provides a natural audit trail for when and why an agent's behavior changed, which can support verification and post-hoc analysis.
Where Pith is reading between the lines
- A natural next test, which the paper does not run, is a study with non-expert users who are not told the failure cause: if recovery rates drop to near the autonomous baseline, the 'structurally recoverable' claim would need to be re-scoped from an upper bound to a property that depends on user diagnostic skill.
- The plan artifact as a coordination protocol could generalize beyond GUI automation to other long-horizon agent domains (e.g., data-cleaning pipelines or robotics task plans) where partial progress is valuable and corrections must be localized.
- The System-Driven IR watchdog (behavioral repetition + perceptual-hash stability) is a concrete, model-agnostic component that could be extracted and benchmarked on its own to measure how many agent stalls it catches before a human would notice.
- The paper itself flags that visible plans may inflate user confidence; an empirical study measuring whether users over-accept plan proposals when the system looks confident would be a direct extension.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces Plover, a plan-centric GUI automation system that externalizes task plans as persistent, editable artifacts and supports mixed-initiative repair through natural-language guidance, multimodal annotation, plan edits, and system-driven replanning. The authors report a formative study with six participants, a benchmark repair study on 38 OSWorld-Verified tasks (26 autonomous non-successes re-run with expert interventions, yielding 23 improved, 17 complete successes, 6 partial successes, 3 failures), and a scenario-based stability analysis with trajectory-derived prompts. The central claim is that many GUI-agent failures are structurally repairable when plans remain visible and interventions are localized, and that explicit replanning improves transparency, controllability, and adaptability.
Significance. If the central claim were fully established, the paper would make a useful contribution to human-agent interaction for GUI automation: it proposes a concrete design space (persistent plans as coordination artifacts, localized repair, visible replanning) and provides a failure taxonomy that could guide future interface design. The paper is honest in labeling the benchmark repair study as an upper bound on plan-centric recoverability, and the appendices contain substantial implementation and evaluation detail, including per-task results and a formative study summary. These are real strengths. However, the main empirical evidence does not currently separate the effect of the plan-centric interface from the effect of an expert oracle user, so the causal design conclusions (DG1–DG5, 'explicit replanning helps') are not yet established by the data.
major comments (3)
- [Section 5.1, Table 1; Abstract] The load-bearing empirical claim—23/26 non-success cases improved, 88% recovery—is measured with the first author, who knows each failure cause and the correct target, supplying all interventions. There is no control condition that strips away the plan-centric affordances (plan panel, plan editing, annotation, visible replanning) while keeping the same underlying agent and the same expert. As written, the 88% figure is an upper bound on recoverability by an informed oracle, not evidence that plan visibility and localized repair cause the improvement. The paper's own limitation statement in Section 6 ('our design assumes users generally know the intended path well enough to correct the agent') does not carry through to the abstract/conclusion, which assert a causal role for visible plans. Add an ablation or control (e.g., same expert corrections issued as text prompts to the same base age
- [Section 5.2, 'Trajectory Sampling and Prompt Reconstruction' and Table 2] The scenario stability analysis synthesizes task prompts from sampled interaction trajectories and then compares the replayed plans and final states against the same trajectories. This creates a circularity: the prompt is derived from the reference trajectory, so plan alignment metrics (coverage 0.62, order 0.41, actionability 0.97) partly measure reconstruction from a trajectory-derived instruction rather than general plan quality. The browser-vs-desktop visual fidelity differences are still informative, but the plan alignment results should be presented as a property of this reverse-synthesis setup, not as evidence about Plover's planning under independently authored user instructions. A sanity check with manually authored prompts, or a clear caveat, is needed before these numbers are used to support the 'opportunity for users to inspect and correct' argument.
- [Section 5.3, 'Characterizing Repairable Failures'] The failure-mode analysis is based on the same 26 cases and the same expert interventions. Counts such as 'Execution Drift appeared in 46% (n=12/26)' and 'NL Guidance resolved 11 cases' are reported without any uncertainty or sensitivity analysis. With n=26 and intervention choices made by a single expert who already knows the failure causes, small counts can easily flip; the recovered vs. unrecovered distinction is not robust enough to support the strong claim that 'compound failures' are fundamentally harder. At minimum, report bootstrap or exact binomial confidence intervals and clarify that all recovery counts are conditional on the expert's choice of intervention.
minor comments (4)
- [Appendix C, Algorithm 1] System-Driven IR relies on hardcoded thresholds (REPEAT_SEQ_L3_R3, dHash Hamming distance > 40) with no sensitivity analysis. Since Section 5.3 attributes 9 successful recoveries to System-Driven IR, the threshold choices can materially affect the results; report how varying them changes detection and downstream recovery.
- [Section 5.1, 'no regressions'] The statement 'no regressions were observed' only covers the 26 autonomous non-success cases; it does not address whether the mixed-initiative interaction could degrade autonomous successes, since those were not re-run in the mixed-initiative condition. Please state this scope explicitly.
- [Table 1 and Table 2] Table 1's 'Improv. Rate' is not formally defined; for Multi-App, 6S+2P out of 10 corresponds to 80%, but the reader must infer the denominator. In Table 2, the 'Overall Average' row for MSE (939.57) is the mean of scenario averages, not the mean over all trials; clarify the aggregation.
- [Throughout] The paper uses 'Conference’17' in the ACM reference format and several placeholder-style citations (e.g., the DOI is 'XXXXXXX.XXXXXXX'). Please update the formatting to the final venue style and correct minor typographical issues such as the 'MI (a)' label in Figure 5.
Circularity Check
No significant circularity: the 88% recoverability result is an externally grounded upper-bound measurement with the expert-oracle caveat explicitly acknowledged; remaining concerns are validity limitations, not circular reductions.
full rationale
Plover's central empirical claim is an upper-bound measurement on the external OSWorld-Verified benchmark, not a quantity derived from a fitted parameter, an ansatz, or a load-bearing self-citation. The paper states the intervention condition explicitly: 'This setup establishes an upper bound on plan-centric recoverability rather than typical user performance' (Section 5.1), and the corresponding limitation is restated in Section 6 ('our design assumes users generally know the intended path well enough to correct the agent when needed'). Because the expert interventions are acknowledged as ideal rather than typical, the 88% recovery figure is an honest conditional result, not a hidden assumption presented as a finding. The scenario-based stability analysis (Section 5.2) synthesizes prompts from reference trajectories and then compares generated plans against those trajectories, which introduces non-independence; however the reported alignment metrics are moderate (coverage 0.62, order 0.41), so the result is not forced by construction. The only self-citation is a related-work mention ([10], multimodal interaction) and is not load-bearing. No uniqueness theorem, no ansatz smuggled via citation, and no equation reduces the conclusion to its input. The absence of a chat-only or invisible-plan control is a causal-identification limitation, but under the circularity criteria it does not constitute circularity.
Axiom & Free-Parameter Ledger
free parameters (4)
- Visual non-progress dHash threshold τ =
40 (Hamming distance)
- Repetition pattern REPEAT_SEQ_L3_R3 =
repeated action subsequence of length 3 seen 3 times
- Per-scenario similarity thresholds (SSIM/MSE/dHash) =
e.g., Firefox Fillable Form High: SSIM≥0.98, MSE≤100, dHash≤2; LibreOffice thresholds differ
- Selected trajectories per scenario =
5 per scenario from 100 exploration trials, after manual verification
axioms (6)
- domain assumption Claude 4.5 Sonnet via computer-use can decompose tasks into deterministic UI instructions and ground them in the executor.
- domain assumption Image-similarity metrics (SSIM, MSE, dHash) are valid proxies for GUI execution stability and task-state equivalence.
- domain assumption The 38 OSWorld-Verified tasks that failed for Claude 4.5 Sonnet are representative of GUI-agent failures in general.
- domain assumption An expert who knows failure causes can stand in for a typical user to establish 'structural recoverability'.
- domain assumption GPA phase taxonomy and trajectory-derived phases are a meaningful ground truth for plan alignment.
- domain assumption The invariant C_{t+1}=C_t (completed steps immutable) is a sound design constraint for repair.
read the original abstract
Graphical user interface (GUI) automation remains challenging in real-world environments, where dynamic layouts, unexpected dialogs, and evolving interface states can cause autonomous agents to drift from user intent. Recent vision-based multimodal agents improve flexibility by operating directly over screenshots and natural language instructions, but planning and adaptation often remain internal, limiting users' ability to inspect, supervise, or correct system behavior. We present Plover, a plan-centric vision-based GUI automation system that externalizes task plans and replanning as persistent, inspectable, and revisable artifacts. Through a planner--executor architecture, Plover supports explicit supervision of evolving execution, localized correction through editable plans, natural-language guidance, and screenshot-grounded interventions, while preserving prior progress during repair. A formative study with six participants informed the interaction design. We then evaluate Plover through benchmark failure-case repair and scenario-based workflow analyses. Our results show that many autonomous GUI-agent failures are structurally repairable when plans remain visible and interventions are localized, and that explicit replanning helps make GUI automation more transparent, controllable, and adaptable.
Figures
Reference graph
Works this paper leans on
-
[1]
Mohamed Aghzal, Gregory J Stein, and Ziyu Yao. 2026. Why do LLM-based web agents fail? A hierarchical planning perspective. InProceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 32157–32180
2026
-
[2]
Saleema Amershi, Dan Weld, Mihaela Vorvoreanu, Adam Fourney, Besmira Nushi, Penny Collisson, Jina Suh, Shamsi Iqbal, Paul N Bennett, Kori Inkpen, et al. 2019. Guidelines for human-AI interaction. InProceedings of the 2019 CHI Conference on Human Factors in Computing Systems. 1–13
2019
-
[3]
Nishant Balepur, Matthew Shu, Yoo Yeon Sung, Seraphina Goldfarb-Tarrant, Shi Feng, Fumeng Yang, Rachel Rudinger, and Jordan Lee Boyd-Graber. 2025. A Good Plan is Hard to Find: Aligning Models with Preferences is Misaligned with What Helps Users. InProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, EMNLP 2025, Suzhou, C...
2025
-
[4]
Tanvir Bhathal and Asanshay Gupta. 2025. Websight: A vision-first architecture for robust web agents.arXiv preprint arXiv:2508.16987(2025)
Pith/arXiv arXiv 2025
-
[5]
Léo Boisvert, Megh Thakkar, Maxime Gasse, Massimo Caccia, Thibault L De Chezelles, Quentin Cappart, Nicolas Chapados, Alexandre Lacoste, and Alexandre Drouin. 2024. Workarena++: Towards compositional planning and reasoning-based common knowledge work tasks.Advances in Neural Information Processing Systems37 (2024), 5996–6051
2024
-
[6]
Sacha Brisset, Romain Rouvoy, Lionel Seinturier, and Renaud Pawlak. 2022. Er- ratum: Leveraging Flexible Tree Matching to repair broken locators in web automation scripts.Inf. Softw. Technol.144 (2022), 106754. doi:10.1016/J.INFSOF. 2021.106754
arXiv 2022
-
[7]
Meyer, Yuning Chai, Dennis Park, and Yong Jae Lee
Mu Cai, Haotian Liu, Siva Karthik Mustikovela, Gregory P. Meyer, Yuning Chai, Dennis Park, and Yong Jae Lee. 2024. ViP-LLaVA: Making Large Multimodal Models Understand Arbitrary Visual Prompts. InIEEE/CVF Conference on Com- puter Vision and Pattern Recognition, CVPR 2024, Seattle, W A, USA, June 16-22,
2024
-
[8]
Fadnis, Kartik Talamadupula, Mishal Dholakia, Biplav Srivastava, Jeffrey O
Tathagata Chakraborti, Kshitij P. Fadnis, Kartik Talamadupula, Mishal Dholakia, Biplav Srivastava, Jeffrey O. Kephart, and Rachel K. E. Bellamy. 2018. Visualiza- tions for an Explainable Planning Agent. InProceedings of the Twenty-Seventh International Joint Conference on Artificial Intelligence, IJCAI 2018, July 13-19, 2018, Stockholm, Sweden, Jérôme Lan...
2018
-
[9]
Tathagata Chakraborti, Sarath Sreedharan, Sachin Grover, and Subbarao Kamb- hampati. 2019. Plan explanations as model reconciliation–an empirical study. In 2019 14th ACM/IEEE International Conference on Human-Robot Interaction (HRI). Ieee, 258–266
2019
-
[10]
Juntong Chen, Jiang Wu, Jiajing Guo, Vikram Mohanty, Xueming Li, Jorge Pi- azentin Ono, Wenbin He, Liu Ren, and Dongyu Liu. 2025. InterChat: Enhancing generative visual analytics using multimodal interactions. InComputer Graphics Forum, Vol. 44. Wiley Online Library, e70112
2025
-
[11]
Weihao Chen, Chun Yu, Huadong Wang, Zheng Wang, Lichen Yang, Yukun Wang, Weinan Shi, and Yuanchun Shi. 2023. From gap to synergy: Enhancing contextual understanding through human-machine collaboration in personalized systems. InProceedings of the 36th Annual ACM Symposium on User Interface Software and Technology. 1–15
2023
-
[12]
Ziming Cheng, Zhiyuan Huang, Junting Pan, Zhaohui Hou, and Mingjie Zhan
-
[13]
Riccardo Coppola, Luca Ardito, and Marco Torchiano. 2019. Fragility of layout-based and visual GUI test scripts: an assessment study on a hybrid mo- bile application. InProceedings of the 10th ACM SIGSOFT International Work- shop on Automating TEST Case Design, Selection, and Evaluation. ACM, 28–34. doi:10.1145/3340433.3342824
arXiv 2019
-
[14]
Laradji, Manuel Del Verme, Tom Marty, David Vázquez, Nicolas Chapados, and Alexandre Lacoste
Alexandre Drouin, Maxime Gasse, Massimo Caccia, Issam H. Laradji, Manuel Del Verme, Tom Marty, David Vázquez, Nicolas Chapados, and Alexandre Lacoste
-
[15]
Anurag Dwarakanath, Neville Dubash, and Sanjay Podder. 2018. Machines that test Software like Humans.CoRRabs/1809.09455 (2018). arXiv:1809.09455 http://arxiv.org/abs/1809.09455
Pith/arXiv arXiv 2018
-
[16]
Lutfi Eren Erdogan, Nicholas Lee, Sehoon Kim, Suhong Moon, Hiroki Furuta, Gopala Anumanchipalli, Kurt Keutzer, and Amir Gholami. 2025. Plan-and-Act: Improving Planning of Agents for Long-Horizon Tasks. InProceedings of the 42nd International Conference on Machine Learning (Proceedings of Machine Learning Research, Vol. 267), Aarti Singh, Maryam Fazel, Dan...
2025
-
[17]
WorkArena: How Capable are Web Agents at Solving Common Knowledge Work Tasks?Forty-first International Conference on Machine Learning, ICML 2024, Vienna, Austria, July 21-27, 2024235 (2024), 11642–11662
2024
-
[18]
Boyu Gou, Demi Ruohan Wang, Boyuan Zheng, Yanan Xie, Cheng Chang, Yiheng Shu, Huan Sun, and Yu Su. 2025. Navigating the digital world as humans do: Universal visual grounding for gui agents. InInternational Conference on Learning Representations, Vol. 2025. 30851–30883
2025
-
[19]
Nitesh Goyal, Minsuk Chang, and Michael Terry. 2024. Designing for Human- Agent Alignment: Understanding what humans want from their agents. InEx- tended Abstracts of the CHI Conference on Human Factors in Computing Systems. 1–6
2024
-
[20]
KJ Kevin Feng, Kevin Pu, Matt Latzke, Tal August, Pao Siangliulue, Jonathan Bragg, Daniel S Weld, Amy X Zhang, and Joseph Chee Chang. 2026. Cocoa: Co-planning and co-execution with ai agents. InProceedings of the 2026 CHI Conference on Human Factors in Computing Systems. ACM, 1–23
2026
-
[21]
Grigorev, Alexey K
Danil S. Grigorev, Alexey K. Kovalev, and Aleksandr I. Panov. 2025. VerifyLLM: LLM-Based Pre-Execution Task Plan Verification for Robots. InIEEE/RSJ Interna- tional Conference on Intelligent Robots and Systems, IROS 2025, Hangzhou, China, October 19-25, 2025. IEEE, 18489–18496
2025
-
[22]
Xiangwu Guo, Difei Gao, and Mike Zheng Shou. 2025. AUTO-Explorer: Automated Data Collection for GUI Agent.CoRRabs/2511.06417 (2025). arXiv:2511.06417 doi:10.48550/ARXIV.2511.06417
-
[23]
Maria Fernanda Granda, Otto Parra, and Bryan Alba-Sarango. 2021. Towards a Model-Driven Testing Framework for GUI Test Cases Generation from User Stories.. InENASE. 453–460
2021
-
[24]
Hongliang He, Wenlin Yao, Kaixin Ma, Wenhao Yu, Yong Dai, Hongming Zhang, Zhenzhong Lan, and Dong Yu. 2024. Webvoyager: Building an end-to-end web agent with large multimodal models. InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 6864–6890
2024
-
[25]
Theodore D Hellmann and Frank Maurer. 2011. Rule-based exploratory testing of graphical user interfaces. In2011 Agile Conference. IEEE, 107–116
2011
-
[26]
Chao Hao, Shuai Wang, and Kaiwen Zhou. 2025. Uncertainty-aware gui agent: Adaptive perception through component recommendation and human-in-the- loop refinement.arXiv preprint arXiv:2508.04025(2025)
Pith/arXiv arXiv 2025
-
[27]
Eric Horvitz. 1999. Principles of Mixed-Initiative User Interfaces. InProceeding of the CHI ’99 Conference on Human Factors in Computing Systems: The CHI is the Limit, Pittsburgh, PA, USA, May 15-20, 1999, Marian G. Williams and Mark W. Altom (Eds.). ACM, 159–166. doi:10.1145/302979.303030
arXiv 1999
-
[28]
Siyuan Hu, Mingyu Ouyang, Difei Gao, and Mike Zheng Shou. 2024. The dawn of gui agent: A preliminary case study with claude 3.5 computer use.arXiv preprint arXiv:2411.10323(2024)
Pith/arXiv arXiv 2024
-
[29]
Peter Hofmann, Caroline Samp, and Nils Urbach. 2020. Robotic process automa- tion.Electronic markets30, 1 (2020), 99–106
2020
-
[30]
Zhiyuan Huang, Ziming Cheng, Junting Pan, Zhaohui Hou, and Mingjie Zhan
-
[31]
Faria Huq, Zora Zhiruo Wang, Frank F Xu, Tianyue Ou, Shuyan Zhou, Jeffrey P Bigham, and Graham Neubig. 2025. Cowpilot: a framework for autonomous and human-agent collaborative web navigation. InProceedings of the 2025 Confer- ence of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (System D...
2025
-
[32]
Wenyue Hua, Mengting Wan, Jagannath Vadrevu, Ryan Nadel, Yongfeng Zhang, and Chi Wang. 2025. Interactive speculative planning: Enhance agent efficiency through co-design of system and user interface. InInternational Conference on Learning Representations, Vol. 2025. 14256–14283
2025
-
[33]
Arushi Jain, Shubham Paliwal, Monika Sharma, Lovekesh Vig, and Gau- tam Shroff. 2024. SmartFlow: Robotic Process Automation using LLMs. arXiv:2405.12842 [cs.RO] https://arxiv.org/abs/2405.12842
Pith/arXiv arXiv 2024
-
[34]
InProceedings of the Computer Vision and Pattern Recognition Conference
Spiritsight agent: Advanced gui agent with one look. InProceedings of the Computer Vision and Pattern Recognition Conference. 29490–29500
-
[35]
Linjia Kang, Zhimin Wang, Yongkang Zhang, Duo Wu, Jinghe Wang, Ming Ma, Haopeng Yan, and Zhi Wang. 2026. Learning with Challenges: Adaptive Difficulty-Aware Data Generation for Mobile GUI Agent Training.arXiv preprint arXiv:2601.22781(2026)
arXiv 2026
-
[36]
Alayt Issak, Jeba Rezwana, and Casper Harteveld. 2025. MOSAAIC: Managing Optimization towards Shared Autonomy, Authority, and Initiative in Co-creation. arXiv:2505.11481 [cs.AI] https://arxiv.org/abs/2505.11481
Pith/arXiv arXiv 2025
-
[37]
Jing Yu Koh, Robert Lo, Lawrence Jang, Vikram Duvvur, Ming Lim, Po-Yu Huang, Graham Neubig, Shuyan Zhou, Russ Salakhutdinov, and Daniel Fried. 2024. Vi- sualwebarena: Evaluating multimodal agents on realistic visual web tasks. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 881–905
2024
-
[38]
Allison Sihan Jia, Daniel Huang, Nikhil Vytla, Nirvika Choudhury, Shayak Sen, John C. Mitchell, and Anupam Datta. 2025. What Is Your Agent’s GPA? A Frame- work for Evaluating Agent Goal-Plan-Action Alignment.CoRRabs/2510.08847 (2025)
arXiv 2025
-
[39]
Kevin Qinghong Lin, Linjie Li, Difei Gao, Zhengyuan Yang, Shiwei Wu, Zechen Bai, Stan Weixian Lei, Lijuan Wang, and Mike Zheng Shou. 2025. Showui: One vision-language-action model for gui visual agent. InProceedings of the Computer Vision and Pattern Recognition Conference. 19498–19508
2025
-
[40]
Anjali Khurana, Xiaotian Su, April Yi Wang, and Parmit K. Chilana. 2025. Do It For Me vs. Do It With Me: Investigating User Perceptions of Different Paradigms of Automation in Copilots for Feature-Rich Software. InProceedings of the 2025 CHI Conference on Human Factors in Computing Systems, CHI 2025, Yokohama- Japan, 26 April 2025- 1 May 2025, Naomi Yamas...
arXiv 2025
-
[41]
Yadong Lu, Jianwei Yang, Yelong Shen, and Ahmed Hassan Awadallah. 2025. OmniParser for Pure Vision Based GUI Agent. https://openreview.net/forum? id=C6hUK6Q1Pi
2025
-
[42]
Toby Jia-Jun Li, Amos Azaria, and Brad A Myers. 2017. SUGILITE: creating multimodal smartphone automation by demonstration. InProceedings of the 2017 CHI conference on human factors in computing systems. 6038–6049
2017
-
[43]
Jordan Madden, Moxanki Bhavsar, Lhamo Dorje, and Xiaohua Li. 2024. Ro- bustness of Practical Perceptual Hashing Algorithms to Hash-Evasion and Hash- Inversion Attacks. InThe Third Workshop on New Frontiers in Adversarial Machine Learning. https://openreview.net/forum?id=hraOxsleRl
2024
-
[44]
Tao Long, Xuanming Zhang, Sitong Wang, Zhou Yu, and Lydia B Chilton. 2025. DoubleAgents: Interactive Simulations for Alignment in Agentic AI.arXiv preprint arXiv:2509.12626(2025)
Pith/arXiv arXiv 2025
-
[45]
Hussein Mozannar, Gagan Bansal, Cheng Tan, Adam Fourney, Victor Dibia, Jingya Chen, Jack Gerrits, Tyler Payne, Matheus Kunzler Maldaner, Madeleine Grunde-McLaughlin, et al. 2025. Magentic-ui: Towards human-in-the-loop agen- tic systems.arXiv preprint arXiv:2507.22358(2025)
Pith/arXiv arXiv 2025
-
[46]
Shang Ma, Xusheng Xiao, and Yanfang Ye. 2025. Agent+ P: Guiding UI Agents via Symbolic Planning.arXiv preprint arXiv:2510.06042(2025)
arXiv 2025
-
[47]
Rodriguez, Montek Kalsi, Nicolas Chapados, M
Shravan Nayak, Xiangru Jian, Kevin Qinghong Lin, Juan A. Rodriguez, Montek Kalsi, Nicolas Chapados, M. Tamer Özsu, Aishwarya Agrawal, David Vazquez, Christopher Pal, Perouz Taslakian, Spandana Gella, and Sai Rajeswar. 2025. UI- Vision: A Desktop-centric GUI Benchmark for Visual Perception and Interaction. InProceedings of the 42nd International Conference...
2025
-
[48]
Valérie Maquil, Dimitra Anastasiou, Hoorieh Afkari, Adrien Coppens, Johannes Hermen, and Lou Schwartz. 2023. Establishing Awareness through Pointing Gestures during Collaborative Decision-Making in a Wall-Display Environment. InExtended Abstracts of the 2023 CHI Conference on Human Factors in Computing Systems, CHI EA 2023, Hamburg, Germany, April 23-28, ...
2023
-
[49]
Yi-Hao Peng, Dingzeyu Li, Jeffrey P. Bigham, and Amy Pavel. 2025. Morae: Proactively Pausing UI Agents for User Choices. InProceedings of the 38th Annual ACM Symposium on User Interface Software and Technology, UIST 2025, Busan, Korea, 28 September 2025 - 1 October 2025, Andrea Bianchi, Elena L. Glassman, Wendy E. Mackay, Shengdong Zhao, Jeeeun Kim, and I...
arXiv 2025
-
[50]
Arpit Narechania, Shunan Guo, Eunyee Koh, Alex Endert, and Jane Hoffswell
-
[51]
Utilizing Provenance as an Attribute for Visual Data Analysis: A Design Probe With ProvenanceLens.IEEE Trans. Vis. Comput. Graph.31, 10 (2025), 8452–8465
2025
-
[52]
Ju Qian, Zhengyu Shang, Shuoyan Yan, Yan Wang, and Lin Chen. 2020. Roscript: a visual script driven truly non-intrusive robotic testing system for touch screen applications. InProceedings of the ACM/IEEE 42nd International Conference on Software Engineering. 297–308
2020
-
[53]
Mehrab Tanjim, Nesreen K
Dang Nguyen, Jian Chen, Yu Wang, Gang Wu, Namyong Park, Zhengmian Hu, Hanjia Lyu, Junda Wu, Ryan Aponte, Yu Xia, Xintong Li, Jing Shi, Hongjie Chen, Viet Dac Lai, Zhouhang Xie, Sungchul Kim, Ruiyi Zhang, Tong Yu, Md. Mehrab Tanjim, Nesreen K. Ahmed, Puneet Mathur, Seunghyun Yoon, Lina Yao, Branislav Kveton, Jihyung Kil, Thien Huu Nguyen, Trung Bui, Tianyi...
2025
-
[54]
Stephan Rabanser, Sayash Kapoor, Peter Kirgis, Kangheng Liu, Saiteja Utpala, and Arvind Narayanan. 2026. Towards a Science of AI Agent Reliability.arXiv preprint arXiv:2602.16666(2026)
Pith/arXiv arXiv 2026
-
[55]
Christopher Potts and Moritz Sudhof. 2026. Invisible failures in human-AI inter- actions.arXiv preprint arXiv:2603.15423(2026)
Pith/arXiv arXiv 2026
-
[56]
Petr Průcha, Michaela Matoušková, and Jan Strnad. 2025. Are LLM Agents the New RPA? A Comparative Study with RPA Across Enterprise Workflows.arXiv preprint arXiv:2509.04198(2025)
Pith/arXiv arXiv 2025
-
[57]
Minjie Shen, Yanshu Li, Lulu Chen, and Qikai Yang. 2025. From mind to ma- chine: The rise of manus ai as a fully autonomous digital agent.arXiv preprint arXiv:2505.02024(2025)
Pith/arXiv arXiv 2025
-
[58]
Yujia Qin, Yining Ye, Junjie Fang, Haoming Wang, Shihao Liang, Shizuo Tian, Junda Zhang, Jiahao Li, Yunxin Li, Shijue Huang, Wanjun Zhong, Kuanye Li, Jiale Yang, Yu Miao, Woyu Lin, Longxiang Liu, Xu Jiang, Qianli Ma, Jingyu Li, Xiaojun Xiao, Kai Cai, Chuang Li, Yaowei Zheng, Chaolin Jin, Chen Li, Xiao Zhou, Minchao Wang, Haoli Chen, Zhaojian Li, Haihua Ya...
-
[59]
When to Hand Off, When to Work Together
Kihoon Son, Hyewon Lee, DaEun Choi, Yoonsu Kim, Tae Soo Kim, Yoonjoo Lee, John Joon Young Chung, HyunJoon Jung, and Juho Kim. 2026. " When to Hand Off, When to Work Together": Expanding Human-Agent Co-Creative Collaboration through Concurrent Interaction.arXiv preprint arXiv:2603.02050 (2026)
Pith/arXiv arXiv 2026
-
[60]
Peter Shaw, Mandar Joshi, James Cohan, Jonathan Berant, Panupong Pasupat, Hexiang Hu, Urvashi Khandelwal, Kenton Lee, and Kristina N Toutanova. 2023. From pixels to ui actions: Learning to follow instructions via graphical user interfaces.Advances in Neural Information Processing Systems36 (2023), 34354– 34370
2023
-
[61]
Erfan Shayegani, Keegan Hines, Yue Dong, Nael Abu-Ghazaleh, Roman Lutz, Spencer Whitehead, Vidhisha Balachandran, Besmira Nushi, and Vibhav Vineet
-
[62]
Qiushi Sun, Kanzhi Cheng, Zichen Ding, Chuanyang Jin, Yian Wang, Fangzhi Xu, Zhenyu Wu, Chengyou Jia, Liheng Chen, Zhoumianze Liu, Ben Kao, Guohao Li, Junxian He, Yu Qiao, and Zhiyong Wu. 2025. OS-Genesis: Automating GUI Agent Trajectory Construction via Reverse Task Synthesis. InProceedings of the 63rd Annual Meeting of the Association for Computational ...
2025
-
[63]
Fei Tang, Haolei Xu, Hang Zhang, Siqi Chen, Xingyu Wu, Yongliang Shen, Wenqi Zhang, Guiyang Hou, Zeqi Tan, Yuchen Yan, et al. 2025. A survey on (m) llm-based gui agents.arXiv preprint arXiv:2504.13865(2025)
Pith/arXiv arXiv 2025
-
[64]
Aleksandar Shtedritski, Christian Rupprecht, and Andrea Vedaldi. 2023. What does CLIP know about a red circle? Visual prompt engineering for VLMs. In IEEE/CVF International Conference on Computer Vision, ICCV 2023, Paris, France, October 1-6, 2023. IEEE, 11953–11963. doi:10.1109/ICCV51070.2023.01101
arXiv 2023
-
[65]
Yuanrong Tang, Huiling Peng, Bingxi Zhao, Hengyang Ding, Hanchao Song, Tianhong Wang, Chen Zhong, and Jiangtao Gong. 2026. Human Tool: An MCP- Style Framework for Human-Agent Collaboration.arXiv preprint arXiv:2602.12953 (2026)
arXiv 2026
-
[66]
Yunpeng Song, Yiheng Bian, Yongtao Tang, Guiyu Ma, and Zhongmin Cai. 2024. Visiontasker: Mobile task automation using vision based ui understanding and llm task planning. InProceedings of the 37th Annual ACM Symposium on User Interface Software and Technology. 1–17
2024
-
[67]
Zihe Song, S. M. Hasan Mansur, Ravishka Rathnasuriya, Yumna Fatima, Wei Yang, Kevin Moran, and Wing Lam. 2025. Can You Mimic Me? Exploring the Use of Android Record & Replay Tools in Debugging. In12th IEEE/ACM In- ternational Conference on Mobile Software Engineering and Systems, MOBILE- Soft@ICSE 2025, Ottawa, ON, Canada, April 27-28, 2025. IEEE, 32–43. ...
arXiv 2025
-
[68]
Vanshika Vats, Marzia Binta Nizam, Minghao Liu, Ziyuan Wang, Richard Ho, Mohnish Sai Prasad, Vincent Titterton, Sai Venkat Malreddy, Riya Aggarwal, Yan- wen Xu, et al. 2024. A Survey on Human-AI Collaboration with Large Foundation Models.arXiv preprint arXiv:2403.04931(2024)
Pith/arXiv arXiv 2024
-
[69]
Bowen Wang, Xinyuan Wang, Jiaqi Deng, Tianbao Xie, Ryan Li, Yanzhe Zhang, Junli Wang, Dunjie Lu, Zicheng Gong, Gavin Li, Toh Jing Hua, Wei-Lin Chiang, Ion Stoica, Diyi Yang, Yu Su, Yi Zhang, Zhiguo Wang, Victor Zhong, and Tao Yu
-
[70]
Jingyu Tang, Chaoran Chen, Jiawen Li, Zhiping Zhang, Bingcan Guo, Ibrahim Khalilov, Simret Araya Gebreegziabher, Bingsheng Yao, Dakuo Wang, Yanfang Ye, et al. 2026. Dark patterns meet gui agents: Llm agent susceptibility to manip- ulative interfaces and the role of human oversight.Proceedings of the 2026 CHI Conference on Human Factors in Computing System...
2026
-
[71]
Hui Wei, Zihao Zhang, Shenghua He, Tian Xia, Shijia Pan, and Fei Liu. 2025. PlanGenLLMs: A Modern Survey of LLM Planning Capabilities. InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2025, Vienna, Austria, July 27 - August 1, 2025, Wanxiang Che, Joyce Nabende, Ekaterina Shutova, and Mo...
2025
-
[72]
Kimi Team, Tongtong Bai, Yifan Bai, Yiping Bao, SH Cai, Yuan Cao, Y Charles, HS Che, Cheng Chen, Guanduo Chen, et al . 2026. Kimi K2. 5: Visual Agentic Intelligence.arXiv preprint arXiv:2602.02276(2026)
Pith/arXiv arXiv 2026
-
[73]
Ching-Yi Tsai, Nicole Tacconi, Andrew D Wilson, and Parastoo Abtahi. 2026. Uncertain Pointer: Situated Feedforward Visualizations for Ambiguity-Aware AR Target Selection.arXiv preprint arXiv:2602.13433(2026)
arXiv 2026
-
[74]
Judith Wewerka and Manfred Reichert. 2020. Robotic Process Automation - A Systematic Literature Review and Assessment Framework.CoRRabs/2012.11951 (2020). arXiv:2012.11951 https://arxiv.org/abs/2012.11951
Pith/arXiv arXiv 2020
-
[75]
Rossi, Ruiyi Zhang, Subrata Mitra, Dimitris N
Junda Wu, Zhehao Zhang, Yu Xia, Xintong Li, Zhaoyang Xia, Aaron Chang, Tong Yu, Sungchul Kim, Ryan A. Rossi, Ruiyi Zhang, Subrata Mitra, Dimitris N. Metaxas, Lina Yao, Jingbo Shang, and Julian J. McAuley. 2024. Visual Prompting in Multimodal Large Language Models: A Survey.CoRRabs/2409.15310 (2024). arXiv:2409.15310 doi:10.48550/ARXIV.2409.15310
-
[76]
InThe Fourteenth International Conference on Learning Representations
Computer Agent Arena: Toward Human-Centric Evaluation and Analysis of Computer-Use Agents. InThe Fourteenth International Conference on Learning Representations. https://openreview.net/forum?id=3x4SDbXbgl
-
[77]
Bovik, Hamid R
Zhou Wang, Alan C. Bovik, Hamid R. Sheikh, and Eero P. Simoncelli. 2004. Image quality assessment: from error visibility to structural similarity.IEEE Trans. Image Process.13, 4 (2004), 600–612
2004
-
[78]
Tianbao Xie, Danyang Zhang, Jixuan Chen, Xiaochuan Li, Siheng Zhao, Ruisheng Cao, Toh Jing Hua, Zhoujun Cheng, Dongchan Shin, Fangyu Lei, Yitao Liu, Yiheng Xu, Shuyan Zhou, Silvio Savarese, Caiming Xiong, Victor Zhong, and Tao Yu
-
[79]
Hao Wen, Yuanchun Li, Guohong Liu, Shanhui Zhao, Tao Yu, Toby Jia-Jun Li, Shiqi Jiang, Yunhao Liu, Yaqin Zhang, and Yunxin Liu. 2024. Autodroid: Llm-powered task automation in android. InProceedings of the 30th Annual International Con- ference on Mobile Computing and Networking. 543–557
2024
-
[80]
Zhen Wen, Luoxuan Weng, Yinghao Tang, Runjin Zhang, Yuxin Liu, Bo Pan, Minfeng Zhu, and Wei Chen. 2025. Exploring Multimodal Prompt for Visual- ization Authoring with Large Language Models.CoRRabs/2504.13700 (2025). doi:10.48550/ARXIV.2504.13700
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.