Pith. sign in

REVIEW 3 major objections 4 minor 102 references

Plover shows that many GUI-agent failures become repairable when the task plan stays visible and corrections stay localized, with 23 of 26 benchmark failures improved by mixed-initiative interaction.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-01 23:52 UTC pith:KG7UY24M

load-bearing objection Plover is a credible systems paper with an honest upper-bound recovery result; just don't let the abstract sell the 88% as proof that plan visibility is what rescues failures. the 3 major comments →

arxiv 2607.15193 v1 pith:KG7UY24M submitted 2026-07-16 cs.AI

Plover: Steering GUI Agents through Plan-Centric Interaction

classification cs.AI
keywords GUI agentsmixed-initiative systemshuman-AI interactionplan-centric interactionintelligent replanningmultimodal annotationfailure recoveryGUI automation
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The paper introduces Plover, a plan-centric GUI automation system that externalizes task plans as persistent, inspectable, and revisable artifacts. Its central claim is that many autonomous GUI-agent failures are not terminal errors but localized breakdowns in grounding, state interpretation, or execution continuity that become structurally recoverable when users can see the plan, intervene surgically, and preserve completed work. In a repair study over 26 autonomous non-success cases from the OSWorld-Verified benchmark, mixed-initiative interaction with an expert user improved 23 cases, converting 17 to complete success and 6 to partial success, with an average of 2.04 interventions per task and no regressions. The paper argues that reliable GUI automation should be treated as an interaction problem as much as a modeling problem: making replanning visible and localized turns silent drift into a collaborative, repairable process.

Core claim

Plover's core claim is that GUI-agent failures are structurally recoverable when plans are externalized and repairs are localized. The system keeps a versioned plan artifact, separates immutable executed steps from an editable pending suffix, and supports user-driven interventions (natural-language guidance, plan edits, and screenshot annotations) plus system-driven replanning triggered by non-progress detection. In the benchmark repair analysis, 26 autonomous failures were re-run in a mixed-initiative setting; 23 improved (17 complete successes, 6 partial), only 3 remained failures, and all 10 autonomous partial successes became complete successes. The paper also characterizes which failure

What carries the argument

The central mechanism is the persistent plan artifact, a shared state representation with the invariant that executed steps are immutable and only the pending suffix can be revised (C_{t+1}=C_t). On top of this, Plover implements Intelligent Replanning in two modes: User-Driven IR, where natural-language guidance or multimodal annotations (strokes, shapes, text overlays captured as primitives with a bounding box) generate localized plan proposals, and System-Driven IR, where a watchdog detects behavioral loop repetition and visual non-progress (via dHash Hamming distance) and injects a structured failure message that prompts the model to propose a recovery step with rationale. The versioned

Load-bearing premise

The 88% recovery figure depends on an expert user who already knows what went wrong and what the correct target is; if ordinary users cannot detect drift or formulate correct localized corrections, the recoverability claim may not transfer to practice.

What would settle it

Run the same 26 benchmark failures with naive participants who are not told the failure cause, providing only the Plover interface as the intervention channel. If the mixed-initiative recovery rate falls to near the autonomous baseline (no significant improvement over the 0% success on these originally failed tasks), the claim that failures are structurally repairable through visible plans is not supported for realistic users.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • If recoverability holds beyond the expert setting, GUI agents can be deployed in long-horizon, high-friction workflows with a human steering loop instead of requiring near-perfect autonomy.
  • The plan-invariant design means corrections preserve executed history, so each intervention is cheaper and less disruptive than re-prompting or restarting the whole task.
  • System-driven non-progress detection (repeated semantic actions plus visual stability) can act as a reusable watchdog that catches drift before it propagates, independent of the specific planner or executor.
  • The failure taxonomy suggests that perception errors and state misinterpretations are cheaply repairable with language or annotations, while compound failures require catching the initial planning error earlier.
  • Exposing plans as versioned, diffable artifacts provides a natural audit trail for when and why an agent's behavior changed, which can support verification and post-hoc analysis.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • A natural next test, which the paper does not run, is a study with non-expert users who are not told the failure cause: if recovery rates drop to near the autonomous baseline, the 'structurally recoverable' claim would need to be re-scoped from an upper bound to a property that depends on user diagnostic skill.
  • The plan artifact as a coordination protocol could generalize beyond GUI automation to other long-horizon agent domains (e.g., data-cleaning pipelines or robotics task plans) where partial progress is valuable and corrections must be localized.
  • The System-Driven IR watchdog (behavioral repetition + perceptual-hash stability) is a concrete, model-agnostic component that could be extracted and benchmarked on its own to measure how many agent stalls it catches before a human would notice.
  • The paper itself flags that visible plans may inflate user confidence; an empirical study measuring whether users over-accept plan proposals when the system looks confident would be a direct extension.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. The paper introduces Plover, a plan-centric GUI automation system that externalizes task plans as persistent, editable artifacts and supports mixed-initiative repair through natural-language guidance, multimodal annotation, plan edits, and system-driven replanning. The authors report a formative study with six participants, a benchmark repair study on 38 OSWorld-Verified tasks (26 autonomous non-successes re-run with expert interventions, yielding 23 improved, 17 complete successes, 6 partial successes, 3 failures), and a scenario-based stability analysis with trajectory-derived prompts. The central claim is that many GUI-agent failures are structurally repairable when plans remain visible and interventions are localized, and that explicit replanning improves transparency, controllability, and adaptability.

Significance. If the central claim were fully established, the paper would make a useful contribution to human-agent interaction for GUI automation: it proposes a concrete design space (persistent plans as coordination artifacts, localized repair, visible replanning) and provides a failure taxonomy that could guide future interface design. The paper is honest in labeling the benchmark repair study as an upper bound on plan-centric recoverability, and the appendices contain substantial implementation and evaluation detail, including per-task results and a formative study summary. These are real strengths. However, the main empirical evidence does not currently separate the effect of the plan-centric interface from the effect of an expert oracle user, so the causal design conclusions (DG1–DG5, 'explicit replanning helps') are not yet established by the data.

major comments (3)
  1. [Section 5.1, Table 1; Abstract] The load-bearing empirical claim—23/26 non-success cases improved, 88% recovery—is measured with the first author, who knows each failure cause and the correct target, supplying all interventions. There is no control condition that strips away the plan-centric affordances (plan panel, plan editing, annotation, visible replanning) while keeping the same underlying agent and the same expert. As written, the 88% figure is an upper bound on recoverability by an informed oracle, not evidence that plan visibility and localized repair cause the improvement. The paper's own limitation statement in Section 6 ('our design assumes users generally know the intended path well enough to correct the agent') does not carry through to the abstract/conclusion, which assert a causal role for visible plans. Add an ablation or control (e.g., same expert corrections issued as text prompts to the same base age
  2. [Section 5.2, 'Trajectory Sampling and Prompt Reconstruction' and Table 2] The scenario stability analysis synthesizes task prompts from sampled interaction trajectories and then compares the replayed plans and final states against the same trajectories. This creates a circularity: the prompt is derived from the reference trajectory, so plan alignment metrics (coverage 0.62, order 0.41, actionability 0.97) partly measure reconstruction from a trajectory-derived instruction rather than general plan quality. The browser-vs-desktop visual fidelity differences are still informative, but the plan alignment results should be presented as a property of this reverse-synthesis setup, not as evidence about Plover's planning under independently authored user instructions. A sanity check with manually authored prompts, or a clear caveat, is needed before these numbers are used to support the 'opportunity for users to inspect and correct' argument.
  3. [Section 5.3, 'Characterizing Repairable Failures'] The failure-mode analysis is based on the same 26 cases and the same expert interventions. Counts such as 'Execution Drift appeared in 46% (n=12/26)' and 'NL Guidance resolved 11 cases' are reported without any uncertainty or sensitivity analysis. With n=26 and intervention choices made by a single expert who already knows the failure causes, small counts can easily flip; the recovered vs. unrecovered distinction is not robust enough to support the strong claim that 'compound failures' are fundamentally harder. At minimum, report bootstrap or exact binomial confidence intervals and clarify that all recovery counts are conditional on the expert's choice of intervention.
minor comments (4)
  1. [Appendix C, Algorithm 1] System-Driven IR relies on hardcoded thresholds (REPEAT_SEQ_L3_R3, dHash Hamming distance > 40) with no sensitivity analysis. Since Section 5.3 attributes 9 successful recoveries to System-Driven IR, the threshold choices can materially affect the results; report how varying them changes detection and downstream recovery.
  2. [Section 5.1, 'no regressions'] The statement 'no regressions were observed' only covers the 26 autonomous non-success cases; it does not address whether the mixed-initiative interaction could degrade autonomous successes, since those were not re-run in the mixed-initiative condition. Please state this scope explicitly.
  3. [Table 1 and Table 2] Table 1's 'Improv. Rate' is not formally defined; for Multi-App, 6S+2P out of 10 corresponds to 80%, but the reader must infer the denominator. In Table 2, the 'Overall Average' row for MSE (939.57) is the mean of scenario averages, not the mean over all trials; clarify the aggregation.
  4. [Throughout] The paper uses 'Conference’17' in the ACM reference format and several placeholder-style citations (e.g., the DOI is 'XXXXXXX.XXXXXXX'). Please update the formatting to the final venue style and correct minor typographical issues such as the 'MI (a)' label in Figure 5.

Circularity Check

0 steps flagged

No significant circularity: the 88% recoverability result is an externally grounded upper-bound measurement with the expert-oracle caveat explicitly acknowledged; remaining concerns are validity limitations, not circular reductions.

full rationale

Plover's central empirical claim is an upper-bound measurement on the external OSWorld-Verified benchmark, not a quantity derived from a fitted parameter, an ansatz, or a load-bearing self-citation. The paper states the intervention condition explicitly: 'This setup establishes an upper bound on plan-centric recoverability rather than typical user performance' (Section 5.1), and the corresponding limitation is restated in Section 6 ('our design assumes users generally know the intended path well enough to correct the agent when needed'). Because the expert interventions are acknowledged as ideal rather than typical, the 88% recovery figure is an honest conditional result, not a hidden assumption presented as a finding. The scenario-based stability analysis (Section 5.2) synthesizes prompts from reference trajectories and then compares generated plans against those trajectories, which introduces non-independence; however the reported alignment metrics are moderate (coverage 0.62, order 0.41), so the result is not forced by construction. The only self-citation is a related-work mention ([10], multimodal interaction) and is not load-bearing. No uniqueness theorem, no ansatz smuggled via citation, and no equation reduces the conclusion to its input. The absence of a chat-only or invisible-plan control is a causal-identification limitation, but under the circularity criteria it does not constitute circularity.

Axiom & Free-Parameter Ledger

4 free parameters · 6 axioms · 0 invented entities

No invented physical or formal entities. The central numerical result is an empirical recovery rate, not a derivation. The main free quantities are hand-set detection/evaluation thresholds and manual trajectory curation, none of which are validated independently.

free parameters (4)
  • Visual non-progress dHash threshold τ = 40 (Hamming distance)
    Algorithm 1 line 14: if any Hamming distance > 40, execution is considered non-static and not stuck. This value is hand-set, not validated; changing it changes which failures trigger System-Driven IR.
  • Repetition pattern REPEAT_SEQ_L3_R3 = repeated action subsequence of length 3 seen 3 times
    Appendix C, Figure 8: used to detect behavioral loops; hand-set. Higher/lower values alter how many drift cases are classified as stuck.
  • Per-scenario similarity thresholds (SSIM/MSE/dHash) = e.g., Firefox Fillable Form High: SSIM≥0.98, MSE≤100, dHash≤2; LibreOffice thresholds differ
    Table 8: thresholds defining High/Partial/Low replay match are set per scenario, so the visual-fidelity conclusions are partly products of chosen cutoffs.
  • Selected trajectories per scenario = 5 per scenario from 100 exploration trials, after manual verification
    Section 5.2: manual curation of 'diverse' trajectories is a subjective filter; it determines which prompts and reference plans are used for both replay and plan alignment.
axioms (6)
  • domain assumption Claude 4.5 Sonnet via computer-use can decompose tasks into deterministic UI instructions and ground them in the executor.
    Sections 4.4-4.5: if the planner produces non-executable steps or the executor cannot act on them, both autonomous execution and repair fail regardless of plan visibility.
  • domain assumption Image-similarity metrics (SSIM, MSE, dHash) are valid proxies for GUI execution stability and task-state equivalence.
    Section 5.2 uses these to conclude browser workflows are stable and desktop workflows are not; no validation that image similarity corresponds to task success/failure.
  • domain assumption The 38 OSWorld-Verified tasks that failed for Claude 4.5 Sonnet are representative of GUI-agent failures in general.
    Section 5.1 selects failure cases from a single model/benchmark; the failure distribution (e.g., compound failures 4/26) may not generalize across agents and environments.
  • domain assumption An expert who knows failure causes can stand in for a typical user to establish 'structural recoverability'.
    Section 5.1: 'This setup establishes an upper bound on plan-centric recoverability rather than typical user performance.' The entire headline rate depends on this premise.
  • domain assumption GPA phase taxonomy and trajectory-derived phases are a meaningful ground truth for plan alignment.
    Section 5.2 maps plans and trajectories to a shared interface-agnostic taxonomy; coverage/order/redundancy/actionability are computed relative to this lossy mapping.
  • domain assumption The invariant C_{t+1}=C_t (completed steps immutable) is a sound design constraint for repair.
    Section 4.4: preserving completed history assumes completed steps are correct; if an early error was committed to history, the invariant freezes it and can prevent full recovery, as seen in compound failures.

pith-pipeline@v1.3.0-alltime-deepseek · 28875 in / 14302 out tokens · 118705 ms · 2026-08-01T23:52:27.268038+00:00 · methodology

0 comments
read the original abstract

Graphical user interface (GUI) automation remains challenging in real-world environments, where dynamic layouts, unexpected dialogs, and evolving interface states can cause autonomous agents to drift from user intent. Recent vision-based multimodal agents improve flexibility by operating directly over screenshots and natural language instructions, but planning and adaptation often remain internal, limiting users' ability to inspect, supervise, or correct system behavior. We present Plover, a plan-centric vision-based GUI automation system that externalizes task plans and replanning as persistent, inspectable, and revisable artifacts. Through a planner--executor architecture, Plover supports explicit supervision of evolving execution, localized correction through editable plans, natural-language guidance, and screenshot-grounded interventions, while preserving prior progress during repair. A formative study with six participants informed the interaction design. We then evaluate Plover through benchmark failure-case repair and scenario-based workflow analyses. Our results show that many autonomous GUI-agent failures are structurally repairable when plans remain visible and interventions are localized, and that explicit replanning helps make GUI automation more transparent, controllable, and adaptable.

Figures

Figures reproduced from arXiv: 2607.15193 by Dongyu Liu, Jiajing Guo, Jorge Piazentin Ono, Liu Ren, Madhumitha Venkatesan, Shicheng Wen.

Figure 1
Figure 1. Figure 1: The Plover system architecture. it was making progress, and apply corrections without losing prior work. We synthesize these findings into three design questions: Q1) How can systems maintain alignment between evolving user intent and execution across multi-step workflows? Q2) How can in￾terfaces support precise, visually grounded correction when execution deviates? Q3) How can replanning be surfaced as a … view at source ↗
Figure 2
Figure 2. Figure 2: Plover Agentic Interface. (a) System Status Bar shows execution phase and plan state. (b) Planner Chat supports [PITH_FULL_IMAGE:figures/full_fig_p005_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Provenance Bar visualizing branching plan revi [PITH_FULL_IMAGE:figures/full_fig_p006_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: Multimodal Annotation workflow for resolving [PITH_FULL_IMAGE:figures/full_fig_p006_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: Representative failure archetypes and recovery pathways in [PITH_FULL_IMAGE:figures/full_fig_p008_5.png] view at source ↗
Figure 6
Figure 6. Figure 6: Early Plover prototype used in Formative Study. The interface supported prompt authoring, plan inspection/editing, and execution monitoring. Observations from this prototype informed the redesign of Plover (DG1–DG5) [PITH_FULL_IMAGE:figures/full_fig_p016_6.png] view at source ↗
Figure 7
Figure 7. Figure 7: Early Annotation Panel. Users could provide visual annotations during execution; feedback from this component [PITH_FULL_IMAGE:figures/full_fig_p016_7.png] view at source ↗
Figure 9
Figure 9. Figure 9: Proposal-mode prompt suffix used during System-Driven IR. When execution drift is detected, this suffix is appended to the system prompt to constrain the model to produce a structured recovery proposal consisting of a short next-step summary and rationale. D Implementation Details This section provides additional technical details regarding the prompt engineering, technical implementation and cross-platfor… view at source ↗
Figure 8
Figure 8. Figure 8: Structured failure message injected into the con [PITH_FULL_IMAGE:figures/full_fig_p018_8.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

102 extracted references · 2 canonical work pages

  1. [1]

    Mohamed Aghzal, Gregory J Stein, and Ziyu Yao. 2026. Why do LLM-based web agents fail? A hierarchical planning perspective. InProceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 32157–32180

  2. [2]

    Saleema Amershi, Dan Weld, Mihaela Vorvoreanu, Adam Fourney, Besmira Nushi, Penny Collisson, Jina Suh, Shamsi Iqbal, Paul N Bennett, Kori Inkpen, et al. 2019. Guidelines for human-AI interaction. InProceedings of the 2019 CHI Conference on Human Factors in Computing Systems. 1–13

  3. [3]

    Nishant Balepur, Matthew Shu, Yoo Yeon Sung, Seraphina Goldfarb-Tarrant, Shi Feng, Fumeng Yang, Rachel Rudinger, and Jordan Lee Boyd-Graber. 2025. A Good Plan is Hard to Find: Aligning Models with Preferences is Misaligned with What Helps Users. InProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, EMNLP 2025, Suzhou, C...

  4. [4]

    Tanvir Bhathal and Asanshay Gupta. 2025. Websight: A vision-first architecture for robust web agents.arXiv preprint arXiv:2508.16987(2025)

  5. [5]

    Léo Boisvert, Megh Thakkar, Maxime Gasse, Massimo Caccia, Thibault L De Chezelles, Quentin Cappart, Nicolas Chapados, Alexandre Lacoste, and Alexandre Drouin. 2024. Workarena++: Towards compositional planning and reasoning-based common knowledge work tasks.Advances in Neural Information Processing Systems37 (2024), 5996–6051

  6. [6]

    Sacha Brisset, Romain Rouvoy, Lionel Seinturier, and Renaud Pawlak. 2022. Er- ratum: Leveraging Flexible Tree Matching to repair broken locators in web automation scripts.Inf. Softw. Technol.144 (2022), 106754. doi:10.1016/J.INFSOF. 2021.106754

  7. [7]

    Meyer, Yuning Chai, Dennis Park, and Yong Jae Lee

    Mu Cai, Haotian Liu, Siva Karthik Mustikovela, Gregory P. Meyer, Yuning Chai, Dennis Park, and Yong Jae Lee. 2024. ViP-LLaVA: Making Large Multimodal Models Understand Arbitrary Visual Prompts. InIEEE/CVF Conference on Com- puter Vision and Pattern Recognition, CVPR 2024, Seattle, W A, USA, June 16-22,

  8. [8]

    Fadnis, Kartik Talamadupula, Mishal Dholakia, Biplav Srivastava, Jeffrey O

    Tathagata Chakraborti, Kshitij P. Fadnis, Kartik Talamadupula, Mishal Dholakia, Biplav Srivastava, Jeffrey O. Kephart, and Rachel K. E. Bellamy. 2018. Visualiza- tions for an Explainable Planning Agent. InProceedings of the Twenty-Seventh International Joint Conference on Artificial Intelligence, IJCAI 2018, July 13-19, 2018, Stockholm, Sweden, Jérôme Lan...

  9. [9]

    Tathagata Chakraborti, Sarath Sreedharan, Sachin Grover, and Subbarao Kamb- hampati. 2019. Plan explanations as model reconciliation–an empirical study. In 2019 14th ACM/IEEE International Conference on Human-Robot Interaction (HRI). Ieee, 258–266

  10. [10]

    Juntong Chen, Jiang Wu, Jiajing Guo, Vikram Mohanty, Xueming Li, Jorge Pi- azentin Ono, Wenbin He, Liu Ren, and Dongyu Liu. 2025. InterChat: Enhancing generative visual analytics using multimodal interactions. InComputer Graphics Forum, Vol. 44. Wiley Online Library, e70112

  11. [11]

    Weihao Chen, Chun Yu, Huadong Wang, Zheng Wang, Lichen Yang, Yukun Wang, Weinan Shi, and Yuanchun Shi. 2023. From gap to synergy: Enhancing contextual understanding through human-machine collaboration in personalized systems. InProceedings of the 36th Annual ACM Symposium on User Interface Software and Technology. 1–15

  12. [12]

    Ziming Cheng, Zhiyuan Huang, Junting Pan, Zhaohui Hou, and Mingjie Zhan

  13. [13]

    Riccardo Coppola, Luca Ardito, and Marco Torchiano. 2019. Fragility of layout-based and visual GUI test scripts: an assessment study on a hybrid mo- bile application. InProceedings of the 10th ACM SIGSOFT International Work- shop on Automating TEST Case Design, Selection, and Evaluation. ACM, 28–34. doi:10.1145/3340433.3342824

  14. [14]

    Laradji, Manuel Del Verme, Tom Marty, David Vázquez, Nicolas Chapados, and Alexandre Lacoste

    Alexandre Drouin, Maxime Gasse, Massimo Caccia, Issam H. Laradji, Manuel Del Verme, Tom Marty, David Vázquez, Nicolas Chapados, and Alexandre Lacoste

  15. [15]

    Anurag Dwarakanath, Neville Dubash, and Sanjay Podder. 2018. Machines that test Software like Humans.CoRRabs/1809.09455 (2018). arXiv:1809.09455 http://arxiv.org/abs/1809.09455

  16. [16]

    Lutfi Eren Erdogan, Nicholas Lee, Sehoon Kim, Suhong Moon, Hiroki Furuta, Gopala Anumanchipalli, Kurt Keutzer, and Amir Gholami. 2025. Plan-and-Act: Improving Planning of Agents for Long-Horizon Tasks. InProceedings of the 42nd International Conference on Machine Learning (Proceedings of Machine Learning Research, Vol. 267), Aarti Singh, Maryam Fazel, Dan...

  17. [17]

    WorkArena: How Capable are Web Agents at Solving Common Knowledge Work Tasks?Forty-first International Conference on Machine Learning, ICML 2024, Vienna, Austria, July 21-27, 2024235 (2024), 11642–11662

  18. [18]

    Boyu Gou, Demi Ruohan Wang, Boyuan Zheng, Yanan Xie, Cheng Chang, Yiheng Shu, Huan Sun, and Yu Su. 2025. Navigating the digital world as humans do: Universal visual grounding for gui agents. InInternational Conference on Learning Representations, Vol. 2025. 30851–30883

  19. [19]

    Nitesh Goyal, Minsuk Chang, and Michael Terry. 2024. Designing for Human- Agent Alignment: Understanding what humans want from their agents. InEx- tended Abstracts of the CHI Conference on Human Factors in Computing Systems. 1–6

  20. [20]

    KJ Kevin Feng, Kevin Pu, Matt Latzke, Tal August, Pao Siangliulue, Jonathan Bragg, Daniel S Weld, Amy X Zhang, and Joseph Chee Chang. 2026. Cocoa: Co-planning and co-execution with ai agents. InProceedings of the 2026 CHI Conference on Human Factors in Computing Systems. ACM, 1–23

  21. [21]

    Grigorev, Alexey K

    Danil S. Grigorev, Alexey K. Kovalev, and Aleksandr I. Panov. 2025. VerifyLLM: LLM-Based Pre-Execution Task Plan Verification for Robots. InIEEE/RSJ Interna- tional Conference on Intelligent Robots and Systems, IROS 2025, Hangzhou, China, October 19-25, 2025. IEEE, 18489–18496

  22. [22]

    Xiangwu Guo, Difei Gao, and Mike Zheng Shou. 2025. AUTO-Explorer: Automated Data Collection for GUI Agent.CoRRabs/2511.06417 (2025). arXiv:2511.06417 doi:10.48550/ARXIV.2511.06417

  23. [23]

    Maria Fernanda Granda, Otto Parra, and Bryan Alba-Sarango. 2021. Towards a Model-Driven Testing Framework for GUI Test Cases Generation from User Stories.. InENASE. 453–460

  24. [24]

    Hongliang He, Wenlin Yao, Kaixin Ma, Wenhao Yu, Yong Dai, Hongming Zhang, Zhenzhong Lan, and Dong Yu. 2024. Webvoyager: Building an end-to-end web agent with large multimodal models. InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 6864–6890

  25. [25]

    Theodore D Hellmann and Frank Maurer. 2011. Rule-based exploratory testing of graphical user interfaces. In2011 Agile Conference. IEEE, 107–116

  26. [26]

    Chao Hao, Shuai Wang, and Kaiwen Zhou. 2025. Uncertainty-aware gui agent: Adaptive perception through component recommendation and human-in-the- loop refinement.arXiv preprint arXiv:2508.04025(2025)

  27. [27]

    Eric Horvitz. 1999. Principles of Mixed-Initiative User Interfaces. InProceeding of the CHI ’99 Conference on Human Factors in Computing Systems: The CHI is the Limit, Pittsburgh, PA, USA, May 15-20, 1999, Marian G. Williams and Mark W. Altom (Eds.). ACM, 159–166. doi:10.1145/302979.303030

  28. [28]

    Siyuan Hu, Mingyu Ouyang, Difei Gao, and Mike Zheng Shou. 2024. The dawn of gui agent: A preliminary case study with claude 3.5 computer use.arXiv preprint arXiv:2411.10323(2024)

  29. [29]

    Peter Hofmann, Caroline Samp, and Nils Urbach. 2020. Robotic process automa- tion.Electronic markets30, 1 (2020), 99–106

  30. [30]

    Zhiyuan Huang, Ziming Cheng, Junting Pan, Zhaohui Hou, and Mingjie Zhan

  31. [31]

    Faria Huq, Zora Zhiruo Wang, Frank F Xu, Tianyue Ou, Shuyan Zhou, Jeffrey P Bigham, and Graham Neubig. 2025. Cowpilot: a framework for autonomous and human-agent collaborative web navigation. InProceedings of the 2025 Confer- ence of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (System D...

  32. [32]

    Wenyue Hua, Mengting Wan, Jagannath Vadrevu, Ryan Nadel, Yongfeng Zhang, and Chi Wang. 2025. Interactive speculative planning: Enhance agent efficiency through co-design of system and user interface. InInternational Conference on Learning Representations, Vol. 2025. 14256–14283

  33. [33]

    Arushi Jain, Shubham Paliwal, Monika Sharma, Lovekesh Vig, and Gau- tam Shroff. 2024. SmartFlow: Robotic Process Automation using LLMs. arXiv:2405.12842 [cs.RO] https://arxiv.org/abs/2405.12842

  34. [34]

    InProceedings of the Computer Vision and Pattern Recognition Conference

    Spiritsight agent: Advanced gui agent with one look. InProceedings of the Computer Vision and Pattern Recognition Conference. 29490–29500

  35. [35]

    Linjia Kang, Zhimin Wang, Yongkang Zhang, Duo Wu, Jinghe Wang, Ming Ma, Haopeng Yan, and Zhi Wang. 2026. Learning with Challenges: Adaptive Difficulty-Aware Data Generation for Mobile GUI Agent Training.arXiv preprint arXiv:2601.22781(2026)

  36. [36]

    Alayt Issak, Jeba Rezwana, and Casper Harteveld. 2025. MOSAAIC: Managing Optimization towards Shared Autonomy, Authority, and Initiative in Co-creation. arXiv:2505.11481 [cs.AI] https://arxiv.org/abs/2505.11481

  37. [37]

    Jing Yu Koh, Robert Lo, Lawrence Jang, Vikram Duvvur, Ming Lim, Po-Yu Huang, Graham Neubig, Shuyan Zhou, Russ Salakhutdinov, and Daniel Fried. 2024. Vi- sualwebarena: Evaluating multimodal agents on realistic visual web tasks. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 881–905

  38. [38]

    Mitchell, and Anupam Datta

    Allison Sihan Jia, Daniel Huang, Nikhil Vytla, Nirvika Choudhury, Shayak Sen, John C. Mitchell, and Anupam Datta. 2025. What Is Your Agent’s GPA? A Frame- work for Evaluating Agent Goal-Plan-Action Alignment.CoRRabs/2510.08847 (2025)

  39. [39]

    Kevin Qinghong Lin, Linjie Li, Difei Gao, Zhengyuan Yang, Shiwei Wu, Zechen Bai, Stan Weixian Lei, Lijuan Wang, and Mike Zheng Shou. 2025. Showui: One vision-language-action model for gui visual agent. InProceedings of the Computer Vision and Pattern Recognition Conference. 19498–19508

  40. [40]

    Anjali Khurana, Xiaotian Su, April Yi Wang, and Parmit K. Chilana. 2025. Do It For Me vs. Do It With Me: Investigating User Perceptions of Different Paradigms of Automation in Copilots for Feature-Rich Software. InProceedings of the 2025 CHI Conference on Human Factors in Computing Systems, CHI 2025, Yokohama- Japan, 26 April 2025- 1 May 2025, Naomi Yamas...

  41. [41]

    Yadong Lu, Jianwei Yang, Yelong Shen, and Ahmed Hassan Awadallah. 2025. OmniParser for Pure Vision Based GUI Agent. https://openreview.net/forum? id=C6hUK6Q1Pi

  42. [42]

    Toby Jia-Jun Li, Amos Azaria, and Brad A Myers. 2017. SUGILITE: creating multimodal smartphone automation by demonstration. InProceedings of the 2017 CHI conference on human factors in computing systems. 6038–6049

  43. [43]

    Jordan Madden, Moxanki Bhavsar, Lhamo Dorje, and Xiaohua Li. 2024. Ro- bustness of Practical Perceptual Hashing Algorithms to Hash-Evasion and Hash- Inversion Attacks. InThe Third Workshop on New Frontiers in Adversarial Machine Learning. https://openreview.net/forum?id=hraOxsleRl

  44. [44]

    Tao Long, Xuanming Zhang, Sitong Wang, Zhou Yu, and Lydia B Chilton. 2025. DoubleAgents: Interactive Simulations for Alignment in Agentic AI.arXiv preprint arXiv:2509.12626(2025)

  45. [45]

    Hussein Mozannar, Gagan Bansal, Cheng Tan, Adam Fourney, Victor Dibia, Jingya Chen, Jack Gerrits, Tyler Payne, Matheus Kunzler Maldaner, Madeleine Grunde-McLaughlin, et al. 2025. Magentic-ui: Towards human-in-the-loop agen- tic systems.arXiv preprint arXiv:2507.22358(2025)

  46. [46]

    Shang Ma, Xusheng Xiao, and Yanfang Ye. 2025. Agent+ P: Guiding UI Agents via Symbolic Planning.arXiv preprint arXiv:2510.06042(2025)

  47. [47]

    Rodriguez, Montek Kalsi, Nicolas Chapados, M

    Shravan Nayak, Xiangru Jian, Kevin Qinghong Lin, Juan A. Rodriguez, Montek Kalsi, Nicolas Chapados, M. Tamer Özsu, Aishwarya Agrawal, David Vazquez, Christopher Pal, Perouz Taslakian, Spandana Gella, and Sai Rajeswar. 2025. UI- Vision: A Desktop-centric GUI Benchmark for Visual Perception and Interaction. InProceedings of the 42nd International Conference...

  48. [48]

    Valérie Maquil, Dimitra Anastasiou, Hoorieh Afkari, Adrien Coppens, Johannes Hermen, and Lou Schwartz. 2023. Establishing Awareness through Pointing Gestures during Collaborative Decision-Making in a Wall-Display Environment. InExtended Abstracts of the 2023 CHI Conference on Human Factors in Computing Systems, CHI EA 2023, Hamburg, Germany, April 23-28, ...

  49. [49]

    Bigham, and Amy Pavel

    Yi-Hao Peng, Dingzeyu Li, Jeffrey P. Bigham, and Amy Pavel. 2025. Morae: Proactively Pausing UI Agents for User Choices. InProceedings of the 38th Annual ACM Symposium on User Interface Software and Technology, UIST 2025, Busan, Korea, 28 September 2025 - 1 October 2025, Andrea Bianchi, Elena L. Glassman, Wendy E. Mackay, Shengdong Zhao, Jeeeun Kim, and I...

  50. [50]

    Arpit Narechania, Shunan Guo, Eunyee Koh, Alex Endert, and Jane Hoffswell

  51. [51]

    Utilizing Provenance as an Attribute for Visual Data Analysis: A Design Probe With ProvenanceLens.IEEE Trans. Vis. Comput. Graph.31, 10 (2025), 8452–8465

  52. [52]

    Ju Qian, Zhengyu Shang, Shuoyan Yan, Yan Wang, and Lin Chen. 2020. Roscript: a visual script driven truly non-intrusive robotic testing system for touch screen applications. InProceedings of the ACM/IEEE 42nd International Conference on Software Engineering. 297–308

  53. [53]

    Mehrab Tanjim, Nesreen K

    Dang Nguyen, Jian Chen, Yu Wang, Gang Wu, Namyong Park, Zhengmian Hu, Hanjia Lyu, Junda Wu, Ryan Aponte, Yu Xia, Xintong Li, Jing Shi, Hongjie Chen, Viet Dac Lai, Zhouhang Xie, Sungchul Kim, Ruiyi Zhang, Tong Yu, Md. Mehrab Tanjim, Nesreen K. Ahmed, Puneet Mathur, Seunghyun Yoon, Lina Yao, Branislav Kveton, Jihyung Kil, Thien Huu Nguyen, Trung Bui, Tianyi...

  54. [54]

    Stephan Rabanser, Sayash Kapoor, Peter Kirgis, Kangheng Liu, Saiteja Utpala, and Arvind Narayanan. 2026. Towards a Science of AI Agent Reliability.arXiv preprint arXiv:2602.16666(2026)

  55. [55]

    Christopher Potts and Moritz Sudhof. 2026. Invisible failures in human-AI inter- actions.arXiv preprint arXiv:2603.15423(2026)

  56. [56]

    Petr Průcha, Michaela Matoušková, and Jan Strnad. 2025. Are LLM Agents the New RPA? A Comparative Study with RPA Across Enterprise Workflows.arXiv preprint arXiv:2509.04198(2025)

  57. [57]

    Minjie Shen, Yanshu Li, Lulu Chen, and Qikai Yang. 2025. From mind to ma- chine: The rise of manus ai as a fully autonomous digital agent.arXiv preprint arXiv:2505.02024(2025)

  58. [58]

    Yujia Qin, Yining Ye, Junjie Fang, Haoming Wang, Shihao Liang, Shizuo Tian, Junda Zhang, Jiahao Li, Yunxin Li, Shijue Huang, Wanjun Zhong, Kuanye Li, Jiale Yang, Yu Miao, Woyu Lin, Longxiang Liu, Xu Jiang, Qianli Ma, Jingyu Li, Xiaojun Xiao, Kai Cai, Chuang Li, Yaowei Zheng, Chaolin Jin, Chen Li, Xiao Zhou, Minchao Wang, Haoli Chen, Zhaojian Li, Haihua Ya...

  59. [59]

    When to Hand Off, When to Work Together

    Kihoon Son, Hyewon Lee, DaEun Choi, Yoonsu Kim, Tae Soo Kim, Yoonjoo Lee, John Joon Young Chung, HyunJoon Jung, and Juho Kim. 2026. " When to Hand Off, When to Work Together": Expanding Human-Agent Co-Creative Collaboration through Concurrent Interaction.arXiv preprint arXiv:2603.02050 (2026)

  60. [60]

    Peter Shaw, Mandar Joshi, James Cohan, Jonathan Berant, Panupong Pasupat, Hexiang Hu, Urvashi Khandelwal, Kenton Lee, and Kristina N Toutanova. 2023. From pixels to ui actions: Learning to follow instructions via graphical user interfaces.Advances in Neural Information Processing Systems36 (2023), 34354– 34370

  61. [61]

    Erfan Shayegani, Keegan Hines, Yue Dong, Nael Abu-Ghazaleh, Roman Lutz, Spencer Whitehead, Vidhisha Balachandran, Besmira Nushi, and Vibhav Vineet

  62. [62]

    Qiushi Sun, Kanzhi Cheng, Zichen Ding, Chuanyang Jin, Yian Wang, Fangzhi Xu, Zhenyu Wu, Chengyou Jia, Liheng Chen, Zhoumianze Liu, Ben Kao, Guohao Li, Junxian He, Yu Qiao, and Zhiyong Wu. 2025. OS-Genesis: Automating GUI Agent Trajectory Construction via Reverse Task Synthesis. InProceedings of the 63rd Annual Meeting of the Association for Computational ...

  63. [63]

    Fei Tang, Haolei Xu, Hang Zhang, Siqi Chen, Xingyu Wu, Yongliang Shen, Wenqi Zhang, Guiyang Hou, Zeqi Tan, Yuchen Yan, et al. 2025. A survey on (m) llm-based gui agents.arXiv preprint arXiv:2504.13865(2025)

  64. [64]

    Aleksandar Shtedritski, Christian Rupprecht, and Andrea Vedaldi. 2023. What does CLIP know about a red circle? Visual prompt engineering for VLMs. In IEEE/CVF International Conference on Computer Vision, ICCV 2023, Paris, France, October 1-6, 2023. IEEE, 11953–11963. doi:10.1109/ICCV51070.2023.01101

  65. [65]

    Yuanrong Tang, Huiling Peng, Bingxi Zhao, Hengyang Ding, Hanchao Song, Tianhong Wang, Chen Zhong, and Jiangtao Gong. 2026. Human Tool: An MCP- Style Framework for Human-Agent Collaboration.arXiv preprint arXiv:2602.12953 (2026)

  66. [66]

    Yunpeng Song, Yiheng Bian, Yongtao Tang, Guiyu Ma, and Zhongmin Cai. 2024. Visiontasker: Mobile task automation using vision based ui understanding and llm task planning. InProceedings of the 37th Annual ACM Symposium on User Interface Software and Technology. 1–17

  67. [67]

    Zihe Song, S. M. Hasan Mansur, Ravishka Rathnasuriya, Yumna Fatima, Wei Yang, Kevin Moran, and Wing Lam. 2025. Can You Mimic Me? Exploring the Use of Android Record & Replay Tools in Debugging. In12th IEEE/ACM In- ternational Conference on Mobile Software Engineering and Systems, MOBILE- Soft@ICSE 2025, Ottawa, ON, Canada, April 27-28, 2025. IEEE, 32–43. ...

  68. [68]

    Vanshika Vats, Marzia Binta Nizam, Minghao Liu, Ziyuan Wang, Richard Ho, Mohnish Sai Prasad, Vincent Titterton, Sai Venkat Malreddy, Riya Aggarwal, Yan- wen Xu, et al. 2024. A Survey on Human-AI Collaboration with Large Foundation Models.arXiv preprint arXiv:2403.04931(2024)

  69. [69]

    Bowen Wang, Xinyuan Wang, Jiaqi Deng, Tianbao Xie, Ryan Li, Yanzhe Zhang, Junli Wang, Dunjie Lu, Zicheng Gong, Gavin Li, Toh Jing Hua, Wei-Lin Chiang, Ion Stoica, Diyi Yang, Yu Su, Yi Zhang, Zhiguo Wang, Victor Zhong, and Tao Yu

  70. [70]

    Jingyu Tang, Chaoran Chen, Jiawen Li, Zhiping Zhang, Bingcan Guo, Ibrahim Khalilov, Simret Araya Gebreegziabher, Bingsheng Yao, Dakuo Wang, Yanfang Ye, et al. 2026. Dark patterns meet gui agents: Llm agent susceptibility to manip- ulative interfaces and the role of human oversight.Proceedings of the 2026 CHI Conference on Human Factors in Computing System...

  71. [71]

    Hui Wei, Zihao Zhang, Shenghua He, Tian Xia, Shijia Pan, and Fei Liu. 2025. PlanGenLLMs: A Modern Survey of LLM Planning Capabilities. InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2025, Vienna, Austria, July 27 - August 1, 2025, Wanxiang Che, Joyce Nabende, Ekaterina Shutova, and Mo...

  72. [72]

    Kimi Team, Tongtong Bai, Yifan Bai, Yiping Bao, SH Cai, Yuan Cao, Y Charles, HS Che, Cheng Chen, Guanduo Chen, et al . 2026. Kimi K2. 5: Visual Agentic Intelligence.arXiv preprint arXiv:2602.02276(2026)

  73. [73]

    Ching-Yi Tsai, Nicole Tacconi, Andrew D Wilson, and Parastoo Abtahi. 2026. Uncertain Pointer: Situated Feedforward Visualizations for Ambiguity-Aware AR Target Selection.arXiv preprint arXiv:2602.13433(2026)

  74. [74]

    Judith Wewerka and Manfred Reichert. 2020. Robotic Process Automation - A Systematic Literature Review and Assessment Framework.CoRRabs/2012.11951 (2020). arXiv:2012.11951 https://arxiv.org/abs/2012.11951

  75. [75]

    Rossi, Ruiyi Zhang, Subrata Mitra, Dimitris N

    Junda Wu, Zhehao Zhang, Yu Xia, Xintong Li, Zhaoyang Xia, Aaron Chang, Tong Yu, Sungchul Kim, Ryan A. Rossi, Ruiyi Zhang, Subrata Mitra, Dimitris N. Metaxas, Lina Yao, Jingbo Shang, and Julian J. McAuley. 2024. Visual Prompting in Multimodal Large Language Models: A Survey.CoRRabs/2409.15310 (2024). arXiv:2409.15310 doi:10.48550/ARXIV.2409.15310

  76. [76]

    InThe Fourteenth International Conference on Learning Representations

    Computer Agent Arena: Toward Human-Centric Evaluation and Analysis of Computer-Use Agents. InThe Fourteenth International Conference on Learning Representations. https://openreview.net/forum?id=3x4SDbXbgl

  77. [77]

    Bovik, Hamid R

    Zhou Wang, Alan C. Bovik, Hamid R. Sheikh, and Eero P. Simoncelli. 2004. Image quality assessment: from error visibility to structural similarity.IEEE Trans. Image Process.13, 4 (2004), 600–612

  78. [78]

    Tianbao Xie, Danyang Zhang, Jixuan Chen, Xiaochuan Li, Siheng Zhao, Ruisheng Cao, Toh Jing Hua, Zhoujun Cheng, Dongchan Shin, Fangyu Lei, Yitao Liu, Yiheng Xu, Shuyan Zhou, Silvio Savarese, Caiming Xiong, Victor Zhong, and Tao Yu

  79. [79]

    Hao Wen, Yuanchun Li, Guohong Liu, Shanhui Zhao, Tao Yu, Toby Jia-Jun Li, Shiqi Jiang, Yunhao Liu, Yaqin Zhang, and Yunxin Liu. 2024. Autodroid: Llm-powered task automation in android. InProceedings of the 30th Annual International Con- ference on Mobile Computing and Networking. 543–557

  80. [80]

    Zhen Wen, Luoxuan Weng, Yinghao Tang, Runjin Zhang, Yuxin Liu, Bo Pan, Minfeng Zhu, and Wei Chen. 2025. Exploring Multimodal Prompt for Visual- ization Authoring with Large Language Models.CoRRabs/2504.13700 (2025). doi:10.48550/ARXIV.2504.13700

Showing first 80 references.