Pith. sign in

REVIEW 3 major objections 4 minor 102 references

Plover: Steering GUI Agents through Plan-Centric Interaction

T0 review · 3 major / 4 minor · reviewed 2026-08-01 · deepseek-v4-flash

Pith's one-line read Plover shows that many GUI-agent failures become repairable when the task plan stays visible and corrections stay localized, with 23 of 26 benchmark failures improved by mixed-initiative interaction.

desk verdict Plover is a credible systems paper with an honest upper-bound recovery result; just don't let the abstract sell the 88% as proof that plan visibility is what rescues failures. read the letter →

arxiv 2607.15193 v1 pith:KG7UY24M submitted 2026-07-16 cs.AI

classification cs.AI
keywords GUIagentsmixed-initiativesystemshuman-AIinteractionplan-centricintelligentreplanningmultimodalannotationfailurerecoveryautomation
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper introduces Plover, a plan-centric GUI automation system that externalizes task plans as persistent, inspectable, and revisable artifacts. Its central claim is that many autonomous GUI-agent failures are not terminal errors but localized breakdowns in grounding, state interpretation, or execution continuity that become structurally recoverable when users can see the plan, intervene surgically, and preserve completed work. In a repair study over 26 autonomous non-success cases from the OSWorld-Verified benchmark, mixed-initiative interaction with an expert user improved 23 cases, converting 17 to complete success and 6 to partial success, with an average of 2.04 interventions per task and no regressions. The paper argues that reliable GUI automation should be treated as an interaction problem as much as a modeling problem: making replanning visible and localized turns silent drift into a collaborative, repairable process.

What carries the argument

The central mechanism is the persistent plan artifact, a shared state representation with the invariant that executed steps are immutable and only the pending suffix can be revised (C_{t+1}=C_t). On top of this, Plover implements Intelligent Replanning in two modes: User-Driven IR, where natural-language guidance or multimodal annotations (strokes, shapes, text overlays captured as primitives with a bounding box) generate localized plan proposals, and System-Driven IR, where a watchdog detects behavioral loop repetition and visual non-progress (via dHash Hamming distance) and injects a structured failure message that prompts the model to propose a recovery step with rationale. The versioned

What would settle it

Run the same 26 benchmark failures with naive participants who are not told the failure cause, providing only the Plover interface as the intervention channel. If the mixed-initiative recovery rate falls to near the autonomous baseline (no significant improvement over the 0% success on these originally failed tasks), the claim that failures are structurally repairable through visible plans is not supported for realistic users.

Watch

Extended reading notes

Core claim

Plover's core claim is that GUI-agent failures are structurally recoverable when plans are externalized and repairs are localized. The system keeps a versioned plan artifact, separates immutable executed steps from an editable pending suffix, and supports user-driven interventions (natural-language guidance, plan edits, and screenshot annotations) plus system-driven replanning triggered by non-progress detection. In the benchmark repair analysis, 26 autonomous failures were re-run in a mixed-initiative setting; 23 improved (17 complete successes, 6 partial), only 3 remained failures, and all 10 autonomous partial successes became complete successes. The paper also characterizes which failure

Load-bearing premise

The 88% recovery figure depends on an expert user who already knows what went wrong and what the correct target is; if ordinary users cannot detect drift or formulate correct localized corrections, the recoverability claim may not transfer to practice.

Editorial extensions

If this is right

  • If recoverability holds beyond the expert setting, GUI agents can be deployed in long-horizon, high-friction workflows with a human steering loop instead of requiring near-perfect autonomy.
  • The plan-invariant design means corrections preserve executed history, so each intervention is cheaper and less disruptive than re-prompting or restarting the whole task.
  • System-driven non-progress detection (repeated semantic actions plus visual stability) can act as a reusable watchdog that catches drift before it propagates, independent of the specific planner or executor.
  • The failure taxonomy suggests that perception errors and state misinterpretations are cheaply repairable with language or annotations, while compound failures require catching the initial planning error earlier.
  • Exposing plans as versioned, diffable artifacts provides a natural audit trail for when and why an agent's behavior changed, which can support verification and post-hoc analysis.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A natural next test, which the paper does not run, is a study with non-expert users who are not told the failure cause: if recovery rates drop to near the autonomous baseline, the 'structurally recoverable' claim would need to be re-scoped from an upper bound to a property that depends on user diagnostic skill.
  • The plan artifact as a coordination protocol could generalize beyond GUI automation to other long-horizon agent domains (e.g., data-cleaning pipelines or robotics task plans) where partial progress is valuable and corrections must be localized.
  • The System-Driven IR watchdog (behavioral repetition + perceptual-hash stability) is a concrete, model-agnostic component that could be extracted and benchmarked on its own to measure how many agent stalls it catches before a human would notice.
  • The paper itself flags that visible plans may inflate user confidence; an empirical study measuring whether users over-accept plan proposals when the system looks confident would be a direct extension.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. The paper introduces Plover, a plan-centric GUI automation system that externalizes task plans as persistent, editable artifacts and supports mixed-initiative repair through natural-language guidance, multimodal annotation, plan edits, and system-driven replanning. The authors report a formative study with six participants, a benchmark repair study on 38 OSWorld-Verified tasks (26 autonomous non-successes re-run with expert interventions, yielding 23 improved, 17 complete successes, 6 partial successes, 3 failures), and a scenario-based stability analysis with trajectory-derived prompts. The central claim is that many GUI-agent failures are structurally repairable when plans remain visible and interventions are localized, and that explicit replanning improves transparency, controllability, and adaptability.

Significance. If the central claim were fully established, the paper would make a useful contribution to human-agent interaction for GUI automation: it proposes a concrete design space (persistent plans as coordination artifacts, localized repair, visible replanning) and provides a failure taxonomy that could guide future interface design. The paper is honest in labeling the benchmark repair study as an upper bound on plan-centric recoverability, and the appendices contain substantial implementation and evaluation detail, including per-task results and a formative study summary. These are real strengths. However, the main empirical evidence does not currently separate the effect of the plan-centric interface from the effect of an expert oracle user, so the causal design conclusions (DG1–DG5, 'explicit replanning helps') are not yet established by the data.

major comments (3)
  1. [Section 5.1, Table 1; Abstract] The load-bearing empirical claim—23/26 non-success cases improved, 88% recovery—is measured with the first author, who knows each failure cause and the correct target, supplying all interventions. There is no control condition that strips away the plan-centric affordances (plan panel, plan editing, annotation, visible replanning) while keeping the same underlying agent and the same expert. As written, the 88% figure is an upper bound on recoverability by an informed oracle, not evidence that plan visibility and localized repair cause the improvement. The paper's own limitation statement in Section 6 ('our design assumes users generally know the intended path well enough to correct the agent') does not carry through to the abstract/conclusion, which assert a causal role for visible plans. Add an ablation or control (e.g., same expert corrections issued as text prompts to the same base age
  2. [Section 5.2, 'Trajectory Sampling and Prompt Reconstruction' and Table 2] The scenario stability analysis synthesizes task prompts from sampled interaction trajectories and then compares the replayed plans and final states against the same trajectories. This creates a circularity: the prompt is derived from the reference trajectory, so plan alignment metrics (coverage 0.62, order 0.41, actionability 0.97) partly measure reconstruction from a trajectory-derived instruction rather than general plan quality. The browser-vs-desktop visual fidelity differences are still informative, but the plan alignment results should be presented as a property of this reverse-synthesis setup, not as evidence about Plover's planning under independently authored user instructions. A sanity check with manually authored prompts, or a clear caveat, is needed before these numbers are used to support the 'opportunity for users to inspect and correct' argument.
  3. [Section 5.3, 'Characterizing Repairable Failures'] The failure-mode analysis is based on the same 26 cases and the same expert interventions. Counts such as 'Execution Drift appeared in 46% (n=12/26)' and 'NL Guidance resolved 11 cases' are reported without any uncertainty or sensitivity analysis. With n=26 and intervention choices made by a single expert who already knows the failure causes, small counts can easily flip; the recovered vs. unrecovered distinction is not robust enough to support the strong claim that 'compound failures' are fundamentally harder. At minimum, report bootstrap or exact binomial confidence intervals and clarify that all recovery counts are conditional on the expert's choice of intervention.
minor comments (4)
  1. [Appendix C, Algorithm 1] System-Driven IR relies on hardcoded thresholds (REPEAT_SEQ_L3_R3, dHash Hamming distance > 40) with no sensitivity analysis. Since Section 5.3 attributes 9 successful recoveries to System-Driven IR, the threshold choices can materially affect the results; report how varying them changes detection and downstream recovery.
  2. [Section 5.1, 'no regressions'] The statement 'no regressions were observed' only covers the 26 autonomous non-success cases; it does not address whether the mixed-initiative interaction could degrade autonomous successes, since those were not re-run in the mixed-initiative condition. Please state this scope explicitly.
  3. [Table 1 and Table 2] Table 1's 'Improv. Rate' is not formally defined; for Multi-App, 6S+2P out of 10 corresponds to 80%, but the reader must infer the denominator. In Table 2, the 'Overall Average' row for MSE (939.57) is the mean of scenario averages, not the mean over all trials; clarify the aggregation.
  4. [Throughout] The paper uses 'Conference’17' in the ACM reference format and several placeholder-style citations (e.g., the DOI is 'XXXXXXX.XXXXXXX'). Please update the formatting to the final venue style and correct minor typographical issues such as the 'MI (a)' label in Figure 5.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the 88% recoverability result is an externally grounded upper-bound measurement with the expert-oracle caveat explicitly acknowledged; remaining concerns are validity limitations, not circular reductions.

full rationale

Plover's central empirical claim is an upper-bound measurement on the external OSWorld-Verified benchmark, not a quantity derived from a fitted parameter, an ansatz, or a load-bearing self-citation. The paper states the intervention condition explicitly: 'This setup establishes an upper bound on plan-centric recoverability rather than typical user performance' (Section 5.1), and the corresponding limitation is restated in Section 6 ('our design assumes users generally know the intended path well enough to correct the agent when needed'). Because the expert interventions are acknowledged as ideal rather than typical, the 88% recovery figure is an honest conditional result, not a hidden assumption presented as a finding. The scenario-based stability analysis (Section 5.2) synthesizes prompts from reference trajectories and then compares generated plans against those trajectories, which introduces non-independence; however the reported alignment metrics are moderate (coverage 0.62, order 0.41), so the result is not forced by construction. The only self-citation is a related-work mention ([10], multimodal interaction) and is not load-bearing. No uniqueness theorem, no ansatz smuggled via citation, and no equation reduces the conclusion to its input. The absence of a chat-only or invisible-plan control is a causal-identification limitation, but under the circularity criteria it does not constitute circularity.

Assumptions & free parameters 4 free parameters · 6 assumptions · 0 invented entities

No invented physical or formal entities. The central numerical result is an empirical recovery rate, not a derivation. The main free quantities are hand-set detection/evaluation thresholds and manual trajectory curation, none of which are validated independently.

free parameters (4)
  • Visual non-progress dHash threshold τ = 40 (Hamming distance)
    Algorithm 1 line 14: if any Hamming distance > 40, execution is considered non-static and not stuck. This value is hand-set, not validated; changing it changes which failures trigger System-Driven IR.
  • Repetition pattern REPEAT_SEQ_L3_R3 = repeated action subsequence of length 3 seen 3 times
    Appendix C, Figure 8: used to detect behavioral loops; hand-set. Higher/lower values alter how many drift cases are classified as stuck.
  • Per-scenario similarity thresholds (SSIM/MSE/dHash) = e.g., Firefox Fillable Form High: SSIM≥0.98, MSE≤100, dHash≤2; LibreOffice thresholds differ
    Table 8: thresholds defining High/Partial/Low replay match are set per scenario, so the visual-fidelity conclusions are partly products of chosen cutoffs.
  • Selected trajectories per scenario = 5 per scenario from 100 exploration trials, after manual verification
    Section 5.2: manual curation of 'diverse' trajectories is a subjective filter; it determines which prompts and reference plans are used for both replay and plan alignment.
assumptions (6)
  • domain assumption Claude 4.5 Sonnet via computer-use can decompose tasks into deterministic UI instructions and ground them in the executor.
    Sections 4.4-4.5: if the planner produces non-executable steps or the executor cannot act on them, both autonomous execution and repair fail regardless of plan visibility.
  • domain assumption Image-similarity metrics (SSIM, MSE, dHash) are valid proxies for GUI execution stability and task-state equivalence.
    Section 5.2 uses these to conclude browser workflows are stable and desktop workflows are not; no validation that image similarity corresponds to task success/failure.
  • domain assumption The 38 OSWorld-Verified tasks that failed for Claude 4.5 Sonnet are representative of GUI-agent failures in general.
    Section 5.1 selects failure cases from a single model/benchmark; the failure distribution (e.g., compound failures 4/26) may not generalize across agents and environments.
  • domain assumption An expert who knows failure causes can stand in for a typical user to establish 'structural recoverability'.
    Section 5.1: 'This setup establishes an upper bound on plan-centric recoverability rather than typical user performance.' The entire headline rate depends on this premise.
  • domain assumption GPA phase taxonomy and trajectory-derived phases are a meaningful ground truth for plan alignment.
    Section 5.2 maps plans and trajectories to a shared interface-agnostic taxonomy; coverage/order/redundancy/actionability are computed relative to this lossy mapping.
  • domain assumption The invariant C_{t+1}=C_t (completed steps immutable) is a sound design constraint for repair.
    Section 4.4: preserving completed history assumes completed steps are correct; if an early error was committed to history, the invariant freezes it and can prevent full recovery, as seen in compound failures.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Plover: Steering GUI Agents through Plan-Centric Interaction." pith.science (2026). https://pith.science/paper/KG7UY24M

@misc{pith2026260715193,
  author       = {Pith},
  title        = {Pith review of: Plover: Steering GUI Agents through Plan-Centric Interaction},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/KG7UY24M}},
  note         = {Machine review of arXiv:2607.15193}
}
read the original abstract

Graphical user interface (GUI) automation remains challenging in real-world environments, where dynamic layouts, unexpected dialogs, and evolving interface states can cause autonomous agents to drift from user intent. Recent vision-based multimodal agents improve flexibility by operating directly over screenshots and natural language instructions, but planning and adaptation often remain internal, limiting users' ability to inspect, supervise, or correct system behavior. We present Plover, a plan-centric vision-based GUI automation system that externalizes task plans and replanning as persistent, inspectable, and revisable artifacts. Through a planner--executor architecture, Plover supports explicit supervision of evolving execution, localized correction through editable plans, natural-language guidance, and screenshot-grounded interventions, while preserving prior progress during repair. A formative study with six participants informed the interaction design. We then evaluate Plover through benchmark failure-case repair and scenario-based workflow analyses. Our results show that many autonomous GUI-agent failures are structurally repairable when plans remain visible and interventions are localized, and that explicit replanning helps make GUI automation more transparent, controllable, and adaptable.

Figures

Figures reproduced from arXiv: 2607.15193 by the authors.

Figure 1
Figure 1. The Plover system architecture. it was making progress, and apply corrections without losing prior work. We synthesize these findings into three design questions: Q1) How can systems maintain alignment between evolving user intent and execution across multi-step workflows? Q2) How can in￾terfaces support precise, visually grounded correction when execution deviates? Q3) How can replanning be surfaced as a transparen… view at source ↗
Figure 2
Figure 2. Plover Agentic Interface. (a) System Status Bar shows execution phase and plan state. (b) Planner Chat supports [PITH_FULL_IMAGE:figures/full_fig_p005_2.png] view at source ↗
Figure 3
Figure 3. Provenance Bar visualizing branching plan revi [PITH_FULL_IMAGE:figures/full_fig_p006_3.png] view at source ↗
Figures from the paper (6 more)
Figure 4
Figure 4. Figure 4: Multimodal Annotation workflow for resolving [PITH_FULL_IMAGE:figures/full_fig_p006_4.png]
Figure 5
Figure 5. Figure 5: Representative failure archetypes and recovery pathways in [PITH_FULL_IMAGE:figures/full_fig_p008_5.png]
Figure 6
Figure 6. Figure 6: Early Plover prototype used in Formative Study. The interface supported prompt authoring, plan inspection/editing, and execution monitoring. Observations from this prototype informed the redesign of Plover (DG1–DG5) [PITH_FULL_IMAGE:figures/full_fig_p016_6.png]
Figure 7
Figure 7. Figure 7: Early Annotation Panel. Users could provide visual annotations during execution; feedback from this component [PITH_FULL_IMAGE:figures/full_fig_p016_7.png]
Figure 9
Figure 9. Figure 9: Proposal-mode prompt suffix used during System-Driven IR. When execution drift is detected, this suffix is appended to the system prompt to constrain the model to produce a structured recovery proposal consisting of a short next-step summary and rationale. D Implementa…
Figure 8
Figure 8. Figure 8: Structured failure message injected into the con [PITH_FULL_IMAGE:figures/full_fig_p018_8.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

102 extracted references · 2 canonical work pages

  1. [1]

    Mohamed Aghzal, Gregory J Stein, and Ziyu Yao. 2026. Why do LLM-based web agents fail? A hierarchical planning perspective. InProceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 32157–32180

  2. [2]

    Saleema Amershi, Dan Weld, Mihaela Vorvoreanu, Adam Fourney, Besmira Nushi, Penny Collisson, Jina Suh, Shamsi Iqbal, Paul N Bennett, Kori Inkpen, et al. 2019. Guidelines for human-AI interaction. InProceedings of the 2019 CHI Conference on Human Factors in Computing Systems. 1–13

  3. [3]

    Nishant Balepur, Matthew Shu, Yoo Yeon Sung, Seraphina Goldfarb-Tarrant, Shi Feng, Fumeng Yang, Rachel Rudinger, and Jordan Lee Boyd-Graber. 2025. A Good Plan is Hard to Find: Aligning Models with Preferences is Misaligned with What Helps Users. InProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, EMNLP 2025, Suzhou, C...

  4. [4]

    Tanvir Bhathal and Asanshay Gupta. 2025. Websight: A vision-first architecture for robust web agents.arXiv preprint arXiv:2508.16987(2025)

  5. [5]

    Léo Boisvert, Megh Thakkar, Maxime Gasse, Massimo Caccia, Thibault L De Chezelles, Quentin Cappart, Nicolas Chapados, Alexandre Lacoste, and Alexandre Drouin. 2024. Workarena++: Towards compositional planning and reasoning-based common knowledge work tasks.Advances in Neural Information Processing Systems37 (2024), 5996–6051

  6. [6]

    Sacha Brisset, Romain Rouvoy, Lionel Seinturier, and Renaud Pawlak. 2022. Er- ratum: Leveraging Flexible Tree Matching to repair broken locators in web automation scripts.Inf. Softw. Technol.144 (2022), 106754. doi:10.1016/J.INFSOF. 2021.106754

  7. [7]

    Meyer, Yuning Chai, Dennis Park, and Yong Jae Lee

    Mu Cai, Haotian Liu, Siva Karthik Mustikovela, Gregory P. Meyer, Yuning Chai, Dennis Park, and Yong Jae Lee. 2024. ViP-LLaVA: Making Large Multimodal Models Understand Arbitrary Visual Prompts. InIEEE/CVF Conference on Com- puter Vision and Pattern Recognition, CVPR 2024, Seattle, W A, USA, June 16-22,

  8. [8]

    Fadnis, Kartik Talamadupula, Mishal Dholakia, Biplav Srivastava, Jeffrey O

    Tathagata Chakraborti, Kshitij P. Fadnis, Kartik Talamadupula, Mishal Dholakia, Biplav Srivastava, Jeffrey O. Kephart, and Rachel K. E. Bellamy. 2018. Visualiza- tions for an Explainable Planning Agent. InProceedings of the Twenty-Seventh International Joint Conference on Artificial Intelligence, IJCAI 2018, July 13-19, 2018, Stockholm, Sweden, Jérôme Lan...

Show all 102 references
  1. [9]

    Tathagata Chakraborti, Sarath Sreedharan, Sachin Grover, and Subbarao Kamb- hampati. 2019. Plan explanations as model reconciliation–an empirical study. In 2019 14th ACM/IEEE International Conference on Human-Robot Interaction (HRI). Ieee, 258–266

  2. [10]

    Juntong Chen, Jiang Wu, Jiajing Guo, Vikram Mohanty, Xueming Li, Jorge Pi- azentin Ono, Wenbin He, Liu Ren, and Dongyu Liu. 2025. InterChat: Enhancing generative visual analytics using multimodal interactions. InComputer Graphics Forum, Vol. 44. Wiley Online Library, e70112

  3. [11]

    Weihao Chen, Chun Yu, Huadong Wang, Zheng Wang, Lichen Yang, Yukun Wang, Weinan Shi, and Yuanchun Shi. 2023. From gap to synergy: Enhancing contextual understanding through human-machine collaboration in personalized systems. InProceedings of the 36th Annual ACM Symposium on U...

  4. [12]

    Ziming Cheng, Zhiyuan Huang, Junting Pan, Zhaohui Hou, and Mingjie Zhan

  5. [13]

    Riccardo Coppola, Luca Ardito, and Marco Torchiano. 2019. Fragility of layout-based and visual GUI test scripts: an assessment study on a hybrid mo- bile application. InProceedings of the 10th ACM SIGSOFT International Work- shop on Automating TEST Case Design, Selection, and ...

  6. [14]

    Laradji, Manuel Del Verme, Tom Marty, David Vázquez, Nicolas Chapados, and Alexandre Lacoste

    Alexandre Drouin, Maxime Gasse, Massimo Caccia, Issam H. Laradji, Manuel Del Verme, Tom Marty, David Vázquez, Nicolas Chapados, and Alexandre Lacoste

  7. [15]

    Anurag Dwarakanath, Neville Dubash, and Sanjay Podder. 2018. Machines that test Software like Humans.CoRRabs/1809.09455 (2018). arXiv:1809.09455 http://arxiv.org/abs/1809.09455

  8. [16]

    Lutfi Eren Erdogan, Nicholas Lee, Sehoon Kim, Suhong Moon, Hiroki Furuta, Gopala Anumanchipalli, Kurt Keutzer, and Amir Gholami. 2025. Plan-and-Act: Improving Planning of Agents for Long-Horizon Tasks. InProceedings of the 42nd International Conference on Machine Learning (Pro...

  9. [17]

    WorkArena: How Capable are Web Agents at Solving Common Knowledge Work Tasks?Forty-first International Conference on Machine Learning, ICML 2024, Vienna, Austria, July 21-27, 2024235 (2024), 11642–11662

  10. [18]

    Boyu Gou, Demi Ruohan Wang, Boyuan Zheng, Yanan Xie, Cheng Chang, Yiheng Shu, Huan Sun, and Yu Su. 2025. Navigating the digital world as humans do: Universal visual grounding for gui agents. InInternational Conference on Learning Representations, Vol. 2025. 30851–30883

  11. [19]

    Nitesh Goyal, Minsuk Chang, and Michael Terry. 2024. Designing for Human- Agent Alignment: Understanding what humans want from their agents. InEx- tended Abstracts of the CHI Conference on Human Factors in Computing Systems. 1–6

  12. [20]

    KJ Kevin Feng, Kevin Pu, Matt Latzke, Tal August, Pao Siangliulue, Jonathan Bragg, Daniel S Weld, Amy X Zhang, and Joseph Chee Chang. 2026. Cocoa: Co-planning and co-execution with ai agents. InProceedings of the 2026 CHI Conference on Human Factors in Computing Systems. ACM, 1–23

  13. [21]

    Grigorev, Alexey K

    Danil S. Grigorev, Alexey K. Kovalev, and Aleksandr I. Panov. 2025. VerifyLLM: LLM-Based Pre-Execution Task Plan Verification for Robots. InIEEE/RSJ Interna- tional Conference on Intelligent Robots and Systems, IROS 2025, Hangzhou, China, October 19-25, 2025. IEEE, 18489–18496

  14. [22]

    Xiangwu Guo, Difei Gao, and Mike Zheng Shou. 2025. AUTO-Explorer: Automated Data Collection for GUI Agent.CoRRabs/2511.06417 (2025). arXiv:2511.06417 doi:10.48550/ARXIV.2511.06417

  15. [23]

    Maria Fernanda Granda, Otto Parra, and Bryan Alba-Sarango. 2021. Towards a Model-Driven Testing Framework for GUI Test Cases Generation from User Stories.. InENASE. 453–460

  16. [24]

    Hongliang He, Wenlin Yao, Kaixin Ma, Wenhao Yu, Yong Dai, Hongming Zhang, Zhenzhong Lan, and Dong Yu. 2024. Webvoyager: Building an end-to-end web agent with large multimodal models. InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Vol...

  17. [25]

    Theodore D Hellmann and Frank Maurer. 2011. Rule-based exploratory testing of graphical user interfaces. In2011 Agile Conference. IEEE, 107–116

  18. [26]

    Chao Hao, Shuai Wang, and Kaiwen Zhou. 2025. Uncertainty-aware gui agent: Adaptive perception through component recommendation and human-in-the- loop refinement.arXiv preprint arXiv:2508.04025(2025)

  19. [27]

    Eric Horvitz. 1999. Principles of Mixed-Initiative User Interfaces. InProceeding of the CHI ’99 Conference on Human Factors in Computing Systems: The CHI is the Limit, Pittsburgh, PA, USA, May 15-20, 1999, Marian G. Williams and Mark W. Altom (Eds.). ACM, 159–166. doi:10.1145/...

  20. [28]

    Siyuan Hu, Mingyu Ouyang, Difei Gao, and Mike Zheng Shou. 2024. The dawn of gui agent: A preliminary case study with claude 3.5 computer use.arXiv preprint arXiv:2411.10323(2024)

  21. [29]

    Peter Hofmann, Caroline Samp, and Nils Urbach. 2020. Robotic process automa- tion.Electronic markets30, 1 (2020), 99–106

  22. [30]

    Zhiyuan Huang, Ziming Cheng, Junting Pan, Zhaohui Hou, and Mingjie Zhan

  23. [31]

    Faria Huq, Zora Zhiruo Wang, Frank F Xu, Tianyue Ou, Shuyan Zhou, Jeffrey P Bigham, and Graham Neubig. 2025. Cowpilot: a framework for autonomous and human-agent collaborative web navigation. InProceedings of the 2025 Confer- ence of the Nations of the Americas Chapter of the ...

  24. [32]

    Wenyue Hua, Mengting Wan, Jagannath Vadrevu, Ryan Nadel, Yongfeng Zhang, and Chi Wang. 2025. Interactive speculative planning: Enhance agent efficiency through co-design of system and user interface. InInternational Conference on Learning Representations, Vol. 2025. 14256–14283

  25. [33]

    Arushi Jain, Shubham Paliwal, Monika Sharma, Lovekesh Vig, and Gau- tam Shroff. 2024. SmartFlow: Robotic Process Automation using LLMs. arXiv:2405.12842 [cs.RO] https://arxiv.org/abs/2405.12842

  26. [34]

    InProceedings of the Computer Vision and Pattern Recognition Conference

    Spiritsight agent: Advanced gui agent with one look. InProceedings of the Computer Vision and Pattern Recognition Conference. 29490–29500

  27. [35]

    Linjia Kang, Zhimin Wang, Yongkang Zhang, Duo Wu, Jinghe Wang, Ming Ma, Haopeng Yan, and Zhi Wang. 2026. Learning with Challenges: Adaptive Difficulty-Aware Data Generation for Mobile GUI Agent Training.arXiv preprint arXiv:2601.22781(2026)

  28. [36]

    Alayt Issak, Jeba Rezwana, and Casper Harteveld. 2025. MOSAAIC: Managing Optimization towards Shared Autonomy, Authority, and Initiative in Co-creation. arXiv:2505.11481 [cs.AI] https://arxiv.org/abs/2505.11481

  29. [37]

    Jing Yu Koh, Robert Lo, Lawrence Jang, Vikram Duvvur, Ming Lim, Po-Yu Huang, Graham Neubig, Shuyan Zhou, Russ Salakhutdinov, and Daniel Fried. 2024. Vi- sualwebarena: Evaluating multimodal agents on realistic visual web tasks. In Proceedings of the 62nd Annual Meeting of the A...

  30. [38]

    Mitchell, and Anupam Datta

    Allison Sihan Jia, Daniel Huang, Nikhil Vytla, Nirvika Choudhury, Shayak Sen, John C. Mitchell, and Anupam Datta. 2025. What Is Your Agent’s GPA? A Frame- work for Evaluating Agent Goal-Plan-Action Alignment.CoRRabs/2510.08847 (2025)

  31. [39]

    Kevin Qinghong Lin, Linjie Li, Difei Gao, Zhengyuan Yang, Shiwei Wu, Zechen Bai, Stan Weixian Lei, Lijuan Wang, and Mike Zheng Shou. 2025. Showui: One vision-language-action model for gui visual agent. InProceedings of the Computer Vision and Pattern Recognition Conference. 19...

  32. [40]

    Anjali Khurana, Xiaotian Su, April Yi Wang, and Parmit K. Chilana. 2025. Do It For Me vs. Do It With Me: Investigating User Perceptions of Different Paradigms of Automation in Copilots for Feature-Rich Software. InProceedings of the 2025 CHI Conference on Human Factors in Comp...

  33. [41]

    Yadong Lu, Jianwei Yang, Yelong Shen, and Ahmed Hassan Awadallah. 2025. OmniParser for Pure Vision Based GUI Agent. https://openreview.net/forum? id=C6hUK6Q1Pi

  34. [42]

    Toby Jia-Jun Li, Amos Azaria, and Brad A Myers. 2017. SUGILITE: creating multimodal smartphone automation by demonstration. InProceedings of the 2017 CHI conference on human factors in computing systems. 6038–6049

  35. [43]

    Jordan Madden, Moxanki Bhavsar, Lhamo Dorje, and Xiaohua Li. 2024. Ro- bustness of Practical Perceptual Hashing Algorithms to Hash-Evasion and Hash- Inversion Attacks. InThe Third Workshop on New Frontiers in Adversarial Machine Learning. https://openreview.net/forum?id=hraOxsleRl

  36. [44]

    Tao Long, Xuanming Zhang, Sitong Wang, Zhou Yu, and Lydia B Chilton. 2025. DoubleAgents: Interactive Simulations for Alignment in Agentic AI.arXiv preprint arXiv:2509.12626(2025)

  37. [45]

    Hussein Mozannar, Gagan Bansal, Cheng Tan, Adam Fourney, Victor Dibia, Jingya Chen, Jack Gerrits, Tyler Payne, Matheus Kunzler Maldaner, Madeleine Grunde-McLaughlin, et al. 2025. Magentic-ui: Towards human-in-the-loop agen- tic systems.arXiv preprint arXiv:2507.22358(2025)

  38. [46]

    Shang Ma, Xusheng Xiao, and Yanfang Ye. 2025. Agent+ P: Guiding UI Agents via Symbolic Planning.arXiv preprint arXiv:2510.06042(2025)

  39. [47]

    Rodriguez, Montek Kalsi, Nicolas Chapados, M

    Shravan Nayak, Xiangru Jian, Kevin Qinghong Lin, Juan A. Rodriguez, Montek Kalsi, Nicolas Chapados, M. Tamer Özsu, Aishwarya Agrawal, David Vazquez, Christopher Pal, Perouz Taslakian, Spandana Gella, and Sai Rajeswar. 2025. UI- Vision: A Desktop-centric GUI Benchmark for Visua...

  40. [48]

    Valérie Maquil, Dimitra Anastasiou, Hoorieh Afkari, Adrien Coppens, Johannes Hermen, and Lou Schwartz. 2023. Establishing Awareness through Pointing Gestures during Collaborative Decision-Making in a Wall-Display Environment. InExtended Abstracts of the 2023 CHI Conference on ...

  41. [49]

    Bigham, and Amy Pavel

    Yi-Hao Peng, Dingzeyu Li, Jeffrey P. Bigham, and Amy Pavel. 2025. Morae: Proactively Pausing UI Agents for User Choices. InProceedings of the 38th Annual ACM Symposium on User Interface Software and Technology, UIST 2025, Busan, Korea, 28 September 2025 - 1 October 2025, Andre...

  42. [50]

    Arpit Narechania, Shunan Guo, Eunyee Koh, Alex Endert, and Jane Hoffswell

  43. [51]

    Utilizing Provenance as an Attribute for Visual Data Analysis: A Design Probe With ProvenanceLens.IEEE Trans. Vis. Comput. Graph.31, 10 (2025), 8452–8465

  44. [52]

    Ju Qian, Zhengyu Shang, Shuoyan Yan, Yan Wang, and Lin Chen. 2020. Roscript: a visual script driven truly non-intrusive robotic testing system for touch screen applications. InProceedings of the ACM/IEEE 42nd International Conference on Software Engineering. 297–308

  45. [53]

    Mehrab Tanjim, Nesreen K

    Dang Nguyen, Jian Chen, Yu Wang, Gang Wu, Namyong Park, Zhengmian Hu, Hanjia Lyu, Junda Wu, Ryan Aponte, Yu Xia, Xintong Li, Jing Shi, Hongjie Chen, Viet Dac Lai, Zhouhang Xie, Sungchul Kim, Ruiyi Zhang, Tong Yu, Md. Mehrab Tanjim, Nesreen K. Ahmed, Puneet Mathur, Seunghyun Yo...

  46. [54]

    Stephan Rabanser, Sayash Kapoor, Peter Kirgis, Kangheng Liu, Saiteja Utpala, and Arvind Narayanan. 2026. Towards a Science of AI Agent Reliability.arXiv preprint arXiv:2602.16666(2026)

  47. [55]

    Christopher Potts and Moritz Sudhof. 2026. Invisible failures in human-AI inter- actions.arXiv preprint arXiv:2603.15423(2026)

  48. [56]

    Petr Průcha, Michaela Matoušková, and Jan Strnad. 2025. Are LLM Agents the New RPA? A Comparative Study with RPA Across Enterprise Workflows.arXiv preprint arXiv:2509.04198(2025)

  49. [57]

    Minjie Shen, Yanshu Li, Lulu Chen, and Qikai Yang. 2025. From mind to ma- chine: The rise of manus ai as a fully autonomous digital agent.arXiv preprint arXiv:2505.02024(2025)

  50. [58]

    Yujia Qin, Yining Ye, Junjie Fang, Haoming Wang, Shihao Liang, Shizuo Tian, Junda Zhang, Jiahao Li, Yunxin Li, Shijue Huang, Wanjun Zhong, Kuanye Li, Jiale Yang, Yu Miao, Woyu Lin, Longxiang Liu, Xu Jiang, Qianli Ma, Jingyu Li, Xiaojun Xiao, Kai Cai, Chuang Li, Yaowei Zheng, C...

  51. [59]

    When to Hand Off, When to Work Together

    Kihoon Son, Hyewon Lee, DaEun Choi, Yoonsu Kim, Tae Soo Kim, Yoonjoo Lee, John Joon Young Chung, HyunJoon Jung, and Juho Kim. 2026. " When to Hand Off, When to Work Together": Expanding Human-Agent Co-Creative Collaboration through Concurrent Interaction.arXiv preprint arXiv:2...

  52. [60]

    Peter Shaw, Mandar Joshi, James Cohan, Jonathan Berant, Panupong Pasupat, Hexiang Hu, Urvashi Khandelwal, Kenton Lee, and Kristina N Toutanova. 2023. From pixels to ui actions: Learning to follow instructions via graphical user interfaces.Advances in Neural Information Process...

  53. [61]

    Erfan Shayegani, Keegan Hines, Yue Dong, Nael Abu-Ghazaleh, Roman Lutz, Spencer Whitehead, Vidhisha Balachandran, Besmira Nushi, and Vibhav Vineet

  54. [62]

    Qiushi Sun, Kanzhi Cheng, Zichen Ding, Chuanyang Jin, Yian Wang, Fangzhi Xu, Zhenyu Wu, Chengyou Jia, Liheng Chen, Zhoumianze Liu, Ben Kao, Guohao Li, Junxian He, Yu Qiao, and Zhiyong Wu. 2025. OS-Genesis: Automating GUI Agent Trajectory Construction via Reverse Task Synthesis...

  55. [63]

    Fei Tang, Haolei Xu, Hang Zhang, Siqi Chen, Xingyu Wu, Yongliang Shen, Wenqi Zhang, Guiyang Hou, Zeqi Tan, Yuchen Yan, et al. 2025. A survey on (m) llm-based gui agents.arXiv preprint arXiv:2504.13865(2025)

  56. [64]

    Aleksandar Shtedritski, Christian Rupprecht, and Andrea Vedaldi. 2023. What does CLIP know about a red circle? Visual prompt engineering for VLMs. In IEEE/CVF International Conference on Computer Vision, ICCV 2023, Paris, France, October 1-6, 2023. IEEE, 11953–11963. doi:10.11...

  57. [65]

    Yuanrong Tang, Huiling Peng, Bingxi Zhao, Hengyang Ding, Hanchao Song, Tianhong Wang, Chen Zhong, and Jiangtao Gong. 2026. Human Tool: An MCP- Style Framework for Human-Agent Collaboration.arXiv preprint arXiv:2602.12953 (2026)

  58. [66]

    Yunpeng Song, Yiheng Bian, Yongtao Tang, Guiyu Ma, and Zhongmin Cai. 2024. Visiontasker: Mobile task automation using vision based ui understanding and llm task planning. InProceedings of the 37th Annual ACM Symposium on User Interface Software and Technology. 1–17

  59. [67]

    Zihe Song, S. M. Hasan Mansur, Ravishka Rathnasuriya, Yumna Fatima, Wei Yang, Kevin Moran, and Wing Lam. 2025. Can You Mimic Me? Exploring the Use of Android Record & Replay Tools in Debugging. In12th IEEE/ACM In- ternational Conference on Mobile Software Engineering and Syste...

  60. [68]

    Vanshika Vats, Marzia Binta Nizam, Minghao Liu, Ziyuan Wang, Richard Ho, Mohnish Sai Prasad, Vincent Titterton, Sai Venkat Malreddy, Riya Aggarwal, Yan- wen Xu, et al. 2024. A Survey on Human-AI Collaboration with Large Foundation Models.arXiv preprint arXiv:2403.04931(2024)

  61. [69]

    Bowen Wang, Xinyuan Wang, Jiaqi Deng, Tianbao Xie, Ryan Li, Yanzhe Zhang, Junli Wang, Dunjie Lu, Zicheng Gong, Gavin Li, Toh Jing Hua, Wei-Lin Chiang, Ion Stoica, Diyi Yang, Yu Su, Yi Zhang, Zhiguo Wang, Victor Zhong, and Tao Yu

  62. [70]

    Jingyu Tang, Chaoran Chen, Jiawen Li, Zhiping Zhang, Bingcan Guo, Ibrahim Khalilov, Simret Araya Gebreegziabher, Bingsheng Yao, Dakuo Wang, Yanfang Ye, et al. 2026. Dark patterns meet gui agents: Llm agent susceptibility to manip- ulative interfaces and the role of human overs...

  63. [71]

    Hui Wei, Zihao Zhang, Shenghua He, Tian Xia, Shijia Pan, and Fei Liu. 2025. PlanGenLLMs: A Modern Survey of LLM Planning Capabilities. InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2025, Vienna, Austria, ...

  64. [72]

    Kimi Team, Tongtong Bai, Yifan Bai, Yiping Bao, SH Cai, Yuan Cao, Y Charles, HS Che, Cheng Chen, Guanduo Chen, et al . 2026. Kimi K2. 5: Visual Agentic Intelligence.arXiv preprint arXiv:2602.02276(2026)

  65. [73]

    Ching-Yi Tsai, Nicole Tacconi, Andrew D Wilson, and Parastoo Abtahi. 2026. Uncertain Pointer: Situated Feedforward Visualizations for Ambiguity-Aware AR Target Selection.arXiv preprint arXiv:2602.13433(2026)

  66. [74]

    Judith Wewerka and Manfred Reichert. 2020. Robotic Process Automation - A Systematic Literature Review and Assessment Framework.CoRRabs/2012.11951 (2020). arXiv:2012.11951 https://arxiv.org/abs/2012.11951

  67. [75]

    Rossi, Ruiyi Zhang, Subrata Mitra, Dimitris N

    Junda Wu, Zhehao Zhang, Yu Xia, Xintong Li, Zhaoyang Xia, Aaron Chang, Tong Yu, Sungchul Kim, Ryan A. Rossi, Ruiyi Zhang, Subrata Mitra, Dimitris N. Metaxas, Lina Yao, Jingbo Shang, and Julian J. McAuley. 2024. Visual Prompting in Multimodal Large Language Models: A Survey.CoR...

  68. [76]

    InThe Fourteenth International Conference on Learning Representations

    Computer Agent Arena: Toward Human-Centric Evaluation and Analysis of Computer-Use Agents. InThe Fourteenth International Conference on Learning Representations. https://openreview.net/forum?id=3x4SDbXbgl

  69. [77]

    Bovik, Hamid R

    Zhou Wang, Alan C. Bovik, Hamid R. Sheikh, and Eero P. Simoncelli. 2004. Image quality assessment: from error visibility to structural similarity.IEEE Trans. Image Process.13, 4 (2004), 600–612

  70. [78]

    Tianbao Xie, Danyang Zhang, Jixuan Chen, Xiaochuan Li, Siheng Zhao, Ruisheng Cao, Toh Jing Hua, Zhoujun Cheng, Dongchan Shin, Fangyu Lei, Yitao Liu, Yiheng Xu, Shuyan Zhou, Silvio Savarese, Caiming Xiong, Victor Zhong, and Tao Yu

  71. [79]

    Hao Wen, Yuanchun Li, Guohong Liu, Shanhui Zhao, Tao Yu, Toby Jia-Jun Li, Shiqi Jiang, Yunhao Liu, Yaqin Zhang, and Yunxin Liu. 2024. Autodroid: Llm-powered task automation in android. InProceedings of the 30th Annual International Con- ference on Mobile Computing and Networki...

  72. [80]

    Zhen Wen, Luoxuan Weng, Yinghao Tang, Runjin Zhang, Yuxin Liu, Bo Pan, Minfeng Zhu, and Wei Chen. 2025. Exploring Multimodal Prompt for Visual- ization Authoring with Large Language Models.CoRRabs/2504.13700 (2025). doi:10.48550/ARXIV.2504.13700

  73. [81]

    Yuhao Yang, Zhen Yang, Zi-Yi Dou, Anh Nguyen, Keen You, Omar Attia, Andrew Szot, Michael Feng, Ram Ramrakhya, Alexander Toshev, et al. 2025. Ultracua: A foundation model for computer use agents with hybrid action.arXiv preprint arXiv:2510.17790(2025)

  74. [82]

    Shunyu Yao, Howard Chen, John Yang, and Karthik Narasimhan. 2022. Webshop: Towards scalable real-world web interaction with grounded language agents. Advances in Neural Information Processing Systems35 (2022), 20744–20757

  75. [83]

    Penghao Wu, Shengnan Ma, Bo Wang, Jiaheng Yu, Lewei Lu, and Ziwei Liu. 2026. Gui-reflection: Empowering multimodal gui models with self-reflection behavior. Advances in Neural Information Processing Systems38 (2026), 101861–101896

  76. [84]

    Zhiyong Wu, Zhenyu Wu, Fangzhi Xu, Yian Wang, Qiushi Sun, Chengyou Jia, Kanzhi Cheng, Zichen Ding, Liheng Chen, Paul Pu Liang, and Yu Qiao. 2025. OS- ATLAS: Foundation Action Model for Generalist GUI Agents. InThe Thirteenth International Conference on Learning Representations...

  77. [85]

    Ryan Yen, Jian Zhao, and Daniel Vogel. 2025. Code Shaping: Iterative Code Editing with Free-form AI-Interpreted Sketching. InProceedings of the 2025 CHI Conference on Human Factors in Computing Systems, CHI 2025, YokohamaJapan, 26 Conference’17, July 2017, Washington, DC, USA ...

  78. [86]

    OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments. InAdvances in Neural Information Processing Systems 38: Annual Conference on Neural Information Processing Systems 2024, NeurIPS 2024, Vancouver, BC, Canada, December 10 - 15, 2024, Ami...

  79. [87]

    Yuan Xu, Shaowen Xiang, Yizhi Song, Ruoting Sun, and Xin Tong. 2026. DuetUI: A Bidirectional Context Loop for Human-Agent Co-Generation of Task-Oriented Interfaces.Proceedings of the 2026 CHI Conference on Human Factors in Computing Systems, CHI 2026, Barcelona, Spain, April 1...

  80. [88]

    John Yang, Carlos E Jimenez, Alexander Wettig, Kilian Lieret, Shunyu Yao, Karthik Narasimhan, and Ofir Press. 2024. Swe-agent: Agent-computer interfaces enable automated software engineering.Advances in Neural Information Processing Systems37 (2024), 50528–50652

  81. [89]

    Yichun Zhang, Xiangwu Guo, Yauhong Goh, Jessica Hu, Zhiheng Chen, Xin Wang, Difei Gao, and Mike Zheng Shou. 2026. ShowUI-Aloha: Human-Taught GUI Agent.CoRRabs/2601.07181 (2026)

  82. [90]

    Zhou Zhao, Shengyu Zhang, Liang Wang, Xiangxin Zhou, Zhaokai Wang, Kun Kuang, Fei Wu, Wangchunshu Zhou, Shuofei Qiao, Jiwei Li, Guoyin Wang, Ziyu Zhao, Hongxia Yang, Fan Wu, Jiasheng Ye, Shenzhi Wang, Ruixuan Xiao, Tieyong Zeng, Yuhuai Li, Yuchen Eleanor Jiang, Meiling Tao, Xu...

  83. [91]

    Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik R Narasimhan, and Yuan Cao. 2022. React: Synergizing reasoning and acting in language models. InThe eleventh international conference on learning representations

  84. [92]

    Tom Yeh, Tsung-Hsiang Chang, and Robert C. Miller. 2009. Sikuli: using GUI screenshots for search and automation. InProceedings of the 22nd Annual ACM Symposium on User Interface Software and Technology, Victoria, BC, Canada, Octo- ber 4-7, 2009, Andrew D. Wilson and François ...

  85. [94]

    Shengcheng Yu, Chunrong Fang, Mingzhe Du, Yuchen Ling, Zhenyu Chen, and Zhendong Su. 2024. Practical Non-Intrusive GUI Exploration Testing with Visual- based Robotic Arms. InProceedings of the 46th IEEE/ACM International Conference on Software Engineering, ICSE 2024, Lisbon, P...

  86. [95]

    Hyeonggeun Yun and Jinkyu Jang. 2025. Interaction-Driven Browsing: A Human- in-the-Loop Conceptual Framework Informed by Human Web Browsing for Browser-Using Agents.CoRRabs/2509.12049 (2025). doi:10.48550/ARXIV.2509. 12049

  87. [96]

    Shaojie Zhang, Ruoceng Zhang, Pei Fu, Shaokang Wang, Jiahui Yang, Xin Du, Bin Qin, Ying Huang, Zhenbo Luo, and Jian Luan. 2026. Btl-ui: Blink-think-link reasoning model for gui agent.Advances in Neural Information Processing Systems 38 (2026), 56035–56056

  88. [99]

    Shuyan Zhou, Frank F Xu, Hao Zhu, Xuhui Zhou, Robert Lo, Abishek Sridhar, Xianyi Cheng, Tianyue Ou, Yonatan Bisk, Daniel Fried, et al. 2024. Webarena: A realistic web environment for building autonomous agents. InInternational Conference on Learning Representations, Vol. 2024....

  89. [100]

    Added” and “Changed

    Henry Peng Zou, Wei-Chieh Huang, Yaozu Wu, Jizhou Guo, Yankai Chen, Chunyu Miao, Hoang H Nguyen, Yue Zhou, Weizhi Zhang, Liancheng Fang, et al. 2026. Llm-based human-agent collaboration and interaction systems: A survey.Findings of the Association for Computational Linguistics...

  90. [101]

    SUMMARY: <one short imperative step sentence>

  91. [102]

    I will”, “Let’s

    RATIONALE: <1–2 sentences explaining the detected failure and why the proposed next action helps> - The SUMMARY must: •start with a strong action verb •be written as a standalone executable step •not contain “I will”, “Let’s”, or future tense •not mention internal tool names -...

  92. [2024]

    doi:10.1109/CVPR52733.2024.01227

    IEEE, 12914–12923. doi:10.1109/CVPR52733.2024.01227

  93. [2025]

    Conference’17, July 2017, Washington, DC, USA Venkatesan et al

    Navi-plus: Managing Ambiguous GUI Navigation Tasks with Follow-up Questions.arXiv preprint arXiv:2503.24180(2025). Conference’17, July 2017, Washington, DC, USA Venkatesan et al

  94. [2026]

    In The Fourteenth International Conference on Learning Representations

    Just Do It!? Computer-Use Agents Exhibit Blind Goal-Directedness. In The Fourteenth International Conference on Learning Representations. https: //openreview.net/forum?id=9W4bPRsEIT

Pith tools

Reviewed August 1, 2026 · model on record in the stance chip above.