Pith. sign in

REVIEW 3 major objections 5 minor 63 references

SkillLens: Visual Skill Cards for Retrieval-Augmented GUI Action Prediction and On-Policy Distillation

T0 review · 3 major / 5 minor · reviewed 2026-08-12 · deepseek-v4-flash

Pith's one-line read Visual Skill Cards give frozen GUI agents reusable visual procedural memory and lift step accuracy by up to 11.6 points on Multimodal-Mind2Web.

desk verdict A well-ablated, plausible memory interface for GUI agents; the headline numbers need a clearer statement that evaluation episodes are excluded from the card library. read the letter →

arxiv 2608.10775 v1 pith:QAWZ7HAW submitted 2026-08-11 cs.AI

classification cs.AI
keywords VisualSkillCardsGUIactionpredictionretrieval-augmentedinferenceon-policydistillationproceduralmemorycomputer-useagentsvisual-languagemodelsgrounding
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper argues that computer-using agents lack visual procedural memory: they can see controls but cannot tell which familiar workflow is active, which control matters next, or what evidence confirms progress. It introduces Visual Skill Cards (VSCs), a state-conditioned memory object that bundles a reusable procedure with applicability cues, visual evidence, and verification signals. The paper claims that retrieving and selectively expanding the relevant card lets a frozen visual-language model executor choose grounded GUI actions more accurately, and that the same cards can serve as privileged teacher context to distill that behavior into a smaller student that runs without retrieval. If true, external visual skill libraries would offer a practical, auditable upgrade path for existing GUI agents without retraining them.

What carries the argument

The load-bearing object is the Visual Skill Card, written $s_i = (p_i, z_i, v_i, \kappa_i)$: a procedure $p_i$, state cues $z_i$ (applicability and verification), visual evidence views $v_i$, and optional auxiliary fields $\kappa_i$. Trace-to-VSC converts heterogeneous interaction traces into this common schema, segmenting them into reusable units and auditing that held-out evaluations do not expose answer coordinates as templates. At inference, SkillLens uses a lightweight context-aware selector that scores cards by token overlap with the task query, then reranks and expands only the top candidates' evidence, so retrieval decides where to look while expansion controls how much visual evidence the frozen executor sees. For distillation, CardDistill gives the teacher the card bundle and the student only the benchmark-native context, optimizing a teacher-confidence-weighted reverse KL over student-generated action prefixes.

What would settle it

Build the card library from episodes on one set of sites and evaluate on a disjoint set of sites with the same task types, or strip every after-state and verification crop from the cards: if the improvement mostly collapses, the cards are leaking benchmark-specific outcome templates rather than supplying transferable visual procedures.

Watch

Extended reading notes

Core claim

The central discovery is that a VSC — a card pairing a procedure with when-it-applies cues, before/after visual evidence, and verification signals — works both as external runtime memory and as training-time privilege. On Multimodal-Mind2Web and WebLINX-BrowserGym, adding retrieved VSCs to a frozen GPT-5.4-mini executor raises Step SR by +11.6 points and Overall by +2.9 points, and the same evidence used in CardDistill raises the student-only Qwen3-VL-2B metrics by +12.0 and +3.2 points. The paper further shows that the gains require relevance: random or irrelevant cards do not reproduce them and can actively hurt grounding. The design separates a lightweight retrieval step from selective high-resolution evidence expansion, keeping the live screen as the final grounding source and bounding runtime evidence.

Load-bearing premise

The gains depend on the before/after visual evidence in a card being reusable procedural memory rather than a stored picture of the correct answer; the paper does not say how after-state screenshots and verification cues are kept from encoding the target outcome for the episodes that were used to build the library.

Editorial extensions

If this is right

  • Frozen GUI executors can be upgraded without parameter updates: the paper shows positive Step SR / Overall gains across Qwen, Gemini, and GPT models when relevant VSCs are retrieved and expanded.
  • CardDistill shows the retrieved behavior can be internalized: a Qwen3-VL-2B student improves by +12.0 Step SR on Mind2Web and +3.2 Overall on WebLINX-BG while running without any runtime card retrieval.
  • Relevance is necessary: negative controls with random or irrelevant cards fail to reproduce the gains and can reduce grounding accuracy, so the effect is not simply extra images or longer prompts.
  • The retrieve-then-expand cost split bounds runtime evidence while preserving high-resolution visual detail, making the approach tractable for step-by-step interactive agents.
  • Verification cues give a lightweight contract for checking progress after an action, which could be used for self-monitoring as well as action prediction.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A direct cross-site transfer test — build cards from site A, evaluate on site B — would sharpen the claim that the cards store reusable procedures rather than site-specific layouts; the paper does not report such a test.
  • If the card library is the only source of task knowledge, the same framework could in principle be pointed at new domains (mobile UI, design tools, enterprise software) by swapping in new traces, without retraining the executor.
  • The teacher-confidence-weighted reverse KL objective could be applied to other privileged signals beyond VSCs, such as ground-truth element boxes or future-state crops, to test how general the distillation recipe is.
  • The board-based selector diagnostics suggest an alternative route: a VLM reading a rendered candidate board can improve retrieval-only Hit@1 even though it underperforms in end-to-end execution, so a better board layout or hybrid scoring might convert that retrieval gain into downstream accuracy.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper introduces Visual Skill Cards (VSCs), a memory representation that binds a reusable procedure with applicability cues, visual before/target/after evidence, and verification signals. Trace-to-VSC converts heterogeneous interaction traces into this schema; SkillLens retrieves and selectively expands VSC evidence for a frozen VLM executor; CardDistill uses the same evidence as privileged teacher context for on-policy distillation into a student that runs without retrieval. Experiments on Multimodal-Mind2Web, WebLINX-BG, and OSWorld-G report consistent improvements over no-skill baselines across several frozen executors, plus student-only gains for Qwen3-VL-2B, with negative controls and modality ablations.

Significance. If the results hold, this is a useful contribution: VSCs offer a common, auditable unit for heterogeneous GUI experience; the separation of low-cost retrieval from high-resolution evidence expansion is pragmatic; and CardDistill comes with confidence intervals and a shuffled-VSC control. The negative-control experiments in Table 5 and the retrieval-coverage analysis in Table 3 are genuine strengths. However, because the main quantitative claims hinge on a single table without error bars and on the assumption that the VSC library does not contain evaluation-episode answers, the current evidence does not yet establish that the gains come from reusable visual procedural memory rather than from episode-specific answer lookup or from text-only procedures.

major comments (3)
  1. [Trace-to-VSC Construction / Eq. (7)] The VSC schema stores after-state screenshots and verification cues, and the executor is conditioned on this evidence before action selection (Eq. (14)). The paper evaluates on the same benchmarks used to build the library and only asserts that 'held-out evaluations do not expose answer coordinates as templates.' This does not rule out the after-state screenshot or verification cue acting as an episode-specific answer lookup: for a card built from a Mind2Web or WebLINX trace, the after-state can be the ground-truth next observation of that very episode. Please specify exactly how the construction/evaluation split is managed, and add a control in which the library is built from sources provably disjoint from the evaluation episodes (or from the training split with a statement that no evaluation episode contributed). Without such a control, the +11.6 and +2.9 headline gains could be explained by answer lookup rather than reusable procedural memory.
  2. [Table 4] Table 4's modality ablation directly undercuts the 'visual' half of the central claim. On Mind2Web, removing visual evidence improves over Full SkillLens for both Qwen3-VL-2B (Step SR 14.0 vs 10.3; Elem. Acc. 67.5 vs 62.2) and Gemini 2.5 Flash (Step SR 75.0 vs 66.2; Elem. Acc. 86.0 vs 77.7); for Qwen3-VL-2B, image-only cards (w/o procedure) give Step SR 4.5, essentially identical to the 4.6 no-skill baseline. The Step SR gain on Mind2Web is therefore carried by procedure text, not visual evidence. Please either explain this pattern or reframe the contribution as mixed procedural memory, and report the headline result with the text-only variant disambiguated.
  3. [Table 1 / Experimental Setup] The main results in Table 1 are single-run percentages with no repeated-seed or error-bar information, while the CardDistill results are reported with 95% CIs and p-values. Since the abstract's headline numbers come from Table 1, please report variance estimates (e.g., multiple evaluation subsets or seeds) or state why the evaluation is deterministic and why variance is negligible. This is needed to assess whether the +11.6, +2.9, +12.0, and +3.2 deltas are statistically meaningful.
minor comments (5)
  1. [Eq. (15)] The objective is written as KL(P_theta || P_phi) and called 'reverse KL'; the usual reverse-KL direction is KL(teacher || student). Please clarify the intended direction or adjust the terminology.
  2. [Table 3] All selectors report Oracle Cov. 1.00 on every benchmark. Please discuss whether this means retrieval recall is saturated in these settings and whether the observed gains should be attributed primarily to evidence expansion rather than to retrieval.
  3. [Trace-to-VSC Construction] The construction pipeline (trace normalization, segmentation, summarization, evidence selection, audit) is described at a high level. Please provide concrete adapter details for Mind2Web and WebLINX, such as which LLM or VLM performs summarization and what segmentation rules are used, so the method is reproducible.
  4. [Table 2] The latency numbers (1.5s vs 4.0s) are reported without hardware details. Please specify the inference stack and hardware used for the timing comparison.
  5. [Figure 3] The caption mentions 'training dynamics' but the text does not describe the plotted curves. Either describe the figure in the body or point to the supplementary material for a full explanation.

Circularity Check

1 steps flagged · score 4.0 of 10

VSC 'After' evidence can be the target observation from the same benchmark trace, and the audit only excludes coordinate templates, so the central gains may partly be answer lookup rather than reusable procedure.

  1. fitted input called prediction [Method, Trace-to-VSC Construction; VSC Representation (Eq. 7); Grounded Action Prediction (Eq. 14); Figure 1 Before-Target-After evidence]
    "Trace-to-VSC converts prior interaction records into reusable visual procedures... a structured benchmark can bind targets deterministically... This audit checks that the card is self-contained and that held-out evaluations do not expose answer coordinates as templates. Verification cues define observable postconditions for a successful step. The visual evidence may include a full interface view or a focused target crop... Expanded evidence is reference material, not a coordinate template."

    The VSC schema stores vi as visual evidence, and the Before-Target-After construction (Figure 1) places the postcondition after-state ('Added to cart!') in the card. When a structured benchmark like Mind2Web 'binds targets deterministically,' that after-state can be the ground-truth next observation of the very episode being predicted. Eq. (14) then conditions the frozen executor on Et expanded from the retrieved card, so the 'prediction' at step t is made with the correct outcome already in context. The stated audit only excludes 'answer coordinates as templates'; it does not exclude after-state images or verification cues, and no disjoint-source control is reported.

full rationale

The paper's retrieval and expansion machinery is explicit and mostly self-contained, and its negative controls (random/irrelevant VSCs) plus the shuffled-VSC CardDistill control show that card relevance and alignment, not just prompt length, drive the effect. There is no load-bearing self-citation or imported uniqueness theorem. The central risk is a provenance gap: the VSC library is built from Mind2Web and WebLINX traces, the same benchmarks used for evaluation, and VSC evidence includes postcondition after-state screenshots and verification cues. If any evaluation episode contributed a card, or if cards from the same episode distribution encode the target next observation, then Eq. (14)'s prediction is conditioned on answer-bearing content and the +11.6/+2.9 gains are partly fitted, not predicted. The paper's audit sentence is an assertion about 'answer coordinates as templates' and does not describe a mechanism or measurement covering after-state images; Table 4 even shows text-only cards sometimes beat full visual cards, which leaves the provenance question unresolved rather than resolved. This is not a formal circularity of equations, but it is a partial reduction of the central claim to its own benchmark-derived inputs, so the score is 4.

Assumptions & free parameters 3 free parameters · 4 assumptions · 1 invented entities

The paper introduces one main construct, the VSC, and relies on several unverified assumptions about leakage, audit effectiveness, and executor behavior. Free parameters are hand-chosen hyperparameters: rerank weight, sampling budgets, and distillation temperature. The most consequential assumption is that after-state evidence does not leak the target answer, which is not fully mitigated by the stated audit.

free parameters (3)
  • Rerank blend weight = 0.1
    The rerank score in Equation 11, r(1)_i = |M_{t,i}| + 0.1 r(0)_i, uses a hand-chosen weight between the exact-overlap term and the coarse score; no sensitivity analysis or tuning study is provided.
  • Retrieval and expansion budgets Kc, Ks, Ke, Kv = not specified in text
    The top-K budgets control how many candidate cards, selected cards, expanded cards, and evidence images reach the executor (Equations 4, 10, 11). Values are declared fixed per run but never stated, and no ablation on them is reported.
  • Distillation temperature T = 1.0
    The CardDistill objective in Equation 15 uses temperature T=1.0 in the token distributions; the paper states this value without reporting a sweep.
assumptions (4)
  • domain assumption The VSC library constructed from public traces does not leak held-out answer coordinates; the paper states an audit checks this.
    Trace-to-VSC Construction: 'This audit checks that the card is self-contained and that held-out evaluations do not expose answer coordinates as templates.' This is asserted, not demonstrated with released artifacts.
  • domain assumption Feeding postcondition ('After') screenshots as reference evidence is legitimate and does not reveal the target action.
    Figure 1 shows Before-Target-After evidence, and the runtime feeds this evidence to the frozen executor before action prediction. The after-state could act as an answer key for tasks on the same sites.
  • domain assumption Frozen VLM executors can usefully consume retrieved VSC evidence and still ground actions on the live screen.
    The inference formulation in Equations 6 and 14 assumes the fixed executor conditions on (xi_t, S_t, E_t) without retraining and that the live screen remains the final grounding source.
  • domain assumption Reverse KL with teacher-confidence weighting transfers card-conditioned behavior to the student.
    CardDistill objective in Equation 15 assumes imitation of teacher token distributions, with weights from Equation 16, improves student-only metrics at evaluation time.
invented entities (1)
  • Visual Skill Card (VSC) independent evidence
    purpose: State-conditioned memory object binding procedure text, applicability cues, visual evidence, and verification signals for GUI skill reuse.
    The card is a new representational unit. Independent evidence comes from negative controls (Table 5) and ablations (Table 4), which show relevance matters, though the after-state evidence may partially encode outcomes.

how reviews work

0 comments
Cite this review

Pith. "Pith review of SkillLens: Visual Skill Cards for Retrieval-Augmented GUI Action Prediction and On-Policy Distillation." pith.science (2026). https://pith.science/paper/QAWZ7HAW

@misc{pith2026260810775,
  author       = {Pith},
  title        = {Pith review of: SkillLens: Visual Skill Cards for Retrieval-Augmented GUI Action Prediction and On-Policy Distillation},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/QAWZ7HAW}},
  note         = {Machine review of arXiv:2608.10775}
}
read the original abstract

Computer-using agents can perceive rich software interfaces, yet their decisions often lack visual procedural memory: they may recognize individual controls without identifying which familiar workflow is active, which control matters next, or what evidence would confirm progress. Raw interaction traces preserve such information but are long and noisy to condition on, whereas text-only skills often omit the visual state that makes a procedure applicable. We introduce Visual Skill Cards (VSCs), a state-conditioned memory representation that binds reusable procedures with applicability cues, visual evidence, and verification signals. SkillLens constructs VSCs from heterogeneous interaction experience through Trace-to-Visual-Skill-Card and, at inference time, retrieves relevant cards and selectively expands only the evidence needed by a fixed visual-language model executor for grounded GUI action prediction. The same representation also supports CardDistill, which uses VSC evidence as privileged teacher context to train a student that acts without runtime card retrieval. Across Multimodal-Mind2Web and WebLINX-BrowserGym, SkillLens improves the frozen GPT-5.4-mini executor by +11.6 points in Step SR and +2.9 points in Overall, respectively; CardDistill further improves the corresponding student-only Qwen3-VL-2B metrics by +12.0 and +3.2 points.

Figures

Figures reproduced from arXiv: 2608.10775 by the authors.

Figure 1
Figure 1. Representative SkillLens case. Visual Skill Card [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. Overview of SkillLens. Source-specific adapters convert public traces and annotations into Visual Skill Cards (VSCs); [PITH_FULL_IMAGE:figures/full_fig_p002_2.png] view at source ↗
Figure 3
Figure 3. CardDistill training dynamics and student-only [PITH_FULL_IMAGE:figures/full_fig_p005_3.png] view at source ↗
Figures from the paper (2 more)
Figure 4
Figure 4. Figure 4: Diagnostic grounding reference with external GUI [PITH_FULL_IMAGE:figures/full_fig_p007_4.png]
Figure 5
Figure 5. Figure 5: VSC corpus statistics. The profile summarizes cor [PITH_FULL_IMAGE:figures/full_fig_p007_5.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

63 extracted references · 43 canonical work pages

  1. [2]

    doi:10.48550/arXiv.2601.21123 , url=

    Chen, Tianyi and Li, Yinheng and Solodko, Michael and Wang, Sen and Jiang, Nan and Cui, Tingyuan and Hao, Junheng and Ko, Jongwoo and Abdali, Sara and Xu, Leon and Zheng, Suzhen and Fan, Hao and Cameron, Pashmina and Wagle, Justin and Koishida, Kazuhito , year=. doi:10.48550/arXiv.2601.21123 , url=. 2601.21123 , archivePrefix=

  2. [4]

    Learn where to Click from Yourself: On-Policy Self-Distillation for

    Zhang, Yan and Wu, Daiqing and Shen, Huawen and Ma, Can and Zhou, Yu , year=. Learn where to Click from Yourself: On-Policy Self-Distillation for. doi:10.48550/arXiv.2605.00642 , url=. 2605.00642 , archivePrefix=

  3. [5]

    Vision-OPD: Learning to See Fine Details for Multimodal

    Yuan, Qianhao and Lou, Jie and Yu, Xing and Lin, Hongyu and Sun, Le and Han, Xianpei and Lu, Yaojie , journal=. Vision-OPD: Learning to See Fine Details for Multimodal

  4. [7]

    Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=

    OCR-Memory: Optical Context Retrieval for Long-Horizon Agent Memory , author=. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=. 2026 , doi=

  5. [8]

    LensVLM: Selective Context Expansion for Compressed Visual Representation of Text

    Xie, Roy and Friedman, Dan and Yu, Donghan and Pan, Bowen and Fifty, Christopher and Kim, Jang-Hyun and Du, Xianzhi and Gan, Zhe and Rathod, Vivek and Dhingra, Bhuwan , journal=. 2026 , eprint=. doi:10.48550/arXiv.2605.07019 , url=

  6. [9]

    Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=

    AgentOCR: Reimagining Agent History via Optical Self-Compression , author=. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=. 2026 , doi=

  7. [10]

    Proceedings of the 43rd International Conference on Machine Learning , year=

    Shi, Yaorui and Liu, Shugui and Yang, Yu and Mao, Wenyu and Chen, Yuxin and Qi. Proceedings of the 43rd International Conference on Machine Learning , year=. doi:10.48550/arXiv.2601.21468 , url=. 2601.21468 , archivePrefix=

  8. [11]

    Advances in Neural Information Processing Systems , volume=

    VideoGUI: A Benchmark for GUI Automation from Instructional Videos , author=. Advances in Neural Information Processing Systems , volume=. 2024 , doi=

Show all 63 references
  1. [12]

    Advances in Neural Information Processing Systems , volume=

    OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments , author=. Advances in Neural Information Processing Systems , volume=. 2024 , doi=

  2. [13]

    Proceedings of the 42nd International Conference on Machine Learning , series=

    Windows Agent Arena: Evaluating Multi-Modal OS Agents at Scale , author=. Proceedings of the 42nd International Conference on Machine Learning , series=. 2025 , url=

  3. [14]

    2025 , eprint=

    Wang, Haoming and Zou, Haoyang and Song, Huatong and Feng, Jiazhan and Fang, Junjie and Lu, Junting and Liu, Longxiang and Luo, Qinyu and Liang, Shihao and Huang, Shijue and Zhong, Wanjun and Ye, Yining and Qin, Yujia and Xiong, Yuwen and Song, Yuxin and Wu, Zhiyong and Li, Ao...

  4. [15]

    2026 , address=

    Liang, Weijie and Song, Yuanfeng and Chen, Xing and Cao, Caleb Chen and Han, Sirui and Guo, Yike , booktitle=. 2026 , address=

  5. [16]

    Advances in Neural Information Processing Systems , volume=

    Mind2Web: Towards a Generalist Agent for the Web , author=. Advances in Neural Information Processing Systems , volume=. 2023 , url=

  6. [17]

    2024 , url=

    Zheng, Boyuan and Gou, Boyu and Kil, Jihyung and Sun, Huan and Su, Yu , booktitle=. 2024 , url=. 2401.01614 , archivePrefix=

  7. [18]

    Proceedings of the 41st International Conference on Machine Learning , series=

    L. Proceedings of the 41st International Conference on Machine Learning , series=. 2024 , url=. 2402.05930 , archivePrefix=

  8. [19]

    Xu and Siva Reddy and Graham Neubig and Quentin Cappart and Russ Salakhutdinov and Nicolas Chapados , journal=

    Thibault Le Sellier de Chezelles and Maxime Gasse and Alexandre Lacoste and Massimo Caccia and Alexandre Drouin and Léo Boisvert and Megh Thakkar and Tom Marty and Rim Assouel and Sahar Omidi Shayegan and Lawrence Keunho Jang and Xing Han Lù and Ori Yoran and Dehan Kong and Fr...

  9. [20]

    Advances in Neural Information Processing Systems , volume=

    Scaling Computer-Use Grounding via User Interface Decomposition and Synthesis , author=. Advances in Neural Information Processing Systems , volume=. 2025 , url=

  10. [21]

    2024 , address=

    Cheng, Kanzhi and Sun, Qiushi and Chu, Yougang and Xu, Fangzhi and Li, Yantao and Zhang, Jianbing and Wu, Zhiyong , booktitle=. 2024 , address=. doi:10.18653/v1/2024.acl-long.505 , url=. 2401.10935 , archivePrefix=

  11. [22]

    ScreenSpot-Pro:

    Li, Kaixin and Meng, Ziyang and Lin, Hongzhan and Luo, Ziyang and Tian, Yuchen and Ma, Jing and Huang, Zhiyong and Chua, Tat-Seng , booktitle=. ScreenSpot-Pro:. 2025 , doi=

  12. [23]

    2025 , eprint=

    Qin, Yujia and Ye, Yining and Fang, Junjie and Wang, Haoming and Liang, Shihao and Tian, Shizuo and Zhang, Junda and Li, Jiahao and Li, Yunxin and Huang, Shijue and Zhong, Wanjun and Li, Kuanye and Yang, Jiale and Miao, Yu and Lin, Woyu and Liu, Longxiang and Jiang, Xu and Ma,...

  13. [24]

    2025 , url=

    Wu, Zhiyong and Wu, Zhenyu and Xu, Fangzhi and Wang, Yian and Sun, Qiushi and Jia, Chengyou and Cheng, Kanzhi and Ding, Zichen and Chen, Liheng and Liang, Paul Pu and Qiao, Yu , booktitle=. 2025 , url=

  14. [25]

    Navigating the Digital World as Humans Do: Universal Visual Grounding for

    Gou, Boyu and Wang, Ruohan and Zheng, Boyuan and Xie, Yanan and Chang, Cheng and Shu, Yiheng and Sun, Huan and Su, Yu , booktitle=. Navigating the Digital World as Humans Do: Universal Visual Grounding for. 2025 , url=

  15. [26]

    2025 , url=

    Wu, Qianhui and Cheng, Kanzhi and Yang, Rui and Zhang, Chaoyun and Yang, Jianwei and Jiang, Huiqiang and Mu, Jian and Peng, Baolin and Qiao, Bo and Tan, Reuben and Qin, Si and Liden, Lars and Lin, Qingwei and Zhang, Huan and Zhang, Tong and Zhang, Jianbing and Zhang, Dongmei a...

  16. [27]

    2024 , eprint=

    Wang, Peng and Bai, Shuai and Tan, Sinan and Wang, Shijie and Fan, Zhihao and Bai, Jinze and Chen, Keqin and Liu, Xuejing and Wang, Jialin and Ge, Wenbin and Fan, Yang and Dang, Kai and Du, Mengfei and Ren, Xuancheng and Men, Rui and Liu, Dayiheng and Zhou, Chang and Zhou, Jin...

  17. [28]

    2025 , eprint=

    Bai, Shuai and Chen, Keqin and Liu, Xuejing and Wang, Jialin and Ge, Wenbin and Song, Sibo and Dang, Kai and Wang, Peng and Wang, Shijie and Tang, Jun and Zhong, Humen and Zhu, Yuanzhi and Yang, Mingkun and Li, Zhaohai and Wan, Jianqiang and Wang, Pengfei and Ding, Wei and Fu,...

  18. [29]

    2025 , eprint=

    Bai, Shuai and Cai, Yuxuan and Chen, Ruizhe and Chen, Keqin and Chen, Xionghui and Cheng, Zesen and Deng, Lianghao and Ding, Wei and Gao, Chang and Ge, Chunjiang and Ge, Wenbin and Guo, Zhifang and Huang, Qidong and Huang, Jie and Huang, Fei and Hui, Binyuan and Jiang, Shutong...

  19. [30]

    2025 , type=

    Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities , author=. 2025 , type=

  20. [31]

    and Yang, Zhilin and Yu, Tao , booktitle=

    Wang, Xinyuan and Wang, Bowen and Lu, Dunjie and Yang, Junlin and Xie, Tianbao and Wang, Junli and Deng, Jiaqi and Guo, Xiaole and Xu, Yiheng and Wu, Chen and Shen, Zhennan and Li, Zhuokai and Li, Ryan and Li, Xiaochuan and Chen, Junda and Zheng, Boyuan and Li, Peihang and Lei...

  21. [33]

    ToolCUA: Towards Optimal

    Hu, Xuhao and Zhang, Xi and Xu, Haiyang and Qiao, Kyle and Yang, Jingyi and Huang, Xuanjing and Shao, Jing and Yan, Ming and Ye, Jieping , journal=. ToolCUA: Towards Optimal

  22. [35]

    Bai, S.; Cai, Y.; Chen, R.; Chen, K.; Chen, X.; Cheng, Z.; Deng, L.; Ding, W.; Gao, C.; Ge, C.; Ge, W.; Guo, Z.; Huang, Q.; Huang, J.; Huang, F.; Hui, B.; Jiang, S.; Li, Z.; Li, M.; Li, M.; Li, K.; Lin, Z.; Lin, J.; Liu, X.; Liu, J.; Liu, C.; Liu, Y.; Liu, D.; Liu, S.; Lu, D.;...

  23. [36]

    Bai, S.; Chen, K.; Liu, X.; Wang, J.; Ge, W.; Song, S.; Dang, K.; Wang, P.; Wang, S.; Tang, J.; Zhong, H.; Zhu, Y.; Yang, M.; Li, Z.; Wan, J.; Wang, P.; Ding, W.; Fu, Z.; Xu, Y.; Ye, J.; Zhang, X.; Xie, T.; Cheng, Z.; Zhang, H.; Yang, Z.; Xu, H.; and Lin, J. 2025 b . Qwen2.5-V...

  24. [37]

    K.; and Hui, Z

    Bonatti, R.; Zhao, D.; Bonacci, F.; Dupont, D.; Abdali, S.; Li, Y.; Lu, Y.; Wagle, J.; Koishida, K.; Bucker, A.; Jang, L. K.; and Hui, Z. 2025. Windows Agent Arena: Evaluating Multi-Modal OS Agents at Scale. In Proceedings of the 42nd International Conference on Machine Learni...

  25. [38]

    Cai, Y.; Liu, J.; Liu, Y.; Deng, H.; Yao, L.; Zheng, Y.; Ouyang, K.; Li, Z.; Wang, Z.; Sun, X.; Bai, H.; and Li, X. 2026. Thinking Without Images: Internalizing Visual Manipulation with On-Policy Self-Distillation. arXiv preprint arXiv:2606.08719

  26. [39]

    Chen, T.; Li, Y.; Solodko, M.; Wang, S.; Jiang, N.; Cui, T.; Hao, J.; Ko, J.; Abdali, S.; Xu, L.; Zheng, S.; Fan, H.; Cameron, P.; Wagle, J.; and Koishida, K. 2026. CUA -Skill: Develop Skills for Computer Using Agent. arXiv:2601.21123

  27. [40]

    Cheng, K.; Sun, Q.; Chu, Y.; Xu, F.; Li, Y.; Zhang, J.; and Wu, Z. 2024. SeeClick : Harnessing GUI Grounding for Advanced Visual GUI Agents. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 9313--9332. Bangkok,...

  28. [41]

    de Chezelles, T. L. S.; Gasse, M.; Lacoste, A.; Caccia, M.; Drouin, A.; Boisvert, L.; Thakkar, M.; Marty, T.; Assouel, R.; Shayegan, S. O.; Jang, L. K.; Lù, X. H.; Yoran, O.; Kong, D.; Xu, F. F.; Reddy, S.; Neubig, G.; Cappart, Q.; Salakhutdinov, R.; and Chapados, N. 2025. The...

  29. [42]

    Deng, X.; Gu, Y.; Zheng, B.; Chen, S.; Stevens, S.; Wang, B.; Sun, H.; and Su, Y. 2023. Mind2Web: Towards a Generalist Agent for the Web. In Advances in Neural Information Processing Systems, volume 36, 28091--28114

  30. [43]

    Feng, L.; Yang, F.; Chen, F.; Cheng, X.; Xu, H.; Wan, Z.; Yan, M.; and An, B. 2026. AgentOCR: Reimagining Agent History via Optical Self-Compression. In Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 5067--5086....

  31. [44]

    Gemini Team, Google . 2025. Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities. Technical report, Google

  32. [45]

    Gou, B.; Wang, R.; Zheng, B.; Xie, Y.; Chang, C.; Shu, Y.; Sun, H.; and Su, Y. 2025. Navigating the Digital World as Humans Do: Universal Visual Grounding for GUI Agents. In The Thirteenth International Conference on Learning Representations

  33. [46]

    Hu, X.; Zhang, X.; Xu, H.; Qiao, K.; Yang, J.; Huang, X.; Shao, J.; Yan, M.; and Ye, J. 2026. ToolCUA: Towards Optimal GUI -Tool Path Orchestration for Computer Use Agents. arXiv preprint arXiv:2605.12481

  34. [47]

    Jiang, Z.; An, L.; Liu, Y.; Ji, J.; Wu, Q.; Andreas, J.; Zhang, Y.; and Chang, S. 2026. VISUALSKILL: Multimodal Skills for Computer-Use Agents. arXiv preprint arXiv:2606.18448

  35. [48]

    Li, J.; Zhang, Y.; Yang, X.; QU, J.; Xu, J.; Yang, S.; Ding, J.; and Ngai, E. C.-H. 2026. OCR-Memory: Optical Context Retrieval for Long-Horizon Agent Memory. In Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 10...

  36. [49]

    C.; Han, S.; and Guo, Y

    Liang, W.; Song, Y.; Chen, X.; Cao, C. C.; Han, S.; and Guo, Y. 2026. V izo M em: A Visual-Textual Memory Framework for Efficient Long-Horizon Reasoning. In Findings of the Association for Computational Linguistics: ACL 2026, 7399--7422. San Diego, California, United States: A...

  37. [50]

    Q.; Li, L.; Gao, D.; Wu, Q.; Yan, M.; Yang, Z.; Wang, L.; and Shou, M

    Lin, K. Q.; Li, L.; Gao, D.; Wu, Q.; Yan, M.; Yang, Z.; Wang, L.; and Shou, M. Z. 2024. VideoGUI: A Benchmark for GUI Automation from Instructional Videos. In Advances in Neural Information Processing Systems, volume 37, 69329--69360

  38. [51]

    H.; Kasner, Z.; and Reddy, S

    L \`u , X. H.; Kasner, Z.; and Reddy, S. 2024. WebLINX : Real-World Website Navigation with Multi-Turn Dialogue. In Proceedings of the 41st International Conference on Machine Learning, volume 235 of Proceedings of Machine Learning Research, 33007--33056. PMLR

  39. [52]

    OpenAI . 2024. GPT-4o System Card

  40. [53]

    Qin, Y.; Ye, Y.; Fang, J.; Wang, H.; Liang, S.; Tian, S.; Zhang, J.; Li, J.; Li, Y.; Huang, S.; Zhong, W.; Li, K.; Yang, J.; Miao, Y.; Lin, W.; Liu, L.; Jiang, X.; Ma, Q.; Li, J.; Xiao, X.; Cai, K.; Li, C.; Zheng, Y.; Jin, C.; Li, C.; Zhou, X.; Wang, M.; Chen, H.; Li, Z.; Yang...

  41. [54]

    Shi, Y.; Liu, S.; Yang, Y.; Mao, W.; Chen, Y.; GU , Q.; Su, H.; Cai, X.; Wang, X.; and Zhang, A. 2026. MemOCR : Layout-Aware Visual Memory for Efficient Long-Horizon Reasoning. In Proceedings of the 43rd International Conference on Machine Learning

  42. [55]

    Wang, H.; Zou, H.; Song, H.; Feng, J.; Fang, J.; Lu, J.; Liu, L.; Luo, Q.; Liang, S.; Huang, S.; Zhong, W.; Ye, Y.; Qin, Y.; Xiong, Y.; Song, Y.; Wu, Z.; Li, A.; Li, B.; Dun, C.; Liu, C.; Zan, D.; Leng, F.; Wang, H.; Yu, H.; Chen, H.; Guo, H.; Su, J.; Huang, J.; Shen, K.; Shi,...

  43. [56]

    Wang, P.; Bai, S.; Tan, S.; Wang, S.; Fan, Z.; Bai, J.; Chen, K.; Liu, X.; Wang, J.; Ge, W.; Fan, Y.; Dang, K.; Du, M.; Ren, X.; Men, R.; Liu, D.; Zhou, C.; Zhou, J.; and Lin, J. 2024. Qwen2-VL : Enhancing Vision-Language Model's Perception of the World at Any Resolution. arXi...

  44. [57]

    Wang, X.; Wang, B.; Lu, D.; Yang, J.; Xie, T.; Wang, J.; Deng, J.; Guo, X.; Xu, Y.; Wu, C.; Shen, Z.; Li, Z.; Li, R.; Li, X.; Chen, J.; Zheng, B.; Li, P.; Lei, F.; Cao, R.; Fu, Y.; Shin, D.; Shin, M.; Hu, J.; Wang, Y.; Chen, J.; Ye, Y.; Zhang, D.; Wang, Y.; Wang, H.; Yang, D.;...

  45. [58]

    Wei, H.; Sun, Y.; and Li, Y. 2025. DeepSeek-OCR: Contexts Optical Compression. arXiv preprint arXiv:2510.18234

  46. [59]

    Wu, Q.; Cheng, K.; Yang, R.; Zhang, C.; Yang, J.; Jiang, H.; Mu, J.; Peng, B.; Qiao, B.; Tan, R.; Qin, S.; Liden, L.; Lin, Q.; Zhang, H.; Zhang, T.; Zhang, J.; Zhang, D.; and Gao, J. 2025 a . GUI-Actor : Coordinate-Free Visual Grounding for GUI Agents. In Advances in Neural In...

  47. [60]

    P.; and Qiao, Y

    Wu, Z.; Wu, Z.; Xu, F.; Wang, Y.; Sun, Q.; Jia, C.; Cheng, K.; Ding, Z.; Chen, L.; Liang, P. P.; and Qiao, Y. 2025 b . OS-ATLAS : Foundation Action Model for Generalist GUI Agents. In The Thirteenth International Conference on Learning Representations

  48. [61]

    Xie, R.; Friedman, D.; Yu, D.; Pan, B.; Fifty, C.; Kim, J.-H.; Du, X.; Gan, Z.; Rathod, V.; and Dhingra, B. 2026. LensVLM : Selective Context Expansion for Compressed Visual Representation of Text. arXiv preprint arXiv:2605.07019

  49. [62]

    Xie, T.; Deng, J.; Li, X.; Yang, J.; Wu, H.; Chen, J.; Hu, W.; Wang, X.; Xu, Y.; Wang, Z.; Xu, Y.; Wang, J.; Sahoo, D.; Yu, T.; and Xiong, C. 2025 a . Scaling Computer-Use Grounding via User Interface Decomposition and Synthesis. In Advances in Neural Information Processing Sy...

  50. [63]

    H.; Cheng, Z.; Shin, D.; Lei, F.; Liu, Y.; Xu, Y.; Zhou, S.; Savarese, S.; Xiong, C.; Zhong, V.; and Yu, T

    Xie, T.; Zhang, D.; Chen, J.; Li, X.; Zhao, S.; Cao, R.; Toh, J. H.; Cheng, Z.; Shin, D.; Lei, F.; Liu, Y.; Xu, Y.; Zhou, S.; Savarese, S.; Xiong, C.; Zhong, V.; and Yu, T. 2024. OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments. In Adv...

  51. [64]

    Xie, Y.; Li, Z.; Shao, R.; Chen, G.; Zhou, K.; Li, Y.; Jiang, D.; and Nie, L. 2025 b . Mirage-1: Augmenting and Updating GUI Agent with Hierarchical Multimodal Skills. arXiv preprint arXiv:2506.10387

  52. [65]

    Yuan, Q.; Lou, J.; Yu, X.; Lin, H.; Sun, L.; Han, X.; and Lu, Y. 2026. Vision-OPD: Learning to See Fine Details for Multimodal LLM s via On-Policy Self-Distillation. arXiv preprint arXiv:2605.18740

  53. [66]

    Zhang, K.; Shao, S.; Li, Q.; Lin, J.; Fu, L.; Wang, S.; Jiao, W.; Lu, Y.; Liu, W.; Zhang, W.; and Yu, Y. 2026 a . MMSkills: Towards Multimodal Skills for General Visual Agents. arXiv preprint arXiv:2605.13527

  54. [67]

    Zhang, Y.; Wu, D.; Shen, H.; Ma, C.; and Zhou, Y. 2026 b . Learn where to Click from Yourself: On-Policy Self-Distillation for GUI Grounding. arXiv:2605.00642

  55. [68]

    Zheng, B.; Gou, B.; Kil, J.; Sun, H.; and Su, Y. 2024. GPT -4 V (ision) is a Generalist Web Agent, if Grounded. In Proceedings of the 41st International Conference on Machine Learning, volume 235 of Proceedings of Machine Learning Research, 61349--61385. PMLR

Pith tools

Reviewed August 12, 2026 · model on record in the stance chip above.