REVIEW 3 major objections 5 minor 63 references
SkillLens: Visual Skill Cards for Retrieval-Augmented GUI Action Prediction and On-Policy Distillation
T0 review · 3 major / 5 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read Visual Skill Cards give frozen GUI agents reusable visual procedural memory and lift step accuracy by up to 11.6 points on Multimodal-Mind2Web.
desk verdict A well-ablated, plausible memory interface for GUI agents; the headline numbers need a clearer statement that evaluation episodes are excluded from the card library. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the Visual Skill Card, written $s_i = (p_i, z_i, v_i, \kappa_i)$: a procedure $p_i$, state cues $z_i$ (applicability and verification), visual evidence views $v_i$, and optional auxiliary fields $\kappa_i$. Trace-to-VSC converts heterogeneous interaction traces into this common schema, segmenting them into reusable units and auditing that held-out evaluations do not expose answer coordinates as templates. At inference, SkillLens uses a lightweight context-aware selector that scores cards by token overlap with the task query, then reranks and expands only the top candidates' evidence, so retrieval decides where to look while expansion controls how much visual evidence the frozen executor sees. For distillation, CardDistill gives the teacher the card bundle and the student only the benchmark-native context, optimizing a teacher-confidence-weighted reverse KL over student-generated action prefixes.
What would settle it
Build the card library from episodes on one set of sites and evaluate on a disjoint set of sites with the same task types, or strip every after-state and verification crop from the cards: if the improvement mostly collapses, the cards are leaking benchmark-specific outcome templates rather than supplying transferable visual procedures.
Extended reading notes
Core claim
The central discovery is that a VSC — a card pairing a procedure with when-it-applies cues, before/after visual evidence, and verification signals — works both as external runtime memory and as training-time privilege. On Multimodal-Mind2Web and WebLINX-BrowserGym, adding retrieved VSCs to a frozen GPT-5.4-mini executor raises Step SR by +11.6 points and Overall by +2.9 points, and the same evidence used in CardDistill raises the student-only Qwen3-VL-2B metrics by +12.0 and +3.2 points. The paper further shows that the gains require relevance: random or irrelevant cards do not reproduce them and can actively hurt grounding. The design separates a lightweight retrieval step from selective high-resolution evidence expansion, keeping the live screen as the final grounding source and bounding runtime evidence.
Load-bearing premise
The gains depend on the before/after visual evidence in a card being reusable procedural memory rather than a stored picture of the correct answer; the paper does not say how after-state screenshots and verification cues are kept from encoding the target outcome for the episodes that were used to build the library.
Editorial extensions
If this is right
- Frozen GUI executors can be upgraded without parameter updates: the paper shows positive Step SR / Overall gains across Qwen, Gemini, and GPT models when relevant VSCs are retrieved and expanded.
- CardDistill shows the retrieved behavior can be internalized: a Qwen3-VL-2B student improves by +12.0 Step SR on Mind2Web and +3.2 Overall on WebLINX-BG while running without any runtime card retrieval.
- Relevance is necessary: negative controls with random or irrelevant cards fail to reproduce the gains and can reduce grounding accuracy, so the effect is not simply extra images or longer prompts.
- The retrieve-then-expand cost split bounds runtime evidence while preserving high-resolution visual detail, making the approach tractable for step-by-step interactive agents.
- Verification cues give a lightweight contract for checking progress after an action, which could be used for self-monitoring as well as action prediction.
Reading between the lines
- A direct cross-site transfer test — build cards from site A, evaluate on site B — would sharpen the claim that the cards store reusable procedures rather than site-specific layouts; the paper does not report such a test.
- If the card library is the only source of task knowledge, the same framework could in principle be pointed at new domains (mobile UI, design tools, enterprise software) by swapping in new traces, without retraining the executor.
- The teacher-confidence-weighted reverse KL objective could be applied to other privileged signals beyond VSCs, such as ground-truth element boxes or future-state crops, to test how general the distillation recipe is.
- The board-based selector diagnostics suggest an alternative route: a VLM reading a rendered candidate board can improve retrieval-only Hit@1 even though it underperforms in end-to-end execution, so a better board layout or hybrid scoring might convert that retrieval gain into downstream accuracy.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces Visual Skill Cards (VSCs), a memory representation that binds a reusable procedure with applicability cues, visual before/target/after evidence, and verification signals. Trace-to-VSC converts heterogeneous interaction traces into this schema; SkillLens retrieves and selectively expands VSC evidence for a frozen VLM executor; CardDistill uses the same evidence as privileged teacher context for on-policy distillation into a student that runs without retrieval. Experiments on Multimodal-Mind2Web, WebLINX-BG, and OSWorld-G report consistent improvements over no-skill baselines across several frozen executors, plus student-only gains for Qwen3-VL-2B, with negative controls and modality ablations.
Significance. If the results hold, this is a useful contribution: VSCs offer a common, auditable unit for heterogeneous GUI experience; the separation of low-cost retrieval from high-resolution evidence expansion is pragmatic; and CardDistill comes with confidence intervals and a shuffled-VSC control. The negative-control experiments in Table 5 and the retrieval-coverage analysis in Table 3 are genuine strengths. However, because the main quantitative claims hinge on a single table without error bars and on the assumption that the VSC library does not contain evaluation-episode answers, the current evidence does not yet establish that the gains come from reusable visual procedural memory rather than from episode-specific answer lookup or from text-only procedures.
major comments (3)
- [Trace-to-VSC Construction / Eq. (7)] The VSC schema stores after-state screenshots and verification cues, and the executor is conditioned on this evidence before action selection (Eq. (14)). The paper evaluates on the same benchmarks used to build the library and only asserts that 'held-out evaluations do not expose answer coordinates as templates.' This does not rule out the after-state screenshot or verification cue acting as an episode-specific answer lookup: for a card built from a Mind2Web or WebLINX trace, the after-state can be the ground-truth next observation of that very episode. Please specify exactly how the construction/evaluation split is managed, and add a control in which the library is built from sources provably disjoint from the evaluation episodes (or from the training split with a statement that no evaluation episode contributed). Without such a control, the +11.6 and +2.9 headline gains could be explained by answer lookup rather than reusable procedural memory.
- [Table 4] Table 4's modality ablation directly undercuts the 'visual' half of the central claim. On Mind2Web, removing visual evidence improves over Full SkillLens for both Qwen3-VL-2B (Step SR 14.0 vs 10.3; Elem. Acc. 67.5 vs 62.2) and Gemini 2.5 Flash (Step SR 75.0 vs 66.2; Elem. Acc. 86.0 vs 77.7); for Qwen3-VL-2B, image-only cards (w/o procedure) give Step SR 4.5, essentially identical to the 4.6 no-skill baseline. The Step SR gain on Mind2Web is therefore carried by procedure text, not visual evidence. Please either explain this pattern or reframe the contribution as mixed procedural memory, and report the headline result with the text-only variant disambiguated.
- [Table 1 / Experimental Setup] The main results in Table 1 are single-run percentages with no repeated-seed or error-bar information, while the CardDistill results are reported with 95% CIs and p-values. Since the abstract's headline numbers come from Table 1, please report variance estimates (e.g., multiple evaluation subsets or seeds) or state why the evaluation is deterministic and why variance is negligible. This is needed to assess whether the +11.6, +2.9, +12.0, and +3.2 deltas are statistically meaningful.
minor comments (5)
- [Eq. (15)] The objective is written as KL(P_theta || P_phi) and called 'reverse KL'; the usual reverse-KL direction is KL(teacher || student). Please clarify the intended direction or adjust the terminology.
- [Table 3] All selectors report Oracle Cov. 1.00 on every benchmark. Please discuss whether this means retrieval recall is saturated in these settings and whether the observed gains should be attributed primarily to evidence expansion rather than to retrieval.
- [Trace-to-VSC Construction] The construction pipeline (trace normalization, segmentation, summarization, evidence selection, audit) is described at a high level. Please provide concrete adapter details for Mind2Web and WebLINX, such as which LLM or VLM performs summarization and what segmentation rules are used, so the method is reproducible.
- [Table 2] The latency numbers (1.5s vs 4.0s) are reported without hardware details. Please specify the inference stack and hardware used for the timing comparison.
- [Figure 3] The caption mentions 'training dynamics' but the text does not describe the plotted curves. Either describe the figure in the body or point to the supplementary material for a full explanation.
Circularity Check
VSC 'After' evidence can be the target observation from the same benchmark trace, and the audit only excludes coordinate templates, so the central gains may partly be answer lookup rather than reusable procedure.
-
fitted input called prediction
[Method, Trace-to-VSC Construction; VSC Representation (Eq. 7); Grounded Action Prediction (Eq. 14); Figure 1 Before-Target-After evidence]
"Trace-to-VSC converts prior interaction records into reusable visual procedures... a structured benchmark can bind targets deterministically... This audit checks that the card is self-contained and that held-out evaluations do not expose answer coordinates as templates. Verification cues define observable postconditions for a successful step. The visual evidence may include a full interface view or a focused target crop... Expanded evidence is reference material, not a coordinate template."
The VSC schema stores vi as visual evidence, and the Before-Target-After construction (Figure 1) places the postcondition after-state ('Added to cart!') in the card. When a structured benchmark like Mind2Web 'binds targets deterministically,' that after-state can be the ground-truth next observation of the very episode being predicted. Eq. (14) then conditions the frozen executor on Et expanded from the retrieved card, so the 'prediction' at step t is made with the correct outcome already in context. The stated audit only excludes 'answer coordinates as templates'; it does not exclude after-state images or verification cues, and no disjoint-source control is reported.
full rationale
The paper's retrieval and expansion machinery is explicit and mostly self-contained, and its negative controls (random/irrelevant VSCs) plus the shuffled-VSC CardDistill control show that card relevance and alignment, not just prompt length, drive the effect. There is no load-bearing self-citation or imported uniqueness theorem. The central risk is a provenance gap: the VSC library is built from Mind2Web and WebLINX traces, the same benchmarks used for evaluation, and VSC evidence includes postcondition after-state screenshots and verification cues. If any evaluation episode contributed a card, or if cards from the same episode distribution encode the target next observation, then Eq. (14)'s prediction is conditioned on answer-bearing content and the +11.6/+2.9 gains are partly fitted, not predicted. The paper's audit sentence is an assertion about 'answer coordinates as templates' and does not describe a mechanism or measurement covering after-state images; Table 4 even shows text-only cards sometimes beat full visual cards, which leaves the provenance question unresolved rather than resolved. This is not a formal circularity of equations, but it is a partial reduction of the central claim to its own benchmark-derived inputs, so the score is 4.
Assumptions & free parameters
free parameters (3)
- Rerank blend weight =
0.1
- Retrieval and expansion budgets Kc, Ks, Ke, Kv =
not specified in text
- Distillation temperature T =
1.0
assumptions (4)
- domain assumption The VSC library constructed from public traces does not leak held-out answer coordinates; the paper states an audit checks this.
- domain assumption Feeding postcondition ('After') screenshots as reference evidence is legitimate and does not reveal the target action.
- domain assumption Frozen VLM executors can usefully consume retrieved VSC evidence and still ground actions on the live screen.
- domain assumption Reverse KL with teacher-confidence weighting transfers card-conditioned behavior to the student.
invented entities (1)
-
Visual Skill Card (VSC)
independent evidence
Cite this review
Pith. "Pith review of SkillLens: Visual Skill Cards for Retrieval-Augmented GUI Action Prediction and On-Policy Distillation." pith.science (2026). https://pith.science/paper/QAWZ7HAW
@misc{pith2026260810775,
author = {Pith},
title = {Pith review of: SkillLens: Visual Skill Cards for Retrieval-Augmented GUI Action Prediction and On-Policy Distillation},
year = {2026},
howpublished = {\url{https://pith.science/paper/QAWZ7HAW}},
note = {Machine review of arXiv:2608.10775}
}
read the original abstract
Computer-using agents can perceive rich software interfaces, yet their decisions often lack visual procedural memory: they may recognize individual controls without identifying which familiar workflow is active, which control matters next, or what evidence would confirm progress. Raw interaction traces preserve such information but are long and noisy to condition on, whereas text-only skills often omit the visual state that makes a procedure applicable. We introduce Visual Skill Cards (VSCs), a state-conditioned memory representation that binds reusable procedures with applicability cues, visual evidence, and verification signals. SkillLens constructs VSCs from heterogeneous interaction experience through Trace-to-Visual-Skill-Card and, at inference time, retrieves relevant cards and selectively expands only the evidence needed by a fixed visual-language model executor for grounded GUI action prediction. The same representation also supports CardDistill, which uses VSC evidence as privileged teacher context to train a student that acts without runtime card retrieval. Across Multimodal-Mind2Web and WebLINX-BrowserGym, SkillLens improves the frozen GPT-5.4-mini executor by +11.6 points in Step SR and +2.9 points in Overall, respectively; CardDistill further improves the corresponding student-only Qwen3-VL-2B metrics by +12.0 and +3.2 points.
Figures
Reference graph
Works this paper leans on
-
[2]
doi:10.48550/arXiv.2601.21123 , url=
Chen, Tianyi and Li, Yinheng and Solodko, Michael and Wang, Sen and Jiang, Nan and Cui, Tingyuan and Hao, Junheng and Ko, Jongwoo and Abdali, Sara and Xu, Leon and Zheng, Suzhen and Fan, Hao and Cameron, Pashmina and Wagle, Justin and Koishida, Kazuhito , year=. doi:10.48550/arXiv.2601.21123 , url=. 2601.21123 , archivePrefix=
-
[4]
Learn where to Click from Yourself: On-Policy Self-Distillation for
Zhang, Yan and Wu, Daiqing and Shen, Huawen and Ma, Can and Zhou, Yu , year=. Learn where to Click from Yourself: On-Policy Self-Distillation for. doi:10.48550/arXiv.2605.00642 , url=. 2605.00642 , archivePrefix=
-
[5]
Vision-OPD: Learning to See Fine Details for Multimodal
Yuan, Qianhao and Lou, Jie and Yu, Xing and Lin, Hongyu and Sun, Le and Han, Xianpei and Lu, Yaojie , journal=. Vision-OPD: Learning to See Fine Details for Multimodal
-
[7]
OCR-Memory: Optical Context Retrieval for Long-Horizon Agent Memory , author=. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=. 2026 , doi=
work page 2026
-
[8]
LensVLM: Selective Context Expansion for Compressed Visual Representation of Text
Xie, Roy and Friedman, Dan and Yu, Donghan and Pan, Bowen and Fifty, Christopher and Kim, Jang-Hyun and Du, Xianzhi and Gan, Zhe and Rathod, Vivek and Dhingra, Bhuwan , journal=. 2026 , eprint=. doi:10.48550/arXiv.2605.07019 , url=
work page Pith review arXiv doi:10.48550/arxiv.2605.07019 2026
-
[9]
AgentOCR: Reimagining Agent History via Optical Self-Compression , author=. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=. 2026 , doi=
work page 2026
-
[10]
Proceedings of the 43rd International Conference on Machine Learning , year=
Shi, Yaorui and Liu, Shugui and Yang, Yu and Mao, Wenyu and Chen, Yuxin and Qi. Proceedings of the 43rd International Conference on Machine Learning , year=. doi:10.48550/arXiv.2601.21468 , url=. 2601.21468 , archivePrefix=
-
[11]
Advances in Neural Information Processing Systems , volume=
VideoGUI: A Benchmark for GUI Automation from Instructional Videos , author=. Advances in Neural Information Processing Systems , volume=. 2024 , doi=
work page 2024
Show all 63 references
-
[12]
Advances in Neural Information Processing Systems , volume=
OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments , author=. Advances in Neural Information Processing Systems , volume=. 2024 , doi=
2024
-
[13]
Proceedings of the 42nd International Conference on Machine Learning , series=
Windows Agent Arena: Evaluating Multi-Modal OS Agents at Scale , author=. Proceedings of the 42nd International Conference on Machine Learning , series=. 2025 , url=
2025
-
[14]
2025 , eprint=
Wang, Haoming and Zou, Haoyang and Song, Huatong and Feng, Jiazhan and Fang, Junjie and Lu, Junting and Liu, Longxiang and Luo, Qinyu and Liang, Shihao and Huang, Shijue and Zhong, Wanjun and Ye, Yining and Qin, Yujia and Xiong, Yuwen and Song, Yuxin and Wu, Zhiyong and Li, Ao...
-
[15]
2026 , address=
Liang, Weijie and Song, Yuanfeng and Chen, Xing and Cao, Caleb Chen and Han, Sirui and Guo, Yike , booktitle=. 2026 , address=
2026
-
[16]
Advances in Neural Information Processing Systems , volume=
Mind2Web: Towards a Generalist Agent for the Web , author=. Advances in Neural Information Processing Systems , volume=. 2023 , url=
2023
-
[17]
2024 , url=
Zheng, Boyuan and Gou, Boyu and Kil, Jihyung and Sun, Huan and Su, Yu , booktitle=. 2024 , url=. 2401.01614 , archivePrefix=
2024 arXiv
-
[18]
Proceedings of the 41st International Conference on Machine Learning , series=
L. Proceedings of the 41st International Conference on Machine Learning , series=. 2024 , url=. 2402.05930 , archivePrefix=
2024
-
[19]
Xu and Siva Reddy and Graham Neubig and Quentin Cappart and Russ Salakhutdinov and Nicolas Chapados , journal=
Thibault Le Sellier de Chezelles and Maxime Gasse and Alexandre Lacoste and Massimo Caccia and Alexandre Drouin and Léo Boisvert and Megh Thakkar and Tom Marty and Rim Assouel and Sahar Omidi Shayegan and Lawrence Keunho Jang and Xing Han Lù and Ori Yoran and Dehan Kong and Fr...
2025
-
[20]
Advances in Neural Information Processing Systems , volume=
Scaling Computer-Use Grounding via User Interface Decomposition and Synthesis , author=. Advances in Neural Information Processing Systems , volume=. 2025 , url=
2025
-
[21]
2024 , address=
Cheng, Kanzhi and Sun, Qiushi and Chu, Yougang and Xu, Fangzhi and Li, Yantao and Zhang, Jianbing and Wu, Zhiyong , booktitle=. 2024 , address=. doi:10.18653/v1/2024.acl-long.505 , url=. 2401.10935 , archivePrefix=
2024 arXiv
-
[22]
ScreenSpot-Pro:
Li, Kaixin and Meng, Ziyang and Lin, Hongzhan and Luo, Ziyang and Tian, Yuchen and Ma, Jing and Huang, Zhiyong and Chua, Tat-Seng , booktitle=. ScreenSpot-Pro:. 2025 , doi=
2025
-
[23]
2025 , eprint=
Qin, Yujia and Ye, Yining and Fang, Junjie and Wang, Haoming and Liang, Shihao and Tian, Shizuo and Zhang, Junda and Li, Jiahao and Li, Yunxin and Huang, Shijue and Zhong, Wanjun and Li, Kuanye and Yang, Jiale and Miao, Yu and Lin, Woyu and Liu, Longxiang and Jiang, Xu and Ma,...
2025
-
[24]
2025 , url=
Wu, Zhiyong and Wu, Zhenyu and Xu, Fangzhi and Wang, Yian and Sun, Qiushi and Jia, Chengyou and Cheng, Kanzhi and Ding, Zichen and Chen, Liheng and Liang, Paul Pu and Qiao, Yu , booktitle=. 2025 , url=
2025
-
[25]
Navigating the Digital World as Humans Do: Universal Visual Grounding for
Gou, Boyu and Wang, Ruohan and Zheng, Boyuan and Xie, Yanan and Chang, Cheng and Shu, Yiheng and Sun, Huan and Su, Yu , booktitle=. Navigating the Digital World as Humans Do: Universal Visual Grounding for. 2025 , url=
2025
-
[26]
2025 , url=
Wu, Qianhui and Cheng, Kanzhi and Yang, Rui and Zhang, Chaoyun and Yang, Jianwei and Jiang, Huiqiang and Mu, Jian and Peng, Baolin and Qiao, Bo and Tan, Reuben and Qin, Si and Liden, Lars and Lin, Qingwei and Zhang, Huan and Zhang, Tong and Zhang, Jianbing and Zhang, Dongmei a...
2025
-
[27]
2024 , eprint=
Wang, Peng and Bai, Shuai and Tan, Sinan and Wang, Shijie and Fan, Zhihao and Bai, Jinze and Chen, Keqin and Liu, Xuejing and Wang, Jialin and Ge, Wenbin and Fan, Yang and Dang, Kai and Du, Mengfei and Ren, Xuancheng and Men, Rui and Liu, Dayiheng and Zhou, Chang and Zhou, Jin...
2024
-
[28]
2025 , eprint=
Bai, Shuai and Chen, Keqin and Liu, Xuejing and Wang, Jialin and Ge, Wenbin and Song, Sibo and Dang, Kai and Wang, Peng and Wang, Shijie and Tang, Jun and Zhong, Humen and Zhu, Yuanzhi and Yang, Mingkun and Li, Zhaohai and Wan, Jianqiang and Wang, Pengfei and Ding, Wei and Fu,...
2025
-
[29]
2025 , eprint=
Bai, Shuai and Cai, Yuxuan and Chen, Ruizhe and Chen, Keqin and Chen, Xionghui and Cheng, Zesen and Deng, Lianghao and Ding, Wei and Gao, Chang and Ge, Chunjiang and Ge, Wenbin and Guo, Zhifang and Huang, Qidong and Huang, Jie and Huang, Fei and Hui, Binyuan and Jiang, Shutong...
2025
-
[30]
2025 , type=
Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities , author=. 2025 , type=
2025
-
[31]
and Yang, Zhilin and Yu, Tao , booktitle=
Wang, Xinyuan and Wang, Bowen and Lu, Dunjie and Yang, Junlin and Xie, Tianbao and Wang, Junli and Deng, Jiaqi and Guo, Xiaole and Xu, Yiheng and Wu, Chen and Shen, Zhennan and Li, Zhuokai and Li, Ryan and Li, Xiaochuan and Chen, Junda and Zheng, Boyuan and Li, Peihang and Lei...
2025
-
[33]
ToolCUA: Towards Optimal
Hu, Xuhao and Zhang, Xi and Xu, Haiyang and Qiao, Kyle and Yang, Jingyi and Huang, Xuanjing and Shao, Jing and Yan, Ming and Ye, Jieping , journal=. ToolCUA: Towards Optimal
-
[35]
Bai, S.; Cai, Y.; Chen, R.; Chen, K.; Chen, X.; Cheng, Z.; Deng, L.; Ding, W.; Gao, C.; Ge, C.; Ge, W.; Guo, Z.; Huang, Q.; Huang, J.; Huang, F.; Hui, B.; Jiang, S.; Li, Z.; Li, M.; Li, M.; Li, K.; Lin, Z.; Lin, J.; Liu, X.; Liu, J.; Liu, C.; Liu, Y.; Liu, D.; Liu, S.; Lu, D.;...
2025 arXiv
-
[36]
Bai, S.; Chen, K.; Liu, X.; Wang, J.; Ge, W.; Song, S.; Dang, K.; Wang, P.; Wang, S.; Tang, J.; Zhong, H.; Zhu, Y.; Yang, M.; Li, Z.; Wan, J.; Wang, P.; Ding, W.; Fu, Z.; Xu, Y.; Ye, J.; Zhang, X.; Xie, T.; Cheng, Z.; Zhang, H.; Yang, Z.; Xu, H.; and Lin, J. 2025 b . Qwen2.5-V...
2025 arXiv
-
[37]
K.; and Hui, Z
Bonatti, R.; Zhao, D.; Bonacci, F.; Dupont, D.; Abdali, S.; Li, Y.; Lu, Y.; Wagle, J.; Koishida, K.; Bucker, A.; Jang, L. K.; and Hui, Z. 2025. Windows Agent Arena: Evaluating Multi-Modal OS Agents at Scale. In Proceedings of the 42nd International Conference on Machine Learni...
2025
-
[38]
Cai, Y.; Liu, J.; Liu, Y.; Deng, H.; Yao, L.; Zheng, Y.; Ouyang, K.; Li, Z.; Wang, Z.; Sun, X.; Bai, H.; and Li, X. 2026. Thinking Without Images: Internalizing Visual Manipulation with On-Policy Self-Distillation. arXiv preprint arXiv:2606.08719
2026 arXiv
-
[39]
Chen, T.; Li, Y.; Solodko, M.; Wang, S.; Jiang, N.; Cui, T.; Hao, J.; Ko, J.; Abdali, S.; Xu, L.; Zheng, S.; Fan, H.; Cameron, P.; Wagle, J.; and Koishida, K. 2026. CUA -Skill: Develop Skills for Computer Using Agent. arXiv:2601.21123
2026
-
[40]
Cheng, K.; Sun, Q.; Chu, Y.; Xu, F.; Li, Y.; Zhang, J.; and Wu, Z. 2024. SeeClick : Harnessing GUI Grounding for Advanced Visual GUI Agents. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 9313--9332. Bangkok,...
2024
-
[41]
de Chezelles, T. L. S.; Gasse, M.; Lacoste, A.; Caccia, M.; Drouin, A.; Boisvert, L.; Thakkar, M.; Marty, T.; Assouel, R.; Shayegan, S. O.; Jang, L. K.; Lù, X. H.; Yoran, O.; Kong, D.; Xu, F. F.; Reddy, S.; Neubig, G.; Cappart, Q.; Salakhutdinov, R.; and Chapados, N. 2025. The...
2025
-
[42]
Deng, X.; Gu, Y.; Zheng, B.; Chen, S.; Stevens, S.; Wang, B.; Sun, H.; and Su, Y. 2023. Mind2Web: Towards a Generalist Agent for the Web. In Advances in Neural Information Processing Systems, volume 36, 28091--28114
2023
-
[43]
Feng, L.; Yang, F.; Chen, F.; Cheng, X.; Xu, H.; Wan, Z.; Yan, M.; and An, B. 2026. AgentOCR: Reimagining Agent History via Optical Self-Compression. In Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 5067--5086....
2026
-
[44]
Gemini Team, Google . 2025. Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities. Technical report, Google
2025
-
[45]
Gou, B.; Wang, R.; Zheng, B.; Xie, Y.; Chang, C.; Shu, Y.; Sun, H.; and Su, Y. 2025. Navigating the Digital World as Humans Do: Universal Visual Grounding for GUI Agents. In The Thirteenth International Conference on Learning Representations
2025
-
[46]
Hu, X.; Zhang, X.; Xu, H.; Qiao, K.; Yang, J.; Huang, X.; Shao, J.; Yan, M.; and Ye, J. 2026. ToolCUA: Towards Optimal GUI -Tool Path Orchestration for Computer Use Agents. arXiv preprint arXiv:2605.12481
2026 arXiv
-
[47]
Jiang, Z.; An, L.; Liu, Y.; Ji, J.; Wu, Q.; Andreas, J.; Zhang, Y.; and Chang, S. 2026. VISUALSKILL: Multimodal Skills for Computer-Use Agents. arXiv preprint arXiv:2606.18448
2026 arXiv
-
[48]
Li, J.; Zhang, Y.; Yang, X.; QU, J.; Xu, J.; Yang, S.; Ding, J.; and Ngai, E. C.-H. 2026. OCR-Memory: Optical Context Retrieval for Long-Horizon Agent Memory. In Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 10...
2026
-
[49]
C.; Han, S.; and Guo, Y
Liang, W.; Song, Y.; Chen, X.; Cao, C. C.; Han, S.; and Guo, Y. 2026. V izo M em: A Visual-Textual Memory Framework for Efficient Long-Horizon Reasoning. In Findings of the Association for Computational Linguistics: ACL 2026, 7399--7422. San Diego, California, United States: A...
2026
-
[50]
Q.; Li, L.; Gao, D.; Wu, Q.; Yan, M.; Yang, Z.; Wang, L.; and Shou, M
Lin, K. Q.; Li, L.; Gao, D.; Wu, Q.; Yan, M.; Yang, Z.; Wang, L.; and Shou, M. Z. 2024. VideoGUI: A Benchmark for GUI Automation from Instructional Videos. In Advances in Neural Information Processing Systems, volume 37, 69329--69360
2024
-
[51]
H.; Kasner, Z.; and Reddy, S
L \`u , X. H.; Kasner, Z.; and Reddy, S. 2024. WebLINX : Real-World Website Navigation with Multi-Turn Dialogue. In Proceedings of the 41st International Conference on Machine Learning, volume 235 of Proceedings of Machine Learning Research, 33007--33056. PMLR
2024
-
[52]
OpenAI . 2024. GPT-4o System Card
2024
-
[53]
Qin, Y.; Ye, Y.; Fang, J.; Wang, H.; Liang, S.; Tian, S.; Zhang, J.; Li, J.; Li, Y.; Huang, S.; Zhong, W.; Li, K.; Yang, J.; Miao, Y.; Lin, W.; Liu, L.; Jiang, X.; Ma, Q.; Li, J.; Xiao, X.; Cai, K.; Li, C.; Zheng, Y.; Jin, C.; Li, C.; Zhou, X.; Wang, M.; Chen, H.; Li, Z.; Yang...
2025 arXiv
-
[54]
Shi, Y.; Liu, S.; Yang, Y.; Mao, W.; Chen, Y.; GU , Q.; Su, H.; Cai, X.; Wang, X.; and Zhang, A. 2026. MemOCR : Layout-Aware Visual Memory for Efficient Long-Horizon Reasoning. In Proceedings of the 43rd International Conference on Machine Learning
2026
-
[55]
Wang, H.; Zou, H.; Song, H.; Feng, J.; Fang, J.; Lu, J.; Liu, L.; Luo, Q.; Liang, S.; Huang, S.; Zhong, W.; Ye, Y.; Qin, Y.; Xiong, Y.; Song, Y.; Wu, Z.; Li, A.; Li, B.; Dun, C.; Liu, C.; Zan, D.; Leng, F.; Wang, H.; Yu, H.; Chen, H.; Guo, H.; Su, J.; Huang, J.; Shen, K.; Shi,...
2025 arXiv
-
[56]
Wang, P.; Bai, S.; Tan, S.; Wang, S.; Fan, Z.; Bai, J.; Chen, K.; Liu, X.; Wang, J.; Ge, W.; Fan, Y.; Dang, K.; Du, M.; Ren, X.; Men, R.; Liu, D.; Zhou, C.; Zhou, J.; and Lin, J. 2024. Qwen2-VL : Enhancing Vision-Language Model's Perception of the World at Any Resolution. arXi...
2024 arXiv
-
[57]
Wang, X.; Wang, B.; Lu, D.; Yang, J.; Xie, T.; Wang, J.; Deng, J.; Guo, X.; Xu, Y.; Wu, C.; Shen, Z.; Li, Z.; Li, R.; Li, X.; Chen, J.; Zheng, B.; Li, P.; Lei, F.; Cao, R.; Fu, Y.; Shin, D.; Shin, M.; Hu, J.; Wang, Y.; Chen, J.; Ye, Y.; Zhang, D.; Wang, Y.; Wang, H.; Yang, D.;...
2025
-
[58]
Wei, H.; Sun, Y.; and Li, Y. 2025. DeepSeek-OCR: Contexts Optical Compression. arXiv preprint arXiv:2510.18234
2025 arXiv
-
[59]
Wu, Q.; Cheng, K.; Yang, R.; Zhang, C.; Yang, J.; Jiang, H.; Mu, J.; Peng, B.; Qiao, B.; Tan, R.; Qin, S.; Liden, L.; Lin, Q.; Zhang, H.; Zhang, T.; Zhang, J.; Zhang, D.; and Gao, J. 2025 a . GUI-Actor : Coordinate-Free Visual Grounding for GUI Agents. In Advances in Neural In...
2025
-
[60]
P.; and Qiao, Y
Wu, Z.; Wu, Z.; Xu, F.; Wang, Y.; Sun, Q.; Jia, C.; Cheng, K.; Ding, Z.; Chen, L.; Liang, P. P.; and Qiao, Y. 2025 b . OS-ATLAS : Foundation Action Model for Generalist GUI Agents. In The Thirteenth International Conference on Learning Representations
2025
-
[61]
Xie, R.; Friedman, D.; Yu, D.; Pan, B.; Fifty, C.; Kim, J.-H.; Du, X.; Gan, Z.; Rathod, V.; and Dhingra, B. 2026. LensVLM : Selective Context Expansion for Compressed Visual Representation of Text. arXiv preprint arXiv:2605.07019
2026 arXiv
-
[62]
Xie, T.; Deng, J.; Li, X.; Yang, J.; Wu, H.; Chen, J.; Hu, W.; Wang, X.; Xu, Y.; Wang, Z.; Xu, Y.; Wang, J.; Sahoo, D.; Yu, T.; and Xiong, C. 2025 a . Scaling Computer-Use Grounding via User Interface Decomposition and Synthesis. In Advances in Neural Information Processing Sy...
2025
-
[63]
H.; Cheng, Z.; Shin, D.; Lei, F.; Liu, Y.; Xu, Y.; Zhou, S.; Savarese, S.; Xiong, C.; Zhong, V.; and Yu, T
Xie, T.; Zhang, D.; Chen, J.; Li, X.; Zhao, S.; Cao, R.; Toh, J. H.; Cheng, Z.; Shin, D.; Lei, F.; Liu, Y.; Xu, Y.; Zhou, S.; Savarese, S.; Xiong, C.; Zhong, V.; and Yu, T. 2024. OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments. In Adv...
2024
-
[64]
Xie, Y.; Li, Z.; Shao, R.; Chen, G.; Zhou, K.; Li, Y.; Jiang, D.; and Nie, L. 2025 b . Mirage-1: Augmenting and Updating GUI Agent with Hierarchical Multimodal Skills. arXiv preprint arXiv:2506.10387
2025 arXiv
-
[65]
Yuan, Q.; Lou, J.; Yu, X.; Lin, H.; Sun, L.; Han, X.; and Lu, Y. 2026. Vision-OPD: Learning to See Fine Details for Multimodal LLM s via On-Policy Self-Distillation. arXiv preprint arXiv:2605.18740
2026 arXiv
-
[66]
Zhang, K.; Shao, S.; Li, Q.; Lin, J.; Fu, L.; Wang, S.; Jiao, W.; Lu, Y.; Liu, W.; Zhang, W.; and Yu, Y. 2026 a . MMSkills: Towards Multimodal Skills for General Visual Agents. arXiv preprint arXiv:2605.13527
2026 arXiv
-
[67]
Zhang, Y.; Wu, D.; Shen, H.; Ma, C.; and Zhou, Y. 2026 b . Learn where to Click from Yourself: On-Policy Self-Distillation for GUI Grounding. arXiv:2605.00642
2026 arXiv
-
[68]
Zheng, B.; Gou, B.; Kil, J.; Sun, H.; and Su, Y. 2024. GPT -4 V (ision) is a Generalist Web Agent, if Grounded. In Proceedings of the 41st International Conference on Machine Learning, volume 235 of Proceedings of Machine Learning Research, 61349--61385. PMLR
2024
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.