{"total":31,"items":[{"citing_arxiv_id":"2606.28715","ref_index":18,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"SEATauBench: Adapting Tool-Agent-User Evaluation Into Low-Resource Southeast Asian Languages","primary_cat":"cs.CL","submitted_at":"2026-06-27T03:44:00+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":7.0,"formal_verification":"none","one_line_summary":"SEATauBench is the first agent benchmark for SEA languages, finding that performance holds for language-only changes but degrades sharply with full domain localization.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2606.23049","ref_index":22,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"PhoneBuddy: Training Open Models for Agentic Phone Use","primary_cat":"cs.CL","submitted_at":"2026-06-22T08:57:54+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":6.0,"formal_verification":"none","one_line_summary":"PhoneBuddy combines real-app and mock-app RL after shared SFT, raising real-phone task success from 36.67% to 45.33% and AndroidWorld from 60.3% to 83.2%.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2606.17929","ref_index":3,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"PreAct: Computer-Using Agents that Get Faster on Repeated Tasks","primary_cat":"cs.AI","submitted_at":"2026-06-16T13:40:21+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":7.0,"formal_verification":"none","one_line_summary":"PreAct compiles successful agent executions into verifiable state-machine programs for 8.5-13x faster replay on repeated tasks, with an independent evaluator check before storing each program.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2606.17321","ref_index":1,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"ProCUA-SFT Technical Report","primary_cat":"cs.LG","submitted_at":"2026-06-15T22:04:11+00:00","verdict":"CONDITIONAL","verdict_confidence":"MODERATE","novelty_score":7.0,"formal_verification":"none","one_line_summary":"ProCUA-SFT is a 3.1M-sample SFT dataset from 93K verified synthetic trajectories that lifts UI-TARS 7B OSWorld score from 26.3% to 45%.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2606.12195","ref_index":96,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning","primary_cat":"cs.CV","submitted_at":"2026-06-10T15:17:08+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":4.0,"formal_verification":"none","one_line_summary":"InternVideo3 introduces Multimodal Contextual Reasoning and M^2LA attention to enable closed-loop evidence accumulation in long-video understanding and agentic tool use, reporting strong benchmark results.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2606.12191","ref_index":55,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"Agentic Environment Engineering for Large Language Models: A Survey of Environment Modeling, Synthesis, Evaluation, and Application","primary_cat":"cs.CL","submitted_at":"2026-06-10T15:15:01+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":5.0,"formal_verification":"none","one_line_summary":"This survey categorizes agentic environments for LLMs by eight attributes and domains, introduces symbolic and neural synthesis paradigms with evaluation, and outlines four agent evolution pathways plus three environment evolution paradigms.","context_count":1,"top_context_role":"dataset","top_context_polarity":"background","context_text":"ment learning simulators which operate over predefined state-action spaces and fixed transition dynamics, agentic JOURNAL OF LATEX CLASS FILES, JANUARY 2025 4 Environment EnvironmentDomain (§4) GUI (§4.1) Desktop GUI (§4.1.1)e.g.,WorkArena [27], OSWorld [28], WindowsAgentArena [52], OSWorld-MCP [53],etc. Mobile GUI (§4.1.2)e.g.,Mobile-Env [54], AitW [55], AndroidWorld [56], MobileWorld [57], Mobile-Bench [58],etc. Web GUI (§4.1.3) e.g.,WebShop [15], Mind2Web [59], WebArena [60], VisualWebArena [61],etc. Deep Research (§4.2) Information Search (§4.2.1)e.g.,SimpleQA [29], WideSearch [62], InfoDeepSeek [63], InfoSeek [64],etc. Multi-Source Reasoning(§4.2.2) e.g.,MMDR-Bench [65], WebWalker [66], BrowseComp [31], Conflicts [67], BrowseComp-ZH [68],LiveDRBench [69], OmniGAIA [70], GAIA [30],etc."},{"citing_arxiv_id":"2606.01936","ref_index":31,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"What to Format and How: A Benchmark and Workflow Approach for Document Formatting","primary_cat":"cs.CL","submitted_at":"2026-06-01T09:02:33+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":6.0,"formal_verification":"none","one_line_summary":"Presents DocFormBench benchmark and DocFormFlow workflow for content-aware LLM document formatting, claiming higher accuracy and lower token use via decoupled localization and modification.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2605.29115","ref_index":1,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"unix-ctf: Procedural Environments for Unix-Competence Reinforcement Learning","primary_cat":"cs.CR","submitted_at":"2026-05-27T21:23:00+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":6.0,"formal_verification":"none","one_line_summary":"unix-ctf procedurally generates 656 Unix CTF tasks across 155 techniques; fine-tuning Qwen3-8B on them raises solve rate from 11.6% to 43.6% on a 15-skill holdout and yields +33 pp in Forensics on InterCode-CTF.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2605.28775","ref_index":4,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"Learn from Weaknesses: Automated Domain Specialization for Small Computer-Use Agents","primary_cat":"cs.LG","submitted_at":"2026-05-27T17:37:00+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":6.0,"formal_verification":"none","one_line_summary":"LearnWeak specializes small CUAs via weakness detection by a reference agent, targeted task synthesis, and error-aware training, delivering 11+ point gains on OSWorld.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2605.27761","ref_index":7,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"AndroidDaily: A Verifiable Benchmark for Mobile GUI Agents on Real-World Closed-Source Applications","primary_cat":"cs.CV","submitted_at":"2026-05-26T23:19:42+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":7.0,"formal_verification":"none","one_line_summary":"AndroidDaily supplies 350 verifiable tasks on 94 closed-source Android apps evaluated by GRADE (87.37% human agreement), with the strongest model achieving 62% success.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2605.25343","ref_index":134,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"Toward Native Multimodal Modeling: A Roadmap","primary_cat":"cs.CV","submitted_at":"2026-05-25T01:57:43+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":3.0,"formal_verification":"none","one_line_summary":"A roadmap that defines architectural nativity for multimodal models and categorizes them into Multi-to-Text, Multi-to-Target, and Multi-to-Multi types while outlining an industrial pipeline toward unified transformer-based native multimodal modeling.","context_count":1,"top_context_role":"background","top_context_polarity":"background","context_text":"environmental sound generation. Interact Web Interaction WebShop [123], Mind2Web [124], We- bArena [125], VisualWebArena [126], WebLINX [127], WebV oyager [128] T, I (GUI) Goal-driven web navigation: searching, clicking, form filling on real/simulated websites. Mobile & Desktop GUI AITW [ 129], RICO [ 130], ScreenAI [ 131], SeeClick [ 132], OSWorld [ 133], Windows Agent Arena [134] T, I (GUI) Screenshot/UI-tree to action (tap, type, drag); covers mobile and OS environments. Embodied Interaction ALFWorld [ 135], BridgeData V2 [136], Open X-Embodiment [137], Magma [138] T, I, V (robot) Language-conditioned manipulation from visual observations and robot states. Align Hallucination & Faithfulness LLaV A-RLHF [139], RLHF-V [ 140],"},{"citing_arxiv_id":"2605.25160","ref_index":20,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"ScaleWoB: Guiding GUI Agents with Coding Agents via Large-Scale Environmental Synthesis","primary_cat":"cs.AI","submitted_at":"2026-05-24T16:33:14+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":7.0,"formal_verification":"none","one_line_summary":"ScaleWoB generates 100+ synthetic interactive GUI environments and 1000+ verifiable tasks as web pages, releasing a 120-task mobile benchmark where state-of-the-art agents achieve 27.92% success (17.82% on long-horizon tasks) versus 92.08% for humans, with synthetic results generalizing to real apps","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2605.24785","ref_index":2,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"PANDO: Efficient Multimodal AI Agents via Online Skill Distillation","primary_cat":"cs.AI","submitted_at":"2026-05-24T00:07:25+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":7.0,"formal_verification":"none","one_line_summary":"PANDO introduces an online skill-distillation method with a structured library, reflection, demotion, routing, compression, and cache-aware prompting that reaches 58.3% success on 910 VisualWebArena tasks using 58-61% fewer tokens than prior methods.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2605.19769","ref_index":3,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"OpenComputer: Verifiable Software Worlds for Computer-Use Agents","primary_cat":"cs.AI","submitted_at":"2026-05-19T12:40:29+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":6.0,"formal_verification":"none","one_line_summary":"OpenComputer introduces a verifier-grounded framework with state verifiers, self-evolving layers, task synthesis, and auditable evaluation for 33 desktop apps and 1000 tasks to support computer-use AI agents.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2605.16402","ref_index":2,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"WinDeskGround: A Benchmark for Robust GUI Grounding in Complex Multi-Window Desktop Environments","primary_cat":"cs.CV","submitted_at":"2026-05-13T02:48:52+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":7.0,"formal_verification":"none","one_line_summary":"WinDeskGround is a parametrically generated benchmark of 1,356 instruction-target pairs that reveals accuracy declines in state-of-the-art MLLMs under partial occlusion in multi-window GUI settings.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2605.12501","ref_index":35,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"Covering Human Action Space for Computer Use: Data Synthesis and Benchmark","primary_cat":"cs.CV","submitted_at":"2026-05-12T17:59:58+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":7.0,"formal_verification":"none","one_line_summary":"Presents CUActSpot benchmark and renderer-LLM data synthesis that lets a 4B model outperform larger open-source models on complex computer interactions.","context_count":1,"top_context_role":"background","top_context_polarity":"background","context_text":"InFindings of the Association for Computational Linguistics: ACL 2025, pages 22418-22433, 2025. [34] Haoming Wang, Haoyang Zou, Huatong Song, Jiazhan Feng, Junjie Fang, Junting Lu, Longxi- ang Liu, Qinyu Luo, Shihao Liang, Shijue Huang, et al. Ui-tars-2 technical report: Advancing gui agent with multi-turn reinforcement learning.arXiv preprint arXiv:2509.02544, 2025. [35] Rogerio Bonatti, Dan Zhao, Francesco Bonacci, Dillon Dupont, Sara Abdali, Yinheng Li, Yadong Lu, Justin Wagle, Kazuhito Koishida, Arthur Bucker, et al. Windows agent arena: Evaluating multi-modal os agents at scale.arXiv preprint arXiv:2409.08264, 2024. [36] Christopher Rawles, Sarah Clinckemaillie, Yifan Chang, Jonathan Waltz, Gabrielle Lau, Mary-"},{"citing_arxiv_id":"2605.12481","ref_index":5,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"ToolCUA: Towards Optimal GUI-Tool Path Orchestration for Computer Use Agents","primary_cat":"cs.AI","submitted_at":"2026-05-12T17:57:04+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":6.0,"formal_verification":"none","one_line_summary":"ToolCUA introduces a trajectory scaling pipeline and staged RL to optimize GUI-tool switching, reaching 46.85% accuracy on OSWorld-MCP for a 66% relative gain over baseline.","context_count":1,"top_context_role":"dataset","top_context_polarity":"use_dataset","context_text":"settings, and ToolCUA demonstrates a +3.9% improvement compared with pure GUI actions, demonstrating successful orchestration of GUI and Tool actions in optimal path selection. Additionally, ToolCUA shows out-of-distribution generalization across tasks and platforms, reaching 23.9% on unseenmulti_appsLinux tasks and achieving 33.8% on unseenWindowsdesktop apps in WindowsAgentArena [5]. These results confirm that operating in a hybrid GUI-Tool action space is essential for achieving generalizable and efficient real-world digital automation. Our main contributions are summarized as follows: • We propose anInterleaved GUI-Tool trajectory scaling pipelinethat repurposes existing pure GUI corpora into scalable hybrid-action training data through tool synthesis, obviating the need for manual"},{"citing_arxiv_id":"2605.10912","ref_index":5,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"WildClawBench: A Benchmark for Real-World, Long-Horizon Agent Evaluation","primary_cat":"cs.CL","submitted_at":"2026-05-11T17:49:43+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":8.0,"formal_verification":"none","one_line_summary":"A new native-runtime benchmark reveals that current frontier AI agents succeed on at most 62 percent of realistic long-horizon CLI tasks.","context_count":1,"top_context_role":"background","top_context_polarity":"background","context_text":"evaluation. 2. Related Work Agent Benchmarks across Environments. Agent benchmarks have largely been organized by interaction sur- face: software engineering (SWE-bench [ 17], Terminal-Bench [24], LiveCodeBench [ 16]), web and GUI con- trol (WebArena [59], WebShop [48], VisualWebArena [20]), OS and mobile control (OSWorld [45], Windows Agent Arena [ 5], AndroidWorld [ 33]), enterprise knowledge work (WorkArena [ 11], OdysseyBench [ 38]), interactive coding (AppWorld [ 37]), browsing-centric research (BrowseComp [ 40]), and tool orchestration (ToolBench [31], τ -bench [ 50]). Broader suites such as GAIA [ 25] and TheAgentCompany [ 46] widen task coverage, but most prior benchmarks remain restricted along one or more of the axes summarized in Tab."},{"citing_arxiv_id":"2605.10754","ref_index":4,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"The Agent Use of Agent Beings: Agent Cybernetics Is the Missing Science of Foundation Agents","primary_cat":"cs.AI","submitted_at":"2026-05-11T15:53:54+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":4.0,"formal_verification":"none","one_line_summary":"Agent Cybernetics reframes foundation agent design by adapting classical cybernetics laws into three engineering desiderata for reliable, long-running, self-improving agents.","context_count":1,"top_context_role":"background","top_context_polarity":"background","context_text":"of an autonomous loop that perceives, reasons, acts, and revises its own behavior over extended horizons. Unlike single-shot inference, such agents accumulate state, invoke external tools, and operate across thousands of reasoning steps in pursuit of complex, open-ended goals. The celebrating progress of foundation agents spans across software engineering [8, 18], computer use [4, 42], math proof [15, 35], and scientific discovery [10, 11, 50]. Engineering practice has converged on useful primitives: tool loops [25, 43], memory banks [6, 30], harnesses [20, 51], and reflection steps [29, 32]. However, these successes are assembled by empirical trial and error rather than by principled design. Questions that any mature engineering discipline would consider fundamental remain open: Under"},{"citing_arxiv_id":"2604.27996","ref_index":11,"ref_count":2,"confidence":0.9,"is_internal_anchor":false,"paper_title":"Exploring LLM Agent Designs and Interaction Modalities for Scientific Visualization","primary_cat":"cs.AI","submitted_at":"2026-04-30T15:22:28+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":5.0,"formal_verification":"none","one_line_summary":"Empirical comparison of domain-specific, computer-use, and general-purpose LLM agents plus CLI/GUI modalities on SciVis tasks reveals general-purpose agents highest in success rate but costliest, domain-specific agents more efficient, and persistent memory beneficial depending on mode.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2604.27955","ref_index":8,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"GUI Agents with Reinforcement Learning: Toward Digital Inhabitants","primary_cat":"cs.AI","submitted_at":"2026-04-30T14:51:49+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":5.0,"formal_verification":"none","one_line_summary":"The paper delivers the first comprehensive overview of RL for GUI agents, organizing methods into offline, online, and hybrid strategies while analyzing trends in rewards, efficiency, and deliberation to outline a future roadmap.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2604.21375","ref_index":11,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"VLAA-GUI: Knowing When to Stop, Recover, and Search, A Modular Framework for GUI Automation","primary_cat":"cs.CL","submitted_at":"2026-04-23T07:42:37+00:00","verdict":"CONDITIONAL","verdict_confidence":"MODERATE","novelty_score":6.0,"formal_verification":"none","one_line_summary":"VLAA-GUI adds mandatory visual verifiers, multi-tier loop breakers, and on-demand search to GUI agents, reaching 77.5% on OSWorld and 61.0% on WindowsAgentArena with some models exceeding human performance.","context_count":1,"top_context_role":"dataset","top_context_polarity":"background","context_text":"Han, H. Tu et al. 1 Introduction The rapid advancement of multimodal large language models (MLLMs) [79] has catalyzed a new generation of autonomous Graphical User Interface (GUI) agents [1,33,49,65,66] capable of performing desktop tasks by observing screen- shots and executing mouse and keyboard actions. Systems such as OSWorld [67], WindowsAgentArena [11], and related benchmarks [72,79] have established stan- dardized evaluation environments spanning Linux, Windows, and macOS, re- vealing both promises and persistent limitations of current approaches. De- spite steady progress, two fundamental problems remain largely unsolved. First, agents do not reliably knowwhen a task is finished: they routinely declare suc-"},{"citing_arxiv_id":"2512.13564","ref_index":295,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"Memory in the Age of AI Agents","primary_cat":"cs.CL","submitted_at":"2025-12-15T17:22:34+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":6.0,"formal_verification":"none","one_line_summary":"The paper maps agent memory research via three forms (token-level, parametric, latent), three functions (factual, experiential, working), and dynamics of formation/evolution/retrieval, plus benchmarks and future directions.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2510.05307","ref_index":14,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"When Should Users Check? Modeling Confirmation Frequency inMulti-Step Agentic AI Tasks","primary_cat":"cs.HC","submitted_at":"2025-10-06T19:18:56+00:00","verdict":"CONDITIONAL","verdict_confidence":"LOW","novelty_score":6.0,"formal_verification":"none","one_line_summary":"A decision-theoretic model based on the observed Confirmation-Diagnosis-Correction-Redo user pattern places intermediate confirmations in AI agent tasks, yielding 81% user preference and 13.54% faster completion versus confirm-at-end.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2509.21816","ref_index":6,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"From Task to Tutorial: An Automated GUI Framework for Excel Tutorial Document and Video Creation","primary_cat":"cs.SE","submitted_at":"2025-09-26T03:21:39+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":7.0,"formal_verification":"none","one_line_summary":"An AI framework automates Excel tutorial and video creation from task descriptions via an Execution Agent, achieving 8.5% higher task success and 1/20th the authoring time of experts.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2508.18265","ref_index":7,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency","primary_cat":"cs.CV","submitted_at":"2025-08-25T17:58:17+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":6.0,"formal_verification":"none","one_line_summary":"InternVL3.5 advances open-source multimodal models with Cascade RL for +16% reasoning gains and ViR for 4x inference speedup, with the 241B model reaching SOTA among open-source MLLMs on multimodal, reasoning, and agentic tasks.","context_count":1,"top_context_role":"dataset","top_context_polarity":"use_dataset","context_text":"To assess the GUI agent capabilities of InternVL3.5, we conducted evaluations on a diverse set of platforms. We evaluate GUI grounding capabilities across 3 benchmarks including ScreenSpot [16], ScreenSpot-v2 [150] and OSWorld-G [155]. For online agentic evaluation, our assessment covers Ubuntu, Windows, and Web utilizing the OSWorld [ 156], WindowsAgentArena [7], and WebArena-Lite-v2 [145]. Models with symbols † and ‡ are evaluated under 100 and 200 steps, respectively, while all other results were evaluated under 50 steps. 3.11 GUI Agent Tasks To validate the GUI agent capabilities of InternVL3.5, we conduct detailed experiments in Table 10. In particular, we evaluate InternVL3.5 on six GUI grounding and agent tasks, namely ScreenSpot [16], ScreenSpot-v2 [150],"},{"citing_arxiv_id":"2506.02387","ref_index":10,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"VS-Bench: Evaluating VLMs for Strategic Abilities in Multi-Agent Environments","primary_cat":"cs.AI","submitted_at":"2025-06-03T02:57:38+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":6.0,"formal_verification":"none","one_line_summary":"VS-Bench is a new benchmark of ten visual multi-agent environments that measures VLMs on element recognition, next-action prediction, and normalized episode return, showing strong perception but large gaps in reasoning and decision-making with the best model at 46.6% prediction accuracy and 31.4% of","context_count":1,"top_context_role":"background","top_context_polarity":"background","context_text":"environment: An evaluation platform for general agents. Journal of artificial intelligence research, 47:253-279, 2013. [9] Christopher Berner, Greg Brockman, Brooke Chan, Vicki Cheung, Przemysław D˛ ebiak, Christy Dennison, David Farhi, Quirin Fischer, Shariq Hashme, Chris Hesse, et al. Dota 2 with large scale deep reinforcement learning. arXiv preprint arXiv:1912.06680, 2019. [10] Rogerio Bonatti, Dan Zhao, Francesco Bonacci, Dillon Dupont, Sara Abdali, Yinheng Li, Yadong Lu, Justin Wagle, Kazuhito Koishida, Arthur Bucker, et al. Windows agent arena: Evaluating multi-modal os agents at scale. arXiv preprint arXiv:2409.08264, 2024. [11] Noam Brown and Tuomas Sandholm. Superhuman ai for multiplayer poker. Science, 365(6456):885-890, 2019."},{"citing_arxiv_id":"2505.07062","ref_index":11,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"Seed1.5-VL Technical Report","primary_cat":"cs.CV","submitted_at":"2025-05-11T17:28:30+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":4.0,"formal_verification":"none","one_line_summary":"Seed1.5-VL is a compact multimodal model that sets new records on dozens of vision-language benchmarks and outperforms prior systems on agent-style tasks.","context_count":1,"top_context_role":"baseline","top_context_polarity":"baseline","context_text":"FPS for Dream-1K, and 1 FPS for all other datasets. Capability Benchmark Seed OpenAI Claude UI-TARS Kimi Qwen 2.5 1.5-VL CUA [98] 3.7 Sonnet [6] 1.5 [116] VL-A3B [130] VL 72B [7] GUI Grounding ScreenSpot-V2 [149] 95.2 87.9 87 .6 94 .2 92.8 - ScreenSpot-Pro [72] 60.9 23.4 27 .7 61.6 34.5 43 .6 Computer Use OSWorld [152] 36.7 38 .1 28.0 42.5 8.2 8 .8 Windows Agent Arena [11] 39.6 - 38.9 42.1 10.4 - Browser Use WebVoyager [42] 87.2 87.0 84.1 84 .8 - - Online-Mind2Web [158] 76.4 71.0 62 .9 75 .8 - - Phone Use Android World [111] 62.1 - - 64.2 - 35.0 Table 8 Seed1.5-VL performance on public GUI online benchmarks compared to previous models. 26 Game Seed1.5-VL UI-TARS-1.5 OpenAI CUA Claude 3.7 Sonnet 2048 (score) 870.6 721."},{"citing_arxiv_id":"2504.14239","ref_index":11,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"InfiGUI-R1: Advancing Multimodal GUI Agents from Reactive Actors to Deliberative Reasoners","primary_cat":"cs.AI","submitted_at":"2025-04-19T09:25:55+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":6.0,"formal_verification":"none","one_line_summary":"InfiGUI-R1 uses Reasoning Injection via spatial distillation followed by Deliberation Enhancement via RL to evolve GUI agents from reactive actors to deliberative reasoners, reporting strong performance on grounding and trajectory tasks.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2501.16150","ref_index":10,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"A Comprehensive Survey of Agents for Computer Use: Foundations, Challenges, and Future Directions","primary_cat":"cs.AI","submitted_at":"2025-01-27T15:44:02+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":5.0,"formal_verification":"none","one_line_summary":"A survey of 87 agents for computer use and 33 datasets that introduces a three-dimensional taxonomy across domain, interaction, and agent perspectives and identifies six research gaps.","context_count":1,"top_context_role":"background","top_context_polarity":"background","context_text":"For instance, a user could instruct an agent embedded in a smartphone to propose meeting dates and send them via email. The agent would then operate the phone through simulated touch actions to fulfill the request, as illustrated in Fig. 1a. We refer to this class of agents asagents for computer use (ACUs). Early research on ACUs focused primarily on the learning methodology, particularly reinforcement learning (RL) techniques [10, 59, 61]. Recently, a shift toward integrating foundation models has accelerated progress, significantly enhancing reasoning capabilities and enabling ACUs to tackle increasingly complex tasks [65, 156]. This transition has stimulated research activity, reflected in a strong increase in publications in the field (see Fig. 2). Concurrently, commercial prototypes of instruction-based agents for computer use have begun to emerge"},{"citing_arxiv_id":"2405.14573","ref_index":2,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"AndroidWorld: A Dynamic Benchmarking Environment for Autonomous Agents","primary_cat":"cs.AI","submitted_at":"2024-05-23T13:48:54+00:00","verdict":"ACCEPT","verdict_confidence":"MODERATE","novelty_score":7.0,"formal_verification":"none","one_line_summary":"AndroidWorld is a dynamic, reproducible Android benchmark that generates unlimited natural-language tasks for autonomous agents and shows current agents succeed on only 30.6 percent of them.","context_count":1,"top_context_role":"background","top_context_polarity":"unclear","context_text":"3% and 33.2% (mean 29.0%), obtained using M3A with GPT-4 Turbo with accessibility trees as input. Note that for consistency with existing literature we maintain the single-seed results in Table 3. To better understand the sources of this variability, we evaluate agent robustness under two condi- tions: (1) identical tasks with the same parameters and (2) tasks with different parameter combina- tions, which change the initial state and task definition. We perform this analysis on a representa- tive subset of ANDROID WORLD tasks that span different interaction patterns and complexity levels (listed in Appendix E.4). Due to computational constraints, we conduct 20 trials for each task using our strongest agent configuration - M3A using the accessibility tree and GPT-4."}],"limit":50,"offset":0}