Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T12:10:40.573941Z
Paper Citation Record · LEDGER
As of 23 August 2026, this Paper Citation Record lists 55 of 55 outbound references and 44 inbound Pith citation observations for arXiv:2411.17465.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T12:10:40.573941Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-16T12:03:32.149993Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-05T02:28:24.338817Z
55 of 55 outbound references displayed
External citation measurements
1
pith, observed 2026-08-05T02:28:24.338817Z
Observation 7c670098-3f5a-406e-a521-3801d2e673a8 · outbound
ShowUI: One Vision-Language-Action Model for GUI Visual Agent https://copilot.microsoft
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 80c8f537-d6a8-44d0-878b-802183dde6a5 · outbound
ShowUI: One Vision-Language-Action Model for GUI Visual Agent Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7fb7be2c-25c5-4172-a117-34960f121956 · outbound
ShowUI: One Vision-Language-Action Model for GUI Visual Agent Covla: Comprehensive vision-language-action dataset for autonomous driving
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 76a79547-cf46-4bdc-8e1d-ad7a2f534ae8 · outbound
ShowUI: One Vision-Language-Action Model for GUI Visual Agent Qwen Technical Report
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 79fe3dca-5c55-4c1d-bf30-6ddb8413037a · outbound
ShowUI: One Vision-Language-Action Model for GUI Visual Agent Introducing our multimodal models, 2023
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 6261f0f2-4bf3-446a-8149-fbfbeab5e8ad · outbound
ShowUI: One Vision-Language-Action Model for GUI Visual Agent Token Merging: Your ViT But Faster
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a76e26a7-79d4-4e8a-83e5-5a2e6a124367 · outbound
ShowUI: One Vision-Language-Action Model for GUI Visual Agent RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9ef7ad54-29af-4c62-b261-660b3f510f8e · outbound
ShowUI: One Vision-Language-Action Model for GUI Visual Agent AMEX: Android Multi-annotation Expo Dataset for Mobile GUI Agents
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 70857531-3874-479e-9ad8-5ff0cbb9bb3a · outbound
ShowUI: One Vision-Language-Action Model for GUI Visual Agent An image is worth 1/2 tokens after layer 2: Plug-and-play inference acceleration for large vision-language models
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 3c3f5ffb-0b26-4706-819e-ea952a77af91 · outbound
ShowUI: One Vision-Language-Action Model for GUI Visual Agent GUICourse: From General Vision Language Models to Versatile GUI Agents
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 976251f9-2a25-4b82-8ae3-93e19f22803f · outbound
ShowUI: One Vision-Language-Action Model for GUI Visual Agent SeeClick: Harnessing GUI Grounding for Advanced Visual GUI Agents
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 57b48757-492b-454a-a4a4-623154301756 · outbound
ShowUI: One Vision-Language-Action Model for GUI Visual Agent Mind2web: Towards a generalist agent for the web
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bb453a4d-e2d7-483e-940c-5da05e68663f · outbound
ShowUI: One Vision-Language-Action Model for GUI Visual Agent Multimodal Web Navigation with Instruction-Finetuned Foundation Models
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 788a59e1-cd6a-4f0e-a07c-0d463d17a68a · outbound
ShowUI: One Vision-Language-Action Model for GUI Visual Agent AssistGPT: A General Multi-modal Assistant that can Plan, Execute, Inspect, and Learn
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 558b8420-bdca-4fee-b51e-7c1d00ecbb96 · outbound
ShowUI: One Vision-Language-Action Model for GUI Visual Agent Navigating the Digital World as Humans Do: Universal Visual Grounding for GUI Agents
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a3318932-0b8b-40bf-9f8c-1d2ef43710fb · outbound
ShowUI: One Vision-Language-Action Model for GUI Visual Agent A Real-World WebAgent with Planning, Long Context Understanding, and Program Synthesis
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1f74c328-1b8d-46aa-b884-fcec144b2bf9 · outbound
ShowUI: One Vision-Language-Action Model for GUI Visual Agent CogAgent: A Visual Language Model for GUI Agents
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 593d74e8-2242-46ad-a6ee-4a8ee57df14c · outbound
ShowUI: One Vision-Language-Action Model for GUI Visual Agent A3VLM: Actionable Articulation-Aware Vision Language Model
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2d90eedd-8a1f-47b6-84c6-e497dd4e6867 · outbound
ShowUI: One Vision-Language-Action Model for GUI Visual Agent A data-driven approach for learning to control computers
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation e80c56fc-9f2e-4c2c-9f64-192bfedc12de · outbound
ShowUI: One Vision-Language-Action Model for GUI Visual Agent Chat-univi: Unified visual representation em- powers large language models with image and video un- derstanding
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 7158d439-6dad-4389-87cd-26953bd0845e · outbound
ShowUI: One Vision-Language-Action Model for GUI Visual Agent OmniACT: A Dataset and Benchmark for Enabling Multimodal Generalist Autonomous Agents for Desktop and Web
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c9d19f04-3d22-4c57-a5d4-2a8b58dcc2d0 · outbound
ShowUI: One Vision-Language-Action Model for GUI Visual Agent OpenVLA: An Open-Source Vision-Language-Action Model
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8ab48d97-4a22-42de-bb29-b8288f30b61d · outbound
ShowUI: One Vision-Language-Action Model for GUI Visual Agent Pix2struct: Screenshot parsing as pretraining for visual lan- guage understanding
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 2257d4b5-c773-4990-ac38-37d73ede5a0c · outbound
ShowUI: One Vision-Language-Action Model for GUI Visual Agent Mapping Natural Language Instructions to Mobile UI Action Sequences
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c6ea514d-6b3c-4c38-86f9-c08941ca8b9a · outbound
ShowUI: One Vision-Language-Action Model for GUI Visual Agent Llama-vid: An image is worth 2 tokens in large language models
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 12a3cd1f-a442-4b5e-af78-8ba2f637784e · outbound
ShowUI: One Vision-Language-Action Model for GUI Visual Agent Ferret-UI 2: Mastering Universal User Interface Understanding Across Platforms
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 46498d31-b026-4853-9157-230c44b2a370 · outbound
ShowUI: One Vision-Language-Action Model for GUI Visual Agent Steve-1: A generative model for text-to- behavior in minecraft
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 0c05b605-94bc-40ca-84ca-f59af5c86c7f · outbound
ShowUI: One Vision-Language-Action Model for GUI Visual Agent VideoGUI: A Benchmark for GUI Automation from Instructional Videos
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e9b71702-94be-4214-8b94-2043f001864a · outbound
ShowUI: One Vision-Language-Action Model for GUI Visual Agent GUIOdyssey: A Comprehensive Dataset for Cross-App GUI Navigation on Mobile Devices
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 75f95677-6055-4f9a-8375-d79ffdb9397c · outbound
ShowUI: One Vision-Language-Action Model for GUI Visual Agent OmniParser for Pure Vision Based GUI Agent
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation de9fcea8-f743-4b36-a642-a14d074b9347 · outbound
ShowUI: One Vision-Language-Action Model for GUI Visual Agent Gpt-4 technical report, 2023
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 76ea9d6a-b5ea-472d-b70a-8612b48d9591 · outbound
ShowUI: One Vision-Language-Action Model for GUI Visual Agent Open X-Embodiment: Robotic Learning Datasets and RT-X Models
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0aa31d74-aaf4-4137-a9af-13152efbb5fd · outbound
ShowUI: One Vision-Language-Action Model for GUI Visual Agent Pyautogui
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation b712552d-a83a-4e6b-bbf5-74d02ba10c0a · outbound
ShowUI: One Vision-Language-Action Model for GUI Visual Agent Mixture-of-Depths: Dynamically allocating compute in transformer-based language models
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 50bcbe0d-de34-491d-89ba-89e284b78ce9 · outbound
ShowUI: One Vision-Language-Action Model for GUI Visual Agent Android in the Wild: A Large-Scale Dataset for Android Device Control
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6c621243-67bc-4608-b987-8230920a4a49 · outbound
ShowUI: One Vision-Language-Action Model for GUI Visual Agent Android in the wild: a large- scale dataset for android device control
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation a80fbf24-fce4-4b37-a9ec-1a196c6d3eb2 · outbound
ShowUI: One Vision-Language-Action Model for GUI Visual Agent Llava-prumerge: Adaptive token reduction for efficient large multimodal models
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2359f70b-b22d-4b3b-af6b-a86545c8b2d0 · outbound
ShowUI: One Vision-Language-Action Model for GUI Visual Agent From pixels to ui ac- tions: Learning to follow instructions via graphical user in- terfaces
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 4c417f3e-9090-4070-9fe7-bec873bea0b5 · outbound
ShowUI: One Vision-Language-Action Model for GUI Visual Agent World of bits: An open-domain plat- form for web-based agents
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 5dbef036-307a-4762-855d-2f8f14e32864 · outbound
ShowUI: One Vision-Language-Action Model for GUI Visual Agent Qwen2-vl: Enhancing vision-language model’s perception of the world at any resolution, 2024
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 8336ed07-9009-4f6b-a3b8-71520e6978b3 · outbound
ShowUI: One Vision-Language-Action Model for GUI Visual Agent Omnijarvis: Unified vision-language-action to- kenization enables open-world instruction following agents, 2024
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 0525132e-f123-486f-a338-5b0f6deedf7f · outbound
ShowUI: One Vision-Language-Action Model for GUI Visual Agent Chain-of-thought prompting elicits reasoning in large lan- guage models
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dd2ce77e-ec5e-424a-977b-ead6099935b3 · outbound
ShowUI: One Vision-Language-Action Model for GUI Visual Agent Webui: A dataset for en- hancing visual ui understanding with web semantics
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 4837df9a-c0cc-477f-81bf-59677419f289 · outbound
ShowUI: One Vision-Language-Action Model for GUI Visual Agent Videollm-mod: Efficient video-language streaming with mixture-of-depths vision computation
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 802cad14-f3a6-40c7-ac1f-aeb857e351bd · outbound
ShowUI: One Vision-Language-Action Model for GUI Visual Agent GPT-4V in Wonderland: Large Multimodal Models for Zero-Shot Smartphone GUI Navigation
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bc7cc6d7-8ee8-4c2a-a2a4-b73be0b2d9b4 · outbound
ShowUI: One Vision-Language-Action Model for GUI Visual Agent Set-of-mark prompting unleashes ex- traordinary visual grounding in gpt-4v, 2023
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 5a5bb20b-7599-4fef-b58c-536ca036e1c1 · outbound
ShowUI: One Vision-Language-Action Model for GUI Visual Agent Gpt4tools: Teaching large language model to use tools via self-instruction
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 63e3b143-8fa3-49a3-a4cf-b7779c7eb18e · outbound
ShowUI: One Vision-Language-Action Model for GUI Visual Agent React: Synergizing rea- soning and acting in language models
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 4f21432b-e5c7-4b46-809b-8134bc812366 · outbound
ShowUI: One Vision-Language-Action Model for GUI Visual Agent Ferret-UI: Grounded Mobile UI Understanding with Multimodal LLMs
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e6e3edac-4335-44a3-8dcb-0934636c89d1 · outbound
ShowUI: One Vision-Language-Action Model for GUI Visual Agent So- cratic models: Composing zero-shot multimodal reasoning with language
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 1bbda601-87c9-46ac-b059-ed25d56d7e62 · outbound
ShowUI: One Vision-Language-Action Model for GUI Visual Agent xLAM: A Family of Large Action Models to Empower AI Agent Systems
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 72b6d15f-8d44-480c-a07c-9ef99f016fea · outbound
ShowUI: One Vision-Language-Action Model for GUI Visual Agent You only look at screens: Multimodal chain-of-action agents
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 9c8a6d4a-033b-4043-b8b2-31194363a28f · outbound
ShowUI: One Vision-Language-Action Model for GUI Visual Agent 3D-VLA: A 3D Vision-Language-Action Generative World Model
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ae096c46-5a6a-4be8-ab23-fbc791c74ad0 · outbound
ShowUI: One Vision-Language-Action Model for GUI Visual Agent Gpt-4v(ision) is a generalist web agent, if grounded
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 87da435c-a45e-4b4b-a759-1df2e0a9aef3 · outbound
ShowUI: One Vision-Language-Action Model for GUI Visual Agent WebArena: A Realistic Web Environment for Building Autonomous Agents
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5d7febd8-656a-443a-a059-5faad6a0eeab · inbound
Improved GUI Grounding via Iterative Narrowing ShowUI: One Vision-Language-Action Model for GUI Visual Agent
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f94b606f-5219-435c-8c71-da418617373b · inbound
Large Language Model-Brained GUI Agents: A Survey ShowUI: One Vision-Language-Action Model for GUI Visual Agent
Reference 246
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 2890b80e-82f0-4366-88df-721edd3d8bd6 · inbound
WorldGUI: An Interactive Benchmark for Desktop GUI Automation from Any Starting Point ShowUI: One Vision-Language-Action Model for GUI Visual Agent
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2311856d-2514-42ac-87a7-3d399d889f4b · inbound
LearnAct: Few-Shot Mobile GUI Agent with a Unified Demonstration Benchmark ShowUI: One Vision-Language-Action Model for GUI Visual Agent
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1ae04baf-0455-4b5a-bee0-d3446246659e · inbound
InfantAgent-Next: A Multimodal Generalist Agent for Automated Computer Interaction ShowUI: One Vision-Language-Action Model for GUI Visual Agent
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 028484cb-2496-47b1-a36c-1d4c57e55cc7 · inbound
ReGUIDE: Data Efficient GUI Grounding via Spatial Reasoning and Search ShowUI: One Vision-Language-Action Model for GUI Visual Agent
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d1892b3a-10fd-4e5a-8075-68b746c4343a · inbound
GUI-G1: Understanding R1-Zero-Like Training for Visual Grounding in GUI Agents ShowUI: One Vision-Language-Action Model for GUI Visual Agent
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 82470402-8cf8-40f0-8292-4cdc19026485 · inbound
BacktrackAgent: Enhancing GUI Agent with Error Detection and Backtracking Mechanism ShowUI: One Vision-Language-Action Model for GUI Visual Agent
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 01f5e6b7-1193-492b-b807-c30c44dc20db · inbound
Grounded Reinforcement Learning for Visual Reasoning ShowUI: One Vision-Language-Action Model for GUI Visual Agent
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 4fe73d73-f7ee-41c9-94e3-c0534fbcf123 · inbound
ZeroGUI: Automating Online GUI Learning at Zero Human Cost ShowUI: One Vision-Language-Action Model for GUI Visual Agent
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 038741e4-74ad-4c5c-9a4a-56f7f7ac5b94 · inbound
RiOSWorld: Benchmarking the Risk of Multimodal Computer-Use Agents ShowUI: One Vision-Language-Action Model for GUI Visual Agent
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 19d78170-56f1-40a2-b7d3-5490c874e824 · inbound
GUI-Actor: Coordinate-Free Visual Grounding for GUI Agents ShowUI: One Vision-Language-Action Model for GUI Visual Agent
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 83b75b66-3873-4d89-8e09-65c9ab7b1b8b · inbound
What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities ShowUI: One Vision-Language-Action Model for GUI Visual Agent
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 604549a2-8239-4c56-9bb1-6062220719f0 · inbound
Atomic-to-Compositional Generalization for Mobile Agents with A New Benchmark and Scheduling System ShowUI: One Vision-Language-Action Model for GUI Visual Agent
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a9b674cc-0ab6-4c3a-8888-a49c58b29e6f · inbound
GUI-Robust: A Comprehensive Dataset for Testing GUI Agent Robustness in Real-World Anomalies ShowUI: One Vision-Language-Action Model for GUI Visual Agent
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ed1bf49c-f7f2-4d16-b271-6706a3f5e214 · inbound
Understanding GUI Agent Localization Biases through Logit Sharpness ShowUI: One Vision-Language-Action Model for GUI Visual Agent
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7f4c0137-4cdd-40bf-8b5b-2e0533244c88 · inbound
Learning, Reasoning, Refinement: A Framework for Kahneman's Dual-System Intelligence in GUI Agents ShowUI: One Vision-Language-Action Model for GUI Visual Agent
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 44f450de-7495-43d0-9caa-5f74dfb736c0 · inbound
ZonUI-3B: A Lightweight Vision-Language Model for Cross-Resolution GUI Grounding ShowUI: One Vision-Language-Action Model for GUI Visual Agent
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2fd21acb-63aa-437f-97d1-f4854ec9f7cb · inbound
PresentAgent: Multimodal Agent for Presentation Video Generation ShowUI: One Vision-Language-Action Model for GUI Visual Agent
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f1cc3105-4402-434e-b218-a93148939593 · inbound
DreamVLA: A Vision-Language-Action Model Dreamed with Comprehensive World Knowledge ShowUI: One Vision-Language-Action Model for GUI Visual Agent
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 502ea755-018e-4996-ba99-3175a3291e2c · inbound
GUI-G$^2$: Gaussian Reward Modeling for GUI Grounding ShowUI: One Vision-Language-Action Model for GUI Visual Agent
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fb3aec4d-05fc-4eab-ba62-092c7b43bebb · inbound
Screen2AX: Vision-Based Approach for Automatic macOS Accessibility Generation ShowUI: One Vision-Language-Action Model for GUI Visual Agent
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1faa229f-368a-4280-a1e4-b9a998812fdc · inbound
MMBench-GUI: Hierarchical Multi-Platform Evaluation Framework for GUI Agents ShowUI: One Vision-Language-Action Model for GUI Visual Agent
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cc18a95a-d7f0-4b16-a168-4c714c8654e2 · inbound
SEAgent: Self-Evolving Computer Use Agent with Autonomous Learning from Experience ShowUI: One Vision-Language-Action Model for GUI Visual Agent
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 96710486-0104-4b5f-8a91-a545d55cb11b · inbound
CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning ShowUI: One Vision-Language-Action Model for GUI Visual Agent
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f93935eb-d7b6-4d83-823a-06fb7832e71c · inbound
MobiAgent: A Systematic Framework for Customizable Mobile Agents ShowUI: One Vision-Language-Action Model for GUI Visual Agent
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3619084a-7ace-4ad1-92f4-9831062dc879 · inbound
PG-Agent: An Agent Powered by Page Graph ShowUI: One Vision-Language-Action Model for GUI Visual Agent
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4573b04a-c972-4ab2-b201-0eee654c3d57 · inbound
Mitigating Coordinate Prediction Bias from Positional Encoding Failures ShowUI: One Vision-Language-Action Model for GUI Visual Agent
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation c09c03e6-46bb-4a3a-b492-3c3eccfe00d0 · inbound
GUI-AIMA: Aligning Intrinsic Multimodal Attention with a Context Anchor for GUI Grounding ShowUI: One Vision-Language-Action Model for GUI Visual Agent
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8a3e9fe5-f353-494a-8a6d-e3914f42405a · inbound
Grounding Computer Use Agents on Human Demonstrations ShowUI: One Vision-Language-Action Model for GUI Visual Agent
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 37e5ef16-a407-4fb9-bb37-3c778e521058 · inbound
GUI-Libra: Training Native GUI Agents to Reason and Act with Action-aware Supervision and Partially Verifiable RL ShowUI: One Vision-Language-Action Model for GUI Visual Agent
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d346d563-3f51-40b0-85a8-67c4c7a9f2db · inbound
GUIDE: Resolving Domain Bias in GUI Agents through Real-Time Web Video Retrieval and Plug-and-Play Annotation ShowUI: One Vision-Language-Action Model for GUI Visual Agent
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d261bd46-3e49-43de-841f-377eba841c26 · inbound
UI-Zoomer: Uncertainty-Driven Adaptive Zoom-In for GUI Grounding ShowUI: One Vision-Language-Action Model for GUI Visual Agent
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 5a0e84a0-0dd1-4ce2-a42e-d12a6024b301 · inbound
VLAA-GUI: Knowing When to Stop, Recover, and Search, A Modular Framework for GUI Automation ShowUI: One Vision-Language-Action Model for GUI Visual Agent
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation bcb010b6-2ea1-43bd-a68e-3974373b6513 · inbound
Do Vision-Language-Models show human-like logical problem-solving capability in point and click puzzle games? ShowUI: One Vision-Language-Action Model for GUI Visual Agent
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 1bd235fc-16b4-42df-99dd-ddaf0219dbca · inbound
Skim: Speculative Execution for Fast and Efficient Web Agents ShowUI: One Vision-Language-Action Model for GUI Visual Agent
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 7cd99c0c-668e-45a6-a8d9-0625225c3568 · inbound
GUI-C$^2$: Coarse-to-Fine GUI Grounding via Difficulty-Aware Reinforcement Learning ShowUI: One Vision-Language-Action Model for GUI Visual Agent
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 45d438a7-4e55-43d2-8f87-3af3cc421c45 · inbound
OSWorld 2.0: Benchmarking Computer Use Agents on Long-Horizon Real-World Tasks ShowUI: One Vision-Language-Action Model for GUI Visual Agent
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 33244b6a-53c8-4fce-b7e7-641a23df1245 · inbound
OSWorld 2.0: Benchmarking Computer Use Agents on Long-Horizon Real-World Tasks ShowUI: One Vision-Language-Action Model for GUI Visual Agent
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cd1725ad-d3e7-4fe6-bb89-4f6507c86a23 · inbound
Do GUI Agents Believe Their Eyes? Diagnosing State-Belief Reliance on Pixels versus Structure ShowUI: One Vision-Language-Action Model for GUI Visual Agent
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6ee9538c-55f7-44f6-a56c-cf5914914214 · inbound
Do GUI Agents Believe Their Eyes? Diagnosing State-Belief Reliance on Pixels versus Structure ShowUI: One Vision-Language-Action Model for GUI Visual Agent
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9a08a3d2-bf56-40cb-a8e0-2c9632d1a57e · inbound
Vision as Unified Multimodal Generation ShowUI: One Vision-Language-Action Model for GUI Visual Agent
Reference 108
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation db8235ac-0579-45c3-84ea-dcf60d4b40ca · inbound
StepX-Edge: An On-Device UI Vision-Language Model via Architecture-Training-Deployment Co-Design ShowUI: One Vision-Language-Action Model for GUI Visual Agent
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ba010f80-452c-400c-b9ce-9af0e7cc0593 · inbound
Rethinking Inference-Time Scaling in Local Computer-Use Agents: Failure Modes and Compute Tradeoffs ShowUI: One Vision-Language-Action Model for GUI Visual Agent
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.