Pith. sign in

Paper Citation Record · LEDGER

ShowUI: One Vision-Language-Action Model for GUI Visual Agent

As of 23 August 2026, this Paper Citation Record lists 55 of 55 outbound references and 44 inbound Pith citation observations for arXiv:2411.17465.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.17465 v1

Coverage vector

measured 55 of 55 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T12:10:40.573941Z

measured 99 of 99 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 44 of 44 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T12:03:32.149993Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

55 of 55 outbound references displayed

  • verified exact0
  • verified fuzzy22
  • unresolved33
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

1
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 7c670098-3f5a-406e-a521-3801d2e673a8 · outbound

This paper cites https://copilot.microsoft.

ShowUI: One Vision-Language-Action Model for GUI Visual Agent https://copilot.microsoft

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:10:41.584014Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T12:10:40.282949Z digest=sha256:bc47891c424d79c3b3d3fe76bb215f3ec73001e94e4d18fb939069cb1875dc49

Observation 80c8f537-d6a8-44d0-878b-802183dde6a5 · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

ShowUI: One Vision-Language-Action Model for GUI Visual Agent Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T12:10:40.289530Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:10:40.289530Z digest=sha256:4babb79334951845aa3baf80cfd8092378bf6c146cc17b2cdc8be6dc0bbe4b0f

Observation 7fb7be2c-25c5-4172-a117-34960f121956 · outbound

This paper cites Covla: Comprehensive vision-language-action dataset for autonomous driving.

ShowUI: One Vision-Language-Action Model for GUI Visual Agent Covla: Comprehensive vision-language-action dataset for autonomous driving

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T12:10:40.296273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:10:40.296273Z digest=sha256:13626c14a28478925060a45ca7b7162fdcb2aeabb3b2c60f91b3a76167512e93

Observation 76a79547-cf46-4bdc-8e1d-ad7a2f534ae8 · outbound

This paper cites Qwen Technical Report.

ShowUI: One Vision-Language-Action Model for GUI Visual Agent Qwen Technical Report

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T12:10:40.301795Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:10:40.301795Z digest=sha256:b35d1414c0d94592265d3106742228acbf55fb91336017d4f758a7c7b1bcc1de

Observation 79fe3dca-5c55-4c1d-bf30-6ddb8413037a · outbound

This paper cites Introducing our multimodal models, 2023.

ShowUI: One Vision-Language-Action Model for GUI Visual Agent Introducing our multimodal models, 2023

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:10:41.568843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T12:10:40.308046Z digest=sha256:255911776a2e223ddd57e6e5f7e49e8910f02d015faec9cf32b11d51dbd36d86

Observation 6261f0f2-4bf3-446a-8149-fbfbeab5e8ad · outbound

This paper cites Token Merging: Your ViT But Faster.

ShowUI: One Vision-Language-Action Model for GUI Visual Agent Token Merging: Your ViT But Faster

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T12:10:40.313340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:10:40.313340Z digest=sha256:bdc93cfde64fddf373e6d8bf9037897d832c281160faaaeab2e03120e2161103

Observation a76e26a7-79d4-4e8a-83e5-5a2e6a124367 · outbound

This paper cites RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control.

ShowUI: One Vision-Language-Action Model for GUI Visual Agent RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T12:10:40.319251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:10:40.319251Z digest=sha256:f9e62a2f30c379a6d6982bbc6e118a1e1938492e12488aa75a75e24fe1ab9f0f

Observation 9ef7ad54-29af-4c62-b261-660b3f510f8e · outbound

This paper cites AMEX: Android Multi-annotation Expo Dataset for Mobile GUI Agents.

ShowUI: One Vision-Language-Action Model for GUI Visual Agent AMEX: Android Multi-annotation Expo Dataset for Mobile GUI Agents

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T12:10:40.324753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:10:40.324753Z digest=sha256:43c57c2367beed310052141175c319a7e1eef4e6625dc44460d2d4ae4e3597bb

Observation 70857531-3874-479e-9ad8-5ff0cbb9bb3a · outbound

This paper cites An image is worth 1/2 tokens after layer 2: Plug-and-play inference acceleration for large vision-language models.

ShowUI: One Vision-Language-Action Model for GUI Visual Agent An image is worth 1/2 tokens after layer 2: Plug-and-play inference acceleration for large vision-language models

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:10:41.553089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T12:10:40.329666Z digest=sha256:ee7a1150576d3b0636240fa54f06ced89e49a9109d2535a87aac4b776829334a

Observation 3c3f5ffb-0b26-4706-819e-ea952a77af91 · outbound

This paper cites GUICourse: From General Vision Language Models to Versatile GUI Agents.

ShowUI: One Vision-Language-Action Model for GUI Visual Agent GUICourse: From General Vision Language Models to Versatile GUI Agents

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T12:10:40.334993Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:10:40.334993Z digest=sha256:b03393c69b5975e937eaece7a7c926d751eb0319b13b9c46b2ae7b208824865e

Observation 976251f9-2a25-4b82-8ae3-93e19f22803f · outbound

This paper cites SeeClick: Harnessing GUI Grounding for Advanced Visual GUI Agents.

ShowUI: One Vision-Language-Action Model for GUI Visual Agent SeeClick: Harnessing GUI Grounding for Advanced Visual GUI Agents

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T12:10:40.340201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:10:40.340201Z digest=sha256:74995d318974ad51b82ba99770391d456b5000e7d88ac57be1a751266edd6eee

Observation 57b48757-492b-454a-a4a4-623154301756 · outbound

This paper cites Mind2web: Towards a generalist agent for the web.

ShowUI: One Vision-Language-Action Model for GUI Visual Agent Mind2web: Towards a generalist agent for the web

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T12:10:40.344940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:10:40.344940Z digest=sha256:3520a9942fe08b806776c6e0a9dee22c848400bba5490c9ef52d394dd172599a

Observation bb453a4d-e2d7-483e-940c-5da05e68663f · outbound

This paper cites Multimodal Web Navigation with Instruction-Finetuned Foundation Models.

ShowUI: One Vision-Language-Action Model for GUI Visual Agent Multimodal Web Navigation with Instruction-Finetuned Foundation Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T12:10:40.350697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:10:40.350697Z digest=sha256:2cac55b4b9ec34b342a3a51851c8e4c4cbfc1166fa407da7594a65ad961e4320

Observation 788a59e1-cd6a-4f0e-a07c-0d463d17a68a · outbound

This paper cites AssistGPT: A General Multi-modal Assistant that can Plan, Execute, Inspect, and Learn.

ShowUI: One Vision-Language-Action Model for GUI Visual Agent AssistGPT: A General Multi-modal Assistant that can Plan, Execute, Inspect, and Learn

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T12:10:40.355704Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:10:40.355704Z digest=sha256:485fad98e0a4224100d6fb6c95bcbefdf1cffc4d852139c8fd0d2e4d588e6968

Observation 558b8420-bdca-4fee-b51e-7c1d00ecbb96 · outbound

This paper cites Navigating the Digital World as Humans Do: Universal Visual Grounding for GUI Agents.

ShowUI: One Vision-Language-Action Model for GUI Visual Agent Navigating the Digital World as Humans Do: Universal Visual Grounding for GUI Agents

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T12:10:40.360940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:10:40.360940Z digest=sha256:6a3bf4632c19aee9745abf4f5d731b3b22cb5e4883c94d48cba42b8d448e58ed

Observation a3318932-0b8b-40bf-9f8c-1d2ef43710fb · outbound

This paper cites A Real-World WebAgent with Planning, Long Context Understanding, and Program Synthesis.

ShowUI: One Vision-Language-Action Model for GUI Visual Agent A Real-World WebAgent with Planning, Long Context Understanding, and Program Synthesis

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T12:10:40.366402Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:10:40.366402Z digest=sha256:c906634d99a6b2643a20129d7182f98adccfbf03d1b43cf328c654f036b77aa3

Observation 1f74c328-1b8d-46aa-b884-fcec144b2bf9 · outbound

This paper cites CogAgent: A Visual Language Model for GUI Agents.

ShowUI: One Vision-Language-Action Model for GUI Visual Agent CogAgent: A Visual Language Model for GUI Agents

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T12:10:40.371481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:10:40.371481Z digest=sha256:97b1a219ee5d25e6be4aa49a4df62f039e4ae5c52dd19c1a8eba44a83f5de943

Observation 593d74e8-2242-46ad-a6ee-4a8ee57df14c · outbound

This paper cites A3VLM: Actionable Articulation-Aware Vision Language Model.

ShowUI: One Vision-Language-Action Model for GUI Visual Agent A3VLM: Actionable Articulation-Aware Vision Language Model

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T12:10:40.376762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:10:40.376762Z digest=sha256:deb3a2c455b30eb8f2b011760a29a89eeb6ae9bbe4405b8778d081bbe835829d

Observation 2d90eedd-8a1f-47b6-84c6-e497dd4e6867 · outbound

This paper cites A data-driven approach for learning to control computers.

ShowUI: One Vision-Language-Action Model for GUI Visual Agent A data-driven approach for learning to control computers

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:10:41.525946Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T12:10:40.381875Z digest=sha256:c4cb3e2c0a7cd7a45bcb608f1eb1e048809d50a693057cd8221cdd9e339b6018

Observation e80c56fc-9f2e-4c2c-9f64-192bfedc12de · outbound

This paper cites Chat-univi: Unified visual representation em- powers large language models with image and video un- derstanding.

ShowUI: One Vision-Language-Action Model for GUI Visual Agent Chat-univi: Unified visual representation em- powers large language models with image and video un- derstanding

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:10:41.510795Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T12:10:40.391161Z digest=sha256:0e6915554731f4724946d111e5b679bae120fc4b5d1398dc7d2a4465a7bdbf04

Observation 7158d439-6dad-4389-87cd-26953bd0845e · outbound

This paper cites OmniACT: A Dataset and Benchmark for Enabling Multimodal Generalist Autonomous Agents for Desktop and Web.

ShowUI: One Vision-Language-Action Model for GUI Visual Agent OmniACT: A Dataset and Benchmark for Enabling Multimodal Generalist Autonomous Agents for Desktop and Web

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T12:10:40.395915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:10:40.395915Z digest=sha256:603a875e99e67575be9f3c62e546792ad71a229484b6a7c3203584534842c158

Observation c9d19f04-3d22-4c57-a5d4-2a8b58dcc2d0 · outbound

This paper cites OpenVLA: An Open-Source Vision-Language-Action Model.

ShowUI: One Vision-Language-Action Model for GUI Visual Agent OpenVLA: An Open-Source Vision-Language-Action Model

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T12:10:40.401290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:10:40.401290Z digest=sha256:0e7a70c05f3782276291eb0c98d239b45099dee4f3de69b55958ae387267748c

Observation 8ab48d97-4a22-42de-bb29-b8288f30b61d · outbound

This paper cites Pix2struct: Screenshot parsing as pretraining for visual lan- guage understanding.

ShowUI: One Vision-Language-Action Model for GUI Visual Agent Pix2struct: Screenshot parsing as pretraining for visual lan- guage understanding

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:10:41.495692Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T12:10:40.407954Z digest=sha256:a09fcb709108350d6d23d10d63b1c47ebb03dbfe8e72ee538a421f7f79935510

Observation 2257d4b5-c773-4990-ac38-37d73ede5a0c · outbound

This paper cites Mapping Natural Language Instructions to Mobile UI Action Sequences.

ShowUI: One Vision-Language-Action Model for GUI Visual Agent Mapping Natural Language Instructions to Mobile UI Action Sequences

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T12:10:40.413986Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:10:40.413986Z digest=sha256:aaeb0fc3d6abfef6c8c845d585c756955ff90498f7f8e853e1958bb95c99e195

Observation c6ea514d-6b3c-4c38-86f9-c08941ca8b9a · outbound

This paper cites Llama-vid: An image is worth 2 tokens in large language models.

ShowUI: One Vision-Language-Action Model for GUI Visual Agent Llama-vid: An image is worth 2 tokens in large language models

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:10:41.479859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T12:10:40.419379Z digest=sha256:92f1a93949cbd70cbc25959053f5b4ea2050472df6c2046ec86f093f3946ed8d

Observation 12a3cd1f-a442-4b5e-af78-8ba2f637784e · outbound

This paper cites Ferret-UI 2: Mastering Universal User Interface Understanding Across Platforms.

ShowUI: One Vision-Language-Action Model for GUI Visual Agent Ferret-UI 2: Mastering Universal User Interface Understanding Across Platforms

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T12:10:40.424171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:10:40.424171Z digest=sha256:26e1cdd489f910b90d5e66b367bb6139305b065a52916e96c4074f085099918d

Observation 46498d31-b026-4853-9157-230c44b2a370 · outbound

This paper cites Steve-1: A generative model for text-to- behavior in minecraft.

ShowUI: One Vision-Language-Action Model for GUI Visual Agent Steve-1: A generative model for text-to- behavior in minecraft

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:10:41.463433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T12:10:40.429577Z digest=sha256:ee8627e11093281ee96be65c68024720f51f54cd06d2e43d7b98b31f0ab54424

Observation 0c05b605-94bc-40ca-84ca-f59af5c86c7f · outbound

This paper cites VideoGUI: A Benchmark for GUI Automation from Instructional Videos.

ShowUI: One Vision-Language-Action Model for GUI Visual Agent VideoGUI: A Benchmark for GUI Automation from Instructional Videos

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T12:10:40.435268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:10:40.435268Z digest=sha256:ea1f867698228ea64449162d91c22e3752a008a024defc0c3cc82d0e1c931f7a

Observation e9b71702-94be-4214-8b94-2043f001864a · outbound

This paper cites GUIOdyssey: A Comprehensive Dataset for Cross-App GUI Navigation on Mobile Devices.

ShowUI: One Vision-Language-Action Model for GUI Visual Agent GUIOdyssey: A Comprehensive Dataset for Cross-App GUI Navigation on Mobile Devices

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-12T12:10:40.441116Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:10:40.441116Z digest=sha256:352b73d92e75d49db73b8be42d5b1582b3ee1de6fc6e1d41329038337efb9040

Observation 75f95677-6055-4f9a-8375-d79ffdb9397c · outbound

This paper cites OmniParser for Pure Vision Based GUI Agent.

ShowUI: One Vision-Language-Action Model for GUI Visual Agent OmniParser for Pure Vision Based GUI Agent

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-12T12:10:40.447651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:10:40.447651Z digest=sha256:3573d7ef4d8184a87d3b2880a003d763cf38521de6b4a1ae6b0a46465d3b73f6

Observation de9fcea8-f743-4b36-a642-a14d074b9347 · outbound

This paper cites Gpt-4 technical report, 2023.

ShowUI: One Vision-Language-Action Model for GUI Visual Agent Gpt-4 technical report, 2023

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-12T12:10:40.453387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:10:40.453387Z digest=sha256:52ebd68ff4ceb421071ce7b3cf2e66045877cb4d78bd538e5c70773013c5f4b1

Observation 76ea9d6a-b5ea-472d-b70a-8612b48d9591 · outbound

This paper cites Open X-Embodiment: Robotic Learning Datasets and RT-X Models.

ShowUI: One Vision-Language-Action Model for GUI Visual Agent Open X-Embodiment: Robotic Learning Datasets and RT-X Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T12:10:40.457894Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:10:40.457894Z digest=sha256:60252a34722013860cb8b1ef8fa3e8bedb7eea417668b14b30c119cd393db305

Observation 0aa31d74-aaf4-4137-a9af-13152efbb5fd · outbound

This paper cites Pyautogui.

ShowUI: One Vision-Language-Action Model for GUI Visual Agent Pyautogui

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:10:41.437727Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T12:10:40.464555Z digest=sha256:067b14ca5f02663c049c8ce690a9e4d90599ee02eb4d5e41013bd383cd0b9007

Observation b712552d-a83a-4e6b-bbf5-74d02ba10c0a · outbound

This paper cites Mixture-of-Depths: Dynamically allocating compute in transformer-based language models.

ShowUI: One Vision-Language-Action Model for GUI Visual Agent Mixture-of-Depths: Dynamically allocating compute in transformer-based language models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-12T12:10:40.469361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:10:40.469361Z digest=sha256:e0cce997bb1a945e45b5c1da2675871c5ac00858fbeae687f7ede2d434500c53

Observation 50bcbe0d-de34-491d-89ba-89e284b78ce9 · outbound

This paper cites Android in the Wild: A Large-Scale Dataset for Android Device Control.

ShowUI: One Vision-Language-Action Model for GUI Visual Agent Android in the Wild: A Large-Scale Dataset for Android Device Control

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T12:10:40.474066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:10:40.474066Z digest=sha256:a218db54fe75caab82d31e42677cf8be61daf0f6a7d8741d3c46cc11a42625db

Observation 6c621243-67bc-4608-b987-8230920a4a49 · outbound

This paper cites Android in the wild: a large- scale dataset for android device control.

ShowUI: One Vision-Language-Action Model for GUI Visual Agent Android in the wild: a large- scale dataset for android device control

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:10:41.420661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T12:10:40.479014Z digest=sha256:5f145841a45e6914434d28234a5899af9a7a6cdb109c3c5dd50fe70d85d09952

Observation a80fbf24-fce4-4b37-a9ec-1a196c6d3eb2 · outbound

This paper cites Llava-prumerge: Adaptive token reduction for efficient large multimodal models.

ShowUI: One Vision-Language-Action Model for GUI Visual Agent Llava-prumerge: Adaptive token reduction for efficient large multimodal models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-12T12:10:40.483829Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:10:40.483829Z digest=sha256:9bfe2c686e2aabff9b5be2b5e903b0ced41a74d647538d77d517cfdc5b22108f

Observation 2359f70b-b22d-4b3b-af6b-a86545c8b2d0 · outbound

This paper cites From pixels to ui ac- tions: Learning to follow instructions via graphical user in- terfaces.

ShowUI: One Vision-Language-Action Model for GUI Visual Agent From pixels to ui ac- tions: Learning to follow instructions via graphical user in- terfaces

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:10:41.405170Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T12:10:40.488992Z digest=sha256:2f299f8a5c12aaa883fd630c49de09843de6cf660456742b9a6b38dd377bc08e

Observation 4c417f3e-9090-4070-9fe7-bec873bea0b5 · outbound

This paper cites World of bits: An open-domain plat- form for web-based agents.

ShowUI: One Vision-Language-Action Model for GUI Visual Agent World of bits: An open-domain plat- form for web-based agents

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:10:41.390536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T12:10:40.493778Z digest=sha256:010e0c5746206f4017603b1072de303b223590343b490d65251de131e7d9919a

Observation 5dbef036-307a-4762-855d-2f8f14e32864 · outbound

This paper cites Qwen2-vl: Enhancing vision-language model’s perception of the world at any resolution, 2024.

ShowUI: One Vision-Language-Action Model for GUI Visual Agent Qwen2-vl: Enhancing vision-language model’s perception of the world at any resolution, 2024

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:10:41.375667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T12:10:40.498748Z digest=sha256:2f4519552be7a37d22300e6fa75d0ddd2366a2bcfed3df78e51c55da7d8f0df3

Observation 8336ed07-9009-4f6b-a3b8-71520e6978b3 · outbound

This paper cites Omnijarvis: Unified vision-language-action to- kenization enables open-world instruction following agents, 2024.

ShowUI: One Vision-Language-Action Model for GUI Visual Agent Omnijarvis: Unified vision-language-action to- kenization enables open-world instruction following agents, 2024

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:10:41.359720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T12:10:40.503590Z digest=sha256:f0356972a121f74d6cddf44201d060b525d346cbaff621d8357e52431d615e88

Observation 0525132e-f123-486f-a338-5b0f6deedf7f · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large lan- guage models.

ShowUI: One Vision-Language-Action Model for GUI Visual Agent Chain-of-thought prompting elicits reasoning in large lan- guage models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-12T12:10:40.508250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:10:40.508250Z digest=sha256:bddb85bfe9029223792ec4749733de7f669d8b0ecfe6d9bc954db77e765caabb

Observation dd2ce77e-ec5e-424a-977b-ead6099935b3 · outbound

This paper cites Webui: A dataset for en- hancing visual ui understanding with web semantics.

ShowUI: One Vision-Language-Action Model for GUI Visual Agent Webui: A dataset for en- hancing visual ui understanding with web semantics

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:10:41.333569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T12:10:40.513328Z digest=sha256:b45e7bcc33d7a48061838aaab89490bf398e5945efa74d6ffb666a836f82748c

Observation 4837df9a-c0cc-477f-81bf-59677419f289 · outbound

This paper cites Videollm-mod: Efficient video-language streaming with mixture-of-depths vision computation.

ShowUI: One Vision-Language-Action Model for GUI Visual Agent Videollm-mod: Efficient video-language streaming with mixture-of-depths vision computation

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:10:41.318329Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T12:10:40.518365Z digest=sha256:f6d1aecddf93b331ce062d250d53ad930767c1ddc58e0661acf44e8db8d74f1d

Observation 802cad14-f3a6-40c7-ac1f-aeb857e351bd · outbound

This paper cites GPT-4V in Wonderland: Large Multimodal Models for Zero-Shot Smartphone GUI Navigation.

ShowUI: One Vision-Language-Action Model for GUI Visual Agent GPT-4V in Wonderland: Large Multimodal Models for Zero-Shot Smartphone GUI Navigation

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-12T12:10:40.522993Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:10:40.522993Z digest=sha256:352801f7f3d327499a32f81d209efe6495a4819ec725a3986a654d39adc8c770

Observation bc7cc6d7-8ee8-4c2a-a2a4-b73be0b2d9b4 · outbound

This paper cites Set-of-mark prompting unleashes ex- traordinary visual grounding in gpt-4v, 2023.

ShowUI: One Vision-Language-Action Model for GUI Visual Agent Set-of-mark prompting unleashes ex- traordinary visual grounding in gpt-4v, 2023

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:10:41.302655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T12:10:40.527953Z digest=sha256:33c9d3931d942e24bd51cd7d5d0e039b73918140535d9d8683a34bb3e1029ff3

Observation 5a5bb20b-7599-4fef-b58c-536ca036e1c1 · outbound

This paper cites Gpt4tools: Teaching large language model to use tools via self-instruction.

ShowUI: One Vision-Language-Action Model for GUI Visual Agent Gpt4tools: Teaching large language model to use tools via self-instruction

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:10:41.285802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T12:10:40.532717Z digest=sha256:9f4077ac1b8c62b0fd684717059460b5a7525cb7ca8ddf8ea3f92d58e01622e1

Observation 63e3b143-8fa3-49a3-a4cf-b7779c7eb18e · outbound

This paper cites React: Synergizing rea- soning and acting in language models.

ShowUI: One Vision-Language-Action Model for GUI Visual Agent React: Synergizing rea- soning and acting in language models

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:10:41.269792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T12:10:40.538019Z digest=sha256:87775d1db142f55da2bd36956ce34ec51fd79a7365e6f4ad02a067b81f6145a4

Observation 4f21432b-e5c7-4b46-809b-8134bc812366 · outbound

This paper cites Ferret-UI: Grounded Mobile UI Understanding with Multimodal LLMs.

ShowUI: One Vision-Language-Action Model for GUI Visual Agent Ferret-UI: Grounded Mobile UI Understanding with Multimodal LLMs

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-12T12:10:40.543012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:10:40.543012Z digest=sha256:7c7102ae2e58599d8f467d828ef51a431a70a1f971da7ffbc8451b34716d152f

Observation e6e3edac-4335-44a3-8dcb-0934636c89d1 · outbound

This paper cites So- cratic models: Composing zero-shot multimodal reasoning with language.

ShowUI: One Vision-Language-Action Model for GUI Visual Agent So- cratic models: Composing zero-shot multimodal reasoning with language

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:10:41.254618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T12:10:40.548260Z digest=sha256:f7b881367ff743a6d603f45285bf9c9dafc4830667648ecfcf003240342d8b71

Observation 1bbda601-87c9-46ac-b059-ed25d56d7e62 · outbound

This paper cites xLAM: A Family of Large Action Models to Empower AI Agent Systems.

ShowUI: One Vision-Language-Action Model for GUI Visual Agent xLAM: A Family of Large Action Models to Empower AI Agent Systems

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-12T12:10:40.552947Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:10:40.552947Z digest=sha256:7e81df3dce787505c19c2d63ff01324fe28f48b7a6f1c5ea6501f2e495afef5b

Observation 72b6d15f-8d44-480c-a07c-9ef99f016fea · outbound

This paper cites You only look at screens: Multimodal chain-of-action agents.

ShowUI: One Vision-Language-Action Model for GUI Visual Agent You only look at screens: Multimodal chain-of-action agents

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:10:41.238270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T12:10:40.558165Z digest=sha256:4533918983c0201ce3bdcf55d5e121df2878fbc104c31487d31b72781812e7ac

Observation 9c8a6d4a-033b-4043-b8b2-31194363a28f · outbound

This paper cites 3D-VLA: A 3D Vision-Language-Action Generative World Model.

ShowUI: One Vision-Language-Action Model for GUI Visual Agent 3D-VLA: A 3D Vision-Language-Action Generative World Model

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-12T12:10:40.564363Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:10:40.564363Z digest=sha256:a2cfaad7d0c63ed722504882d28f706d16af360ba031b781f3f1996bf9502f5d

Observation ae096c46-5a6a-4be8-ab23-fbc791c74ad0 · outbound

This paper cites Gpt-4v(ision) is a generalist web agent, if grounded.

ShowUI: One Vision-Language-Action Model for GUI Visual Agent Gpt-4v(ision) is a generalist web agent, if grounded

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:10:41.223010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T12:10:40.569129Z digest=sha256:55b5a8cf9f018f47e3a99ab109e4b7103011159cc5e371f336ce3aefafc90230

Observation 87da435c-a45e-4b4b-a759-1df2e0a9aef3 · outbound

This paper cites WebArena: A Realistic Web Environment for Building Autonomous Agents.

ShowUI: One Vision-Language-Action Model for GUI Visual Agent WebArena: A Realistic Web Environment for Building Autonomous Agents

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-12T12:10:40.573941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:10:40.573941Z digest=sha256:aabde71914141e4e085a9eed10659c9620ce6d9e1fa270ea1d5b60c2ac3615d4

Pith citing papers

Observation 5d7febd8-656a-443a-a059-5faad6a0eeab · inbound

Improved GUI Grounding via Iterative Narrowing cites this paper.

Improved GUI Grounding via Iterative Narrowing ShowUI: One Vision-Language-Action Model for GUI Visual Agent

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T18:44:24.161425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:44:24.161425Z digest=sha256:739c467a214e420e7070ee491f8768c2451a2813d8c3d596f035fdc41f22bbfa

Observation f94b606f-5219-435c-8c71-da418617373b · inbound

Large Language Model-Brained GUI Agents: A Survey cites this paper.

Large Language Model-Brained GUI Agents: A Survey ShowUI: One Vision-Language-Action Model for GUI Visual Agent

Reference 246

Resolution
verified exact
arxiv_id, observed 2026-05-19T11:08:27.838859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-19T11:08:27.472508Z digest=sha256:ebac1ce933d343bdb61a01b2895af4107897ce567b1354b37e792275657c3fd4

Observation 2890b80e-82f0-4366-88df-721edd3d8bd6 · inbound

WorldGUI: An Interactive Benchmark for Desktop GUI Automation from Any Starting Point cites this paper.

WorldGUI: An Interactive Benchmark for Desktop GUI Automation from Any Starting Point ShowUI: One Vision-Language-Action Model for GUI Visual Agent

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-08T11:05:06.126655Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T11:05:06.126655Z digest=sha256:9d26f5c7c0c4955fc27707529229746bf705d2cf7a07ec1763a626bb1bf3ef93

Observation 2311856d-2514-42ac-87a7-3d399d889f4b · inbound

LearnAct: Few-Shot Mobile GUI Agent with a Unified Demonstration Benchmark cites this paper.

LearnAct: Few-Shot Mobile GUI Agent with a Unified Demonstration Benchmark ShowUI: One Vision-Language-Action Model for GUI Visual Agent

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-16T12:03:32.149993Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:03:32.149993Z digest=sha256:debadcf4af21b433428ab311227cda487306523d4106a314f2bff3c3d31e2275

Observation 1ae04baf-0455-4b5a-bee0-d3446246659e · inbound

InfantAgent-Next: A Multimodal Generalist Agent for Automated Computer Interaction cites this paper.

InfantAgent-Next: A Multimodal Generalist Agent for Automated Computer Interaction ShowUI: One Vision-Language-Action Model for GUI Visual Agent

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-22T15:21:45.076189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-22T15:18:15.475294Z digest=sha256:2f348de8a049648b3bbc2ccc4b156045f31622e78a8eca8921a56ba8c2fc4c04

Observation 028484cb-2496-47b1-a36c-1d4c57e55cc7 · inbound

ReGUIDE: Data Efficient GUI Grounding via Spatial Reasoning and Search cites this paper.

ReGUIDE: Data Efficient GUI Grounding via Spatial Reasoning and Search ShowUI: One Vision-Language-Action Model for GUI Visual Agent

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T15:25:31.922315Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:25:31.922315Z digest=sha256:fa251c93f366f9af43bad36d5fe7be022ef12e5ceeabd1125521e876974e3bef

Observation d1892b3a-10fd-4e5a-8075-68b746c4343a · inbound

GUI-G1: Understanding R1-Zero-Like Training for Visual Grounding in GUI Agents cites this paper.

GUI-G1: Understanding R1-Zero-Like Training for Visual Grounding in GUI Agents ShowUI: One Vision-Language-Action Model for GUI Visual Agent

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T15:16:02.648058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:16:02.648058Z digest=sha256:43523a3af43007d7e4ecc52bbf1ffdb5216cf58aae487dcd489164694af6284e

Observation 82470402-8cf8-40f0-8292-4cdc19026485 · inbound

BacktrackAgent: Enhancing GUI Agent with Error Detection and Backtracking Mechanism cites this paper.

BacktrackAgent: Enhancing GUI Agent with Error Detection and Backtracking Mechanism ShowUI: One Vision-Language-Action Model for GUI Visual Agent

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:39.067010Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:39.067010Z digest=sha256:40ed101c8f7ebfaac1ca9fd8d55b3d05b9affd332b2f7c2e3749761c403393d0

Observation 01f5e6b7-1193-492b-b807-c30c44dc20db · inbound

Grounded Reinforcement Learning for Visual Reasoning cites this paper.

Grounded Reinforcement Learning for Visual Reasoning ShowUI: One Vision-Language-Action Model for GUI Visual Agent

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-22T01:05:51.951989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-22T01:05:18.801388Z digest=sha256:3ea8255230403c75064a4a545f39b3828990bd02243beb83723598a7c8011bf7

Observation 4fe73d73-f7ee-41c9-94e3-c0534fbcf123 · inbound

ZeroGUI: Automating Online GUI Learning at Zero Human Cost cites this paper.

ZeroGUI: Automating Online GUI Learning at Zero Human Cost ShowUI: One Vision-Language-Action Model for GUI Visual Agent

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T12:41:22.745953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:41:22.745953Z digest=sha256:049e6b8b2e8073d45fa28304eb07a1c924b7b437fd3606f09e62b569152fbb66

Observation 038741e4-74ad-4c5c-9a4a-56f7f7ac5b94 · inbound

RiOSWorld: Benchmarking the Risk of Multimodal Computer-Use Agents cites this paper.

RiOSWorld: Benchmarking the Risk of Multimodal Computer-Use Agents ShowUI: One Vision-Language-Action Model for GUI Visual Agent

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T12:06:24.763166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:06:24.763166Z digest=sha256:0f5a92ed6b79fd6a462cac0505c67ece2216dbdc6595fb1153ceede6bf1b0193

Observation 19d78170-56f1-40a2-b7d3-5490c874e824 · inbound

GUI-Actor: Coordinate-Free Visual Grounding for GUI Agents cites this paper.

GUI-Actor: Coordinate-Free Visual Grounding for GUI Agents ShowUI: One Vision-Language-Action Model for GUI Visual Agent

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T11:14:51.672228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:14:51.672228Z digest=sha256:ba7464f724be56005066443df666d0b0b324f1060a66657e0e9618031f643b99

Observation 83b75b66-3873-4d89-8e09-65c9ab7b1b8b · inbound

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities cites this paper.

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities ShowUI: One Vision-Language-Action Model for GUI Visual Agent

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T05:02:29.221294Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:02:29.221294Z digest=sha256:8b2fd707f880e7a419f200de076ffce9f8ab815e6486e95add41d8122c7d0632

Observation 604549a2-8239-4c56-9bb1-6062220719f0 · inbound

Atomic-to-Compositional Generalization for Mobile Agents with A New Benchmark and Scheduling System cites this paper.

Atomic-to-Compositional Generalization for Mobile Agents with A New Benchmark and Scheduling System ShowUI: One Vision-Language-Action Model for GUI Visual Agent

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T05:04:04.513377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:04:04.513377Z digest=sha256:9842a653aecac1ea09b0f31b84e21848f6e854bef6f3dc57b45bb3f108cbb3a9

Observation a9b674cc-0ab6-4c3a-8888-a49c58b29e6f · inbound

GUI-Robust: A Comprehensive Dataset for Testing GUI Agent Robustness in Real-World Anomalies cites this paper.

GUI-Robust: A Comprehensive Dataset for Testing GUI Agent Robustness in Real-World Anomalies ShowUI: One Vision-Language-Action Model for GUI Visual Agent

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T00:22:52.869086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:22:52.869086Z digest=sha256:4736b096e84e9aeed3c1c6c5c3484fd0392eb9014299cc40278991dada667d8b

Observation ed1bf49c-f7f2-4d16-b271-6706a3f5e214 · inbound

Understanding GUI Agent Localization Biases through Logit Sharpness cites this paper.

Understanding GUI Agent Localization Biases through Logit Sharpness ShowUI: One Vision-Language-Action Model for GUI Visual Agent

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T19:39:59.717577Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:39:59.717577Z digest=sha256:08fb2b6a4a0e3be4c9d77987c488e0d1277e06f22473d11829d580ce7664f99b

Observation 7f4c0137-4cdd-40bf-8b5b-2e0533244c88 · inbound

Learning, Reasoning, Refinement: A Framework for Kahneman's Dual-System Intelligence in GUI Agents cites this paper.

Learning, Reasoning, Refinement: A Framework for Kahneman's Dual-System Intelligence in GUI Agents ShowUI: One Vision-Language-Action Model for GUI Visual Agent

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T19:05:16.074321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:05:16.074321Z digest=sha256:1e14666e9febf415dc6f1c29590d32a2f29c1bf17c7af3b6e433bbae44c74454

Observation 44f450de-7495-43d0-9caa-5f74dfb736c0 · inbound

ZonUI-3B: A Lightweight Vision-Language Model for Cross-Resolution GUI Grounding cites this paper.

ZonUI-3B: A Lightweight Vision-Language Model for Cross-Resolution GUI Grounding ShowUI: One Vision-Language-Action Model for GUI Visual Agent

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T21:44:51.512452Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:44:51.512452Z digest=sha256:fc0451e6fa2968513f7ebaed44cf27dedac8ccbef38b0e94fabe01cd67a61c5a

Observation 2fd21acb-63aa-437f-97d1-f4854ec9f7cb · inbound

PresentAgent: Multimodal Agent for Presentation Video Generation cites this paper.

PresentAgent: Multimodal Agent for Presentation Video Generation ShowUI: One Vision-Language-Action Model for GUI Visual Agent

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T20:01:29.076779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:01:29.076779Z digest=sha256:c258e07bf52464f1191112de3b8ede5ea7100be64696b5ec965d0091990ac9a8

Observation f1cc3105-4402-434e-b218-a93148939593 · inbound

DreamVLA: A Vision-Language-Action Model Dreamed with Comprehensive World Knowledge cites this paper.

DreamVLA: A Vision-Language-Action Model Dreamed with Comprehensive World Knowledge ShowUI: One Vision-Language-Action Model for GUI Visual Agent

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-16T15:42:41.472571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-16T15:42:41.363422Z digest=sha256:958ee23c95b01ab7a9e6e633dcbbbc6febe2570bf1be081656bd3364688c8323

Observation 502ea755-018e-4996-ba99-3175a3291e2c · inbound

GUI-G$^2$: Gaussian Reward Modeling for GUI Grounding cites this paper.

GUI-G$^2$: Gaussian Reward Modeling for GUI Grounding ShowUI: One Vision-Language-Action Model for GUI Visual Agent

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T15:29:11.808864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:29:11.808864Z digest=sha256:40bb27df7fdf43d4c3b13e35b0d58d97c46364ffa190a8e520267e9cae695b30

Observation fb3aec4d-05fc-4eab-ba62-092c7b43bebb · inbound

Screen2AX: Vision-Based Approach for Automatic macOS Accessibility Generation cites this paper.

Screen2AX: Vision-Based Approach for Automatic macOS Accessibility Generation ShowUI: One Vision-Language-Action Model for GUI Visual Agent

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T15:08:35.653738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:08:35.653738Z digest=sha256:78ea3ad4b98349ca7615c26a3d22695049f91a8ea22b3a7075a569909f61edd0

Observation 1faa229f-368a-4280-a1e4-b9a998812fdc · inbound

MMBench-GUI: Hierarchical Multi-Platform Evaluation Framework for GUI Agents cites this paper.

MMBench-GUI: Hierarchical Multi-Platform Evaluation Framework for GUI Agents ShowUI: One Vision-Language-Action Model for GUI Visual Agent

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T14:25:32.876994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:25:32.876994Z digest=sha256:d83d718e585f435abdab49d6f7b3ac3b37d772c6023f38f94d015b3381cb3458

Observation cc18a95a-d7f0-4b16-a168-4c714c8654e2 · inbound

SEAgent: Self-Evolving Computer Use Agent with Autonomous Learning from Experience cites this paper.

SEAgent: Self-Evolving Computer Use Agent with Autonomous Learning from Experience ShowUI: One Vision-Language-Action Model for GUI Visual Agent

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-05T23:55:49.349932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:55:49.349932Z digest=sha256:feb1e871409a1a05b79bd0ccf028aaad8117cb08bbe744103184c189b05ec2f7

Observation 96710486-0104-4b5f-8a91-a545d55cb11b · inbound

CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning cites this paper.

CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning ShowUI: One Vision-Language-Action Model for GUI Visual Agent

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-05T15:19:50.672939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:19:50.672939Z digest=sha256:9da05574477567bfa55ec75318de63625b2fa327e32f6b1d93a4a0ecc6561e46

Observation f93935eb-d7b6-4d83-823a-06fb7832e71c · inbound

MobiAgent: A Systematic Framework for Customizable Mobile Agents cites this paper.

MobiAgent: A Systematic Framework for Customizable Mobile Agents ShowUI: One Vision-Language-Action Model for GUI Visual Agent

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-05T13:35:00.146538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:35:00.146538Z digest=sha256:dd84723055c1c560dd1f65b710ba062f4d53b842ccd5ebcb5abd3d0f9ca4f274

Observation 3619084a-7ace-4ad1-92f4-9831062dc879 · inbound

PG-Agent: An Agent Powered by Page Graph cites this paper.

PG-Agent: An Agent Powered by Page Graph ShowUI: One Vision-Language-Action Model for GUI Visual Agent

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-05T15:29:42.495412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:29:42.495412Z digest=sha256:7fb40bdb6a583e64fbb245efa26f8f6b11cf498ed475347bbe83f334b856c205

Observation 4573b04a-c972-4ab2-b201-0eee654c3d57 · inbound

Mitigating Coordinate Prediction Bias from Positional Encoding Failures cites this paper.

Mitigating Coordinate Prediction Bias from Positional Encoding Failures ShowUI: One Vision-Language-Action Model for GUI Visual Agent

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-18T05:12:23.537619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-18T05:11:30.734633Z digest=sha256:27e7c03a5707dc95971fae0e34955d50400a0ea9ab9d94f19bec2fead26f6489

Observation c09c03e6-46bb-4a3a-b492-3c3eccfe00d0 · inbound

GUI-AIMA: Aligning Intrinsic Multimodal Attention with a Context Anchor for GUI Grounding cites this paper.

GUI-AIMA: Aligning Intrinsic Multimodal Attention with a Context Anchor for GUI Grounding ShowUI: One Vision-Language-Action Model for GUI Visual Agent

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T00:30:24.114175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:30:24.114175Z digest=sha256:2208463cfb8ab772712290bb794f9e2b6102eb605c3a0ed8edd05a23bc412c83

Observation 8a3e9fe5-f353-494a-8a6d-e3914f42405a · inbound

Grounding Computer Use Agents on Human Demonstrations cites this paper.

Grounding Computer Use Agents on Human Demonstrations ShowUI: One Vision-Language-Action Model for GUI Visual Agent

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-03T23:06:04.903445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:06:04.903445Z digest=sha256:60a677358dc87e2d81a75154a8c3893dee0deb44ffaa3acc47dd47d1757082b0

Observation 37e5ef16-a407-4fb9-bb37-3c778e521058 · inbound

GUI-Libra: Training Native GUI Agents to Reason and Act with Action-aware Supervision and Partially Verifiable RL cites this paper.

GUI-Libra: Training Native GUI Agents to Reason and Act with Action-aware Supervision and Partially Verifiable RL ShowUI: One Vision-Language-Action Model for GUI Visual Agent

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-02T20:51:42.283887Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T20:51:42.283887Z digest=sha256:6c3ba5ad051077d4db705dbea562e34a7f54c0fa82e1f399765a03858f6a4de6

Observation d346d563-3f51-40b0-85a8-67c4c7a9f2db · inbound

GUIDE: Resolving Domain Bias in GUI Agents through Real-Time Web Video Retrieval and Plug-and-Play Annotation cites this paper.

GUIDE: Resolving Domain Bias in GUI Agents through Real-Time Web Video Retrieval and Plug-and-Play Annotation ShowUI: One Vision-Language-Action Model for GUI Visual Agent

Reference 15

Resolution
unresolved
no resolver link, observed 2026-07-13T17:40:03.471339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T17:40:03.471339Z digest=sha256:068e90f308b272680fa89f7635c3c0ba38cd27a59bbc9f02288cf4c9a4998759

Observation d261bd46-3e49-43de-841f-377eba841c26 · inbound

UI-Zoomer: Uncertainty-Driven Adaptive Zoom-In for GUI Grounding cites this paper.

UI-Zoomer: Uncertainty-Driven Adaptive Zoom-In for GUI Grounding ShowUI: One Vision-Language-Action Model for GUI Visual Agent

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-10T14:10:28.917044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-10T14:06:55.472857Z digest=sha256:9a11020c326df1ccf7990ead12bea9835ce7205d453d5ff529276748e2ee500e

Observation 5a0e84a0-0dd1-4ce2-a42e-d12a6024b301 · inbound

VLAA-GUI: Knowing When to Stop, Recover, and Search, A Modular Framework for GUI Automation cites this paper.

VLAA-GUI: Knowing When to Stop, Recover, and Search, A Modular Framework for GUI Automation ShowUI: One Vision-Language-Action Model for GUI Visual Agent

Reference 40

Resolution
metadata mismatch
arxiv_id, observed 2026-05-09T22:34:07.779203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-09T22:24:45.045405Z digest=sha256:e9dd92077ab2718f186b83467add21f38bc55d56d563751c778a0f86378e09b8

Observation bcb010b6-2ea1-43bd-a68e-3974373b6513 · inbound

Do Vision-Language-Models show human-like logical problem-solving capability in point and click puzzle games? cites this paper.

Do Vision-Language-Models show human-like logical problem-solving capability in point and click puzzle games? ShowUI: One Vision-Language-Action Model for GUI Visual Agent

Reference 21

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T02:02:05.445608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-13T01:58:39.476408Z digest=sha256:8225ba3f843a9a593ac84e4d44d9ed78c699f1d1f610c304ff7472e31a475648

Observation 1bd235fc-16b4-42df-99dd-ddaf0219dbca · inbound

Skim: Speculative Execution for Fast and Efficient Web Agents cites this paper.

Skim: Speculative Execution for Fast and Efficient Web Agents ShowUI: One Vision-Language-Action Model for GUI Visual Agent

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-20T18:33:38.039247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T18:31:18.061863Z digest=sha256:904714ef5c55d0682c522a108bf7a3899c654e110524fd32182f3eadf1d6e43a

Observation 7cd99c0c-668e-45a6-a8d9-0625225c3568 · inbound

GUI-C$^2$: Coarse-to-Fine GUI Grounding via Difficulty-Aware Reinforcement Learning cites this paper.

GUI-C$^2$: Coarse-to-Fine GUI Grounding via Difficulty-Aware Reinforcement Learning ShowUI: One Vision-Language-Action Model for GUI Visual Agent

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-06-28T22:52:44.698903Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-06-28T22:52:04.755523Z digest=sha256:85b0a03e3745befd7c1b1e38367f6f153ec5f3f7d0f031cac578f92f218019f2

Observation 45d438a7-4e55-43d2-8f87-3af3cc421c45 · inbound

OSWorld 2.0: Benchmarking Computer Use Agents on Long-Horizon Real-World Tasks cites this paper.

OSWorld 2.0: Benchmarking Computer Use Agents on Long-Horizon Real-World Tasks ShowUI: One Vision-Language-Action Model for GUI Visual Agent

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-06-30T07:14:21.126877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-30T07:10:38.909339Z digest=sha256:126ecc3e3eb3a78865087c153fad853dbc97acf9fb67cb2db10ac7296bfb51d4

Observation 33244b6a-53c8-4fce-b7e7-641a23df1245 · inbound

OSWorld 2.0: Benchmarking Computer Use Agents on Long-Horizon Real-World Tasks cites this paper.

OSWorld 2.0: Benchmarking Computer Use Agents on Long-Horizon Real-World Tasks ShowUI: One Vision-Language-Action Model for GUI Visual Agent

Reference 52

Resolution
unresolved
no resolver link, observed 2026-07-15T10:24:53.345620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T10:24:53.345620Z digest=sha256:24f0dc1546736faeac27c113987b735adab98e5ac409d1bdc5853543221f8aeb

Observation cd1725ad-d3e7-4fe6-bb89-4f6507c86a23 · inbound

Do GUI Agents Believe Their Eyes? Diagnosing State-Belief Reliance on Pixels versus Structure cites this paper.

Do GUI Agents Believe Their Eyes? Diagnosing State-Belief Reliance on Pixels versus Structure ShowUI: One Vision-Language-Action Model for GUI Visual Agent

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-11T20:00:19.123214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T20:00:19.123214Z digest=sha256:86466592b737031498a61fd4245a5fe8135b500b22c96bd1733f55e1b408a56e

Observation 6ee9538c-55f7-44f6-a56c-cf5914914214 · inbound

Do GUI Agents Believe Their Eyes? Diagnosing State-Belief Reliance on Pixels versus Structure cites this paper.

Do GUI Agents Believe Their Eyes? Diagnosing State-Belief Reliance on Pixels versus Structure ShowUI: One Vision-Language-Action Model for GUI Visual Agent

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-02T08:46:59.431501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:46:59.431501Z digest=sha256:7ea925b822d7e905bbcdd5bcce4923b6509793f6f1e29b7e979e99cf02ba5cdc

Observation 9a08a3d2-bf56-40cb-a8e0-2c9632d1a57e · inbound

Vision as Unified Multimodal Generation cites this paper.

Vision as Unified Multimodal Generation ShowUI: One Vision-Language-Action Model for GUI Visual Agent

Reference 108

Resolution
verified exact
local_arxiv, observed 2026-07-08T02:04:26.431023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-08T01:54:30.649092Z digest=sha256:177582f8bca4cd4240d702b65bcc51f57035e40d6342ee4ab1d23b8b3c8e0568

Observation db8235ac-0579-45c3-84ea-dcf60d4b40ca · inbound

StepX-Edge: An On-Device UI Vision-Language Model via Architecture-Training-Deployment Co-Design cites this paper.

StepX-Edge: An On-Device UI Vision-Language Model via Architecture-Training-Deployment Co-Design ShowUI: One Vision-Language-Action Model for GUI Visual Agent

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-01T16:31:11.287336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:31:11.287336Z digest=sha256:ca7be4dada5637b50315be49c0015093d6a4acd7fdbd32b0b4b8f4f6842b2c17

Observation ba010f80-452c-400c-b9ce-9af0e7cc0593 · inbound

Rethinking Inference-Time Scaling in Local Computer-Use Agents: Failure Modes and Compute Tradeoffs cites this paper.

Rethinking Inference-Time Scaling in Local Computer-Use Agents: Failure Modes and Compute Tradeoffs ShowUI: One Vision-Language-Action Model for GUI Visual Agent

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-31T03:24:11.977217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T03:24:11.977217Z digest=sha256:d3846fcc05e1a3a48a82c53ce7cd4d9d82d2397c6ba38d3b1c58bb4d8eb87472