Pith. sign in

Paper Citation Record · LEDGER

ShowUI: One Vision-Language-Action Model for GUI Visual Agent

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 40 inbound Pith citation observations for arXiv:2411.17465.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.17465 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 40 of 40 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 40 of 40 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T11:05:06.126655Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

1
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation f94b606f-5219-435c-8c71-da418617373b · inbound

Large Language Model-Brained GUI Agents: A Survey cites this paper.

Large Language Model-Brained GUI Agents: A Survey ShowUI: One Vision-Language-Action Model for GUI Visual Agent

Reference 246

Resolution
verified exact
arxiv_id, observed 2026-05-19T11:08:27.838859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-19T11:08:27.472508Z digest=sha256:4c659d92d7ed6b53ecd364b8ccc3406a2232a557aaf02f60e616148dcd142ade

Observation 2890b80e-82f0-4366-88df-721edd3d8bd6 · inbound

WorldGUI: An Interactive Benchmark for Desktop GUI Automation from Any Starting Point cites this paper.

WorldGUI: An Interactive Benchmark for Desktop GUI Automation from Any Starting Point ShowUI: One Vision-Language-Action Model for GUI Visual Agent

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-08T11:05:06.126655Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T11:05:06.126655Z digest=sha256:391f7520516c86ac5a4685b0a533ae3e9fca70b8b016a06b1bad75fde5393d81

Observation 1ae04baf-0455-4b5a-bee0-d3446246659e · inbound

InfantAgent-Next: A Multimodal Generalist Agent for Automated Computer Interaction cites this paper.

InfantAgent-Next: A Multimodal Generalist Agent for Automated Computer Interaction ShowUI: One Vision-Language-Action Model for GUI Visual Agent

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-22T15:21:45.076189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-22T15:18:15.475294Z digest=sha256:07951f85957bb09cce84ce1dcd9e33cb7429f1924379d827d5160ffc4717fad0

Observation 028484cb-2496-47b1-a36c-1d4c57e55cc7 · inbound

ReGUIDE: Data Efficient GUI Grounding via Spatial Reasoning and Search cites this paper.

ReGUIDE: Data Efficient GUI Grounding via Spatial Reasoning and Search ShowUI: One Vision-Language-Action Model for GUI Visual Agent

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T15:25:31.922315Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:25:31.922315Z digest=sha256:99e9582088ac6b0ba90c0c2e2dda5d51fe63f5cdb6573597727f9fc18c262d2f

Observation d1892b3a-10fd-4e5a-8075-68b746c4343a · inbound

GUI-G1: Understanding R1-Zero-Like Training for Visual Grounding in GUI Agents cites this paper.

GUI-G1: Understanding R1-Zero-Like Training for Visual Grounding in GUI Agents ShowUI: One Vision-Language-Action Model for GUI Visual Agent

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T15:16:02.648058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:16:02.648058Z digest=sha256:701acecdc1ec4378beb0ff241f54f18fc8c833f3053f525244ea28d80fc7b299

Observation 82470402-8cf8-40f0-8292-4cdc19026485 · inbound

BacktrackAgent: Enhancing GUI Agent with Error Detection and Backtracking Mechanism cites this paper.

BacktrackAgent: Enhancing GUI Agent with Error Detection and Backtracking Mechanism ShowUI: One Vision-Language-Action Model for GUI Visual Agent

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:39.067010Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:39.067010Z digest=sha256:8d18c935c79d23479370c55e11f0fc32d085ae08f9b5749ab3f608483da4f10e

Observation 01f5e6b7-1193-492b-b807-c30c44dc20db · inbound

Grounded Reinforcement Learning for Visual Reasoning cites this paper.

Grounded Reinforcement Learning for Visual Reasoning ShowUI: One Vision-Language-Action Model for GUI Visual Agent

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-22T01:05:51.951989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-22T01:05:18.801388Z digest=sha256:193e5b082cb09436131804e048d966c5dfd5707f7659e04824df8e05a60080e1

Observation 4fe73d73-f7ee-41c9-94e3-c0534fbcf123 · inbound

ZeroGUI: Automating Online GUI Learning at Zero Human Cost cites this paper.

ZeroGUI: Automating Online GUI Learning at Zero Human Cost ShowUI: One Vision-Language-Action Model for GUI Visual Agent

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T12:41:22.745953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:41:22.745953Z digest=sha256:9d51013e5da2f994489d4f86f82c16d257a524db94de94c66809a97958a55394

Observation 038741e4-74ad-4c5c-9a4a-56f7f7ac5b94 · inbound

RiOSWorld: Benchmarking the Risk of Multimodal Computer-Use Agents cites this paper.

RiOSWorld: Benchmarking the Risk of Multimodal Computer-Use Agents ShowUI: One Vision-Language-Action Model for GUI Visual Agent

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T12:06:24.763166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:06:24.763166Z digest=sha256:770cb229b176a010fc3998f0a172fe69b94cf7fd480ac9786c117d371117140a

Observation 19d78170-56f1-40a2-b7d3-5490c874e824 · inbound

GUI-Actor: Coordinate-Free Visual Grounding for GUI Agents cites this paper.

GUI-Actor: Coordinate-Free Visual Grounding for GUI Agents ShowUI: One Vision-Language-Action Model for GUI Visual Agent

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T11:14:51.672228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:14:51.672228Z digest=sha256:4efaeaf9407aec9221d301e0733bf18b911b5d28540e004bec5184b0b9edf0a9

Observation 83b75b66-3873-4d89-8e09-65c9ab7b1b8b · inbound

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities cites this paper.

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities ShowUI: One Vision-Language-Action Model for GUI Visual Agent

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T05:02:29.221294Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:02:29.221294Z digest=sha256:3cfd2f2f3f89c3334f4ca57fb4f10368701ba4f31f77412b4c8291a68f192ce9

Observation 604549a2-8239-4c56-9bb1-6062220719f0 · inbound

Atomic-to-Compositional Generalization for Mobile Agents with A New Benchmark and Scheduling System cites this paper.

Atomic-to-Compositional Generalization for Mobile Agents with A New Benchmark and Scheduling System ShowUI: One Vision-Language-Action Model for GUI Visual Agent

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T05:04:04.513377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:04:04.513377Z digest=sha256:599bf48590d073f44123312d644a09d28c131796770ecb52d0a702c51e1e3cbb

Observation a9b674cc-0ab6-4c3a-8888-a49c58b29e6f · inbound

GUI-Robust: A Comprehensive Dataset for Testing GUI Agent Robustness in Real-World Anomalies cites this paper.

GUI-Robust: A Comprehensive Dataset for Testing GUI Agent Robustness in Real-World Anomalies ShowUI: One Vision-Language-Action Model for GUI Visual Agent

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T00:22:52.869086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:22:52.869086Z digest=sha256:00b064b27cfaaba861e4cb4d24d15ad584e6961053716e7d8f0bbfbef75a7b08

Observation 44f450de-7495-43d0-9caa-5f74dfb736c0 · inbound

ZonUI-3B: A Lightweight Vision-Language Model for Cross-Resolution GUI Grounding cites this paper.

ZonUI-3B: A Lightweight Vision-Language Model for Cross-Resolution GUI Grounding ShowUI: One Vision-Language-Action Model for GUI Visual Agent

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T21:44:51.512452Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:44:51.512452Z digest=sha256:a9242aab28ef846bebaa23eab60c430dbec7442ef1643d537ff29ea40e2b9da2

Observation 2fd21acb-63aa-437f-97d1-f4854ec9f7cb · inbound

PresentAgent: Multimodal Agent for Presentation Video Generation cites this paper.

PresentAgent: Multimodal Agent for Presentation Video Generation ShowUI: One Vision-Language-Action Model for GUI Visual Agent

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T20:01:29.076779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:01:29.076779Z digest=sha256:af2d1a7e86dc34ceeb5bb395f8f7dc8cd96ac4b1bfe995d0e7aa8d057a44242d

Observation f1cc3105-4402-434e-b218-a93148939593 · inbound

DreamVLA: A Vision-Language-Action Model Dreamed with Comprehensive World Knowledge cites this paper.

DreamVLA: A Vision-Language-Action Model Dreamed with Comprehensive World Knowledge ShowUI: One Vision-Language-Action Model for GUI Visual Agent

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-16T15:42:41.472571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T15:42:41.363422Z digest=sha256:c2a62186bf4aa90c10b5b97d50e20ff054519db2950da4d572a7625e0d238a7f

Observation 502ea755-018e-4996-ba99-3175a3291e2c · inbound

GUI-G$^2$: Gaussian Reward Modeling for GUI Grounding cites this paper.

GUI-G$^2$: Gaussian Reward Modeling for GUI Grounding ShowUI: One Vision-Language-Action Model for GUI Visual Agent

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T15:29:11.808864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:29:11.808864Z digest=sha256:9d4062012eda7af77426301454d4844bf60fdfe4d5832d4bfdafb39901e8da61

Observation fb3aec4d-05fc-4eab-ba62-092c7b43bebb · inbound

Screen2AX: Vision-Based Approach for Automatic macOS Accessibility Generation cites this paper.

Screen2AX: Vision-Based Approach for Automatic macOS Accessibility Generation ShowUI: One Vision-Language-Action Model for GUI Visual Agent

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T15:08:35.653738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:08:35.653738Z digest=sha256:3c14bd17459054fd38566d0d4f479f73a8e4419215b85f3db55a7da56cb60609

Observation 1faa229f-368a-4280-a1e4-b9a998812fdc · inbound

MMBench-GUI: Hierarchical Multi-Platform Evaluation Framework for GUI Agents cites this paper.

MMBench-GUI: Hierarchical Multi-Platform Evaluation Framework for GUI Agents ShowUI: One Vision-Language-Action Model for GUI Visual Agent

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T14:25:32.876994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:25:32.876994Z digest=sha256:6065bb77c50f3569ab48af182bc5a8b62114be4d8f90bd30b860763b16c65c8a

Observation cc18a95a-d7f0-4b16-a168-4c714c8654e2 · inbound

SEAgent: Self-Evolving Computer Use Agent with Autonomous Learning from Experience cites this paper.

SEAgent: Self-Evolving Computer Use Agent with Autonomous Learning from Experience ShowUI: One Vision-Language-Action Model for GUI Visual Agent

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-05T23:55:49.349932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:55:49.349932Z digest=sha256:dbc529129c15f08b0d295e03a542fc0c6392dec8663ca613946a1ea16678cb8f

Observation 96710486-0104-4b5f-8a91-a545d55cb11b · inbound

CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning cites this paper.

CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning ShowUI: One Vision-Language-Action Model for GUI Visual Agent

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-05T15:19:50.672939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:19:50.672939Z digest=sha256:043d90c0fb71b60df58369839a502e91ba97df3f96f4a34920c32a2afa2d5906

Observation f93935eb-d7b6-4d83-823a-06fb7832e71c · inbound

MobiAgent: A Systematic Framework for Customizable Mobile Agents cites this paper.

MobiAgent: A Systematic Framework for Customizable Mobile Agents ShowUI: One Vision-Language-Action Model for GUI Visual Agent

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-05T13:35:00.146538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:35:00.146538Z digest=sha256:ff395a961f18c24966fa6d2f8d69adf59cdfddbfd00df258e6e2da9ad6954116

Observation 3619084a-7ace-4ad1-92f4-9831062dc879 · inbound

PG-Agent: An Agent Powered by Page Graph cites this paper.

PG-Agent: An Agent Powered by Page Graph ShowUI: One Vision-Language-Action Model for GUI Visual Agent

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-05T15:29:42.495412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:29:42.495412Z digest=sha256:6d5d3286e18176ee10fe3966d3ef4acbde467349154f1ca56e0abf4e5f63d4ea

Observation 4573b04a-c972-4ab2-b201-0eee654c3d57 · inbound

Mitigating Coordinate Prediction Bias from Positional Encoding Failures cites this paper.

Mitigating Coordinate Prediction Bias from Positional Encoding Failures ShowUI: One Vision-Language-Action Model for GUI Visual Agent

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-18T05:12:23.537619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-18T05:11:30.734633Z digest=sha256:806d4865c5f8c5bf7f6afb683016b48248a666acde21a8c57b98f6efd6d613ab

Observation c09c03e6-46bb-4a3a-b492-3c3eccfe00d0 · inbound

GUI-AIMA: Aligning Intrinsic Multimodal Attention with a Context Anchor for GUI Grounding cites this paper.

GUI-AIMA: Aligning Intrinsic Multimodal Attention with a Context Anchor for GUI Grounding ShowUI: One Vision-Language-Action Model for GUI Visual Agent

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T00:30:24.114175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:30:24.114175Z digest=sha256:526ca281764613402a453a82c54d77b58f2fe0cc185ed6aeaa5f321a7179866f

Observation 8a3e9fe5-f353-494a-8a6d-e3914f42405a · inbound

Grounding Computer Use Agents on Human Demonstrations cites this paper.

Grounding Computer Use Agents on Human Demonstrations ShowUI: One Vision-Language-Action Model for GUI Visual Agent

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-03T23:06:04.903445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:06:04.903445Z digest=sha256:30fc0bae1b18178e75b8c7dcfc7434b18e5d9bf2693bfb292bd6714a6293fc24

Observation 37e5ef16-a407-4fb9-bb37-3c778e521058 · inbound

GUI-Libra: Training Native GUI Agents to Reason and Act with Action-aware Supervision and Partially Verifiable RL cites this paper.

GUI-Libra: Training Native GUI Agents to Reason and Act with Action-aware Supervision and Partially Verifiable RL ShowUI: One Vision-Language-Action Model for GUI Visual Agent

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-02T20:51:42.283887Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T20:51:42.283887Z digest=sha256:ceb79b88d9a47e8ac6bc0af024850f690e39c53eef0fd16b4aba29002fc2f694

Observation d346d563-3f51-40b0-85a8-67c4c7a9f2db · inbound

GUIDE: Resolving Domain Bias in GUI Agents through Real-Time Web Video Retrieval and Plug-and-Play Annotation cites this paper.

GUIDE: Resolving Domain Bias in GUI Agents through Real-Time Web Video Retrieval and Plug-and-Play Annotation ShowUI: One Vision-Language-Action Model for GUI Visual Agent

Reference 15

Resolution
unresolved
no resolver link, observed 2026-07-13T17:40:03.471339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T17:40:03.471339Z digest=sha256:e1e474aebf0cb79eaef70c2a7342bbcc00727e0739a8c441aadc9af61ff94c4c

Observation d261bd46-3e49-43de-841f-377eba841c26 · inbound

UI-Zoomer: Uncertainty-Driven Adaptive Zoom-In for GUI Grounding cites this paper.

UI-Zoomer: Uncertainty-Driven Adaptive Zoom-In for GUI Grounding ShowUI: One Vision-Language-Action Model for GUI Visual Agent

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-10T14:10:28.917044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T14:06:55.472857Z digest=sha256:d17f4ba50a2e352403b11c990551b32361ed4af678a5ae97b5ee8a0148323be9

Observation 5a0e84a0-0dd1-4ce2-a42e-d12a6024b301 · inbound

VLAA-GUI: Knowing When to Stop, Recover, and Search, A Modular Framework for GUI Automation cites this paper.

VLAA-GUI: Knowing When to Stop, Recover, and Search, A Modular Framework for GUI Automation ShowUI: One Vision-Language-Action Model for GUI Visual Agent

Reference 40

Resolution
metadata mismatch
arxiv_id, observed 2026-05-09T22:34:07.779203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-09T22:24:45.045405Z digest=sha256:97e1dcf9301e471863c466dfb63c94ec2d432ec68c0e2e4001e28a7dad0e7717

Observation bcb010b6-2ea1-43bd-a68e-3974373b6513 · inbound

Do Vision-Language-Models show human-like logical problem-solving capability in point and click puzzle games? cites this paper.

Do Vision-Language-Models show human-like logical problem-solving capability in point and click puzzle games? ShowUI: One Vision-Language-Action Model for GUI Visual Agent

Reference 21

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T02:02:05.445608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-13T01:58:39.476408Z digest=sha256:c582f0a2d2140aeaf6744ff29dd1478fbbbbd245a7495e6ccb67ba29ca25119b

Observation 1bd235fc-16b4-42df-99dd-ddaf0219dbca · inbound

Skim: Speculative Execution for Fast and Efficient Web Agents cites this paper.

Skim: Speculative Execution for Fast and Efficient Web Agents ShowUI: One Vision-Language-Action Model for GUI Visual Agent

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-20T18:33:38.039247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-20T18:31:18.061863Z digest=sha256:b51a500ed9a498545aadef8c81d7d684542b1971da9e8c1a532f0e8a84cd2bc6

Observation 7cd99c0c-668e-45a6-a8d9-0625225c3568 · inbound

GUI-C$^2$: Coarse-to-Fine GUI Grounding via Difficulty-Aware Reinforcement Learning cites this paper.

GUI-C$^2$: Coarse-to-Fine GUI Grounding via Difficulty-Aware Reinforcement Learning ShowUI: One Vision-Language-Action Model for GUI Visual Agent

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-06-28T22:52:44.698903Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-28T22:52:04.755523Z digest=sha256:91ac8d3a48d8e2d1de50f4d499a564d035e27fa6297993823aeacccd3297b610

Observation 45d438a7-4e55-43d2-8f87-3af3cc421c45 · inbound

OSWorld 2.0: Benchmarking Computer Use Agents on Long-Horizon Real-World Tasks cites this paper.

OSWorld 2.0: Benchmarking Computer Use Agents on Long-Horizon Real-World Tasks ShowUI: One Vision-Language-Action Model for GUI Visual Agent

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-06-30T07:14:21.126877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T07:10:38.909339Z digest=sha256:3ba09dd1358a004ac7893c8da0c6671aa9867250521a115c2b855ce67ce27508

Observation 33244b6a-53c8-4fce-b7e7-641a23df1245 · inbound

OSWorld 2.0: Benchmarking Computer Use Agents on Long-Horizon Real-World Tasks cites this paper.

OSWorld 2.0: Benchmarking Computer Use Agents on Long-Horizon Real-World Tasks ShowUI: One Vision-Language-Action Model for GUI Visual Agent

Reference 52

Resolution
unresolved
no resolver link, observed 2026-07-15T10:24:53.345620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T10:24:53.345620Z digest=sha256:05a1e0c93df3df9bb71bd06fbaa3381a047f0a38ae29ed808a6ee098cc0a1246

Observation cd1725ad-d3e7-4fe6-bb89-4f6507c86a23 · inbound

Do GUI Agents Believe Their Eyes? Diagnosing State-Belief Reliance on Pixels versus Structure cites this paper.

Do GUI Agents Believe Their Eyes? Diagnosing State-Belief Reliance on Pixels versus Structure ShowUI: One Vision-Language-Action Model for GUI Visual Agent

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-11T20:00:19.123214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T20:00:19.123214Z digest=sha256:41f4214eb6cfe771fdf9403e91d8516d15f5124e874d3080381919fd7862a6d0

Observation 6ee9538c-55f7-44f6-a56c-cf5914914214 · inbound

Do GUI Agents Believe Their Eyes? Diagnosing State-Belief Reliance on Pixels versus Structure cites this paper.

Do GUI Agents Believe Their Eyes? Diagnosing State-Belief Reliance on Pixels versus Structure ShowUI: One Vision-Language-Action Model for GUI Visual Agent

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-02T08:46:59.431501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:46:59.431501Z digest=sha256:747d83f1e54b2f43b2c380abe1b88244cdc37bcd47dea61ebbbba89a4e669522

Observation 9a08a3d2-bf56-40cb-a8e0-2c9632d1a57e · inbound

Vision as Unified Multimodal Generation cites this paper.

Vision as Unified Multimodal Generation ShowUI: One Vision-Language-Action Model for GUI Visual Agent

Reference 108

Resolution
verified exact
local_arxiv, observed 2026-07-08T02:04:26.431023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-08T01:54:30.649092Z digest=sha256:1110813d3ef2dc98c588cfa4eab8ad70ede1599ac39bfdd54a8c256aecb772e9

Observation db8235ac-0579-45c3-84ea-dcf60d4b40ca · inbound

StepX-Edge: An On-Device UI Vision-Language Model via Architecture-Training-Deployment Co-Design cites this paper.

StepX-Edge: An On-Device UI Vision-Language Model via Architecture-Training-Deployment Co-Design ShowUI: One Vision-Language-Action Model for GUI Visual Agent

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-01T16:31:11.287336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:31:11.287336Z digest=sha256:ff089baaa785a8c27ca7beeacc47ed1c3875234bfc137183c9631be99d7859ae

Observation ba010f80-452c-400c-b9ce-9af0e7cc0593 · inbound

Rethinking Inference-Time Scaling in Local Computer-Use Agents: Failure Modes and Compute Tradeoffs cites this paper.

Rethinking Inference-Time Scaling in Local Computer-Use Agents: Failure Modes and Compute Tradeoffs ShowUI: One Vision-Language-Action Model for GUI Visual Agent

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-31T03:24:11.977217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T03:24:11.977217Z digest=sha256:7ebcd6c22e25b5b07d1e698b8e391e278d7cd6382e8de6e4e6b9553c4d509fc3