Pith. sign in

Paper Citation Record · LEDGER

ZonUI-3B: A Lightweight Vision-Language Model for Cross-Resolution GUI Grounding

As of 8 August 2026, this Paper Citation Record lists 12 of 12 outbound references and 0 inbound Pith citation observations for arXiv:2506.23491.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.23491 v3

Coverage vector

measured 12 of 12 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T21:44:51.870568Z

measured 12 of 12 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

12 of 12 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved12
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 992594d9-e92b-4bc9-97ea-9ebde3e7b141 · outbound

This paper cites Qwen2.5-VL Technical Report.

ZonUI-3B: A Lightweight Vision-Language Model for Cross-Resolution GUI Grounding Qwen2.5-VL Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T21:44:50.843200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:44:50.843200Z digest=sha256:95732cd8965e407052317846e719d20c984e7cd4f9b3a12e1b1b1234771d91ee

Observation 4324f7a0-8bdd-4710-b96e-c080e1e20730 · outbound

This paper cites AMEX: Android Multi-annotation Expo Dataset for Mobile GUI Agents.

ZonUI-3B: A Lightweight Vision-Language Model for Cross-Resolution GUI Grounding AMEX: Android Multi-annotation Expo Dataset for Mobile GUI Agents

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T21:44:50.910470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:44:50.910470Z digest=sha256:ca3f8f91354b03da3688ea9f0671555833c3bd5cdae6ce8319a5099c1c0c1a00

Observation 6e7e588f-5aa1-4c2b-95f4-f77288f202db · outbound

This paper cites SeeClick: Harnessing GUI Grounding for Advanced Visual GUI Agents.

ZonUI-3B: A Lightweight Vision-Language Model for Cross-Resolution GUI Grounding SeeClick: Harnessing GUI Grounding for Advanced Visual GUI Agents

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T21:44:51.037486Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:44:51.037486Z digest=sha256:1d2f23b04e55efa4995face8f0b175dfba8d9e677abb825328a7163d245a8bd8

Observation 10bf4088-d97f-444f-af54-c6f7aa581c79 · outbound

This paper cites Navigating the Digital World as Humans Do: Universal Visual Grounding for GUI Agents.

ZonUI-3B: A Lightweight Vision-Language Model for Cross-Resolution GUI Grounding Navigating the Digital World as Humans Do: Universal Visual Grounding for GUI Agents

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T21:44:51.133054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:44:51.133054Z digest=sha256:fb323c3ab45e661170051455d23b5207f4b204baca44cbbe490d7330aca519da

Observation cc4f8ab6-75d4-4a7d-bf8b-aa198f685dbe · outbound

This paper cites CogAgent: A Visual Language Model for GUI Agents.

ZonUI-3B: A Lightweight Vision-Language Model for Cross-Resolution GUI Grounding CogAgent: A Visual Language Model for GUI Agents

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T21:44:51.212777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:44:51.212777Z digest=sha256:2a3913aa1aae98bb85ea9d922e290a82535f15a7e4d702154a01da083a447229

Observation 943b2969-9783-4047-8f10-92d7dd04fe24 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

ZonUI-3B: A Lightweight Vision-Language Model for Cross-Resolution GUI Grounding LoRA: Low-Rank Adaptation of Large Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T21:44:51.300123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:44:51.300123Z digest=sha256:ca4cbdc16434b8b600449f5ffb2218ea4ebaa1cabec2d8c4237bb56f601c92e4

Observation e3cbe5b5-0cc5-473a-8312-81e95ddd2e71 · outbound

This paper cites ScreenSpot-Pro: GUI Grounding for Professional High-Resolution Computer Use.

ZonUI-3B: A Lightweight Vision-Language Model for Cross-Resolution GUI Grounding ScreenSpot-Pro: GUI Grounding for Professional High-Resolution Computer Use

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T21:44:51.404442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:44:51.404442Z digest=sha256:73f2741de2a9b4278e45588c95b3af286e141ae4bdbb8a8fc25c9f79eec9f99f

Observation 44f450de-7495-43d0-9caa-5f74dfb736c0 · outbound

This paper cites ShowUI: One Vision-Language-Action Model for GUI Visual Agent.

ZonUI-3B: A Lightweight Vision-Language Model for Cross-Resolution GUI Grounding ShowUI: One Vision-Language-Action Model for GUI Visual Agent

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T21:44:51.512452Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:44:51.512452Z digest=sha256:a9242aab28ef846bebaa23eab60c430dbec7442ef1643d537ff29ea40e2b9da2

Observation 5db9d50c-133d-4525-9eb1-e66ec95fc4c8 · outbound

This paper cites UI-TARS: Pioneering Automated GUI Interaction with Native Agents.

ZonUI-3B: A Lightweight Vision-Language Model for Cross-Resolution GUI Grounding UI-TARS: Pioneering Automated GUI Interaction with Native Agents

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T21:44:51.637524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:44:51.637524Z digest=sha256:65b8a1e7ca76e039d43182640d4d58f474144f62de2a077a53d5d923e225fe8d

Observation 1df72db2-919b-4610-876e-e429fe62d5e5 · outbound

This paper cites OS-ATLAS: A Foundation Action Model for Generalist GUI Agents.

ZonUI-3B: A Lightweight Vision-Language Model for Cross-Resolution GUI Grounding OS-ATLAS: A Foundation Action Model for Generalist GUI Agents

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T21:44:51.722301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:44:51.722301Z digest=sha256:2e258881c4c9705077ea77f9176e84dcaf58078f9cd4e5faab919d37f44e2a10

Observation ffa7ece6-be0e-4642-ae6d-871d955bb696 · outbound

This paper cites Aguvis: Unified Pure Vision Agents for Autonomous GUI Interaction.

ZonUI-3B: A Lightweight Vision-Language Model for Cross-Resolution GUI Grounding Aguvis: Unified Pure Vision Agents for Autonomous GUI Interaction

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T21:44:51.798796Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:44:51.798796Z digest=sha256:cc10446c0b17b984929315c704e5a38c6668bf14460c4afea35692c47a9c15f6

Observation 30a0a296-4d0d-4cc7-8987-ab2dfd049432 · outbound

This paper cites InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models.

ZonUI-3B: A Lightweight Vision-Language Model for Cross-Resolution GUI Grounding InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T21:44:51.870568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:44:51.870568Z digest=sha256:9202b981670762db0a4015bebf3563755095e37042b00e5084edb614dad5ede1

Pith citing papers

No inbound Pith citation observations are available.