Pith. sign in

Paper Citation Record · LEDGER

ZonUI-3B: A Lightweight Vision-Language Model for Cross-Resolution GUI Grounding

As of 21 August 2026, this Paper Citation Record lists 12 of 12 outbound references and 0 inbound Pith citation observations for arXiv:2506.23491.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.23491 v3

Coverage vector

measured 12 of 12 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T21:44:51.870568Z

measured 12 of 12 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

12 of 12 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved12
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 992594d9-e92b-4bc9-97ea-9ebde3e7b141 · outbound

This paper cites Qwen2.5-VL Technical Report.

ZonUI-3B: A Lightweight Vision-Language Model for Cross-Resolution GUI Grounding Qwen2.5-VL Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T21:44:50.843200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:44:50.843200Z digest=sha256:f350d2c13aef6119aba82706dbb970bb65b76707616688ba6078f1ccdb8e1d05

Observation 4324f7a0-8bdd-4710-b96e-c080e1e20730 · outbound

This paper cites AMEX: Android Multi-annotation Expo Dataset for Mobile GUI Agents.

ZonUI-3B: A Lightweight Vision-Language Model for Cross-Resolution GUI Grounding AMEX: Android Multi-annotation Expo Dataset for Mobile GUI Agents

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T21:44:50.910470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:44:50.910470Z digest=sha256:168b6f19b7aeec68c70df8dc1a37d4e7e4e12e292ee4ebdf0bc2cb0eef47de74

Observation 6e7e588f-5aa1-4c2b-95f4-f77288f202db · outbound

This paper cites SeeClick: Harnessing GUI Grounding for Advanced Visual GUI Agents.

ZonUI-3B: A Lightweight Vision-Language Model for Cross-Resolution GUI Grounding SeeClick: Harnessing GUI Grounding for Advanced Visual GUI Agents

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T21:44:51.037486Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:44:51.037486Z digest=sha256:1f99532c4b013e165520a94d5d45492a8d0903893bcf4d093e77ba2763e184c5

Observation 10bf4088-d97f-444f-af54-c6f7aa581c79 · outbound

This paper cites Navigating the Digital World as Humans Do: Universal Visual Grounding for GUI Agents.

ZonUI-3B: A Lightweight Vision-Language Model for Cross-Resolution GUI Grounding Navigating the Digital World as Humans Do: Universal Visual Grounding for GUI Agents

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T21:44:51.133054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:44:51.133054Z digest=sha256:3bdc0d9215762ad3d4e25f09eb729930c68a27c3b33fab60d6cec72f84e863ef

Observation cc4f8ab6-75d4-4a7d-bf8b-aa198f685dbe · outbound

This paper cites CogAgent: A Visual Language Model for GUI Agents.

ZonUI-3B: A Lightweight Vision-Language Model for Cross-Resolution GUI Grounding CogAgent: A Visual Language Model for GUI Agents

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T21:44:51.212777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:44:51.212777Z digest=sha256:ae7d8ffd7b01707947a9fe3c4ecb3ae2c3fc1ae3ded26977161501a85c13c6b5

Observation 943b2969-9783-4047-8f10-92d7dd04fe24 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

ZonUI-3B: A Lightweight Vision-Language Model for Cross-Resolution GUI Grounding LoRA: Low-Rank Adaptation of Large Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T21:44:51.300123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:44:51.300123Z digest=sha256:deb7d22d98071a3fcc2b997cd34f359c3b4fc56903614d95479ebcbe8d6b1db0

Observation e3cbe5b5-0cc5-473a-8312-81e95ddd2e71 · outbound

This paper cites ScreenSpot-Pro: GUI Grounding for Professional High-Resolution Computer Use.

ZonUI-3B: A Lightweight Vision-Language Model for Cross-Resolution GUI Grounding ScreenSpot-Pro: GUI Grounding for Professional High-Resolution Computer Use

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T21:44:51.404442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:44:51.404442Z digest=sha256:7844a0093564acdd4e5ed7c22501fcd5135fc9d3d58363da0da5be7a10b0dc78

Observation 44f450de-7495-43d0-9caa-5f74dfb736c0 · outbound

This paper cites ShowUI: One Vision-Language-Action Model for GUI Visual Agent.

ZonUI-3B: A Lightweight Vision-Language Model for Cross-Resolution GUI Grounding ShowUI: One Vision-Language-Action Model for GUI Visual Agent

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T21:44:51.512452Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:44:51.512452Z digest=sha256:0b7346470a5bb8cdbc92d6abfd970de0c71cb07410fc111dd9195f7ae9290332

Observation 5db9d50c-133d-4525-9eb1-e66ec95fc4c8 · outbound

This paper cites UI-TARS: Pioneering Automated GUI Interaction with Native Agents.

ZonUI-3B: A Lightweight Vision-Language Model for Cross-Resolution GUI Grounding UI-TARS: Pioneering Automated GUI Interaction with Native Agents

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T21:44:51.637524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:44:51.637524Z digest=sha256:e6ce33e182719031ea53a5e14277bf688a16427678f767b93a577c71279c34af

Observation 1df72db2-919b-4610-876e-e429fe62d5e5 · outbound

This paper cites OS-ATLAS: A Foundation Action Model for Generalist GUI Agents.

ZonUI-3B: A Lightweight Vision-Language Model for Cross-Resolution GUI Grounding OS-ATLAS: A Foundation Action Model for Generalist GUI Agents

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T21:44:51.722301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:44:51.722301Z digest=sha256:2015b8d626f73cea1c687c3b9e28a2e6885b65d1602876c154ef74be13aacca4

Observation ffa7ece6-be0e-4642-ae6d-871d955bb696 · outbound

This paper cites Aguvis: Unified Pure Vision Agents for Autonomous GUI Interaction.

ZonUI-3B: A Lightweight Vision-Language Model for Cross-Resolution GUI Grounding Aguvis: Unified Pure Vision Agents for Autonomous GUI Interaction

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T21:44:51.798796Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:44:51.798796Z digest=sha256:919a6ddbaa5a83a4a83fb7b0efae40b8e7bafa480e0b710fee8ea4f5c9b1694e

Observation 30a0a296-4d0d-4cc7-8987-ab2dfd049432 · outbound

This paper cites InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models.

ZonUI-3B: A Lightweight Vision-Language Model for Cross-Resolution GUI Grounding InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T21:44:51.870568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:44:51.870568Z digest=sha256:e7faf6fc2a6bd176a30669d08ffd0d43df7dc5db7f5d61b3c6fc5a0d20636aac

Pith citing papers

No inbound Pith citation observations are available.