Pith. sign in

Paper Citation Record · LEDGER

GUI-AIMA: Aligning Intrinsic Multimodal Attention with a Context Anchor for GUI Grounding

As of 5 August 2026, this Paper Citation Record lists 29 of 29 outbound references and 1 inbound Pith citation observation for arXiv:2511.00810.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2511.00810 v4

Coverage vector

measured 29 of 29 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T00:30:26.196169Z

measured 30 of 30 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-30T06:58:01.734975Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-30T07:04:21.780384Z

Reference resolution

29 of 29 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved29
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e39fffa7-6dcf-47d7-91b1-44da382ec99c · outbound

This paper cites Qwen2.5-VL Technical Report.

GUI-AIMA: Aligning Intrinsic Multimodal Attention with a Context Anchor for GUI Grounding Qwen2.5-VL Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-04T00:30:23.309580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:30:23.309580Z digest=sha256:974e17474959ba822fe436d6d7e3e363a4cf5b4678d6eefd5098184db411c417

Observation 1e202318-0dce-4686-ba21-2209760da269 · outbound

This paper cites What Does BERT Look At? An Analysis of BERT's Attention.

GUI-AIMA: Aligning Intrinsic Multimodal Attention with a Context Anchor for GUI Grounding What Does BERT Look At? An Analysis of BERT's Attention

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-04T00:30:23.591433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:30:23.591433Z digest=sha256:5d78d8253196b23eb8e0a326da38c98fd8bb57144d9fb68c0d8219368e05dd28

Observation ee7cf3f2-2050-42e6-a0a9-a7781fdf09f4 · outbound

This paper cites GUI-Bee: Align GUI Action Grounding to Novel Environments via Autonomous Exploration.

GUI-AIMA: Aligning Intrinsic Multimodal Attention with a Context Anchor for GUI Grounding GUI-Bee: Align GUI Action Grounding to Novel Environments via Autonomous Exploration

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T00:30:23.808488Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:30:23.808488Z digest=sha256:dc326db141d479be2ac72f2388c84708e22021dbea485226b215f3be5566f02a

Observation 13f1da45-8b1e-4948-8401-922d1044be0b · outbound

This paper cites Navigating the Digital World as Humans Do: Universal Visual Grounding for GUI Agents.

GUI-AIMA: Aligning Intrinsic Multimodal Attention with a Context Anchor for GUI Grounding Navigating the Digital World as Humans Do: Universal Visual Grounding for GUI Agents

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T00:30:23.903148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:30:23.903148Z digest=sha256:8c01327f897c95f8403fa869944cd12a7c16239ac471dae95f68e81f02957906

Observation a9f0c8f0-48a9-4953-b389-041c6ce702f9 · outbound

This paper cites The Dawn of GUI Agent: A Preliminary Case Study with Claude 3.5 Computer Use.

GUI-AIMA: Aligning Intrinsic Multimodal Attention with a Context Anchor for GUI Grounding The Dawn of GUI Agent: A Preliminary Case Study with Claude 3.5 Computer Use

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T00:30:23.975168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:30:23.975168Z digest=sha256:9a0f3d29f21397c0fdbb4fc0a438df79cd551b220cef6bca0a09da09ac7e0fff

Observation 427bfb5f-5655-4bdb-b990-456100610178 · outbound

This paper cites ScreenSpot-Pro: GUI Grounding for Professional High-Resolution Computer Use.

GUI-AIMA: Aligning Intrinsic Multimodal Attention with a Context Anchor for GUI Grounding ScreenSpot-Pro: GUI Grounding for Professional High-Resolution Computer Use

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T00:30:24.040770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:30:24.040770Z digest=sha256:b163a217b42acfe670d45fc5c84b496995aa63c09fd2d3e4fd0edc3fde25ecae

Observation c09c03e6-46bb-4a3a-b492-3c3eccfe00d0 · outbound

This paper cites ShowUI: One Vision-Language-Action Model for GUI Visual Agent.

GUI-AIMA: Aligning Intrinsic Multimodal Attention with a Context Anchor for GUI Grounding ShowUI: One Vision-Language-Action Model for GUI Visual Agent

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T00:30:24.114175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:30:24.114175Z digest=sha256:68dca249776aea3d624c9d0b498d52ed5c3c10f92ef3161788466f2eea1174c6

Observation ff7587c0-23c4-45d5-8d79-755ec649c23e · outbound

This paper cites Zhengxi Lu, Yuxiang Chai, Yaxuan Guo, Xi Yin, Liang Liu, Hao Wang, Han Xiao, Shuai Ren, Guanjing Xiong, and Hongsheng Li.

GUI-AIMA: Aligning Intrinsic Multimodal Attention with a Context Anchor for GUI Grounding Zhengxi Lu, Yuxiang Chai, Yaxuan Guo, Xi Yin, Liang Liu, Hao Wang, Han Xiao, Shuai Ren, Guanjing Xiong, and Hongsheng Li

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T00:30:24.190521Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:30:24.190521Z digest=sha256:f3e6e0ad01306c3cdd0bb2f926669c331b604c9176f83af26ff53ee5337b792a

Observation 98aae9bd-f7d0-4022-9cd3-65f84eca9e4c · outbound

This paper cites GUI-R1 : A Generalist R1-Style Vision-Language Action Model For GUI Agents.

GUI-AIMA: Aligning Intrinsic Multimodal Attention with a Context Anchor for GUI Grounding GUI-R1 : A Generalist R1-Style Vision-Language Action Model For GUI Agents

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T00:30:24.329429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:30:24.329429Z digest=sha256:ff68ce97e54ffab5a191b70e39fd65596f25751e7a3bd248dad1db9e11cb0ac1

Observation dd2ccb4c-3cbe-41e9-8a43-82d05f607ef7 · outbound

This paper cites In-context Learning and Induction Heads.

GUI-AIMA: Aligning Intrinsic Multimodal Attention with a Context Anchor for GUI Grounding In-context Learning and Induction Heads

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T00:30:24.418295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:30:24.418295Z digest=sha256:55e237e537b1d41b051e4f37e58f0801ca9ef6ead724e9467f8a5149621d98de

Observation d9be6992-341a-4c3e-a02e-f43c2a27a6aa · outbound

This paper cites UI-TARS: Pioneering Automated GUI Interaction with Native Agents.

GUI-AIMA: Aligning Intrinsic Multimodal Attention with a Context Anchor for GUI Grounding UI-TARS: Pioneering Automated GUI Interaction with Native Agents

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T00:30:24.495470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:30:24.495470Z digest=sha256:c15659bfb69215f086936700d9ecef57031441f127a3fd032524ac17860365e8

Observation ccb000bc-a6d2-4d22-b89b-f127dbd777ba · outbound

This paper cites AndroidWorld: A Dynamic Benchmarking Environment for Autonomous Agents.

GUI-AIMA: Aligning Intrinsic Multimodal Attention with a Context Anchor for GUI Grounding AndroidWorld: A Dynamic Benchmarking Environment for Autonomous Agents

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T00:30:24.641300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:30:24.641300Z digest=sha256:5b5c6b43f38818ecd5a9ec4f7c9546c59f4c146ef8e76af08d29ffea6524e306

Observation abf08a80-530f-4f86-9830-16c0557a9ab9 · outbound

This paper cites GUI-G$^2$: Gaussian Reward Modeling for GUI Grounding.

GUI-AIMA: Aligning Intrinsic Multimodal Attention with a Context Anchor for GUI Grounding GUI-G$^2$: Gaussian Reward Modeling for GUI Grounding

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T00:30:24.763113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:30:24.763113Z digest=sha256:0ca09c6b70d0dbc35ff9aa778f14a2b14ce6024d8991b6b3d501d0e454d44953

Observation 84c6474e-f102-42f1-9cd7-b00ee6f97686 · outbound

This paper cites Mobile-Agent: Autonomous Multi-Modal Mobile Device Agent with Visual Perception.

GUI-AIMA: Aligning Intrinsic Multimodal Attention with a Context Anchor for GUI Grounding Mobile-Agent: Autonomous Multi-Modal Mobile Device Agent with Visual Perception

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-04T00:30:24.934146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:30:24.934146Z digest=sha256:4a51f8c879df7af326502d4e8dd19097915d8e2e0837b13efb157fc21364f12a

Observation a88d6600-0a8e-4b79-bbb5-dfd5dfde4b35 · outbound

This paper cites OS-ATLAS: A Foundation Action Model for Generalist GUI Agents.

GUI-AIMA: Aligning Intrinsic Multimodal Attention with a Context Anchor for GUI Grounding OS-ATLAS: A Foundation Action Model for Generalist GUI Agents

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T00:30:25.003122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:30:25.003122Z digest=sha256:7f8ebf0ac5642aa468baac23a1c965d85b34b110c2c258af65facd0c0ed30d45

Observation f5892f10-c9c4-4b96-bf55-80ce1735b92d · outbound

This paper cites Scaling computer-use grounding via user interface decomposition and synthesis.arXiv preprint arXiv:2505.13227, 2025b.

GUI-AIMA: Aligning Intrinsic Multimodal Attention with a Context Anchor for GUI Grounding Scaling computer-use grounding via user interface decomposition and synthesis.arXiv preprint arXiv:2505.13227, 2025b

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-04T00:30:25.173013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:30:25.173013Z digest=sha256:bfb17678129b04a36ddee377aa834eeed80b828fa491f6ad1d22ec4f91742b89

Observation b851987c-5a10-40ab-85e2-bc1f1bce42a2 · outbound

This paper cites GTA1: GUI Test-time Scaling Agent.

GUI-AIMA: Aligning Intrinsic Multimodal Attention with a Context Anchor for GUI Grounding GTA1: GUI Test-time Scaling Agent

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-04T00:30:25.286424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:30:25.286424Z digest=sha256:8fb8b1a43b4f432f1445a9052452387f022b484c8ac6c4947dc531c090386a5a

Observation 89eacc2e-d49f-4c54-a7c7-71a20fce3f37 · outbound

This paper cites Mobile-Agent-v3: Fundamental Agents for GUI Automation.

GUI-AIMA: Aligning Intrinsic Multimodal Attention with a Context Anchor for GUI Grounding Mobile-Agent-v3: Fundamental Agents for GUI Automation

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-04T00:30:25.438320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:30:25.438320Z digest=sha256:35dde4f252f435189e8c66a672723e7a6973b9afffcd6aef74cdfc7a964e6075

Observation 1082df2f-537e-4cea-aac1-077ec8bce7d1 · outbound

This paper cites OmniParser V2: Structured-Points-of-Thought for Unified Visual Text Parsing and Its Generality to Multimodal Large Language Models.

GUI-AIMA: Aligning Intrinsic Multimodal Attention with a Context Anchor for GUI Grounding OmniParser V2: Structured-Points-of-Thought for Unified Visual Text Parsing and Its Generality to Multimodal Large Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-04T00:30:25.554263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:30:25.554263Z digest=sha256:0ef3b50e2adcd3abf62af5f9192c5bf31b63d5d33097c2aff60dae80e5487eb4

Observation 15ac2636-7501-4b68-9aee-81bae9f4d19e · outbound

This paper cites UFO: A UI-Focused Agent for Windows OS Interaction.

GUI-AIMA: Aligning Intrinsic Multimodal Attention with a Context Anchor for GUI Grounding UFO: A UI-Focused Agent for Windows OS Interaction

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-04T00:30:25.679125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:30:25.679125Z digest=sha256:e37a3f2b5254f38b04330364426db17028dd019b74db8bd8402a8b56068b74cc

Observation 49fb783d-6a20-40e6-9382-abb78a06a4e4 · outbound

This paper cites Appagent: Multimodal agents as smartphone users.

GUI-AIMA: Aligning Intrinsic Multimodal Attention with a Context Anchor for GUI Grounding Appagent: Multimodal agents as smartphone users

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-04T00:30:25.804623Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:30:25.804623Z digest=sha256:7e0cd228a4b8e2b8c54acbc271d1b6698f58c5e807118cc0436d402228a917a6

Observation deebc45c-d2ae-48f4-bdf4-c09549858777 · outbound

This paper cites GPT-4V(ision) is a Generalist Web Agent, if Grounded.

GUI-AIMA: Aligning Intrinsic Multimodal Attention with a Context Anchor for GUI Grounding GPT-4V(ision) is a Generalist Web Agent, if Grounded

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-04T00:30:25.933515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:30:25.933515Z digest=sha256:8860b3ba85da647a67b9aba5b797855c755ada65a3b7cc75e28b8efd6cb31405

Observation 0eb3ef86-65d2-4543-99cc-91c9732d39b1 · outbound

This paper cites GUI-G1: Understanding R1-Zero-Like Training for Visual Grounding in GUI Agents.

GUI-AIMA: Aligning Intrinsic Multimodal Attention with a Context Anchor for GUI Grounding GUI-G1: Understanding R1-Zero-Like Training for Visual Grounding in GUI Agents

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-04T00:30:26.090264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:30:26.090264Z digest=sha256:71e4a0c84373e9749c4e5f82703fb764276f05e634043935c34ce2312f07510d

Observation 185d94af-22eb-4742-88fa-e626f9f0a1f4 · outbound

This paper cites Embedding.

GUI-AIMA: Aligning Intrinsic Multimodal Attention with a Context Anchor for GUI Grounding Embedding

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-04T00:30:26.196169Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:30:26.196169Z digest=sha256:8a14c141f488c787991bf196b43731612134c0e8c5304a683ac7efcd300175a8

Observation fff5cb83-329d-4f33-bff5-9a11c6eaba45 · outbound

This paper cites Analyzing Multi-Head Self-Attention: Specialized Heads Do the Heavy Lifting, the Rest Can Be Pruned.

GUI-AIMA: Aligning Intrinsic Multimodal Attention with a Context Anchor for GUI Grounding Analyzing Multi-Head Self-Attention: Specialized Heads Do the Heavy Lifting, the Rest Can Be Pruned

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-04T00:30:24.831905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:30:24.831905Z digest=sha256:ce3b1c96efab089d82d49246a4240bad8894c772b21b0c4010eba56bcad9929a

Observation 64bb1954-a65f-4d53-bb5c-ccb54adf5d7e · outbound

This paper cites Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities.

GUI-AIMA: Aligning Intrinsic Multimodal Attention with a Context Anchor for GUI Grounding Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-04T00:30:23.665333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:30:23.665333Z digest=sha256:db559316cf27d1130a6994acbebc81dac233cdcbf81ebef712c5e38fe645f62f

Observation 94d3d990-f9e7-487f-900f-0a0747732e6a · outbound

This paper cites Inferring Functionality of Attention Heads from their Parameters.

GUI-AIMA: Aligning Intrinsic Multimodal Attention with a Context Anchor for GUI Grounding Inferring Functionality of Attention Heads from their Parameters

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-04T00:30:23.736838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:30:23.736838Z digest=sha256:577ca88cb3c407de37aeb9f803165406db91f15e6fbb75ea734ffbd8f93b4944

Observation 0499cf22-1ab0-42e7-acbe-a8bf0d715f14 · outbound

This paper cites GUICourse: From General Vision Language Models to Versatile GUI Agents.

GUI-AIMA: Aligning Intrinsic Multimodal Attention with a Context Anchor for GUI Grounding GUICourse: From General Vision Language Models to Versatile GUI Agents

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-04T00:30:23.490027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:30:23.490027Z digest=sha256:4c240474bc7222ec999b5a21f5cf6cffbd77892bf4cdd4ef30785c205fe90941

Observation e2007cee-e0a0-4394-aef7-f7a0f268f474 · outbound

This paper cites AMEX: Android Multi-annotation Expo Dataset for Mobile GUI Agents.

GUI-AIMA: Aligning Intrinsic Multimodal Attention with a Context Anchor for GUI Grounding AMEX: Android Multi-annotation Expo Dataset for Mobile GUI Agents

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-04T00:30:23.389647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:30:23.389647Z digest=sha256:08fcc2fb729e1a34e8516aba1ebdde2ff44ae991fc4d7af95bca378bb3f0f267

Pith citing papers

Observation 785fbea9-5e6e-4eca-a103-25d691629df6 · inbound

One Forward Beats Two: InnerZoom for Accurate and Efficient GUI Grounding cites this paper.

One Forward Beats Two: InnerZoom for Accurate and Efficient GUI Grounding GUI-AIMA: Aligning Intrinsic Multimodal Attention with a Context Anchor for GUI Grounding

Reference 110

Resolution
verified exact
arxiv_id, observed 2026-07-01T02:17:17.481434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-30T06:58:01.734975Z digest=sha256:54fc580958745d58a0944bd6b4c8082a671c1998dd9a51a48b68f3e15bd1a156