Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 23 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 27 inbound Pith citation observations for arXiv:2404.05955.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-15T23:21:13.438310Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-08T16:15:06.153368Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation 4e1379bf-4769-48f9-97d2-52098f1f293d · inbound
Personal LLM Agents: Insights and Survey about the Capability, Efficiency and Security VisualWebBench: How Far Have Multimodal LLMs Evolved in Web Page Understanding and Grounding?
Reference 136
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 012d0a51-fcba-4c5d-a8b4-f3b6f5d65818 · inbound
Large Language Model-Brained GUI Agents: A Survey VisualWebBench: How Far Have Multimodal LLMs Evolved in Web Page Understanding and Grounding?
Reference 218
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation c24df63a-9a42-40c3-b1d7-218e41340dab · inbound
Iris: Breaking GUI Complexity with Adaptive Focus and Self-Refining VisualWebBench: How Far Have Multimodal LLMs Evolved in Web Page Understanding and Grounding?
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 66533dcb-d213-4fb2-b63a-1e7b8fa93d16 · inbound
Attention-driven GUI Grounding: Leveraging Pretrained Multimodal Large Language Models without Fine-Tuning VisualWebBench: How Far Have Multimodal LLMs Evolved in Web Page Understanding and Grounding?
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cbf8049c-5d5b-4b04-ac59-83bd09176a9c · inbound
InSTA: Towards Internet-Scale Training For Agents VisualWebBench: How Far Have Multimodal LLMs Evolved in Web Page Understanding and Grounding?
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 094a4728-86ac-42c7-a5dd-af30dd98fdb7 · inbound
TRISHUL: Towards Region Identification and Screen Hierarchy Understanding for Large VLM based GUI Agents VisualWebBench: How Far Have Multimodal LLMs Evolved in Web Page Understanding and Grounding?
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8b400764-19b1-4e5e-a4ee-45093d0d9ca4 · inbound
Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models VisualWebBench: How Far Have Multimodal LLMs Evolved in Web Page Understanding and Grounding?
Reference 291
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 45680f7d-05e7-4bf0-8c52-88abf138915e · inbound
Seed1.5-VL Technical Report VisualWebBench: How Far Have Multimodal LLMs Evolved in Web Page Understanding and Grounding?
Reference 80
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation c2c3d76d-d054-49be-a60d-cadda5bc6f84 · inbound
FullFront: Benchmarking MLLMs Across the Full Front-End Engineering Workflow VisualWebBench: How Far Have Multimodal LLMs Evolved in Web Page Understanding and Grounding?
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1a4b7967-1c27-43b6-ab93-12049c63494b · inbound
ZeroGUI: Automating Online GUI Learning at Zero Human Cost VisualWebBench: How Far Have Multimodal LLMs Evolved in Web Page Understanding and Grounding?
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 58f43906-2a00-43a5-9296-71b30caad99b · inbound
MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments VisualWebBench: How Far Have Multimodal LLMs Evolved in Web Page Understanding and Grounding?
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation db0c7814-3e4a-4b6f-86b6-8e5a0d247c6d · inbound
MiMo-VL Technical Report VisualWebBench: How Far Have Multimodal LLMs Evolved in Web Page Understanding and Grounding?
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b4bc93dd-f677-4116-8ae0-4a539e227d7b · inbound
Atomic-to-Compositional Generalization for Mobile Agents with A New Benchmark and Scheduling System VisualWebBench: How Far Have Multimodal LLMs Evolved in Web Page Understanding and Grounding?
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d71e48b4-7a6a-4fe1-8633-4b0352734e33 · inbound
Build the web for agents, not agents for the web VisualWebBench: How Far Have Multimodal LLMs Evolved in Web Page Understanding and Grounding?
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0a562665-edbc-40bf-b491-b13ba4ea8fd5 · inbound
Argus Inspection: Do Multimodal Large Language Models Possess the Eye of Panoptes? VisualWebBench: How Far Have Multimodal LLMs Evolved in Web Page Understanding and Grounding?
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation af8737ca-a8eb-4f6f-adec-119ddcb46c29 · inbound
OS-MAP: How Far Can Computer-Using Agents Go in Breadth and Depth? VisualWebBench: How Far Have Multimodal LLMs Evolved in Web Page Understanding and Grounding?
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3419e5f9-f23f-4dee-9cbc-ad7f3ed5bbb0 · inbound
SEAgent: Self-Evolving Computer Use Agent with Autonomous Learning from Experience VisualWebBench: How Far Have Multimodal LLMs Evolved in Web Page Understanding and Grounding?
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 490dfc8c-26ba-4655-9663-14c5316641ee · inbound
WebMMU: A Benchmark for Multimodal Multilingual Website Understanding and Code Generation VisualWebBench: How Far Have Multimodal LLMs Evolved in Web Page Understanding and Grounding?
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5edcc879-91b4-4fd9-928a-13aad209714d · inbound
UItron: Foundational GUI Agent with Advanced Perception and Planning VisualWebBench: How Far Have Multimodal LLMs Evolved in Web Page Understanding and Grounding?
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 161fceba-5e13-4c5c-bea0-3232da40bec5 · inbound
Robix: A Unified Model for Robot Interaction, Reasoning and Planning VisualWebBench: How Far Have Multimodal LLMs Evolved in Web Page Understanding and Grounding?
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9bc3b8ab-3acb-4cf4-80a6-49d71ba73a65 · inbound
DocOS: Towards Proactive Document-Guided Actions in GUI Agents VisualWebBench: How Far Have Multimodal LLMs Evolved in Web Page Understanding and Grounding?
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 7ba40c9d-fe48-49ab-b9fb-f30ea15acccd · inbound
PACE: A Proxy for Agentic Capability Evaluation VisualWebBench: How Far Have Multimodal LLMs Evolved in Web Page Understanding and Grounding?
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 94f74fcc-0954-4848-b120-156b7c4611e5 · inbound
PACE: A Proxy for Agentic Capability Evaluation VisualWebBench: How Far Have Multimodal LLMs Evolved in Web Page Understanding and Grounding?
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8e37b665-4417-4cd4-acbf-bbb52ec8eb24 · inbound
WebRetriever: A Large-Scale Comprehensive Benchmark for Efficient Web Agent Evaluation VisualWebBench: How Far Have Multimodal LLMs Evolved in Web Page Understanding and Grounding?
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation c705f643-7e56-4759-9253-c5959683b8ad · inbound
Scaling GUI Agents with Visual State Transitions VisualWebBench: How Far Have Multimodal LLMs Evolved in Web Page Understanding and Grounding?
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e9beb5ea-b3ec-47a3-b7b2-d89a5d70fffd · inbound
XL-DocBench: Benchmarking Evidence-Grounded Extra-Long Document Understanding VisualWebBench: How Far Have Multimodal LLMs Evolved in Web Page Understanding and Grounding?
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 359f681f-19fb-4464-88d7-4ca2c2253f49 · inbound
Software Engineering for and with GUI Agent VisualWebBench: How Far Have Multimodal LLMs Evolved in Web Page Understanding and Grounding?
Reference 140
Source-reported events for the cited work
Unavailable: canonical work link unavailable.