Pith. sign in

Paper Citation Record · LEDGER

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities

As of 7 August 2026, this Paper Citation Record lists 46 of 46 outbound references and 1 inbound Pith citation observation for arXiv:2506.08933.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.08933 v1

Coverage vector

measured 46 of 46 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:02:29.300721Z

measured 47 of 47 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-20T11:13:48.529796Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-20T11:18:13.822352Z

Reference resolution

46 of 46 outbound references displayed

  • verified exact0
  • verified fuzzy14
  • unresolved30
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8f8405fa-a978-4069-8133-d9e233067f62 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T05:02:29.127878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:02:29.127878Z digest=sha256:90938aba279b0f024bc38ba6b3e649a8b13614473ac30c46c6282fc6890fead6

Observation 5c44b633-0590-4036-a2e2-3fd8e01b96ee · outbound

This paper cites an unresolved cited work.

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:02:34.084103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T05:02:29.132482Z digest=sha256:7138cf9a90653ce864d4d520c5dbac2f3bc46270d923823ca6f27148b6f42ac3

Observation 1e40ed52-f436-4e95-88cd-94d4f9cc67fb · outbound

This paper cites Spider2-V: How Far Are Multimodal Agents From Automating Data Science and Engineering Workflows?.

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities Spider2-V: How Far Are Multimodal Agents From Automating Data Science and Engineering Workflows?

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T05:02:29.136593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:02:29.136593Z digest=sha256:9b87236674417d682e44e8854ccf7e125e10ac3e61e7dc1923ea2f255813ccd2

Observation c585176a-3b37-4fc7-ad61-f9bb97bedd82 · outbound

This paper cites Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks.

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:02:33.919293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T05:02:29.140682Z digest=sha256:3d87d44f49db5d7d1f03da44c9d3ab49c1f4f23f49c8bc73a5a4037072c154b4

Observation e72355a2-aa45-4aa6-83bc-121b7dbb4d8f · outbound

This paper cites SeeClick: Harnessing GUI Grounding for Advanced Visual GUI Agents.

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities SeeClick: Harnessing GUI Grounding for Advanced Visual GUI Agents

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T05:02:29.144415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:02:29.144415Z digest=sha256:b6cc53a7d28c76c87af8481b463e41966698da1593e9b38b6f5df4610cf7b387

Observation 9d9c051e-a413-4673-9b0f-7ebe379f7d5b · outbound

This paper cites Mind2web: Towards a generalist agent for the web.

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities Mind2web: Towards a generalist agent for the web

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:02:33.734383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T05:02:29.148148Z digest=sha256:a7c86166ced88abc4cf869dd6ef4ac2dc0a682c98605e7207e1a06dd797f5d9c

Observation 67c688ee-3875-47df-a4da-3c9fce6c7863 · outbound

This paper cites Dysen-vdm: Empowering dynamics-aware text-to-video diffusion with llms.

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities Dysen-vdm: Empowering dynamics-aware text-to-video diffusion with llms

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:02:33.528030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T05:02:29.152268Z digest=sha256:b182c33042b63bf63e7b45c4a411ec2593b751ec76aa7c9d6632839435a2ed68

Observation 50aed938-2aa1-4eab-a811-b90eb574a63f · outbound

This paper cites Video-of-thought: Step-by-step video reasoning from perception to cognition.

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities Video-of-thought: Step-by-step video reasoning from perception to cognition

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:02:33.349259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T05:02:29.155750Z digest=sha256:0bd51c1d288e8d0a7210bb8bcea80b80d1ba2aad40fae42a400c3d60cf67daac

Observation 1bf59cfd-e644-4886-8d98-1efac1ee8971 · outbound

This paper cites Vitron: A unified pixel-level vision llm for understanding, generating, segmenting, editing.

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities Vitron: A unified pixel-level vision llm for understanding, generating, segmenting, editing

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:02:33.167945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T05:02:29.159175Z digest=sha256:f9db97c43de84b3d95ef6cec5a545496389c2834ed145f6e89be860322f55aba

Observation 5cc1e063-7d1b-4e3b-8cfe-2036738d51d2 · outbound

This paper cites Enhancing video-language representations with structural spatio-temporal alignment.

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities Enhancing video-language representations with structural spatio-temporal alignment

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:02:32.995079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T05:02:29.163035Z digest=sha256:65626a252498a65b267043f1b809bae11934f6daa3c71d72f6f7356a957d9ad4

Observation 0fcf5a49-4059-4727-be64-f8e5bbd3ba15 · outbound

This paper cites Generalist Virtual Agents: A Survey on Autonomous Agents Across Digital Platforms.

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities Generalist Virtual Agents: A Survey on Autonomous Agents Across Digital Platforms

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T05:02:29.166735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:02:29.166735Z digest=sha256:3d0e68ead9a6dd25cd3bfb20a1336ce635d91ddfd4cfbbf0eb9cdd4f85216ca0

Observation b3fe1d08-9cfd-4472-b003-c31dc1ce5fa1 · outbound

This paper cites De-fine: Decomposing and refining visual programs with auto-feedback.

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities De-fine: Decomposing and refining visual programs with auto-feedback

Reference 12

Resolution
metadata mismatch
raw_fallback, observed 2026-08-07T05:02:30.062241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T05:02:29.170447Z digest=sha256:e246e6c33e1951cbd53e3aa22aca9cf3a8f6e5455def315649287eaee86f25d6

Observation 37941295-095e-4a19-ba50-e32226ef34bb · outbound

This paper cites Benchmarking Multimodal CoT Reward Model Stepwise by Visual Program.

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities Benchmarking Multimodal CoT Reward Model Stepwise by Visual Program

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T05:02:29.173944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:02:29.173944Z digest=sha256:f7b287aee354fa4da707decfda7db29f1837df7cb380080cad7f8f4d836cbdcd

Observation 47c6f886-57fb-4877-b41b-63b287dd429e · outbound

This paper cites Iris: Breaking GUI Complexity with Adaptive Focus and Self-Refining.

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities Iris: Breaking GUI Complexity with Adaptive Focus and Self-Refining

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T05:02:29.178376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:02:29.178376Z digest=sha256:1488efaf038aef6e1b9ac0041ade86639d3889ab6b432275db85836808c1ac89

Observation fe01b38f-2a86-416b-8679-b17ba9bfccf1 · outbound

This paper cites Navigating the Digital World as Humans Do: Universal Visual Grounding for GUI Agents.

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities Navigating the Digital World as Humans Do: Universal Visual Grounding for GUI Agents

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T05:02:29.182555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:02:29.182555Z digest=sha256:6a29e344fcdef205d7f00fcadb6034bd8c7d8610a57be8df388ccbb48a5a8424

Observation 4b378d0a-8832-4d3f-9430-80b301def70a · outbound

This paper cites Cogagent: A visual language model for gui agents.

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities Cogagent: A visual language model for gui agents

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:02:32.731848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T05:02:29.186253Z digest=sha256:ff797d525b5e5da67ee8119b2a605a7aa35578dbf32460ab336727c6beccb0b3

Observation 6eeeba8a-e01f-4a5f-962a-9c4bf82099da · outbound

This paper cites The Dawn of GUI Agent: A Preliminary Case Study with Claude 3.5 Computer Use.

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities The Dawn of GUI Agent: A Preliminary Case Study with Claude 3.5 Computer Use

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T05:02:29.189552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:02:29.189552Z digest=sha256:fdf40f9aed4300934600ef854290edaf6f96c343e212ba2f304662d1a50e9027

Observation dc06c457-3cfb-4c4f-a88c-688f5b10b629 · outbound

This paper cites GPT-4o System Card.

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities GPT-4o System Card

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T05:02:29.193533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:02:29.193533Z digest=sha256:928d4d0c1c71a000f864e152bdd1e28e48d3914bdd3f6d27f192748d0fc22f1d

Observation 4e48966b-d59f-4071-9c24-86c6d746d836 · outbound

This paper cites P., Russak, M., Koh, J.

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities P., Russak, M., Koh, J

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:02:32.475848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T05:02:29.197603Z digest=sha256:6443f44e0b80aa610f42fd9523c86a4e39e4cff0edf8c80c08fccaba97dfbb4f

Observation f121153e-fd89-45f5-93ad-6929d8b236cc · outbound

This paper cites VisualWebArena: Evaluating Multimodal Agents on Realistic Visual Web Tasks.

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities VisualWebArena: Evaluating Multimodal Agents on Realistic Visual Web Tasks

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T05:02:29.201086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:02:29.201086Z digest=sha256:6375d92be71ddfaecc5c81743360f3a3f045f90dc1e1e78f46b411f55315e68c

Observation 09b81ed4-503d-4aec-9e09-16341fd9932b · outbound

This paper cites an unresolved cited work.

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:02:32.148264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T05:02:29.204810Z digest=sha256:30dfa38d43e1d1d83845820e5d9b27caba800f318ec656d52843e509450e4153

Observation 876b291c-883c-4483-9127-a323e8e9a1f2 · outbound

This paper cites Fine-grained semantically aligned vision-language pre-training.

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities Fine-grained semantically aligned vision-language pre-training

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:02:31.782608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T05:02:29.208063Z digest=sha256:484c69288a2e230edf3bff7307cb38ad41ed0a9703058280b82f8d04364ae4a0

Observation 00160e87-6b8f-499d-9ed5-1656767e577f · outbound

This paper cites Fine-tuning multimodal llms to follow zero-shot demonstrative instructions.

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities Fine-tuning multimodal llms to follow zero-shot demonstrative instructions

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:02:31.509796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T05:02:29.211153Z digest=sha256:7bb4f572ead52584115ab9f6c21564f7960f3e1f40c78316607d85c32354a6f0

Observation 3f872814-3c4a-4412-8a69-c117157f7ba5 · outbound

This paper cites Variational cross-graph reasoning and adaptive structured semantics learning for compositional temporal grounding.

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities Variational cross-graph reasoning and adaptive structured semantics learning for compositional temporal grounding

Reference 24

Resolution
metadata mismatch
raw_fallback, observed 2026-08-07T05:02:29.764865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T05:02:29.214658Z digest=sha256:d325e1cac9eac415353c85af694aa2d99a9028d45d1e5d795f9f0f3dd0f622d1

Observation 0347396e-3fa2-4774-9a3d-65449bb06670 · outbound

This paper cites Mapping Natural Language Instructions to Mobile UI Action Sequences.

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities Mapping Natural Language Instructions to Mobile UI Action Sequences

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T05:02:29.217639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:02:29.217639Z digest=sha256:42cdefca2db335c2e0e200cce492c81a58b874985aec4ae40e139d7ae87c65f5

Observation 83b75b66-3873-4d89-8e09-65c9ab7b1b8b · outbound

This paper cites ShowUI: One Vision-Language-Action Model for GUI Visual Agent.

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities ShowUI: One Vision-Language-Action Model for GUI Visual Agent

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T05:02:29.221294Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:02:29.221294Z digest=sha256:3cfd2f2f3f89c3334f4ca57fb4f10368701ba4f31f77412b4c8291a68f192ce9

Observation 089d71e4-e170-48f0-81f2-235ddf624bfd · outbound

This paper cites GUIOdyssey: A Comprehensive Dataset for Cross-App GUI Navigation on Mobile Devices.

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities GUIOdyssey: A Comprehensive Dataset for Cross-App GUI Navigation on Mobile Devices

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T05:02:29.228691Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:02:29.228691Z digest=sha256:40c9885bbde779976f5339bfb4c9e607bb410e7ada3d17f39c3c4e10454b3d4c

Observation ae985b10-a123-412a-a6b7-b4d7fac1a85a · outbound

This paper cites Boosting Virtual Agent Learning and Reasoning: A Step-Wise, Multi-Dimensional, and Generalist Reward Model with Benchmark.

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities Boosting Virtual Agent Learning and Reasoning: A Step-Wise, Multi-Dimensional, and Generalist Reward Model with Benchmark

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T05:02:29.232405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:02:29.232405Z digest=sha256:b1900996f3a338c07b3b29addd4e6c59bf2e4176b1a9a403de2337a421e40be1

Observation 35383744-47a4-4262-8a60-eea83d4ecbb3 · outbound

This paper cites Self-supervised Meta-Prompt Learning with Meta-Gradient Regularization for Few-shot Generalization.

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities Self-supervised Meta-Prompt Learning with Meta-Gradient Regularization for Few-shot Generalization

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T05:02:29.236247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:02:29.236247Z digest=sha256:0e4eb753d192da93bb9a4f81c554202d428986f524bc7e702cd0119a529f4594

Observation 22983b81-c386-4972-8cf2-d3a86314cd39 · outbound

This paper cites Towards unified multimodal editing with enhanced knowledge collaboration.

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities Towards unified multimodal editing with enhanced knowledge collaboration

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:02:31.207365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T05:02:29.239831Z digest=sha256:94ac5e217d820331246f6642da40a5328af084bf4ec5643c982ed3f0dffd317a

Observation 84353a44-ec40-4be6-9285-b5c2ae7fab47 · outbound

This paper cites I3: I ntent-i ntrospective retrieval conditioned on i nstructions.

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities I3: I ntent-i ntrospective retrieval conditioned on i nstructions

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:02:30.908383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T05:02:29.243088Z digest=sha256:693fc8840683e3527f81813d7198c34a75bff7dff3ef0ba66c3d813f1edefb96

Observation 17e9e6e2-2b33-4e2e-8b5b-99c661c349f5 · outbound

This paper cites Auto-Encoding Morph-Tokens for Multimodal LLM.

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities Auto-Encoding Morph-Tokens for Multimodal LLM

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T05:02:29.246334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:02:29.246334Z digest=sha256:008ea562b425d1851a16f146ae0a8a93d0da73c1346b66e68bcd0abd4e8dd5b8

Observation 9ae4e593-e82f-407f-bafa-0a8250e4ea4f · outbound

This paper cites Generative Multimodal Pretraining with Discrete Diffusion Timestep Tokens.

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities Generative Multimodal Pretraining with Discrete Diffusion Timestep Tokens

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T05:02:29.249982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:02:29.249982Z digest=sha256:e3a34847b64ba199136b2a006f056cfdabc1b98d518e34458691d217af6943e7

Observation 2d9bacd5-bd51-44a3-be3e-09a722465bc2 · outbound

This paper cites Unlocking aha moments via reinforcement learning: Advancing collaborative visual comprehension and generation.

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities Unlocking aha moments via reinforcement learning: Advancing collaborative visual comprehension and generation

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T05:02:29.254238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:02:29.254238Z digest=sha256:be9e909cf45417f4789f8cb4cff20bd424b79bdbce15594b7f90322cb9a750ff

Observation 4a7d1a39-3a00-4c29-a01b-fa13b471f726 · outbound

This paper cites Androidinthewild: A large-scale dataset for android device control.

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities Androidinthewild: A large-scale dataset for android device control

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:02:30.570748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T05:02:29.258032Z digest=sha256:6c4814e23490dd89ae0ffea692721079b9e7b529997cbffee8d5e29528481928

Observation 215b55d8-7e85-4afa-aed2-eeb7a5b6b12d · outbound

This paper cites ScribeAgent: Towards Specialized Web Agents Using Production-Scale Workflow Data.

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities ScribeAgent: Towards Specialized Web Agents Using Production-Scale Workflow Data

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T05:02:29.261946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:02:29.261946Z digest=sha256:ed883f879e08ad6ec70645f6209736f2f0b2d3d34f3e185860d173a88a0ec140

Observation 06084d52-bd1f-4843-ba0b-22d00827e86e · outbound

This paper cites TaskBench: Benchmarking Large Language Models for Task Automation.

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities TaskBench: Benchmarking Large Language Models for Task Automation

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T05:02:29.265475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:02:29.265475Z digest=sha256:756ab8e8e3ae08e85fcf88c5b2da1c997663cb3fdf15f481bdf1d6449aee43f0

Observation dc08f448-c10b-4727-b331-adce39e1ed8a · outbound

This paper cites MobileAgentBench: An Efficient and User-Friendly Benchmark for Mobile LLM Agents.

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities MobileAgentBench: An Efficient and User-Friendly Benchmark for Mobile LLM Agents

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T05:02:29.269282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:02:29.269282Z digest=sha256:cfd6b48a0804b59f36c985f54fa4e98255b0bc6b443072f816baeff7b6a6e669

Observation 11c0cfee-b8d0-4148-a46d-5aec8bc47488 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T05:02:29.273032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:02:29.273032Z digest=sha256:f58ae939e2877c69210f08bb049ee417a6010b98a162c0747efa0d06e22f0bc7

Observation 68938c66-c337-432b-a812-8693b38d13aa · outbound

This paper cites NE x T - GPT : Any-to-any multimodal LLM.

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities NE x T - GPT : Any-to-any multimodal LLM

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:02:30.239250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T05:02:29.276630Z digest=sha256:fc6b0c4aaee10d1bcc8e91c77a9e1bf9a38804955d264406a62c37c6c8fdb2e7

Observation 08be7f69-11fe-49c2-9877-76c527e51f68 · outbound

This paper cites OS-ATLAS: A Foundation Action Model for Generalist GUI Agents.

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities OS-ATLAS: A Foundation Action Model for Generalist GUI Agents

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T05:02:29.280843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:02:29.280843Z digest=sha256:bd0a9bb1f033c3de0cfe1871f5b41ed1bed8b108e28acb2f731ddd0b1ea6ed82

Observation bb747de6-b1f3-453e-8a2f-ccdbc3835f84 · outbound

This paper cites OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments.

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T05:02:29.285064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:02:29.285064Z digest=sha256:a7675c6fed974d038cbdb4f7e13febf5000006772173ba390db3fcdafb09b3c9

Observation c5611f9f-14ea-4cb8-95f4-037ad2065bbb · outbound

This paper cites CRAB: Cross-environment Agent Benchmark for Multimodal Language Model Agents.

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities CRAB: Cross-environment Agent Benchmark for Multimodal Language Model Agents

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T05:02:29.289598Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:02:29.289598Z digest=sha256:35c3812104c87048c81a5bd3b42761cbe6249a0565197af87d0c16ae6164abcc

Observation 9f40b8c8-a93a-4918-a518-b39157dca0fd · outbound

This paper cites Aguvis: Unified Pure Vision Agents for Autonomous GUI Interaction.

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities Aguvis: Unified Pure Vision Agents for Autonomous GUI Interaction

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T05:02:29.293289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:02:29.293289Z digest=sha256:e9f32ddca59340d1d3b8d30d0201ea7444597735e5ba1e37d86e299c58eff86f

Observation 6febf858-5e25-45ba-986e-e62d8cbc97f9 · outbound

This paper cites Webshop: Towards scalable real-world web interaction with grounded language agents.

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities Webshop: Towards scalable real-world web interaction with grounded language agents

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T05:02:29.297322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:02:29.297322Z digest=sha256:b82fb51aeb37717a015d8ca459eacb4a2a27fdd1fccb2d71482eed2425fd3869

Observation fd8ab3c2-cc57-474f-b695-c15d6a274724 · outbound

This paper cites WebArena: A Realistic Web Environment for Building Autonomous Agents.

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities WebArena: A Realistic Web Environment for Building Autonomous Agents

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T05:02:29.300721Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:02:29.300721Z digest=sha256:787920bb47895d9ea3ecef0c58c73baea72ced06a778623cc5fb7bb9e59d3f08

Pith citing papers

Observation 337ed338-7df9-4130-b2f5-f60bf7e91408 · inbound

DocOS: Towards Proactive Document-Guided Actions in GUI Agents cites this paper.

DocOS: Towards Proactive Document-Guided Actions in GUI Agents What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-20T11:18:13.823868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-20T11:13:48.529796Z digest=sha256:e7348e07660b6e5c6efd0ae0a892e34dbd29cc02530089a92d54a638fd24ae37