Pith. sign in

Paper Citation Record · LEDGER

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities

As of 23 August 2026, this Paper Citation Record lists 46 of 46 outbound references and 2 inbound Pith citation observations for arXiv:2506.08933.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.08933 v1

Coverage vector

measured 46 of 46 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:02:29.300721Z

measured 48 of 48 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T19:54:51.906676Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-20T11:18:13.822352Z

Reference resolution

46 of 46 outbound references displayed

  • verified exact0
  • verified fuzzy14
  • unresolved30
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8f8405fa-a978-4069-8133-d9e233067f62 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T05:02:29.127878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:02:29.127878Z digest=sha256:61c51c2942bfd3eb80d84538443163d4dd610bfb26f6de39f80b8348fb754acd

Observation 5c44b633-0590-4036-a2e2-3fd8e01b96ee · outbound

This paper cites an unresolved cited work.

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:02:34.084103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-07T05:02:29.132482Z digest=sha256:06779f3ac38f8ecd4660f96a06a9324aa50aeeb22c6646568f77d77e56485780

Observation 1e40ed52-f436-4e95-88cd-94d4f9cc67fb · outbound

This paper cites Spider2-V: How Far Are Multimodal Agents From Automating Data Science and Engineering Workflows?.

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities Spider2-V: How Far Are Multimodal Agents From Automating Data Science and Engineering Workflows?

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T05:02:29.136593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:02:29.136593Z digest=sha256:1d26ad51b9743bf10c0e9748aac3a4cc750828a13b329903b43cae1ad6efa410

Observation c585176a-3b37-4fc7-ad61-f9bb97bedd82 · outbound

This paper cites Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks.

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:02:33.919293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-07T05:02:29.140682Z digest=sha256:4064edaaee4cb99ac174e11cb50d13f7688e1bc2b94b1cf5fca42c151337edf2

Observation e72355a2-aa45-4aa6-83bc-121b7dbb4d8f · outbound

This paper cites SeeClick: Harnessing GUI Grounding for Advanced Visual GUI Agents.

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities SeeClick: Harnessing GUI Grounding for Advanced Visual GUI Agents

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T05:02:29.144415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:02:29.144415Z digest=sha256:c8b692bbf9804afc4ad63a1974221d921f6151568966212e241789e04ef1ffb9

Observation 9d9c051e-a413-4673-9b0f-7ebe379f7d5b · outbound

This paper cites Mind2web: Towards a generalist agent for the web.

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities Mind2web: Towards a generalist agent for the web

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:02:33.734383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-07T05:02:29.148148Z digest=sha256:4904e683d113eea92c2399afee82466451e9c620a3f648fb89a3884832ccb7fa

Observation 67c688ee-3875-47df-a4da-3c9fce6c7863 · outbound

This paper cites Dysen-vdm: Empowering dynamics-aware text-to-video diffusion with llms.

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities Dysen-vdm: Empowering dynamics-aware text-to-video diffusion with llms

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:02:33.528030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-07T05:02:29.152268Z digest=sha256:90ed21352d55687a60d136a5e251c615202ae6e51a38e301bae914f6f64ae5b2

Observation 50aed938-2aa1-4eab-a811-b90eb574a63f · outbound

This paper cites Video-of-thought: Step-by-step video reasoning from perception to cognition.

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities Video-of-thought: Step-by-step video reasoning from perception to cognition

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:02:33.349259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-07T05:02:29.155750Z digest=sha256:7a2e5d43b5f5cb4a2bfe8ad9901e7df371020d040a08ef238abf3f404523be4f

Observation 1bf59cfd-e644-4886-8d98-1efac1ee8971 · outbound

This paper cites Vitron: A unified pixel-level vision llm for understanding, generating, segmenting, editing.

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities Vitron: A unified pixel-level vision llm for understanding, generating, segmenting, editing

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:02:33.167945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-07T05:02:29.159175Z digest=sha256:da6639308d4b34f4fc5d580576c3d215ac5306d02a8ac3fb33951b7c3b9f7ca1

Observation 5cc1e063-7d1b-4e3b-8cfe-2036738d51d2 · outbound

This paper cites Enhancing video-language representations with structural spatio-temporal alignment.

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities Enhancing video-language representations with structural spatio-temporal alignment

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:02:32.995079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-07T05:02:29.163035Z digest=sha256:e66ddb43b8ea0bc1d8d63063fd5873b15d1a218b5bf9a246fb1e398a7daef07b

Observation 0fcf5a49-4059-4727-be64-f8e5bbd3ba15 · outbound

This paper cites Generalist Virtual Agents: A Survey on Autonomous Agents Across Digital Platforms.

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities Generalist Virtual Agents: A Survey on Autonomous Agents Across Digital Platforms

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T05:02:29.166735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:02:29.166735Z digest=sha256:316f62537e8fc6a0dc84fc460954cb7ab16f9772487d5eb38b1c865bd75100f7

Observation b3fe1d08-9cfd-4472-b003-c31dc1ce5fa1 · outbound

This paper cites De-fine: Decomposing and refining visual programs with auto-feedback.

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities De-fine: Decomposing and refining visual programs with auto-feedback

Reference 12

Resolution
metadata mismatch
raw_fallback, observed 2026-08-07T05:02:30.062241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-07T05:02:29.170447Z digest=sha256:8953a1516ee5ddc9406f7760426bb6dbf746098e824279102b78c894107c27c5

Observation 37941295-095e-4a19-ba50-e32226ef34bb · outbound

This paper cites Benchmarking Multimodal CoT Reward Model Stepwise by Visual Program.

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities Benchmarking Multimodal CoT Reward Model Stepwise by Visual Program

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T05:02:29.173944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:02:29.173944Z digest=sha256:0ab861be536410c5636f40fa3b72d54185a7df9e17822eaa51606b9fcf0be854

Observation 47c6f886-57fb-4877-b41b-63b287dd429e · outbound

This paper cites Iris: Breaking GUI Complexity with Adaptive Focus and Self-Refining.

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities Iris: Breaking GUI Complexity with Adaptive Focus and Self-Refining

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T05:02:29.178376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:02:29.178376Z digest=sha256:1c72eda35da2d61e9277d60844a3c3fcd93c6444b517906b974e9af6d0e2602a

Observation fe01b38f-2a86-416b-8679-b17ba9bfccf1 · outbound

This paper cites Navigating the Digital World as Humans Do: Universal Visual Grounding for GUI Agents.

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities Navigating the Digital World as Humans Do: Universal Visual Grounding for GUI Agents

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T05:02:29.182555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:02:29.182555Z digest=sha256:61aea85e87c5331a252a639c9800b436f26e762e1eba4a26e22b3cfedaf1b901

Observation 4b378d0a-8832-4d3f-9430-80b301def70a · outbound

This paper cites Cogagent: A visual language model for gui agents.

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities Cogagent: A visual language model for gui agents

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:02:32.731848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-07T05:02:29.186253Z digest=sha256:06dde0add800abb14733ce8ec53586a06146d87b910658212087fadb07766f65

Observation 6eeeba8a-e01f-4a5f-962a-9c4bf82099da · outbound

This paper cites The Dawn of GUI Agent: A Preliminary Case Study with Claude 3.5 Computer Use.

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities The Dawn of GUI Agent: A Preliminary Case Study with Claude 3.5 Computer Use

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T05:02:29.189552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:02:29.189552Z digest=sha256:71e59534c27f615aab914cd7aced758fb7b7e280425cad4e91516f3b2a7f1b79

Observation dc06c457-3cfb-4c4f-a88c-688f5b10b629 · outbound

This paper cites GPT-4o System Card.

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities GPT-4o System Card

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T05:02:29.193533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:02:29.193533Z digest=sha256:576927b9dc73bea34b604b63e6a93f5cdcdd6b474383bc85c964c57c364d3e74

Observation 4e48966b-d59f-4071-9c24-86c6d746d836 · outbound

This paper cites P., Russak, M., Koh, J.

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities P., Russak, M., Koh, J

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:02:32.475848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-07T05:02:29.197603Z digest=sha256:e8b7596f11f8ccc814183f6fdecda1c7d0e480b18415e77f64f24d8a2502ac7d

Observation f121153e-fd89-45f5-93ad-6929d8b236cc · outbound

This paper cites VisualWebArena: Evaluating Multimodal Agents on Realistic Visual Web Tasks.

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities VisualWebArena: Evaluating Multimodal Agents on Realistic Visual Web Tasks

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T05:02:29.201086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:02:29.201086Z digest=sha256:3bcf3809afc74f8b768552510d9eb0ea5f3c3bcd65716aad504d96d1fe1747e0

Observation 09b81ed4-503d-4aec-9e09-16341fd9932b · outbound

This paper cites an unresolved cited work.

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:02:32.148264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-07T05:02:29.204810Z digest=sha256:17603e17604f8a0565f97ffdcdd4c385949682092a0c9a131ef55ec9ca533b43

Observation 876b291c-883c-4483-9127-a323e8e9a1f2 · outbound

This paper cites Fine-grained semantically aligned vision-language pre-training.

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities Fine-grained semantically aligned vision-language pre-training

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:02:31.782608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-07T05:02:29.208063Z digest=sha256:a37dab9f5715d6ffe701e1421b6dd8976aebb0448c110801aae6edd2bbd14e09

Observation 00160e87-6b8f-499d-9ed5-1656767e577f · outbound

This paper cites Fine-tuning multimodal llms to follow zero-shot demonstrative instructions.

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities Fine-tuning multimodal llms to follow zero-shot demonstrative instructions

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:02:31.509796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-07T05:02:29.211153Z digest=sha256:6a95852c34b71fc3b389c6f936f1a592eb61501fd789ce03908a9b32e3fe7568

Observation 3f872814-3c4a-4412-8a69-c117157f7ba5 · outbound

This paper cites Variational cross-graph reasoning and adaptive structured semantics learning for compositional temporal grounding.

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities Variational cross-graph reasoning and adaptive structured semantics learning for compositional temporal grounding

Reference 24

Resolution
metadata mismatch
raw_fallback, observed 2026-08-07T05:02:29.764865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-07T05:02:29.214658Z digest=sha256:f4599543fdc20800d3c15af279f135f0a6575bae6bd8cc769da5725f0931fbb7

Observation 0347396e-3fa2-4774-9a3d-65449bb06670 · outbound

This paper cites Mapping Natural Language Instructions to Mobile UI Action Sequences.

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities Mapping Natural Language Instructions to Mobile UI Action Sequences

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T05:02:29.217639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:02:29.217639Z digest=sha256:a297abc793f1502b3a04d5f3f2ceb4adb2855d1ea2cf58e8a6b5b0f8801e9832

Observation 83b75b66-3873-4d89-8e09-65c9ab7b1b8b · outbound

This paper cites ShowUI: One Vision-Language-Action Model for GUI Visual Agent.

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities ShowUI: One Vision-Language-Action Model for GUI Visual Agent

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T05:02:29.221294Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:02:29.221294Z digest=sha256:8b2fd707f880e7a419f200de076ffce9f8ab815e6486e95add41d8122c7d0632

Observation 089d71e4-e170-48f0-81f2-235ddf624bfd · outbound

This paper cites GUIOdyssey: A Comprehensive Dataset for Cross-App GUI Navigation on Mobile Devices.

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities GUIOdyssey: A Comprehensive Dataset for Cross-App GUI Navigation on Mobile Devices

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T05:02:29.228691Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:02:29.228691Z digest=sha256:b9b7b11bf78e9d481ab94962199931e28f4d61bd190f18265abfa041847080f8

Observation ae985b10-a123-412a-a6b7-b4d7fac1a85a · outbound

This paper cites Boosting Virtual Agent Learning and Reasoning: A Step-Wise, Multi-Dimensional, and Generalist Reward Model with Benchmark.

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities Boosting Virtual Agent Learning and Reasoning: A Step-Wise, Multi-Dimensional, and Generalist Reward Model with Benchmark

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T05:02:29.232405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:02:29.232405Z digest=sha256:c6a85708a3b55604b10fcd16f401d42de5f288214ab06aefed2c7bda752de901

Observation 35383744-47a4-4262-8a60-eea83d4ecbb3 · outbound

This paper cites Self-supervised Meta-Prompt Learning with Meta-Gradient Regularization for Few-shot Generalization.

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities Self-supervised Meta-Prompt Learning with Meta-Gradient Regularization for Few-shot Generalization

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T05:02:29.236247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:02:29.236247Z digest=sha256:1a36b1adb20606f1892b6ae572865a60d7c623df12d21eef5c7d13137ea63640

Observation 22983b81-c386-4972-8cf2-d3a86314cd39 · outbound

This paper cites Towards unified multimodal editing with enhanced knowledge collaboration.

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities Towards unified multimodal editing with enhanced knowledge collaboration

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:02:31.207365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-07T05:02:29.239831Z digest=sha256:f27a0d2d6a4299f5cad4f3443fe439832bbe04798aadf6f00099177da5ceeec9

Observation 84353a44-ec40-4be6-9285-b5c2ae7fab47 · outbound

This paper cites I3: I ntent-i ntrospective retrieval conditioned on i nstructions.

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities I3: I ntent-i ntrospective retrieval conditioned on i nstructions

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:02:30.908383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-07T05:02:29.243088Z digest=sha256:2cdda73a13235c29f9d12efa9e06b492132a15889fbd87466a66200fdcbe57ca

Observation 17e9e6e2-2b33-4e2e-8b5b-99c661c349f5 · outbound

This paper cites Auto-Encoding Morph-Tokens for Multimodal LLM.

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities Auto-Encoding Morph-Tokens for Multimodal LLM

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T05:02:29.246334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:02:29.246334Z digest=sha256:191f86662d2f7f72fc465923ec4cbbe35dce9f7d709f0ba967a4c13ef2ee4a08

Observation 9ae4e593-e82f-407f-bafa-0a8250e4ea4f · outbound

This paper cites Generative Multimodal Pretraining with Discrete Diffusion Timestep Tokens.

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities Generative Multimodal Pretraining with Discrete Diffusion Timestep Tokens

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T05:02:29.249982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:02:29.249982Z digest=sha256:500448e118243a9a93db384e9842864a3b460604ad89c03f22e3a40aad3ab2a1

Observation 2d9bacd5-bd51-44a3-be3e-09a722465bc2 · outbound

This paper cites Unlocking aha moments via reinforcement learning: Advancing collaborative visual comprehension and generation.

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities Unlocking aha moments via reinforcement learning: Advancing collaborative visual comprehension and generation

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T05:02:29.254238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:02:29.254238Z digest=sha256:8398e64ed788c6a3355e98238cade1872504ae7c806a86efafa5b754ec1fad50

Observation 4a7d1a39-3a00-4c29-a01b-fa13b471f726 · outbound

This paper cites Androidinthewild: A large-scale dataset for android device control.

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities Androidinthewild: A large-scale dataset for android device control

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:02:30.570748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-07T05:02:29.258032Z digest=sha256:1c61d7a5c1f5f00323587d9019d2ae12d960fb9c59dad164d9f27d1cd8ee4d64

Observation 215b55d8-7e85-4afa-aed2-eeb7a5b6b12d · outbound

This paper cites ScribeAgent: Towards Specialized Web Agents Using Production-Scale Workflow Data.

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities ScribeAgent: Towards Specialized Web Agents Using Production-Scale Workflow Data

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T05:02:29.261946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:02:29.261946Z digest=sha256:bbec2541b98556372d79f2ad58bca88b5ade6e64e8033106c7c42160cb415b33

Observation 06084d52-bd1f-4843-ba0b-22d00827e86e · outbound

This paper cites TaskBench: Benchmarking Large Language Models for Task Automation.

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities TaskBench: Benchmarking Large Language Models for Task Automation

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T05:02:29.265475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:02:29.265475Z digest=sha256:c3808f3cffd67791a889ab552ebcb3b9f765ca8b261d2e364844906a12ffe429

Observation dc08f448-c10b-4727-b331-adce39e1ed8a · outbound

This paper cites MobileAgentBench: An Efficient and User-Friendly Benchmark for Mobile LLM Agents.

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities MobileAgentBench: An Efficient and User-Friendly Benchmark for Mobile LLM Agents

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T05:02:29.269282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:02:29.269282Z digest=sha256:8f11876dbe90a4b75670f575f681e474d036fef82ccdfc6c65a24b29078df1ef

Observation 11c0cfee-b8d0-4148-a46d-5aec8bc47488 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T05:02:29.273032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:02:29.273032Z digest=sha256:1d5ba66c364ad7b9dfd7f897029663605323bc45b36d7e5be83a34d7b8cc872c

Observation 68938c66-c337-432b-a812-8693b38d13aa · outbound

This paper cites NE x T - GPT : Any-to-any multimodal LLM.

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities NE x T - GPT : Any-to-any multimodal LLM

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:02:30.239250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-07T05:02:29.276630Z digest=sha256:f96587651f99dc4b4a1fd9a4438cac6c461a1830d3c6e8878844fa894bd248b1

Observation 08be7f69-11fe-49c2-9877-76c527e51f68 · outbound

This paper cites OS-ATLAS: A Foundation Action Model for Generalist GUI Agents.

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities OS-ATLAS: A Foundation Action Model for Generalist GUI Agents

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T05:02:29.280843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:02:29.280843Z digest=sha256:e37682d7aa9f06438ace00d77be4280e8f4b14a55e5e3978ac00c3e7f3566d2e

Observation bb747de6-b1f3-453e-8a2f-ccdbc3835f84 · outbound

This paper cites OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments.

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T05:02:29.285064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:02:29.285064Z digest=sha256:4299da3eef549b3270083b39bc691eeed9a2e630caf879976eee02249c4f7f0f

Observation c5611f9f-14ea-4cb8-95f4-037ad2065bbb · outbound

This paper cites CRAB: Cross-environment Agent Benchmark for Multimodal Language Model Agents.

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities CRAB: Cross-environment Agent Benchmark for Multimodal Language Model Agents

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T05:02:29.289598Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:02:29.289598Z digest=sha256:2bc34bbe700950f0d8bb6b6c9ae086f3abf56cf4a38a74361cd29862985fdeed

Observation 9f40b8c8-a93a-4918-a518-b39157dca0fd · outbound

This paper cites Aguvis: Unified Pure Vision Agents for Autonomous GUI Interaction.

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities Aguvis: Unified Pure Vision Agents for Autonomous GUI Interaction

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T05:02:29.293289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:02:29.293289Z digest=sha256:ee9d628e120cd7ef916dbff5e36e1df3ab8bedb6d61108098b60ac41a827757e

Observation 6febf858-5e25-45ba-986e-e62d8cbc97f9 · outbound

This paper cites Webshop: Towards scalable real-world web interaction with grounded language agents.

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities Webshop: Towards scalable real-world web interaction with grounded language agents

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T05:02:29.297322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:02:29.297322Z digest=sha256:d01c43b7866023ac623c4e48baeaa989b56115bb3890d13378ee63ad47f743a0

Observation fd8ab3c2-cc57-474f-b695-c15d6a274724 · outbound

This paper cites WebArena: A Realistic Web Environment for Building Autonomous Agents.

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities WebArena: A Realistic Web Environment for Building Autonomous Agents

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T05:02:29.300721Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:02:29.300721Z digest=sha256:a7403629f30b69f013cc675d3d8aa60e0872fe61e2c15b9bd4cea258cea83e91

Pith citing papers

Observation 553be00a-773f-471d-b90d-b411981a531c · inbound

Mastering Collaborative Multi-modal Data Selection: A Focus on Informativeness, Uniqueness, and Representativeness cites this paper.

Mastering Collaborative Multi-modal Data Selection: A Focus on Informativeness, Uniqueness, and Representativeness What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T19:54:51.906676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:54:51.906676Z digest=sha256:1b1744162ea457878ac3a513bd9609834d732828409097567e9bea3216aee283

Observation 337ed338-7df9-4130-b2f5-f60bf7e91408 · inbound

DocOS: Towards Proactive Document-Guided Actions in GUI Agents cites this paper.

DocOS: Towards Proactive Document-Guided Actions in GUI Agents What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-20T11:18:13.823868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-20T11:13:48.529796Z digest=sha256:fa88bea5360ddc1cab6361e0633310f4ae80c2c165a9ddc5f6347073e7b86d9e