Pith. sign in

REVIEW 14 cited by

GUI-World: A Video Benchmark and Dataset for Multimodal GUI-oriented Understanding

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.10819 v2 pith:5WLNJWNQ submitted 2024-06-16 cs.CV cs.AIcs.CL

classification cs.CVcs.AIcs.CL
keywords understandingvideocontentllmsdatasetdynamicgui-orientedagents
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Recently, Multimodal Large Language Models (MLLMs) have been used as agents to control keyboard and mouse inputs by directly perceiving the Graphical User Interface (GUI) and generating corresponding commands. However, current agents primarily demonstrate strong understanding capabilities in static environments and are mainly applied to relatively simple domains, such as Web or mobile interfaces. We argue that a robust GUI agent should be capable of perceiving temporal information on the GUI, including dynamic Web content and multi-step tasks. Additionally, it should possess a comprehensive understanding of various GUI scenarios, including desktop software and multi-window interactions. To this end, this paper introduces a new dataset, termed GUI-World, which features meticulously crafted Human-MLLM annotations, extensively covering six GUI scenarios and eight types of GUI-oriented questions in three formats. We evaluate the capabilities of current state-of-the-art MLLMs, including Image LLMs and Video LLMs, in understanding various types of GUI content, especially dynamic and sequential content. Our findings reveal that current models struggle with dynamic GUI content without manually annotated keyframes or operation history. On the other hand, Video LLMs fall short in all GUI-oriented tasks given the sparse GUI video dataset. Therefore, we take the initial step of leveraging a fine-tuned Video LLM, GUI-Vid, as a GUI-oriented assistant, demonstrating an improved understanding of various GUI tasks. However, due to the limitations in the performance of base LLMs, we conclude that using video LLMs as GUI agents remains a significant challenge. We believe our work provides valuable insights for future research in dynamic GUI content understanding. All the dataset and code are publicly available at: https://gui-world.github.io.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 14 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Do GUI Agents Believe Their Eyes? Diagnosing State-Belief Reliance on Pixels versus Structure

    cs.AI 2026-07 accept novelty 7.0 of 10

    Across four open-weight models and three OpenAI models, textual state beliefs of GUI agents follow a conflicting structured value on 30–75% of probes where the same model reads the pixels correctly, so agent state bel...

  2. Chain-of-Models: Cross-Model Auditing for Bias-Robust LLM Judges

    cs.CL 2026-05 conditional novelty 6.0 of 10

    Per-bias selection of a cross-family LLM auditor lifts biased-judgment accuracy from 0.805/0.824 baselines to 0.884.

  3. WebMMU: A Benchmark for Multimodal Multilingual Website Understanding and Code Generation

    cs.CV 2025-08 conditional novelty 6.0 of 10

    WebMMU introduces a multilingual, three-task benchmark for website understanding and code generation, and finds current MLLMs underperform on reasoning, grounding, and functional code editing.

  4. Chain-of-Memory: Enhancing GUI Agents for Cross-Application Navigation

    cs.AI 2025-06 conditional novelty 6.0 of 10

    Explicitly storing short-term and long-term text memories, instead of raw screenshots, improves GUI agent accuracy on cross-app tasks, and a new annotated dataset helps 7B models approach 72B-level memory use.

  5. TransBench: Breaking Barriers for Transferable Graphical User Interface Agents in Dynamic Digital Environments

    cs.HC 2025-05 conditional novelty 6.0 of 10

    TransBench is a new benchmark of 1,459 screenshots and 22,000 grounding instructions for measuring how well GUI agents transfer across app versions, platforms, and applications.

  6. Scalable Video-to-Dataset Generation for Cross-Platform Mobile Agents

    cs.CV 2025-05 conditional novelty 6.0 of 10

    MONDAY automatically converts 20,000 YouTube mobile instructional videos into 313K annotated navigation frames and reports that models trained on it generalize better on an unseen mobile OS.

  7. Media Content Atlas: A Pipeline to Explore and Investigate Multidimensional Media Space using Multimodal LLMs

    cs.HC 2025-04 conditional novelty 6.0 of 10

    A pipeline combining CLIP, LLaVA, and BERTopic clusters and retrieves content from 1.12 million smartphone screenshots, with expert-rated topic relevance of 96%.

  8. Software Engineering for and with GUI Agent

    cs.SE 2026-08 conditional novelty 5.0 of 10

    A survey of 336 GUI-agent papers finds rapid growth alongside weak engineering support for recovery, human oversight, maintainability, and privacy, and calls for lifecycle-centered testing and governance.

  9. StepX-Edge: An On-Device UI Vision-Language Model via Architecture-Training-Deployment Co-Design

    cs.CV 2026-07 conditional novelty 5.0 of 10

    A co-designed 0.9B UI vision-language model tops small-model benchmarks on screen Q&A and Chinese OCR and runs on a Snapdragon 8 Gen5 at 98 tokens/second.

  10. Benchmarking Vision, Language, & Action Models in Procedurally Generated, Open Ended Action Environments

    cs.CV 2025-05 conditional novelty 5.0 of 10

    MultiNet v0.2 shows that current vision-language and vision-language-action models achieve near-chance zero-shot next-action accuracy on Procgen and collapse onto a small set of preferred actions.

  11. Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models

    cs.CV 2025-05 conditional novelty 4.0 of 10

    A survey organizes multimodal reasoning research into a staged roadmap and proposes native large multimodal reasoning models that unify perception, generation, and agentic planning.

  12. Artificial Intelligence in Spectroscopy: Advancing Chemistry from Prediction to Generation and Beyond

    cs.AI 2025-02 conditional novelty 4.0 of 10

    A unified review of AI applications in spectroscopy, organizing forward and inverse tasks across MS, NMR, IR, Raman, and UV-Vis, with a curated resource repository.

  13. Large language models for artificial general intelligence (AGI): A survey of foundational principles and approaches

    cs.AI 2025-01 conditional novelty 3.0 of 10

    This survey argues that embodiment, symbol grounding, causality, and memory are the foundational principles needed to make large language models achieve artificial general intelligence.

  14. Generalist Virtual Agents: A Survey on Autonomous Agents Across Digital Platforms

    cs.MA 2024-11 conditional novelty 3.0 of 10

    A survey that proposes the Generalist Virtual Agent concept and taxonomies for agent environments, tasks, perceptions, actions, models, and evaluation, concluding that real-world-like environments favor human-like int...

Pith tools