Pith. sign in

REVIEW 4 cited by

MobileVLM: A Vision-Language Model for Better Intra- and Inter-UI Understanding

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2409.14818 v2 pith:5WYMZP6N submitted 2024-09-23 cs.CL cs.AI

classification cs.CLcs.AI
keywords mobilevlmselementsinter-uilackmobilevlmpre-trainingunderstanding
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Recently, mobile AI agents based on VLMs have been gaining increasing attention. These works typically utilize VLM as a foundation, fine-tuning it with instruction-based mobile datasets. However, these VLMs are typically pre-trained on general-domain data, which often results in a lack of fundamental capabilities specific to the mobile domain. Therefore, they may struggle to recognize specific UI elements and understand intra-UI fine-grained information. In addition, the current fine-tuning task focuses on interacting with the most relevant element for the given instruction. These fine-tuned VLMs may still ignore the relationships between UI pages, neglect the roles of elements in page transitions and lack inter-UI understanding. To address issues, we propose a VLM called MobileVLM, which includes two additional pre-training stages to enhance both intra- and inter-UI understanding. We defined four UI-based pre-training tasks, enabling the model to better perceive fine-grained elements and capture page transition actions. To address the lack of mobile pre-training data, we built a large Chinese mobile dataset Mobile3M from scratch, which contains 3 million UI pages, and real-world transition actions, forming a directed graph structure. Experimental results show MobileVLM excels on both our test set and public mobile benchmarks, outperforming existing VLMs.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. MobiBench: Multi-Branch, Modular Benchmark for Mobile GUI Agents

    cs.AI 2025-12 conditional novelty 8.0 of 10

    MobiBench reaches near-human offline evaluation fidelity for mobile GUI agents by accepting any valid action at each step, and enables modular attribution of performance to agent components.

  2. Da Yu: Towards USV-Based Image Captioning for Waterway Surveillance and Scene Understanding

    cs.CV 2025-06 conditional novelty 6.0 of 10

    WaterCaption adds 20.2k waterway images with long, multi-region captions, and Da Yu with its Nano Transformer Adaptor produces competitive captions at a smaller computational cost.

  3. TransBench: Breaking Barriers for Transferable Graphical User Interface Agents in Dynamic Digital Environments

    cs.HC 2025-05 conditional novelty 6.0 of 10

    TransBench is a new benchmark of 1,459 screenshots and 22,000 grounding instructions for measuring how well GUI agents transfer across app versions, platforms, and applications.

  4. Agentic Services Computing

    cs.SE 2025-09 conditional novelty 5.0 of 10

    A position and survey paper that defines Agentic Services Computing, a lifecycle-based framework for engineering LLM agents as governed, first-class services.

Pith tools