Pith. sign in

Mobiledreamer: Generative sketch world model for gui agent

6 Pith papers cite this work. Polarity classification is still indexing.

6 Pith papers citing it

citation-role summary

background 1

citation-polarity summary

years

2026 6

roles

background 1

polarities

background 1

representative citing papers

A History-Aware Visually Grounded Critic for Computer Use Agents

cs.AI · 2026-06-09 · unverdicted · novelty 7.0

HiViG is a test-time critic that combines macro-action history summarization with visual grounding of execution coordinates to reduce short-sighted and visually erroneous actions in long-horizon GUI agents.

Qwen-AgentWorld: Language World Models for General Agents

cs.CL · 2026-06-23 · unverdicted · novelty 6.0

Qwen-AgentWorld are language world models that simulate multi-domain agent environments and boost general agent capabilities via decoupled RL simulation and unified foundation model training.

How Mobile World Model Guides GUI Agents?

cs.AI · 2026-05-11 · unverdicted · novelty 4.0 · 2 refs

World models trained on delta text, full text, diffusion images, and renderable code achieve SoTA on two benchmarks and improve downstream GUI agent performance on three mobile datasets with modality-specific strengths.

citing papers explorer

Showing 6 of 6 citing papers.

  • A History-Aware Visually Grounded Critic for Computer Use Agents cs.AI · 2026-06-09 · unverdicted · none · ref 36

    HiViG is a test-time critic that combines macro-action history summarization with visual grounding of execution coordinates to reduce short-sighted and visually erroneous actions in long-horizon GUI agents.

  • Qwen-AgentWorld: Language World Models for General Agents cs.CL · 2026-06-23 · unverdicted · none · ref 5

    Qwen-AgentWorld are language world models that simulate multi-domain agent environments and boost general agent capabilities via decoupled RL simulation and unified foundation model training.

  • VisCritic: Visual State Comparison as Process Reward for GUI Agents cs.CV · 2026-06-23 · unverdicted · none · ref 2

    VisCritic uses visual comparison of pre- and post-action GUI screenshots via a Siamese vision transformer and Action-Aware Critic Head to provide process rewards, improving agent performance on benchmarks.

  • MIRAGE: Mobile Agents with Implicit Reasoning and Generative World Models cs.AI · 2026-06-03 · unverdicted · none · ref 23

    MIRAGE compresses explicit chain-of-thought into latent vectors and adds a generative world model to predict future interface states, matching explicit reasoning performance with 3-5x fewer tokens on Android benchmarks.

  • How Mobile World Model Guides GUI Agents? cs.AI · 2026-05-11 · unverdicted · none · ref 9 · 2 links

    World models trained on delta text, full text, diffusion images, and renderable code achieve SoTA on two benchmarks and improve downstream GUI agent performance on three mobile datasets with modality-specific strengths.

  • Agentic World Modeling: Foundations, Capabilities, Laws, and Beyond cs.AI · 2026-04-24 · conditional · none · ref 44 · 2 links

    A survey proposing a three-level capability taxonomy (L1 Predictor, L2 Simulator, L3 Evolver) for world models across physical, digital, social, and scientific domains.