Pith. sign in

REVIEW 2 cited by

From Pixels to UI Actions: Learning to Follow Instructions via Graphical User Interfaces

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2306.00245 v2 pith:FMB22R3D submitted 2023-05-31 cs.LG cs.CLcs.CVcs.HC

classification cs.LGcs.CLcs.CVcs.HC
keywords agentsactionactionsdigitalgraphicalinterfacespixel-basedrepresentations
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Much of the previous work towards digital agents for graphical user interfaces (GUIs) has relied on text-based representations (derived from HTML or other structured data sources), which are not always readily available. These input representations have been often coupled with custom, task-specific action spaces. This paper focuses on creating agents that interact with the digital world using the same conceptual interface that humans commonly use -- via pixel-based screenshots and a generic action space corresponding to keyboard and mouse actions. Building upon recent progress in pixel-based pretraining, we show, for the first time, that it is possible for such agents to outperform human crowdworkers on the MiniWob++ benchmark of GUI-based instruction following tasks.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. OSWorld 2.0: Benchmarking Computer Use Agents on Long-Horizon Real-World Tasks

    cs.AI 2026-06 unverdicted novelty 7.0 of 10

    OSWorld 2.0 is a benchmark of 108 realistic long-horizon computer-use tasks where current agents achieve only 20.6% binary completion, struggling with state inference and constraint tracking.

  2. WebArXiv: Evaluating Multimodal Agents on Time-Invariant arXiv Tasks

    cs.IR 2025-07 conditional novelty 5.0 of 10

    WebArXiv is a time-invariant 275-task benchmark for multimodal web agents on arXiv, plus a dynamic-reflection prompting method that modestly improves success rates.

Pith tools