Pith. sign in

REVIEW 3 major objections 6 minor 82 references

Vision-language judges of computer-using agents systematically accept failed runs as successes, and only expensive frontier models hold up—until open reward models trained on a new human-gold benchmark match them at a fraction of the cost.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-31 02:16 UTC pith:GU67LITU

load-bearing objection Solid human-gold judge benchmark plus usable open reward models; the main soft spot is ensemble-distilled training labels, not the measurement claims. the 3 major comments →

arxiv 2607.28609 v1 pith:GU67LITU submitted 2026-07-30 cs.AI cs.CLcs.CV

OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models

classification cs.AI cs.CLcs.CV
keywords computer-using agentsVLM-as-judgereward modelstrajectory evaluationleniency biascross-platform GUI agentsOSRewardOS-Shepherd
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Computer-using agents leave long trajectories of screenshots, actions, and thoughts, and the field has started treating vision-language models as automatic judges of whether those runs actually completed the task. This paper shows that those judges are not yet trustworthy: even the strongest ones share a leniency bias, especially on hard false successes where the agent claims done but the environment never reached the goal. On a fresh cross-platform benchmark with human gold labels, frontier accuracy collapses from near 90% on the full set to below 70% on the hard subset, while cheap open models fare worse. The authors release that benchmark, a 100K reasoning-annotated training corpus, and open 9B/35B reward models that recover commercial-level judging at roughly 30–60× lower cost than the frontier and keep the de-biasing on held-out agent benchmarks. The practical stake is a reward signal that evaluation, data curation, and reinforcement learning can actually afford to run at scale.

Core claim

VLM judges of CUA trajectories fall short of an ideal judge and share one dominant failure mode—over-accepting incomplete tasks as successes—driven more by the agent’s text history than by the screenshots; the few judges accurate enough to trust cost too much for training-scale use, while open OS-Shepherd models trained on a new agreement-filtered corpus close most of that gap at 30–60× lower cost and transfer leniency resistance out of distribution.

What carries the argument

OSReward: a cross-platform benchmark of 1019 human-gold CUA trajectories (with Hard and Multi subsets) that measures judges themselves, plus OS-Shepherd-100K and the two-stage SFT+RL OS-Shepherd reward models that target false successes directly.

Load-bearing premise

That high-agreement labels from strong vision-language judges—after dropping ambiguous cases—are reliable enough training targets even though those judges herd and share the same leniency bias the models are meant to unlearn.

What would settle it

Hold OS-Shepherd and the frontier judges to a new batch of long-horizon desktop trajectories with independent multi-annotator human gold, and check whether fail-recall on false successes stays near the paper’s ~60% hard-set level or collapses back toward the untuned base’s near-zero catch rate.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Training-time CUA reward no longer has to be bought only at frontier API prices; a small self-hostable judge can sit near the cost–accuracy frontier.
  • CUA evaluation and data filtering should treat false successes as the primary error mode and measure fail-recall, not only binary accuracy.
  • Judging prompts and reward pipelines should keep full action/thought text; screenshot count and click markers move aggregate accuracy little.
  • A single learned judge can begin to replace per-task human-written verifiers across mobile, web, and desktop once de-biasing transfers.
  • Fine-grained alignment and efficiency grading remain much weaker than binary outcome judging and need separate calibration work.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • If text narrative dominates the verdict, agents that learn to write confident closing claims could systematically game reward models unless screen-verification is forced in the objective.
  • Agreement filtering may quietly discard the hardest genuine failures, so the next corpus gains may come from adversarial mining of judge disagreement rather than more unanimous easy labels.
  • The platform gap (desktop hardest, mobile easiest) suggests reward-model progress will stall on long GUI+CLI desktop runs unless those are overweighted in training.
  • Soft, confidence-weighted labels look more promising than majority vote, since judges herd on the same hard cases.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. The paper introduces OSReward, a cross-platform benchmark of 1019 human-gold CUA trajectories for evaluating VLM judges, plus OSReward-Hard and OSReward-Multi. Across 27 judges it reports a shared leniency bias (over-accepting incomplete tasks), a sharp accuracy drop on Hard (best ~70%), and a cost–accuracy frontier on which only expensive frontier models remain usable. It then releases OS-Shepherd-100K (ensemble-labeled, agreement-filtered) and trains open 9B/35B reward models that approach mid-tier commercial accuracy at much lower cost and transfer fail-recall gains to OSWorld, WebArena, and AndroidWorld. Primary measurement claims rest on multi-stage human labels; the open models rest on distilled VLM-ensemble supervision plus a short GRPO stage targeting false successes.

Significance. If the human-gold measurements hold, the paper supplies the first standardized, multi-platform stress test of CUA trajectory judges and documents a concrete, shared failure mode (narrative-driven false successes) that matters for evaluation, curation, and RL. The released benchmark, reasoning-annotated corpus, and self-hostable 9B/35B checkpoints are immediately usable artifacts; the cost–accuracy framing and held-out transfer results strengthen the case that reliable CUA reward need not be frontier-priced. The human pipeline (peer-screened instructions, diverse agent backbones, triple annotation plus meta-review, Hard re-verification, reported agreement) is a genuine strength relative to reused-trajectory judge studies.

major comments (3)
  1. [§6.1–6.2] §6.1–6.2 and Fig. 10/14: OS-Shepherd-100K labels are produced by the same VLM-judge class shown in §4.3 and §5.3 to herd and share leniency, with agreement filtering and a short GRPO pass on mined false successes as the main mitigations. This does not undermine the human-gold measurement claims on OSReward, but it is load-bearing for any claim that the open models are independently reliable rather than distilled ensemble imitators. The manuscript should state this limitation more explicitly (e.g., in §6.3 or the conclusion), quantify residual disagreement between the ensemble label and a human audit on a held-out training subsample if feasible, and avoid language that equates ensemble agreement with ground truth.
  2. [§7, Fig. 11] §7 / Fig. 11: Transfer is reported as agreement with each benchmark’s human-written verifier, which the paper correctly notes are imperfect (citing Xie et al., 2025). The text should more clearly separate “agreement with verifier” from “accuracy,” and, where possible, break out false-positive vs. false-negative shifts so readers can see that the transferred gain is specifically fail-catching rather than overall verifier mimicry. This is especially important on OSWorld, where §E.5 already shows FPs dominate judge errors.
  3. [Fig. 2, §5.4] Fig. 2, §5.4, §E.1: Cost figures mix official API list prices with “May 2026 market rates” for open-weight models and report 30–60× savings vs. frontier. The comparison is directionally convincing for OS-Shepherd-9B vs. Opus/GPT-5.5, but the paper should fix the pricing date/source table, state whether self-hosted GPU amortization is included in the “API-equivalent” $1.36 figure, and ensure the abstract/body multiplier is computed on a single, documented basket (full-set vs. Hard, same token assumptions).
minor comments (6)
  1. [Abstract] Abstract vs. body: the provided abstract text elsewhere says “30–60% lower cost” while the manuscript body and Fig. 2 claim “30–60×”; unify the multiplier everywhere.
  2. [§3.1] §3.1 / §A.4: Mobile environment description is duplicated (“The mobile environment is hosted in an Android emulator…” appears twice in close succession).
  3. [§4.5] §4.5 / Table 2: OSReward-Multi alignment is effectively two-level after removing the single 0-scored run (§B.4); state this in the main Multi discussion so macro-recall is not over-interpreted as three-class grading.
  4. [Table 1] Fig. 1 caption and Table 1: several model marketing names (GPT-5.5, Claude-Opus-4-8, etc.) will date quickly; keep API identifiers from Table 9 adjacent in the main results for reproducibility.
  5. [§5.2] §5.2 takeaway claims actions carry “roughly more signal” than CoT; the −1.8pp vs −7.2pp ablations support text>vision more cleanly than actions>CoT—soften the wording.
  6. [§3.4] Minor typos/style: “with detailed definitions in with detailed definitions in §B.4” (§3.4); “envel⌢pe” artifacts in the author block; ensure Hard-set N=284 platform counts in Fig. 8 match the released metadata.

Circularity Check

0 steps flagged

No significant circularity: primary claims rest on held-out multi-stage human-gold labels and external verifiers, not on quantities defined by the training ensemble.

full rationale

OSReward is an empirical measurement and systems paper, not a first-principles derivation. The load-bearing reliability claims (frontier judges share leniency; accuracy collapses on OSReward-Hard; OS-Shepherd matches commercial judges at much lower cost and transfers de-biasing) are scored against 1019 trajectories with multi-annotator human gold plus meta-review, fully disjoint from OS-Shepherd-100K. Training labels are agreement-filtered VLM ensemble judgments—a standard distillation setup with a known shared-bias risk—but the paper does not redefine success by those labels, does not fit a parameter on the evaluation set and call it a prediction, and does not import a self-authored uniqueness theorem to force the result. Held-out agreement with OSWorld/WebArena/AndroidWorld human-written verifiers and the reported base→SFT→RL fail-recall shift on human-gold Hard further separate measured performance from training-label construction. Self-citations (OS-Genesis, OpenMobile, infrastructure papers) supply data-collection methods and reused open trajectories, not the gold verdicts or the accuracy numbers. No step reduces a claimed prediction to its inputs by construction.

Axiom & Free-Parameter Ledger

4 free parameters · 4 axioms · 2 invented entities

Empirical systems paper. Load-bearing premises are methodological: human multi-annotator verdicts as gold; fixed judging protocol (last-N screenshots + text); ensemble agreement as scalable proxy labels; list/market API prices for cost claims. No physical constants or fitted scientific laws. Invented items are named datasets/models, not unobserved theoretical entities.

free parameters (4)
  • N trailing screenshots (default 5) = 5
    Main protocol fixes N=5; ablations show aggregate accuracy is flat in N but individual verdicts flip (§4.1, §5.1, Fig. 19).
  • Ensemble agreement filter threshold = ~85% of judged trajectories retained
    Trajectories enter OS-Shepherd-100K only under high cross-judge agreement; exact keep/drop rule shapes the training distribution (§6.1).
  • GRPO RL hyperparameters = lr=1e-6, KL=0.001, ~150 steps
    Learning rate 1e-6, KL 0.001, 8 rollouts, ~150 steps, reward 1.0/0.1/0.0 scheme relocate the operating point on false successes (§D.2).
  • API/market unit prices for cost frontier = e.g. OS-Shepherd-9B ~$1.36 per full-set pass
    Cost–accuracy claims (30–60×) depend on May-2026 list/market prices and average trajectory token mix (Fig. 2, §E.1).
axioms (4)
  • domain assumption Multi-stage human annotation (3-way + meta-review) yields ground-truth success/fail for CUA trajectories.
    Entire benchmark validity rests on this; stated in §3.3–3.4 and §B.
  • domain assumption A trajectory is FAIL if the agent did not obtain/verify the answer through the environment, even if the answer is factually correct.
    Explicit judging rule shared by annotators and prompt (§3.3, §C.2 grounding rule).
  • ad hoc to paper High-agreement votes among strong VLM judges are sufficiently clean training labels after dropping the ambiguous middle.
    Core scalable-supervision choice in §6.1; motivated by herding analysis in §5.3.
  • domain assumption Standard supervised fine-tuning plus GRPO improves judge calibration without needing new human labels.
    Training recipe §6.2 / §D.2; common RLHF/GRPO practice applied to this setting.
invented entities (2)
  • OSReward / OSReward-Hard / OSReward-Multi independent evidence
    purpose: Standardized human-gold evaluation of CUA trajectory judges across platforms and difficulty.
    Named benchmark artifacts constructed by the authors; falsifiable via released labels and re-annotation.
  • OS-Shepherd-100K and OS-Shepherd 9B/35B independent evidence
    purpose: Open training corpus and self-hostable reward models targeting false-success bias at low cost.
    Released models/data; performance claims are testable on OSReward and external benchmarks.

pith-pipeline@v1.2.0-daily-grok45 · 49881 in / 3310 out tokens · 60212 ms · 2026-07-31T02:16:21.366544+00:00 · methodology

0 comments
read the original abstract

Computer-using agents (CUAs) are advancing rapidly across the digital world. A CUA trajectory records the agent's actions, states, and reasoning. Verifying whether it fulfilled the task instruction is central to CUA evaluation, data curation, and reinforcement learning. Neither human-written verifiers nor human annotators can provide such verification at scale, so the field increasingly turns to vision-language models (VLMs) as judges of CUA trajectories. But a fundamental question has long gone unexamined: are these VLM judges reliable enough? To study it systematically, we introduce OSReward, a realistic, high-quality benchmark that evaluates VLM judges on CUA trajectories. The trajectories come from diverse agent backbones executing human-verified instructions across platforms, then rigorously labeled with ground-truth verdicts through multi-stage human annotation. Building on it, we derive OSReward-Hard, a challenge set concentrating genuinely hard cases, and OSReward-Multi for fine-grained efficiency and alignment scoring. The most comprehensive evaluation of VLM judges to date finds even state-of-the-art models fall short of an ideal judge, sharing a systematic leniency bias that mislabels failed runs as successes. The few reliable enough to trust are too expensive to run at scale, while affordable open models trail far behind. To close this gap, we construct and release OS-Shepherd-100K, an open corpus of reasoning-annotated trajectory judgments for the CUA community. On it, we train OS-Shepherd (9B and 35B), open reward models that supply low-cost, stable, and reliable reward signals, matching commercial judges at 30-60% lower cost than the frontier. Extensive analyses further inform the design of reliable CUA reward at scale. Our code, benchmark, dataset, and model checkpoints are available at https://os-copilot.github.io/OSReward-Home/.

Figures

Figures reproduced from arXiv: 2607.28609 by Ben Kao, Bowen Yang, Fangzhi Xu, Hang Yan, Jiahui Gao, Jialin Cao, Jianbing Zhang, Jingyang Gong, Kaiming Jin, Kanzhi Cheng, Liheng Chen, Lingpeng Kong, Nuo Chen, Qiushi Sun, Tianbao Xie, Xinfeng Yuan, Xingdong Gong, Yian Wang, Zehao Li, Zhangyue Yin, Zhiyong Wu, Zhoumianze Liu, Zichen Ding.

Figure 1
Figure 1. Figure 1: How well VLM judges score CUA trajectories (a), and the strict–lenient bias they share (b). Contact author(s): qiushisun@connect.hku.hk, chengkz@smail.nju.edu.cn, zjb@nju.edu.cn * Equal contribution † Project co-lead # Corresponding author arXiv:2607.28609v1 [cs.AI] 30 Jul 2026 [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: Cost against binary accuracy on OSReward-Hard: reliable judges are expensive, and the OS-Shepherd models come nearest their accuracy at a fraction of the cost. collection pipeline, process and screen its output, and curate it together with filtered public corpora into OS-Shepherd-100K, an open corpus of 100K reasoning-annotated trajectory judgments that spans more trajectories, agent backbones, scenarios, … view at source ↗
Figure 3
Figure 3. Figure 3: From realistic environments to raw trajectories: annotators prepare the environments and write grounded instructions on them, agents from four model families execute the instructions. 3. OSReward Our goal is a corpus of realistic, cross-platform CUA trajectories paired with gold verdicts trustworthy enough to measure the judges themselves. Reusing existing benchmarks’ rollouts cannot supply it: their runs … view at source ↗
Figure 4
Figure 4. Figure 4: The annotation pipeline. Each pre-filtered trajectory is labeled by three independent annotators; disagreements go to meta review, and the verified gold set is read as three views instruction is run by one to three of them, spanning the Claude, Gemini, Kimi, and Qwen families. Backbones differ in action idiom, thought verbosity, and failure modes, so spreading the rollouts across the pool keeps the benchma… view at source ↗
Figure 5
Figure 5. Figure 5: OSReward at a glance: outcome composition of the full and Hard sets, platform mix, and trajectory length; failed runs are markedly longer. latter two are nested subsets of the full set. The Full Set (OSReward). The full OSReward set holds 1019 trajectories spanning all four platforms, roughly balanced between successes (43%) and failures (57%) ( [PITH_FULL_IMAGE:figures/full_fig_p008_5.png] view at source ↗
Figure 6
Figure 6. Figure 6: Judge bias on OSReward (left) and OSReward-Hard (right). Judges below the diagonal skew lenient; most of the field sits there, and the skew widens on the hard set. 4.4. OSReward-Hard: The Challenge Set The Collapse. The aggregate ∼90% is optimistic. OSReward-Hard, the failure-heavy challenge set of §3.4 (30/70 success/fail; selection details in §E.1), drops every judge by 20–43 pp ( [PITH_FULL_IMAGE:figur… view at source ↗
Figure 7
Figure 7. Figure 7: Per-judge error composition, seven representative judges: over-accepting an incomplete task (warm colors) dominates every family. 0 10 20 30 40 50 60 Mean judge binary acc (%) windows (n=29) web (n=60) ubuntu (n=132) mobile (n=63) 42.4% 51.9% 52.1% 58.3% OSReward-Hard: accuracy by platform 0 10 20 30 40 50 60 Mean judge binary acc (%) Perception (n=51) Action (n=62) Planning (n=187) Memory (n=20) 41.6% 43.… view at source ↗
Figure 8
Figure 8. Figure 8: Mean per-judge binary accuracy on OSReward-Hard, by platform and by failure type. Failure types are multi-label, so their counts sum to more than the number of failed trajectories. Observations from the Hard Set. OSReward-Hard concentrates on the cases the shared lenient mode of § 4.3 misses, failed runs whose records read like completed tasks. On these deceptive cases, the leading judges finally separate,… view at source ↗
Figure 9
Figure 9. Figure 9: Overview of input ablations: Δ binary accuracy per (setting × model) vs. the main setting. This is the mechanism behind the leniency bias of §4.3. A judge that leans on the agent’s own narrative is exactly a judge that a confident closing claim can fool. This suggests a recipe for CUA reward-model labeling: keep the full text history, drop the marker, and set the screenshot count per model. Takeaway: Even … view at source ↗
Figure 10
Figure 10. Figure 10: The OS-Shepherd-100K pipeline: self-collected and open-source trajectories are ensemble￾judged and distilled into the training set. Band widths ∝ trajectory counts. kept trajectory is scored by an ensemble of strong VLM-as-a-Judge runs under the protocol of § 4.1, varying both the judge model (Gemini-3.1-Pro, Kimi-K2.5, Gemini-3-Pro, and others) and the screenshot setting across runs. Aggregating these vo… view at source ↗
Figure 11
Figure 11. Figure 11: Judges on three existing CUA benchmarks against each benchmark’s human-written verifier (matched subsets): accuracy and failure recall; means are computed before rounding. up to Qwen3.5-397B-A17B (∼44× the 9B’s size) on all three ( [PITH_FULL_IMAGE:figures/full_fig_p019_11.png] view at source ↗
Figure 12
Figure 12. Figure 12: The twenty most frequent task types among the roughly 32K web instructions sampled for collection, together about 65% of the pool. Fact-finding lookups dominate, followed by article, academic￾paper, and shopping tasks. A.2. Windows Environment. Our Windows environment extends existing works’ virtual machine setting (Bonatti et al., 2024) with about twenty everyday applications and common command-line tool… view at source ↗
Figure 13
Figure 13. Figure 13: Failure-type profile over OSReward’s fail trajectories (multi-label shares; the catch-all others tag is excluded). A single run can carry several tags. benchmark was checked for personally identifiable information, and none enters the release. B.3. Failure-Type Taxonomy Every trajectory judged fail is tagged by the annotators with one or more failure types (§3.3); a long run can accumulate several. Reason… view at source ↗
Figure 14
Figure 14. Figure 14: The de-biasing trajectory: base→ SFT→SFT+RL moves OS-Shepherd-9B from the lenient corner toward the balanced diagonal, with RL supplying the largest hard-set step. SFT RL (both sizes) 9B 35B-A3B Base model Qwen3.5-9B / Qwen3.5-35B-A3B 9B SFT ckpt 35B SFT ckpt Samples 96.6K 3.1K (shared) Rollouts / sample — 8 (at 𝑇=1.0, top-𝑝 1.0) Batch size — 16 Learning rate — 1e−6 KL to SFT ref. — 0.001 (low-variance, a… view at source ↗
Figure 15
Figure 15. Figure 15: gives the cost-accuracy view on the full set. Costs are official API list prices where available; open-weight models with no official pricing are charged at May-2026 market rates for models of similar size [PITH_FULL_IMAGE:figures/full_fig_p042_15.png] view at source ↗
Figure 16
Figure 16. Figure 16: Open-weight vs. closed-source binary accuracy (group means over the 27-judge field). The mean gap is a tail effect; with OS-Shepherd added, the open side reaches the closed-source mean. Macro-recall is the balanced accuracy of the emitted levels: threshold-dependent, it asks whether the scores are usable as they come out. AUC is the pairwise ranking accuracy (ties count 0.5): threshold-free, it asks wheth… view at source ↗
Figure 17
Figure 17. Figure 17: Pairwise agreement. (a) binary-verdict 𝜅; (b) 3-class alignment 𝜅, both hierarchically clustered on 1 − 𝜅. E.4. Inter-Judge Agreement [PITH_FULL_IMAGE:figures/full_fig_p044_17.png] view at source ↗
Figure 18
Figure 18. Figure 18: OSWorld deep dive (external benchmark), pooled over eight judges. Left: accuracy by application domain. Right: accuracy and false-positive rate vs. trajectory length [PITH_FULL_IMAGE:figures/full_fig_p045_18.png] view at source ↗
Figure 19
Figure 19. Figure 19: Binary accuracy vs. the number of trailing screenshots 𝑁. No judge trends with 𝑁; the shaded band marks the 𝑁 = 5–9 regime the main setting draws from. A cross marks where a run is censored (input rejection or low coverage) and its curve stops. 45 [PITH_FULL_IMAGE:figures/full_fig_p045_19.png] view at source ↗
Figure 20
Figure 20. Figure 20: Self-consistency at 𝑇=0.7 over five trials. (a) aggregate accuracy vs. greedy; (b) per-trajectory flip rate (lower = steadier labels). 46 [PITH_FULL_IMAGE:figures/full_fig_p046_20.png] view at source ↗
Figure 21
Figure 21. Figure 21: A recency-grounding failure: the selected documentary visibly dates to 2017, but the agent presents it as recently popular without supporting evidence. 47 [PITH_FULL_IMAGE:figures/full_fig_p047_21.png] view at source ↗
Figure 22
Figure 22. Figure 22: A false-success case in which self-narration overrides terminal evidence: the required git status command never appears and produces no visible output. TRAJECTORY EVIDENCE INSTRUCTION DESKTOP Navigate to "finance.yahoo.com" using the address bar, search for "NVDA" to find Nvidia Corporation’s stock data, and locate the current "Beta" value. STEP 0 STEP 7 NVDA result identified STEP 32 Statistics page reac… view at source ↗
Figure 23
Figure 23. Figure 23: A fine-grained perception failure on Yahoo Finance: after reaching NVDA’s Statistics page, the agent misreads both the Beta window and its displayed value. 48 [PITH_FULL_IMAGE:figures/full_fig_p048_23.png] view at source ↗
Figure 24
Figure 24. Figure 24: A fine-grained visual miss in Audacity: the agent follows convincing action and export cues without independently verifying whether the waveform satisfies the requested first-ten-second Fade Out. TRAJECTORY EVIDENCE INSTRUCTION DESKTOP In Krita, use the Crop Tool to remove the excess white space around the drawn question mark, then export the cropped image as ‘~/Desktop/reference.jpg‘. Next, open LibreCAD… view at source ↗
Figure 25
Figure 25. Figure 25: A long-horizon Ubuntu planning failure: repeated image-insertion errors are obscured by a successful final save, causing a visually incorrect LibreCAD document to be mistaken for a completed artifact. 49 [PITH_FULL_IMAGE:figures/full_fig_p049_25.png] view at source ↗
Figure 26
Figure 26. Figure 26: A successful long-horizon desktop case requiring sustained state tracking and verification of the final spreadsheet. 50 [PITH_FULL_IMAGE:figures/full_fig_p050_26.png] view at source ↗
Figure 27
Figure 27. Figure 27: A mobile hard case requiring the judge to verify the nearest open hospital and a fuel stop in the final route. TRAJECTORY EVIDENCE INSTRUCTION WEB Find a hotel in Rome under 180 euros per night for next weekend. Then search for Italian restaurants within walking distance of that hotel with ratings above 4 stars. STEP 0 STEP 20 Exact hotel ad￾dress recovered STEP 26 Yelp anti-bot block .. . .. . .. . .. . … view at source ↗
Figure 28
Figure 28. Figure 28: A successful Web hard case requiring cross-site constraint tracking and grounded recovery from anti-bot blocks. 51 [PITH_FULL_IMAGE:figures/full_fig_p051_28.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

82 extracted references · 1 canonical work pages

  1. [1]

    Introducing Claude Haiku 4.5

    Anthropic. Introducing Claude Haiku 4.5. https://www.anthropic.com/news/ claude-haiku-4-5, October 2025

  2. [2]

    Introducing Claude Opus 4.6

    Anthropic. Introducing Claude Opus 4.6. https://www.anthropic.com/news/ claude-opus-4-6, February 2026a

  3. [3]

    Introducing Claude Opus 4.8

    Anthropic. Introducing Claude Opus 4.8. https://www.anthropic.com/news/ claude-opus-4-8, May 2026b

  4. [4]

    Introducing Claude Sonnet 4.6

    Anthropic. Introducing Claude Sonnet 4.6. https://www.anthropic.com/news/ claude-sonnet-4-6, February 2026c

  5. [5]

    Digirl: Training in-the-wild device-control agents with autonomous reinforcement learning, 2024

    Hao Bai, Yifei Zhou, Mert Cemri, Jiayi Pan, Alane Suhr, Sergey Levine, and Aviral Kumar. Digirl: Training in-the-wild device-control agents with autonomous reinforcement learning, 2024. URL https://arxiv.org/abs/2406.11896

  6. [6]

    Qwen3-vl technical report.arXiv preprint arXiv:2511.21631, 2025

    Shuai Bai, Yuxuan Cai, Ruizhe Chen, Keqin Chen, Xionghui Chen, Zesen Cheng, Lianghao Deng, Wei Ding, Chang Gao, Chunjiang Ge, et al. Qwen3-vl technical report.arXiv preprint arXiv:2511.21631, 2025

  7. [7]

    Windowsagentarena: Evaluating multi-modal os agents at scale, 2024

    Rogerio Bonatti, Dan Zhao, Francesco Bonacci, Dillon Dupont, Sara Abdali, Yinheng Li, Yadong Lu, JustinWagle,KazuhitoKoishida,ArthurBucker,LawrenceJang,andZackHui. Windowsagentarena: Evaluating multi-modal os agents at scale, 2024. URLhttps://arxiv.org/abs/2409.08264

  8. [8]

    Seed2.0 model card: Towards intelligence frontier for real-world complexity

    ByteDance Seed Team. Seed2.0 model card: Towards intelligence frontier for real-world complexity. Technical report, ByteDance, February 2026

  9. [9]

    Web-shepherd: Advancing PRMs for reinforcing web agents

    Hyungjoo Chae, Sunghwan Kim, Junhee Cho, Seungone Kim, Seungjun Moon, Gyeom Hwangbo, Dongha Lim, Minjin Kim, Yeonjun Hwang, Minju Gwak, Dongwook Choi, Minseok Kang, Gwanhoon Im, ByeongUng Cho, Hyojun Kim, Jun Hee Han, Taeyoon Kwon, Minju Kim, Beong woo Kwak, Dongjin Kang, and Jinyoung Yeo. Web-shepherd: Advancing PRMs for reinforcing web agents. InThe Thi...

  10. [10]

    Gui-shepherd: Reliableprocessrewardand verification for long-sequence gui tasks, 2025a

    Cong Chen, Kaixiang Ji, Hao Zhong, Muzhi Zhu, Anzhou Li, Guo Gan, Ziyuan Huang, Cheng Zou, JiajiaLiu,JingdongChen,HaoChen,andChunhuaShen. Gui-shepherd: Reliableprocessrewardand verification for long-sequence gui tasks, 2025a. URLhttps://arxiv.org/abs/2509.23738

  11. [11]

    Mllm-as-a-judge: Assessing multimodal llm-as-a-judge with vision-language benchmark, 2024

    Dongping Chen, Ruoxi Chen, Shilin Zhang, Yinuo Liu, Yaochen Wang, Huichi Zhou, Qihui Zhang, Yao Wan, Pan Zhou, and Lichao Sun. Mllm-as-a-judge: Assessing multimodal llm-as-a-judge with vision-language benchmark, 2024. URLhttps://arxiv.org/abs/2402.04788. 20 OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models

  12. [12]

    Os-map: How far can computer-using agents go in breadth and depth?arXiv preprint arXiv:2507.19132, 2025b

    Xuetian Chen, Yinghao Chen, Xinfeng Yuan, Zhuo Peng, Lu Chen, Yuekeng Li, Zhoujia Zhang, Yingqian Huang, Leyan Huang, Jiaqing Liang, et al. Os-map: How far can computer-using agents go in breadth and depth?arXiv preprint arXiv:2507.19132, 2025b

  13. [13]

    SeeClick: Harnessing GUI grounding for advanced visual GUI agents

    Kanzhi Cheng, Qiushi Sun, Yougang Chu, Fangzhi Xu, Li YanTao, Jianbing Zhang, and Zhiyong Wu. SeeClick: Harnessing GUI grounding for advanced visual GUI agents. InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 9313–9332, Bangkok, Thailand, August 2024. Association for Computational Li...

  14. [14]

    Openmobile: Building open mobile agents with task and trajectory synthesis.arXiv preprint arXiv:2604.15093, 2026

    Kanzhi Cheng, Zehao Li, Zheng Ma, Nuo Chen, Jialin Cao, Qiushi Sun, Zichen Ding, Fangzhi Xu, Hang Yan, Jiajun Chen, et al. Openmobile: Building open mobile agents with task and trajectory synthesis.arXiv preprint arXiv:2604.15093, 2026

  15. [15]

    Training verifiers to solve math word problems, 2021

    Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, Christopher Hesse, and John Schulman. Training verifiers to solve math word problems, 2021. URLhttps://arxiv.org/ abs/2110.14168

  16. [16]

    A coefficient of agreement for nominal scales.Educational and Psychological Measure- ment, 20(1):37–46, 1960

    Jacob Cohen. A coefficient of agreement for nominal scales.Educational and Psychological Measure- ment, 20(1):37–46, 1960

  17. [17]

    Gemini 2.5: Pushing the frontier with advanced reasoning, multimodality, long context, and next generation agentic capabilities

    Gheorghe Comanici, Eric Bieber, Mike Schaekermann, Ice Pasupat, Noveen Sachdeva, Inderjit Dhillon, Marcel Blistein, Ori Ram, Dan Zhang, Evan Rosen, et al. Gemini 2.5: Pushing the frontier with advanced reasoning, multimodality, long context, and next generation agentic capabilities. arXiv preprint arXiv:2507.06261, 2025

  18. [18]

    Gemini 3 pro model card, November 2025

    Gemini Team. Gemini 3 pro model card, November 2025. URLhttps://deepmind.google/ models/gemini/

  19. [19]

    Navigating the digital world as humans do: Universal visual grounding for GUI agents

    Boyu Gou, Ruohan Wang, Boyuan Zheng, Yanan Xie, Cheng Chang, Yiheng Shu, Huan Sun, and Yu Su. Navigating the digital world as humans do: Universal visual grounding for GUI agents. InThe Thirteenth International Conference on Learning Representations, 2025. URLhttps: //openreview.net/forum?id=kxnoqaisCT

  20. [20]

    GPT-4o system card.arXiv preprint arXiv:2410.21276, 2024

    Aaron Hurst, Adam Lerer, Adam P Goucher, Adam Perelman, Aditya Ramesh, Aidan Clark, AJ Os- trow, Akila Welihinda, Alan Hayes, Alec Radford, et al. GPT-4o system card.arXiv preprint arXiv:2410.21276, 2024

  21. [21]

    AgentStore: Scalable integration of heterogeneous agents as specialized generalist computer assistant

    Chengyou Jia, Minnan Luo, Zhuohang Dang, Qiushi Sun, Fangzhi Xu, Junlin Hu, Tianbao Xie, and Zhiyong Wu. AgentStore: Scalable integration of heterogeneous agents as specialized generalist computer assistant. InFindings of the Association for Computational Linguistics: ACL 2025, pages 8908–8934, Vienna, Austria, July 2025. Association for Computational Lin...

  22. [22]

    TreeCUA: Efficiently scaling GUI automation with tree-structured verifiable evolution

    Deyang Jiang, Jing Huang, Xuanle Zhao, Lei Chen, Liming Zheng, Fanfan Liu, Haibo Qiu, Peng Shi, and Zhixiong Zeng. TreeCUA: Efficiently scaling GUI automation with tree-structured verifiable evolution. InForty-third International Conference on Machine Learning, 2026. URL https:// openreview.net/forum?id=KBCWS6OBnD. 21 OSReward: Instituting Standardized Ev...

  23. [23]

    Kimi Team, Tongtong Bai, Yifan Bai, Yiping Bao, SH Cai, Yuan Cao, Y Charles, HS Che, Cheng Chen, Guanduo Chen, et al. Kimi k2. 5: Visual agentic intelligence.arXiv preprint arXiv:2602.02276, 2026

  24. [24]

    Mobileworld: Benchmarking autonomous mobile agents in agent-user interactive and mcp-augmented environments

    Quyu Kong, Xu Zhang, Zhenyu Yang, Nolan Gao, Chen Liu, Panrong Tong, Chenglin Cai, Hanzhang Zhou, Jianan Zhang, Liangyu Chen, et al. Mobileworld: Benchmarking autonomous mobile agents in agent-user interactive and mcp-augmented environments. InProceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ...

  25. [25]

    Computing Krippendorff’s alpha-reliability

    Klaus Krippendorff. Computing Krippendorff’s alpha-reliability. Technical report, University of Pennsylvania, Annenberg School for Communication, 2011

  26. [26]

    Vl-rewardbench: a challenging benchmark for vision-language generative reward models

    Lei Li, Yuancheng Wei, Zhihui Xie, Xuqing Yang, Yifan Song, Peiyi Wang, Chenxin An, Tianyu Liu, Sujian Li, Bill Yuchen Lin, et al. Vl-rewardbench: a challenging benchmark for vision-language generative reward models. InProceedings of the Computer Vision and Pattern Recognition Conference, pages 24657–24668, 2025

  27. [27]

    Os- themis: A scalable critic framework for generalist gui rewards, 2026

    Zehao Li, Zhenyu Wu, Yibo Zhao, Bowen Yang, Jingjing Xie, Zhaoyang Liu, Zhoumianze Liu, Kaiming Jin, Jianze Liang, Zonglin Li, Feng Wu, Bowen Zhou, Zun Wang, and Zichen Ding. Os- themis: A scalable critic framework for generalist gui rewards, 2026. URLhttps://arxiv.org/ abs/2603.19191

  28. [28]

    Cuarewardbench: A benchmark for evaluating reward models on computer-using agent, 2025

    Haojia Lin, Xiaoyu Tan, Yulei Qin, Zihan Xu, Yuchen Shi, Zongyi Li, Gang Li, Shaofei Cai, Siqi Cai, Chaoyou Fu, Ke Li, and Xing Sun. Cuarewardbench: A benchmark for evaluating reward models on computer-using agent, 2025. URLhttps://arxiv.org/abs/2510.18596

  29. [29]

    ScaleCUA: Scaling open- source computer use agents with cross-platform data

    Zhaoyang Liu, JingJing Xie, Zichen Ding, Zehao Li, Bowen Yang, Zhenyu Wu, Xuehui Wang, Qiushi Sun, Shi Liu, Weiyun Wang, Shenglong Ye, Qingyun Li, Zeyue Tian, Gen Luo, Xiangyu Yue, Biqing Qi, Kai Chen, Bowen Zhou, Yu Qiao, Qifeng Chen, and Wenhai Wang. ScaleCUA: Scaling open- source computer use agents with cross-platform data. InThe Fourteenth Internatio...

  30. [30]

    Pal, and Siva Reddy

    Xing Han Lù, Amirhossein Kazemnejad, Nicholas Meade, Arkil Patel, Dongchan Shin, Alejandra Zambrano, Karolina Stańczak, Peter Shaw, Christopher J. Pal, and Siva Reddy. Agentrewardbench: Evaluating automatic evaluations of web agent trajectories, 2025. URLhttps://arxiv.org/ abs/2504.08942

  31. [31]

    Computer-using agent: Introducing a universal interface for ai to interact with the digital world, 2025a

    OpenAI. Computer-using agent: Introducing a universal interface for ai to interact with the digital world, 2025a. URLhttps://openai.com/index/computer-using-agent

  32. [32]

    GPT-5 System Card

    OpenAI. GPT-5 System Card. Technical report, OpenAI, August 2025b. URLhttps://openai. com/index/gpt-5-system-card/

  33. [33]

    GPT-5.4 Thinking System Card

    OpenAI. GPT-5.4 Thinking System Card. Technical report, OpenAI, March 2026a. URLhttps: //openai.com/index/gpt-5-4-thinking-system-card/

  34. [34]

    GPT-5.5 System Card

    OpenAI. GPT-5.5 System Card. Technical report, OpenAI, April 2026b. URLhttps://openai. com/index/gpt-5-5-system-card/

  35. [35]

    Autonomous evaluation and refinement of digital agents

    Jiayi Pan, Yichi Zhang, Nicholas Tomlin, Yifei Zhou, Sergey Levine, and Alane Suhr. Autonomous evaluation and refinement of digital agents. InFirst Conference on Language Modeling, 2024. URL https://openreview.net/forum?id=NPAQ6FKSmK. 22 OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models

  36. [36]

    Webrl: Training llm web agents via self-evolving online curriculum reinforcement learning, 2024

    Zehan Qi, Xiao Liu, Iat Long Iong, Hanyu Lai, Xueqiao Sun, Wenyi Zhao, Yu Yang, Xinyue Yang, Jiadai Sun, Shuntian Yao, Tianjie Zhang, Wei Xu, Jie Tang, and Yuxiao Dong. Webrl: Training llm web agents via self-evolving online curriculum reinforcement learning, 2024. URLhttps: //arxiv.org/abs/2411.02337

  37. [37]

    Ui-tars: Pioneering automated gui interaction with native agents

    Yujia Qin, Yining Ye, Junjie Fang, Haoming Wang, Shihao Liang, Shizuo Tian, Junda Zhang, Jiahao Li, Yunxin Li, Shijue Huang, et al. Ui-tars: Pioneering automated gui interaction with native agents. arXiv preprint arXiv:2501.12326, 2025

  38. [38]

    Qwen3.5: Towards native multimodal agents, February 2026

    Qwen Team. Qwen3.5: Towards native multimodal agents, February 2026. URLhttps://qwen. ai/blog?id=qwen3.5

  39. [39]

    Androidworld: A dynamic benchmarking environment for autonomous agents

    Christopher Rawles, Sarah Clinckemaillie, Yifan Chang, Jonathan Waltz, Gabrielle Lau, Marybeth Fair, Alice Li, William E Bishop, Wei Li, Folawiyo Campbell-Ajala, Daniel Kenji Toyama, Robert James Berry, Divya Tyamagundlu, Timothy P Lillicrap, and Oriana Riva. Androidworld: A dynamic benchmarking environment for autonomous agents. InThe Thirteenth Internat...

  40. [40]

    Deepseekmath: Pushing the limits of mathematical reasoning in open language models.arXiv preprint arXiv:2402.03300, 2024

    ZhihongShao,PeiyiWang,QihaoZhu,RunxinXu,JunxiaoSong,XiaoBi,HaoweiZhang,Mingchuan Zhang, YK Li, Yang Wu, et al. Deepseekmath: Pushing the limits of mathematical reasoning in open language models.arXiv preprint arXiv:2402.03300, 2024

  41. [41]

    Hybridflow: A flexible and efficient RLHF framework.arXiv preprint arXiv:2409.19256, 2024

    Guangming Sheng, Chi Zhang, Zilingfeng Ye, Xibin Wu, Wang Zhang, Ru Zhang, Yanghua Peng, Haibin Lin, and Chuan Wu. Hybridflow: A flexible and efficient RLHF framework.arXiv preprint arXiv:2409.19256, 2024

  42. [42]

    A survey of neural code intelligence: Paradigms, advances and beyond.arXiv preprint arXiv:2403.14734, 2024

    Qiushi Sun, Zhirui Chen, Fangzhi Xu, Kanzhi Cheng, Chang Ma, Zhangyue Yin, Jianing Wang, Chengcheng Han, Renyu Zhu, Shuai Yuan, et al. A survey of neural code intelligence: Paradigms, advances and beyond.arXiv preprint arXiv:2403.14734, 2024

  43. [43]

    Os-genesis: Automating gui agent trajectory construction via reverse task synthesis

    Qiushi Sun, Kanzhi Cheng, Zichen Ding, Chuanyang Jin, Yian Wang, Fangzhi Xu, Zhenyu Wu, Chengyou Jia, Liheng Chen, Zhoumianze Liu, et al. Os-genesis: Automating gui agent trajectory construction via reverse task synthesis. InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 5555–5579, 2025a

  44. [44]

    OS- sentinel: Towards safety-enhanced mobile GUI agents via hybrid validation in realistic workflows

    Qiushi Sun, Mukai Li, Zhoumianze Liu, Zhihui Xie, Fangzhi Xu, Zhangyue Yin, Kanzhi Cheng, Zehao Li, Zichen Ding, Qi Liu, Zhiyong Wu, Zhuosheng Zhang, Ben Kao, and Lingpeng Kong. OS- sentinel: Towards safety-enhanced mobile GUI agents via hybrid validation in realistic workflows. InProceedings of the 64th Annual Meeting of the Association for Computational...

  45. [45]

    Scienceboard: Evaluating multimodal autonomous agents in realistic scientific workflows

    Qiushi Sun, Zhoumianze Liu, Chang Ma, Zichen Ding, Fangzhi Xu, Zhangyue Yin, Haiteng Zhao, Zhenyu Wu, Kanzhi Cheng, Zhaoyang Liu, Jianing Wang, Qintong Li, Xiangru Tang, Tianbao Xie, Xiachong Feng, Xiang Li, Ben Kao, Wenhai Wang, Biqing Qi, Lingpeng Kong, and Zhiyong Wu. Scienceboard: Evaluating multimodal autonomous agents in realistic scientific workflo...

  46. [46]

    Coda: Coordinating the cerebrum and cerebellum for a dual- brain computer use agent with decoupled reinforcement learning.arXiv preprint arXiv:2508.20096, 2025b

    Zeyi Sun, Yuhang Cao, Jianze Liang, Qiushi Sun, Ziyu Liu, Zhixiong Zhang, Yuhang Zang, Xiaoyi Dong, Kai Chen, Dahua Lin, et al. Coda: Coordinating the cerebrum and cerebellum for a dual- brain computer use agent with decoupled reinforcement learning.arXiv preprint arXiv:2508.20096, 2025b

  47. [47]

    InSTA: Towards internet-scale training for agents.arXiv preprint arXiv:2502.06776, 2025

    Brandon Trabucco, Gunnar Sigurdsson, Robinson Piramuthu, and Ruslan Salakhutdinov. InSTA: Towards internet-scale training for agents.arXiv preprint arXiv:2502.06776, 2025

  48. [48]

    Math-shepherd: Verify and reinforce LLMs step-by-step without human annotations

    Peiyi Wang, Lei Li, Zhihong Shao, Runxin Xu, Damai Dai, Yifei Li, Deli Chen, Yu Wu, and Zhifang Sui. Math-shepherd: Verify and reinforce LLMs step-by-step without human annotations. InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 9426–9439, Bangkok, Thailand, August 2024. Association...

  49. [49]

    Charles, Zhilin Yang, and Tao Yu

    Xinyuan Wang, Bowen Wang, Dunjie Lu, Junlin Yang, Tianbao Xie, Junli Wang, Jiaqi Deng, Xiaole Guo, Yiheng Xu, Chen Henry Wu, Zhennan Shen, Zhuokai Li, Ryan Li, Xiaochuan Li, Junda Chen, Zheng Boyuan, LI PEIHANG, Fangyu Lei, Ruisheng Cao, Yeqiao Fu, Dongchan Shin, Martin Shin, Hu Jiarui, Yuyan Wang, Jixuan Chen, Yuxiao Ye, Danyang Zhang, Yipu Wang, Heng Wa...

  50. [50]

    SynthAgent: Adapting web agents with synthetic supervision

    Zhaoyang Wang, Yiming Liang, Xuchao Zhang, Qianhui Wu, Siwei Han, Anson Bastos, Rujia Wang, Chetan Bansal, Baolin Peng, Jianfeng Gao, Saravan Rajmohan, and Huaxiu Yao. SynthAgent: Adapting web agents with synthetic supervision. InProceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 15730–15...

  51. [51]

    Gui-actor: Coordinate-free visual grounding for gui agents

    Qianhui Wu, Kanzhi Cheng, Rui Yang, Chaoyun Zhang, Jianwei Yang, Huiqiang Jiang, Jian Mu, Baolin Peng, Bo Qiao, Reuben Tan, et al. Gui-actor: Coordinate-free visual grounding for gui agents. arXiv preprint arXiv:2506.03143, 2025a

  52. [52]

    Os-oracle: A comprehensive framework for cross-platform gui critic models

    Zhenyu Wu, Jingjing Xie, Zehao Li, Bowen Yang, Qiushi Sun, Zhaoyang Liu, Zhoumianze Liu, Yu Qiao, Xiangyu Yue, Zun Wang, et al. Os-oracle: A comprehensive framework for cross-platform gui critic models. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 27514–27524, 2026

  53. [53]

    Os-copilot: Towards generalist computer agents with self-improvement,

    Zhiyong Wu, Chengcheng Han, Zichen Ding, Zhenmin Weng, Zhoumianze Liu, Shunyu Yao, Tao Yu, and Lingpeng Kong. Os-copilot: Towards generalist computer agents with self-improvement,

  54. [54]

    OS-ATLAS: Foundation action model for generalist GUI agents

    Zhiyong Wu, Zhenyu Wu, Fangzhi Xu, Yian Wang, Qiushi Sun, Chengyou Jia, Kanzhi Cheng, Zichen Ding, Liheng Chen, Paul Pu Liang, and Yu Qiao. OS-ATLAS: Foundation action model for generalist GUI agents. InThe Thirteenth International Conference on Learning Representations, 2025b. URL https://openreview.net/forum?id=n9PDaFNi8t

  55. [55]

    UI- genie: A self-improving approach for iteratively boosting MLLM-based mobile GUI agents

    Han Xiao, Guozhi Wang, Yuxiang Chai, Zimu Lu, Weifeng Lin, Hao He, Lue Fan, Liuyang Bian, Rui Hu, Liang Liu, Shuai Ren, Yafei Wen, Xiaoxin Chen, Aojun Zhou, and Hongsheng Li. UI- genie: A self-improving approach for iteratively boosting MLLM-based mobile GUI agents. In 24 OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward...

  56. [56]

    OSWorld: Benchmarking multimodal agents for open-ended tasks in real computer environments

    Tianbao Xie, Danyang Zhang, Jixuan Chen, Xiaochuan Li, Siheng Zhao, Ruisheng Cao, Toh Jing Hua, Zhoujun Cheng, Dongchan Shin, Fangyu Lei, Yitao Liu, Yiheng Xu, Shuyan Zhou, Silvio Savarese, Caiming Xiong, Victor Zhong, and Tao Yu. OSWorld: Benchmarking multimodal agents for open-ended tasks in real computer environments. InThe Thirty-eight Conference on N...

  57. [57]

    Introducing osworld-verified.xlang.ai, July 2025

    TianbaoXie, MengqiYuan, DanyangZhang, XinzhuangXiong, ZhennanShen, ZilongZhou, Xinyuan Wang, Yanxu Chen, Jiaqi Deng, Junda Chen, Bowen Wang, Haoyuan Wu, Jixuan Chen, Junli Wang, Dunjie Lu, Hao Hu, and Tao Yu. Introducing osworld-verified.xlang.ai, July 2025. URL https://xlang.ai/blog/osworld-verified

  58. [58]

    Odysseyarena: Benchmarking large language models for long-horizon, active and inductive interactions.arXiv preprint arXiv:2602.05843, 2026

    Fangzhi Xu, Hang Yan, Qiushi Sun, Jinyang Wu, Zixian Huang, Muye Huang, Jingyang Gong, Zichen Ding, Kanzhi Cheng, Yian Wang, et al. Odysseyarena: Benchmarking large language models for long-horizon, active and inductive interactions.arXiv preprint arXiv:2602.05843, 2026

  59. [59]

    Agenttrek: Agenttrajectorysynthesisviaguidingreplaywithwebtutorials

    Yiheng Xu, Dunjie Lu, Zhennan Shen, Junli Wang, Zekun Wang, Yuchen Mao, Caiming Xiong, and TaoYu. Agenttrek: Agenttrajectorysynthesisviaguidingreplaywithwebtutorials. InTheThirteenth International Conference on Learning Representations, 2025a. URLhttps://openreview.net/ forum?id=EEgYUccwsV

  60. [60]

    Aguvis: Unified pure vision agents for autonomous GUI interaction

    Yiheng Xu, Zekun Wang, Junli Wang, Dunjie Lu, Tianbao Xie, Amrita Saha, Doyen Sahoo, Tao Yu, and Caiming Xiong. Aguvis: Unified pure vision agents for autonomous GUI interaction. In Forty-second International Conference on Machine Learning, 2025b. URLhttps://openreview. net/forum?id=PlihOwfx4r

  61. [61]

    EvoCUA: Evolving computer use agents via learning from scalable synthetic experience

    Taofeng Xue, Chong Peng, Mianqiu Huang, Linsen Guo, Tiancheng Han, Haozhe Wang, Xiaocheng Zhang, Xin Yang, Dengchang Zhao, Jinrui Ding, Xiandi Ma, Yuchen Xie, Peng Pei, Xunliang Cai, and Xipeng Qiu. EvoCUA: Evolving computer use agents via learning from scalable synthetic experience. InSecond Workshop on Agents in the Wild: Safety, Security, and Beyond, 2...

  62. [62]

    An illusion of progress? assessing the current state of web agents

    Tianci Xue, Weijian Qi, Tianneng Shi, Chan Hee Song, Boyu Gou, Dawn Song, Huan Sun, and Yu Su. An illusion of progress? assessing the current state of web agents. InSecond Conference on Language Modeling, 2025. URLhttps://openreview.net/forum?id=6jZi4HSs6o

  63. [63]

    Autonomous continual learning for environment adaptation of computer-use agents, 2026b

    Tianci Xue, Zeyi Liao, Tianneng Shi, Zilu Wang, Kai Zhang, Dawn Song, Yu Su, and Huan Sun. Autonomous continual learning for environment adaptation of computer-use agents, 2026b. URL https://arxiv.org/abs/2602.10356

  64. [64]

    OS-symphony: A holistic framework for robust and generalist computer-using agents

    Bowen Yang, Kaiming Jin, Zhenyu Wu, Zhaoyang Liu, Qiushi Sun, Zehao Li, JingJing Xie, Zhoumianze Liu, Fangzhi Xu, Kanzhi Cheng, Yian Wang, Qingyun Li, Yu Qiao, Zun Wang, and Zichen Ding. OS-symphony: A holistic framework for robust and generalist computer-using agents. InProceedings of the 64th Annual Meeting of the Association for Computational Lin- guis...

  65. [65]

    Set-of-mark prompting unleashes extraordinary visual grounding in gpt-4v.arXiv preprint arXiv:2310.11441, 2023

    Jianwei Yang, Hao Zhang, Feng Li, Xueyan Zou, Chunyuan Li, and Jianfeng Gao. Set-of-mark prompting unleashes extraordinary visual grounding in gpt-4v.arXiv preprint arXiv:2310.11441, 2023

  66. [66]

    Breaking the data barrier – building GUI agents through task generalization

    Junlei Zhang, Zichen Ding, Chang Ma, Zijie Chen, Qiushi Sun, Zhenzhong Lan, and Junxian He. Breaking the data barrier – building GUI agents through task generalization. InSecond Conference on Language Modeling, 2025. URLhttps://openreview.net/forum?id=QDtORaZt8K

  67. [67]

    Agent learning via early experience

    Kai Zhang, Xiangchao Chen, Bo Liu, Tianci Xue, Zeyi Liao, Zhihan Liu, Xiyao Wang, Yuting Ning, Zhaorun Chen, Xiaohan Fu, Jian Xie, Yuxuan Sun, Boyu Gou, Qi Qi, Zihang Meng, Jianwei Yang, Ning Zhang, Xian Li, Ashish Shah, Dat Huynh, Hengduo Li, Zi Yang, Xuefei Cao, Lawrence Keunho Jang, Shuyan Zhou, Jiacheng Zhu, Huan Sun, Jason E Weston, Yu Su, and Yifan ...

  68. [68]

    Gpt-4v(ision) is a generalist web agent, if grounded

    Boyuan Zheng, Boyu Gou, Jihyung Kil, Huan Sun, and Yu Su. Gpt-4v(ision) is a generalist web agent, if grounded. InForty-first International Conference on Machine Learning, 2024. URL https://openreview.net/forum?id=piecKJ2DlB

  69. [69]

    Xing, Hao Zhang, Joseph E

    Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric P. Xing, Hao Zhang, Joseph E. Gonzalez, and Ion Stoica. Judging llm-as-a-judge with mt-bench and chatbot arena. InProceedings of the 37th International Conference on Neural Information Processing Systems, NIPS ’23, Red Hook, NY, USA,...

  70. [70]

    Gonzalez, Clark Barrett, and Ying Sheng

    Lianmin Zheng, Liangsheng Yin, Zhiqiang Xie, Chuyue Sun, Jeff Huang, Cody Hao Yu, Shiyi Cao, Christos Kozyrakis, Ion Stoica, Joseph E. Gonzalez, Clark Barrett, and Ying Sheng. SGLang: Efficient execution of structured language model programs.arXiv preprint arXiv:2312.07104, 2023b

  71. [71]

    RMB: Comprehensively benchmarking reward models in LLM alignment

    Enyu Zhou, Guodong Zheng, Binghai Wang, Zhiheng Xi, Shihan Dou, Rong Bao, Wei Shen, Limao Xiong, Jessica Fan, Yurong Mou, Rui Zheng, Tao Gui, Qi Zhang, and Xuanjing Huang. RMB: Comprehensively benchmarking reward models in LLM alignment. InThe Thirteenth International Conference on Learning Representations, 2025. URL https://openreview.net/forum?id= kmgrlG9TR0

  72. [72]

    Xu, Hao Zhu, Xuhui Zhou, Robert Lo, Abishek Sridhar, Xianyi Cheng, Tianyue Ou, Yonatan Bisk, Daniel Fried, Uri Alon, and Graham Neubig

    Shuyan Zhou, Frank F. Xu, Hao Zhu, Xuhui Zhou, Robert Lo, Abishek Sridhar, Xianyi Cheng, Tianyue Ou, Yonatan Bisk, Daniel Fried, Uri Alon, and Graham Neubig. Webarena: A realistic web environment for building autonomous agents. InThe Twelfth International Conference on Learning Representations, 2024. URLhttps://openreview.net/forum?id=oKn9c6ytLx

  73. [73]

    Ashley, Wenyi Wang, Dmitrii Khizbullin, Yunyang Xiong, Zechun Liu, Ernie Chang, Raghuraman Krishnamoorthi, Yuandong Tian, Yangyang Shi, Vikas Chandra, and Jürgen Schmidhuber

    Mingchen Zhuge, Changsheng Zhao, Dylan R. Ashley, Wenyi Wang, Dmitrii Khizbullin, Yunyang Xiong, Zechun Liu, Ernie Chang, Raghuraman Krishnamoorthi, Yuandong Tian, Yangyang Shi, Vikas Chandra, and Jürgen Schmidhuber. Agent-as-a-judge: Evaluate agents with agents. InForty-second International Conference on Machine Learning, 2025. URLhttps://openreview.net/...

  74. [74]

    Intern-s1-pro: Scientific multimodal foundation model at trillion scale.arXiv preprint arXiv:2603.25040, 2026

    Yicheng Zou, Dongsheng Zhu, Lin Zhu, Tong Zhu, Yunhua Zhou, Peiheng Zhou, Xinyu Zhou, Dongzhan Zhou, Zhiwang Zhou, Yuhao Zhou, et al. Intern-s1-pro: Scientific multimodal foundation model at trillion scale.arXiv preprint arXiv:2603.25040, 2026. 26 OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models Appendix Contents...

  75. [76]

    User Instruction: The task to be completed

  76. [77]

    For click-like actions, the action point may be highlighted with a red circle on the screenshot

    Visual States: Screenshots from selected steps of the trajectory (e.g., the final few states, or a mix of initial and final states). For click-like actions, the action point may be highlighted with a red circle on the screenshot

  77. [78]

    [EVALUATION GOAL] Your task is to synthesize all the evidence above and determine whether the agent reasonably completed the task according to the user’s instruction

    Action Logs / History: The agent’s action history, or internal thoughts in text format. [EVALUATION GOAL] Your task is to synthesize all the evidence above and determine whether the agent reasonably completed the task according to the user’s instruction. The action history may incorrectly claim success or failure, and actions listed in the history are not...

  78. [79]

    Explicit answer requirement: - For general tasks (e.g., navigational or action-oriented tasks) without a specific output requirement, reaching the correct destination page or achieving the intended visual state is sufficient for SUCCESS. - If the instruction explicitly requests a text-based answer, such as answering a question, providing a filename, or st...

  79. [80]

    Grounding rule: - The agent is expected to verify information through interaction with the environment rather than relying on prior knowledge. - If the final answer contains specific facts such as numbers, names, dates, prices, rankings, titles, or claims, these facts should be obtained or verified through the agent’s interaction with the environment rath...

  80. [81]

    - Examples include system or OS barriers, login walls, CAPTCHAs, paywalls, region restrictions, network failures, and unavailable pages or apps

    Blocked / impossible rule: - If the task fails because it is persistently blocked by external constraints, the final judgment must be FAIL, even if the agent behaved logically. - Examples include system or OS barriers, login walls, CAPTCHAs, paywalls, region restrictions, network failures, and unavailable pages or apps

Showing first 80 references.