REVIEW 3 major objections 6 minor 82 references
Vision-language judges of computer-using agents systematically accept failed runs as successes, and only expensive frontier models hold up—until open reward models trained on a new human-gold benchmark match them at a fraction of the cost.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-31 02:16 UTC pith:GU67LITU
load-bearing objection Solid human-gold judge benchmark plus usable open reward models; the main soft spot is ensemble-distilled training labels, not the measurement claims. the 3 major comments →
OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
VLM judges of CUA trajectories fall short of an ideal judge and share one dominant failure mode—over-accepting incomplete tasks as successes—driven more by the agent’s text history than by the screenshots; the few judges accurate enough to trust cost too much for training-scale use, while open OS-Shepherd models trained on a new agreement-filtered corpus close most of that gap at 30–60× lower cost and transfer leniency resistance out of distribution.
What carries the argument
OSReward: a cross-platform benchmark of 1019 human-gold CUA trajectories (with Hard and Multi subsets) that measures judges themselves, plus OS-Shepherd-100K and the two-stage SFT+RL OS-Shepherd reward models that target false successes directly.
Load-bearing premise
That high-agreement labels from strong vision-language judges—after dropping ambiguous cases—are reliable enough training targets even though those judges herd and share the same leniency bias the models are meant to unlearn.
What would settle it
Hold OS-Shepherd and the frontier judges to a new batch of long-horizon desktop trajectories with independent multi-annotator human gold, and check whether fail-recall on false successes stays near the paper’s ~60% hard-set level or collapses back toward the untuned base’s near-zero catch rate.
If this is right
- Training-time CUA reward no longer has to be bought only at frontier API prices; a small self-hostable judge can sit near the cost–accuracy frontier.
- CUA evaluation and data filtering should treat false successes as the primary error mode and measure fail-recall, not only binary accuracy.
- Judging prompts and reward pipelines should keep full action/thought text; screenshot count and click markers move aggregate accuracy little.
- A single learned judge can begin to replace per-task human-written verifiers across mobile, web, and desktop once de-biasing transfers.
- Fine-grained alignment and efficiency grading remain much weaker than binary outcome judging and need separate calibration work.
Where Pith is reading between the lines
- If text narrative dominates the verdict, agents that learn to write confident closing claims could systematically game reward models unless screen-verification is forced in the objective.
- Agreement filtering may quietly discard the hardest genuine failures, so the next corpus gains may come from adversarial mining of judge disagreement rather than more unanimous easy labels.
- The platform gap (desktop hardest, mobile easiest) suggests reward-model progress will stall on long GUI+CLI desktop runs unless those are overweighted in training.
- Soft, confidence-weighted labels look more promising than majority vote, since judges herd on the same hard cases.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces OSReward, a cross-platform benchmark of 1019 human-gold CUA trajectories for evaluating VLM judges, plus OSReward-Hard and OSReward-Multi. Across 27 judges it reports a shared leniency bias (over-accepting incomplete tasks), a sharp accuracy drop on Hard (best ~70%), and a cost–accuracy frontier on which only expensive frontier models remain usable. It then releases OS-Shepherd-100K (ensemble-labeled, agreement-filtered) and trains open 9B/35B reward models that approach mid-tier commercial accuracy at much lower cost and transfer fail-recall gains to OSWorld, WebArena, and AndroidWorld. Primary measurement claims rest on multi-stage human labels; the open models rest on distilled VLM-ensemble supervision plus a short GRPO stage targeting false successes.
Significance. If the human-gold measurements hold, the paper supplies the first standardized, multi-platform stress test of CUA trajectory judges and documents a concrete, shared failure mode (narrative-driven false successes) that matters for evaluation, curation, and RL. The released benchmark, reasoning-annotated corpus, and self-hostable 9B/35B checkpoints are immediately usable artifacts; the cost–accuracy framing and held-out transfer results strengthen the case that reliable CUA reward need not be frontier-priced. The human pipeline (peer-screened instructions, diverse agent backbones, triple annotation plus meta-review, Hard re-verification, reported agreement) is a genuine strength relative to reused-trajectory judge studies.
major comments (3)
- [§6.1–6.2] §6.1–6.2 and Fig. 10/14: OS-Shepherd-100K labels are produced by the same VLM-judge class shown in §4.3 and §5.3 to herd and share leniency, with agreement filtering and a short GRPO pass on mined false successes as the main mitigations. This does not undermine the human-gold measurement claims on OSReward, but it is load-bearing for any claim that the open models are independently reliable rather than distilled ensemble imitators. The manuscript should state this limitation more explicitly (e.g., in §6.3 or the conclusion), quantify residual disagreement between the ensemble label and a human audit on a held-out training subsample if feasible, and avoid language that equates ensemble agreement with ground truth.
- [§7, Fig. 11] §7 / Fig. 11: Transfer is reported as agreement with each benchmark’s human-written verifier, which the paper correctly notes are imperfect (citing Xie et al., 2025). The text should more clearly separate “agreement with verifier” from “accuracy,” and, where possible, break out false-positive vs. false-negative shifts so readers can see that the transferred gain is specifically fail-catching rather than overall verifier mimicry. This is especially important on OSWorld, where §E.5 already shows FPs dominate judge errors.
- [Fig. 2, §5.4] Fig. 2, §5.4, §E.1: Cost figures mix official API list prices with “May 2026 market rates” for open-weight models and report 30–60× savings vs. frontier. The comparison is directionally convincing for OS-Shepherd-9B vs. Opus/GPT-5.5, but the paper should fix the pricing date/source table, state whether self-hosted GPU amortization is included in the “API-equivalent” $1.36 figure, and ensure the abstract/body multiplier is computed on a single, documented basket (full-set vs. Hard, same token assumptions).
minor comments (6)
- [Abstract] Abstract vs. body: the provided abstract text elsewhere says “30–60% lower cost” while the manuscript body and Fig. 2 claim “30–60×”; unify the multiplier everywhere.
- [§3.1] §3.1 / §A.4: Mobile environment description is duplicated (“The mobile environment is hosted in an Android emulator…” appears twice in close succession).
- [§4.5] §4.5 / Table 2: OSReward-Multi alignment is effectively two-level after removing the single 0-scored run (§B.4); state this in the main Multi discussion so macro-recall is not over-interpreted as three-class grading.
- [Table 1] Fig. 1 caption and Table 1: several model marketing names (GPT-5.5, Claude-Opus-4-8, etc.) will date quickly; keep API identifiers from Table 9 adjacent in the main results for reproducibility.
- [§5.2] §5.2 takeaway claims actions carry “roughly more signal” than CoT; the −1.8pp vs −7.2pp ablations support text>vision more cleanly than actions>CoT—soften the wording.
- [§3.4] Minor typos/style: “with detailed definitions in with detailed definitions in §B.4” (§3.4); “envel⌢pe” artifacts in the author block; ensure Hard-set N=284 platform counts in Fig. 8 match the released metadata.
Circularity Check
No significant circularity: primary claims rest on held-out multi-stage human-gold labels and external verifiers, not on quantities defined by the training ensemble.
full rationale
OSReward is an empirical measurement and systems paper, not a first-principles derivation. The load-bearing reliability claims (frontier judges share leniency; accuracy collapses on OSReward-Hard; OS-Shepherd matches commercial judges at much lower cost and transfers de-biasing) are scored against 1019 trajectories with multi-annotator human gold plus meta-review, fully disjoint from OS-Shepherd-100K. Training labels are agreement-filtered VLM ensemble judgments—a standard distillation setup with a known shared-bias risk—but the paper does not redefine success by those labels, does not fit a parameter on the evaluation set and call it a prediction, and does not import a self-authored uniqueness theorem to force the result. Held-out agreement with OSWorld/WebArena/AndroidWorld human-written verifiers and the reported base→SFT→RL fail-recall shift on human-gold Hard further separate measured performance from training-label construction. Self-citations (OS-Genesis, OpenMobile, infrastructure papers) supply data-collection methods and reused open trajectories, not the gold verdicts or the accuracy numbers. No step reduces a claimed prediction to its inputs by construction.
Axiom & Free-Parameter Ledger
free parameters (4)
- N trailing screenshots (default 5) =
5
- Ensemble agreement filter threshold =
~85% of judged trajectories retained
- GRPO RL hyperparameters =
lr=1e-6, KL=0.001, ~150 steps
- API/market unit prices for cost frontier =
e.g. OS-Shepherd-9B ~$1.36 per full-set pass
axioms (4)
- domain assumption Multi-stage human annotation (3-way + meta-review) yields ground-truth success/fail for CUA trajectories.
- domain assumption A trajectory is FAIL if the agent did not obtain/verify the answer through the environment, even if the answer is factually correct.
- ad hoc to paper High-agreement votes among strong VLM judges are sufficiently clean training labels after dropping the ambiguous middle.
- domain assumption Standard supervised fine-tuning plus GRPO improves judge calibration without needing new human labels.
invented entities (2)
-
OSReward / OSReward-Hard / OSReward-Multi
independent evidence
-
OS-Shepherd-100K and OS-Shepherd 9B/35B
independent evidence
read the original abstract
Computer-using agents (CUAs) are advancing rapidly across the digital world. A CUA trajectory records the agent's actions, states, and reasoning. Verifying whether it fulfilled the task instruction is central to CUA evaluation, data curation, and reinforcement learning. Neither human-written verifiers nor human annotators can provide such verification at scale, so the field increasingly turns to vision-language models (VLMs) as judges of CUA trajectories. But a fundamental question has long gone unexamined: are these VLM judges reliable enough? To study it systematically, we introduce OSReward, a realistic, high-quality benchmark that evaluates VLM judges on CUA trajectories. The trajectories come from diverse agent backbones executing human-verified instructions across platforms, then rigorously labeled with ground-truth verdicts through multi-stage human annotation. Building on it, we derive OSReward-Hard, a challenge set concentrating genuinely hard cases, and OSReward-Multi for fine-grained efficiency and alignment scoring. The most comprehensive evaluation of VLM judges to date finds even state-of-the-art models fall short of an ideal judge, sharing a systematic leniency bias that mislabels failed runs as successes. The few reliable enough to trust are too expensive to run at scale, while affordable open models trail far behind. To close this gap, we construct and release OS-Shepherd-100K, an open corpus of reasoning-annotated trajectory judgments for the CUA community. On it, we train OS-Shepherd (9B and 35B), open reward models that supply low-cost, stable, and reliable reward signals, matching commercial judges at 30-60% lower cost than the frontier. Extensive analyses further inform the design of reliable CUA reward at scale. Our code, benchmark, dataset, and model checkpoints are available at https://os-copilot.github.io/OSReward-Home/.
Figures
Reference graph
Works this paper leans on
-
[1]
Introducing Claude Haiku 4.5
Anthropic. Introducing Claude Haiku 4.5. https://www.anthropic.com/news/ claude-haiku-4-5, October 2025
2025
-
[2]
Introducing Claude Opus 4.6
Anthropic. Introducing Claude Opus 4.6. https://www.anthropic.com/news/ claude-opus-4-6, February 2026a
-
[3]
Introducing Claude Opus 4.8
Anthropic. Introducing Claude Opus 4.8. https://www.anthropic.com/news/ claude-opus-4-8, May 2026b
-
[4]
Introducing Claude Sonnet 4.6
Anthropic. Introducing Claude Sonnet 4.6. https://www.anthropic.com/news/ claude-sonnet-4-6, February 2026c
-
[5]
Digirl: Training in-the-wild device-control agents with autonomous reinforcement learning, 2024
Hao Bai, Yifei Zhou, Mert Cemri, Jiayi Pan, Alane Suhr, Sergey Levine, and Aviral Kumar. Digirl: Training in-the-wild device-control agents with autonomous reinforcement learning, 2024. URL https://arxiv.org/abs/2406.11896
Pith/arXiv arXiv 2024
-
[6]
Qwen3-vl technical report.arXiv preprint arXiv:2511.21631, 2025
Shuai Bai, Yuxuan Cai, Ruizhe Chen, Keqin Chen, Xionghui Chen, Zesen Cheng, Lianghao Deng, Wei Ding, Chang Gao, Chunjiang Ge, et al. Qwen3-vl technical report.arXiv preprint arXiv:2511.21631, 2025
Pith/arXiv arXiv 2025
-
[7]
Windowsagentarena: Evaluating multi-modal os agents at scale, 2024
Rogerio Bonatti, Dan Zhao, Francesco Bonacci, Dillon Dupont, Sara Abdali, Yinheng Li, Yadong Lu, JustinWagle,KazuhitoKoishida,ArthurBucker,LawrenceJang,andZackHui. Windowsagentarena: Evaluating multi-modal os agents at scale, 2024. URLhttps://arxiv.org/abs/2409.08264
Pith/arXiv arXiv 2024
-
[8]
Seed2.0 model card: Towards intelligence frontier for real-world complexity
ByteDance Seed Team. Seed2.0 model card: Towards intelligence frontier for real-world complexity. Technical report, ByteDance, February 2026
2026
-
[9]
Web-shepherd: Advancing PRMs for reinforcing web agents
Hyungjoo Chae, Sunghwan Kim, Junhee Cho, Seungone Kim, Seungjun Moon, Gyeom Hwangbo, Dongha Lim, Minjin Kim, Yeonjun Hwang, Minju Gwak, Dongwook Choi, Minseok Kang, Gwanhoon Im, ByeongUng Cho, Hyojun Kim, Jun Hee Han, Taeyoon Kwon, Minju Kim, Beong woo Kwak, Dongjin Kang, and Jinyoung Yeo. Web-shepherd: Advancing PRMs for reinforcing web agents. InThe Thi...
2025
-
[10]
Gui-shepherd: Reliableprocessrewardand verification for long-sequence gui tasks, 2025a
Cong Chen, Kaixiang Ji, Hao Zhong, Muzhi Zhu, Anzhou Li, Guo Gan, Ziyuan Huang, Cheng Zou, JiajiaLiu,JingdongChen,HaoChen,andChunhuaShen. Gui-shepherd: Reliableprocessrewardand verification for long-sequence gui tasks, 2025a. URLhttps://arxiv.org/abs/2509.23738
-
[11]
Mllm-as-a-judge: Assessing multimodal llm-as-a-judge with vision-language benchmark, 2024
Dongping Chen, Ruoxi Chen, Shilin Zhang, Yinuo Liu, Yaochen Wang, Huichi Zhou, Qihui Zhang, Yao Wan, Pan Zhou, and Lichao Sun. Mllm-as-a-judge: Assessing multimodal llm-as-a-judge with vision-language benchmark, 2024. URLhttps://arxiv.org/abs/2402.04788. 20 OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models
Pith/arXiv arXiv 2024
-
[12]
Xuetian Chen, Yinghao Chen, Xinfeng Yuan, Zhuo Peng, Lu Chen, Yuekeng Li, Zhoujia Zhang, Yingqian Huang, Leyan Huang, Jiaqing Liang, et al. Os-map: How far can computer-using agents go in breadth and depth?arXiv preprint arXiv:2507.19132, 2025b
-
[13]
SeeClick: Harnessing GUI grounding for advanced visual GUI agents
Kanzhi Cheng, Qiushi Sun, Yougang Chu, Fangzhi Xu, Li YanTao, Jianbing Zhang, and Zhiyong Wu. SeeClick: Harnessing GUI grounding for advanced visual GUI agents. InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 9313–9332, Bangkok, Thailand, August 2024. Association for Computational Li...
2024
-
[14]
Kanzhi Cheng, Zehao Li, Zheng Ma, Nuo Chen, Jialin Cao, Qiushi Sun, Zichen Ding, Fangzhi Xu, Hang Yan, Jiajun Chen, et al. Openmobile: Building open mobile agents with task and trajectory synthesis.arXiv preprint arXiv:2604.15093, 2026
Pith/arXiv arXiv 2026
-
[15]
Training verifiers to solve math word problems, 2021
Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, Christopher Hesse, and John Schulman. Training verifiers to solve math word problems, 2021. URLhttps://arxiv.org/ abs/2110.14168
Pith/arXiv arXiv 2021
-
[16]
A coefficient of agreement for nominal scales.Educational and Psychological Measure- ment, 20(1):37–46, 1960
Jacob Cohen. A coefficient of agreement for nominal scales.Educational and Psychological Measure- ment, 20(1):37–46, 1960
1960
-
[17]
Gheorghe Comanici, Eric Bieber, Mike Schaekermann, Ice Pasupat, Noveen Sachdeva, Inderjit Dhillon, Marcel Blistein, Ori Ram, Dan Zhang, Evan Rosen, et al. Gemini 2.5: Pushing the frontier with advanced reasoning, multimodality, long context, and next generation agentic capabilities. arXiv preprint arXiv:2507.06261, 2025
Pith/arXiv arXiv 2025
-
[18]
Gemini 3 pro model card, November 2025
Gemini Team. Gemini 3 pro model card, November 2025. URLhttps://deepmind.google/ models/gemini/
2025
-
[19]
Navigating the digital world as humans do: Universal visual grounding for GUI agents
Boyu Gou, Ruohan Wang, Boyuan Zheng, Yanan Xie, Cheng Chang, Yiheng Shu, Huan Sun, and Yu Su. Navigating the digital world as humans do: Universal visual grounding for GUI agents. InThe Thirteenth International Conference on Learning Representations, 2025. URLhttps: //openreview.net/forum?id=kxnoqaisCT
2025
-
[20]
GPT-4o system card.arXiv preprint arXiv:2410.21276, 2024
Aaron Hurst, Adam Lerer, Adam P Goucher, Adam Perelman, Aditya Ramesh, Aidan Clark, AJ Os- trow, Akila Welihinda, Alan Hayes, Alec Radford, et al. GPT-4o system card.arXiv preprint arXiv:2410.21276, 2024
Pith/arXiv arXiv 2024
-
[21]
Chengyou Jia, Minnan Luo, Zhuohang Dang, Qiushi Sun, Fangzhi Xu, Junlin Hu, Tianbao Xie, and Zhiyong Wu. AgentStore: Scalable integration of heterogeneous agents as specialized generalist computer assistant. InFindings of the Association for Computational Linguistics: ACL 2025, pages 8908–8934, Vienna, Austria, July 2025. Association for Computational Lin...
-
[22]
TreeCUA: Efficiently scaling GUI automation with tree-structured verifiable evolution
Deyang Jiang, Jing Huang, Xuanle Zhao, Lei Chen, Liming Zheng, Fanfan Liu, Haibo Qiu, Peng Shi, and Zhixiong Zeng. TreeCUA: Efficiently scaling GUI automation with tree-structured verifiable evolution. InForty-third International Conference on Machine Learning, 2026. URL https:// openreview.net/forum?id=KBCWS6OBnD. 21 OSReward: Instituting Standardized Ev...
2026
-
[23]
Kimi Team, Tongtong Bai, Yifan Bai, Yiping Bao, SH Cai, Yuan Cao, Y Charles, HS Che, Cheng Chen, Guanduo Chen, et al. Kimi k2. 5: Visual agentic intelligence.arXiv preprint arXiv:2602.02276, 2026
Pith/arXiv arXiv 2026
-
[24]
Mobileworld: Benchmarking autonomous mobile agents in agent-user interactive and mcp-augmented environments
Quyu Kong, Xu Zhang, Zhenyu Yang, Nolan Gao, Chen Liu, Panrong Tong, Chenglin Cai, Hanzhang Zhou, Jianan Zhang, Liangyu Chen, et al. Mobileworld: Benchmarking autonomous mobile agents in agent-user interactive and mcp-augmented environments. InProceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ...
2026
-
[25]
Computing Krippendorff’s alpha-reliability
Klaus Krippendorff. Computing Krippendorff’s alpha-reliability. Technical report, University of Pennsylvania, Annenberg School for Communication, 2011
2011
-
[26]
Vl-rewardbench: a challenging benchmark for vision-language generative reward models
Lei Li, Yuancheng Wei, Zhihui Xie, Xuqing Yang, Yifan Song, Peiyi Wang, Chenxin An, Tianyu Liu, Sujian Li, Bill Yuchen Lin, et al. Vl-rewardbench: a challenging benchmark for vision-language generative reward models. InProceedings of the Computer Vision and Pattern Recognition Conference, pages 24657–24668, 2025
2025
-
[27]
Os- themis: A scalable critic framework for generalist gui rewards, 2026
Zehao Li, Zhenyu Wu, Yibo Zhao, Bowen Yang, Jingjing Xie, Zhaoyang Liu, Zhoumianze Liu, Kaiming Jin, Jianze Liang, Zonglin Li, Feng Wu, Bowen Zhou, Zun Wang, and Zichen Ding. Os- themis: A scalable critic framework for generalist gui rewards, 2026. URLhttps://arxiv.org/ abs/2603.19191
arXiv 2026
-
[28]
Cuarewardbench: A benchmark for evaluating reward models on computer-using agent, 2025
Haojia Lin, Xiaoyu Tan, Yulei Qin, Zihan Xu, Yuchen Shi, Zongyi Li, Gang Li, Shaofei Cai, Siqi Cai, Chaoyou Fu, Ke Li, and Xing Sun. Cuarewardbench: A benchmark for evaluating reward models on computer-using agent, 2025. URLhttps://arxiv.org/abs/2510.18596
arXiv 2025
-
[29]
ScaleCUA: Scaling open- source computer use agents with cross-platform data
Zhaoyang Liu, JingJing Xie, Zichen Ding, Zehao Li, Bowen Yang, Zhenyu Wu, Xuehui Wang, Qiushi Sun, Shi Liu, Weiyun Wang, Shenglong Ye, Qingyun Li, Zeyue Tian, Gen Luo, Xiangyu Yue, Biqing Qi, Kai Chen, Bowen Zhou, Yu Qiao, Qifeng Chen, and Wenhai Wang. ScaleCUA: Scaling open- source computer use agents with cross-platform data. InThe Fourteenth Internatio...
2026
-
[30]
Xing Han Lù, Amirhossein Kazemnejad, Nicholas Meade, Arkil Patel, Dongchan Shin, Alejandra Zambrano, Karolina Stańczak, Peter Shaw, Christopher J. Pal, and Siva Reddy. Agentrewardbench: Evaluating automatic evaluations of web agent trajectories, 2025. URLhttps://arxiv.org/ abs/2504.08942
arXiv 2025
-
[31]
Computer-using agent: Introducing a universal interface for ai to interact with the digital world, 2025a
OpenAI. Computer-using agent: Introducing a universal interface for ai to interact with the digital world, 2025a. URLhttps://openai.com/index/computer-using-agent
-
[32]
GPT-5 System Card
OpenAI. GPT-5 System Card. Technical report, OpenAI, August 2025b. URLhttps://openai. com/index/gpt-5-system-card/
-
[33]
GPT-5.4 Thinking System Card
OpenAI. GPT-5.4 Thinking System Card. Technical report, OpenAI, March 2026a. URLhttps: //openai.com/index/gpt-5-4-thinking-system-card/
-
[34]
GPT-5.5 System Card
OpenAI. GPT-5.5 System Card. Technical report, OpenAI, April 2026b. URLhttps://openai. com/index/gpt-5-5-system-card/
-
[35]
Autonomous evaluation and refinement of digital agents
Jiayi Pan, Yichi Zhang, Nicholas Tomlin, Yifei Zhou, Sergey Levine, and Alane Suhr. Autonomous evaluation and refinement of digital agents. InFirst Conference on Language Modeling, 2024. URL https://openreview.net/forum?id=NPAQ6FKSmK. 22 OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models
2024
-
[36]
Webrl: Training llm web agents via self-evolving online curriculum reinforcement learning, 2024
Zehan Qi, Xiao Liu, Iat Long Iong, Hanyu Lai, Xueqiao Sun, Wenyi Zhao, Yu Yang, Xinyue Yang, Jiadai Sun, Shuntian Yao, Tianjie Zhang, Wei Xu, Jie Tang, and Yuxiao Dong. Webrl: Training llm web agents via self-evolving online curriculum reinforcement learning, 2024. URLhttps: //arxiv.org/abs/2411.02337
Pith/arXiv arXiv 2024
-
[37]
Ui-tars: Pioneering automated gui interaction with native agents
Yujia Qin, Yining Ye, Junjie Fang, Haoming Wang, Shihao Liang, Shizuo Tian, Junda Zhang, Jiahao Li, Yunxin Li, Shijue Huang, et al. Ui-tars: Pioneering automated gui interaction with native agents. arXiv preprint arXiv:2501.12326, 2025
Pith/arXiv arXiv 2025
-
[38]
Qwen3.5: Towards native multimodal agents, February 2026
Qwen Team. Qwen3.5: Towards native multimodal agents, February 2026. URLhttps://qwen. ai/blog?id=qwen3.5
2026
-
[39]
Androidworld: A dynamic benchmarking environment for autonomous agents
Christopher Rawles, Sarah Clinckemaillie, Yifan Chang, Jonathan Waltz, Gabrielle Lau, Marybeth Fair, Alice Li, William E Bishop, Wei Li, Folawiyo Campbell-Ajala, Daniel Kenji Toyama, Robert James Berry, Divya Tyamagundlu, Timothy P Lillicrap, and Oriana Riva. Androidworld: A dynamic benchmarking environment for autonomous agents. InThe Thirteenth Internat...
2025
-
[40]
ZhihongShao,PeiyiWang,QihaoZhu,RunxinXu,JunxiaoSong,XiaoBi,HaoweiZhang,Mingchuan Zhang, YK Li, Yang Wu, et al. Deepseekmath: Pushing the limits of mathematical reasoning in open language models.arXiv preprint arXiv:2402.03300, 2024
Pith/arXiv arXiv 2024
-
[41]
Hybridflow: A flexible and efficient RLHF framework.arXiv preprint arXiv:2409.19256, 2024
Guangming Sheng, Chi Zhang, Zilingfeng Ye, Xibin Wu, Wang Zhang, Ru Zhang, Yanghua Peng, Haibin Lin, and Chuan Wu. Hybridflow: A flexible and efficient RLHF framework.arXiv preprint arXiv:2409.19256, 2024
Pith/arXiv arXiv 2024
-
[42]
Qiushi Sun, Zhirui Chen, Fangzhi Xu, Kanzhi Cheng, Chang Ma, Zhangyue Yin, Jianing Wang, Chengcheng Han, Renyu Zhu, Shuai Yuan, et al. A survey of neural code intelligence: Paradigms, advances and beyond.arXiv preprint arXiv:2403.14734, 2024
Pith/arXiv arXiv 2024
-
[43]
Os-genesis: Automating gui agent trajectory construction via reverse task synthesis
Qiushi Sun, Kanzhi Cheng, Zichen Ding, Chuanyang Jin, Yian Wang, Fangzhi Xu, Zhenyu Wu, Chengyou Jia, Liheng Chen, Zhoumianze Liu, et al. Os-genesis: Automating gui agent trajectory construction via reverse task synthesis. InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 5555–5579, 2025a
-
[44]
OS- sentinel: Towards safety-enhanced mobile GUI agents via hybrid validation in realistic workflows
Qiushi Sun, Mukai Li, Zhoumianze Liu, Zhihui Xie, Fangzhi Xu, Zhangyue Yin, Kanzhi Cheng, Zehao Li, Zichen Ding, Qi Liu, Zhiyong Wu, Zhuosheng Zhang, Ben Kao, and Lingpeng Kong. OS- sentinel: Towards safety-enhanced mobile GUI agents via hybrid validation in realistic workflows. InProceedings of the 64th Annual Meeting of the Association for Computational...
2026
-
[45]
Scienceboard: Evaluating multimodal autonomous agents in realistic scientific workflows
Qiushi Sun, Zhoumianze Liu, Chang Ma, Zichen Ding, Fangzhi Xu, Zhangyue Yin, Haiteng Zhao, Zhenyu Wu, Kanzhi Cheng, Zhaoyang Liu, Jianing Wang, Qintong Li, Xiangru Tang, Tianbao Xie, Xiachong Feng, Xiang Li, Ben Kao, Wenhai Wang, Biqing Qi, Lingpeng Kong, and Zhiyong Wu. Scienceboard: Evaluating multimodal autonomous agents in realistic scientific workflo...
-
[46]
Zeyi Sun, Yuhang Cao, Jianze Liang, Qiushi Sun, Ziyu Liu, Zhixiong Zhang, Yuhang Zang, Xiaoyi Dong, Kai Chen, Dahua Lin, et al. Coda: Coordinating the cerebrum and cerebellum for a dual- brain computer use agent with decoupled reinforcement learning.arXiv preprint arXiv:2508.20096, 2025b
-
[47]
InSTA: Towards internet-scale training for agents.arXiv preprint arXiv:2502.06776, 2025
Brandon Trabucco, Gunnar Sigurdsson, Robinson Piramuthu, and Ruslan Salakhutdinov. InSTA: Towards internet-scale training for agents.arXiv preprint arXiv:2502.06776, 2025
Pith/arXiv arXiv 2025
-
[48]
Math-shepherd: Verify and reinforce LLMs step-by-step without human annotations
Peiyi Wang, Lei Li, Zhihong Shao, Runxin Xu, Damai Dai, Yifei Li, Deli Chen, Yu Wu, and Zhifang Sui. Math-shepherd: Verify and reinforce LLMs step-by-step without human annotations. InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 9426–9439, Bangkok, Thailand, August 2024. Association...
-
[49]
Charles, Zhilin Yang, and Tao Yu
Xinyuan Wang, Bowen Wang, Dunjie Lu, Junlin Yang, Tianbao Xie, Junli Wang, Jiaqi Deng, Xiaole Guo, Yiheng Xu, Chen Henry Wu, Zhennan Shen, Zhuokai Li, Ryan Li, Xiaochuan Li, Junda Chen, Zheng Boyuan, LI PEIHANG, Fangyu Lei, Ruisheng Cao, Yeqiao Fu, Dongchan Shin, Martin Shin, Hu Jiarui, Yuyan Wang, Jixuan Chen, Yuxiao Ye, Danyang Zhang, Yipu Wang, Heng Wa...
2025
-
[50]
SynthAgent: Adapting web agents with synthetic supervision
Zhaoyang Wang, Yiming Liang, Xuchao Zhang, Qianhui Wu, Siwei Han, Anson Bastos, Rujia Wang, Chetan Bansal, Baolin Peng, Jianfeng Gao, Saravan Rajmohan, and Huaxiu Yao. SynthAgent: Adapting web agents with synthetic supervision. InProceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 15730–15...
2026
-
[51]
Gui-actor: Coordinate-free visual grounding for gui agents
Qianhui Wu, Kanzhi Cheng, Rui Yang, Chaoyun Zhang, Jianwei Yang, Huiqiang Jiang, Jian Mu, Baolin Peng, Bo Qiao, Reuben Tan, et al. Gui-actor: Coordinate-free visual grounding for gui agents. arXiv preprint arXiv:2506.03143, 2025a
-
[52]
Os-oracle: A comprehensive framework for cross-platform gui critic models
Zhenyu Wu, Jingjing Xie, Zehao Li, Bowen Yang, Qiushi Sun, Zhaoyang Liu, Zhoumianze Liu, Yu Qiao, Xiangyu Yue, Zun Wang, et al. Os-oracle: A comprehensive framework for cross-platform gui critic models. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 27514–27524, 2026
2026
-
[53]
Os-copilot: Towards generalist computer agents with self-improvement,
Zhiyong Wu, Chengcheng Han, Zichen Ding, Zhenmin Weng, Zhoumianze Liu, Shunyu Yao, Tao Yu, and Lingpeng Kong. Os-copilot: Towards generalist computer agents with self-improvement,
-
[54]
OS-ATLAS: Foundation action model for generalist GUI agents
Zhiyong Wu, Zhenyu Wu, Fangzhi Xu, Yian Wang, Qiushi Sun, Chengyou Jia, Kanzhi Cheng, Zichen Ding, Liheng Chen, Paul Pu Liang, and Yu Qiao. OS-ATLAS: Foundation action model for generalist GUI agents. InThe Thirteenth International Conference on Learning Representations, 2025b. URL https://openreview.net/forum?id=n9PDaFNi8t
-
[55]
UI- genie: A self-improving approach for iteratively boosting MLLM-based mobile GUI agents
Han Xiao, Guozhi Wang, Yuxiang Chai, Zimu Lu, Weifeng Lin, Hao He, Lue Fan, Liuyang Bian, Rui Hu, Liang Liu, Shuai Ren, Yafei Wen, Xiaoxin Chen, Aojun Zhou, and Hongsheng Li. UI- genie: A self-improving approach for iteratively boosting MLLM-based mobile GUI agents. In 24 OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward...
2025
-
[56]
OSWorld: Benchmarking multimodal agents for open-ended tasks in real computer environments
Tianbao Xie, Danyang Zhang, Jixuan Chen, Xiaochuan Li, Siheng Zhao, Ruisheng Cao, Toh Jing Hua, Zhoujun Cheng, Dongchan Shin, Fangyu Lei, Yitao Liu, Yiheng Xu, Shuyan Zhou, Silvio Savarese, Caiming Xiong, Victor Zhong, and Tao Yu. OSWorld: Benchmarking multimodal agents for open-ended tasks in real computer environments. InThe Thirty-eight Conference on N...
2024
-
[57]
Introducing osworld-verified.xlang.ai, July 2025
TianbaoXie, MengqiYuan, DanyangZhang, XinzhuangXiong, ZhennanShen, ZilongZhou, Xinyuan Wang, Yanxu Chen, Jiaqi Deng, Junda Chen, Bowen Wang, Haoyuan Wu, Jixuan Chen, Junli Wang, Dunjie Lu, Hao Hu, and Tao Yu. Introducing osworld-verified.xlang.ai, July 2025. URL https://xlang.ai/blog/osworld-verified
2025
-
[58]
Fangzhi Xu, Hang Yan, Qiushi Sun, Jinyang Wu, Zixian Huang, Muye Huang, Jingyang Gong, Zichen Ding, Kanzhi Cheng, Yian Wang, et al. Odysseyarena: Benchmarking large language models for long-horizon, active and inductive interactions.arXiv preprint arXiv:2602.05843, 2026
Pith/arXiv arXiv 2026
-
[59]
Agenttrek: Agenttrajectorysynthesisviaguidingreplaywithwebtutorials
Yiheng Xu, Dunjie Lu, Zhennan Shen, Junli Wang, Zekun Wang, Yuchen Mao, Caiming Xiong, and TaoYu. Agenttrek: Agenttrajectorysynthesisviaguidingreplaywithwebtutorials. InTheThirteenth International Conference on Learning Representations, 2025a. URLhttps://openreview.net/ forum?id=EEgYUccwsV
-
[60]
Aguvis: Unified pure vision agents for autonomous GUI interaction
Yiheng Xu, Zekun Wang, Junli Wang, Dunjie Lu, Tianbao Xie, Amrita Saha, Doyen Sahoo, Tao Yu, and Caiming Xiong. Aguvis: Unified pure vision agents for autonomous GUI interaction. In Forty-second International Conference on Machine Learning, 2025b. URLhttps://openreview. net/forum?id=PlihOwfx4r
-
[61]
EvoCUA: Evolving computer use agents via learning from scalable synthetic experience
Taofeng Xue, Chong Peng, Mianqiu Huang, Linsen Guo, Tiancheng Han, Haozhe Wang, Xiaocheng Zhang, Xin Yang, Dengchang Zhao, Jinrui Ding, Xiandi Ma, Yuchen Xie, Peng Pei, Xunliang Cai, and Xipeng Qiu. EvoCUA: Evolving computer use agents via learning from scalable synthetic experience. InSecond Workshop on Agents in the Wild: Safety, Security, and Beyond, 2...
-
[62]
An illusion of progress? assessing the current state of web agents
Tianci Xue, Weijian Qi, Tianneng Shi, Chan Hee Song, Boyu Gou, Dawn Song, Huan Sun, and Yu Su. An illusion of progress? assessing the current state of web agents. InSecond Conference on Language Modeling, 2025. URLhttps://openreview.net/forum?id=6jZi4HSs6o
2025
-
[63]
Autonomous continual learning for environment adaptation of computer-use agents, 2026b
Tianci Xue, Zeyi Liao, Tianneng Shi, Zilu Wang, Kai Zhang, Dawn Song, Yu Su, and Huan Sun. Autonomous continual learning for environment adaptation of computer-use agents, 2026b. URL https://arxiv.org/abs/2602.10356
-
[64]
OS-symphony: A holistic framework for robust and generalist computer-using agents
Bowen Yang, Kaiming Jin, Zhenyu Wu, Zhaoyang Liu, Qiushi Sun, Zehao Li, JingJing Xie, Zhoumianze Liu, Fangzhi Xu, Kanzhi Cheng, Yian Wang, Qingyun Li, Yu Qiao, Zun Wang, and Zichen Ding. OS-symphony: A holistic framework for robust and generalist computer-using agents. InProceedings of the 64th Annual Meeting of the Association for Computational Lin- guis...
2026
-
[65]
Jianwei Yang, Hao Zhang, Feng Li, Xueyan Zou, Chunyuan Li, and Jianfeng Gao. Set-of-mark prompting unleashes extraordinary visual grounding in gpt-4v.arXiv preprint arXiv:2310.11441, 2023
Pith/arXiv arXiv 2023
-
[66]
Breaking the data barrier – building GUI agents through task generalization
Junlei Zhang, Zichen Ding, Chang Ma, Zijie Chen, Qiushi Sun, Zhenzhong Lan, and Junxian He. Breaking the data barrier – building GUI agents through task generalization. InSecond Conference on Language Modeling, 2025. URLhttps://openreview.net/forum?id=QDtORaZt8K
2025
-
[67]
Agent learning via early experience
Kai Zhang, Xiangchao Chen, Bo Liu, Tianci Xue, Zeyi Liao, Zhihan Liu, Xiyao Wang, Yuting Ning, Zhaorun Chen, Xiaohan Fu, Jian Xie, Yuxuan Sun, Boyu Gou, Qi Qi, Zihang Meng, Jianwei Yang, Ning Zhang, Xian Li, Ashish Shah, Dat Huynh, Hengduo Li, Zi Yang, Xuefei Cao, Lawrence Keunho Jang, Shuyan Zhou, Jiacheng Zhu, Huan Sun, Jason E Weston, Yu Su, and Yifan ...
2026
-
[68]
Gpt-4v(ision) is a generalist web agent, if grounded
Boyuan Zheng, Boyu Gou, Jihyung Kil, Huan Sun, and Yu Su. Gpt-4v(ision) is a generalist web agent, if grounded. InForty-first International Conference on Machine Learning, 2024. URL https://openreview.net/forum?id=piecKJ2DlB
2024
-
[69]
Xing, Hao Zhang, Joseph E
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric P. Xing, Hao Zhang, Joseph E. Gonzalez, and Ion Stoica. Judging llm-as-a-judge with mt-bench and chatbot arena. InProceedings of the 37th International Conference on Neural Information Processing Systems, NIPS ’23, Red Hook, NY, USA,...
-
[70]
Gonzalez, Clark Barrett, and Ying Sheng
Lianmin Zheng, Liangsheng Yin, Zhiqiang Xie, Chuyue Sun, Jeff Huang, Cody Hao Yu, Shiyi Cao, Christos Kozyrakis, Ion Stoica, Joseph E. Gonzalez, Clark Barrett, and Ying Sheng. SGLang: Efficient execution of structured language model programs.arXiv preprint arXiv:2312.07104, 2023b
-
[71]
RMB: Comprehensively benchmarking reward models in LLM alignment
Enyu Zhou, Guodong Zheng, Binghai Wang, Zhiheng Xi, Shihan Dou, Rong Bao, Wei Shen, Limao Xiong, Jessica Fan, Yurong Mou, Rui Zheng, Tao Gui, Qi Zhang, and Xuanjing Huang. RMB: Comprehensively benchmarking reward models in LLM alignment. InThe Thirteenth International Conference on Learning Representations, 2025. URL https://openreview.net/forum?id= kmgrlG9TR0
2025
-
[72]
Xu, Hao Zhu, Xuhui Zhou, Robert Lo, Abishek Sridhar, Xianyi Cheng, Tianyue Ou, Yonatan Bisk, Daniel Fried, Uri Alon, and Graham Neubig
Shuyan Zhou, Frank F. Xu, Hao Zhu, Xuhui Zhou, Robert Lo, Abishek Sridhar, Xianyi Cheng, Tianyue Ou, Yonatan Bisk, Daniel Fried, Uri Alon, and Graham Neubig. Webarena: A realistic web environment for building autonomous agents. InThe Twelfth International Conference on Learning Representations, 2024. URLhttps://openreview.net/forum?id=oKn9c6ytLx
2024
-
[73]
Ashley, Wenyi Wang, Dmitrii Khizbullin, Yunyang Xiong, Zechun Liu, Ernie Chang, Raghuraman Krishnamoorthi, Yuandong Tian, Yangyang Shi, Vikas Chandra, and Jürgen Schmidhuber
Mingchen Zhuge, Changsheng Zhao, Dylan R. Ashley, Wenyi Wang, Dmitrii Khizbullin, Yunyang Xiong, Zechun Liu, Ernie Chang, Raghuraman Krishnamoorthi, Yuandong Tian, Yangyang Shi, Vikas Chandra, and Jürgen Schmidhuber. Agent-as-a-judge: Evaluate agents with agents. InForty-second International Conference on Machine Learning, 2025. URLhttps://openreview.net/...
2025
-
[74]
Yicheng Zou, Dongsheng Zhu, Lin Zhu, Tong Zhu, Yunhua Zhou, Peiheng Zhou, Xinyu Zhou, Dongzhan Zhou, Zhiwang Zhou, Yuhao Zhou, et al. Intern-s1-pro: Scientific multimodal foundation model at trillion scale.arXiv preprint arXiv:2603.25040, 2026. 26 OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models Appendix Contents...
arXiv 2026
-
[76]
User Instruction: The task to be completed
-
[77]
For click-like actions, the action point may be highlighted with a red circle on the screenshot
Visual States: Screenshots from selected steps of the trajectory (e.g., the final few states, or a mix of initial and final states). For click-like actions, the action point may be highlighted with a red circle on the screenshot
-
[78]
[EVALUATION GOAL] Your task is to synthesize all the evidence above and determine whether the agent reasonably completed the task according to the user’s instruction
Action Logs / History: The agent’s action history, or internal thoughts in text format. [EVALUATION GOAL] Your task is to synthesize all the evidence above and determine whether the agent reasonably completed the task according to the user’s instruction. The action history may incorrectly claim success or failure, and actions listed in the history are not...
-
[79]
Explicit answer requirement: - For general tasks (e.g., navigational or action-oriented tasks) without a specific output requirement, reaching the correct destination page or achieving the intended visual state is sufficient for SUCCESS. - If the instruction explicitly requests a text-based answer, such as answering a question, providing a filename, or st...
-
[80]
Grounding rule: - The agent is expected to verify information through interaction with the environment rather than relying on prior knowledge. - If the final answer contains specific facts such as numbers, names, dates, prices, rankings, titles, or claims, these facts should be obtained or verified through the agent’s interaction with the environment rath...
-
[81]
- Examples include system or OS barriers, login walls, CAPTCHAs, paywalls, region restrictions, network failures, and unavailable pages or apps
Blocked / impossible rule: - If the task fails because it is persistently blocked by external constraints, the final judgment must be FAIL, even if the agent behaved logically. - Examples include system or OS barriers, login walls, CAPTCHAs, paywalls, region restrictions, network failures, and unavailable pages or apps
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.