{"record_type":"pith_number_record","schema_url":"https://pith.science/schemas/pith-number/v1.json","pith_number":"pith:2025:XPQEDQ52F3GJ7IPFWOLIUFU6AK","short_pith_number":"pith:XPQEDQ52","schema_version":"1.0","canonical_sha256":"bbe041c3ba2ecc9fa1e5b3968a169e028df195688e9d15e88e2b358805e66801","source":{"kind":"arxiv","id":"2507.19478","version":1},"attestation_state":"computed","paper":{"title":"MMBench-GUI: Hierarchical Multi-Platform Evaluation Framework for GUI Agents","license":"http://creativecommons.org/licenses/by/4.0/","headline":"","cross_cats":["cs.CL"],"primary_cat":"cs.CV","authors_text":"Bowen Yang, Chenyu Yang, Haodong Duan, Jifeng Dai, Jingjing Xie, Jixuan Chen, Qingyun Li, Shiqian Su, Tianbao Xie, Weijie Su, Wei Shen, Weiyun Wang, Wenhai Wang, Xiangyu Yue, Xiangyu Zhao, Xiao Zhang, Xizhou Zhu, Xuan Dong, Xuehui Wang, Yanting Zhang, Yiqian Liu, Yuan Huang, Yue Yu, Zehao Li, Zhaoyang Liu, Zhe Chen, Zhenyu Wu, Zichen Ding","submitted_at":"2025-07-25T17:59:26Z","abstract_excerpt":"We introduce MMBench-GUI, a hierarchical benchmark for evaluating GUI automation agents across Windows, macOS, Linux, iOS, Android, and Web platforms. It comprises four levels: GUI Content Understanding, Element Grounding, Task Automation, and Task Collaboration, covering essential skills for GUI agents. In addition, we propose a novel Efficiency-Quality Area (EQA) metric to assess GUI agent execution efficiency in online automation scenarios. Through MMBench-GUI, we identify accurate visual grounding as a critical determinant of overall task success, emphasizing the substantial benefits of mo"},"verification_status":{"content_addressed":true,"pith_receipt":true,"author_attested":false,"weak_author_claims":0,"strong_author_claims":0,"externally_anchored":false,"storage_verified":false,"citation_signatures":0,"replication_records":0,"graph_snapshot":true,"references_resolved":false,"formal_links_present":false},"canonical_record":{"source":{"id":"2507.19478","kind":"arxiv","version":1},"metadata":{"license":"http://creativecommons.org/licenses/by/4.0/","primary_cat":"cs.CV","submitted_at":"2025-07-25T17:59:26Z","cross_cats_sorted":["cs.CL"],"title_canon_sha256":"ad345452a5d62325efbf4b63942310fe5c7f167b099132aad1f37da668e7553a","abstract_canon_sha256":"53be7f39310da7aa9920551e61651262a5b1aff55c35bf486a45e4a10a5ced07"},"schema_version":"1.0"},"receipt":{"kind":"pith_receipt","key_id":"pith-v1-2026-05","algorithm":"ed25519","signed_at":"2026-07-05T11:43:26.084649Z","signature_b64":"QnLSdQaRJHyg6nbxleqdeYo57voU6GKLtwyD5aTRU0e7vRF4Sob9McUlEfyTC+Dvu1aqA9hflTSlF5YoZWsVDA==","signed_message":"canonical_sha256_bytes","builder_version":"pith-number-builder-2026-05-17-v1","receipt_version":"0.3","canonical_sha256":"bbe041c3ba2ecc9fa1e5b3968a169e028df195688e9d15e88e2b358805e66801","last_reissued_at":"2026-07-05T11:43:26.084084Z","signature_status":"signed_v1","first_computed_at":"2026-07-05T11:43:26.084084Z","public_key_fingerprint":"8d4b5ee74e4693bcd1df2446408b0d54"},"graph_snapshot":{"paper":{"title":"MMBench-GUI: Hierarchical Multi-Platform Evaluation Framework for GUI Agents","license":"http://creativecommons.org/licenses/by/4.0/","headline":"","cross_cats":["cs.CL"],"primary_cat":"cs.CV","authors_text":"Bowen Yang, Chenyu Yang, Haodong Duan, Jifeng Dai, Jingjing Xie, Jixuan Chen, Qingyun Li, Shiqian Su, Tianbao Xie, Weijie Su, Wei Shen, Weiyun Wang, Wenhai Wang, Xiangyu Yue, Xiangyu Zhao, Xiao Zhang, Xizhou Zhu, Xuan Dong, Xuehui Wang, Yanting Zhang, Yiqian Liu, Yuan Huang, Yue Yu, Zehao Li, Zhaoyang Liu, Zhe Chen, Zhenyu Wu, Zichen Ding","submitted_at":"2025-07-25T17:59:26Z","abstract_excerpt":"We introduce MMBench-GUI, a hierarchical benchmark for evaluating GUI automation agents across Windows, macOS, Linux, iOS, Android, and Web platforms. It comprises four levels: GUI Content Understanding, Element Grounding, Task Automation, and Task Collaboration, covering essential skills for GUI agents. In addition, we propose a novel Efficiency-Quality Area (EQA) metric to assess GUI agent execution efficiency in online automation scenarios. Through MMBench-GUI, we identify accurate visual grounding as a critical determinant of overall task success, emphasizing the substantial benefits of mo"},"claims":{"count":0,"items":[],"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"source":{"id":"2507.19478","kind":"arxiv","version":1},"verdict":{"id":null,"model_set":{},"created_at":null,"strongest_claim":"","one_line_summary":"","pipeline_version":null,"weakest_assumption":"","pith_extraction_headline":""},"integrity":{"clean":true,"summary":{"advisory":0,"critical":0,"by_detector":{},"informational":0},"endpoint":"/pith/2507.19478/integrity.json","findings":[],"available":true,"detectors_run":[],"snapshot_sha256":"c28c3603d3b5d939e8dc4c7e95fa8dfce3d595e45f758748cecf8e644a296938"},"references":{"count":0,"sample":[],"resolved_work":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57","internal_anchors":0},"formal_canon":{"evidence_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"},"aliases":[{"alias_kind":"arxiv","alias_value":"2507.19478","created_at":"2026-07-05T11:43:26.084147+00:00"},{"alias_kind":"arxiv_version","alias_value":"2507.19478v1","created_at":"2026-07-05T11:43:26.084147+00:00"},{"alias_kind":"doi","alias_value":"10.48550/arxiv.2507.19478","created_at":"2026-07-05T11:43:26.084147+00:00"},{"alias_kind":"pith_short_12","alias_value":"XPQEDQ52F3GJ","created_at":"2026-07-05T11:43:26.084147+00:00"},{"alias_kind":"pith_short_16","alias_value":"XPQEDQ52F3GJ7IPF","created_at":"2026-07-05T11:43:26.084147+00:00"},{"alias_kind":"pith_short_8","alias_value":"XPQEDQ52","created_at":"2026-07-05T11:43:26.084147+00:00"}],"events":[],"event_summary":{},"paper_claims":[],"inbound_citations":{"count":12,"internal_anchor_count":0,"sample":[{"citing_arxiv_id":"2606.30084","citing_title":"One Forward Beats Two: InnerZoom for Accurate and Efficient GUI Grounding","ref_index":121,"is_internal_anchor":false},{"citing_arxiv_id":"2605.27761","citing_title":"AndroidDaily: A Verifiable Benchmark for Mobile GUI Agents on Real-World Closed-Source Applications","ref_index":45,"is_internal_anchor":false},{"citing_arxiv_id":"2605.19260","citing_title":"AQuaUI: Visual Token Reduction for GUI Agents with Adaptive Quadtrees","ref_index":29,"is_internal_anchor":false},{"citing_arxiv_id":"2509.07553","citing_title":"VeriOS: Query-Driven Proactive Human-Agent-GUI Interaction for Trustworthy OS Agents","ref_index":50,"is_internal_anchor":false},{"citing_arxiv_id":"2602.11724","citing_title":"WebTestPilot: Agentic End-to-End Web Testing against Natural Language Specification by Inferring Oracles with Symbolized GUI Elements","ref_index":72,"is_internal_anchor":false},{"citing_arxiv_id":"2605.12549","citing_title":"What Happens Before Decoding? Prefill Determines GUI Grounding in VLMs","ref_index":22,"is_internal_anchor":false},{"citing_arxiv_id":"2605.00642","citing_title":"Learn where to Click from Yourself: On-Policy Self-Distillation for GUI Grounding","ref_index":37,"is_internal_anchor":false},{"citing_arxiv_id":"2604.24348","citing_title":"OS-SPEAR: A Toolkit for the Safety, Performance,Efficiency, and Robustness Analysis of OS Agents","ref_index":44,"is_internal_anchor":false},{"citing_arxiv_id":"2605.00642","citing_title":"Learn where to Click from Yourself: On-Policy Self-Distillation for GUI Grounding","ref_index":37,"is_internal_anchor":false},{"citing_arxiv_id":"2604.21268","citing_title":"Measure Twice, Click Once: Co-evolving Proposer and Visual Critic via Reinforcement Learning for GUI Grounding","ref_index":9,"is_internal_anchor":false},{"citing_arxiv_id":"2605.06761","citing_title":"Weblica: Scalable and Reproducible Training Environments for Visual Web Agents","ref_index":37,"is_internal_anchor":false},{"citing_arxiv_id":"2508.18265","citing_title":"InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency","ref_index":146,"is_internal_anchor":false}]},"formal_canon":{"evidence_count":0,"sample":[],"anchors":[]},"links":{"html":"https://pith.science/pith/XPQEDQ52F3GJ7IPFWOLIUFU6AK","json":"https://pith.science/pith/XPQEDQ52F3GJ7IPFWOLIUFU6AK.json","graph_json":"https://pith.science/api/pith-number/XPQEDQ52F3GJ7IPFWOLIUFU6AK/graph.json","events_json":"https://pith.science/api/pith-number/XPQEDQ52F3GJ7IPFWOLIUFU6AK/events.json","paper":"https://pith.science/paper/XPQEDQ52"},"agent_actions":{"view_html":"https://pith.science/pith/XPQEDQ52F3GJ7IPFWOLIUFU6AK","download_json":"https://pith.science/pith/XPQEDQ52F3GJ7IPFWOLIUFU6AK.json","view_paper":"https://pith.science/paper/XPQEDQ52","resolve_alias":"https://pith.science/api/pith-number/resolve?arxiv=2507.19478&json=true","fetch_graph":"https://pith.science/api/pith-number/XPQEDQ52F3GJ7IPFWOLIUFU6AK/graph.json","fetch_events":"https://pith.science/api/pith-number/XPQEDQ52F3GJ7IPFWOLIUFU6AK/events.json","actions":{"anchor_timestamp":"https://pith.science/pith/XPQEDQ52F3GJ7IPFWOLIUFU6AK/action/timestamp_anchor","attest_storage":"https://pith.science/pith/XPQEDQ52F3GJ7IPFWOLIUFU6AK/action/storage_attestation","attest_author":"https://pith.science/pith/XPQEDQ52F3GJ7IPFWOLIUFU6AK/action/author_attestation","sign_citation":"https://pith.science/pith/XPQEDQ52F3GJ7IPFWOLIUFU6AK/action/citation_signature","submit_replication":"https://pith.science/pith/XPQEDQ52F3GJ7IPFWOLIUFU6AK/action/replication_record"}},"created_at":"2026-07-05T11:43:26.084147+00:00","updated_at":"2026-07-05T11:43:26.084147+00:00"}