{"record_type":"pith_number_record","schema_url":"https://pith.science/schemas/pith-number/v1.json","pith_number":"pith:2025:U7WVJYRVWGXVMEOFR2ZBNH2C7W","short_pith_number":"pith:U7WVJYRV","schema_version":"1.0","canonical_sha256":"a7ed54e235b1af5611c58eb2169f42fda975a3d671e2ec21519e73c469f7fc11","source":{"kind":"arxiv","id":"2506.07672","version":1},"attestation_state":"computed","paper":{"title":"MCPWorld: A Unified Benchmarking Testbed for API, GUI, and Hybrid Computer Use Agents","license":"http://creativecommons.org/licenses/by/4.0/","headline":"","cross_cats":[],"primary_cat":"cs.AI","authors_text":"Chen Peng, Jiajun Du, Mao Qin, Mengwei Xu, Qichen Qiu, Shangguang Wang, Shihe Wang, Xianqing Jia, Xinge Wang, Xin Yuan, Xu Han, Yexuan Yang, Yinxiao Chen, Yunhe Yan, Yuxuan Shan","submitted_at":"2025-06-09T11:50:33Z","abstract_excerpt":"(M)LLM-powered computer use agents (CUA) are emerging as a transformative technique to automate human-computer interaction. However, existing CUA benchmarks predominantly target GUI agents, whose evaluation methods are susceptible to UI changes and ignore function interactions exposed by application APIs, e.g., Model Context Protocol (MCP). To this end, we propose MCPWorld, the first automatic CUA testbed for API, GUI, and API-GUI hybrid agents. A key principle of MCPWorld is the use of \"white-box apps\", i.e., those with source code availability and can be revised/re-compiled as needed (e.g., "},"verification_status":{"content_addressed":true,"pith_receipt":true,"author_attested":false,"weak_author_claims":0,"strong_author_claims":0,"externally_anchored":false,"storage_verified":false,"citation_signatures":0,"replication_records":0,"graph_snapshot":true,"references_resolved":false,"formal_links_present":false},"canonical_record":{"source":{"id":"2506.07672","kind":"arxiv","version":1},"metadata":{"license":"http://creativecommons.org/licenses/by/4.0/","primary_cat":"cs.AI","submitted_at":"2025-06-09T11:50:33Z","cross_cats_sorted":[],"title_canon_sha256":"1a37e0c4e28bf068f3d78fb87cca7c9fb5c32b7c527d5309ebbf31135c9063a7","abstract_canon_sha256":"7894a46d152f30d8ce2fc407102a983a2b98a826442e96e67659893c2ec2841e"},"schema_version":"1.0"},"receipt":{"kind":"pith_receipt","key_id":"pith-v1-2026-05","algorithm":"ed25519","signed_at":"2026-07-05T11:18:31.314831Z","signature_b64":"yzTKzVP51ayuesWiq9atc3GQ2gkZlq2Jtocu37bM8k73s4rS3bloV9kMEIXxKFqxJUArZ8Nxk0g1WOuu9tdaDQ==","signed_message":"canonical_sha256_bytes","builder_version":"pith-number-builder-2026-05-17-v1","receipt_version":"0.3","canonical_sha256":"a7ed54e235b1af5611c58eb2169f42fda975a3d671e2ec21519e73c469f7fc11","last_reissued_at":"2026-07-05T11:18:31.314340Z","signature_status":"signed_v1","first_computed_at":"2026-07-05T11:18:31.314340Z","public_key_fingerprint":"8d4b5ee74e4693bcd1df2446408b0d54"},"graph_snapshot":{"paper":{"title":"MCPWorld: A Unified Benchmarking Testbed for API, GUI, and Hybrid Computer Use Agents","license":"http://creativecommons.org/licenses/by/4.0/","headline":"","cross_cats":[],"primary_cat":"cs.AI","authors_text":"Chen Peng, Jiajun Du, Mao Qin, Mengwei Xu, Qichen Qiu, Shangguang Wang, Shihe Wang, Xianqing Jia, Xinge Wang, Xin Yuan, Xu Han, Yexuan Yang, Yinxiao Chen, Yunhe Yan, Yuxuan Shan","submitted_at":"2025-06-09T11:50:33Z","abstract_excerpt":"(M)LLM-powered computer use agents (CUA) are emerging as a transformative technique to automate human-computer interaction. However, existing CUA benchmarks predominantly target GUI agents, whose evaluation methods are susceptible to UI changes and ignore function interactions exposed by application APIs, e.g., Model Context Protocol (MCP). To this end, we propose MCPWorld, the first automatic CUA testbed for API, GUI, and API-GUI hybrid agents. A key principle of MCPWorld is the use of \"white-box apps\", i.e., those with source code availability and can be revised/re-compiled as needed (e.g., "},"claims":{"count":0,"items":[],"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"source":{"id":"2506.07672","kind":"arxiv","version":1},"verdict":{"id":null,"model_set":{},"created_at":null,"strongest_claim":"","one_line_summary":"","pipeline_version":null,"weakest_assumption":"","pith_extraction_headline":""},"integrity":{"clean":true,"summary":{"advisory":0,"critical":0,"by_detector":{},"informational":0},"endpoint":"/pith/2506.07672/integrity.json","findings":[],"available":true,"detectors_run":[],"snapshot_sha256":"c28c3603d3b5d939e8dc4c7e95fa8dfce3d595e45f758748cecf8e644a296938"},"references":{"count":0,"sample":[],"resolved_work":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57","internal_anchors":0},"formal_canon":{"evidence_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"},"aliases":[{"alias_kind":"arxiv","alias_value":"2506.07672","created_at":"2026-07-05T11:18:31.314397+00:00"},{"alias_kind":"arxiv_version","alias_value":"2506.07672v1","created_at":"2026-07-05T11:18:31.314397+00:00"},{"alias_kind":"doi","alias_value":"10.48550/arxiv.2506.07672","created_at":"2026-07-05T11:18:31.314397+00:00"},{"alias_kind":"pith_short_12","alias_value":"U7WVJYRVWGXV","created_at":"2026-07-05T11:18:31.314397+00:00"},{"alias_kind":"pith_short_16","alias_value":"U7WVJYRVWGXVMEOF","created_at":"2026-07-05T11:18:31.314397+00:00"},{"alias_kind":"pith_short_8","alias_value":"U7WVJYRV","created_at":"2026-07-05T11:18:31.314397+00:00"}],"events":[],"event_summary":{},"paper_claims":[],"inbound_citations":{"count":6,"internal_anchor_count":0,"sample":[{"citing_arxiv_id":"2606.20683","citing_title":"From Question Answering to Task Completion: A Survey on Agent System and Harness Design","ref_index":147,"is_internal_anchor":false},{"citing_arxiv_id":"2605.10787","citing_title":"ComplexMCP: Evaluation of LLM Agents in Dynamic, Interdependent, and Large-Scale Tool Sandbox","ref_index":19,"is_internal_anchor":false},{"citing_arxiv_id":"2602.10139","citing_title":"Anonymization-Enhanced Privacy Protection for Mobile GUI Agents: Available but Invisible","ref_index":17,"is_internal_anchor":false},{"citing_arxiv_id":"2605.12481","citing_title":"ToolCUA: Towards Optimal GUI-Tool Path Orchestration for Computer Use Agents","ref_index":50,"is_internal_anchor":false},{"citing_arxiv_id":"2605.10787","citing_title":"ComplexMCP: Evaluation of LLM Agents in Dynamic, Interdependent, and Large-Scale Tool Sandbox","ref_index":19,"is_internal_anchor":false},{"citing_arxiv_id":"2604.09815","citing_title":"EE-MCP: Self-Evolving MCP-GUI Agents via Automated Environment Generation and Experience Learning","ref_index":13,"is_internal_anchor":false}]},"formal_canon":{"evidence_count":0,"sample":[],"anchors":[]},"links":{"html":"https://pith.science/pith/U7WVJYRVWGXVMEOFR2ZBNH2C7W","json":"https://pith.science/pith/U7WVJYRVWGXVMEOFR2ZBNH2C7W.json","graph_json":"https://pith.science/api/pith-number/U7WVJYRVWGXVMEOFR2ZBNH2C7W/graph.json","events_json":"https://pith.science/api/pith-number/U7WVJYRVWGXVMEOFR2ZBNH2C7W/events.json","paper":"https://pith.science/paper/U7WVJYRV"},"agent_actions":{"view_html":"https://pith.science/pith/U7WVJYRVWGXVMEOFR2ZBNH2C7W","download_json":"https://pith.science/pith/U7WVJYRVWGXVMEOFR2ZBNH2C7W.json","view_paper":"https://pith.science/paper/U7WVJYRV","resolve_alias":"https://pith.science/api/pith-number/resolve?arxiv=2506.07672&json=true","fetch_graph":"https://pith.science/api/pith-number/U7WVJYRVWGXVMEOFR2ZBNH2C7W/graph.json","fetch_events":"https://pith.science/api/pith-number/U7WVJYRVWGXVMEOFR2ZBNH2C7W/events.json","actions":{"anchor_timestamp":"https://pith.science/pith/U7WVJYRVWGXVMEOFR2ZBNH2C7W/action/timestamp_anchor","attest_storage":"https://pith.science/pith/U7WVJYRVWGXVMEOFR2ZBNH2C7W/action/storage_attestation","attest_author":"https://pith.science/pith/U7WVJYRVWGXVMEOFR2ZBNH2C7W/action/author_attestation","sign_citation":"https://pith.science/pith/U7WVJYRVWGXVMEOFR2ZBNH2C7W/action/citation_signature","submit_replication":"https://pith.science/pith/U7WVJYRVWGXVMEOFR2ZBNH2C7W/action/replication_record"}},"created_at":"2026-07-05T11:18:31.314397+00:00","updated_at":"2026-07-05T11:18:31.314397+00:00"}