{"record_type":"pith_number_record","schema_url":"https://pith.science/schemas/pith-number/v1.json","pith_number":"pith:2024:LHWCONZ27XNAONTVSMJMTQGYTM","short_pith_number":"pith:LHWCONZ2","schema_version":"1.0","canonical_sha256":"59ec27373afdda0736759312c9c0d89b32c4945a5436b1fd086da5b7fd1a2ee2","source":{"kind":"arxiv","id":"2411.04890","version":2},"attestation_state":"computed","paper":{"title":"GUI Agents with Foundation Models: A Comprehensive Survey","license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","headline":"","cross_cats":["cs.HC"],"primary_cat":"cs.AI","authors_text":"Bin Wang, Chuhan Wu, Jianye Hao, Jingxuan Chen, Kun Shao, Ruiming Tang, Shuai Wang, Shuai Yu, Weinan Gan, Weiwen Liu, Xingshan Zeng, Xinlong Hao, Yasheng Wang, Yuhan Che, Yuqi Zhou","submitted_at":"2024-11-07T17:28:10Z","abstract_excerpt":"Recent advances in foundation models, particularly Large Language Models (LLMs) and Multimodal Large Language Models (MLLMs), have facilitated the development of intelligent agents capable of performing complex tasks. By leveraging the ability of (M)LLMs to process and interpret Graphical User Interfaces (GUIs), these agents can autonomously execute user instructions, simulating human-like interactions such as clicking and typing. This survey consolidates recent research on (M)LLM-based GUI agents, highlighting key innovations in data resources, frameworks, and applications. We begin by review"},"verification_status":{"content_addressed":true,"pith_receipt":true,"author_attested":false,"weak_author_claims":0,"strong_author_claims":0,"externally_anchored":false,"storage_verified":false,"citation_signatures":0,"replication_records":0,"graph_snapshot":true,"references_resolved":false,"formal_links_present":false},"canonical_record":{"source":{"id":"2411.04890","kind":"arxiv","version":2},"metadata":{"license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","primary_cat":"cs.AI","submitted_at":"2024-11-07T17:28:10Z","cross_cats_sorted":["cs.HC"],"title_canon_sha256":"9a1da207e32ec36925572e1c2db971345e0195236d7ec5713b619250e193d0ce","abstract_canon_sha256":"48182740c277666eff7905e06eaa76061e5418bee2499cc4ab134224af001ab1"},"schema_version":"1.0"},"receipt":{"kind":"pith_receipt","key_id":"pith-v1-2026-05","algorithm":"ed25519","signed_at":"2026-07-05T10:13:44.658475Z","signature_b64":"+Nnqx++dk8ufeZczVihXNrbTx7TtNBD9WMbr63TfzmiUO6NdkZYLj/VrUmmPoRRaZNDFA011uB7l6iSwEdRbDg==","signed_message":"canonical_sha256_bytes","builder_version":"pith-number-builder-2026-05-17-v1","receipt_version":"0.3","canonical_sha256":"59ec27373afdda0736759312c9c0d89b32c4945a5436b1fd086da5b7fd1a2ee2","last_reissued_at":"2026-07-05T10:13:44.657974Z","signature_status":"signed_v1","first_computed_at":"2026-07-05T10:13:44.657974Z","public_key_fingerprint":"8d4b5ee74e4693bcd1df2446408b0d54"},"graph_snapshot":{"paper":{"title":"GUI Agents with Foundation Models: A Comprehensive Survey","license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","headline":"","cross_cats":["cs.HC"],"primary_cat":"cs.AI","authors_text":"Bin Wang, Chuhan Wu, Jianye Hao, Jingxuan Chen, Kun Shao, Ruiming Tang, Shuai Wang, Shuai Yu, Weinan Gan, Weiwen Liu, Xingshan Zeng, Xinlong Hao, Yasheng Wang, Yuhan Che, Yuqi Zhou","submitted_at":"2024-11-07T17:28:10Z","abstract_excerpt":"Recent advances in foundation models, particularly Large Language Models (LLMs) and Multimodal Large Language Models (MLLMs), have facilitated the development of intelligent agents capable of performing complex tasks. By leveraging the ability of (M)LLMs to process and interpret Graphical User Interfaces (GUIs), these agents can autonomously execute user instructions, simulating human-like interactions such as clicking and typing. This survey consolidates recent research on (M)LLM-based GUI agents, highlighting key innovations in data resources, frameworks, and applications. We begin by review"},"claims":{"count":0,"items":[],"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"source":{"id":"2411.04890","kind":"arxiv","version":2},"verdict":{"id":null,"model_set":{},"created_at":null,"strongest_claim":"","one_line_summary":"","pipeline_version":null,"weakest_assumption":"","pith_extraction_headline":""},"integrity":{"clean":true,"summary":{"advisory":0,"critical":0,"by_detector":{},"informational":0},"endpoint":"/pith/2411.04890/integrity.json","findings":[],"available":true,"detectors_run":[],"snapshot_sha256":"c28c3603d3b5d939e8dc4c7e95fa8dfce3d595e45f758748cecf8e644a296938"},"references":{"count":0,"sample":[],"resolved_work":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57","internal_anchors":0},"formal_canon":{"evidence_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"},"aliases":[{"alias_kind":"arxiv","alias_value":"2411.04890","created_at":"2026-07-05T10:13:44.658035+00:00"},{"alias_kind":"arxiv_version","alias_value":"2411.04890v2","created_at":"2026-07-05T10:13:44.658035+00:00"},{"alias_kind":"doi","alias_value":"10.48550/arxiv.2411.04890","created_at":"2026-07-05T10:13:44.658035+00:00"},{"alias_kind":"pith_short_12","alias_value":"LHWCONZ27XNA","created_at":"2026-07-05T10:13:44.658035+00:00"},{"alias_kind":"pith_short_16","alias_value":"LHWCONZ27XNAONTV","created_at":"2026-07-05T10:13:44.658035+00:00"},{"alias_kind":"pith_short_8","alias_value":"LHWCONZ2","created_at":"2026-07-05T10:13:44.658035+00:00"}],"events":[],"event_summary":{},"paper_claims":[],"inbound_citations":{"count":21,"internal_anchor_count":0,"sample":[{"citing_arxiv_id":"2606.27330","citing_title":"Empowering GUI Agents via Autonomous Experience Exploration and Hindsight Experience Utilization for Task Planning","ref_index":12,"is_internal_anchor":false},{"citing_arxiv_id":"2606.26935","citing_title":"Where Do CoT Training Gains Land in LLM based Agents?","ref_index":24,"is_internal_anchor":false},{"citing_arxiv_id":"2606.26918","citing_title":"Diagnosing Task Insensitivity in Language Agents","ref_index":14,"is_internal_anchor":false},{"citing_arxiv_id":"2606.10522","citing_title":"GUI-AC: Enhancing Continual Learning in GUI Agents","ref_index":1,"is_internal_anchor":false},{"citing_arxiv_id":"2606.07027","citing_title":"StainFlow: Entity-Stain Tracking and Evidence Linking for Process Rewards in GUI Agents","ref_index":30,"is_internal_anchor":false},{"citing_arxiv_id":"2605.31365","citing_title":"Learning to Adapt: Self-Improving Web Agent via Cognitive-Aware Exploration","ref_index":26,"is_internal_anchor":false},{"citing_arxiv_id":"2605.28629","citing_title":"Mobile-Aptus: Confidence-Driven Proactive and Robust Interaction in MLLM-based Mobile-Using Agents","ref_index":17,"is_internal_anchor":false},{"citing_arxiv_id":"2501.16150","citing_title":"A Comprehensive Survey of Agents for Computer Use: Foundations, Challenges, and Future Directions","ref_index":162,"is_internal_anchor":false},{"citing_arxiv_id":"2605.18048","citing_title":"DocOS: Towards Proactive Document-Guided Actions in GUI Agents","ref_index":79,"is_internal_anchor":false},{"citing_arxiv_id":"2411.18279","citing_title":"Large Language Model-Brained GUI Agents: A Survey","ref_index":63,"is_internal_anchor":false},{"citing_arxiv_id":"2506.09373","citing_title":"LPO: Towards Accurate GUI Agent Interaction via Location Preference Optimization","ref_index":21,"is_internal_anchor":false},{"citing_arxiv_id":"2510.24168","citing_title":"MGA: Memory-Driven GUI Agent for Observation-Centric Interaction","ref_index":29,"is_internal_anchor":false},{"citing_arxiv_id":"2604.02345","citing_title":"UI-Oceanus: Scaling GUI Agents with Synthetic Environmental Dynamics","ref_index":44,"is_internal_anchor":false},{"citing_arxiv_id":"2603.26041","citing_title":"Rethinking Token Pruning for Historical Screenshots in GUI Visual Agents: Semantic, Spatial, and Temporal Perspectives","ref_index":26,"is_internal_anchor":false},{"citing_arxiv_id":"2605.12549","citing_title":"What Happens Before Decoding? Prefill Determines GUI Grounding in VLMs","ref_index":3,"is_internal_anchor":false},{"citing_arxiv_id":"2509.02544","citing_title":"UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning","ref_index":71,"is_internal_anchor":false},{"citing_arxiv_id":"2604.24348","citing_title":"OS-SPEAR: A Toolkit for the Safety, Performance,Efficiency, and Robustness Analysis of OS Agents","ref_index":18,"is_internal_anchor":false},{"citing_arxiv_id":"2605.01208","citing_title":"Faithful Mobile GUI Agents with Guided Advantage Estimator","ref_index":16,"is_internal_anchor":false},{"citing_arxiv_id":"2604.11259","citing_title":"Mobile GUI Agent Privacy Personalization with Trajectory Induced Preference Optimization","ref_index":32,"is_internal_anchor":false},{"citing_arxiv_id":"2605.07110","citing_title":"Securing Computer-Use Agents: A Unified Architecture-Lifecycle Framework for Deployment-Grounded Reliability","ref_index":24,"is_internal_anchor":false},{"citing_arxiv_id":"2604.05719","citing_title":"Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing","ref_index":111,"is_internal_anchor":false}]},"formal_canon":{"evidence_count":0,"sample":[],"anchors":[]},"links":{"html":"https://pith.science/pith/LHWCONZ27XNAONTVSMJMTQGYTM","json":"https://pith.science/pith/LHWCONZ27XNAONTVSMJMTQGYTM.json","graph_json":"https://pith.science/api/pith-number/LHWCONZ27XNAONTVSMJMTQGYTM/graph.json","events_json":"https://pith.science/api/pith-number/LHWCONZ27XNAONTVSMJMTQGYTM/events.json","paper":"https://pith.science/paper/LHWCONZ2"},"agent_actions":{"view_html":"https://pith.science/pith/LHWCONZ27XNAONTVSMJMTQGYTM","download_json":"https://pith.science/pith/LHWCONZ27XNAONTVSMJMTQGYTM.json","view_paper":"https://pith.science/paper/LHWCONZ2","resolve_alias":"https://pith.science/api/pith-number/resolve?arxiv=2411.04890&json=true","fetch_graph":"https://pith.science/api/pith-number/LHWCONZ27XNAONTVSMJMTQGYTM/graph.json","fetch_events":"https://pith.science/api/pith-number/LHWCONZ27XNAONTVSMJMTQGYTM/events.json","actions":{"anchor_timestamp":"https://pith.science/pith/LHWCONZ27XNAONTVSMJMTQGYTM/action/timestamp_anchor","attest_storage":"https://pith.science/pith/LHWCONZ27XNAONTVSMJMTQGYTM/action/storage_attestation","attest_author":"https://pith.science/pith/LHWCONZ27XNAONTVSMJMTQGYTM/action/author_attestation","sign_citation":"https://pith.science/pith/LHWCONZ27XNAONTVSMJMTQGYTM/action/citation_signature","submit_replication":"https://pith.science/pith/LHWCONZ27XNAONTVSMJMTQGYTM/action/replication_record"}},"created_at":"2026-07-05T10:13:44.658035+00:00","updated_at":"2026-07-05T10:13:44.658035+00:00"}