{"record_type":"pith_number_record","schema_url":"https://pith.science/schemas/pith-number/v1.json","pith_number":"pith:2025:FNFSF2DYPSORWJRDRNE4CDIHZU","short_pith_number":"pith:FNFSF2DY","schema_version":"1.0","canonical_sha256":"2b4b22e8787c9d1b26238b49c10d07cd132f6f8f870590612fc647b47888ac2e","source":{"kind":"arxiv","id":"2506.03143","version":1},"attestation_state":"computed","paper":{"title":"GUI-Actor: Coordinate-Free Visual Grounding for GUI Agents","license":"http://creativecommons.org/licenses/by/4.0/","headline":"","cross_cats":["cs.AI","cs.CV"],"primary_cat":"cs.CL","authors_text":"Baolin Peng, Bo Qiao, Chaoyun Zhang, Dongmei Zhang, Huan Zhang, Huiqiang Jiang, Jianbing Zhang, Jianfeng Gao, Jian Mu, Jianwei Yang, Kanzhi Cheng, Lars Liden, Qianhui Wu, Qingwei Lin, Reuben Tan, Rui Yang, Si Qin, Tong Zhang","submitted_at":"2025-06-03T17:59:08Z","abstract_excerpt":"One of the principal challenges in building VLM-powered GUI agents is visual grounding, i.e., localizing the appropriate screen region for action execution based on both the visual content and the textual plans. Most existing work formulates this as a text-based coordinate generation task. However, these approaches suffer from several limitations: weak spatial-semantic alignment, inability to handle ambiguous supervision targets, and a mismatch between the dense nature of screen coordinates and the coarse, patch-level granularity of visual features extracted by models like Vision Transformers."},"verification_status":{"content_addressed":true,"pith_receipt":true,"author_attested":false,"weak_author_claims":0,"strong_author_claims":0,"externally_anchored":false,"storage_verified":false,"citation_signatures":0,"replication_records":0,"graph_snapshot":true,"references_resolved":false,"formal_links_present":false},"canonical_record":{"source":{"id":"2506.03143","kind":"arxiv","version":1},"metadata":{"license":"http://creativecommons.org/licenses/by/4.0/","primary_cat":"cs.CL","submitted_at":"2025-06-03T17:59:08Z","cross_cats_sorted":["cs.AI","cs.CV"],"title_canon_sha256":"25d0f9eeb0ab44111c578d06d95378387673092f383acd075023cbe5b1ce9049","abstract_canon_sha256":"97ce6041a1b2d9fb9d6b4b29700398d4e562ebcbec0123ad4afbcfbdba166564"},"schema_version":"1.0"},"receipt":{"kind":"pith_receipt","key_id":"pith-v1-2026-05","algorithm":"ed25519","signed_at":"2026-07-05T11:15:15.676523Z","signature_b64":"FZP03/Ra4h0Wzr2UKbZVciUulSM2cBHmATQ2ADeZPR+XfokKohkJkJGGh7j4T5McnqmFp/SwYFtTsnVpY1oaDA==","signed_message":"canonical_sha256_bytes","builder_version":"pith-number-builder-2026-05-17-v1","receipt_version":"0.3","canonical_sha256":"2b4b22e8787c9d1b26238b49c10d07cd132f6f8f870590612fc647b47888ac2e","last_reissued_at":"2026-07-05T11:15:15.676015Z","signature_status":"signed_v1","first_computed_at":"2026-07-05T11:15:15.676015Z","public_key_fingerprint":"8d4b5ee74e4693bcd1df2446408b0d54"},"graph_snapshot":{"paper":{"title":"GUI-Actor: Coordinate-Free Visual Grounding for GUI Agents","license":"http://creativecommons.org/licenses/by/4.0/","headline":"","cross_cats":["cs.AI","cs.CV"],"primary_cat":"cs.CL","authors_text":"Baolin Peng, Bo Qiao, Chaoyun Zhang, Dongmei Zhang, Huan Zhang, Huiqiang Jiang, Jianbing Zhang, Jianfeng Gao, Jian Mu, Jianwei Yang, Kanzhi Cheng, Lars Liden, Qianhui Wu, Qingwei Lin, Reuben Tan, Rui Yang, Si Qin, Tong Zhang","submitted_at":"2025-06-03T17:59:08Z","abstract_excerpt":"One of the principal challenges in building VLM-powered GUI agents is visual grounding, i.e., localizing the appropriate screen region for action execution based on both the visual content and the textual plans. Most existing work formulates this as a text-based coordinate generation task. However, these approaches suffer from several limitations: weak spatial-semantic alignment, inability to handle ambiguous supervision targets, and a mismatch between the dense nature of screen coordinates and the coarse, patch-level granularity of visual features extracted by models like Vision Transformers."},"claims":{"count":0,"items":[],"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"source":{"id":"2506.03143","kind":"arxiv","version":1},"verdict":{"id":null,"model_set":{},"created_at":null,"strongest_claim":"","one_line_summary":"","pipeline_version":null,"weakest_assumption":"","pith_extraction_headline":""},"integrity":{"clean":true,"summary":{"advisory":0,"critical":0,"by_detector":{},"informational":0},"endpoint":"/pith/2506.03143/integrity.json","findings":[],"available":true,"detectors_run":[],"snapshot_sha256":"c28c3603d3b5d939e8dc4c7e95fa8dfce3d595e45f758748cecf8e644a296938"},"references":{"count":0,"sample":[],"resolved_work":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57","internal_anchors":0},"formal_canon":{"evidence_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"},"aliases":[{"alias_kind":"arxiv","alias_value":"2506.03143","created_at":"2026-07-05T11:15:15.676077+00:00"},{"alias_kind":"arxiv_version","alias_value":"2506.03143v1","created_at":"2026-07-05T11:15:15.676077+00:00"},{"alias_kind":"doi","alias_value":"10.48550/arxiv.2506.03143","created_at":"2026-07-05T11:15:15.676077+00:00"},{"alias_kind":"pith_short_12","alias_value":"FNFSF2DYPSOR","created_at":"2026-07-05T11:15:15.676077+00:00"},{"alias_kind":"pith_short_16","alias_value":"FNFSF2DYPSORWJRD","created_at":"2026-07-05T11:15:15.676077+00:00"},{"alias_kind":"pith_short_8","alias_value":"FNFSF2DY","created_at":"2026-07-05T11:15:15.676077+00:00"}],"events":[],"event_summary":{},"paper_claims":[],"inbound_citations":{"count":21,"internal_anchor_count":0,"sample":[{"citing_arxiv_id":"2606.02031","citing_title":"OpenWebRL: Demystifying Online Multi-turn Reinforcement Learning for Visual Web Agents","ref_index":47,"is_internal_anchor":false},{"citing_arxiv_id":"2606.29705","citing_title":"GUICrafter: Weakly-Supervised GUI Agent Leveraging Massive Unannotated Screenshots","ref_index":42,"is_internal_anchor":false},{"citing_arxiv_id":"2605.30884","citing_title":"GUI-C$^2$: Coarse-to-Fine GUI Grounding via Difficulty-Aware Reinforcement Learning","ref_index":37,"is_internal_anchor":false},{"citing_arxiv_id":"2606.06322","citing_title":"DragOn: A Benchmark and Dataset for Drag-Based GUI Interactions","ref_index":19,"is_internal_anchor":false},{"citing_arxiv_id":"2605.17933","citing_title":"AtlasVA: Self-Evolving Visual Skill Memory for Teacher-Free VLM Agents","ref_index":47,"is_internal_anchor":false},{"citing_arxiv_id":"2605.14311","citing_title":"Beyond Binary: Reframing GUI Critique as Continuous Semantic Alignment","ref_index":115,"is_internal_anchor":false},{"citing_arxiv_id":"2605.15542","citing_title":"DRS-GUI: Dynamic Region Search for Training-Free GUI Grounding","ref_index":42,"is_internal_anchor":false},{"citing_arxiv_id":"2509.21816","citing_title":"From Task to Tutorial: An Automated GUI Framework for Excel Tutorial Document and Video Creation","ref_index":24,"is_internal_anchor":false},{"citing_arxiv_id":"2605.14311","citing_title":"Beyond Binary: Reframing GUI Critique as Continuous Semantic Alignment","ref_index":115,"is_internal_anchor":false},{"citing_arxiv_id":"2605.11212","citing_title":"ReVision: Scaling Computer-Use Agents via Temporal Visual Redundancy Reduction","ref_index":22,"is_internal_anchor":false},{"citing_arxiv_id":"2602.12430","citing_title":"Agent Skills for Large Language Models: Architecture, Acquisition, Security, and the Path Forward","ref_index":29,"is_internal_anchor":false},{"citing_arxiv_id":"2605.11212","citing_title":"ReVision: Scaling Computer-Use Agents via Temporal Visual Redundancy Reduction","ref_index":22,"is_internal_anchor":false},{"citing_arxiv_id":"2605.00642","citing_title":"Learn where to Click from Yourself: On-Policy Self-Distillation for GUI Grounding","ref_index":39,"is_internal_anchor":false},{"citing_arxiv_id":"2605.00642","citing_title":"Learn where to Click from Yourself: On-Policy Self-Distillation for GUI Grounding","ref_index":39,"is_internal_anchor":false},{"citing_arxiv_id":"2604.21268","citing_title":"Measure Twice, Click Once: Co-evolving Proposer and Visual Critic via Reinforcement Learning for GUI Grounding","ref_index":75,"is_internal_anchor":false},{"citing_arxiv_id":"2604.13019","citing_title":"PrecisionCUA: Iterative Visual Refinement for Pixel-Precise Cursor Grounding in Code Editors","ref_index":14,"is_internal_anchor":false},{"citing_arxiv_id":"2604.08516","citing_title":"MolmoWeb: Open Visual Web Agent and Open Data for the Open Web","ref_index":57,"is_internal_anchor":false},{"citing_arxiv_id":"2605.07505","citing_title":"LiteGUI: Distilling Compact GUI Agents with Reinforcement Learning","ref_index":20,"is_internal_anchor":false},{"citing_arxiv_id":"2604.15376","citing_title":"Zoom Consistency: A Free Confidence Signal in Multi-Step Visual Grounding Pipelines","ref_index":9,"is_internal_anchor":false},{"citing_arxiv_id":"2604.20940","citing_title":"Sema: Semantic Transport for Real-Time Multimodal Agents","ref_index":48,"is_internal_anchor":false},{"citing_arxiv_id":"2605.02630","citing_title":"AutoFocus: Uncertainty-Aware Active Visual Search for GUI Grounding","ref_index":35,"is_internal_anchor":false}]},"formal_canon":{"evidence_count":0,"sample":[],"anchors":[]},"links":{"html":"https://pith.science/pith/FNFSF2DYPSORWJRDRNE4CDIHZU","json":"https://pith.science/pith/FNFSF2DYPSORWJRDRNE4CDIHZU.json","graph_json":"https://pith.science/api/pith-number/FNFSF2DYPSORWJRDRNE4CDIHZU/graph.json","events_json":"https://pith.science/api/pith-number/FNFSF2DYPSORWJRDRNE4CDIHZU/events.json","paper":"https://pith.science/paper/FNFSF2DY"},"agent_actions":{"view_html":"https://pith.science/pith/FNFSF2DYPSORWJRDRNE4CDIHZU","download_json":"https://pith.science/pith/FNFSF2DYPSORWJRDRNE4CDIHZU.json","view_paper":"https://pith.science/paper/FNFSF2DY","resolve_alias":"https://pith.science/api/pith-number/resolve?arxiv=2506.03143&json=true","fetch_graph":"https://pith.science/api/pith-number/FNFSF2DYPSORWJRDRNE4CDIHZU/graph.json","fetch_events":"https://pith.science/api/pith-number/FNFSF2DYPSORWJRDRNE4CDIHZU/events.json","actions":{"anchor_timestamp":"https://pith.science/pith/FNFSF2DYPSORWJRDRNE4CDIHZU/action/timestamp_anchor","attest_storage":"https://pith.science/pith/FNFSF2DYPSORWJRDRNE4CDIHZU/action/storage_attestation","attest_author":"https://pith.science/pith/FNFSF2DYPSORWJRDRNE4CDIHZU/action/author_attestation","sign_citation":"https://pith.science/pith/FNFSF2DYPSORWJRDRNE4CDIHZU/action/citation_signature","submit_replication":"https://pith.science/pith/FNFSF2DYPSORWJRDRNE4CDIHZU/action/replication_record"}},"created_at":"2026-07-05T11:15:15.676077+00:00","updated_at":"2026-07-05T11:15:15.676077+00:00"}