{"record_type":"pith_number_record","schema_url":"https://pith.science/schemas/pith-number/v1.json","pith_number":"pith:2025:43KOC5XGLVFO53MQODY3PO4Q6M","short_pith_number":"pith:43KOC5XG","schema_version":"1.0","canonical_sha256":"e6d4e176e65d4aeeed9070f1b7bb90f31925a22f8291a2c8993109f523a6833d","source":{"kind":"arxiv","id":"2507.00008","version":2},"attestation_state":"computed","paper":{"title":"DiMo-GUI: Advancing Test-time Scaling in GUI Grounding via Modality-Aware Visual Reasoning","license":"http://creativecommons.org/licenses/by/4.0/","headline":"","cross_cats":["cs.CV","cs.HC"],"primary_cat":"cs.AI","authors_text":"Chang Liu, Hang Wu, Hongkai Chen, Ming-Hsuan Yang, Qingwen Ye, Yiwei Wang, Yujun Cai","submitted_at":"2025-06-12T03:13:21Z","abstract_excerpt":"Grounding natural language queries in graphical user interfaces (GUIs) poses unique challenges due to the diversity of visual elements, spatial clutter, and the ambiguity of language. In this paper, we introduce DiMo-GUI, a training-free framework for GUI grounding that leverages two core strategies: dynamic visual grounding and modality-aware optimization. Instead of treating the GUI as a monolithic image, our method splits the input into textual elements and iconic elements, allowing the model to reason over each modality independently using general-purpose vision-language models. When predi"},"verification_status":{"content_addressed":true,"pith_receipt":true,"author_attested":false,"weak_author_claims":0,"strong_author_claims":0,"externally_anchored":false,"storage_verified":false,"citation_signatures":0,"replication_records":0,"graph_snapshot":true,"references_resolved":false,"formal_links_present":false},"canonical_record":{"source":{"id":"2507.00008","kind":"arxiv","version":2},"metadata":{"license":"http://creativecommons.org/licenses/by/4.0/","primary_cat":"cs.AI","submitted_at":"2025-06-12T03:13:21Z","cross_cats_sorted":["cs.CV","cs.HC"],"title_canon_sha256":"a5533d5e651a8f12987c0640741875c3bfe8f917cda699db8420dbb949fce218","abstract_canon_sha256":"0875240bf4a2c08e2006afbe0398466cbcdd9f5fc62ab27ab5f6d9cfd1179ea1"},"schema_version":"1.0"},"receipt":{"kind":"pith_receipt","key_id":"pith-v1-2026-05","algorithm":"ed25519","signed_at":"2026-07-05T12:05:18.336176Z","signature_b64":"tT+aCyH5h6mg/QIHq8XAv32Ufsv+iMyVA6dmIbRk5/9LYclmg8nlEhDbZMNEbt5QjcwExj7iZZxtYR5io5miBw==","signed_message":"canonical_sha256_bytes","builder_version":"pith-number-builder-2026-05-17-v1","receipt_version":"0.3","canonical_sha256":"e6d4e176e65d4aeeed9070f1b7bb90f31925a22f8291a2c8993109f523a6833d","last_reissued_at":"2026-07-05T12:05:18.334719Z","signature_status":"signed_v1","first_computed_at":"2026-07-05T12:05:18.334719Z","public_key_fingerprint":"8d4b5ee74e4693bcd1df2446408b0d54"},"graph_snapshot":{"paper":{"title":"DiMo-GUI: Advancing Test-time Scaling in GUI Grounding via Modality-Aware Visual Reasoning","license":"http://creativecommons.org/licenses/by/4.0/","headline":"","cross_cats":["cs.CV","cs.HC"],"primary_cat":"cs.AI","authors_text":"Chang Liu, Hang Wu, Hongkai Chen, Ming-Hsuan Yang, Qingwen Ye, Yiwei Wang, Yujun Cai","submitted_at":"2025-06-12T03:13:21Z","abstract_excerpt":"Grounding natural language queries in graphical user interfaces (GUIs) poses unique challenges due to the diversity of visual elements, spatial clutter, and the ambiguity of language. In this paper, we introduce DiMo-GUI, a training-free framework for GUI grounding that leverages two core strategies: dynamic visual grounding and modality-aware optimization. Instead of treating the GUI as a monolithic image, our method splits the input into textual elements and iconic elements, allowing the model to reason over each modality independently using general-purpose vision-language models. When predi"},"claims":{"count":0,"items":[],"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"source":{"id":"2507.00008","kind":"arxiv","version":2},"verdict":{"id":null,"model_set":{},"created_at":null,"strongest_claim":"","one_line_summary":"","pipeline_version":null,"weakest_assumption":"","pith_extraction_headline":""},"integrity":{"clean":true,"summary":{"advisory":0,"critical":0,"by_detector":{},"informational":0},"endpoint":"/pith/2507.00008/integrity.json","findings":[],"available":true,"detectors_run":[],"snapshot_sha256":"c28c3603d3b5d939e8dc4c7e95fa8dfce3d595e45f758748cecf8e644a296938"},"references":{"count":0,"sample":[],"resolved_work":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57","internal_anchors":0},"formal_canon":{"evidence_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"},"aliases":[{"alias_kind":"arxiv","alias_value":"2507.00008","created_at":"2026-07-05T12:05:18.334795+00:00"},{"alias_kind":"arxiv_version","alias_value":"2507.00008v2","created_at":"2026-07-05T12:05:18.334795+00:00"},{"alias_kind":"doi","alias_value":"10.48550/arxiv.2507.00008","created_at":"2026-07-05T12:05:18.334795+00:00"},{"alias_kind":"pith_short_12","alias_value":"43KOC5XGLVFO","created_at":"2026-07-05T12:05:18.334795+00:00"},{"alias_kind":"pith_short_16","alias_value":"43KOC5XGLVFO53MQ","created_at":"2026-07-05T12:05:18.334795+00:00"},{"alias_kind":"pith_short_8","alias_value":"43KOC5XG","created_at":"2026-07-05T12:05:18.334795+00:00"}],"events":[],"event_summary":{},"paper_claims":[],"inbound_citations":{"count":8,"internal_anchor_count":0,"sample":[{"citing_arxiv_id":"2606.08231","citing_title":"Test-Time Scaling in Multimodal Foundation Models: A Comprehensive Survey of Generation and Reasoning","ref_index":89,"is_internal_anchor":false},{"citing_arxiv_id":"2606.04046","citing_title":"Dive into the Scene: Breaking the Perceptual Bottleneck in Vision-Language Decision Making via Focus Plan Generation","ref_index":30,"is_internal_anchor":false},{"citing_arxiv_id":"2605.30884","citing_title":"GUI-C$^2$: Coarse-to-Fine GUI Grounding via Difficulty-Aware Reinforcement Learning","ref_index":36,"is_internal_anchor":false},{"citing_arxiv_id":"2605.15542","citing_title":"DRS-GUI: Dynamic Region Search for Training-Free GUI Grounding","ref_index":40,"is_internal_anchor":false},{"citing_arxiv_id":"2509.07553","citing_title":"VeriOS: Query-Driven Proactive Human-Agent-GUI Interaction for Trustworthy OS Agents","ref_index":53,"is_internal_anchor":false},{"citing_arxiv_id":"2605.06664","citing_title":"BAMI: Training-Free Bias Mitigation in GUI Grounding","ref_index":31,"is_internal_anchor":false},{"citing_arxiv_id":"2604.21268","citing_title":"Measure Twice, Click Once: Co-evolving Proposer and Visual Critic via Reinforcement Learning for GUI Grounding","ref_index":83,"is_internal_anchor":false},{"citing_arxiv_id":"2605.02630","citing_title":"AutoFocus: Uncertainty-Aware Active Visual Search for GUI Grounding","ref_index":34,"is_internal_anchor":false}]},"formal_canon":{"evidence_count":0,"sample":[],"anchors":[]},"links":{"html":"https://pith.science/pith/43KOC5XGLVFO53MQODY3PO4Q6M","json":"https://pith.science/pith/43KOC5XGLVFO53MQODY3PO4Q6M.json","graph_json":"https://pith.science/api/pith-number/43KOC5XGLVFO53MQODY3PO4Q6M/graph.json","events_json":"https://pith.science/api/pith-number/43KOC5XGLVFO53MQODY3PO4Q6M/events.json","paper":"https://pith.science/paper/43KOC5XG"},"agent_actions":{"view_html":"https://pith.science/pith/43KOC5XGLVFO53MQODY3PO4Q6M","download_json":"https://pith.science/pith/43KOC5XGLVFO53MQODY3PO4Q6M.json","view_paper":"https://pith.science/paper/43KOC5XG","resolve_alias":"https://pith.science/api/pith-number/resolve?arxiv=2507.00008&json=true","fetch_graph":"https://pith.science/api/pith-number/43KOC5XGLVFO53MQODY3PO4Q6M/graph.json","fetch_events":"https://pith.science/api/pith-number/43KOC5XGLVFO53MQODY3PO4Q6M/events.json","actions":{"anchor_timestamp":"https://pith.science/pith/43KOC5XGLVFO53MQODY3PO4Q6M/action/timestamp_anchor","attest_storage":"https://pith.science/pith/43KOC5XGLVFO53MQODY3PO4Q6M/action/storage_attestation","attest_author":"https://pith.science/pith/43KOC5XGLVFO53MQODY3PO4Q6M/action/author_attestation","sign_citation":"https://pith.science/pith/43KOC5XGLVFO53MQODY3PO4Q6M/action/citation_signature","submit_replication":"https://pith.science/pith/43KOC5XGLVFO53MQODY3PO4Q6M/action/replication_record"}},"created_at":"2026-07-05T12:05:18.334795+00:00","updated_at":"2026-07-05T12:05:18.334795+00:00"}