A training-free agentic framework, combining SAM-based region parsing with spatially enhanced element descriptions, improves LVLM action grounding and GUI referring on multiple benchmarks.
org/CorpusID:267211622
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CV 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
TRISHUL: Towards Region Identification and Screen Hierarchy Understanding for Large VLM based GUI Agents
A training-free agentic framework, combining SAM-based region parsing with spatially enhanced element descriptions, improves LVLM action grounding and GUI referring on multiple benchmarks.