FALCON embeds learned register tokens in a vision transformer to compress high-resolution visual representations and enable cross-crop interaction, achieving strong benchmark scores with 9 times fewer visual tokens.
Less is more: Empowering gui agent with context- aware simplification
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
background 1
citation-polarity summary
fields
cs.CV 1years
2025 1verdicts
CONDITIONAL 1roles
background 1polarities
unclear 1representative citing papers
citing papers explorer
-
FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers
FALCON embeds learned register tokens in a vision transformer to compress high-resolution visual representations and enable cross-crop interaction, achieving strong benchmark scores with 9 times fewer visual tokens.