Pith. sign in

Texthawk: Exploring efficient fine- grained perception of multimodal large language models

7 Pith papers cite this work. Polarity classification is still indexing.

7 Pith papers citing it

verdicts

UNVERDICTED 7

representative citing papers

Grounded 3D-Aware Spatial Vision-Language Modeling

cs.CV · 2026-05-28 · unverdicted · novelty 5.0

GR3D is a VLM that combines explicit 2D, implicit 2D, and monocular 3D grounding mechanisms to improve performance on spatial understanding benchmarks.

citing papers explorer

Showing 7 of 7 citing papers.