TinyGiantVLM, a 64M-parameter RGB-D vision-language model with two-phase training, reached 5th place on the AI City Challenge 2025 warehouse spatial reasoning track.
Chat-scene: Bridging 3d scene and large language models with object identifiers
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
background 1
citation-polarity summary
fields
cs.CV 1years
2025 1verdicts
CONDITIONAL 1roles
background 1polarities
unclear 1representative citing papers
citing papers explorer
-
TinyGiantVLM: A Lightweight Vision-Language Architecture for Spatial Reasoning under Resource Constraints
TinyGiantVLM, a 64M-parameter RGB-D vision-language model with two-phase training, reached 5th place on the AI City Challenge 2025 warehouse spatial reasoning track.