HiSC compresses visual tokens for 3D VLMs by graph-based merging before inference and hierarchical spatial clustering pruning inside the LLM, achieving about 90% token reduction while retaining roughly 92% of original performance.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CV 1years
2026 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
HiSC: Hierarchical Spatial Clustering Token Compression for Efficient 3D Scene Understanding
HiSC compresses visual tokens for 3D VLMs by graph-based merging before inference and hierarchical spatial clustering pruning inside the LLM, achieving about 90% token reduction while retaining roughly 92% of original performance.