OpenMaskDINO3D reports state-of-the-art 3D reasoning segmentation with a LISA-style SEG token and object identifiers, but uses Mask3D pseudo-labels as ground truth and lacks released code.
VL-Fields: Towards Language-Grounded Neural Implicit Spatial Representations
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
We present Visual-Language Fields (VL-Fields), a neural implicit spatial representation that enables open-vocabulary semantic queries. Our model encodes and fuses the geometry of a scene with vision-language trained latent features by distilling information from a language-driven segmentation model. VL-Fields is trained without requiring any prior knowledge of the scene object classes, which makes it a promising representation for the field of robotics. Our model outperformed the similar CLIP-Fields model in the task of semantic segmentation by almost 10%.
citation-role summary
citation-polarity summary
fields
cs.CV 1years
2025 1verdicts
REJECT 1roles
background 1polarities
background 1representative citing papers
citing papers explorer
-
OpenMaskDINO3D : Reasoning 3D Segmentation via Large Language Model
OpenMaskDINO3D reports state-of-the-art 3D reasoning segmentation with a LISA-style SEG token and object identifiers, but uses Mask3D pseudo-labels as ground truth and lacks released code.