Distillation from visual foundation models to lidar enables frame-wise indoor semantic segmentation without manual annotations, achieving up to 56% mIoU on pseudo labels and 36% on real labels.
Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pages=
5 Pith papers cite this work. Polarity classification is still indexing.
fields
cs.CV 5years
2026 5representative citing papers
Coarse four-group semantic color coding (RGBB) appended to point clouds before tokenization improves LLM-based structured indoor prediction on Structured3D, SpatialLM, and ARKitScenes, especially for openings and furniture instances.
MambaPanoptic is a fully Mamba-based panoptic segmentation model that uses MambaFPN for multi-scale features and a QuadMamba kernel generator to outperform PanopticDeepLab and PanopticFCN on Cityscapes and COCO while using fewer parameters than Mask2Former.
Key vectors of concept text tokens in multimodal diffusion transformers encode an 'omission signal'; amplifying the omission-minus-presence direction during early denoising reduces object omission on FLUX.1-Dev and SD3.5-Medium.
A ray-tracing pipeline aligns ground-level image pixels to outdated DEM rasters for real-time 3D terrain reconstruction in wildfire zones, validated primarily through a custom simulator.
citing papers explorer
-
Feasibility of Indoor Frame-Wise Lidar Semantic Segmentation via Distillation from Visual Foundation Model
Distillation from visual foundation models to lidar enables frame-wise indoor semantic segmentation without manual annotations, achieving up to 56% mIoU on pseudo labels and 36% on real labels.
-
Coarse Semantic Injection for LLM-Conditioned Structured Indoor Prediction
Coarse four-group semantic color coding (RGBB) appended to point clouds before tokenization improves LLM-based structured indoor prediction on Structured3D, SpatialLM, and ARKitScenes, especially for openings and furniture instances.
-
MambaPanoptic: A Vision Mamba-based Structured State Space Framework for Panoptic Segmentation
MambaPanoptic is a fully Mamba-based panoptic segmentation model that uses MambaFPN for multi-scale features and a QuadMamba kernel generator to outperform PanopticDeepLab and PanopticFCN on Cityscapes and COCO while using fewer parameters than Mask2Former.
-
Diagnosing and Correcting Concept Omission in Multimodal Diffusion Transformers
Key vectors of concept text tokens in multimodal diffusion transformers encode an 'omission signal'; amplifying the omission-minus-presence direction during early denoising reduces object omission on FLUX.1-Dev and SD3.5-Medium.
-
LTM: Large-scale Terrain Model for Wildfire-prone Landscapes
A ray-tracing pipeline aligns ground-level image pixels to outdated DEM rasters for real-time 3D terrain reconstruction in wildfire zones, validated primarily through a custom simulator.