Topology-aware attention over hierarchical scene graphs lets a 3D-LLM ground, caption, and answer questions across multi-room homes, with large gains on a new HM3D benchmark.
3ur-llm: An end-to-end multimodal large language model for 3d scene understanding.IEEE Transactions on Multimedia, 27: 2899–2911, 2025
1 Pith paper cite this work, alongside 10 external citations. Polarity classification is still indexing.
1
Pith paper citing it
10
external citations · external index
fields
cs.CV 1years
2026 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
CAIRN: Cross-Room 3D Scene Understanding with Topology-Aware Large Multimodal Models
Topology-aware attention over hierarchical scene graphs lets a 3D-LLM ground, caption, and answer questions across multi-room homes, with large gains on a new HM3D benchmark.