{"work":{"id":"9a869478-2848-4347-9eb4-d2e778992d25","openalex_id":null,"doi":null,"arxiv_id":"2403.11401","raw_key":null,"title":"Scene-LLM: Extending Language Model for 3D Visual Understanding and Reasoning","authors":null,"authors_text":"Fu, R","year":2024,"venue":"cs.CV","abstract":"This paper introduces Scene-LLM, a 3D-visual-language model that enhances embodied agents' abilities in interactive 3D indoor environments by integrating the reasoning strengths of Large Language Models (LLMs). Scene-LLM adopts a hybrid 3D visual feature representation, that incorporates dense spatial information and supports scene state updates. The model employs a projection layer to efficiently project these features in the pre-trained textual embedding space, enabling effective interpretation of 3D visual information. Unique to our approach is the integration of both scene-level and ego-centric 3D information. This combination is pivotal for interactive planning, where scene-level data supports global planning and ego-centric data is important for localization. Notably, we use ego-centric 3D frame features for feature alignment, an efficient technique that enhances the model's ability to align features of small objects within the scene. Our experiments with Scene-LLM demonstrate its strong capabilities in dense captioning, question answering, and interactive planning. We believe Scene-LLM advances the field of 3D visual understanding and reasoning, offering new possibilities for sophisticated agent interactions in indoor settings.","external_url":"https://arxiv.org/abs/2403.11401","cited_by_count":null,"metadata_source":"pith","metadata_fetched_at":"2026-07-09T22:06:35.185955+00:00","pith_arxiv_id":"2403.11401","created_at":"2026-05-09T23:04:18.051456+00:00","updated_at":"2026-07-09T22:06:35.185955+00:00","title_quality_ok":true,"display_title":"Scene-llm: Extending language model for 3d visual understanding and reasoning","render_title":"Scene-llm: Extending language model for 3d visual understanding and reasoning"},"hub":{"state":{"work_id":"9a869478-2848-4347-9eb4-d2e778992d25","tier":"hub","tier_reason":"10+ Pith inbound or 1,000+ external citations","pith_inbound_count":24,"external_cited_by_count":null,"distinct_field_count":5,"first_pith_cited_at":"2025-01-27T07:34:33+00:00","last_pith_cited_at":"2026-07-08T04:51:24+00:00","author_build_status":"not_needed","summary_status":"needed","contexts_status":"needed","graph_status":"needed","ask_index_status":"not_needed","reader_status":"not_needed","recognition_status":"not_needed","updated_at":"2026-08-22T11:39:40.082031+00:00","tier_text":"hub"},"tier":"hub","role_counts":[{"context_role":"background","n":2},{"context_role":"baseline","n":1}],"polarity_counts":[{"context_polarity":"background","n":2},{"context_polarity":"baseline","n":1}],"runs":{},"summary":{},"graph":{},"authors":[]}}