REVIEW 37 cited by
LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Recent advancements in Large Multimodal Models (LMMs) have greatly enhanced their proficiency in 2D visual understanding tasks, enabling them to effectively process and understand images and videos. However, the development of LMMs with 3D scene understanding capabilities has been hindered by the lack of large-scale 3D vision-language datasets and powerful 3D encoders. In this paper, we introduce a simple yet effective framework called LLaVA-3D. Leveraging the strong 2D visual understanding priors from LLaVA, our LLaVA-3D efficiently adapts LLaVA for 3D scene understanding without compromising 2D understanding capabilities. To achieve this, we utilize the 3D position embeddings to enhance the 2D CLIP Patches with 3D spatial context information and construct 3D patches. By integrating the 3D position embeddings into 2D LMMs and employing joint 2D and 3D vision-language instruction tuning, we establish a unified architecture for both 2D visual understanding and 3D scene understanding. In contrast to previous 3D LMMs, LLaVA-3D supports decoding accurate 3D spatial perception outputs, e.g., 3D bounding boxes, directly from these 3D patches, without relying on the time-consuming off-the-shelf 3D segmentors. Experimental results show that LLaVA-3D converges 3.5x faster than existing 3D LMMs when trained on 3D vision-language datasets. Moreover, LLaVA-3D not only achieves state-of-the-art performance across various 3D tasks but also maintains comparable 2D visual understanding and vision-language conversation capabilities with LLaVA.
Forward citations
Cited by 37 Pith papers
-
GAE: Unleashing Physical Potential of VLM with Generalizable Action Expert
A frozen, large-scale pretrained diffusion policy converts sparse 3D waypoints from a VLM into dense robot actions, enabling zero-shot reuse of the action expert on new tasks and environments.
-
SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing
SAVVY-Bench tests audio-visual LLMs on dynamic 3D spatial questions, and the SAVVY pipeline, combining visual tracks with spatial audio and global mapping, lifts Gemini-2.5-pro accuracy from 50.9% to 58.0%.
-
Out of Sight, Not Out of Context? Egocentric Spatial Reasoning in VLMs Across Disjoint Frames
A new egocentric benchmark shows vision-language models fail at spatial reasoning across disjoint frames, falling 28 points behind humans and only improving sharply when handed ground-truth 3D coordinates.
-
LLaVA-4D: Embedding SpatioTemporal Prompt into LMMs for 4D Scene Understanding
A vision-language model that embeds 3D coordinates plus time into visual and linguistic tokens beats 3D-only models on dynamic scene captioning, grounding, and QA.
-
Seeing from Another Perspective: Evaluating Multi-View Understanding in MLLMs
All-Angles Bench, a 2,132-question benchmark across 90 real scenes, shows current MLLMs score around 60% on multi-view understanding while humans score 82%, with the largest gaps in camera pose estimation and cross-vi...
-
Hypo3D: Exploring Hypothetical Reasoning in 3D
Hypo3D, a 3D VQA benchmark with 14,885 QAs over 700 indoor scenes, tests whether models can reason about imagined scene changes and shows they fall far behind humans.
-
GPT4Scene: Understand 3D Scenes from Videos with Vision-Language Models
Consistent object IDs across video frames and a reconstructed bird's-eye view let large VLMs, and a fine-tuned 7B model, reach state-of-the-art results on ScanNet 3D question answering, captioning, and grounding.
-
VIEW2SPACE: Studying Multi-View Visual Reasoning from Sparse Observations
Sparse multi-view reasoning is largely unsolved for VLMs; grounded CoT with visual evidence improves moderate cases and transfers to real data, but deep multi-hop reasoning still scales poorly.
-
Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models
Multi-view relational distillation transfers geometric knowledge to vision-language models by matching cross-view patch similarity matrices, improving spatial reasoning with minimal overhead.
-
ViSTR-Bench: Can MLLMs Reason from Continuous Visual Cues in Dynamic Scenes?
A new benchmark shows the best multimodal LLM reaches 62% versus 91% human accuracy on qualitative spatial-temporal reasoning from videos.
-
Self in Space: Benchmarking Self-Awareness and Spatial Cognition in UAV Embodied Intelligence
SIS-Bench shows MLLMs lag on UAV self-awareness versus spatial tasks, and motion-aware optical-flow fusion improves perception and memory.
-
EgoMind: Activating Spatial Cognition through Linguistic Reasoning in MLLMs
EgoMind uses Role-Play Caption and Progressive Spatial Analysis to give MLLMs competitive multi-frame spatial reasoning without 3D priors, using only 5K SFT and 20K RL samples.
-
GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models
A sparse pointmap-plus-semantics pre-training stage with multi-level gated fusion of geometric and visual tokens improves RGB-only 3D perception in multimodal LLMs.
-
OpenGround: Planning-based Online Perception for Open-World 3D Visual Grounding
OpenGround grounds open-world 3D targets by planning a task chain and dynamically expanding the object lookup table through online 2D segmentation and 3D lifting, achieving SOTA zero-shot ScanRefer accuracy and 46.2% ...
-
SpaceDrive: Infusing Spatial Awareness into VLM-based Autonomous Driving
SpaceDrive replaces textual coordinate tokens with shared 3D positional encodings in a VLM driving planner, achieving state-of-the-art open-loop planning on nuScenes and 78.02 Driving Score on Bench2Drive.
-
GeoVLA: Empowering 3D Representations in Vision-Language-Action Models
A robot policy that combines 2D vision-language features with a point-cloud encoder and a mixture-of-experts diffusion action head reports SOTA manipulation success in simulation and robust real-world behavior under h...
-
3D-R1: Enhancing Reasoning in 3D VLMs for Unified Scene Understanding
3D-R1 uses a synthetic chain-of-thought cold start plus GRPO reinforcement learning with perception, semantic, and format rewards, and reports best published results across 3D dense captioning, QA, grounding, dialogue...
-
IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering
A benchmark that scores vision-language models by reconstructing the 3D scene behind an image as executable Blender code finds the models fail mainly on spatial precision, not tool usage.
-
AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making
AntiGrounding lifts candidate robot trajectories into the VLM's visual space via multi-view rendering and structured VQA, and reports 57.5% average success across eight manipulation tasks, beating three intermediate-r...
-
Pts3D-LLM: Studying the Impact of Token Structure for 3D Scene Understanding With Large Language Models
A 3D vision-language model that merges point-cloud features with image tokens and uses a new view-aware point sampling method (FPS6D) achieves state-of-the-art normalized scores on ScanNet benchmarks.
-
Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces
VeBrain unifies perception, spatial reasoning, and robot control in one MLLM by representing control as keypoint detection plus skill selection, with a robotic adapter for deployment.
-
Dynam3D: Dynamic Layered 3D Tokens Empower VLM for Vision-and-Language Navigation
Dynam3D represents scenes as patch, instance, and zone tokens that update dynamically, and feeds them to a 3.8B vision-language model to improve action prediction in vision-and-language navigation.
-
Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding
A contrastive pre-training method aligns 3D point features with language at the entity level and achieves state-of-the-art open-vocabulary semantic segmentation on ScanNet.
-
Revisiting 3D LLM Benchmarks: Are We Really Testing 3D Capabilities?
Vision-language models given rendered 2D images of point clouds can outperform specialized 3D LLMs on object-level benchmarks, showing these benchmarks do not isolate 3D understanding.
-
Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding
Adding 3D coordinate encodings to sampled RGB-D video frames lets a video LLM beat prior 3D scene understanding models on five benchmarks.
-
DepthCues: Evaluating Monocular Depth Perception in Large Vision Models
A six-task benchmark shows newer large vision models encode human-like monocular depth cues, and cue understanding strongly correlates with their depth estimation performance.
-
IGGT4D: Streaming 4D Instance-Grounded Geometry Transformer
A causal streaming transformer that jointly predicts camera motion, 3D geometry, and persistent object-instance features from video, trained on a new 147K-sequence 4D dataset.
-
Towards Enhancing 3D Spatial Reasoning in Medical Multimodal Large Language Models
A synthetic chain-of-thought dataset generated from CT reports lets a 2D-pretrained medical MLLM improve on 3D CT spatial-reasoning benchmarks.
-
Vision-Language Memory for Spatial Reasoning
A video-based vision-language model with 3D-aligned visual features and bounded dual memory achieves state-of-the-art scores on four spatial reasoning benchmarks.
-
Does Your 3D Encoder Really Work? When Pretrain-SFT from 2D VLMs Meets 3D VLMs
3D scene-centric VLMs underutilize the 3D encoder, overfit to linguistic and answer-frequency shortcuts, and the proposed 3D-RDQA dataset helps expose and mitigate this problem.
-
Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts
A sparse mixture-of-experts 3D multimodal LLM adaptively fuses RGB, RGBD, BEV, point cloud, and voxel tokens, achieving SOTA on several ScanNet-based 3D scene understanding benchmarks.
-
PerLA: Perceptive 3D Language Assistant
PerLA improves 3D question answering and dense captioning by combining Hilbert-curve partitioned local point-cloud details with global context through cross-attention and graph convolution.
-
TinyGiantVLM: A Lightweight Vision-Language Architecture for Spatial Reasoning under Resource Constraints
TinyGiantVLM, a 64M-parameter RGB-D vision-language model with two-phase training, reached 5th place on the AI City Challenge 2025 warehouse spatial reasoning track.
-
PySeizure: A single machine learning classifier framework to detect seizures in diverse datasets
A unified EEG seizure-detection framework with standardized preprocessing and majority voting reaches within-dataset AUC 0.86-0.90 and cross-dataset AUC 0.615-0.762 across CHB-MIT and TUSZ.
-
AdaToken-3D: Dynamic Spatial Gating for Efficient 3D Large Multimodal-Models Reasoning
AdaToken-3D uses attention-derived contribution scores to prune spatial tokens layer by layer in 3D LMMs, achieving about 60 percent FLOPs reduction with roughly unchanged benchmark scores.
-
Foundational Models for 3D Point Clouds: A Survey and Outlook
A structured review of methods that build or adapt 2D foundation models and LLMs for 3D point cloud tasks, with a proposed taxonomy and curated paper list.
-
The Internet of Large Language Models: An Orchestration Framework for LLM Training and Knowledge Exchange Toward Artificial General Intelligence
The paper proposes the Internet of LLM framework for model sharing, unified environments, agent-path optimization, and compute-sharing incentives, but presents no implementation or empirical evidence that it works.
Discussion (0). Continue with ORCID to comment.