Pith. sign in

REVIEW 27 cited by

Chat-3D: Data-efficiently Tuning Large Language Model for Universal Dialogue of 3D Scenes

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2308.08769 v1 pith:37AJNUXF submitted 2023-08-17 cs.CV

classification cs.CV
keywords chat-3dabilityllmsreasoningscenesachievesapplicationsdata
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

3D scene understanding has gained significant attention due to its wide range of applications. However, existing methods for 3D scene understanding are limited to specific downstream tasks, which hinders their practicality in real-world applications. This paper presents Chat-3D, which combines the 3D visual perceptual ability of pre-trained 3D representations and the impressive reasoning and conversation capabilities of advanced LLMs to achieve the first universal dialogue systems for 3D scenes. Specifically, we align 3D representations into the feature space of LLMs, thus enabling LLMs to perceive the 3D world. Given the scarcity of 3D scene-text data, we propose a three-stage training strategy to efficiently utilize the available data for better alignment. To enhance the reasoning ability and develop a user-friendly interaction scheme, we further construct a high-quality object-centric 3D instruction dataset and design an associated object-centric prompt. Our experiments show that Chat-3D achieves an impressive ability to comprehend diverse instructions for 3D scenes, engage in intricate spatial reasoning, and incorporate external knowledge into its responses. Chat-3D achieves a 75.6% relative score compared with GPT-4 on the constructed instruction dataset.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 27 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. GPT4Scene: Understand 3D Scenes from Videos with Vision-Language Models

    cs.CV 2025-01 conditional novelty 7.0 of 10

    Consistent object IDs across video frames and a reconstructed bird's-eye view let large VLMs, and a fine-tuned 7B model, reach state-of-the-art results on ScanNet 3D question answering, captioning, and grounding.

  2. Holo-Captioning: Toward the Text Equivalent of 3D Scenes

    cs.CV 2026-07 conditional novelty 6.5 of 10

    HoloScribe generates pure-text structured captions of all entities, boxes, attributes and relations in 3D indoor scenes and outperforms dense captioners and 3D LLMs on a new 15K-scene benchmark.

  3. 3DZip: Spatial-Aware Feature Diversity-Guided Token Compression for 3D Question Answering

    cs.CV 2026-08 conditional novelty 6.0 of 10

    A feature-diversity-guided token compression pipeline for 3D VLMs keeps 94.7% of original QA accuracy at 128 tokens and runs 1.92x faster than uncompressed LLaVA-3D.

  4. Beyond Single Expert: Harmonizing Diverse Visual Priors in MLLMs for Spatial Understanding

    cs.CV 2026-07 conditional novelty 6.0 of 10

    ViPS fuses five complementary visual priors into an MLLM via lightweight distillation proxies and query-conditioned dynamic weighting, reporting state-of-the-art results on VSI-Bench and ScanNet-series benchmarks.

  5. Self in Space: Benchmarking Self-Awareness and Spatial Cognition in UAV Embodied Intelligence

    cs.CV 2026-07 unverdicted novelty 6.0 of 10

    SIS-Bench shows MLLMs lag on UAV self-awareness versus spatial tasks, and motion-aware optical-flow fusion improves perception and memory.

  6. ELSA3D: Elastic Semantic Anchoring for Unified 3D Understanding and Generation

    cs.CV 2026-07 unverdicted novelty 6.0 of 10

    ELSA3D introduces elastic semantic anchoring via sparse anchor tokens and a scale-aware octree tokenizer to unify 3D generation and captioning at reduced computational cost.

  7. Seeing Once is Enough? Online Geometry-Aware Token Pruning for 3D Question Answering

    cs.CV 2026-07 conditional novelty 6.0 of 10

    Online geometry-aware voxel overlap pruning removes up to 50% of visual tokens from multi-view 3D scenes while improving zero-shot 3D QA on Qwen VL models.

  8. EgoMind: Activating Spatial Cognition through Linguistic Reasoning in MLLMs

    cs.CV 2026-04 unverdicted novelty 6.0 of 10

    EgoMind uses Role-Play Caption and Progressive Spatial Analysis to give MLLMs competitive multi-frame spatial reasoning without 3D priors, using only 5K SFT and 20K RL samples.

  9. Generation Models Know Space: Unleashing Implicit 3D Priors for Scene Understanding

    cs.CV 2026-03 unverdicted novelty 6.0 of 10

    VEGA-3D extracts intermediate spatiotemporal features from a pretrained video diffusion model and fuses them into MLLMs to improve geometric and embodied reasoning without explicit 3D supervision.

  10. SURPRISE3D: A Dataset for Spatial Understanding and Reasoning in Complex 3D Scenes

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A large-scale 3D spatial reasoning segmentation benchmark with human-written queries that avoid object names shows current 3D vision-language models underperform.

  11. Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation

    cs.CV 2025-06 conditional novelty 6.0 of 10

    A two-stage reasoning-segmentation method plus a new LLM-generated 3D dataset improves spatial reasoning in 3D multimodal large language models on several benchmarks.

  12. Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs

    cs.CV 2025-01 conditional novelty 6.0 of 10

    Multimodal LLMs show systematic weaknesses in instance-level visual correspondence, and CoLVA, trained with a fine-grained vision expert and object-level contrastive learning, reaches 49.8% accuracy on the new MMVM be...

  13. LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences

    cs.CV 2024-12 conditional novelty 6.0 of 10

    LSceneLLM chooses task-relevant 3D regions via LLM attention, magnifies their details, and improves large-scene 3D question answering, planning, and captioning.

  14. Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding

    cs.CV 2024-11 conditional novelty 6.0 of 10

    Adding 3D coordinate encodings to sampled RGB-D video frames lets a video LLM beat prior 3D scene understanding models on five benchmarks.

  15. Vision-Language Memory for Spatial Reasoning

    cs.CV 2025-11 conditional novelty 5.0 of 10

    A video-based vision-language model with 3D-aligned visual features and bounded dual memory achieves state-of-the-art scores on four spatial reasoning benchmarks.

  16. SeqVLM: Proposal-Guided Multi-View Sequences Reasoning via VLM for Zero-Shot 3D Visual Grounding

    cs.CV 2025-08 conditional novelty 5.0 of 10

    Proposal-guided multi-view projection with iterative VLM selection achieves 55.6% and 53.2% Acc@0.25 on ScanRefer and Nr3D, a new zero-shot 3D visual grounding state of the art.

  17. Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding

    cs.CV 2025-07 conditional novelty 5.0 of 10

    Fast3D prunes up to 90% of object-centric visual tokens in 3D MLLMs while preserving about 96.8% of original benchmark performance, using a trained attention predictor and adaptive layer-wise pruning.

  18. Does Your 3D Encoder Really Work? When Pretrain-SFT from 2D VLMs Meets 3D VLMs

    cs.CV 2025-06 conditional novelty 5.0 of 10

    3D scene-centric VLMs underutilize the 3D encoder, overfit to linguistic and answer-frequency shortcuts, and the proposed 3D-RDQA dataset helps expose and mitigate this problem.

  19. Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts

    cs.CV 2025-05 conditional novelty 5.0 of 10

    A sparse mixture-of-experts 3D multimodal LLM adaptively fuses RGB, RGBD, BEV, point cloud, and voxel tokens, achieving SOTA on several ScanNet-based 3D scene understanding benchmarks.

  20. LLMER: Crafting Interactive Extended Reality Worlds with JSON Data Generated by Large Language Models

    cs.MM 2025-02 conditional novelty 5.0 of 10

    LLMER uses LLM-generated JSON data instead of code to create interactive XR worlds, cutting token use and task completion time in a small user study.

  21. GaussianProperty: Integrating Physical Properties to 3D Gaussians with LMMs

    cs.RO 2024-12 conditional novelty 5.0 of 10

    A training-free pipeline uses SAM segmentation and GPT-4V material recognition, then votes across views to attach density, elasticity, and friction values to 3D Gaussians for simulation and grasping.

  22. 3D Spatial Understanding in MLLMs: Disambiguation and Evaluation

    cs.CV 2024-12 conditional novelty 5.0 of 10

    A pipeline that adds explicit distractor and relative-position information to an MLLM improves generation of target-exclusive 3D referring instructions, validated partly by training 3D grounding models on the generated text.

  23. RG-SAN: Rule-Guided Spatial Awareness Network for End-to-End 3D Referring Expression Segmentation

    cs.CV 2024-12 conditional novelty 5.0 of 10

    A rule-guided spatial-aware network that localizes all mentioned entities in a 3D scene and uses target-position weak supervision raises ScanRefer 3D-RES mIoU from 39.5 to 44.6.

  24. PerLA: Perceptive 3D Language Assistant

    cs.CV 2024-11 conditional novelty 5.0 of 10

    PerLA improves 3D question answering and dense captioning by combining Hilbert-curve partitioned local point-cloud details with global context through cross-attention and graph convolution.

  25. PySeizure: A single machine learning classifier framework to detect seizures in diverse datasets

    cs.LG 2025-08 conditional novelty 4.0 of 10

    A unified EEG seizure-detection framework with standardized preprocessing and majority voting reaches within-dataset AUC 0.86-0.90 and cross-dataset AUC 0.615-0.762 across CHB-MIT and TUSZ.

  26. CL3DOR: Contrastive Learning for 3D Large Multimodal Models via Odds Ratio on High-Resolution Point Clouds

    cs.CV 2025-01 conditional novelty 4.0 of 10

    CL3DOR pairs 8,192-point inputs, GPT-4o-generated hard-negative response triplets, and an odds-ratio contrastive loss to achieve state-of-the-art results on 3D scene understanding benchmarks.

  27. From 2D to 3D Cognition: A Brief Survey of General World Models

    cs.CV 2025-06 conditional novelty 3.0 of 10

    A survey proposing a two-pillar, three-capability framework that organizes recent AI world models by their transition from 2D visual prediction to 3D cognition.

Pith tools