Pith. sign in

REVIEW 2 cited by

Can 3D Vision-Language Models Truly Understand Natural Language?

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2403.14760 v3 pith:4ETCSH7Y submitted 2024-03-21 cs.CV

classification cs.CV
keywords languagemodelsd-vlexistingrobustnessunderstandvariantshuman
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Rapid advancements in 3D vision-language (3D-VL) tasks have opened up new avenues for human interaction with embodied agents or robots using natural language. Despite this progress, we find a notable limitation: existing 3D-VL models exhibit sensitivity to the styles of language input, struggling to understand sentences with the same semantic meaning but written in different variants. This observation raises a critical question: Can 3D vision-language models truly understand natural language? To test the language understandability of 3D-VL models, we first propose a language robustness task for systematically assessing 3D-VL models across various tasks, benchmarking their performance when presented with different language style variants. Importantly, these variants are commonly encountered in applications requiring direct interaction with humans, such as embodied robotics, given the diversity and unpredictability of human language. We propose a 3D Language Robustness Dataset, designed based on the characteristics of human language, to facilitate the systematic study of robustness. Our comprehensive evaluation uncovers a significant drop in the performance of all existing models across various 3D-VL tasks. Even the state-of-the-art 3D-LLM fails to understand some variants of the same sentences. Further in-depth analysis suggests that the existing models have a fragile and biased fusion module, which stems from the low diversity of the existing dataset. Finally, we propose a training-free module driven by LLM, which improves language robustness. Datasets and code will be available at github.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Do large language vision models understand 3D shapes?

    cs.CV 2024-12 conditional novelty 6.0 of 10

    A large synthetic benchmark shows vision-language models match 3D shapes well across single changes like rotation or texture, but fail when rotation and texture change together, trailing humans by a wide margin.

  2. Solving Zero-Shot 3D Visual Grounding as Constraint Satisfaction Problems

    cs.CV 2024-11 conditional novelty 6.0 of 10

    CSVG solves 3D visual grounding by converting the query into a constraint satisfaction program and finding an assignment satisfying all spatial relations at once, improving zero-shot accuracy on ScanRefer and Nr3D.

Pith tools