REVIEW 7 cited by
LLaVA-Grounding: Grounded Visual Chat with Large Multimodal Models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
With the recent significant advancements in large multi-modal models (LMMs), the importance of their grounding capability in visual chat is increasingly recognized. Despite recent efforts to enable LMMs to support grounding, their capabilities for grounding and chat are usually separate, and their chat performance drops dramatically when asked to ground. The problem is the lack of a dataset for grounded visual chat (GVC). Existing grounding datasets only contain short captions. To address this issue, we have created GVC data that allows for the combination of grounding and chat capabilities. To better evaluate the GVC capabilities, we have introduced a benchmark called Grounding-Bench. Additionally, we have proposed a model design that can support GVC and various types of visual prompts by connecting segmentation models with language models. Experimental results demonstrate that our model outperforms other LMMs on Grounding-Bench. Furthermore, our model achieves competitive performance on classic grounding benchmarks like RefCOCO/+/g and Flickr30K Entities. Our code will be released at https://github.com/UX-Decoder/LLaVA-Grounding .
Forward citations
Cited by 7 Pith papers
-
Synthetic Visual Genome
A GPT-4V/GPT-4o pipeline for completing and refining scene graph annotations yields a dense synthetic dataset that, after instruction tuning, gives a 3B model strong relationship understanding and grounding results.
-
Teaching MLLMs to Say No: Generalized Referring Expression Comprehension via Refusal Calibrated GRPO
A recalibrated GRPO reinforcement learning method lets multimodal LLMs say 'None' for nonexistent referring expressions without sacrificing localization accuracy on objects that do exist.
-
MAG-Nav: Language-Driven Object Navigation Leveraging Memory-Reserved Active Grounding
MAG-Nav uses active viewpoint selection and memory replay with GPT-4o to achieve state-of-the-art 40.8% success in zero-shot language-driven object navigation on GOAT-Bench/HM3D.
-
LMM-Det: Make Large Multimodal Models Excel in Object Detection
LMM-Det makes a 7B LMM detect objects on COCO at 47.5 AP, above prior LMM-based detectors (38.5 AP) but below specialist detectors (57.3 AP), using pseudo-label distillation and per-category inference.
-
Advancing Fine-Grained Visual Understanding with Multi-Scale Alignment in Multi-Modal Models
A three-stage training method that aligns object texts, coordinates, and cropped images, plus a GPT-4V-based data synthesis pipeline, lets 1.5B and 3B models match or exceed larger MLLMs on grounding and VQA benchmarks.
-
LeapVAD: A Leap in Autonomous Driving via Cognitive Perception and Dual-Process Thinking
A dual-process knowledge-driven driving framework combining VLM perception, contrastive scene tokens, and a memory bank improves closed-loop driving scores in CARLA and DriveArena simulators.
-
Visual Large Language Models for Generalized and Specialized Applications
This paper reviews and taxonomizes VLLM applications into vision-to-text, vision-to-action, and text-to-vision, adding ethics and future-work discussion.
Discussion (0). Continue with ORCID to comment.