Pith. sign in

REVIEW 12 cited by

BuboGPT: Enabling Visual Grounding in Multi-Modal LLMs

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2307.08581 v1 pith:6IJ53V2P submitted 2023-07-17 cs.CV cs.AI

classification cs.CVcs.AI
keywords llmsvisualbubogptgroundingmulti-modalunderstandingabilitiesimage
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

LLMs have demonstrated remarkable abilities at interacting with humans through language, especially with the usage of instruction-following data. Recent advancements in LLMs, such as MiniGPT-4, LLaVA, and X-LLM, further enlarge their abilities by incorporating multi-modal inputs, including image, video, and speech. Despite their effectiveness at generating precise and detailed language understanding of the given modality signal, these LLMs give up the ability to ground specific parts of inputs, thus only constructing a coarse-grained mapping. However, explicit and informative correspondence between text and other modalities will not only improve the user experience but also help to expand the application scenario of multi-modal LLMs. Therefore, we propose BuboGPT, a multi-modal LLM with visual grounding that can perform cross-modal interaction between vision, audio and language, providing fine-grained understanding of visual objects and other given modalities. As a result, BuboGPT is able to point out the specific location of an object in the image, when it is generating response or description for that object. Our contributions are two-fold: 1) An off-the-shelf visual grounding module based on SAM that extracts entities in a sentence and find corresponding masks in the image. 2) A two-stage training scheme and instruction dataset to endow joint text-image-audio understanding. Our experiments show that BuboGPT achieves impressive multi-modality understanding and visual grounding abilities during the interaction with human. It performs consistently well when provided by arbitrary modality combinations (either aligned or unaligned). Our code, model and dataset are available at https://bubo-gpt.github.io .

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 12 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Sample-efficient Integration of New Modalities into Large Language Models

    cs.CL 2025-09 conditional novelty 7.0 of 10

    A hypernetwork trained on image, audio, and video adapts a shared projector to new, low-resource modalities from as few as 32 examples.

  2. Synthetic Visual Genome

    cs.CV 2025-06 conditional novelty 7.0 of 10

    A GPT-4V/GPT-4o pipeline for completing and refining scene graph annotations yields a dense synthetic dataset that, after instruction tuning, gives a 3B model strong relationship understanding and grounding results.

  3. CAIRN: Cross-Room 3D Scene Understanding with Topology-Aware Large Multimodal Models

    cs.CV 2026-07 conditional novelty 6.0 of 10

    Topology-aware attention over hierarchical scene graphs lets a 3D-LLM ground, caption, and answer questions across multi-room homes, with large gains on a new HM3D benchmark.

  4. Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations

    cs.CV 2025-06 conditional novelty 6.0 of 10

    Video-grounded RL with two-stage 2D grounding and SAM2 lifting enables 3D object localization and QA without dense 3D instance supervision.

  5. Segment Anyword: Mask Prompt Inversion for Open-Set Grounded Segmentation

    cs.CV 2025-05 conditional novelty 6.0 of 10

    A training-free pipeline uses per-image textual inversion in a frozen diffusion model, then feeds linguistic-guided cross-attention prompts to SAM, achieving state-of-the-art open-set grounded segmentation on several ...

  6. GeoPixel: Pixel Grounding Large Multimodal Model in Remote Sensing

    cs.CV 2025-01 conditional novelty 6.0 of 10

    GeoPixel brings pixel-level grounding to remote sensing large multimodal models for the first time, with a new dataset and benchmark built from iSAID.

  7. FiVL: A Framework for Improved Vision-Language Alignment through the Lens of Training, Evaluation and Explainability

    cs.CV 2024-12 conditional novelty 6.0 of 10

    FiVL augments vision-language instruction data with GPT-4o-extracted key expressions and segmentation masks, trains LLaVA with a vision-modeling loss that predicts vocabulary tokens for image patches, and measures vis...

  8. Qwen-Audio-VAE Technical Report

    eess.AS 2026-07 conditional novelty 5.0 of 10

    A 12.5 Hz continuous audio VAE reconstructs speech, music, and sound well while encoding 64×30s clips in 541 ms after latency-aware encoder pruning.

  9. PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs

    cs.CV 2025-06 conditional novelty 5.0 of 10

    MMGrounded-PostAlign trains MLLMs to produce a grounded object token or a rejection token plus selective rationales, improving hallucination and VQA benchmarks.

  10. I'm Spartacus, No, I'm Spartacus: Measuring and Understanding LLM Identity Confusion

    cs.CR 2024-11 reject novelty 5.0 of 10

    Seven of 27 tested LLMs (25.93%) exhibited identity confusion, which the authors link to hallucination and show reduces user trust, especially in critical tasks.

  11. From Visuals to Vocabulary: Establishing Equivalence Between Image and Text Token Through Autoregressive Pre-training in MLLMs

    cs.CV 2025-02 conditional novelty 4.0 of 10

    Adding an L2 loss that pushes the language model's image hidden states back toward the input image embeddings improves LLaVA-style models on several VQA benchmarks, with some benchmarks unaffected or slightly worse.

  12. Visual Large Language Models for Generalized and Specialized Applications

    cs.CV 2025-01 conditional novelty 3.0 of 10

    This paper reviews and taxonomizes VLLM applications into vision-to-text, vision-to-action, and text-to-vision, adding ethics and future-work discussion.

Pith tools