REVIEW 8 cited by
NExT-Chat: An LMM for Chat, Detection and Segmentation
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
The development of large language models (LLMs) has greatly advanced the field of multimodal understanding, leading to the emergence of large multimodal models (LMMs). In order to enhance the level of visual comprehension, recent studies have equipped LMMs with region-level understanding capabilities by representing object bounding box coordinates as a series of text sequences (pix2seq). In this paper, we introduce a novel paradigm for object location modeling called pix2emb method, where we ask the LMM to output the location embeddings and then decode them with different decoders. This paradigm allows us to use different location formats (such as bounding boxes and masks) in multimodal conversations. Leveraging the proposed pix2emb method, we train an LMM named NExT-Chat and demonstrate its capability of handling multiple tasks like visual grounding, region captioning, and grounded reasoning. Comprehensive experiments show the effectiveness of our NExT-Chat on various tasks, e.g., NExT-Chat (87.7) vs. Shikra (86.9) on POPE-Random, NExT-Chat (68.9) vs. LISA (67.9) on referring expression segmentation task, and NExT-Chat (79.6) vs. Kosmos-2 (62.3) on region caption task. The code and model are released at https://github.com/NExT-ChatV/NExT-Chat.
Forward citations
Cited by 8 Pith papers
-
Multimodal Large Language Models for Remote Sensing Image Understanding: Domain-Specific or General-Purpose?
Under one zero-shot protocol, general-purpose MLLMs match or outperform remote-sensing-specific MLLMs on several RS benchmarks, while RS-MLLMs keep advantages in visual grounding, RS-VQA, and ultra-high-resolution und...
-
ReMeREC: Relation-aware and Multi-entity Referring Expression Comprehension
ReMeREC introduces a relation-aware multi-entity referring expression comprehension framework and the ReMeX dataset, reporting state-of-the-art grounding and relation prediction, with some evaluation caveats.
-
HRSeg: High-Resolution Visual Perception and Enhancement for Reasoning Segmentation
HRSeg combines high-resolution image crops, region attention, and cross-attention mask enhancement, improving reasoning segmentation over LLM-Seg by up to 13.2 gIoU points on LLM-Seg40K.
-
Exploring Compositional Generalization of Multimodal LLMs for Medical Imaging
Multimodal LLMs trained on medical images decomposed into modality, anatomy, and task can generalize to unseen combinations of those elements, and this compositional generalization partially explains multi-task traini...
-
MediRound: Multi-Round Entity-Level Reasoning Segmentation in Medical Images
MediRound introduces a multi-round, entity-level medical segmentation task, a 177K-dialogue dataset built from SA-Med2D-20M with GPT-5, and a LLaVA-Med/MedSAM baseline whose inference-time judgment-and-correction modu...
-
CityLoc: 6DoF Pose Distributional Localization for Text Descriptions in Large-Scale Scenes with Gaussian Representation
A text-conditioned diffusion model with 3D Gaussian splatting refinement estimates 6DoF camera pose distributions in city-scale scenes, beating a Monte Carlo dropout baseline on five datasets.
-
Task Preference Optimization: Improving Multimodal Large Language Models with Vision Task Alignment
Adding task-specific heads for tracking, grounding, and segmentation to multimodal LLMs via a three-stage training recipe improves both fine-grained visual tasks and general video understanding benchmarks.
-
Visual Large Language Models for Generalized and Specialized Applications
This paper reviews and taxonomizes VLLM applications into vision-to-text, vision-to-action, and text-to-vision, adding ethics and future-work discussion.
Discussion (0). Continue with ORCID to comment.