REVIEW 30 cited by
LISA: Reasoning Segmentation via Large Language Model
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Although perception systems have made remarkable advancements in recent years, they still rely on explicit human instruction or pre-defined categories to identify the target objects before executing visual recognition tasks. Such systems cannot actively reason and comprehend implicit user intention. In this work, we propose a new segmentation task -- reasoning segmentation. The task is designed to output a segmentation mask given a complex and implicit query text. Furthermore, we establish a benchmark comprising over one thousand image-instruction-mask data samples, incorporating intricate reasoning and world knowledge for evaluation purposes. Finally, we present LISA: large Language Instructed Segmentation Assistant, which inherits the language generation capabilities of multimodal Large Language Models (LLMs) while also possessing the ability to produce segmentation masks. We expand the original vocabulary with a <SEG> token and propose the embedding-as-mask paradigm to unlock the segmentation capability. Remarkably, LISA can handle cases involving complex reasoning and world knowledge. Also, it demonstrates robust zero-shot capability when trained exclusively on reasoning-free datasets. In addition, fine-tuning the model with merely 239 reasoning segmentation data samples results in further performance enhancement. Both quantitative and qualitative experiments show our method effectively unlocks new reasoning segmentation capabilities for multimodal LLMs. Code, models, and data are available at https://github.com/dvlab-research/LISA.
Forward citations
Cited by 30 Pith papers
-
AIpparel: A Multimodal Foundation Model for Digital Garments
AIpparel fine-tunes a large multimodal model to generate and edit sewing patterns from text and images, outperforming prior single-modality methods.
-
Vision-Language Grounding as Bidirectional Concept Correspondence
Grounding can be treated as recovering all word-span to image-mask correspondences in one pass, and a bridge-token model, ConCor-1, does this better than existing grounding pipelines on the tested benchmarks.
-
Symbol and Footprint Database for Electronic Components by Agentic Recognition and Generation
An MLLM-driven agentic pipeline generates PCB component symbols and footprints from datasheets with reported 86%/80% accuracy and builds a 1,000-component library.
-
CAIRN: Cross-Room 3D Scene Understanding with Topology-Aware Large Multimodal Models
Topology-aware attention over hierarchical scene graphs lets a 3D-LLM ground, caption, and answer questions across multi-room homes, with large gains on a new HM3D benchmark.
-
VisCoP: Visual Probing for Video Domain Adaptation of Vision Language Models
Adding layer-wise learnable visual probes to a frozen VLM vision encoder improves domain adaptation across egocentric, depth, and robot-control domains while largely retaining source-domain performance.
-
Advancing Visual Large Language Model for Multi-granular Versatile Perception
A 1.3B visual language model, MVP-LM, unifies word-based and sentence-based box and mask perception in one architecture and reports competitive benchmark scores.
-
DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding
DynImg represents a video snippet as a keyframe plus four resized neighboring frames as temporal prompts, with a 4D rotary position embedding, and reports improved video QA accuracy.
-
VisionThink: Smart and Efficient Vision Language Model via Reinforcement Learning
A vision-language model learns via reinforcement learning when to upscale a low-resolution image, cutting visual tokens roughly in half while preserving accuracy on most benchmarks.
-
Inter2Former: Dynamic Hybrid Attention for Efficient High-Precision Interactive
A new interactive segmentation decoder that routes computation to boundary regions, using binary quantization attention and mixture-of-experts, achieves state-of-the-art accuracy with CPU-friendly latency.
-
ADIEE: Automatic Dataset Creation and Scorer for Instruction-Guided Image Editing Evaluation
An automatically generated training dataset and a fine-tuned LLaVA-NeXT model produce an image editing evaluation scorer that aligns with human preference and serves as a reward model for improving editing models.
-
VideoMolmo: Spatio-Temporal Grounding Meets Pointing
A video language model that conditions each frame on earlier frames via a temporal attention module, predicts text-requested object points, and uses SAM2-based bidirectional mask fusion to outperform prior models on v...
-
STORM: Benchmarking Visual Rating of MLLMs with a Comprehensive Ordinal Regression Dataset
STORM is a new multi-domain ordinal-regression benchmark with coarse-to-fine Chain-of-Thought prompts that improves MLLM zero-shot visual rating, though the 'universal' claim is bounded by its five curated domains.
-
Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces
VeBrain unifies perception, spatial reasoning, and robot control in one MLLM by representing control as keypoint detection plus skill selection, with a robotic adapter for deployment.
-
On the robustness of multimodal language model towards distractions
Adding irrelevant visual and textual distractions to science questions degrades the accuracy of most vision-language models, and text distractions are more harmful than image distractions.
-
Pixel-Level Reasoning Segmentation via Multi-turn Conversations
PRIST, a benchmark for pixel-level segmentation through multi-turn conversations, and the MIRAS model achieve the best reported scores on this new task.
-
Densely Connected Parameter-Efficient Tuning for Referring Image Segmentation
DETRIS uses dense mixtures of convolutions and cross-attention adapters to tune a frozen DINOv2/CLIP pair, achieving top reported IoU on three referring image segmentation benchmarks while updating only a small fracti...
-
GeoPix: Multi-Modal Large Language Model for Pixel-level Image Understanding in Remote Sensing
GeoPix extends multimodal language models for remote sensing to pixel-level referring segmentation through a mask predictor, a class-wise learnable memory, and a new 65,463-image instruction dataset.
-
Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs
Multimodal LLMs show systematic weaknesses in instance-level visual correspondence, and CoLVA, trained with a fine-grained vision expert and object-level contrastive learning, reaches 49.8% accuracy on the new MMVM be...
-
Agri-LLaVA: Knowledge-Infused Large Multimodal Assistant on Agricultural Pests and Diseases
By fine-tuning LLaVA on a knowledge-infused agricultural dataset, Agri-LLaVA improves agricultural conversation and VQA over general LMMs, with gains of about 5 points over LLaVA on the new benchmark.
-
InsightEdit: Towards Better Instruction Following for Image Editing
InsightEdit uses multimodal language model features in a two-stream adapter to improve complex instruction following and background consistency in image editing.
-
Motion-Grounded Video Reasoning: Understanding and Perceiving Motion at Pixel Level
A new benchmark and baseline for motion-grounded video reasoning, where the answer to a motion question is a spatiotemporal segmentation mask.
-
MedSeg-R: Reasoning Segmentation in Medical Images with Multimodal Large Language Models
A medical reasoning segmentation model built from a multimodal LLM and a SAM-style mask decoder, trained on a newly generated 10,000-pair medical QA-mask dataset.
-
InstructSeg: Unifying Instructed Visual Segmentation with Multi-modal Large Language Models
A single 3B-parameter end-to-end model with object-aware video perceiving and multi-granularity text fusion reports SOTA results across four instructed visual segmentation tasks.
-
EditScout: Locating Forged Regions from Diffusion-based Edited Images with Multimodal LLM
A multimodal LLM with a SAM-based mask decoder localizes diffusion-edited regions better than traditional forensic methods on MagicBrush, CocoGLIDE, and a new BrushNet dataset.
-
ChatRex: Taming Multimodal LLM for Joint Perception and Understanding
ChatRex couples a universal proposal network with an LLM that retrieves box indices, reaching 48.2 mAP on COCO and strong referring and region-level results.
-
HyperSeg: Towards Universal Visual Segmentation with Large Language Model
A single VLLM-based model, HyperSeg, unifies image and video segmentation, including reasoning tasks, and reports SOTA on RefCOCO, ReasonSeg, ReVOS, and panoptic segmentation.
-
ReconVLA: Reconstructive Vision-Language-Action Model as Effective Robot Perceiver
Adding a reconstruction target that redraws the object region makes a vision-language-action model focus its attention on the right object and manipulate more precisely.
-
KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model
KptLLM++ unifies keypoint semantic understanding, visual-prompt detection, and text-prompt detection in a single multimodal LLM, reporting SOTA accuracy on COCO, AP-10K, Human-Art, and other benchmarks.
-
Retrieval Augmented Recipe Generation
A retrieval-augmented LMM with stochastic retrieval sampling and self-consistency voting improves recipe generation from food images on Recipe1M.
-
Instruction-Guided Editing Controls for Images and Multimedia: A Survey in LLM era
A survey that organizes over 100 instruction-guided image and multimedia editing papers into a process-based taxonomy, with an emphasis on LLM and MLLM empowered methods.
Discussion (0). Continue with ORCID to comment.