Pith. sign in

REVIEW 30 cited by

LISA: Reasoning Segmentation via Large Language Model

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2308.00692 v3 pith:D62G5RZC submitted 2023-08-01 cs.CV

classification cs.CV
keywords segmentationreasoninglanguagelisadatalargecapabilitiescapability
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Although perception systems have made remarkable advancements in recent years, they still rely on explicit human instruction or pre-defined categories to identify the target objects before executing visual recognition tasks. Such systems cannot actively reason and comprehend implicit user intention. In this work, we propose a new segmentation task -- reasoning segmentation. The task is designed to output a segmentation mask given a complex and implicit query text. Furthermore, we establish a benchmark comprising over one thousand image-instruction-mask data samples, incorporating intricate reasoning and world knowledge for evaluation purposes. Finally, we present LISA: large Language Instructed Segmentation Assistant, which inherits the language generation capabilities of multimodal Large Language Models (LLMs) while also possessing the ability to produce segmentation masks. We expand the original vocabulary with a <SEG> token and propose the embedding-as-mask paradigm to unlock the segmentation capability. Remarkably, LISA can handle cases involving complex reasoning and world knowledge. Also, it demonstrates robust zero-shot capability when trained exclusively on reasoning-free datasets. In addition, fine-tuning the model with merely 239 reasoning segmentation data samples results in further performance enhancement. Both quantitative and qualitative experiments show our method effectively unlocks new reasoning segmentation capabilities for multimodal LLMs. Code, models, and data are available at https://github.com/dvlab-research/LISA.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 30 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. AIpparel: A Multimodal Foundation Model for Digital Garments

    cs.CV 2024-12 conditional novelty 7.0 of 10

    AIpparel fine-tunes a large multimodal model to generate and edit sewing patterns from text and images, outperforming prior single-modality methods.

  2. Vision-Language Grounding as Bidirectional Concept Correspondence

    cs.CV 2026-08 conditional novelty 6.5 of 10

    Grounding can be treated as recovering all word-span to image-mask correspondences in one pass, and a bridge-token model, ConCor-1, does this better than existing grounding pipelines on the tested benchmarks.

  3. Symbol and Footprint Database for Electronic Components by Agentic Recognition and Generation

    cs.AI 2026-07 conditional novelty 6.0 of 10

    An MLLM-driven agentic pipeline generates PCB component symbols and footprints from datasheets with reported 86%/80% accuracy and builds a 1,000-component library.

  4. CAIRN: Cross-Room 3D Scene Understanding with Topology-Aware Large Multimodal Models

    cs.CV 2026-07 conditional novelty 6.0 of 10

    Topology-aware attention over hierarchical scene graphs lets a 3D-LLM ground, caption, and answer questions across multi-room homes, with large gains on a new HM3D benchmark.

  5. VisCoP: Visual Probing for Video Domain Adaptation of Vision Language Models

    cs.CV 2025-10 reject novelty 6.0 of 10

    Adding layer-wise learnable visual probes to a frozen VLM vision encoder improves domain adaptation across egocentric, depth, and robot-control domains while largely retaining source-domain performance.

  6. Advancing Visual Large Language Model for Multi-granular Versatile Perception

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A 1.3B visual language model, MVP-LM, unifies word-based and sentence-based box and mask perception in one architecture and reports competitive benchmark scores.

  7. DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding

    cs.CV 2025-07 conditional novelty 6.0 of 10

    DynImg represents a video snippet as a keyframe plus four resized neighboring frames as temporal prompts, with a 4D rotary position embedding, and reports improved video QA accuracy.

  8. VisionThink: Smart and Efficient Vision Language Model via Reinforcement Learning

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A vision-language model learns via reinforcement learning when to upscale a low-resolution image, cutting visual tokens roughly in half while preserving accuracy on most benchmarks.

  9. Inter2Former: Dynamic Hybrid Attention for Efficient High-Precision Interactive

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A new interactive segmentation decoder that routes computation to boundary regions, using binary quantization attention and mixture-of-experts, achieves state-of-the-art accuracy with CPU-friendly latency.

  10. ADIEE: Automatic Dataset Creation and Scorer for Instruction-Guided Image Editing Evaluation

    cs.CV 2025-07 conditional novelty 6.0 of 10

    An automatically generated training dataset and a fine-tuned LLaVA-NeXT model produce an image editing evaluation scorer that aligns with human preference and serves as a reward model for improving editing models.

  11. VideoMolmo: Spatio-Temporal Grounding Meets Pointing

    cs.CV 2025-06 conditional novelty 6.0 of 10

    A video language model that conditions each frame on earlier frames via a temporal attention module, predicts text-requested object points, and uses SAM2-based bidirectional mask fusion to outperform prior models on v...

  12. STORM: Benchmarking Visual Rating of MLLMs with a Comprehensive Ordinal Regression Dataset

    cs.CV 2025-06 conditional novelty 6.0 of 10

    STORM is a new multi-domain ordinal-regression benchmark with coarse-to-fine Chain-of-Thought prompts that improves MLLM zero-shot visual rating, though the 'universal' claim is bounded by its five curated domains.

  13. Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces

    cs.CV 2025-05 reject novelty 6.0 of 10

    VeBrain unifies perception, spatial reasoning, and robot control in one MLLM by representing control as keypoint detection plus skill selection, with a robotic adapter for deployment.

  14. On the robustness of multimodal language model towards distractions

    cs.CV 2025-02 conditional novelty 6.0 of 10

    Adding irrelevant visual and textual distractions to science questions degrades the accuracy of most vision-language models, and text distractions are more harmful than image distractions.

  15. Pixel-Level Reasoning Segmentation via Multi-turn Conversations

    cs.CV 2025-02 conditional novelty 6.0 of 10

    PRIST, a benchmark for pixel-level segmentation through multi-turn conversations, and the MIRAS model achieve the best reported scores on this new task.

  16. Densely Connected Parameter-Efficient Tuning for Referring Image Segmentation

    cs.CV 2025-01 conditional novelty 6.0 of 10

    DETRIS uses dense mixtures of convolutions and cross-attention adapters to tune a frozen DINOv2/CLIP pair, achieving top reported IoU on three referring image segmentation benchmarks while updating only a small fracti...

  17. GeoPix: Multi-Modal Large Language Model for Pixel-level Image Understanding in Remote Sensing

    cs.CV 2025-01 conditional novelty 6.0 of 10

    GeoPix extends multimodal language models for remote sensing to pixel-level referring segmentation through a mask predictor, a class-wise learnable memory, and a new 65,463-image instruction dataset.

  18. Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs

    cs.CV 2025-01 conditional novelty 6.0 of 10

    Multimodal LLMs show systematic weaknesses in instance-level visual correspondence, and CoLVA, trained with a fine-grained vision expert and object-level contrastive learning, reaches 49.8% accuracy on the new MMVM be...

  19. Agri-LLaVA: Knowledge-Infused Large Multimodal Assistant on Agricultural Pests and Diseases

    cs.CV 2024-12 conditional novelty 6.0 of 10

    By fine-tuning LLaVA on a knowledge-infused agricultural dataset, Agri-LLaVA improves agricultural conversation and VQA over general LMMs, with gains of about 5 points over LLaVA on the new benchmark.

  20. InsightEdit: Towards Better Instruction Following for Image Editing

    cs.CV 2024-11 conditional novelty 6.0 of 10

    InsightEdit uses multimodal language model features in a two-stream adapter to improve complex instruction following and background consistency in image editing.

  21. Motion-Grounded Video Reasoning: Understanding and Perceiving Motion at Pixel Level

    cs.CV 2024-11 conditional novelty 6.0 of 10

    A new benchmark and baseline for motion-grounded video reasoning, where the answer to a motion question is a spatiotemporal segmentation mask.

  22. MedSeg-R: Reasoning Segmentation in Medical Images with Multimodal Large Language Models

    cs.CV 2025-06 conditional novelty 5.0 of 10

    A medical reasoning segmentation model built from a multimodal LLM and a SAM-style mask decoder, trained on a newly generated 10,000-pair medical QA-mask dataset.

  23. InstructSeg: Unifying Instructed Visual Segmentation with Multi-modal Large Language Models

    cs.CV 2024-12 conditional novelty 5.0 of 10

    A single 3B-parameter end-to-end model with object-aware video perceiving and multi-granularity text fusion reports SOTA results across four instructed visual segmentation tasks.

  24. EditScout: Locating Forged Regions from Diffusion-based Edited Images with Multimodal LLM

    cs.CV 2024-12 conditional novelty 5.0 of 10

    A multimodal LLM with a SAM-based mask decoder localizes diffusion-edited regions better than traditional forensic methods on MagicBrush, CocoGLIDE, and a new BrushNet dataset.

  25. ChatRex: Taming Multimodal LLM for Joint Perception and Understanding

    cs.CV 2024-11 conditional novelty 5.0 of 10

    ChatRex couples a universal proposal network with an LLM that retrieves box indices, reaching 48.2 mAP on COCO and strong referring and region-level results.

  26. HyperSeg: Towards Universal Visual Segmentation with Large Language Model

    cs.CV 2024-11 conditional novelty 5.0 of 10

    A single VLLM-based model, HyperSeg, unifies image and video segmentation, including reasoning tasks, and reports SOTA on RefCOCO, ReasonSeg, ReVOS, and panoptic segmentation.

  27. ReconVLA: Reconstructive Vision-Language-Action Model as Effective Robot Perceiver

    cs.RO 2025-08 unverdicted novelty 4.0 of 10

    Adding a reconstruction target that redraws the object region makes a vision-language-action model focus its attention on the right object and manipulate more precisely.

  28. KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model

    cs.CV 2025-07 conditional novelty 4.0 of 10

    KptLLM++ unifies keypoint semantic understanding, visual-prompt detection, and text-prompt detection in a single multimodal LLM, reporting SOTA accuracy on COCO, AP-10K, Human-Art, and other benchmarks.

  29. Retrieval Augmented Recipe Generation

    cs.CV 2024-11 conditional novelty 4.0 of 10

    A retrieval-augmented LMM with stochastic retrieval sampling and self-consistency voting improves recipe generation from food images on Recipe1M.

  30. Instruction-Guided Editing Controls for Images and Multimedia: A Survey in LLM era

    cs.CV 2024-11 unverdicted novelty 3.0 of 10

    A survey that organizes over 100 instruction-guided image and multimedia editing papers into a process-based taxonomy, with an emphasis on LLM and MLLM empowered methods.

Pith tools