Pith. sign in

REVIEW 5 cited by

LaSagnA: Language-based Segmentation Assistant for Complex Queries

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2404.08506 v1 pith:FT66N5C5 submitted 2024-04-12 cs.CV

classification cs.CV
keywords queriessegmentationcomplexvllmsformatlasagnamodelquery
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Recent advancements have empowered Large Language Models for Vision (vLLMs) to generate detailed perceptual outcomes, including bounding boxes and masks. Nonetheless, there are two constraints that restrict the further application of these vLLMs: the incapability of handling multiple targets per query and the failure to identify the absence of query objects in the image. In this study, we acknowledge that the main cause of these problems is the insufficient complexity of training queries. Consequently, we define the general sequence format for complex queries. Then we incorporate a semantic segmentation task in the current pipeline to fulfill the requirements of training data. Furthermore, we present three novel strategies to effectively handle the challenges arising from the direct integration of the proposed format. The effectiveness of our model in processing complex queries is validated by the comparable results with conventional methods on both close-set and open-set semantic segmentation datasets. Additionally, we outperform a series of vLLMs in reasoning and referring segmentation, showcasing our model's remarkable capabilities. We release the code at https://github.com/congvvc/LaSagnA.

Discussion (0). Sign in to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. PixDLM: A Dual-Path Multimodal Language Model for UAV Reasoning Segmentation

    cs.CV 2026-04 unverdicted novelty 7.0 of 10

    The work introduces the UAV Reasoning Segmentation task, the DRSeg benchmark dataset, and PixDLM as a baseline dual-path multimodal language model for reasoning-based segmentation in aerial imagery.

  2. Segmentation before Answering: Pixel Grounding for MLLM Visual Reasoning

    cs.CV 2026-07 conditional novelty 6.0 of 10

    SegAnswer trains an MLLM to generate segmentation masks instead of bounding boxes when zooming into image regions during visual reasoning, yielding consistent improvements across perception and hallucination benchmarks.

  3. Region-based Cluster Discrimination for Visual Representation Learning

    cs.CV 2025-07 conditional novelty 6.0 of 10

    RICE improves vision encoders by applying cluster discrimination at the region level and unifying object and OCR classification targets in one pretraining framework.

  4. Advancing Visual Large Language Model for Multi-granular Versatile Perception

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A 1.3B visual language model, MVP-LM, unifies word-based and sentence-based box and mask perception in one architecture and reports competitive benchmark scores.

  5. LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance

    cs.CV 2025-07 conditional novelty 6.0 of 10

    LIRA improves referring segmentation and reduces hallucination in multimodal LLMs by fusing semantic and pixel features and interleaving local image regions with text descriptions.

Pith tools