Pith. sign in

REVIEW 2 cited by

CoReS: Orchestrating the Dance of Reasoning and Segmentation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2404.05673 v3 pith:2U32M4BI submitted 2024-04-08 cs.CV

classification cs.CV
keywords reasoningsegmentationcoresvisualaccuratelyfindhierarchymllm
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The reasoning segmentation task, which demands a nuanced comprehension of intricate queries to accurately pinpoint object regions, is attracting increasing attention. However, Multi-modal Large Language Models (MLLM) often find it difficult to accurately localize the objects described in complex reasoning contexts. We believe that the act of reasoning segmentation should mirror the cognitive stages of human visual search, where each step is a progressive refinement of thought toward the final object. Thus we introduce the Chains of Reasoning and Segmenting (CoReS) and find this top-down visual hierarchy indeed enhances the visual search process. Specifically, we propose a dual-chain structure that generates multi-modal, chain-like outputs to aid the segmentation process. Furthermore, to steer the MLLM's outputs into this intended hierarchy, we incorporate in-context inputs as guidance. Extensive experiments demonstrate the superior performance of our CoReS, which surpasses the state-of-the-art method by 6.5\% on the ReasonSeg dataset. Project: https://chain-of-reasoning-and-segmentation.github.io/.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model

    cs.CV 2024-11 conditional novelty 6.0 of 10

    ROSE uses patch-wise perception in a large multimodal model to predict dense masks and generate open-set category names without predefined prompts.

  2. EditScout: Locating Forged Regions from Diffusion-based Edited Images with Multimodal LLM

    cs.CV 2024-12 conditional novelty 5.0 of 10

    A multimodal LLM with a SAM-based mask decoder localizes diffusion-edited regions better than traditional forensic methods on MagicBrush, CocoGLIDE, and a new BrushNet dataset.

Pith tools