Pith. sign in

REVIEW 13 cited by

HyperSeg: Towards Universal Visual Segmentation with Large Language Model

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2411.17606 v2 pith:RTS7P2PI submitted 2024-11-26 cs.CV

classification cs.CV
keywords segmentationreasoningtaskshypersegimageperceptionuniversalvideo
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

This paper aims to address universal segmentation for image and video perception with the strong reasoning ability empowered by Visual Large Language Models (VLLMs). Despite significant progress in current unified segmentation methods, limitations in adaptation to both image and video scenarios, as well as the complex reasoning segmentation, make it difficult for them to handle various challenging instructions and achieve an accurate understanding of fine-grained vision-language correlations. We propose HyperSeg, the first VLLM-based universal segmentation model for pixel-level image and video perception, encompassing generic segmentation tasks and more complex reasoning perception tasks requiring powerful reasoning abilities and world knowledge. Besides, to fully leverage the recognition capabilities of VLLMs and the fine-grained visual information, HyperSeg incorporates hybrid entity recognition and fine-grained visual perceiver modules for various segmentation tasks. Combined with the temporal adapter, HyperSeg achieves a comprehensive understanding of temporal information. Experimental results validate the effectiveness of our insights in resolving universal image and video segmentation tasks, including the more complex reasoning perception tasks. Our code is available.

Discussion (0). Sign in to comment.

Forward citations

Cited by 13 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Reason Twice: Segmentation via Candidate Discovery and Comparative Reasoning

    cs.CV 2026-06 unverdicted novelty 7.0 of 10

    Rea2Seg turns image segmentation into candidate mask discovery from MLLM attention followed by MLLM-based comparative scoring and selection, plus a new multi-dimensional reasoning benchmark ReasonSeg-SGDR.

  2. Vision Harnessing Agent for Open Ad-hoc Segmentation

    cs.CV 2026-05 unverdicted novelty 7.0 of 10

    VASA is a vision-guided agent for open ad-hoc segmentation that creates and validates masks through planning, tool use, and error recovery, outperforming baselines on the new PARS benchmark and RefCOCOm.

  3. Tarot-SAM3: Training-free SAM3 for Any Referring Expression Segmentation

    cs.CV 2026-04 unverdicted novelty 7.0 of 10

    Tarot-SAM3 delivers a training-free pipeline for segmenting images from arbitrary referring expressions via expression reasoning prompts and DINOv3-based mask self-refinement.

  4. SAM 3: Segment Anything with Concepts

    cs.CV 2025-11 unverdicted novelty 7.0 of 10

    SAM 3 introduces promptable concept segmentation that doubles accuracy of prior systems on images and videos while improving standard SAM segmentation performance.

  5. Counterfactual Segmentation Reasoning: Diagnosing and Mitigating Pixel-Grounding Hallucination

    cs.CV 2025-06 unverdicted novelty 7.0 of 10

    Proposes CSR task and HalluSegBench using visual counterfactuals to diagnose segmentation hallucinations in VLMs, plus RobustSeg via counterfactual fine-tuning that reduces hallucinations by 30% on FP-RefCOCO.

  6. PixelEyes: Decoupling Perception and Reasoning for Pinpoint Visual Evidence Seeking

    cs.CV 2026-06 unverdicted novelty 6.0 of 10

    PixelEyes decouples reasoning and perception via mask-guided search and semantic BFS, introduces PixelEyes-6K dataset and Pinpoint-Bench benchmark, and open-sources code and models.

  7. InstanceControl: Controllable Complex Image Generation without Instance Labeling

    cs.CV 2026-06 unverdicted novelty 6.0 of 10

    InstanceControl uses VLMs to auto-generate instance masks from text and visual conditions, with adaptive refinement, to enable controllable multi-object image generation without manual labeling.

  8. X2SAM: Any Segmentation in Images and Videos

    cs.CV 2026-04 unverdicted novelty 6.0 of 10

    X2SAM unifies any-segmentation across images and videos in one MLLM by adding a Mask Memory module for temporal consistency and joint training on mixed datasets.

  9. FlowSeg: Dynamic Semantic Guidance for LLM-Conditioned Segmentation

    cs.CV 2026-05 unverdicted novelty 5.0 of 10

    FlowSeg uses bidirectional semantic flow to let LLM conditions guide and be updated by mask generation states, improving alignment on referring expression and reasoning segmentation tasks.

  10. RCoT-Seg: Reinforced Chain-of-Thought for Video Reasoning and Segmentation

    cs.CV 2026-05 unverdicted novelty 4.0 of 10

    RCoT-Seg uses GRPO-reinforced keyframe selection from a CoT-start corpus followed by SAM2 mask propagation to improve video object segmentation under implicit temporal instructions over prior MLLM sampling methods.

  11. APRVOS: 1st Place Winner of 5th PVUW MeViS-Audio Track

    cs.SD 2026-04 unverdicted novelty 3.0 of 10

    A staged pipeline using ASR transcription, visual existence verification, Sa2VA coarse segmentation, and agent-guided SAM3 refinement won first place in the PVUW MeViS-Audio track by decomposing audio-conditioned Ref-...

  12. AgentRVOS for MeViS-Text Track of 5th PVUW Challenge: 3rd Method

    cs.CV 2026-04 unverdicted novelty 3.0 of 10

    An agent-augmented Sa2VA pipeline for referring video object segmentation placed third in the MeViS-Text track of the 5th PVUW Challenge by adding verification, search, and refinement stages.

  13. LMMs Meet Object-Centric Vision: Understanding, Segmentation, Editing and Generation

    cs.CV 2026-04 unverdicted novelty 3.0 of 10

    This review organizes literature on large multimodal models and object-centric vision into four themes—understanding, referring segmentation, editing, and generation—while summarizing paradigms, strategies, and challe...

Pith tools