Pith. sign in

REVIEW 1 cited by

Unveiling the Invisible: Reasoning Complex Occlusions Amodally with AURA

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2503.10225 v2 pith:LON4TFNA submitted 2025-03-13 cs.CV

classification cs.CV
keywords amodalcomplexreasoningsegmentationauradatasetoccludedocclusions
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Amodal segmentation aims to infer the complete shape of occluded objects, even when the occluded region's appearance is unavailable. However, current amodal segmentation methods lack the capability to interact with users through text input and struggle to understand or reason about implicit and complex purposes. While methods like LISA integrate multi-modal large language models (LLMs) with segmentation for reasoning tasks, they are limited to predicting only visible object regions and face challenges in handling complex occlusion scenarios. To address these limitations, we propose a novel task named amodal reasoning segmentation, aiming to predict the complete amodal shape of occluded objects while providing answers with elaborations based on user text input. We develop a generalizable dataset generation pipeline and introduce a new dataset focusing on daily life scenarios, encompassing diverse real-world occlusions. Furthermore, we present AURA (Amodal Understanding and Reasoning Assistant), a novel model with advanced global and spatial-level designs specifically tailored to handle complex occlusions. Extensive experiments validate AURA's effectiveness on the proposed dataset.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. R2SM: Referring and Reasoning for Selective Masks

    cs.CV 2025-06 conditional novelty 6.0 of 10

    R2SM provides the first benchmark pairing modal and amodal text prompts with matching masks, letting models learn when to segment only visible parts versus complete occluded shapes.

Pith tools