Pith. sign in

REVIEW 1 cited by

OCC-MLLM:Empowering Multimodal Large Language Model For the Understanding of Occluded Objects

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.01261 v1 pith:6KEJIQAT submitted 2024-10-02 cs.CV

classification cs.CV
keywords objectsoccludedmodelsmultimodallarge-scaleunderstandingvisualvisual-language
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

There is a gap in the understanding of occluded objects in existing large-scale visual language multi-modal models. Current state-of-the-art multimodal models fail to provide satisfactory results in describing occluded objects for visual-language multimodal models through universal visual encoders. Another challenge is the limited number of datasets containing image-text pairs with a large number of occluded objects. Therefore, we introduce a novel multimodal model that applies a newly designed visual encoder to understand occluded objects in RGB images. We also introduce a large-scale visual-language pair dataset for training large-scale visual-language multimodal models and understanding occluded objects. We start our experiments comparing with the state-of-the-art models.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. CAPTURe: Evaluating Spatial Reasoning in Vision Language Models via Occluded Object Counting

    cs.CV 2025-04 conditional novelty 6.0 of 10

    CAPTURe, a new benchmark for occluded pattern counting, shows that six vision-language models count far worse when objects are hidden, while humans make almost no errors.

Pith tools