Pith. sign in

REVIEW 2 cited by

Osprey: Pixel Understanding with Visual Instruction Tuning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2312.10032 v4 pith:LQOX6UTF submitted 2023-12-15 cs.CV

classification cs.CV
keywords instructionospreyvisualtuningunderstandingmllmsvision-languageachieving
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Multimodal large language models (MLLMs) have recently achieved impressive general-purpose vision-language capabilities through visual instruction tuning. However, current MLLMs primarily focus on image-level or box-level understanding, falling short in achieving fine-grained vision-language alignment at pixel level. Besides, the lack of mask-based instruction data limits their advancements. In this paper, we propose Osprey, a mask-text instruction tuning approach, to extend MLLMs by incorporating fine-grained mask regions into language instruction, aiming at achieving pixel-wise visual understanding. To achieve this goal, we first meticulously curate a mask-based region-text dataset with 724K samples, and then design a vision-language model by injecting pixel-level representation into LLM. Specifically, Osprey adopts a convolutional CLIP backbone as the vision encoder and employs a mask-aware visual extractor to extract precise visual mask features from high resolution input. Experimental results demonstrate Osprey's superiority in various region understanding tasks, showcasing its new capability for pixel-level instruction tuning. In particular, Osprey can be integrated with Segment Anything Model (SAM) seamlessly to obtain multi-granularity semantics. The source code, dataset and demo can be found at https://github.com/CircleRadon/Osprey.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. TrajTok: Learning Trajectory Tokens enables better Video Understanding

    cs.CV 2026-02 unverdicted novelty 7.0 of 10

    TrajTok learns to tokenize video into object-trajectory tokens end-to-end, improving video CLIP, probing, and VLM performance over patch and token-merging baselines.

  2. Reverse Region-to-Entity Annotation for Pixel-Level Visual Entity Linking

    cs.CV 2024-12 conditional novelty 7.0 of 10

    Introduces PL-VEL, a pixel-mask-based visual entity linking task, and MaskOVEN-Wiki, a 5.2M-annotation dataset built via reverse annotation, plus a semantic tokenization method that yields a 5-point accuracy gain.

Pith tools