Pith. sign in

REVIEW 2 cited by

u-LLaVA: Unifying Multi-Modal Tasks via Large Language Model

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2311.05348 v4 pith:YREXNXSE submitted 2023-11-09 cs.CV

classification cs.CV
keywords mllmsmodelu-llavaacrossalignmentfine-grainedframeworkglobal
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Recent advancements in multi-modal large language models (MLLMs) have led to substantial improvements in visual understanding, primarily driven by sophisticated modality alignment strategies. However, predominant approaches prioritize global or regional comprehension, with less focus on fine-grained, pixel-level tasks. To address this gap, we introduce u-LLaVA, an innovative unifying multi-task framework that integrates pixel, regional, and global features to refine the perceptual faculties of MLLMs. We commence by leveraging an efficient modality alignment approach, harnessing both image and video datasets to bolster the model's foundational understanding across diverse visual contexts. Subsequently, a joint instruction tuning method with task-specific projectors and decoders for end-to-end downstream training is presented. Furthermore, this work contributes a novel mask-based multi-task dataset comprising 277K samples, crafted to challenge and assess the fine-grained perception capabilities of MLLMs. The overall framework is simple, effective, and achieves state-of-the-art performance across multiple benchmarks. We also make our model, data, and code publicly accessible at https://github.com/OPPOMKLab/u-LLaVA.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. CROSS: Cascaded Distillation and Dual-Constraint Grounding for Remote Sensing Referring Segmentation

    cs.CV 2026-08 conditional novelty 6.0 of 10

    CROSS improves remote sensing referring segmentation by combining cascaded SAM distillation with contrastive learning, reporting state-of-the-art cIoU on RefSegRS and RRSIS-D.

  2. CLGRPO: Reasoning Ability Enhancement for Small VLMs

    cs.CV 2025-06 conditional novelty 4.0 of 10

    A four-stage training pipeline with a clip-low GRPO variant improves a 1B VLM's emotion recognition accuracy from 78.68% to 81.45% on EmoSet-118K.

Pith tools