Pith. sign in

REVIEW 4 cited by

Reformulating Vision-Language Foundation Models and Datasets Towards Universal Multimodal Assistants

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2310.00653 v1 pith:UCNOS4BH submitted 2023-10-01 cs.CV cs.AI

classification cs.CVcs.AI
keywords modelsdatasetslanguagemultimodalmuffindatasetinstructionmodel
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Recent Multimodal Large Language Models (MLLMs) exhibit impressive abilities to perceive images and follow open-ended instructions. The capabilities of MLLMs depend on two crucial factors: the model architecture to facilitate the feature alignment of visual modules and large language models; the multimodal instruction tuning datasets for human instruction following. (i) For the model architecture, most existing models introduce an external bridge module to connect vision encoders with language models, which needs an additional feature-alignment pre-training. In this work, we discover that compact pre-trained vision language models can inherently serve as ``out-of-the-box'' bridges between vision and language. Based on this, we propose Muffin framework, which directly employs pre-trained vision-language models to act as providers of visual signals. (ii) For the multimodal instruction tuning datasets, existing methods omit the complementary relationship between different datasets and simply mix datasets from different tasks. Instead, we propose UniMM-Chat dataset which explores the complementarities of datasets to generate 1.1M high-quality and diverse multimodal instructions. We merge information describing the same image from diverse datasets and transforms it into more knowledge-intensive conversation data. Experimental results demonstrate the effectiveness of the Muffin framework and UniMM-Chat dataset. Muffin achieves state-of-the-art performance on a wide range of vision-language tasks, significantly surpassing state-of-the-art models like LLaVA and InstructBLIP. Our model and dataset are all accessible at https://github.com/thunlp/muffin.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. CHiP: Cross-modal Hierarchical Direct Preference Optimization for Multimodal LLMs

    cs.CL 2025-01 conditional novelty 6.0 of 10

    CHiP adds image-level and phrase/token-level preference signals to DPO for multimodal LLMs, and the authors report large hallucination-rate reductions on Object HalBench, AMBER, MMHal, and HallusionBench.

  2. Learning to Correction: Explainable Feedback Generation for Visual Commonsense Reasoning Distractor

    cs.CV 2024-12 conditional novelty 6.0 of 10

    A new GPT-4-generated dataset of distractors and corrective feedback for visual commonsense reasoning, plus a compact LMM (PEIFG) that produces explainable corrections and beats existing baselines in automatic and hum...

  3. From Visuals to Vocabulary: Establishing Equivalence Between Image and Text Token Through Autoregressive Pre-training in MLLMs

    cs.CV 2025-02 conditional novelty 4.0 of 10

    Adding an L2 loss that pushes the language model's image hidden states back toward the input image embeddings improves LLaVA-style models on several VQA benchmarks, with some benchmarks unaffected or slightly worse.

  4. ASPO: Adaptive Sentence-Level Preference Optimization for Fine-Grained Multimodal Reasoning

    cs.CL 2025-05 reject novelty 2.0 of 10

    ASPO's adaptive sentence-level loss, by the paper's own definitions, reduces exactly to the standard DPO loss, leaving no difference in the optimization objective.

Pith tools