Pith. sign in

REVIEW 3 cited by

An Empirical Study of Scaling Instruct-Tuned Large Multimodal Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2309.09958 v1 pith:YMRDBEDM submitted 2023-09-18 cs.CV cs.CL

classification cs.CVcs.CL
keywords performancelanguagemodelsscalingtuningcapabilitiesdataempirical
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Visual instruction tuning has recently shown encouraging progress with open-source large multimodal models (LMM) such as LLaVA and MiniGPT-4. However, most existing studies of open-source LMM are performed using models with 13B parameters or smaller. In this paper we present an empirical study of scaling LLaVA up to 33B and 65B/70B, and share our findings from our explorations in image resolution, data mixing and parameter-efficient training methods such as LoRA/QLoRA. These are evaluated by their impact on the multi-modal and language capabilities when completing real-world tasks in the wild. We find that scaling LMM consistently enhances model performance and improves language capabilities, and performance of LoRA/QLoRA tuning of LMM are comparable to the performance of full-model fine-tuning. Additionally, the study highlights the importance of higher image resolutions and mixing multimodal-language data to improve LMM performance, and visual instruction tuning can sometimes improve LMM's pure language capability. We hope that this study makes state-of-the-art LMM research at a larger scale more accessible, thus helping establish stronger baselines for future research. Code and checkpoints will be made public.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Describe, Adapt and Combine: Empowering CLIP Encoders for Open-set 3D Object Retrieval

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A CLIP plus MLLM pipeline with additive-bias Low-Rank adaptation retrieves 3D objects of unseen categories from multi-view images, outperforming prior art by about 10% mAP on average.

  2. Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A dexterous VLA pretrained on a 2.5M-instance human hand motion dataset transfers skills to a real robot hand, outperforming baselines in manipulation tasks.

  3. Transport Discrepancy as a Reliability Signal for Vision-Language-Action Models

    cs.RO 2025-12 conditional novelty 5.0 of 10

    A transport-cost-based gate that modulates observation features in flow-matching VLA policies improves reported success rates on long-horizon and distribution-shifted robot tasks.

Pith tools