Pith. sign in

REVIEW 6 cited by

Mini-InternVL: A Flexible-Transfer Pocket Multimodal Model with 5% Parameters and 90% Performance

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.16261 v3 pith:4VDWM77N submitted 2024-10-21 cs.CV

classification cs.CV
keywords modelsmllmsmini-internvlparametersperformancelargemodelmultimodal
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Multimodal large language models (MLLMs) have demonstrated impressive performance in vision-language tasks across a broad spectrum of domains. However, the large model scale and associated high computational costs pose significant challenges for training and deploying MLLMs on consumer-grade GPUs or edge devices, thereby hindering their widespread application. In this work, we introduce Mini-InternVL, a series of MLLMs with parameters ranging from 1B to 4B, which achieves 90% of the performance with only 5% of the parameters. This significant improvement in efficiency and effectiveness makes our models more accessible and applicable in various real-world scenarios. To further promote the adoption of our models, we develop a unified adaptation framework for Mini-InternVL, which enables our models to transfer and outperform specialized models in downstream tasks, including autonomous driving, medical images, and remote sensing. We believe that our study can provide valuable insights and resources to advance the development of efficient and effective MLLMs. Code is available at https://github.com/OpenGVLab/InternVL.

Discussion (0). Sign in to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Do Video-LLMs Actually Watch? Diagnosing Character-Tracking Failures in Long-Form Video

    cs.CV 2026-07 conditional novelty 6.5 of 10

    Video-LLMs scoring 37–38% on InfiniBench global appearance change answers only 4–31% under character-name swaps, so the score is not character tracking.

  2. In-context Learning of Vision Language Models for Detection of Physical and Digital Attacks against Face Recognition Systems

    cs.CV 2025-07 conditional novelty 6.0 of 10

    An in-context learning framework with open-source vision-language models detects face presentation and morphing attacks without training, beating CLIP-based zero-shot baselines on PAD but with performance highly sensi...

  3. A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets

    cs.CV 2025-07 conditional novelty 6.0 of 10

    Block-based diffusion generation of counterfactual image-text sets, combined with a set-aware loss, improves CLIP's compositional reasoning over several benchmarks, but the paper overstates one benchmark result and sh...

  4. Challenging Vision-Language Models with Surgical Data: A New Dataset and Broad Benchmarking Study

    cs.CV 2025-06 conditional novelty 6.0 of 10

    A new surgical VQA benchmark with 167,384 questions shows generalist VLMs handle basic surgical perception but fall to near-random on medical-knowledge questions, and medical VLMs underperform generalist models.

  5. Docopilot: Improving Multimodal Models for Document-Level Understanding

    cs.CV 2025-07 conditional novelty 5.0 of 10

    A new academic-paper dataset and a retrieval-free fine-tuned InternVL2 model improve multi-page document QA accuracy and latency on several benchmarks.

  6. Pushing the Limits of Safety: A Technical Report on the ATLAS Challenge 2025

    cs.CR 2025-06 conditional novelty 3.0 of 10

    The ATLAS 2025 competition demonstrates that vision-language models remain highly vulnerable to flowchart-based and cross-modal jailbreak attacks, with top scores exceeding 93%.

Pith tools