REVIEW 6 cited by
Mini-InternVL: A Flexible-Transfer Pocket Multimodal Model with 5% Parameters and 90% Performance
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Multimodal large language models (MLLMs) have demonstrated impressive performance in vision-language tasks across a broad spectrum of domains. However, the large model scale and associated high computational costs pose significant challenges for training and deploying MLLMs on consumer-grade GPUs or edge devices, thereby hindering their widespread application. In this work, we introduce Mini-InternVL, a series of MLLMs with parameters ranging from 1B to 4B, which achieves 90% of the performance with only 5% of the parameters. This significant improvement in efficiency and effectiveness makes our models more accessible and applicable in various real-world scenarios. To further promote the adoption of our models, we develop a unified adaptation framework for Mini-InternVL, which enables our models to transfer and outperform specialized models in downstream tasks, including autonomous driving, medical images, and remote sensing. We believe that our study can provide valuable insights and resources to advance the development of efficient and effective MLLMs. Code is available at https://github.com/OpenGVLab/InternVL.
Forward citations
Cited by 6 Pith papers
-
Do Video-LLMs Actually Watch? Diagnosing Character-Tracking Failures in Long-Form Video
Video-LLMs scoring 37–38% on InfiniBench global appearance change answers only 4–31% under character-name swaps, so the score is not character tracking.
-
In-context Learning of Vision Language Models for Detection of Physical and Digital Attacks against Face Recognition Systems
An in-context learning framework with open-source vision-language models detects face presentation and morphing attacks without training, beating CLIP-based zero-shot baselines on PAD but with performance highly sensi...
-
A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets
Block-based diffusion generation of counterfactual image-text sets, combined with a set-aware loss, improves CLIP's compositional reasoning over several benchmarks, but the paper overstates one benchmark result and sh...
-
Challenging Vision-Language Models with Surgical Data: A New Dataset and Broad Benchmarking Study
A new surgical VQA benchmark with 167,384 questions shows generalist VLMs handle basic surgical perception but fall to near-random on medical-knowledge questions, and medical VLMs underperform generalist models.
-
Docopilot: Improving Multimodal Models for Document-Level Understanding
A new academic-paper dataset and a retrieval-free fine-tuned InternVL2 model improve multi-page document QA accuracy and latency on several benchmarks.
-
Pushing the Limits of Safety: A Technical Report on the ATLAS Challenge 2025
The ATLAS 2025 competition demonstrates that vision-language models remain highly vulnerable to flowchart-based and cross-modal jailbreak attacks, with top scores exceeding 93%.
Discussion (0). Sign in to comment.