Pith. sign in

REVIEW 4 cited by

Efficient Large Multi-modal Models via Visual Context Compression

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.20092 v2 pith:TFJ2U5ND submitted 2024-06-28 cs.CV

classification cs.CV
keywords visualtokenstrainingcompressioncontextefficiencymodelswhile
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

While significant advancements have been made in compressed representations for text embeddings in large language models (LLMs), the compression of visual tokens in multi-modal LLMs (MLLMs) has remained a largely overlooked area. In this work, we present the study on the analysis of redundancy concerning visual tokens and efficient training within these models. Our initial experiments show that eliminating up to 70% of visual tokens at the testing stage by simply average pooling only leads to a minimal 3% reduction in visual question answering accuracy on the GQA benchmark, indicating significant redundancy in visual context. Addressing this, we introduce Visual Context Compressor, which reduces the number of visual tokens to enhance training and inference efficiency without sacrificing performance. To minimize information loss caused by the compression on visual tokens while maintaining training efficiency, we develop LLaVolta as a light and staged training scheme that incorporates stage-wise visual context compression to progressively compress the visual tokens from heavily to lightly compression during training, yielding no loss of information when testing. Extensive experiments demonstrate that our approach enhances the performance of MLLMs in both image-language and video-language understanding, while also significantly cutting training costs and improving inference efficiency.

Discussion (0). Sign in to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Efficient Multi-modal Long Context Learning for Training-free Adaptation

    cs.CV 2025-05 conditional novelty 6.0 of 10

    EMLoC prunes and compresses demonstration examples in multimodal long contexts, reducing inference cost up to 77% while matching or slightly beating full-context accuracy.

  2. Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding

    cs.CV 2025-07 conditional novelty 5.0 of 10

    Fast3D prunes up to 90% of object-centric visual tokens in 3D MLLMs while preserving about 96.8% of original benchmark performance, using a trained attention predictor and adaptive layer-wise pruning.

  3. GreedyPrune: Retenting Critical Visual Token Set for Large Vision Language Models

    cs.CV 2025-06 conditional novelty 5.0 of 10

    A training-free token pruning method that combines cosine-similarity saliency with greedy redundancy removal to preserve accuracy at high compression ratios.

  4. Structured Prompting and Multi-Agent Knowledge Distillation for Traffic Video Interpretation and Risk Inference

    cs.CV 2025-08 unverdicted novelty 4.0 of 10

    A 3-billion-parameter vision-language model distilled from GPT-4o and o3-mini pseudo-annotations matches its teachers on traffic-video captioning metrics.

Pith tools