Pith. sign in

REVIEW 2 cited by

MMRL: Multi-Modal Representation Learning for Vision-Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2503.08497 v2 pith:UOSPVVQF submitted 2025-03-11 cs.LG cs.CV

classification cs.LGcs.CV
keywords featuresrepresentationclassmmrltokensknowledgelearningmodels
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Large-scale pre-trained Vision-Language Models (VLMs) have become essential for transfer learning across diverse tasks. However, adapting these models with limited few-shot data often leads to overfitting, diminishing their performance on new tasks. To tackle this issue, we propose a novel Multi-Modal Representation Learning (MMRL) framework that introduces a shared, learnable, and modality-agnostic representation space. MMRL projects the space tokens to text and image representation tokens, facilitating more effective multi-modal interactions. Unlike previous approaches that solely optimize class token features, MMRL integrates representation tokens at higher layers of the encoders--where dataset-specific features are more prominent--while preserving generalized knowledge in the lower layers. During training, both representation and class features are optimized, with trainable projection layer applied to the representation tokens, whereas the class token projection layer remains frozen to retain pre-trained knowledge. Furthermore, a regularization term is introduced to align the class features and text features with the zero-shot features from the frozen VLM, thereby safeguarding the model's generalization capacity. For inference, a decoupling strategy is employed, wherein both representation and class features are utilized for base classes, while only the class features, which retain more generalized knowledge, are used for new tasks. Extensive experiments across 15 datasets demonstrate that MMRL outperforms state-of-the-art methods, achieving a balanced trade-off between task-specific adaptation and generalization. Code is available at https://github.com/yunncheng/MMRL.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. SemPT: Semantic Prompt Tuning for Vision-Language Models

    cs.CV 2025-08 conditional novelty 5.0 of 10

    SemPT, a semantic prompt tuning method, uses shared attribute-level descriptions and adaptive embedding selection to slightly improve base-to-novel, cross-dataset, cross-domain, and few-shot generalization of CLIP pro...

  2. MMRL++: Parameter-Efficient and Interaction-Aware Representation Learning for Vision-Language Models

    cs.CV 2025-05 conditional novelty 4.0 of 10

    MMRL++ inserts shared, learnable representation tokens into the upper layers of CLIP's image and text encoders and uses low-rank shared aligners, achieving state-of-the-art base-to-novel harmonic mean accuracy on 11 d...

Pith tools