Pith. sign in

REVIEW 2 cited by

Cross-Modal Projection in Multimodal LLMs Doesn't Really Project Visual Attributes to Textual Space

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.16832 v2 pith:4EXUYT2K submitted 2024-02-26 cs.CL cs.AIcs.CV

classification cs.CLcs.AIcs.CV
keywords domain-specificvisualattributesprojectioncross-modalfine-tunedlanguagemllms
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Multimodal large language models (MLLMs) like LLaVA and GPT-4(V) enable general-purpose conversations about images with the language modality. As off-the-shelf MLLMs may have limited capabilities on images from domains like dermatology and agriculture, they must be fine-tuned to unlock domain-specific applications. The prevalent architecture of current open-source MLLMs comprises two major modules: an image-language (cross-modal) projection network and a large language model. It is desirable to understand the roles of these two modules in modeling domain-specific visual attributes to inform the design of future models and streamline the interpretability efforts on the current models. To this end, via experiments on 4 datasets and under 2 fine-tuning settings, we find that as the MLLM is fine-tuned, it indeed gains domain-specific visual capabilities, but the updates do not lead to the projection extracting relevant domain-specific visual attributes. Our results indicate that the domain-specific visual attributes are modeled by the LLM, even when only the projection is fine-tuned. Through this study, we offer a potential reinterpretation of the role of cross-modal projections in MLLM architectures. Project webpage: https://claws-lab.github.io/projection-in-MLLMs/

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding

    cs.CV 2026-08 conditional novelty 6.0 of 10

    A unified visual grounding framework combining a broadcast cross-attention head, a JEPA auxiliary loss, and an MLLM-generated caption dataset preserves representation diversity and generalizes across RefCOCO/+/g.

  2. A Review of Multimodal Explainable Artificial Intelligence: Past, Present and Future

    cs.CV 2024-12 conditional novelty 4.0 of 10

    A historical review that organizes multimodal explainability methods into four chronological eras and three explainability types, extending coverage to generative LLMs.

Pith tools