Pith. sign in

REVIEW 1 cited by

Can Linguistic Knowledge Improve Multimodal Alignment in Vision-Language Pretraining?

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2308.12898 v2 pith:EPOMEZ4E submitted 2023-08-24 cs.MM cs.AIcs.CLcs.CV

classification cs.MMcs.AIcs.CLcs.CV
keywords knowledgelinguisticmultimodalalignmentsemanticsyntaxbenchmarkcombinations
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The multimedia community has shown a significant interest in perceiving and representing the physical world with multimodal pretrained neural network models, and among them, the visual-language pertaining (VLP) is, currently, the most captivating topic. However, there have been few endeavors dedicated to the exploration of 1) whether essential linguistic knowledge (e.g., semantics and syntax) can be extracted during VLP, and 2) how such linguistic knowledge impact or enhance the multimodal alignment. In response, here we aim to elucidate the impact of comprehensive linguistic knowledge, including semantic expression and syntactic structure, on multimodal alignment. Specifically, we design and release the SNARE, the first large-scale multimodal alignment probing benchmark, to detect the vital linguistic components, e.g., lexical, semantic, and syntax knowledge, containing four tasks: Semantic structure, Negation logic, Attribute ownership, and Relationship composition. Based on our proposed probing benchmarks, our holistic analyses of five advanced VLP models illustrate that the VLP model: i) shows insensitivity towards complex syntax structures and relies on content words for sentence comprehension; ii) demonstrates limited comprehension of combinations between sentences and negations; iii) faces challenges in determining the presence of actions or spatial relationships within visual information and struggles with verifying the correctness of triple combinations. We make our benchmark and code available at \url{https://github.com/WangFei-2019/SNARE/}.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Seeing Syntax: Uncovering Syntactic Learning Limitations in Vision-Language Models

    cs.CV 2024-12 conditional novelty 6.0 of 10

    Vision-language text encoders encode less syntactic structure than text-only encoders, and contrastive pre-training rather than model size or data volume accounts for most of the deficit.

Pith tools