Pith. sign in

REVIEW 2 cited by

FLAVA: A Foundational Language And Vision Alignment Model

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2112.04482 v3 pith:GVXXTUZZ submitted 2021-12-08 cs.CV cs.CL

classification cs.CVcs.CL
keywords tasksvisionlanguagemodelmodalitiesflavafoundationgood
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

State-of-the-art vision and vision-and-language models rely on large-scale visio-linguistic pretraining for obtaining good performance on a variety of downstream tasks. Generally, such models are often either cross-modal (contrastive) or multi-modal (with earlier fusion) but not both; and they often only target specific modalities or tasks. A promising direction would be to use a single holistic universal model, as a "foundation", that targets all modalities at once -- a true vision and language foundation model should be good at vision tasks, language tasks, and cross- and multi-modal vision and language tasks. We introduce FLAVA as such a model and demonstrate impressive performance on a wide range of 35 tasks spanning these target modalities.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. NLKI: A lightweight Natural Language Knowledge Integration Framework for Improving Small VLMs in Commonsense VQA Tasks

    cs.CL 2025-08 conditional novelty 6.0 of 10

    NLKI combines fine-tuned dense retrieval, LLM-generated explanations, and noise-robust losses to improve small VLMs on commonsense VQA.

  2. Learning to Detect Cross-Modal Negation: An Analysis of Latent Representations and an Attention-Based Solution

    cs.AI 2026-07 reject novelty 4.0 of 10

    Negation is reported as a non-separable signal in standard VLM embeddings, cross-modal attention recovers up to +7% F1, and the paper claims visual negation depends on textual context.

Pith tools