REVIEW 63 cited by
BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Vision-Language Pre-training (VLP) has advanced the performance for many vision-language tasks. However, most existing pre-trained models only excel in either understanding-based tasks or generation-based tasks. Furthermore, performance improvement has been largely achieved by scaling up the dataset with noisy image-text pairs collected from the web, which is a suboptimal source of supervision. In this paper, we propose BLIP, a new VLP framework which transfers flexibly to both vision-language understanding and generation tasks. BLIP effectively utilizes the noisy web data by bootstrapping the captions, where a captioner generates synthetic captions and a filter removes the noisy ones. We achieve state-of-the-art results on a wide range of vision-language tasks, such as image-text retrieval (+2.7% in average recall@1), image captioning (+2.8% in CIDEr), and VQA (+1.6% in VQA score). BLIP also demonstrates strong generalization ability when directly transferred to video-language tasks in a zero-shot manner. Code, models, and datasets are released at https://github.com/salesforce/BLIP.
Forward citations
Showing 60 of 63 Pith papers that cite this
-
Enhancing Few-Shot Vision-Language Classification with Large Multimodal Model Features
SAVs extract a sparse set of attention head outputs from a frozen large multimodal model and use them as nearest-centroid features, achieving state-of-the-art few-shot vision-language classification without finetuning.
-
WaveZip: Wavelet-Driven Space-Time Decoupling for Video Token Condensation
WaveZip, a training-free wavelet method for video token condensation, reports about 10x token reduction while retaining roughly 98-99% accuracy on multi-benchmark LVLM evaluation.
-
Testing chatbots on the creation of encoders for audio conditioned image generation
All chatbot-designed audio encoders failed to align with CLIP text embeddings and produced incoherent images, while showing a surprising architectural similarity across chatbots.
-
Screen2AX: Vision-Based Approach for Automatic macOS Accessibility Generation
Screen2AX generates hierarchical macOS accessibility metadata from a screenshot and reports improved GPT-4 UI task success compared with native accessibility and OmniParser V2.
-
PoemTale Diffusion: Minimising Information Loss in Poem to Image Generation with Multi-Stage Prompt Refinement
PoemTale Diffusion generates a coherent set of images from a poem by combining emotion-based segmentation, multi-stage LLM prompt refinement, and consistent self-attention, outperforming direct poem-to-image approache...
-
Bridging the Gap in Vision Language Models in Identifying Unsafe Concepts Across Modalities
Vision-language models consistently recognize unsafe content better from text than from images, and a simplified reinforcement learning fine-tune narrows that gap.
-
Can Video LLMs Refuse to Answer? Alignment for Answerability in Video Large Language Models
Video-LLMs can be trained, via SFT or DPO on a new synthetic dataset UVQA, to refuse questions that cannot be answered from the video content, with modest cost to answerable QA performance.
-
RICO: Improving Accuracy and Completeness in Image Recaptioning via Visual Reconstruction
RICO refines image captions by reconstructing them into images with a text-to-image model and asking GPT-4o to fix discrepancies against the original, iteratively, with a DPO-distilled fast variant.
-
Towards Omnidirectional Reasoning with 360-R1: A Dataset, Benchmark, and GRPO-based Method
OmniVQA is a first open-source dataset and benchmark for 360-degree visual question answering, and 360-R1 uses GRPO with three LLM-based rewards to improve an existing multimodal model on it.
-
GeoVLM: Improving Automated Vehicle Geolocalisation Using Vision-Language Matching
GeoVLM reranks top-10 candidates from a pretrained cross-view encoder by fusing image and text embeddings, improving top-1 retrieval on VIGOR, CVUK, and University-1652.
-
Multi-modal Synthetic Data Training and Model Collapse: Insights from VLMs and Diffusion Models
Recursive training on synthetic data in multi-modal VLM and diffusion systems shows distinct collapse: caption variance grows while image variance shrinks, and frozen-model relabeling mitigates it.
-
Multi-Modal Language Models as Text-to-Image Model Evaluators
MT2IE uses a single open-source multimodal LLM to generate 20 progressively harder prompts and score image-text consistency, reproducing the 1,600-prompt GenAIBench ranking of 8 text-to-image models.
-
Target-Augmented Shared Fusion-based Multimodal Sarcasm Explanation Generation
TURBO is a target-aware multimodal sarcasm explanation model that outperforms TEAM on MORE+ automatic metrics, with the largest gains coming from the gold target-of-sarcasm input.
-
RLS3: RL-Based Synthetic Sample Selection to Enhance Spatial Reasoning in Vision-Language Models for Indoor Autonomous Perception
An RL agent generates hard synthetic spatial-reasoning examples to fine-tune VLMs, improving performance on simulated test scenes.
-
Lossy Compression with Pretrained Diffusion Models
A complete, zero-shot implementation of the DiffC algorithm lets pretrained Stable Diffusion models act as lossy image compressors at ultra-low bitrates.
-
Zero-shot Video Moment Retrieval via Off-the-shelf Multimodal Large Language Models
A tuning-free pipeline using frozen LLaMA-3, MiniGPT-v2, and Video-ChatGPT reports state-of-the-art zero-shot video moment retrieval on three benchmarks.
-
ZenSVI: An Open-Source Software for the Integrated Acquisition, Processing and Analysis of Street View Imagery Towards Scalable Urban Science
ZenSVI provides an integrated, documented Python pipeline for acquiring, cleaning, analyzing, and visualizing street view imagery for urban science.
-
SEMANTIC SEE-THROUGH GOGGLES: Wearing Linguistic Virtual Reality in (Artificial) Intelligence
A wearable AI system that turns the live view into one sentence and back into an image lets users experientially confront how linguistic mediation filters and biases perception.
-
Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning
A factorized autoregressive decoder, shared across video segments with cross-segment masking, produces denser, more localized captions online while saving about 20 percent compute versus a global decoder.
-
Image Regeneration: Evaluating Text-to-Image Model via Generating Identical Image with Multimodal Large Language Models
The paper introduces Image Regeneration, an evaluation benchmark where text-to-image models must reproduce a reference image from MLLM-generated prompts, along with the ImageRepainter framework and two new datasets.
-
The Hyperspherical Geometry of CLIP Latent Space: A Semantic Mixture Model
CLIP embeddings are modeled as a mixture of von Mises-Fisher distributions on the unit sphere, improving out-of-distribution detection and semantic decomposition over single-Gaussian baselines.
-
Multilingual Training and Evaluation Resources for Vision-Language Models
Releases regenerated multilingual training data and translated benchmarks for VLMs in five languages and demonstrates consistent benefits from multilingual training over English-only baselines.
-
Reconstructing Content with Collaborative Attention for Universal Multimodal Representation Learning
CoCoA forces an MLLM to reconstruct masked text through a single EOS token, improving multimodal embedding quality on MMEB-V1 and matching MoCa at 3B with far less pretraining data.
-
A Unified Geometric Space for Topological Alignment Between Transformer-Based Models and Human Brain Networks
Transformer models lie along a continuous arc in a new seven-dimensional Brain-like Space, where global-semantic models align with higher-order brain networks and local-reconstruction models align with sensory networks.
-
UniECS: Unified Multimodal E-Commerce Search Framework with Gated Cross-modal Fusion
UniECS, a 0.2B parameter gated multimodal encoder, reports strong Recall@K across nine e-commerce retrieval tasks, a new M-BEER benchmark, and positive online A/B metrics.
-
Invisible Injections: Exploiting Vision-Language Models Through Steganographic Prompt Embedding
Steganographic prompt injection is reported to covertly manipulate vision-language models with up to 31.8% success, but the evidence is not reproducible.
-
Affect-aware Cross-Domain Recommendation for Art Therapy via Music Preference Elicitation
A 200-person study of music-driven cross-domain recommendation for art therapy shows music-based and visual-based engines perform equally, contradicting the paper's 'outperforming' claim.
-
CLIP-RL: Surgical Scene Segmentation Using Contrastive Language-Vision Pretraining & Reinforcement Learning
A CLIP-based encoder with RL residual refinement and curriculum learning reaches 81% mIoU on EndoVis 2018 and 74.12% on EndoVis 2017 surgical segmentation.
-
AIGVE-MACS: Unified Multi-Aspect Commenting and Scoring Model for AI-Generated Video Evaluation
A finetuned vision-language model jointly predicts nine aspect scores and written comments for AI-generated videos, with a new benchmark and claims of state-of-the-art alignment with human judgment.
-
Graph-MLLM: Harnessing Multimodal Large Language Models for Multimodal Graph Learning
A unified comparison across six multimodal graph datasets shows that fine-tuned multimodal LLMs used as direct predictors achieve the highest node classification accuracy, even without graph structure input.
-
EfficientVLA: Training-Free Acceleration and Compression for Vision-Language-Action Models
EfficientVLA combines LLM layer pruning, task-aware visual token selection, and diffusion-head feature caching to cut CogACT's inference cost to 28.9% of baseline FLOPs with a 0.6% SIMPLER success drop.
-
CF-VLM:CounterFactual Vision-Language Fine-tuning
CF-VLM fine-tunes VLMs on counterfactual image-text pairs with three objectives, reporting gains on compositional reasoning benchmarks and modest hallucination reductions.
-
SOLVE: Synergy of Language-Vision and End-to-End Networks for Autonomous Driving
SOLVE couples a vision-language model and an end-to-end planner via a shared encoder and a trajectory chain-of-thought, reporting small but state-of-the-art open-loop planning gains on nuScenes.
-
Mitigating Group-Level Fairness Disparities in Federated Visual Language Models
FVL-FP combines cross-layer fair prompts, orthogonal projection off demographic subspaces, and fairness-weighted prompt fusion to reduce group bias in federated vision-language models.
-
Optimizing Multi-Round Enhanced Training in Diffusion Models for Improved Preference Understanding
The VCA framework uses diversity, consistency, and preference rewards to fine-tune a diffusion model with LoRA over multi-round dialogues, reporting improved intent alignment.
-
Image Embedding Sampling Method for Diverse Captioning
A training-free hierarchical embedding sampling method (HBoP) lets a small BLIP model generate captions as diverse as human ones, beating much larger VLMs on diversity metrics.
-
The AI-Therapist Duo: Exploring the Potential of Human-AI Collaboration in Personalized Art Therapy for PICS Intervention
A human-in-the-loop art recommender system matched expert-curated art therapy outcomes while cutting therapist selection time by more than half, but it did not outperform human-only curation on any measured outcome.
-
NanoVLMs: How small can we go and still make coherent Vision Language Models?
NanoVLMs, 5M to 25M parameter vision-language models trained on simplified GPT-4o captions, are judged by GPT-4o as nearly as coherent as the 50x larger Kosmos-2 on a 25-sample test.
-
Large Models in Dialogue for Active Perception and Anomaly Detection
An LLM and a VQA model converse to steer a simulated drone through a scene, improving descriptions and hazard detection over a static baseline.
-
StreamingRAG: Real-time Contextual Retrieval and Generation Framework
StreamingRAG constructs an evolving temporal knowledge graph from streaming video with lightweight VQA models, enabling real-time anomaly detection at a fraction of the cost of heavy captioning models.
-
How Do Generative Models Draw a Software Engineer? A Case Study on Stable Diffusion Bias
Stable Diffusion 2, XL, and 3 all produce starkly male-dominated and ethnically skewed images of software engineers, with SD3 shifting from White to Asian dominance while still suppressing Black and Arab representation.
-
Eve: Efficient Multimodal Vision Language Models with Elastic Visual Experts
A 1.8B-parameter vision-language model, Eve, uses elastic visual experts and type-aware token routing to reach a 68.87% average on six VLM benchmarks while preserving language performance.
-
VisGraphVar: A Benchmark Generator for Assessing Variability in Graph Analysis Using Large Vision-Language Models
A new graph-image benchmark generator shows that six large vision-language models are sensitive to layout, labeling, and visual defects across seven graph tasks.
-
Diff-ID: Identity Consistent Facial Image Generation and Morphing via Diffusion Models
Diff-ID trades a bit of ArcFace identity score for much lower FID, yielding the best FS/FID trade-off among tested face generators, plus qualitative morphing without per-identity fine-tuning.
-
Effectively obtaining acoustic, visual and textual data from videos
A video-processing pipeline created a 2.24 million-sample audio-image-text dataset, with text captions generated by BLIP from video frames.
-
CLARIFY: A Specialist-Generalist Framework for Accurate and Lightweight Dermatological Visual Question Answering
A specialist classifier feeding a pruned VLM with knowledge-graph grounding reports 82.1% diagnostic accuracy on a 39-image dermatology test set, about 18 percentage points above a fine-tuned VLM baseline.
-
LEARN: A Story-Driven Layout-to-Image Generation Framework for STEM Instruction
A layout-conditioned diffusion model, fine-tuned on book covers with contrastive and semantic losses, generates story-like STEM illustrations.
-
Visual Language Models as Zero-Shot Deepfake Detectors
Zero-shot VLMs scored by normalized yes/no token probabilities beat most trained deepfake detectors on a new SimSwap dataset, and a lightly fine-tuned InstructBLIP is near-perfect on DFDC-P.
-
Tempo-R0: A Video-MLLM for Temporal Video Grounding through Efficient Temporal Sensing Reinforcement Learning
A video-language model that combines adaptive frame sampling, explicit timestamps, and a reinforcement-learning reward for refusing irrelevant queries, beating prior methods on QVHighlights by about 3.5%.
-
From Pixels to Graphs: using Scene and Knowledge Graphs for HD-EPIC VQA Challenge
A challenge report showing that MLLM-generated scene graphs and ConceptNet knowledge graphs each give small accuracy gains over a video-only baseline, and that per-category selection reaches 44.21% on the HD-EPIC VQA ...
-
Seamless and Efficient Interactions within a Mixed-Dimensional Information Space
A thesis that three design strategies, multimodal AI, context-aware placement, and combined 2D/3D views, make mixed-dimensional information spaces seamless and efficient, demonstrated with three systems.
-
A Vision-Language Model for Focal Liver Lesion Classification
A text-guided CLIP-style model with a frozen BERT text encoder and cross-entropy alignment classifies focal liver lesions from multi-phase CT slices with about 79 percent average accuracy, outperforming CLIP and MedCL...
-
MemeBLIP2: A novel lightweight multimodal system to detect harmful memes
MemeBLIP2, built on BLIP-2 with linear projections, adapters, and an MLP classifier, reaches 77.5% accuracy and 79.0% F1 on PrideMM harmful meme detection, but the paper's internal inconsistencies weaken the claim.
-
Visual Language Models as Operator Agents in the Space Domain
Vision-language models can act as spacecraft operators in the KSPDG simulator from screenshots, and fine-tuning OpenVLA on ten episodes shows preliminary promise for robotic satellite inspection.
-
Multilevel Semantic-Aware Model for AI-Generated Video Quality Assessment
MSA-VQA combines CLIP-based prompt checking with cross-attention over frames to predict human quality scores for AI-generated videos, reporting state-of-the-art numbers on the T2VQA-DB benchmark.
-
ErgoChat: a Visual Query System for the Ergonomic Risk Assessment of Construction Workers
Fine-tuning MiniGPT-v2 on a new 1,900-image construction ergonomics dataset improves visual question answering and image captioning of postural risk versus the same model without fine-tuning.
-
SubstationAI: Multimodal Large Model-Based Approaches for Analyzing Substation Equipment Faults
SubstationAI, a fine-tuned LLaVA-1.5-7B model augmented with a fault knowledge base, receives higher expert ratings than GPT-4 for substation fault reports, but suspected train/test overlap makes the result unreliable.
-
Health AI Developer Foundations
Health AI Developer Foundations packages six domain-specific medical embedding models into one platform, claiming large data and compute savings for downstream health ML tasks.
-
Investigating the Invertibility of Multimodal Latent Spaces: Limitations of Optimization-Based Methods
Optimization can force BLIP, Flux, Whisper, and Chatterbox to hit textual targets, but the inverted inputs are perceptually incoherent and the recovered text embeddings are semantically meaningless.
-
E.A.R.T.H.: Structuring Creative Evolution through Model Error in Generative AI
A five-stage pipeline that induces, scores, rewrites, and validates model errors reports large creativity gains that largely arise from selection on the measurement metric itself.
Discussion (0). Continue with ORCID to comment.