Pith. sign in

REVIEW 63 cited by

BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2201.12086 v2 pith:VK76W57U submitted 2022-01-28 cs.CV

classification cs.CV
keywords tasksblipvision-languagenoisybootstrappingcaptionsgenerationimage-text
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Vision-Language Pre-training (VLP) has advanced the performance for many vision-language tasks. However, most existing pre-trained models only excel in either understanding-based tasks or generation-based tasks. Furthermore, performance improvement has been largely achieved by scaling up the dataset with noisy image-text pairs collected from the web, which is a suboptimal source of supervision. In this paper, we propose BLIP, a new VLP framework which transfers flexibly to both vision-language understanding and generation tasks. BLIP effectively utilizes the noisy web data by bootstrapping the captions, where a captioner generates synthetic captions and a filter removes the noisy ones. We achieve state-of-the-art results on a wide range of vision-language tasks, such as image-text retrieval (+2.7% in average recall@1), image captioning (+2.8% in CIDEr), and VQA (+1.6% in VQA score). BLIP also demonstrates strong generalization ability when directly transferred to video-language tasks in a zero-shot manner. Code, models, and datasets are released at https://github.com/salesforce/BLIP.

Discussion (0). Continue with ORCID to comment.

Forward citations

Showing 60 of 63 Pith papers that cite this

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 867 citations worldwide. See all 63 Pith citations

  1. Enhancing Few-Shot Vision-Language Classification with Large Multimodal Model Features

    cs.CV 2024-11 conditional novelty 7.0 of 10

    SAVs extract a sparse set of attention head outputs from a frozen large multimodal model and use them as nearest-centroid features, achieving state-of-the-art few-shot vision-language classification without finetuning.

  2. WaveZip: Wavelet-Driven Space-Time Decoupling for Video Token Condensation

    cs.CV 2026-07 conditional novelty 6.0 of 10

    WaveZip, a training-free wavelet method for video token condensation, reports about 10x token reduction while retaining roughly 98-99% accuracy on multi-benchmark LVLM evaluation.

  3. Testing chatbots on the creation of encoders for audio conditioned image generation

    cs.SD 2025-09 conditional novelty 6.0 of 10

    All chatbot-designed audio encoders failed to align with CLIP text embeddings and produced incoherent images, while showing a surprising architectural similarity across chatbots.

  4. Screen2AX: Vision-Based Approach for Automatic macOS Accessibility Generation

    cs.LG 2025-07 conditional novelty 6.0 of 10

    Screen2AX generates hierarchical macOS accessibility metadata from a screenshot and reports improved GPT-4 UI task success compared with native accessibility and OmniParser V2.

  5. PoemTale Diffusion: Minimising Information Loss in Poem to Image Generation with Multi-Stage Prompt Refinement

    cs.CV 2025-07 conditional novelty 6.0 of 10

    PoemTale Diffusion generates a coherent set of images from a poem by combining emotion-based segmentation, multi-stage LLM prompt refinement, and consistent self-attention, outperforming direct poem-to-image approache...

  6. Bridging the Gap in Vision Language Models in Identifying Unsafe Concepts Across Modalities

    cs.CR 2025-07 conditional novelty 6.0 of 10

    Vision-language models consistently recognize unsafe content better from text than from images, and a simplified reinforcement learning fine-tune narrows that gap.

  7. Can Video LLMs Refuse to Answer? Alignment for Answerability in Video Large Language Models

    cs.CV 2025-07 conditional novelty 6.0 of 10

    Video-LLMs can be trained, via SFT or DPO on a new synthetic dataset UVQA, to refuse questions that cannot be answered from the video content, with modest cost to answerable QA performance.

  8. RICO: Improving Accuracy and Completeness in Image Recaptioning via Visual Reconstruction

    cs.CV 2025-05 conditional novelty 6.0 of 10

    RICO refines image captions by reconstructing them into images with a text-to-image model and asking GPT-4o to fix discrepancies against the original, iteratively, with a DPO-distilled fast variant.

  9. Towards Omnidirectional Reasoning with 360-R1: A Dataset, Benchmark, and GRPO-based Method

    cs.CV 2025-05 conditional novelty 6.0 of 10

    OmniVQA is a first open-source dataset and benchmark for 360-degree visual question answering, and 360-R1 uses GRPO with three LLM-based rewards to improve an existing multimodal model on it.

  10. GeoVLM: Improving Automated Vehicle Geolocalisation Using Vision-Language Matching

    cs.CV 2025-05 conditional novelty 6.0 of 10

    GeoVLM reranks top-10 candidates from a pretrained cross-view encoder by fusing image and text embeddings, improving top-1 retrieval on VIGOR, CVUK, and University-1652.

  11. Multi-modal Synthetic Data Training and Model Collapse: Insights from VLMs and Diffusion Models

    cs.LG 2025-05 conditional novelty 6.0 of 10

    Recursive training on synthetic data in multi-modal VLM and diffusion systems shows distinct collapse: caption variance grows while image variance shrinks, and frozen-model relabeling mitigates it.

  12. Multi-Modal Language Models as Text-to-Image Model Evaluators

    cs.CV 2025-05 conditional novelty 6.0 of 10

    MT2IE uses a single open-source multimodal LLM to generate 20 progressively harder prompts and score image-text consistency, reproducing the 1,600-prompt GenAIBench ranking of 8 text-to-image models.

  13. Target-Augmented Shared Fusion-based Multimodal Sarcasm Explanation Generation

    cs.CL 2025-02 conditional novelty 6.0 of 10

    TURBO is a target-aware multimodal sarcasm explanation model that outperforms TEAM on MORE+ automatic metrics, with the largest gains coming from the gold target-of-sarcasm input.

  14. RLS3: RL-Based Synthetic Sample Selection to Enhance Spatial Reasoning in Vision-Language Models for Indoor Autonomous Perception

    cs.CV 2025-01 conditional novelty 6.0 of 10

    An RL agent generates hard synthetic spatial-reasoning examples to fine-tune VLMs, improving performance on simulated test scenes.

  15. Lossy Compression with Pretrained Diffusion Models

    cs.CV 2025-01 conditional novelty 6.0 of 10

    A complete, zero-shot implementation of the DiffC algorithm lets pretrained Stable Diffusion models act as lossy image compressors at ultra-low bitrates.

  16. Zero-shot Video Moment Retrieval via Off-the-shelf Multimodal Large Language Models

    cs.MM 2025-01 conditional novelty 6.0 of 10

    A tuning-free pipeline using frozen LLaMA-3, MiniGPT-v2, and Video-ChatGPT reports state-of-the-art zero-shot video moment retrieval on three benchmarks.

  17. ZenSVI: An Open-Source Software for the Integrated Acquisition, Processing and Analysis of Street View Imagery Towards Scalable Urban Science

    cs.CV 2024-12 conditional novelty 6.0 of 10

    ZenSVI provides an integrated, documented Python pipeline for acquiring, cleaning, analyzing, and visualizing street view imagery for urban science.

  18. SEMANTIC SEE-THROUGH GOGGLES: Wearing Linguistic Virtual Reality in (Artificial) Intelligence

    cs.HC 2024-12 conditional novelty 6.0 of 10

    A wearable AI system that turns the live view into one sentence and back into an image lets users experientially confront how linguistic mediation filters and biases perception.

  19. Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning

    cs.CV 2024-11 conditional novelty 6.0 of 10

    A factorized autoregressive decoder, shared across video segments with cross-segment masking, produces denser, more localized captions online while saving about 20 percent compute versus a global decoder.

  20. Image Regeneration: Evaluating Text-to-Image Model via Generating Identical Image with Multimodal Large Language Models

    cs.CV 2024-11 conditional novelty 6.0 of 10

    The paper introduces Image Regeneration, an evaluation benchmark where text-to-image models must reproduce a reference image from MLLM-generated prompts, along with the ImageRepainter framework and two new datasets.

  21. The Hyperspherical Geometry of CLIP Latent Space: A Semantic Mixture Model

    cs.LG 2026-07 conditional novelty 5.0 of 10

    CLIP embeddings are modeled as a mixture of von Mises-Fisher distributions on the unit sphere, improving out-of-distribution detection and semantic decomposition over single-Gaussian baselines.

  22. Multilingual Training and Evaluation Resources for Vision-Language Models

    cs.CL 2026-04 conditional novelty 5.0 of 10

    Releases regenerated multilingual training data and translated benchmarks for VLMs in five languages and demonstrates consistent benefits from multilingual training over English-only baselines.

  23. Reconstructing Content with Collaborative Attention for Universal Multimodal Representation Learning

    cs.IR 2026-03 conditional novelty 5.0 of 10

    CoCoA forces an MLLM to reconstruct masked text through a single EOS token, improving multimodal embedding quality on MMEB-V1 and matching MoCa at 3B with far less pretraining data.

  24. A Unified Geometric Space for Topological Alignment Between Transformer-Based Models and Human Brain Networks

    cs.AI 2025-10 conditional novelty 5.0 of 10

    Transformer models lie along a continuous arc in a new seven-dimensional Brain-like Space, where global-semantic models align with higher-order brain networks and local-reconstruction models align with sensory networks.

  25. UniECS: Unified Multimodal E-Commerce Search Framework with Gated Cross-modal Fusion

    cs.IR 2025-08 conditional novelty 5.0 of 10

    UniECS, a 0.2B parameter gated multimodal encoder, reports strong Recall@K across nine e-commerce retrieval tasks, a new M-BEER benchmark, and positive online A/B metrics.

  26. Invisible Injections: Exploiting Vision-Language Models Through Steganographic Prompt Embedding

    cs.CR 2025-07 reject novelty 5.0 of 10

    Steganographic prompt injection is reported to covertly manipulate vision-language models with up to 31.8% success, but the evidence is not reproducible.

  27. Affect-aware Cross-Domain Recommendation for Art Therapy via Music Preference Elicitation

    cs.IR 2025-07 reject novelty 5.0 of 10

    A 200-person study of music-driven cross-domain recommendation for art therapy shows music-based and visual-based engines perform equally, contradicting the paper's 'outperforming' claim.

  28. CLIP-RL: Surgical Scene Segmentation Using Contrastive Language-Vision Pretraining & Reinforcement Learning

    eess.IV 2025-07 conditional novelty 5.0 of 10

    A CLIP-based encoder with RL residual refinement and curriculum learning reaches 81% mIoU on EndoVis 2018 and 74.12% on EndoVis 2017 surgical segmentation.

  29. AIGVE-MACS: Unified Multi-Aspect Commenting and Scoring Model for AI-Generated Video Evaluation

    cs.CV 2025-07 conditional novelty 5.0 of 10

    A finetuned vision-language model jointly predicts nine aspect scores and written comments for AI-generated videos, with a new benchmark and claims of state-of-the-art alignment with human judgment.

  30. Graph-MLLM: Harnessing Multimodal Large Language Models for Multimodal Graph Learning

    cs.LG 2025-06 conditional novelty 5.0 of 10

    A unified comparison across six multimodal graph datasets shows that fine-tuned multimodal LLMs used as direct predictors achieve the highest node classification accuracy, even without graph structure input.

  31. EfficientVLA: Training-Free Acceleration and Compression for Vision-Language-Action Models

    cs.CV 2025-06 conditional novelty 5.0 of 10

    EfficientVLA combines LLM layer pruning, task-aware visual token selection, and diffusion-head feature caching to cut CogACT's inference cost to 28.9% of baseline FLOPs with a 0.6% SIMPLER success drop.

  32. CF-VLM:CounterFactual Vision-Language Fine-tuning

    cs.LG 2025-06 conditional novelty 5.0 of 10

    CF-VLM fine-tunes VLMs on counterfactual image-text pairs with three objectives, reporting gains on compositional reasoning benchmarks and modest hallucination reductions.

  33. SOLVE: Synergy of Language-Vision and End-to-End Networks for Autonomous Driving

    cs.CV 2025-05 conditional novelty 5.0 of 10

    SOLVE couples a vision-language model and an end-to-end planner via a shared encoder and a trajectory chain-of-thought, reporting small but state-of-the-art open-loop planning gains on nuScenes.

  34. Mitigating Group-Level Fairness Disparities in Federated Visual Language Models

    cs.CV 2025-05 conditional novelty 5.0 of 10

    FVL-FP combines cross-layer fair prompts, orthogonal projection off demographic subspaces, and fairness-weighted prompt fusion to reduce group bias in federated vision-language models.

  35. Optimizing Multi-Round Enhanced Training in Diffusion Models for Improved Preference Understanding

    cs.CV 2025-04 reject novelty 5.0 of 10

    The VCA framework uses diversity, consistency, and preference rewards to fine-tune a diffusion model with LoRA over multi-round dialogues, reporting improved intent alignment.

  36. Image Embedding Sampling Method for Diverse Captioning

    cs.CV 2025-02 conditional novelty 5.0 of 10

    A training-free hierarchical embedding sampling method (HBoP) lets a small BLIP model generate captions as diverse as human ones, beating much larger VLMs on diversity metrics.

  37. The AI-Therapist Duo: Exploring the Potential of Human-AI Collaboration in Personalized Art Therapy for PICS Intervention

    cs.HC 2025-02 reject novelty 5.0 of 10

    A human-in-the-loop art recommender system matched expert-curated art therapy outcomes while cutting therapist selection time by more than half, but it did not outperform human-only curation on any measured outcome.

  38. NanoVLMs: How small can we go and still make coherent Vision Language Models?

    cs.CV 2025-02 reject novelty 5.0 of 10

    NanoVLMs, 5M to 25M parameter vision-language models trained on simplified GPT-4o captions, are judged by GPT-4o as nearly as coherent as the 50x larger Kosmos-2 on a 25-sample test.

  39. Large Models in Dialogue for Active Perception and Anomaly Detection

    cs.CV 2025-01 conditional novelty 5.0 of 10

    An LLM and a VQA model converse to steer a simulated drone through a scene, improving descriptions and hazard detection over a static baseline.

  40. StreamingRAG: Real-time Contextual Retrieval and Generation Framework

    cs.CV 2025-01 conditional novelty 5.0 of 10

    StreamingRAG constructs an evolving temporal knowledge graph from streaming video with lightweight VQA models, enabling real-time anomaly detection at a fraction of the cost of heavy captioning models.

  41. How Do Generative Models Draw a Software Engineer? A Case Study on Stable Diffusion Bias

    cs.SE 2025-01 conditional novelty 5.0 of 10

    Stable Diffusion 2, XL, and 3 all produce starkly male-dominated and ethnically skewed images of software engineers, with SD3 shifting from White to Asian dominance while still suppressing Black and Arab representation.

  42. Eve: Efficient Multimodal Vision Language Models with Elastic Visual Experts

    cs.CV 2025-01 conditional novelty 5.0 of 10

    A 1.8B-parameter vision-language model, Eve, uses elastic visual experts and type-aware token routing to reach a 68.87% average on six VLM benchmarks while preserving language performance.

  43. VisGraphVar: A Benchmark Generator for Assessing Variability in Graph Analysis Using Large Vision-Language Models

    cs.CV 2024-11 conditional novelty 5.0 of 10

    A new graph-image benchmark generator shows that six large vision-language models are sensitive to layout, labeling, and visual defects across seven graph tasks.

  44. Diff-ID: Identity Consistent Facial Image Generation and Morphing via Diffusion Models

    cs.CV 2026-07 conditional novelty 4.0 of 10

    Diff-ID trades a bit of ArcFace identity score for much lower FID, yielding the best FS/FID trade-off among tested face generators, plus qualitative morphing without per-identity fine-tuning.

  45. Effectively obtaining acoustic, visual and textual data from videos

    cs.MM 2025-09 conditional novelty 4.0 of 10

    A video-processing pipeline created a 2.24 million-sample audio-image-text dataset, with text captions generated by BLIP from video frames.

  46. CLARIFY: A Specialist-Generalist Framework for Accurate and Lightweight Dermatological Visual Question Answering

    cs.CV 2025-08 reject novelty 4.0 of 10

    A specialist classifier feeding a pruned VLM with knowledge-graph grounding reports 82.1% diagnostic accuracy on a 39-image dermatology test set, about 18 percentage points above a fine-tuned VLM baseline.

  47. LEARN: A Story-Driven Layout-to-Image Generation Framework for STEM Instruction

    cs.CV 2025-08 conditional novelty 4.0 of 10

    A layout-conditioned diffusion model, fine-tuned on book covers with contrastive and semantic losses, generates story-like STEM illustrations.

  48. Visual Language Models as Zero-Shot Deepfake Detectors

    cs.CV 2025-07 conditional novelty 4.0 of 10

    Zero-shot VLMs scored by normalized yes/no token probabilities beat most trained deepfake detectors on a new SimSwap dataset, and a lightly fine-tuned InstructBLIP is near-perfect on DFDC-P.

  49. Tempo-R0: A Video-MLLM for Temporal Video Grounding through Efficient Temporal Sensing Reinforcement Learning

    cs.CV 2025-07 conditional novelty 4.0 of 10

    A video-language model that combines adaptive frame sampling, explicit timestamps, and a reinforcement-learning reward for refusing irrelevant queries, beating prior methods on QVHighlights by about 3.5%.

  50. From Pixels to Graphs: using Scene and Knowledge Graphs for HD-EPIC VQA Challenge

    cs.CV 2025-06 conditional novelty 4.0 of 10

    A challenge report showing that MLLM-generated scene graphs and ConceptNet knowledge graphs each give small accuracy gains over a video-only baseline, and that per-category selection reaches 44.21% on the HD-EPIC VQA ...

  51. Seamless and Efficient Interactions within a Mixed-Dimensional Information Space

    cs.HC 2025-06 conditional novelty 4.0 of 10

    A thesis that three design strategies, multimodal AI, context-aware placement, and combined 2D/3D views, make mixed-dimensional information spaces seamless and efficient, demonstrated with three systems.

  52. A Vision-Language Model for Focal Liver Lesion Classification

    cs.CV 2025-05 conditional novelty 4.0 of 10

    A text-guided CLIP-style model with a frozen BERT text encoder and cross-entropy alignment classifies focal liver lesions from multi-phase CT slices with about 79 percent average accuracy, outperforming CLIP and MedCL...

  53. MemeBLIP2: A novel lightweight multimodal system to detect harmful memes

    cs.CV 2025-04 conditional novelty 4.0 of 10

    MemeBLIP2, built on BLIP-2 with linear projections, adapters, and an MLP classifier, reaches 77.5% accuracy and 79.0% F1 on PrideMM harmful meme detection, but the paper's internal inconsistencies weaken the claim.

  54. Visual Language Models as Operator Agents in the Space Domain

    cs.AI 2025-01 conditional novelty 4.0 of 10

    Vision-language models can act as spacecraft operators in the KSPDG simulator from screenshots, and fine-tuning OpenVLA on ten episodes shows preliminary promise for robotic satellite inspection.

  55. Multilevel Semantic-Aware Model for AI-Generated Video Quality Assessment

    cs.CV 2025-01 conditional novelty 4.0 of 10

    MSA-VQA combines CLIP-based prompt checking with cross-attention over frames to predict human quality scores for AI-generated videos, reporting state-of-the-art numbers on the T2VQA-DB benchmark.

  56. ErgoChat: a Visual Query System for the Ergonomic Risk Assessment of Construction Workers

    cs.CV 2024-12 conditional novelty 4.0 of 10

    Fine-tuning MiniGPT-v2 on a new 1,900-image construction ergonomics dataset improves visual question answering and image captioning of postural risk versus the same model without fine-tuning.

  57. SubstationAI: Multimodal Large Model-Based Approaches for Analyzing Substation Equipment Faults

    cs.AI 2024-12 reject novelty 4.0 of 10

    SubstationAI, a fine-tuned LLaVA-1.5-7B model augmented with a fault knowledge base, receives higher expert ratings than GPT-4 for substation fault reports, but suspected train/test overlap makes the result unreliable.

  58. Health AI Developer Foundations

    cs.LG 2024-11 conditional novelty 4.0 of 10

    Health AI Developer Foundations packages six domain-specific medical embedding models into one platform, claiming large data and compute savings for downstream health ML tasks.

  59. Investigating the Invertibility of Multimodal Latent Spaces: Limitations of Optimization-Based Methods

    cs.LG 2025-07 reject novelty 3.0 of 10

    Optimization can force BLIP, Flux, Whisper, and Chatterbox to hit textual targets, but the inverted inputs are perceptually incoherent and the recovered text embeddings are semantically meaningless.

  60. E.A.R.T.H.: Structuring Creative Evolution through Model Error in Generative AI

    cs.AI 2025-07 reject novelty 3.0 of 10

    A five-stage pipeline that induces, scores, rewrites, and validates model errors reports large creativity gains that largely arise from selection on the measurement metric itself.

See all 63 Pith citations

Pith tools