Pith. sign in

REVIEW 34 cited by

ClipCap: CLIP Prefix for Image Captioning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2111.09734 v1 pith:WE3DIJH5 submitted 2021-11-18 cs.CV

classification cs.CV
keywords modelclipimagecaptioncaptioningcaptionslanguageprefix
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Image captioning is a fundamental task in vision-language understanding, where the model predicts a textual informative caption to a given input image. In this paper, we present a simple approach to address this task. We use CLIP encoding as a prefix to the caption, by employing a simple mapping network, and then fine-tunes a language model to generate the image captions. The recently proposed CLIP model contains rich semantic features which were trained with textual context, making it best for vision-language perception. Our key idea is that together with a pre-trained language model (GPT2), we obtain a wide understanding of both visual and textual data. Hence, our approach only requires rather quick training to produce a competent captioning model. Without additional annotations or pre-training, it efficiently generates meaningful captions for large-scale and diverse datasets. Surprisingly, our method works well even when only the mapping network is trained, while both CLIP and the language model remain frozen, allowing a lighter architecture with less trainable parameters. Through quantitative evaluation, we demonstrate our model achieves comparable results to state-of-the-art methods on the challenging Conceptual Captions and nocaps datasets, while it is simpler, faster, and lighter. Our code is available in https://github.com/rmokady/CLIP_prefix_caption.

Discussion (0). Sign in to comment.

Forward citations

Cited by 34 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. BadBone: Backdoor Attacks Against Backbone Models in Visual Prompt Learning

    cs.CR 2026-05 unverdicted novelty 7.0 of 10

    BadBone backdoors backbone models with bi-level optimization to make prompt learning on downstream tasks vulnerable while preserving model utility.

  2. CB-SLICE: Concept-Based Interpretable Error Slice Discovery

    cs.LG 2026-05 unverdicted novelty 7.0 of 10

    CB-SLICE uses concept mispredictions from CBMs to discover and explain error slices, claiming better performance than existing methods on benchmarks for bias detection.

  3. Semantic Manipulation Localization

    cs.CV 2026-04 unverdicted novelty 7.0 of 10

    Defines SML task for localizing semantic edits and proposes TRACE framework with semantic anchoring, perturbation sensing, and constrained reasoning that outperforms prior IML methods on a custom benchmark.

  4. UIPress: Bringing Optical Token Compression to UI-to-Code Generation

    cs.CL 2026-04 unverdicted novelty 7.0 of 10

    UIPress is the first encoder-side learned optical compression method for UI-to-Code that compresses visual tokens to 256, outperforming the uncompressed baseline by 7.5% CLIP score and the best inference-time baseline...

  5. Sample-efficient Integration of New Modalities into Large Language Models

    cs.CL 2025-09 conditional novelty 7.0 of 10

    A hypernetwork trained on image, audio, and video adapts a shared projector to new, low-resource modalities from as few as 32 examples.

  6. EvoPrompt: Connecting LLMs with Evolutionary Algorithms Yields Powerful Prompt Optimizers

    cs.CL 2023-09 unverdicted novelty 7.0 of 10

    EvoPrompt uses LLMs to run evolutionary operators on populations of prompts, outperforming human-engineered prompts by up to 25% on BIG-Bench Hard tasks across 31 datasets.

  7. LLaMA-Adapter: Efficient Fine-tuning of Language Models with Zero-init Attention

    cs.CV 2023-03 conditional novelty 7.0 of 10

    LLaMA-Adapter turns frozen LLaMA 7B into a capable instruction follower using only 1.2M new parameters and zero-init attention, matching Alpaca while extending to image-conditioned reasoning on ScienceQA and COCO.

  8. LAION-5B: An open large-scale dataset for training next generation image-text models

    cs.CV 2022-10 accept novelty 7.0 of 10

    LAION-5B is an openly released dataset of 5.85 billion CLIP-filtered image-text pairs that enables replication of foundational vision-language models.

  9. Flamingo: a Visual Language Model for Few-Shot Learning

    cs.CV 2022-04 unverdicted novelty 7.0 of 10

    Flamingo models reach new state-of-the-art few-shot results on image and video tasks by bridging frozen vision and language models with cross-attention layers trained on interleaved web-scale data.

  10. Socratic Models: Composing Zero-Shot Multimodal Reasoning with Language

    cs.CV 2022-04 unverdicted novelty 7.0 of 10

    Socratic Models compose zero-shot multimodal reasoning by prompting pretrained language and vision models to exchange information and enable new capabilities without finetuning.

  11. Adjudicated Captioning: Multi-Agent Alignment Scoring and Consensus-Distilled Beam Arbitration for Strict Zero-Shot Image Captioning

    cs.CV 2026-07 conditional novelty 6.0 of 10

    A multi-checkpoint alignment pipeline lifts zero-shot COCO captioning CIDEr from 108.0 to 117.6 by adding a cross-attention verifier and self-supervised beam rerankers to an unchanged IFCap captioner.

  12. REPREC: Representation Driven Parameter-Efficient Recommendation System

    cs.IR 2026-07 conditional novelty 6.0 of 10

    A tiny MLP injector maps a frozen recommender's user embedding into soft tokens, letting a frozen LLM match LoRA-based recommenders at lower training cost.

  13. READ More than What You See: Reinforcement Learning for Accurate and Coherent Audio Description Generations

    cs.CV 2026-06 unverdicted novelty 6.0 of 10

    READ is the first reinforcement-learning framework for training audio-description generators, using sequence-level rewards for reference match, length, format, and context-aware coherence.

  14. Revisiting LLM Adaptation for 3D CT Report Generation: A Study of Scaling and Diagnostic Priors

    cs.CL 2026-06 unverdicted novelty 6.0 of 10

    RAD3D-Prefix is a diagnostic-prior conditioning framework for 3D CT report generation that integrates image embeddings with multi-label classification logits, showing that freezing larger LLMs and training only projec...

  15. Learning to See What You Need: Gaze Attention for Multimodal Large Language Models

    cs.CV 2026-05 unverdicted novelty 6.0 of 10

    Gaze Attention groups visual embeddings into selectable regions and dynamically restricts attention to task-relevant ones, matching dense baselines with up to 90% fewer visual KV entries via added context tokens.

  16. Adversarial Attacks Against MLLMs via Progressive Resolution Processing and Adaptive Feature Alignment

    cs.CV 2026-05 unverdicted novelty 6.0 of 10

    PRAF-Attack improves targeted attack transferability on black-box MLLMs by using multi-scale progressive resolution and adaptive intermediate feature alignment instead of final-layer global features.

  17. Beyond Semantics: Uncovering the Physics of Fakes via Universal Physical Descriptors for Cross-Modal Synthetic Detection

    cs.CV 2026-04 unverdicted novelty 6.0 of 10

    Five universal physical descriptors including Laplacian variance, Sobel statistics, and residual noise variance, when integrated as text encodings with CLIP, achieve up to 99.8% accuracy detecting synthetic images acr...

  18. LatentLens: Revealing Highly Interpretable Visual Tokens in LLMs

    cs.CV 2026-01 conditional novelty 6.0 of 10

    Visual tokens in VLMs look highly interpretable at every layer when decoded by nearest-neighbor retrieval against a corpus of contextualized text embeddings: 72% interpretable on average vs 23–30% for LogitLens and Em...

  19. Unpacking Hateful Memes: Presupposed Context and False Claims

    cs.CL 2025-10 conditional novelty 6.0 of 10

    A hateful-meme detector that combines presupposed-context fusion, LLM-based social perception, and cross-modal reference graphs outperforms prior models on three benchmarks and transfers to fake news.

  20. Sparse and Dense Retrievers Learn Better Together: Joint Sparse-Dense Optimization for Text-Image Retrieval

    cs.CL 2025-08 conditional novelty 6.0 of 10

    A shared integrated teacher improves both sparse and dense text-image retrieval, letting a sparse retriever match or beat dense baselines on MSCOCO and several Flickr30k settings.

  21. RAVID: Retrieval-Augmented Visual Detection: A Knowledge-Driven Approach for AI-Generated Image Identification

    cs.CV 2025-08 conditional novelty 6.0 of 10

    RAVID detects AI-generated images by retrieving similar images from a database and feeding them to a vision-language model, reporting 93.85% average accuracy on UniversalFakeDetect.

  22. Mitigating Information Loss under High Pruning Rates for Efficient Large Vision Language Models

    cs.CV 2025-08 conditional novelty 6.0 of 10

    ACCM recovers information lost in high-rate visual token pruning by generating a question-guided caption from discarded tokens and selecting the best candidate, improving pruned LVLM accuracy with fewer FLOPs.

  23. VL-CLIP: Enhancing Multimodal Recommendations via Visual Grounding and LLM-Augmented CLIP Embeddings

    cs.IR 2025-07 conditional novelty 6.0 of 10

    A production system that combines object-detection-based image cropping and LLM-based text rewriting with CLIP fine-tuning reports large gains in multimodal retrieval and online recommendation metrics at Walmart scale.

  24. Imperceptible and Reversible Adversarial Examples against Vision-Language Models for Privacy Protection

    cs.CV 2026-07 conditional novelty 5.5 of 10

    CloakDiff generates high-fidelity reversible adversarial images that suppress VLM text-query privacy leakage via diffusion attention editing plus invertible steganography.

  25. Cross-Modal Masked Compositional Concept Modeling for Enhancing Visio-Linguistic Compositionality

    cs.CV 2026-06 unverdicted novelty 5.0 of 10

    MACCO applies cross-modal masked reconstruction of compositional concepts with inter- and intra-modal auxiliary objectives to improve visio-linguistic compositionality in VLMs.

  26. Task-Aware Structured Memory for Dynamic Multi-modal In-Context Learning

    cs.CV 2026-06 unverdicted novelty 5.0 of 10

    TASM proposes a task-aware structured memory framework using task-vector compression, bipartite token merging, and a Core Memory plus Latent Bank hierarchy to enable efficient dynamic multi-modal in-context learning.

  27. Modeling Complex Behaviors: Multi-Personality Composition and Dynamic Switching in Vision-Language Models

    cs.CL 2026-06 unverdicted novelty 5.0 of 10

    The work establishes an evaluation framework for personality induction and switching in MLLMs, reporting improved captioning but impaired VQA performance plus balancing and residual effects during multi-trait and dyna...

  28. WRF4CIR: Weight-Regularized Fine-Tuning Network for Composed Image Retrieval

    cs.CV 2026-04 unverdicted novelty 5.0 of 10

    WRF4CIR uses weight-regularized fine-tuning with adversarial perturbations to mitigate overfitting in composed image retrieval and narrows the generalization gap on benchmarks.

  29. MIND: Multi-rationale INtegrated Discriminative Reasoning Framework for Multi-modal Large Models

    cs.AI 2025-12 conditional novelty 5.0 of 10

    MIND improves multimodal reasoning by training on diverse correct and deliberately wrong rationales with two-stage correction and contrastive alignment, reporting SOTA on ScienceQA, A-OKVQA, and M3CoT.

  30. Compression Beyond Pixels: Semantic Compression with Multimodal Foundation Models

    cs.CV 2025-09 conditional novelty 5.0 of 10

    A learned product-quantization VAE with a shared codebook compresses CLIP image features to about 2-3 x 10^-3 bits per pixel with little loss in downstream semantic task accuracy.

  31. LLaMA-Adapter V2: Parameter-Efficient Visual Instruction Model

    cs.CV 2023-04 conditional novelty 5.0 of 10

    LLaMA-Adapter V2 achieves open-ended visual instruction following in LLMs by unlocking more parameters, early fusion of visual tokens, and joint training on disjoint parameter groups with only 14M added parameters.

  32. Beyond Standard Benchmarks: A Systematic Audit of Vision-Language Model's Robustness to Natural Semantic Variation Across Diverse Tasks

    cs.CV 2026-04 unverdicted novelty 4.0 of 10

    Robust CLIP models amplify vulnerabilities to natural adversarial scenarios while standard CLIP shows large performance drops on natural language-induced adversarial examples in zero-shot classification, segmentation,...

  33. From Image Captioning to Visual Storytelling

    cs.CL 2025-07 unverdicted novelty 4.0 of 10

    Visual storytelling improves by treating it as image captioning followed by language-to-language story generation, with a new 'ideality' metric to gauge distance from an oracle.

  34. Multi-Modal LLM based Image Captioning in ICT: Bridging the Gap Between General and Industry Domain

    cs.CV 2026-01 unverdicted novelty 3.0 of 10

    A 7B-parameter domain-specific image captioning model for ICT, trained in three stages on synthesized and annotated data, outperforms 32B-parameter general models on BLEU and expert accuracy metrics.

Pith tools