Pith. sign in

REVIEW 2 cited by

Distilled Dual-Encoder Model for Vision-Language Understanding

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2112.08723 v2 pith:SR2SLFWW submitted 2021-12-16 cs.CL cs.CV

classification cs.CLcs.CV
keywords dual-encodermodelmodelsvisualattentioncross-modaldistillationfusion-encoder
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We propose a cross-modal attention distillation framework to train a dual-encoder model for vision-language understanding tasks, such as visual reasoning and visual question answering. Dual-encoder models have a faster inference speed than fusion-encoder models and enable the pre-computation of images and text during inference. However, the shallow interaction module used in dual-encoder models is insufficient to handle complex vision-language understanding tasks. In order to learn deep interactions of images and text, we introduce cross-modal attention distillation, which uses the image-to-text and text-to-image attention distributions of a fusion-encoder model to guide the training of our dual-encoder model. In addition, we show that applying the cross-modal attention distillation for both pre-training and fine-tuning stages achieves further improvements. Experimental results demonstrate that the distilled dual-encoder model achieves competitive performance for visual reasoning, visual entailment and visual question answering tasks while enjoying a much faster inference speed than fusion-encoder models. Our code and models will be publicly available at https://github.com/kugwzk/Distilled-DualEncoder.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. EffiVLM-BENCH: A Comprehensive Benchmark for Evaluating Training-Free Acceleration in Large Vision-Language Models

    cs.CL 2025-05 conditional novelty 6.0 of 10

    EffiVLM-Bench is a benchmark study showing token compression is task- and model-dependent, KV cache methods are more loyal, and parameter compression preserves accuracy better at typical ratios.

  2. HIT Model: A Hierarchical Interaction-Enhanced Two-Tower Model for Pre-Ranking Systems

    cs.IR 2025-05 conditional novelty 4.0 of 10

    HIT augments the two-tower pre-ranking architecture with class-conditional generators and multi-head MaxSim matching, reporting consistent AUC gains on three public datasets and +1.66% GMV in a Tencent online A/B test.

Pith tools