Pith. sign in

REVIEW 1 cited by

LXMERT Model Compression for Visual Question Answering

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2310.15325 v1 pith:OF2HKIVZ submitted 2023-10-23 cs.CV cs.CLcs.LG

classification cs.CVcs.CLcs.LG
keywords lxmertaccuracylossmodelmodelssizesubnetworksaccording
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Large-scale pretrained models such as LXMERT are becoming popular for learning cross-modal representations on text-image pairs for vision-language tasks. According to the lottery ticket hypothesis, NLP and computer vision models contain smaller subnetworks capable of being trained in isolation to full performance. In this paper, we combine these observations to evaluate whether such trainable subnetworks exist in LXMERT when fine-tuned on the VQA task. In addition, we perform a model size cost-benefit analysis by investigating how much pruning can be done without significant loss in accuracy. Our experiment results demonstrate that LXMERT can be effectively pruned by 40%-60% in size with 3% loss in accuracy.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Can Argus Judge Them All? Comparing VLMs Across Domains

    cs.IR 2025-06 reject novelty 2.0 of 10

    A VLM benchmark paper that proposes a cross-dataset consistency metric, but the abstract and body evaluate different model sets and the metric's bounds are incorrect.

Pith tools