REVIEW 2 major objections 4 minor 87 references
FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression
T0 review · 2 major / 4 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read The paper claims FlashSloth, a 3.2B multimodal model, can cut visual tokens by 80–89%, training memory by 61–80%, and inference computation by 70–98% while shortening response time by 2–5x and staying competitive with advanced tiny MLLMs.
desk verdict Plausible compression design with systematic ablations, but the TFLOPs accounting in Table 1 is inconsistent and the headline efficiency claims need verification. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is a two-part compression pipeline embedded inside the MLLM. Spatial Attention Pooling (SAP) divides the visual token grid into $s \times s$ regions, predicts a softmax weight $\alpha_i = \mathrm{Softmax}(\mathrm{mlp}(F^i_v))$ for each token in a region, and produces one salient token per region via the weighted sum $f^s_v = \sum_i \alpha_i f^i_v$, cutting 729 tokens to 81. The Embedded Query Module (EmbQ) pads 9 learnable query tokens into the LLM input; after a few transformer layers, those queries first cross-attend to text tokens to become instruction-conditioned, then cross-attend to the uncompressed visual tokens, and the resulting features are added back into the query token stream. This lets the model recover task-relevant image details that saliency-only pooling loses, without introducing a separate language model or a dedicated vision-language alignment stage.
What would settle it
Run FlashSloth, FlashSloth-HD, InternVL2-2B, and Qwen2-VL-2B on the same GPU with matched batch size, input resolution, precision, prompts, and maximum generated tokens; if the time-to-first-token gap against InternVL2 does not remain near 2–5x, the architecture-specific efficiency claim fails.
Extended reading notes
Core claim
The central discovery claim is that attention-based pooling combined with an embedded instruction-aware query module can replace a large visual-token sequence almost losslessly. On the paper's terms, visually salient semantics and instruction-related semantics are complementary: SAP captures what stands out in each image region, EmbQ captures what the question cares about, and together they let a 3.2B model use only 90 visual tokens while matching or exceeding advanced tiny MLLMs on many benchmarks. The paper calls this design embedded visual compression and contrasts it with Q-Former-style bridges that require another language model and dedicated alignment pretraining. With the high-resolution FlashSloth-HD variant, the same compression also narrows the gap on OCR-heavy tasks such as DocVQA and ChartQA.
Load-bearing premise
The headline efficiency gains rest on the assumption that every model in Table 1 was measured under identical inference conditions—same hardware, batch size, precision, and decoding length—which the paper does not report.
Editorial extensions
If this is right
- A 3.2B model with 90 visual tokens can match or exceed advanced tiny MLLMs on MMB, GQA, SQA, and AI2D while using far fewer tokens.
- Average time-to-first-token drops to 0.05 seconds, enabling interactive and mobile deployments where response latency is the main constraint.
- Training with the LLaVA-665k split costs only 6.4 GPU-hours for pretraining, substantially lowering the barrier to custom tiny MLLMs.
- FlashSloth-HD recovers most of the OCR and document-understanding gap while still using fewer visual tokens than Qwen2-VL and InternVL2.
- EmbQ is reported as a reusable module: adding it to average pooling, pixel shuffle, or LDP compression consistently improves their benchmark scores.
Reading between the lines
- The reported speedups likely act mainly on prefill and first-token latency; benchmarks with long generated answers would probably show a smaller relative response-time gap, since token-count savings matter less once decoding is dominated by autoregressive generation.
- Combining architectural compression like EmbQ with runtime token pruning could compound the savings, because the paper's compression works at the network-structure level while pruning methods are orthogonal inference-time additions.
- The saliency weights produced by SAP could double as cheap interpretability maps for debugging hallucinations or grounding failures, a use the paper does not explore.
- The near-optimality of 9 query tokens suggests coarse instruction grounding, not fine alignment, is what these benchmarks reward; tasks requiring dense spatial detail may need more queries or a resolution-aware query allocation.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. FlashSloth introduces a 3.2B multimodal LLM that compresses visual tokens with a spatial attention pooling (SAP) module and an embedded query (EmbQ) module. SAP reduces the token count by attention-weighted pooling over spatial regions, while EmbQ uses 9 learnable query tokens to cross-attend to the uncompressed visual tokens at an interior LLM layer, supplying instruction-related visual information. The default model uses Phi-2 2.7B and SigLIP, with 81 visual tokens plus 9 query tokens; a high-resolution variant, FlashSloth-HD, uses 414 tokens. The paper reports accuracy on 14 benchmarks and efficiency comparisons (visual token count, TFLOPs, GPU memory, response time, throughput) against several tiny MLLMs. It claims reductions of 80–89% in visual tokens, 61–80% in training memory, and 70–98% in inference computation, with roughly 2×–5× faster response time while remaining competitive. The central efficiency claim rests on Table 1, and the review focuses on whether the numbers in that table support it.
Significance. The architecture is well motivated and clearly described, and the ablation study in Tables 3–6 is systematic, including comparisons with average pooling, pixel shuffle, and LDP-based compression. The training budget is small (3.7M samples), the code is released, and the accuracy results suggest that a 3.2B MLLM with 90 visual tokens can be competitive with 2–3B MLLMs on many benchmarks. If the efficiency numbers were verified under controlled conditions, the contribution would be meaningful: it would demonstrate that embedded visual compression can preserve accuracy while drastically shortening the visual token sequence. However, the paper's primary claim is about efficiency, and that claim currently rests on Table 1, whose TFLOPs are not reproducible under standard prefill accounting and whose measurement conditions are not stated. The accuracy side is generally sound, but the headline efficiency contribution needs a major correction or clarification.
major comments (2)
- [Sec. 4.3.1, Table 1 and Table 3] The reported first-round TFLOPs for FlashSloth are not reproducible under standard prefill accounting. For a 3.2B-parameter model (Phi-2 ~2.7B plus SigLIP ~0.4B) and a sequence of roughly 110–150 tokens, the usual 2·N·T estimate gives about 0.6 TFLOPs for the LLM alone and about 1.2 TFLOPs if the vision encoder is included; the same estimate applied to the Qwen2-VL rows is consistent with the table, but the FlashSloth rows (0.08–0.10 TFLOPs) are about an order of magnitude too low. FlashSloth-HD reports 2.71 TFLOPs for 414 tokens, so the 30-fold drop in TFLOPs for a 4.6-fold token reduction is internally inconsistent. Table 3's 729-token baseline at 0.30 TFLOPs is similarly below the ~4 TFLOPs expected for that sequence length. Because the abstract and Sec. 4.3.1 use these numbers for the headline '70–98% inference computation reduction', the central efficiency claim is not established. Please state the exact FLOPs counting formula and which components (vision encoder, attention, prefill vs. decode) are included, and re-derive all affected percentages.
- [Sec. 4.3.1, Table 1] The response-time, throughput, and GPU-memory comparisons omit the measurement conditions: GPU model, batch size, numerical precision, maximum generated token length, decoding configuration, and whether numbers are medians over repeated runs. These settings can dominate the reported 2–5× speedups, especially because the benefit of token reduction is concentrated in prefill. Please report a controlled comparison with identical settings across all rows, and separate time-to-first-token from decode throughput.
minor comments (4)
- [Sec. 4.3.2, Table 2] The sentence 'FlashSloth can even achieve new SOTA performance among tiny MLLMs on several benchmarks, such as MMB, GQA and AI2D' is not supported by the table: Qwen2-VL-2B scores 74.9 on MMB, InternVL2 scores 61.6 on GQA and 74.1 on AI2D, all above FlashSloth's 73.0, 61.1, and 72.5.
- [Tables 2–6] Benchmark scores are reported as single point estimates without confidence intervals or repeated-evaluation statistics; given the small margins in some comparisons, please include variance information or state that the differences are within evaluation noise.
- [Sec. 3.1, Fig. 2] The text says the framework is illustrated in Fig. 1, but Fig. 1 is the motivation/efficiency plot and the framework is in Fig. 2; the cross-reference should be corrected.
- [Abstract and Sec. 4.3.1] The claimed reduction ranges (80–89% visual tokens, 61–80% training memory, 70–98% inference computation) do not match the per-model percentages in Table 1; for example, 90 tokens versus InternVL2's 1561 tokens is a 94% token reduction, which is outside the stated range. Please specify the reference model for each range.
Circularity Check
No circularity: FlashSloth's efficiency and accuracy claims are empirical evaluations against external benchmarks, with architecture choices determined by ablations before final comparisons.
full rationale
The paper's central claims are (1) a large reduction in visual token count, (2) reductions in training memory and inference computation, and (3) competitive benchmark accuracy. None of these is derived from its own inputs. The proposed modules are concrete and explicit: SAP is defined by Eq. (3)-(4) as a softmax-weighted pooling over spatial regions, and EmbQ is defined by Eq. (5)-(6) as cross-attention over text and image features. The final model is trained on the LLaVA data split and evaluated on external benchmarks (MMB, GQA, MME, POPE, TextVQA, etc.) via lmms-eval. Hyperparameters such as SAP downsampling rate, number of query tokens, EmbQ dimension, and insertion layer are selected by ablations reported in Tables 3-6 before the final evaluation, not fitted to the benchmark numbers used as evidence. Self-citations to the authors' prior work, e.g., TRAR-style attention [73], gated fusion [45], and FitPrune [62], are either motivational or used as ablated alternatives; the ablation tables directly compare these choices, so the citations are not load-bearing. The only notable concern is the internal plausibility of Table 1's TFLOPs for FlashSloth (0.08-0.10 TFLOPs with 90 tokens), which appears inconsistent with a standard prefill-cost estimate and with the FlashSloth-HD rows; however, this is a measurement-reporting or correctness issue, not a circular derivation. No equation reduces to a fitted quantity, no benchmark result is used to set constants that later define the result, and no prediction is forced by construction. Therefore the paper has no significant circularity.
Assumptions & free parameters
free parameters (4)
- SAP downsampling rate s =
3 (81 visual tokens)
- EmbQ query token count n =
9
- EmbQ embedding dimension =
576
- EmbQ insertion layer =
8
assumptions (3)
- domain assumption Visual tokens from a pretrained ViT encoder carry redundant overlapping semantics, so spatial pooling at s=3 loses little task-critical information.
- domain assumption LLaVA-665k pretraining plus a short SFT stage is a sufficient training budget for a competitive tiny MLLM.
- domain assumption lmms-eval benchmark scores are stable enough to compare models without error bars or repeated runs.
Cite this review
Pith. "Pith review of FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression." pith.science (2026). https://pith.science/paper/CQTKGGO4
@misc{pith2026241204317,
author = {Pith},
title = {Pith review of: FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression},
year = {2026},
howpublished = {\url{https://pith.science/paper/CQTKGGO4}},
note = {Machine review of arXiv:2412.04317}
}
read the original abstract
Despite a big leap forward in capability, multimodal large language models (MLLMs) tend to behave like a sloth in practical use, i.e., slow response and large latency. Recent efforts are devoted to building tiny MLLMs for better efficiency, but the plethora of visual tokens still used limit their actual speedup. In this paper, we propose a powerful and fast tiny MLLM called FlashSloth. Different from previous efforts, FlashSloth focuses on improving the descriptive power of visual tokens in the process of compressing their redundant semantics. In particular, FlashSloth introduces embedded visual compression designs to capture both visually salient and instruction-related image information, so as to achieving superior multimodal performance with fewer visual tokens. Extensive experiments are conducted to validate the proposed FlashSloth, and a bunch of tiny but strong MLLMs are also comprehensively compared, e.g., InternVL2, MiniCPM-V2 and Qwen2-VL. The experimental results show that compared with these advanced tiny MLLMs, our FlashSloth can greatly reduce the number of visual tokens, training memory and computation complexity while retaining high performance on various VL tasks.
Figures
Figures from the paper (6 more)
Reference graph
Works this paper leans on
-
[1]
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ah- mad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al. Gpt-4 technical report. arXiv preprint arXiv:2303.08774 ,
-
[2]
Bottom-up and top-down attention for image captioning and visual question answering
Peter Anderson, Xiaodong He, Chris Buehler, Damien Teney, Mark Johnson, Stephen Gould, and Lei Zhang. Bottom-up and top-down attention for image captioning and visual question answering. In Proceedings of the IEEE con- ference on computer vision and pattern recognition , pages 6077–6086, 2018. 2
2018
-
[3]
Jinze Bai, Shuai Bai, Yunfei Chu, Zeyu Cui, Kai Dang, Xiaodong Deng, Yang Fan, Wenbin Ge, Yu Han, Fei Huang, et al. Qwen technical report. arXiv preprint arXiv:2309.16609, 2023. 2
arXiv 2023
-
[4]
Gemma: Introducing new state-of-the-art open models
Jeanine Banks and Tris Warkentin. Gemma: Introducing new state-of-the-art open models. Google. Available online at: https://blog. google/technology/developers/gemma-open- models/(accessed 6 April, 2024), 2024. 2
2024
-
[5]
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Sub- biah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakan- tan, Pranav Shyam, Girish Sastry, Amanda Askell, Sand- hini Agarwal, Ariel Herbert-V oss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, S...
1901
-
[6]
Honeybee: Locality-enhanced projector for multimodal llm
Junbum Cha, Wooyoung Kang, Jonghwan Mun, and Byungseok Roh. Honeybee: Locality-enhanced projector for multimodal llm. In Proceedings of the IEEE/CVF Con- ference on Computer Vision and Pattern Recognition, pages 13817–13827, 2024. 3
2024
-
[7]
Liang Chen, Haozhe Zhao, Tianyu Liu, Shuai Bai, Jun- yang Lin, Chang Zhou, and Baobao Chang. An image is worth 1/2 tokens after layer 2: Plug-and-play inference ac- celeration for large vision-language models. arXiv preprint arXiv:2403.06764, 2024. 2, 3
arXiv 2024
-
[8]
Lawrence Zit- nick
Xinlei Chen, Hao Fang, Tsung-Yi Lin, Ramakrishna Vedan- tam, Saurabh Gupta, Piotr Dollar, and C. Lawrence Zit- nick. Microsoft coco captions: Data collection and evalu- ation server, 2015. 4
2015
Show all 87 references
-
[9]
Pali: A jointly- scaled multilingual language-image model
Xi Chen, Xiao Wang, Soravit Changpinyo, AJ Piergiovanni, Piotr Padlewski, Daniel Salz, Sebastian Goodman, Adam Grycner, Basil Mustafa, Lucas Beyer, et al. Pali: A jointly- scaled multilingual language-image model. arXiv preprint arXiv:2209.06794, 2022. 3
2022 arXiv
-
[10]
How far are we to gpt-4v? closing the gap to commercial multimodal models with open- source suites, 2024
Zhe Chen, Weiyun Wang, Hao Tian, Shenglong Ye, Zhang- wei Gao, Erfei Cui, Wenwen Tong, Kongzhi Hu, Jiapeng Luo, Zheng Ma, Ji Ma, Jiaqi Wang, Xiaoyi Dong, Hang Yan, Hewei Guo, Conghui He, Botian Shi, Zhenjiang Jin, Chao Xu, Bin Wang, Xingjian Wei, Wei Li, Wenjian Zhang, Bo Zhan...
2024
-
[11]
Mobilevlm: A fast, strong and open vi- sion language assistant for mobile devices
Xiangxiang Chu, Limeng Qiao, Xinyang Lin, Shuang Xu, Yang Yang, Yiming Hu, Fei Wei, Xinyu Zhang, Bo Zhang, Xiaolin Wei, et al. Mobilevlm: A fast, strong and open vi- sion language assistant for mobile devices. arXiv preprint arXiv:2312.16886, 2023. 2, 3
2023 arXiv
-
[12]
Mobilevlm v2: Faster and stronger baseline for vision language model
Xiangxiang Chu, Limeng Qiao, Xinyu Zhang, Shuang Xu, Fei Wei, Yang Yang, Xiaofei Sun, Yiming Hu, Xinyang Lin, Bo Zhang, et al. Mobilevlm v2: Faster and stronger baseline for vision language model. arXiv preprint arXiv:2402.03766, 2024. 1, 2, 3, 7, 13
2024 arXiv
-
[13]
Instructblip: Towards general- purpose vision-language models with instruction tuning,
Wenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong, Junqi Zhao, Weisheng Wang, Boyang Li, Pascale Fung, and Steven Hoi. Instructblip: Towards general- purpose vision-language models with instruction tuning,
-
[14]
Mme: A compre- 9 hensive evaluation benchmark for multimodal large language models, 2024
Chaoyou Fu, Peixian Chen, Yunhang Shen, Yulei Qin, Mengdan Zhang, Xu Lin, Jinrui Yang, Xiawu Zheng, Ke Li, Xing Sun, Yunsheng Wu, and Rongrong Ji. Mme: A compre- 9 hensive evaluation benchmark for multimodal large language models, 2024. 2, 5
2024
-
[15]
Making the v in vqa matter: Elevating the role of image understanding in visual question answer- ing
Yash Goyal, Tejas Khot, Douglas Summers-Stay, Dhruv Ba- tra, and Devi Parikh. Making the v in vqa matter: Elevating the role of image understanding in visual question answer- ing. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 6904–6913...
2017
-
[16]
Minicpm: Un- veiling the potential of small language models with scalable training strategies, 2024
Shengding Hu, Yuge Tu, Xu Han, Chaoqun He, Ganqu Cui, Xiang Long, Zhi Zheng, Yewei Fang, Yuxiang Huang, Weilin Zhao, Xinrong Zhang, Zheng Leng Thai, Kaihuo Zhang, Chongyi Wang, Yuan Yao, Chenyang Zhao, Jie Zhou, Jie Cai, Zhongwu Zhai, Ning Ding, Chao Jia, Guoyang Zeng, Dahai L...
2024
-
[17]
Minicpm: Unveiling the potential of small language models with scalable training strategies
Shengding Hu, Yuge Tu, Xu Han, Chaoqun He, Ganqu Cui, Xiang Long, Zhi Zheng, Yewei Fang, Yuxiang Huang, Weilin Zhao, et al. Minicpm: Unveiling the potential of small language models with scalable training strategies. arXiv preprint arXiv:2404.06395, 2024. 1, 2, 3, 4, 7
2024 arXiv
-
[18]
Token merging for training- free semantic binding in text-to-image synthesis, 2024
Taihang Hu, Linxuan Li, Joost van de Weijer, Hongcheng Gao, Fahad Shahbaz Khan, Jian Yang, Ming-Ming Cheng, Kai Wang, and Yaxing Wang. Token merging for training- free semantic binding in text-to-image synthesis, 2024. 3
2024
-
[19]
Gqa: A new dataset for real-world visual reasoning and compositional question answering
Drew A Hudson and Christopher D Manning. Gqa: A new dataset for real-world visual reasoning and compositional question answering. In Proceedings of the IEEE/CVF con- ference on computer vision and pattern recognition , pages 6700–6709, 2019. 5, 7
2019
-
[20]
Phi-2: The surprising power of small language models
Mojan Javaheripi, S ´ebastien Bubeck, Marah Abdin, Jy- oti Aneja, Sebastien Bubeck, Caio C ´esar Teodoro Mendes, Weizhu Chen, Allie Del Giorno, Ronen Eldan, Sivakanth Gopi, et al. Phi-2: The surprising power of small language models. Microsoft Research Blog, 2023. 1, 2, 5
2023
-
[21]
In defense of grid features for visual question answering
Huaizu Jiang, Ishan Misra, Marcus Rohrbach, Erik Learned- Miller, and Xinlei Chen. In defense of grid features for visual question answering. In Proceedings of the IEEE/CVF con- ference on computer vision and pattern recognition , pages 10267–10276, 2020. 2
2020
-
[22]
Contrast and classify: Training robust vqa models
Yash Kant, Abhinav Moudgil, Dhruv Batra, Devi Parikh, and Harsh Agrawal. Contrast and classify: Training robust vqa models. In Proceedings of the IEEE/CVF International Con- ference on Computer Vision, pages 1604–1613, 2021. 2
2021
-
[23]
Referitgame: Referring to objects in pho- tographs of natural scenes
Sahar Kazemzadeh, Vicente Ordonez, Mark Matten, and Tamara Berg. Referitgame: Referring to objects in pho- tographs of natural scenes. In Proceedings of the 2014 con- ference on empirical methods in natural language processing (EMNLP), pages 787–798, 2014. 4
2014
-
[24]
A diagram is worth a dozen images
Aniruddha Kembhavi, Mike Salvato, Eric Kolve, Minjoon Seo, Hannaneh Hajishirzi, and Ali Farhadi. A diagram is worth a dozen images. In Computer Vision–ECCV 2016: 14th European Conference, Amsterdam, The Netherlands, October 11–14, 2016, Proceedings, Part IV 14 , pages 235–
2016
-
[25]
Seed-bench: Bench- marking multimodal large language models
Bohao Li, Yuying Ge, Yixiao Ge, Guangzhi Wang, Rui Wang, Ruimao Zhang, and Ying Shan. Seed-bench: Bench- marking multimodal large language models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pat- tern Recognition, pages 13299–13308, 2024. 2, 5
2024
-
[26]
Llava-next-interleave: Tackling multi-image, video, and 3d in large multimodal models
Feng Li, Renrui Zhang, Hao Zhang, Yuanhan Zhang, Bo Li, Wei Li, Zejun Ma, and Chunyuan Li. Llava-next-interleave: Tackling multi-image, video, and 3d in large multimodal models. arXiv preprint arXiv:2407.07895, 2024. 5
2024 arXiv
-
[27]
Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation
Junnan Li, Dongxu Li, Caiming Xiong, and Steven Hoi. Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation. In Interna- tional conference on machine learning, pages 12888–12900. PMLR, 2022. 2
2022
-
[28]
Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models
Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi. Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models. In In- ternational conference on machine learning , pages 19730– 19742. PMLR, 2023. 1, 2, 3, 4
2023
-
[29]
Tokenpacker: Effi- cient visual projector for multimodal llm
Wentong Li, Yuqian Yuan, Jian Liu, Dongqi Tang, Song Wang, Jianke Zhu, and Lei Zhang. Tokenpacker: Effi- cient visual projector for multimodal llm. arXiv preprint arXiv:2407.02392, 2024. 3
2024 arXiv
-
[30]
Evaluating object hallucination in large vision-language models
Yifan Li, Yifan Du, Kun Zhou, Jinpeng Wang, Wayne Xin Zhao, and Ji-Rong Wen. Evaluating object hallucination in large vision-language models. In The 2023 Conference on Empirical Methods in Natural Language Processing , 2023. 2, 5
2023
-
[31]
Mini-gemini: Mining the potential of multi-modality vision language models
Yanwei Li, Yuechen Zhang, Chengyao Wang, Zhisheng Zhong, Yixin Chen, Ruihang Chu, Shaoteng Liu, and Jiaya Jia. Mini-gemini: Mining the potential of multi-modality vision language models. arXiv preprint arXiv:2403.18814,
-
[32]
Mon- key: Image resolution and text label are important things for large multi-modal models
Zhang Li, Biao Yang, Qiang Liu, Zhiyin Ma, Shuo Zhang, Jingxu Yang, Yabo Sun, Yuliang Liu, and Xiang Bai. Mon- key: Image resolution and text label are important things for large multi-modal models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recog...
2024
-
[33]
Improved baselines with visual instruction tuning
Haotian Liu, Chunyuan Li, Yuheng Li, and Yong Jae Lee. Improved baselines with visual instruction tuning. In Pro- ceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 26296–26306, 2024. 1, 2, 4, 5, 6, 7
2024
-
[34]
Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024
Haotian Liu, Chunyuan Li, Yuheng Li, Bo Li, Yuanhan Zhang, Sheng Shen, and Yong Jae Lee. Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024. 1, 4, 5, 7
2024
-
[35]
Visual instruction tuning
Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. Visual instruction tuning. Advances in neural information processing systems, 36, 2024. 1, 5
2024
-
[36]
Mmbench: Is your multi-modal model an all-around player?, 2024
Yuan Liu, Haodong Duan, Yuanhan Zhang, Bo Li, Songyang Zhang, Wangbo Zhao, Yike Yuan, Jiaqi Wang, Conghui He, Ziwei Liu, Kai Chen, and Dahua Lin. Mmbench: Is your multi-modal model an all-around player?, 2024. 2, 5, 7
2024
-
[37]
Fixing weight decay regularization in adam
Ilya Loshchilov, Frank Hutter, et al. Fixing weight decay regularization in adam. arXiv preprint arXiv:1711.05101, 5,
-
[38]
Deepseek-vl: towards real-world vision- language understanding
Haoyu Lu, Wen Liu, Bo Zhang, Bingxuan Wang, Kai Dong, Bo Liu, Jingxiang Sun, Tongzheng Ren, Zhuoshu Li, 10 Hao Yang, et al. Deepseek-vl: towards real-world vision- language understanding. arXiv preprint arXiv:2403.05525,
-
[39]
Learn to explain: Multimodal reasoning via thought chains for science question answering
Pan Lu, Swaroop Mishra, Tanglin Xia, Liang Qiu, Kai-Wei Chang, Song-Chun Zhu, Oyvind Tafjord, Peter Clark, and Ashwin Kalyan. Learn to explain: Multimodal reasoning via thought chains for science question answering. Advances in Neural Information Processing Systems , 35:2507–2521,
-
[40]
Mathvista: Evaluating mathemat- ical reasoning of foundation models in visual contexts.arXiv preprint arXiv:2310.02255, 2023
Pan Lu, Hritik Bansal, Tony Xia, Jiacheng Liu, Chunyuan Li, Hannaneh Hajishirzi, Hao Cheng, Kai-Wei Chang, Michel Galley, and Jianfeng Gao. Mathvista: Evaluating mathemat- ical reasoning of foundation models in visual contexts.arXiv preprint arXiv:2310.02255, 2023. 5
-
[41]
To- wards lightweight transformer via group-wise transforma- tion for vision-and-language tasks
Gen Luo, Yiyi Zhou, Xiaoshuai Sun, Yan Wang, Liujuan Cao, Yongjian Wu, Feiyue Huang, and Rongrong Ji. To- wards lightweight transformer via group-wise transforma- tion for vision-and-language tasks. IEEE Transactions on Image Processing, 31:3386–3398, 2022. 2
2022
-
[42]
Cheap and quick: Efficient vision- language instruction tuning for large language models
Gen Luo, Yiyi Zhou, Tianhe Ren, Shengxin Chen, Xiaoshuai Sun, and Rongrong Ji. Cheap and quick: Efficient vision- language instruction tuning for large language models. In Advances in Neural Information Processing Systems , pages 29615–29627. Curran Associates, Inc., 2023. 1, 2
2023
-
[43]
Moil: Momentum imita- tion learning for efficient vision-language adaptation
Gen Luo, Yiyi Zhou, Minglang Huang, Tianhe Ren, Xi- aoshuai Sun, and Rongrong Ji. Moil: Momentum imita- tion learning for efficient vision-language adaptation. IEEE Transactions on Pattern Analysis and Machine Intelligence,
-
[44]
Towards language-guided visual recog- nition via dynamic convolutions
Gen Luo, Yiyi Zhou, Xiaoshuai Sun, Yongjian Wu, Yue Gao, and Rongrong Ji. Towards language-guided visual recog- nition via dynamic convolutions. International Journal of Computer Vision, 132(1):1–19, 2024. 2
2024
-
[45]
Feast your eyes: Mixture-of- resolution adaptation for multimodal large language models
Gen Luo, Yiyi Zhou, Yuxin Zhang, Xiawu Zheng, Xi- aoshuai Sun, and Rongrong Ji. Feast your eyes: Mixture-of- resolution adaptation for multimodal large language models. arXiv preprint arXiv:2403.03003, 2024. 1, 3, 13
2024 arXiv
-
[46]
Chartqa: A benchmark for question answer- ing about charts with visual and logical reasoning, 2022
Ahmed Masry, Do Xuan Long, Jia Qing Tan, Shafiq Joty, and Enamul Hoque. Chartqa: A benchmark for question answer- ing about charts with visual and logical reasoning, 2022. 5, 6
2022
-
[47]
Minesh Mathew, Dimosthenis Karatzas, and C. V . Jawahar. Docvqa: A dataset for vqa on document images, 2021. 5, 6, 7
2021
-
[48]
Dinov2: Learning robust visual features without supervision
Maxime Oquab, Timoth ´ee Darcet, Th ´eo Moutakanni, Huy V o, Marc Szafraniec, Vasil Khalidov, Pierre Fernandez, Daniel Haziza, Francisco Massa, Alaaeldin El-Nouby, et al. Dinov2: Learning robust visual features without supervision. arXiv preprint arXiv:2304.07193, 2023. 2
2023 arXiv
-
[49]
Learning transferable visual models from natural language supervision, 2021
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. Learning transferable visual models from natural language supervision, 2021. 4
2021
-
[50]
Learning transferable visual models from natural language supervi- sion
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning transferable visual models from natural language supervi- sion. In International conference on machine learning, ...
2021
-
[51]
Imp: Highly capable large multimodal models for mobile devices
Zhenwei Shao, Zhou Yu, Jun Yu, Xuecheng Ouyang, Lihao Zheng, Zhenbiao Gai, Mingyang Wang, and Jiajun Ding. Imp: Highly capable large multimodal models for mobile devices. arXiv preprint arXiv:2405.12107, 2024. 1, 2, 3, 4, 5, 6, 7
2024 arXiv
-
[52]
When do we not need larger vision models?, 2024
Baifeng Shi, Ziyang Wu, Maolin Mao, Xin Wang, and Trevor Darrell. When do we not need larger vision models?, 2024. 2
2024
-
[53]
Eagle: Ex- ploring the design space for multimodal llms with mixture of encoders
Min Shi, Fuxiao Liu, Shihao Wang, Shijia Liao, Subhashree Radhakrishnan, De-An Huang, Hongxu Yin, Karan Sapra, Yaser Yacoob, Humphrey Shi, Bryan Catanzaro, Andrew Tao, Jan Kautz, Zhiding Yu, and Guilin Liu. Eagle: Ex- ploring the design space for multimodal llms with mixture o...
2024 arXiv
-
[54]
Towards vqa models that can read
Amanpreet Singh, Vivek Natarajan, Meet Shah, Yu Jiang, Xinlei Chen, Dhruv Batra, Devi Parikh, and Marcus Rohrbach. Towards vqa models that can read. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 8317–8326, 2019. 5, 6
2019
-
[55]
Gemma: Open models based on gemini research and tech- nology
Gemma Team, Thomas Mesnard, Cassidy Hardin, Robert Dadashi, Surya Bhupatiraju, Shreya Pathak, Laurent Sifre, Morgane Rivi `ere, Mihir Sanjay Kale, Juliette Love, et al. Gemma: Open models based on gemini research and tech- nology. arXiv preprint arXiv:2403.08295, 2024. 7
2024 arXiv
-
[56]
Llama 2: Open foundation and fine-tuned chat models.arXiv preprint arXiv:2307.09288, 2023
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. Llama 2: Open foundation and fine-tuned chat models.arXiv preprint arXiv:2307.09288, 2023. 1, 2
2023 arXiv
-
[57]
Well-read students learn better: On the im- portance of pre-training compact models
Iulia Turc, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. Well-read students learn better: On the im- portance of pre-training compact models. arXiv preprint arXiv:1908.08962v2, 2019. 3
1908 arXiv
-
[58]
Qwen2-vl: Enhancing vision-language model’s perception of the world at any resolution
Peng Wang, Shuai Bai, Sinan Tan, Shijie Wang, Zhihao Fan, Jinze Bai, Keqin Chen, Xuejing Liu, Jialin Wang, Wenbin Ge, et al. Qwen2-vl: Enhancing vision-language model’s perception of the world at any resolution. arXiv preprint arXiv:2409.12191, 2024. 1, 2, 3, 4, 5, 6, 7
2024 arXiv
-
[59]
Show, attend and tell: Neural image cap- tion generation with visual attention
Kelvin Xu. Show, attend and tell: Neural image cap- tion generation with visual attention. arXiv preprint arXiv:1502.03044, 2015. 2, 4
2015 arXiv
-
[60]
Stacked attention networks for image question answering
Zichao Yang, Xiaodong He, Jianfeng Gao, Li Deng, and Alex Smola. Stacked attention networks for image question answering. In Proceedings of the IEEE conference on com- puter vision and pattern recognition, pages 21–29, 2016. 2, 4
2016
-
[61]
mplug-owl: Modularization empowers large language mod- els with multimodality, 2024
Qinghao Ye, Haiyang Xu, Guohai Xu, Jiabo Ye, Ming Yan, Yiyang Zhou, Junyang Wang, Anwen Hu, Pengcheng Shi, Yaya Shi, Chenliang Li, Yuanhong Xu, Hehong Chen, Jun- feng Tian, Qi Qian, Ji Zhang, Fei Huang, and Jingren Zhou. mplug-owl: Modularization empowers large language mod- e...
2024
-
[62]
Fit and prune: Fast and training-free visual token pruning 11 for multi-modal large language models
Weihao Ye, Qiong Wu, Wenhao Lin, and Yiyi Zhou. Fit and prune: Fast and training-free visual token pruning 11 for multi-modal large language models. arXiv preprint arXiv:2409.10197, 2024. 2, 3
2024 arXiv
-
[63]
Mm-vet: Evaluating large multimodal models for integrated capabilities, 2023
Weihao Yu, Zhengyuan Yang, Linjie Li, Jianfeng Wang, Kevin Lin, Zicheng Liu, Xinchao Wang, and Lijuan Wang. Mm-vet: Evaluating large multimodal models for integrated capabilities, 2023. 5
2023
-
[64]
Deep modular co-attention networks for visual question an- swering
Zhou Yu, Jun Yu, Yuhao Cui, Dacheng Tao, and Qi Tian. Deep modular co-attention networks for visual question an- swering. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 6281–6290,
-
[65]
Tinygpt-v: Efficient multimodal large lan- guage model via small backbones, 2024
Zhengqing Yuan, Zhaoxu Li, Weiran Huang, Yanfang Ye, and Lichao Sun. Tinygpt-v: Efficient multimodal large lan- guage model via small backbones, 2024. 2
2024
-
[66]
Mmmu: A massive multi-discipline multimodal understand- ing and reasoning benchmark for expert agi
Xiang Yue, Yuansheng Ni, Kai Zhang, Tianyu Zheng, Ruoqi Liu, Ge Zhang, Samuel Stevens, Dongfu Jiang, Weiming Ren, Yuxuan Sun, Cong Wei, Botao Yu, Ruibin Yuan, Ren- liang Sun, Ming Yin, Boyuan Zheng, Zhenzhu Yang, Yibo Liu, Wenhao Huang, Huan Sun, Yu Su, and Wenhu Chen. Mmmu: A...
2024
-
[67]
Sigmoid loss for language image pre-training,
Xiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, and Lucas Beyer. Sigmoid loss for language image pre-training,
-
[68]
Haotian Zhang, Mingfei Gao, Zhe Gan, Philipp Dufter, Nina Wenzel, Forrest Huang, Dhruti Shah, Xianzhi Du, Bowen Zhang, Yanghao Li, et al. Mm1. 5: Methods, analysis & insights from multimodal llm fine-tuning. arXiv preprint arXiv:2409.20566, 2024. 2, 7
2024 arXiv
-
[69]
Lmms- eval: Reality check on the evaluation of large multimodal models, 2024
Kaichen Zhang, Bo Li, Peiyuan Zhang, Fanyi Pu, Joshua Adrian Cahyono, Kairui Hu, Shuai Liu, Yuanhan Zhang, Jingkang Yang, Chunyuan Li, and Ziwei Liu. Lmms- eval: Reality check on the evaluation of large multimodal models, 2024. 5
2024
-
[70]
Vinvl: Revisiting visual representations in vision-language models
Pengchuan Zhang, Xiujun Li, Xiaowei Hu, Jianwei Yang, Lei Zhang, Lijuan Wang, Yejin Choi, and Jianfeng Gao. Vinvl: Revisiting visual representations in vision-language models. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 5579–5588,
-
[71]
Opt: Open pre-trained trans- former language models
Susan Zhang, Stephen Roller, Naman Goyal, Mikel Artetxe, Moya Chen, Shuohui Chen, Christopher Dewan, Mona Diab, Xian Li, Xi Victoria Lin, et al. Opt: Open pre-trained trans- former language models. arXiv preprint arXiv:2205.01068,
-
[72]
Free vqa models from knowledge iner- tia by pairwise inconformity learning
Yiyi Zhou, Rongrong Ji, Jinsong Su, Xiangming Li, and Xiaoshuai Sun. Free vqa models from knowledge iner- tia by pairwise inconformity learning. In Proceedings of the AAAI Conference on Artificial Intelligence, pages 9316– 9323, 2019. 2
2019
-
[73]
Trar: Routing the attention spans in transformer for visual question answering
Yiyi Zhou, Tianhe Ren, Chaoyang Zhu, Xiaoshuai Sun, Jianzhuang Liu, Xinghao Ding, Mingliang Xu, and Ron- grong Ji. Trar: Routing the attention spans in transformer for visual question answering. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV),...
2021
-
[74]
Minigpt-4: Enhancing vision-language understanding with advanced large language models, 2023
Deyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li, and Mo- hamed Elhoseiny. Minigpt-4: Enhancing vision-language understanding with advanced large language models, 2023. 2, 3, 4
2023
-
[75]
airport": xxx,
Minjie Zhu, Yichen Zhu, Xin Liu, Ning Liu, Zhiyuan Xu, Chaomin Shen, Yaxin Peng, Zhicai Ou, Feifei Feng, and Jian Tang. A comprehensive overhaul of multimodal assistant with small language models. arXiv preprint arXiv:2403.06199, 2024. 3, 7 12 A. Quantitative analysis A.1. Imp...
2024 arXiv
-
[76]
Observe the pattern: The sum of the values of the blue circle and the red triangle is 5
-
[77]
The sum of the values of two yellow triangle is 8
-
[78]
The sum of the values of the blue circle and the yellow square is 7
-
[79]
Since the sum of the values of two yellow triangle is 8 , each yellow triangle has a value of 4
-
[80]
From step 3, the sum of the blue circle and the yellow square is 7
-
[81]
Therefore, the value of the square in the figure is 2
From step 1, the sum of the blue circle and the red triangle is 5. Therefore, the value of the square in the figure is 2. FlashSloth-HD: The image shows a brain teaser for an IQ test. The question asks to find the value of the square in the figure. To solve this, we need to lo...
-
[82]
control center
This means that the value of the red square is 7 - 5 = 2. the value of the square in the figure is 2. Figure 5. FlashSloth’s Performance in Ticket OCR Recognition and Mathematical Question Answering. 14 Multi - object Recognition USER: Output the color and number of each box. ...
-
[84]
Increased: 48.8% - This section represents the news directors who reported an increase in their local TV newsroom budgets in 2012
2012
-
[85]
Same: 38.5% - This section represents the news directors who reported that their local TV newsroom budgets remained the same in 2012
2012
-
[86]
Decreased: 5.7% - This section represents the news directors who reported a decrease in their local TV newsroom budgets in 2012
2012
-
[87]
RIGHT",
Not Sure: 7% - This section represents the news directors who were unsure about the change in their local TV newsroom budgets in 2012. Figure 7. FlashSloth’s Performance in Artistic Question Answering and Chart Content Analysis. 16 Code Generation USER: Solve the algorithmic p...
2012
-
[2012]
Decreased
The chart is divided into three sections, each representing a different category of change. The largest section, colored in a dark red hue, represents 48.8% of news directors who reported an increase in their local TV news budgets. This indicates that nearly half of the news d...
2012
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.