REVIEW 3 major objections 8 minor 67 references
B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens
T0 review · 3 major / 8 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read A text question can pick the video frames that matter, keeping long-video QA accurate on a fixed token budget.
desk verdict Solid, well-executed engineering paper on balancing spatio-temporal visual tokens for video LLMs; the central claim holds, but the SOTA framing and CLS-based frame selection need tighter handling. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The mechanism is a two-stage token allocator: a Gumbel-Softmax frame-selection network (a differentiable approximation to discrete sampling) turns the text question and [CLS] tokens into a sparse matrix $S_\tau$ over frames, followed by a spatial Q-Former (a lightweight transformer that pools many tokens into a few query tokens) that samples $R$ question-relevant tokens per selected frame, with optional bipartite token merging to meet a hard token ceiling $\theta$. The [CLS] token is the load-bearing shortcut: it makes whole-video frame selection cheap, so the budget can be spent on fine tokens only where they matter.
What would settle it
Take a video in which the answer appears in a single brief frame whose global CLS token is dominated by irrelevant content, and ask a question that targets that frame; if a variant that selects frames from pooled visual tokens answers correctly while the CLS-based selector does not, the load-bearing shortcut is falsified.
Extended reading notes
Core claim
The central discovery is that a single visual-token budget can support both temporal and spatial understanding if the allocation is conditioned on the text prompt. Each frame contributes a coarse [CLS] token and a set of fine visual tokens; the [CLS] tokens are cheap enough to let a selection network scan all frames, and the selected frames' fine tokens are then reduced by sampling and optional iterative merging to hit the budget. In the paper's experiments, this yields consistent gains over both a baseline that compresses each frame to two tokens and a baseline that uniformly samples eight frames, with the largest gains on long-video benchmarks and a reported ten-percent gain on one multi-purpose video benchmark.
Load-bearing premise
The frame-selection stage trusts a single coarse summary token per frame; if the detail a question depends on survives only in the frame's other visual tokens, that frame can be discarded before it is ever examined.
Editorial extensions
If this is right
- Long-video benchmarks that previously required hour-long context windows can be answered with a fixed, modest number of visual tokens.
- Text-conditioned frame selection can outperform uniform frame sampling, so a video model no longer needs to see every frame to reason about the whole video.
- Heavy per-frame compression (two tokens per frame) is not necessary for temporal reasoning; keeping more spatial tokens on selected frames improves both spatial and temporal benchmarks.
- The same token-sampling machinery transfers to image benchmarks, where the model performs comparably to image-only models while using far fewer visual tokens per image.
- Because the token budget can be set in advance, with performance plateauing rather than degrading, accuracy and compute can be traded deliberately.
Reading between the lines
- A natural next test would feed the model 'needle' videos where the answer lives in a single brief frame whose CLS token is dominated by irrelevant content; such cases would reveal whether coarse CLS semantics suffice or whether the selection stage needs pooled features.
- The same text-conditioned allocation could apply to other long-input settings, such as document retrieval, where a cheap global descriptor filters before expensive local tokens are spent.
- The paper's saturation result around 512 visual tokens suggests a practical rule of thumb: most video question answering can be answered from roughly half a thousand visual tokens if selection is question-conditional.
- The paper itself points toward removing the fixed selected-frame count $L^*$; an adaptive count would let extremely long videos with many relevant moments avoid discarding information.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces B-VLLM, a framework for video large language models that controls the number of visual tokens by combining text-conditioned frame selection (using per-frame CLS tokens and a Q-Former), temporal frame-token merging to remove duplicates, spatial token sampling via another Q-Former, and an optional iterative token-merging step. The method is evaluated on several video QA benchmarks and two image benchmarks, both as a standalone model (with Qwen2) and as an integration on top of LLaMA-VID and VideoLLaMA2. The authors report consistent gains over the two baselines, claim state-of-the-art performance on several benchmarks, and provide ablations for each module and for the main hyperparameters. The paper also includes a transparent limitations section and a supplementary study comparing frame-selection features.
Significance. If the reported results are reproducible, B-VLLM is a practical contribution to efficient long-video understanding: it keeps the visual-token budget fixed and shows that text-conditioned selection and spatial sampling can outperform uniform frame sampling and per-frame compression baselines on several benchmarks. The paper's strengths include a clear architecture description, an extensive set of ablations (Tables 3-4, Figure 3, and supplementary tables), a code release, and a fair comparison of training data amounts. The main weakness is that the central frame-selection mechanism relies on CLS tokens that the authors' own supplementary ablation shows to be inferior to mean pooling, and the unqualified state-of-the-art claim conflicts with entries in the paper's own tables. These issues, together with the use of the test benchmarks for hyperparameter selection, require attention before the central claims can be accepted as stated.
major comments (3)
- [§5.1, Table 1 and Supplementary Table 9] The claim that B-VLLM 'achieves SOTA performance' is contradicted by the paper's own results: in Table 1, VideoChat2 (51.1) and ShareGPT4Video (51.2) outperform B-VLLM (50.8) on MVBench, and Supplementary Table 9 shows Qwen2VL, LLaVA-OV, and InternVL2 well above B-VLLM on VideoMME and EgoSchema. The abstract and introduction repeat the SOTA claim without the qualification that it holds only with respect to the same-training-data baselines. Please scope the claim to the actual comparison set and ensure the headline does not overstate the contribution.
- [§5.3 and Supplementary Table 7] The frame-selection module uses CLS tokens for efficiency, but the paper's own ablation (Table 7) shows that mean pooling gives 51.3 vs 50.8 on MVBench, 54.4 vs 52.9 on VMME, 48.2 vs 46.0 on VMME-OCR, and 36.6 vs 34.0 on VMME-Counting; a Q-Former feature also improves MVBench and EgoSchema. The statement in §5.3 that CLS is 'sufficient' for frame selection is not supported by these numbers. Since adaptive frame selection is a core contribution, please either adopt the better feature in the final model, report main results with the best feature, or explicitly position the reported model as an efficiency-optimized variant and state the performance cost.
- [Figure 3] The hyperparameters L*, R, τ, γ, and θ are analyzed by sweeping on MVBench and VideoMME-Medium, which are the same benchmarks used for the final reported accuracies. This is a form of test-set leakage and the paper's characterization of the evaluation as 'zero-shot' is therefore overstated. Please select hyperparameters on a held-out validation split, or report the tuning protocol explicitly and discuss the risk of overfitting.
minor comments (8)
- [§3.2, Eq. (1)] The notation V* = Sτ·V denotes a soft convex combination, not a hard selection, yet the text refers to 'selected frames' throughout. Please clarify how hard selection is implemented at inference (for example, by taking the argmax row of Sτ or by setting τ to a very small value) and how this interacts with the temporal-order restoration.
- [§3.3, Eq. (4)] Averaging visual-token sets of duplicate frames is not well defined; please specify whether the average is element-wise over spatially aligned tokens or performed using a more invariant pooling scheme.
- [§5.1, Table 1] The -1.7 drop on VideoMME-Short for VideoLLaMA2 w. Ours is not mentioned in the text, which claims consistent short-video gains; please acknowledge this exception in the discussion.
- [Abstract and Introduction] The phrase '10% performance gain on MVBench' is ambiguous (absolute vs. relative) and does not match the numbers in Table 1; please specify the baseline and the type of gain.
- [Figure 3] The plots lack axis labels and benchmark legends, so readers cannot determine which curve corresponds to which benchmark or what the plotted quantity is.
- [Supplementary Table 7] Please report the hardware and software configuration used for the training-time comparison so that the efficiency claim in §5.3 is reproducible.
- [§2.2] The claim of being 'for the first time' in iterative token merging for controllable token counts is not substantiated; please soften the statement or provide a more specific comparison with prior controllable token-reduction methods.
- [References and Supplementary §7.4] There are duplicated citations (for example, MME appears as both [13] and [14]), and the reference to 'EV A-CLIP' in §7.4 appears to be missing its arXiv identifier.
Circularity Check
No significant circularity: the paper's claims are evaluated against external benchmarks and its components are ablated, not derived from the target results.
full rationale
B-VLLM's central claims are empirical: adaptive frame selection, temporal token merging, and spatial token sampling are proposed as mechanisms, and their effectiveness is measured on external video and image benchmarks (MVBench, VideoMME, EgoSchema, MMBench, POPE, etc.). No equation in the paper defines a predicted quantity in terms of the fitted parameters that produced it, nor is any reported accuracy obtained by construction from training data. The frame-selection module uses [CLS] tokens, and Section 5.4 and Supplementary Table 7 openly acknowledge that pooled visual tokens are more informative; this is a stated limitation and a design trade-off, not a circular step, because the downstream evaluation remains independent of that choice. Hyperparameters such as the number of selected frames L*, spatial tokens R, and thresholds are tuned with reference to the same benchmarks later reported, which is a form of selection bias rather than circular reasoning under the definitions used here. Self-citations in the reference list are not load-bearing for the core derivation, and no uniqueness theorem or prior-work assertion is used to force the method's design. The paper is self-contained against external benchmarks, so the appropriate circularity score is 0.
Assumptions & free parameters
free parameters (5)
- L*, number of selected frames =
32 (default; ablation shows saturation at 28)
- R, spatial visual tokens sampled per frame =
32
- theta, maximum visual token count =
512 or 768 in experiments; final default not clearly stated
- tau, Gumbel-Softmax temperature =
0.1
- gamma, duplicate frame similarity threshold =
0.75 or 1.0 in experiments; default not clearly stated
assumptions (3)
- standard math Gumbel-Softmax with temperature tau behaves as a categorical sampler for frame selection.
- domain assumption CLS tokens summarize enough frame content for question-relevant selection.
- domain assumption Public video QA benchmarks accurately measure video understanding and are not contaminated by the training data.
Cite this review
Pith. "Pith review of B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens." pith.science (2026). https://pith.science/paper/UNLEZZ4Z
@misc{pith2026241209919,
author = {Pith},
title = {Pith review of: B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens},
year = {2026},
howpublished = {\url{https://pith.science/paper/UNLEZZ4Z}},
note = {Machine review of arXiv:2412.09919}
}
read the original abstract
Recently, Vision Large Language Models (VLLMs) integrated with vision encoders have shown promising performance in vision understanding. The key of VLLMs is to encode visual content into sequences of visual tokens, enabling VLLMs to simultaneously process both visual and textual content. However, understanding videos, especially long videos, remain a challenge to VLLMs as the number of visual tokens grows rapidly when encoding videos, resulting in the risk of exceeding the context window of VLLMs and introducing heavy computation burden. To restrict the number of visual tokens, existing VLLMs either: (1) uniformly downsample videos into a fixed number of frames or (2) reducing the number of visual tokens encoded from each frame. We argue the former solution neglects the rich temporal cue in videos and the later overlooks the spatial details in each frame. In this work, we present Balanced-VLLM (B-VLLM): a novel VLLM framework that aims to effectively leverage task relevant spatio-temporal cues while restricting the number of visual tokens under the VLLM context window length. At the core of our method, we devise a text-conditioned adaptive frame selection module to identify frames relevant to the visual understanding task. The selected frames are then de-duplicated using a temporal frame token merging technique. The visual tokens of the selected frames are processed through a spatial token sampling module and an optional spatial token merging strategy to achieve precise control over the token count. Experimental results show that B-VLLM is effective in balancing the number of frames and visual tokens in video understanding, yielding superior performance on various video understanding benchmarks. Our code is available at https://github.com/zhuqiangLu/B-VLLM.
Figures
Figures from the paper (7 more)
Reference graph
Works this paper leans on
- [1]
-
[2]
Flamingo: A Visual Language Model For Few-shot Learning
Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katherine Millican, Malcolm Reynolds, et al. Flamingo: A Visual Language Model For Few-shot Learning. InNeurIPS, pages 23716–23736, 2022. 1, 2
work page 2022
-
[3]
Deepspeed-Inference: Enabling Efficient Inference of Trans- former Models at Unprecedented Scale
Reza Yazdani Aminabadi, Samyam Rajbhandari, Am- mar Ahmad Awan, Cheng Li, Du Li, Elton Zheng, Olatunji Ruwase, Shaden Smith, Minjia Zhang, Jeff Rasley, et al. Deepspeed-Inference: Enabling Efficient Inference of Trans- former Models at Unprecedented Scale. In SC, pages 1–15. IEEE, 2022. 2
work page 2022
-
[4]
Jinze Bai, Shuai Bai, Yunfei Chu, Zeyu Cui, Kai Dang, Xiaodong Deng, Yang Fan, Wenbin Ge, Yu Han, Fei Huang, et al. Qwen Technical Report. arXiv preprint arXiv:2309.16609, 2023. 1
arXiv 2023
-
[5]
Frozen in time: A joint video and image encoder for end-to-end retrieval
Max Bain, Arsha Nagrani, G ¨ul Varol, and Andrew Zisser- man. Frozen in time: A joint video and image encoder for end-to-end retrieval. In IEEE International Conference on Computer Vision, 2021. 2
work page 2021
-
[6]
Token Merging for Fast Stable Diffusion
Daniel Bolya and Judy Hoffman. Token Merging for Fast Stable Diffusion. In CVPR, pages 4599–4603, 2023. 2
work page 2023
-
[7]
Token Merging: Your ViT But Faster
Daniel Bolya, Cheng-Yang Fu, Xiaoliang Dai, Peizhao Zhang, Christoph Feichtenhofer, and Judy Hoffman. Token Merging: Your ViT But Faster. In ICLR, 2023. 2, 4
work page 2023
-
[8]
Dif- fusiondet: Diffusion model for object detection
Shoufa Chen, Peize Sun, Yibing Song, and Ping Luo. Dif- fusiondet: Diffusion model for object detection. In CVPR,
Show all 67 references
-
[9]
VideoLLaMA 2: Advancing Spatial- Temporal Modeling and Audio Understanding in Video- LLMs
Zesen Cheng, Sicong Leng, Hang Zhang, Yifei Xin, Xin Li, Guanzheng Chen, Yongxin Zhu, Wenqi Zhang, Ziyang Luo, Deli Zhao, et al. VideoLLaMA 2: Advancing Spatial- Temporal Modeling and Audio Understanding in Video- LLMs. arXiv preprint arXiv:2406.07476, 2024. 2, 3, 5, 6, 8
2024 arXiv
-
[10]
An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.arXiv preprint arXiv:2010.11929, 2020
Alexey Dosovitskiy. An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.arXiv preprint arXiv:2010.11929, 2020. 2
2010 arXiv
-
[11]
The LLaMA 3 Herd of Models
Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, et al. The LLaMA 3 Herd of Models. arXiv preprint arXiv:2407.21783, 2024. 1
2024 arXiv
-
[12]
ActivityNet: A Large-Scale Video Benchmark for Human Activity Understanding
Bernard Ghanem Fabian Caba Heilbron, Victor Escorcia and Juan Carlos Niebles. ActivityNet: A Large-Scale Video Benchmark for Human Activity Understanding. In CVPR, pages 961–970, 2015. 5, 2
2015
-
[13]
MME:A Comprehensive Evaluation Benchmark for Multimodal Large Language Models
Chaoyou Fu, Peixian Chen, Yunhang Shen, Yulei Qin, Mengdan Zhang, Xu Lin, Jinrui Yang, Xiawu Zheng, Ke Li, Xing Sun, et al. MME:A Comprehensive Evaluation Benchmark for Multimodal Large Language Models. arXiv preprint arXiv:2306.13394, 2023. 6
2023 arXiv
-
[14]
MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models, 2024
Chaoyou Fu, Peixian Chen, Yunhang Shen, Yulei Qin, Mengdan Zhang, Xu Lin, Jinrui Yang, Xiawu Zheng, Ke Li, Xing Sun, Yunsheng Wu, and Rongrong Ji. MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models, 2024. 2, 5
2024
-
[15]
Video-MME: The First-Ever Com- prehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis
Chaoyou Fu, Yuhan Dai, Yondong Luo, Lei Li, Shuhuai Ren, Renrui Zhang, Zihan Wang, Chenyu Zhou, Yunhang Shen, Mengdan Zhang, et al. Video-MME: The First-Ever Com- prehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis. arXiv preprint arXiv:2405.21075, 2024. 5, 6, 7, 2
2024 arXiv
-
[16]
Making the V in VQA Matter: Ele- vating the Eole of Image Understanding in Visual Question Answering
Yash Goyal, Tejas Khot, Douglas Summers-Stay, Dhruv Ba- tra, and Devi Parikh. Making the V in VQA Matter: Ele- vating the Eole of Image Understanding in Visual Question Answering. In CVPR, pages 6904–6913, 2017. 5, 6, 2
2017
-
[17]
Vizwiz Grand Challenge: Answering Visual Questions from Blind People
Danna Gurari, Qing Li, Abigale J Stangl, Anhong Guo, Chi Lin, Kristen Grauman, Jiebo Luo, and Jeffrey P Bigham. Vizwiz Grand Challenge: Answering Visual Questions from Blind People. In CVPR, pages 3608–3617, 2018. 5, 6
2018
-
[18]
LoRA: Low-Rank Adaptation of Large Language Models
Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen- Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. LoRA: Low-Rank Adaptation of Large Language Models. arXiv preprint arXiv:2106.09685, 2021. 2
2021 arXiv
-
[19]
Vision-based freezing of gait detection with anatomic patch based representation
Kun Hu, Zhiyong Wang, Kaylena Ehgoetz Martens, and Si- mon Lewis. Vision-based freezing of gait detection with anatomic patch based representation. In ACCV, 2018. 2
2018
-
[20]
GQA: A New Dataset for Real-World Visual Eeasoning and Compositional Question Answering
Drew A Hudson and Christopher D Manning. GQA: A New Dataset for Real-World Visual Eeasoning and Compositional Question Answering. In CVPR, pages 6700–6709, 2019. 5, 6, 2
2019
-
[21]
LLMlingua: Compressing Prompts for Accel- erated Inference of Large Language Models
Huiqiang Jiang, Qianhui Wu, Chin-Yew Lin, Yuqing Yang, and Lili Qiu. LLMlingua: Compressing Prompts for Accel- erated Inference of Large Language Models. arXiv preprint arXiv:2310.05736, 2023. 3
2023 arXiv
-
[22]
Chat-univi: Unified Visual Representation Em- powers Large Language Models with Image and Video Un- derstanding
Peng Jin, Ryuichi Takanobu, Wancai Zhang, Xiaochun Cao, and Li Yuan. Chat-univi: Unified Visual Representation Em- powers Large Language Models with Image and Video Un- derstanding. In CVPR, pages 13700–13710, 2024. 3
2024
-
[23]
ReferItGame: Referring to objects in pho- tographs of natural scenes
Sahar Kazemzadeh, Vicente Ordonez, Mark Matten, and Tamara Berg. ReferItGame: Referring to objects in pho- tographs of natural scenes. In Proceedings of the 2014 Con- ference on Empirical Methods in Natural Language Process- ing (EMNLP), pages 787–798, Doha, Qatar, 2014. Assoc...
2014
-
[24]
Shamma, Michael S
Ranjay Krishna, Yuke Zhu, Oliver Groth, Justin Johnson, Kenji Hata, Joshua Kravitz, Stephanie Chen, Yannis Kalan- tidis, Li-Jia Li, David A. Shamma, Michael S. Bernstein, and Fei-Fei Li. Visual genome: Connecting language and vision using crowdsourced dense image annotations, 2016. 2
2016
-
[25]
Seed-Bench: Benchmarking Multimodal LLMs with Generative Comprehension
Bohao Li, Rui Wang, Guangzhi Wang, Yuying Ge, Yixiao Ge, and Ying Shan. Seed-Bench: Benchmarking Multimodal LLMs with Generative Comprehension. arXiv preprint arXiv:2307.16125, 2023. 5
2023 arXiv
-
[26]
BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models
Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi. BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models. In ICML, pages 19730–19742. PMLR, 2023. 2, 3
2023
-
[27]
MVBench: A Comprehensive Multi-Modal Video Under- standing Benchmark
Kunchang Li, Yali Wang, Yinan He, Yizhuo Li, Yi Wang, Yi Liu, Zun Wang, Jilan Xu, Guo Chen, Ping Luo, et al. MVBench: A Comprehensive Multi-Modal Video Under- standing Benchmark. In CVPR, pages 22195–22206, 2024. 2, 5, 6, 7
2024
-
[28]
VidToMe: Video Token Merging for Zero-Shot Video Edit- ing
Xirui Li, Chao Ma, Xiaokang Yang, and Ming-Hsuan Yang. VidToMe: Video Token Merging for Zero-Shot Video Edit- ing. In CVPR, pages 7486–7495, 2024. 2
2024
-
[29]
Evaluating Object Hallucination in Large Vision-Language Models
Yifan Li, Yifan Du, Kun Zhou, Jinpeng Wang, Wayne Xin Zhao, and Ji Rong Wen. Evaluating Object Hallucination in Large Vision-Language Models. In EMNLP, 2023. 5, 6
2023
-
[30]
LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models
Yanwei Li, Chengyao Wang, and Jiaya Jia. LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models. In ECCV, pages 323–340. Springer, 2025. 2, 3, 5, 6, 7
2025
-
[31]
Video-LLaV A: Learning United Visual Repre- sentation by Alignment Before Projection
Bin Lin, Yang Ye, Bin Zhu, Jiaxi Cui, Munan Ning, Peng Jin, and Li Yuan. Video-LLaV A: Learning United Visual Repre- sentation by Alignment Before Projection. arXiv preprint arXiv:2311.10122, 2023. 2
2023 arXiv
-
[32]
VILA: On Pre-training for Vi- sual Language Models
Ji Lin, Hongxu Yin, Wei Ping, Pavlo Molchanov, Moham- mad Shoeybi, and Song Han. VILA: On Pre-training for Vi- sual Language Models. InCVPR, pages 26689–26699, 2024. 2
2024
-
[33]
Visual Instruction Tuning
Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. Visual Instruction Tuning. In NeurIPS, 2024. 1, 2, 3, 4, 5, 6
2024
-
[34]
MMBench: Is Your Multi-Modal Model an All-Around Player? In ECCV, pages 216–233
Yuan Liu, Haodong Duan, Yuanhan Zhang, Bo Li, Songyang Zhang, Wangbo Zhao, Yike Yuan, Jiaqi Wang, Conghui He, Ziwei Liu, et al. MMBench: Is Your Multi-Modal Model an All-Around Player? In ECCV, pages 216–233. Springer,
-
[35]
Decoupled Weight Decay Regularization
I Loshchilov. Decoupled Weight Decay Regularization. arXiv preprint arXiv:1711.05101, 2017. 2
2017 arXiv
-
[36]
Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question Answering
Pan Lu, Swaroop Mishra, Tanglin Xia, Liang Qiu, Kai-Wei Chang, Song-Chun Zhu, Oyvind Tafjord, Peter Clark, and Ashwin Kalyan. Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question Answering. In NeurIPS, pages 2507–2521, 2022. 5
2022
-
[37]
Autoregressive omni-aware outpainting for open- vocabulary 360-degree image generation
Zhuqiang Lu, Kun Hu, Chaoyue Wang, Lei Bai, and Zhiy- ong Wang. Autoregressive omni-aware outpainting for open- vocabulary 360-degree image generation. In AAAI, 2024. 2
2024
-
[38]
Valley: Video Assistant with Large Language Model Enhanced Ability
Ruipu Luo, Ziwang Zhao, Min Yang, Junwei Dong, Da Li, Pengcheng Lu, Tao Wang, Linmei Hu, Minghui Qiu, and Zhongyu Wei. Valley: Video Assistant with Large Language Model Enhanced Ability. arXiv preprint arXiv:2306.07207,
-
[39]
Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Mod- els
Muhammad Maaz, Hanoona Rasheed, Salman Khan, and Fahad Shahbaz Khan. Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Mod- els. In ACL, 2024. 2
2024
-
[40]
Egoschema: A Diagnostic Benchmark for Bery Long-Form Video Language Understanding
Karttikeya Mangalam, Raiymbek Akshulakov, and Jitendra Malik. Egoschema: A Diagnostic Benchmark for Bery Long-Form Video Language Understanding. In NeurIPS, pages 46212–46244, 2023. 5
2023
-
[41]
Generation and comprehension of unambiguous object descriptions, 2016
Junhua Mao, Jonathan Huang, Alexander Toshev, Oana Camburu, Alan Yuille, and Kevin Murphy. Generation and comprehension of unambiguous object descriptions, 2016. 2
2016
-
[42]
Ocr-vqa: Visual question answering by reading text in images
Anand Mishra, Shashank Shekhar, Ajeet Kumar Singh, and Anirban Chakraborty. Ocr-vqa: Visual question answering by reading text in images. In ICDAR, 2019. 2
2019
-
[43]
Perception Test: A Diagnostic Benchmark for Multimodal Video Models
Viorica P ˘atr˘aucean, Lucas Smaira, Ankush Gupta, Adri`a Re- casens Continente, Larisa Markeeva, Dylan Banarse, Skanda Koppula, Joseph Heyward, Mateusz Malinowski, Yi Yang, Carl Doersch, Tatiana Matejovicova, Yury Sulsky, Antoine Miech, Alex Frechette, Hanna Klimczak, Raphael...
2023
-
[44]
Learning Transferable Visual Models from Natural Language Supervi- sion
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning Transferable Visual Models from Natural Language Supervi- sion. In ICML, pages 8748–8763. PMLR, 2021. 2
2021
-
[45]
A-OKVQA: A Benchmark for Visual Question Answering using World Knowledge, 2022
Dustin Schwenk, Apoorv Khandelwal, Christopher Clark, Kenneth Marino, and Roozbeh Mottaghi. A-OKVQA: A Benchmark for Visual Question Answering using World Knowledge, 2022. 2
2022
-
[46]
Llava-prumerge: Adaptive Token Reduction for Efficient Large Multimodal Models
Yuzhang Shang, Mu Cai, Bingxin Xu, Yong Jae Lee, and Yan Yan. Llava-prumerge: Adaptive Token Reduction for Efficient Large Multimodal Models. arXiv preprint arXiv:2403.15388, 2024. 3
2024
-
[47]
Conceptual captions: A cleaned, hypernymed, im- age alt-text dataset for automatic image captioning
Piyush Sharma, Nan Ding, Sebastian Goodman, and Radu Soricut. Conceptual captions: A cleaned, hypernymed, im- age alt-text dataset for automatic image captioning. In Pro- ceedings of ACL, 2018. 2
2018
-
[48]
TextCaps: a Dataset for Image Captioning with Reading Comprehension, 2020
Oleksii Sidorov, Ronghang Hu, Marcus Rohrbach, and Amanpreet Singh. TextCaps: a Dataset for Image Captioning with Reading Comprehension, 2020. 2
2020
-
[49]
Towards VQA Models That Can Read
Amanpreet Singh, Vivek Natarajan, Meet Shah, Yu Jiang, Xinlei Chen, Dhruv Batra, Devi Parikh, and Marcus Rohrbach. Towards VQA Models That Can Read. In VPR,
-
[50]
Moviechat: From Dense To- ken to Sparse Memory for Long Video Understanding
Enxin Song, Wenhao Chai, Guanhong Wang, Yucheng Zhang, Haoyang Zhou, Feiyang Wu, Haozhe Chi, Xun Guo, Tian Ye, Yanting Zhang, et al. Moviechat: From Dense To- ken to Sparse Memory for Long Video Understanding. In CVPR, pages 18221–18232, 2024. 2
2024
-
[51]
LLaMA 2: Open Foundation and Fine-tuned Chat Models
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. LLaMA 2: Open Foundation and Fine-tuned Chat Models. arXiv preprint arXiv:2307.09288, 2023. 1
2023 arXiv
-
[52]
[cls] token tells everything needed for training-free efficient mllms
Ao Wang, Fengyuan Sun, Hui Chen, Zijia Lin, Jungong Han, and Guiguang Ding. [cls] token tells everything needed for training-free efficient mllms. arXiv preprint arXiv:2412.05819, 2024. 3
2024 arXiv
-
[53]
VisionLLM: Large Language Model is Also An Open-ended Decoder for Vision-Centric Tasks
Wenhai Wang, Zhe Chen, Xiaokang Chen, Jiannan Wu, Xizhou Zhu, Gang Zeng, Ping Luo, Tong Lu, Jie Zhou, Yu Qiao, et al. VisionLLM: Large Language Model is Also An Open-ended Decoder for Vision-Centric Tasks. In NeurIPS,
-
[54]
LongVLM: Efficient Long Video Under- standing Via Large Language Models
Yuetian Weng, Mingfei Han, Haoyu He, Xiaojun Chang, and Bohan Zhuang. LongVLM: Efficient Long Video Under- standing Via Large Language Models. In ECCV, pages 453–
-
[55]
Video Question Answer- ing via Gradually Refined Attention over Appearance and Motion
Dejing Xu, Zhou Zhao, Jun Xiao, Fei Wu, Hanwang Zhang, Xiangnan He, and Yueting Zhuang. Video Question Answer- ing via Gradually Refined Attention over Appearance and Motion. In MM. 5, 2
-
[56]
Qwen2 Technical Report
An Yang, Baosong Yang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Zhou, Chengpeng Li, Chengyuan Li, Dayiheng Liu, Fei Huang, et al. Qwen2 Technical Report. arXiv preprint arXiv:2407.10671, 2024. 1, 5, 2
2024 arXiv
-
[57]
Surgicalpart- sam: Part-to-whole collaborative prompting for surgical in- strument segmentation
Wenxi Yue, Jing Zhang, Kun Hu, Qiuxia Wu, Zongyuan Ge, Yong Xia, Jiebo Luo, and Zhiyong Wang. Surgicalpart- sam: Part-to-whole collaborative prompting for surgical in- strument segmentation. arXiv preprint arXiv:2312.14481 ,
-
[58]
Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding
Hang Zhang, Xin Li, and Lidong Bing. Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding. arXiv preprint arXiv:2306.02858, 2023. 2
2023 arXiv
-
[59]
LLaV A-NeXT: A Strong Zero-shot Video Understanding Model, 2024
Yuanhan Zhang, Bo Li, haotian Liu, Yong jae Lee, Liangke Gui, Di Fu, Jiashi Feng, Ziwei Liu, and Chunyuan Li. LLaV A-NeXT: A Strong Zero-shot Video Understanding Model, 2024. 4, 5
2024
-
[60]
Needle In A Video Haystack: A Scalable Syn- thetic Framework for Benchmarking Video MLLMs
Zijia Zhao, Haoyu Lu, Yuqi Huo, Yifan Du, Tongtian Yue, Longteng Guo, Bingning Wang, Weipeng Chen, and Jing Liu. Needle In A Video Haystack: A Scalable Syn- thetic Framework for Benchmarking Video MLLMs. arXiv preprint, 2024. 5
2024
-
[61]
Clip in medical imaging: A survey
Zihao Zhao, Yuxiao Liu, Han Wu, Mei Wang, Yonghao Li, Sheng Wang, Lin Teng, Disheng Liu, Zhiming Cui, Qian Wang, et al. Clip in medical imaging: A survey. MIA, 2025. 2
2025
-
[62]
Languagebind: Extending Video-Language Pre- training to N-modality by Language-Based Semantic Align- ment
Bin Zhu, Bin Lin, Munan Ning, Yang Yan, Jiaxi Cui, HongFa Wang, Yatian Pang, Wenhao Jiang, Junwu Zhang, Zongwei Li, et al. Languagebind: Extending Video-Language Pre- training to N-modality by Language-Based Semantic Align- ment. arXiv preprint arXiv:2310.01852, 2023. 1
-
[63]
A closer look at the cls token for cross-domain few-shot learning
Yixiong Zou, Shuai Yi, Yuhua Li, and Ruixuan Li. A closer look at the cls token for cross-domain few-shot learning. In NeurIPS, 2024. 3 B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens Supplementary Material Table 5. Statistics of training datasets. S...
2024
-
[65]
Training Details We adopt a two-stage training strategy [9, 30, 33], dividing training into pretraining for modality alignment and fine- tuning for instruction tuning
Additional Implementation Details 7.1. Training Details We adopt a two-stage training strategy [9, 30, 33], dividing training into pretraining for modality alignment and fine- tuning for instruction tuning. During pretraining, the learn- ing rate is set to 1e-3 with a linear l...
-
[66]
Additional Discussion on Different Frame Se- lection Features
Additional Experiment & Discussion 8.1. Additional Discussion on Different Frame Se- lection Features. As reported in Table 7, we additionally report the perfor- mance of selecting frames by using three other feature ex- traction methods: (a) max pooling (b) mean pooling and (...
-
[67]
Game Science
More B-VLLM Qualitative Examples This section provides additional conversation examples with B-VLLM based on videos. Note that the presented videos are accessible only through the specified source. The examples are depicted in Figures 6 to 10. As shown in Fig- ure 6, when a ga...
-
[470]
Springer, 2025. 2, 3
2025
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.