REVIEW 3 major objections 6 minor 80 references
DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding
T0 review · 3 major / 6 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read In long-video question answering, a learned frame selector that spends detailed tokens only on answer-relevant frames can preserve accuracy while cutting memory use sharply.
desk verdict The paper's question-adaptive selection is not in the architecture: the selector never sees the question, so the core claim is unsupported despite a solid token-reduction engineering contribution. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is a learned frame-importance gate. Temporal DPC-KNN clustering produces $L$ event prototypes from pooled frame features; a small MLP score network $U$ assigns each prototype an importance score from its max- and average-pooled features ($s_l = U(\mathrm{Max}(m_l)\|\mathrm{Avg}(m_l))$); and a differentiable perturbed-top-K operator converts those scores into a binary mask $b_t$ over frames. The mask drives a cooperative encoder: selected frames receive a concatenation of event and multi-grained spatial prototypes as detailed tokens, while unselected frames are compressed to two tokens via average pooling plus text-guided attention. The token-allocation rule $O_t = b_t\cdot(U_{t,b_t=1}\|U_{t,b_t=0})+(1-b_t)\cdot U_{t,b_t=0}$ is the identity that carries the argument: it makes the token count a learned, content-dependent quantity rather than a fixed per-frame constant.
What would settle it
Build a benchmark where the same video is paired with questions whose answers live in disjoint frames (e.g., 'what color is the hat at 00:10?' vs 'what happens at 01:20?'), then compare DynFocus with a variant whose selector also receives the question text. If the question-conditioned variant does not clearly beat the content-only selector — or if an oracle mask computed from ground-truth answer frames does not beat the learned mask — the paper's claimed question-dependent correspondence is not what drives the results.
Extended reading notes
Core claim
The central claim is that dynamic, content-dependent token allocation can preserve question-answering accuracy under a tight token budget. Concretely, DynFocus clusters frame features temporally into event prototypes, learns a score function $U$ that ranks them by answer relevance via $s_l = U(\mathrm{Max}(m_l)\|\mathrm{Avg}(m_l))$, and then allocates fine-grained tokens to the selected frames and two compact tokens to every other frame, per $O_t = b_t\cdot(U_{t,b_t=1}\|U_{t,b_t=0}) + (1-b_t)\cdot U_{t,b_t=0}$. Selected frames are encoded with multi-grained spatial prototypes; the rest are reduced to a text-grounded global token plus a content token, providing a coarse temporal view. Trained end-to-end with a differentiable perturbed top-K, the model reaches competitive or state-of-the-art accuracy among the compared open baselines on short-video QA, MLVU, LV-Bench, and VideoMME, and places second among compared open models on the VideoHallucer hallucination diagnosis.
Load-bearing premise
The selection score is computed from frame content alone, with no input from the question, so the method assumes that one learned importance function can identify the frames that matter for any question.
Editorial extensions
If this is right
- Long-video QA becomes feasible under fixed memory: on LV-Bench, DynFocus processes 200 frames with a 7B model and beats open baselines that ingest thousands of frames.
- The token budget becomes a tunable dial: increasing the number of prototypes $L$ improves accuracy up to a point, and lowering the selection ratio $K/L$ helps on longer videos, so practitioners can trade memory for accuracy per domain.
- Frame selection can be trained from the LLM's own response loss without extra frame-level labels, because the perturbed top-K keeps the selector end-to-end differentiable.
- Filtering answer-irrelevant frames reduces visual noise, which the paper connects to better factual correctness scores on VCG-Bench and competitive hallucination diagnoses on VideoHallucer.
Reading between the lines
- The selector never sees the question text (Eq. 4 takes only prototype features), so the reported 'correspondence' may be an emergent property mimicked by content saliency rather than true per-question adaptation; conditioning the selector on the question is the natural next experiment and a sharper test of the paper's motivation.
- The rod/cone split suggests a general principle for multimodal token budgets: a cheap perceptual stream can provide temporal or global context while a costly detailed stream focuses on the answer-relevant region; the same allocation idea could transfer to audio or 3D video.
- Supervising the selector only with the final answer may under-specify which frames to keep; auxiliary objectives such as temporal grounding or frame-relevance labels could stabilize selection and make the masks interpretable.
- The paper's own ablation shows that removing the coarse stream costs several accuracy points even though it contains only two tokens per frame, so the coarse stream is supplying temporal continuity rather than acting as mere compression; token diversity, not just token count, appears to drive the gains.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes DynFocus, a video token-compression pipeline for LLM-based video question answering. A Dynamic Event Prototype Estimation (DPE) module clusters frame features into event prototypes and learns an MLP score to select 'important' prototypes; a Compact Cooperative Encoding (CCE) module then encodes selected frames with fine-grained tokens and the remaining frames with coarse, text-grounded tokens. The method is evaluated on short-video QA benchmarks, three long-video benchmarks, and a video-hallucination benchmark. The paper claims that the selection is question-adaptive and exploits 'correspondence' between frames and questions, but the implementation as written does not condition the selection on the question.
Significance. If the dynamic selection were genuinely question-conditional, DynFocus would offer a useful memory/accuracy trade-off for long-video QA, and the paper is commendable for releasing code, reporting extensive experiments on multiple benchmarks, and including detailed ablations. However, the central mechanism as implemented is not question-conditional, and the evaluation protocol appears to tune hyperparameters on test benchmarks. The descriptive redundancy statistics and the rod/cone analogy are interesting, but they do not compensate for the mismatch between the stated contribution and the actual equations. The paper would need a substantial revision of the selection mechanism and a re-evaluation under a fixed validation protocol before its central claims could be accepted.
major comments (3)
- [Sec. 3.2, Eq. (4)] The selection score is s_l = U(Max(m_l)||Avg(m_l)), where U is an MLP with no question argument; neither the question features Q nor any text-derived conditioning enters the score. Consequently, for a fixed video, the binary mask b in Eq. (11) is identical for every question, so the claimed 'dynamic selection in accordance with answer and question' is not implemented. The end-to-end LLM gradient can only make U prefer frames that are useful on average over the training questions; it cannot produce question-specific selection at inference. This directly contradicts the correspondence motivation in Sec. 1 and Fig. 1(b), and Fig. 5 does not demonstrate question-dependent selection because it shows different videos for different questions. This is the load-bearing mechanism of the paper and must be addressed, for example by conditioning U on Q or otherwise showing how question information affects b, with a re-evaluation after the change.
- [Sec. 3.2, Eqs. (5)-(7)] The linear-programming formulation of the Top-K selection is degenerate. The preceding text defines P as a stack of one-hot vectors with P ∈ {0,1}^{L×K}, but Eq. (5) states P ∈ R^{K×L} with constraint C = {P : P_{k,l} ≥ 0, 1^T P = 1}. If 1 has length K, then every column of P sums to 1, so the objective ⟨P, s1^T⟩ equals ∑_l s_l for any feasible P, independent of K; the optimization does not select K prototypes. If P is instead L×K as defined earlier, the constraint is dimensionally inconsistent, and when corrected to column-stochastic it permits duplicate selections. The resulting P_σ and H = P^T M therefore do not realize the claimed 'Top-K filtered event prototypes', and the token allocation in Eq. (11) is not as described. A correct differentiable top-K operator with a constraint set that provably selects K distinct prototypes, or a removal of the Top-K claim, is required.
- [Tables 3-4 and App. B.7] The hyperparameters L and K/L appear to be selected per evaluation benchmark. Table 3 reports five DynFocus variants with different (L, K/L) values and marks one as 'optimal results', and App. B.7 states that the optimal L shifts with video length and that a smaller K/L is better for longer videos. Because these choices are made on the test benchmarks and no held-out validation protocol is described, the reported numbers are not a fixed-model comparison and are inflated relative to baselines that use a single configuration. The paper should fix L and K/L using a validation split (or report results across a prespecified grid with appropriate multiple-comparison correction) before claiming competitive or SOTA performance.
minor comments (6)
- [Abstract vs. Sec. 1] The abstract says 'five publicly available benchmarks', while the contributions in Sec. 1 mention 'two short video benchmarks, three long video benchmarks, and one diagnosis benchmark', which sums to six; please reconcile this count.
- [Fig. 1(a)] The caption contains an undefined superscript '1' after 'video datasets'; the figure should be self-contained or the footnote should be restored, and the redundancy estimation procedure from App. A should be referenced explicitly.
- [Table 1] The evaluation version of GPT-3.5-Turbo is not fixed across methods; the authors' own note that the default GPT-3.5-Turbo version significantly impacts evaluation makes the comparison ambiguous. Please use one fixed version and state it for all methods.
- [Eq. (8)] The notation h_t in Eq. (8) is not defined; the event prototypes are indexed as h_k with k ∈ [1,K], and the mapping from h_k to the frame-specific mask b_t should be clarified.
- [References] Reference [13] 'John Doe and Jane Smith. EgoQA' appears to be a placeholder and should be replaced with the actual citation.
- [App. B.7, Fig.] The caption for the VideoMME figure mentions 'Medium (K=55)', which suggests per-duration selection of K; if so, this is a test-set selection and should be clearly labeled as such or moved to a validation-based protocol.
Circularity Check
No significant circularity: DynFocus's dynamic selection and cooperative encoding are trained end-to-end against external LLM supervision, and no prediction in the paper reduces by construction to a fitted parameter or to a self-citation.
full rationale
Walking the derivation chain: Sec. 3.2 computes event prototypes from DPC-KNN clustering (Eqs. 1-3) and scores them with an MLP (Eq. 4) whose supervision is the LLM's autoregressive loss via the differentiable top-k operator (Eqs. 5-7); this is a learned selector, not a pre-fitted constant. Sec. 3.3 allocates tokens according to the resulting binary mask (Eqs. 8-11), so the coarse/fine split is determined by the same learned scores, not by the benchmark answers at test time. The redundancy statistics in Appendix A are descriptive (thresholded CLIP similarities computed with q and a), and they are not part of the training loss, so they do not smuggle the target result into the method. The only overlapping-author reference is [17], used in Related Work as an example of dynamic routing; it is not load-bearing for any derivation. One genuinely concerning gap is that Eq. 4's score function U(Max(ml)||Avg(ml)) does not take the question text as input, so the Sec. 3.2 claim of selection 'in accordance with answer and question' is not implemented; likewise, selecting the best L and K/L per benchmark in Table 3 is a test-set selection issue. These are correctness and evaluation-protocol concerns, not circular reasoning, because no claimed output is identical by construction to an input or fitted quantity. Accordingly, no circular step is identified.
Assumptions & free parameters
free parameters (3)
- L (number of initial event prototypes) =
25 default; tuned to 30-80 on long benchmarks
- K/L (filtered prototype ratio) =
0.8 default; 0.4 for LV-Bench
- C (nearest neighbors in DPC-KNN) =
not reported
assumptions (3)
- domain assumption DPC-KNN cluster centers represent meaningful event prototypes
- standard math Perturbed maximum provides a faithful gradient for top-K selection
- ad hoc to paper A static importance function can identify question-relevant frames
Cite this review
Pith. "Pith review of DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding." pith.science (2026). https://pith.science/paper/4IW5PWH7
@misc{pith2026241112355,
author = {Pith},
title = {Pith review of: DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding},
year = {2026},
howpublished = {\url{https://pith.science/paper/4IW5PWH7}},
note = {Machine review of arXiv:2411.12355}
}
read the original abstract
The challenge in LLM-based video understanding lies in preserving visual and semantic information in long videos while maintaining a memory-affordable token count. However, redundancy and correspondence in videos have hindered the performance potential of existing methods. Through statistical learning on current datasets, we observe that redundancy occurs in both repeated and answer-irrelevant frames, and the corresponding frames vary with different questions. This suggests the possibility of adopting dynamic encoding to balance detailed video information preservation with token budget reduction. To this end, we propose a dynamic cooperative network, DynFocus, for memory-efficient video encoding in this paper. Specifically, i) a Dynamic Event Prototype Estimation (DPE) module to dynamically select meaningful frames for question answering; (ii) a Compact Cooperative Encoding (CCE) module that encodes meaningful frames with detailed visual appearance and the remaining frames with sketchy perception separately. We evaluate our method on five publicly available benchmarks, and experimental results consistently demonstrate that our method achieves competitive performance.
Figures
Figures from the paper (5 more)
Reference graph
Works this paper leans on
-
[1]
Perturbation Techniques in Online Learning and Optimiza- tion, pages 233–264. 2017
work page 2017
-
[2]
Textvqa: Towards understanding of visible and invisible text in images
Aishwarya Agrawal, Xinlei Chen, Stefan Lee, Vignesh Ra- manathan, C Lawrence Zitnick, Devi Parikh, and Dhruv Batra. Textvqa: Towards understanding of visible and invisible text in images. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2019
work page 2019
-
[3]
Kirolos Ataallah, Xiaoqian Shen, Eslam Abdelrahman, Es- sam Sleiman, Deyao Zhu, Jian Ding, and Mohamed Elhoseiny. Minigpt4-video: Advancing multimodal llms for video un- derstanding with interleaved visual-textual tokens. CoRR, abs/2404.03413, 2024
arXiv 2024
-
[4]
Frozen in time: A joint video and image encoder for end-to- end retrieval
Max Bain, Arsha Nagrani, G¨ul Varol, and Andrew Zisserman. Frozen in time: A joint video and image encoder for end-to- end retrieval. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 1728–1738, 2021
work page 2021
-
[5]
Frozen in time: A joint video and image encoder for end-to- end retrieval
Max Bain, Arsha Nagrani, G¨ul Varol, and Andrew Zisserman. Frozen in time: A joint video and image encoder for end-to- end retrieval. In 2021 IEEE/CVF International Conference on Computer Vision, ICCV 2021, Montreal, QC, Canada, October 10-17, 2021, pages 1708–1718. IEEE, 2021
work page 2021
-
[6]
Quentin Berthet, Mathieu Blondel, Olivier Teboul, Marco Cu- turi, Jean-Philippe Vert, and Francis R. Bach. Learning with differentiable perturbed optimizers. CoRR, abs/2002.08676, 2020
arXiv 2002
-
[7]
Activitynet: A large-scale video bench- mark for human activity understanding
Fabian Caba Heilbron, Victor Escorcia, Bernard Ghanem, and Juan Carlos Niebles. Activitynet: A large-scale video bench- mark for human activity understanding. In Proceedings of the ieee conference on computer vision and pattern recognition, pages 961–970, 2015
work page 2015
-
[8]
David L. Chen and William B. Dolan. Collecting highly parallel data for paraphrase evaluation. In Dekang Lin, Yuji Matsumoto, and Rada Mihalcea, editors, The 49th Annual Meeting of the Association for Computational Linguistics: Human Language Technologies, Proceedings of the Confer- ence, 19-24 June, 2011, Portland, Oregon, USA, pages 190–
work page 2011
Show all 80 references
-
[9]
Videollama 2: Advancing spatial-temporal modeling and audio understanding in video- llms
Zesen Cheng, Sicong Leng, Hang Zhang, Yifei Xin, Xin Li, Guanzheng Chen, Yongxin Zhu, Wenqi Zhang, Ziyang Luo, Deli Zhao, and Lidong Bing. Videollama 2: Advancing spatial-temporal modeling and audio understanding in video- llms. CoRR, abs/2406.07476, 2024
2024 arXiv
-
[10]
Gonzalez, Ion Stoica, and Eric P
Wei-Lin Chiang, Zhuohan Li, Zi Lin, Ying Sheng, Zhanghao Wu, Hao Zhang, Lianmin Zheng, Siyuan Zhuang, Yonghao Zhuang, Joseph E. Gonzalez, Ion Stoica, and Eric P. Xing. Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality, March 2023
2023
-
[11]
Palm: Scaling language modeling with pathways
Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, et al. Palm: Scaling language modeling with pathways. Jour- nal of Machine Learning Research, 24(240):1–113, 2023
2023
-
[12]
Wenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong, Junqi Zhao, Weisheng Wang, Boyang Li, Pascale Fung, and Steven C. H. Hoi. Instructblip: Towards general- purpose vision-language models with instruction tuning. In Advances in Neural Information Processing Systems 36...
2023
-
[13]
EgoQA: Egocentric question an- swering
John Doe and Jane Smith. EgoQA: Egocentric question an- swering. In Proceedings of the Conference on Computer Vision and Pattern Recognition (CVPR), 2023
2023
-
[14]
Study on density peaks clustering based on k-nearest neighbors and principal component analysis
Mingjing Du, Shifei Ding, and Hongjie Jia. Study on density peaks clustering based on k-nearest neighbors and principal component analysis. Knowl. Based Syst., 99:135–145, 2016
2016
-
[15]
EV A: exploring the limits of masked visual representation learning at scale
Yuxin Fang, Wen Wang, Binhui Xie, Quan Sun, Ledell Wu, Xinggang Wang, Tiejun Huang, Xinlong Wang, and Yue Cao. EV A: exploring the limits of masked visual representation learning at scale. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2023, Vancouver,...
2023
-
[16]
Making the V in VQA matter: Elevating the role of image understanding in visual question answering
Yash Goyal, Tejas Khot, Aishwarya Agrawal, Douglas Summers-Stay, Dhruv Batra, and Devi Parikh. Making the V in VQA matter: Elevating the role of image understanding in visual question answering. Int. J. Comput. Vis., 127(4):398– 414, 2019
2019
-
[17]
Semantic-aware modular capsule routing for visual question answering
Yudong Han, Jianhua Yin, Jianlong Wu, Yinwei Wei, and Liqiang Nie. Semantic-aware modular capsule routing for visual question answering. IEEE Trans. Image Process. , 32:5537–5549, 2023
2023
-
[18]
MA-LMM: memory-augmented large multimodal model for long-term video understanding
Bo He, Hengduo Li, Young Kyun Jang, Menglin Jia, Xuefei Cao, Ashish Shah, Abhinav Shrivastava, and Ser-Nam Lim. MA-LMM: memory-augmented large multimodal model for long-term video understanding. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2024, Seat...
2024
-
[19]
Ma-lmm: Memory-augmented large multimodal model for long-term video understanding
Bo He, Hengduo Li, Young Kyun Jang, Menglin Jia, Xue- fei Cao, Ashish Shah, Abhinav Shrivastava, and Ser-Nam Lim. Ma-lmm: Memory-augmented large multimodal model for long-term video understanding. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recogni...
2024
-
[20]
Activitynet: A large-scale video bench- mark for human activity understanding
Fabian Caba Heilbron, Victor Escorcia, Bernard Ghanem, and Juan Carlos Niebles. Activitynet: A large-scale video bench- mark for human activity understanding. In IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2015, Boston, MA, USA, June 7-12, 2015 , pages 961...
2015
-
[21]
D. C. Hood and M. A. Finkelstein. Rod and cone contri- butions to human brightness perception, pages 7023–7027. 1986
1986
-
[22]
Vtimellm: Empower llm to grasp video moments
Bin Huang, Xin Wang, Hong Chen, Zihan Song, and Wenwu Zhu. Vtimellm: Empower llm to grasp video moments. arXiv preprint arXiv:2311.18445, 2023
2023 arXiv
-
[23]
Ng, Hongqiang Rong, and Zichen Li
Joshua Zhexue Huang, Michael K. Ng, Hongqiang Rong, and Zichen Li. Automated variable weighting in k-means type clustering. IEEE Trans. Pattern Anal. Mach. Intell. , 27(5):657–668, 2005
2005
-
[24]
Gqa: A new dataset for real-world visual reasoning and compositional questions
Drew Hudson and Christopher D Manning. Gqa: A new dataset for real-world visual reasoning and compositional questions. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2019
2019
-
[25]
Chat-univi: Unified visual representation em- powers large language models with image and video under- standing
Peng Jin, Ryuichi Takanobu, Caiwan Zhang, Xiaochun Cao, and Li Yuan. Chat-univi: Unified visual representation em- powers large language models with image and video under- standing. arXiv preprint arXiv:2311.08046, 2023
2023 arXiv
-
[26]
Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B. Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei. Scaling laws for neural language models. CoRR, abs/2001.08361, 2020
2001 arXiv
-
[27]
Sahar Kazemzadeh, Vicente Ordonez, Mark Matten, and Tamara L. Berg. Referitgame: Referring to objects in pho- tographs of natural scenes. In Alessandro Moschitti, Bo Pang, and Walter Daelemans, editors, Proceedings of the 2014 Conference on Empirical Methods in Natural Languag...
2014
-
[28]
Kingma and Jimmy Ba
Diederik P. Kingma and Jimmy Ba. Adam: A method for stochastic optimization. In Yoshua Bengio and Yann LeCun, editors, 3rd International Conference on Learning Represen- tations, ICLR 2015, San Diego, CA, USA, May 7-9, 2015, Conference Track Proceedings, 2015
2015
-
[29]
Visual genome: Connecting language and vision using crowdsourced dense image annotations
Ranjay Krishna, Yuke Zhu, Oliver Groth, Justin Johnson, Kenji Hata, Joshua Kravitz, Stephanie Chen, Yannis Kalan- tidis, Li-Jia Li, David A Shamma, et al. Visual genome: Connecting language and vision using crowdsourced dense image annotations. International journal of compute...
2017
-
[30]
Visual genome: Connecting language and vision using crowdsourced dense image annotations
Ranjay Krishna, Yuke Zhu, Oliver Groth, Justin Johnson, Kenji Hata, Joshua Kravitz, Stephanie Chen, Yannis Kalan- tidis, Li-Jia Li, David A Shamma, et al. Visual genome: Connecting language and vision using crowdsourced dense image annotations. In International Conference on C...
2017
-
[31]
Videochat: Chat-centric video understanding
KunChang Li, Yinan He, Yi Wang, Yizhuo Li, Wenhai Wang, Ping Luo, Yali Wang, Limin Wang, and Yu Qiao. Videochat: Chat-centric video understanding. arXiv preprint arXiv:2305.06355, 2023
2023 arXiv
-
[32]
Mvbench: A comprehensive multi-modal video understand- ing benchmark
Kunchang Li, Yali Wang, Yinan He, Yizhuo Li, Yi Wang, Yi Liu, Zun Wang, Jilan Xu, Guo Chen, Ping Luo, et al. Mvbench: A comprehensive multi-modal video understand- ing benchmark. arXiv preprint arXiv:2311.17005, 2023
2023 arXiv
-
[33]
Scienceqa: A new dataset for science question answering
Yichuan Li, Jing Huang, Xiaoyu Shen, Yizhou Sun, and Yim- ing Yang. Scienceqa: A new dataset for science question answering. In Proceedings of the 2021 Conference on Em- pirical Methods in Natural Language Processing (EMNLP), 2021
2021
-
[34]
Learning dynamic routing for semantic segmentation
Yanwei Li, Lin Song, Yukang Chen, Zeming Li, Xiangyu Zhang, Xingang Wang, and Jian Sun. Learning dynamic routing for semantic segmentation. In 2020 IEEE/CVF Con- ference on Computer Vision and Pattern Recognition, CVPR 2020, Seattle, WA, USA, June 13-19, 2020, pages 8550–8559....
2020
-
[35]
Llama-vid: An image is worth 2 tokens in large language models
Yanwei Li, Chengyao Wang, and Jiaya Jia. Llama-vid: An image is worth 2 tokens in large language models. arXiv preprint arXiv:2311.17043, 2023
2023 arXiv
-
[36]
Video-llava: Learning united visual representa- tion by alignment before projection
Bin Lin, Yang Ye, Bin Zhu, Jiaxi Cui, Munan Ning, Peng Jin, and Li Yuan. Video-llava: Learning united visual representa- tion by alignment before projection. In Yaser Al-Onaizan, Mo- hit Bansal, and Yun-Nung Chen, editors, Proceedings of the 2024 Conference on Empirical Method...
2024
-
[37]
Video-llava: Learning united visual represen- tation by alignment before projection
Bin Lin, Bin Zhu, Yang Ye, Munan Ning, Peng Jin, and Li Yuan. Video-llava: Learning united visual represen- tation by alignment before projection. arXiv preprint arXiv:2311.10122, 2023
2023 arXiv
-
[38]
Vila: On pre-training for visual language models
Ji Lin, Hongxu Yin, Wei Ping, Yao Lu, Pavlo Molchanov, Andrew Tao, Huizi Mao, Jan Kautz, Mohammad Shoeybi, and Song Han. Vila: On pre-training for visual language models. ArXiv preprint, 2023
2023
-
[39]
Lawrence Zitnick
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Doll´ar, and C. Lawrence Zitnick. Microsoft coco: Common objects in context. In European Conference on Computer Vision, pages 740–755. Springer, 2014
2014
-
[40]
Improved baselines with visual instruction tuning, 2023
Haotian Liu, Chunyuan Li, Yuheng Li, and Yong Jae Lee. Improved baselines with visual instruction tuning, 2023
2023
-
[41]
World model on million-length video and language with blockwise ringattention
Hao Liu, Wilson Yan, Matei Zaharia, and Pieter Abbeel. World model on million-length video and language with blockwise ringattention. ArXiv, abs/2402.08268, 2024
2024 arXiv
-
[42]
One for all: Video conversation is feasible without video instruction tuning
Ruyang Liu, Chen Li, Yixiao Ge, Ying Shan, Thomas H Li, and Ge Li. One for all: Video conversation is feasible without video instruction tuning. arXiv preprint arXiv:2309.15785, 2023
2023 arXiv
-
[43]
ST-LLM: large language models are effective temporal learners
Ruyang Liu, Chen Li, Haoran Tang, Yixiao Ge, Ying Shan, and Ge Li. ST-LLM: large language models are effective temporal learners. CoRR, abs/2404.00308, 2024
2024 arXiv
-
[44]
Learn to explain: Multimodal reasoning via thought chains for science question answering
Pan Lu, Swaroop Mishra, Tony Xia, Liang Qiu, Kai-Wei Chang, Song-Chun Zhu, Oyvind Tafjord, Peter Clark, and Ashwin Kalyan. Learn to explain: Multimodal reasoning via thought chains for science question answering. In The 36th Conference on Neural Information Processing Systems ...
2022
-
[45]
Video-chatgpt: Towards detailed video understanding via large vision and language models
Muhammad Maaz, Hanoona Rasheed, Salman Khan, and Fahad Shahbaz Khan. Video-chatgpt: Towards detailed video understanding via large vision and language models. arXiv preprint arXiv:2306.05424, 2023
2023 arXiv
-
[46]
Some methods for classification and anal- ysis of multivariate observations
James MacQueen. Some methods for classification and anal- ysis of multivariate observations. 1:281–297, 1967
1967
-
[47]
Yuille, and Kevin Murphy
Junhua Mao, Jonathan Huang, Alexander Toshev, Oana Cam- buru, Alan L. Yuille, and Kevin Murphy. Generation and comprehension of unambiguous object descriptions. In 2016 IEEE Conference on Computer Vision and Pattern Recogni- tion, CVPR 2016, Las Vegas, NV , USA, June 27-30, 20...
2016
-
[48]
Ocr-vqa: Visual question answering by reading text in images
Anand Mishra, Shashank Shekhar, Ajeet Kumar Singh, and Anirban Chakraborty. Ocr-vqa: Visual question answering by reading text in images. In 2019 international conference on document analysis and recognition (ICDAR), pages 947–952. IEEE, 2019
2019
-
[49]
Webvidqa: A large-scale dataset for video question answering
Yulei Niu, Zhiyao Ma, Lei Ji, Zhe Lin, Yibing Song, Mingkui Tan, and Ling Shao. Webvidqa: A large-scale dataset for video question answering. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2021
2021
-
[50]
Introducing chatgpt
OpenAI. Introducing chatgpt. 2022
2022
-
[51]
GPT-4o system card, 2024
OpenAI. GPT-4o system card, 2024
2024
-
[52]
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. Learning transferable visual models from natural language supervision. In Proceedings of th...
2021
-
[53]
Improving language understanding by gener- ative pre-training
Alec Radford, Karthik Narasimhan, Tim Salimans, Ilya Sutskever, et al. Improving language understanding by gener- ative pre-training. 2018
2018
-
[54]
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Machel Reid, Nikolay Savinov, Denis Teplyashin, Dmitry Lepikhin, Timothy Lillicrap, Jean-baptiste Alayrac, Radu Soricut, Angeliki Lazaridou, Orhan Firat, Julian Schrittwieser, et al. Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context. ArXiv pre...
2024
-
[55]
Timechat: A time-sensitive multimodal large language model for long video understanding
Shuhuai Ren, Linli Yao, Shicheng Li, Xu Sun, and Lu Hou. Timechat: A time-sensitive multimodal large language model for long video understanding. ArXiv, 2023
2023
-
[56]
Massof Sarah L
Robert W. Massof Sarah L. Pardue. Differential roles of rods and cones in visual acuity and motion detection. Investigative Ophthalmology and Visual Science (IOVS), 2021
2021
-
[57]
A-okvqa: A bench- mark for visual question answering using world knowledge
Dustin Schwenk, Apoorv Khandelwal, Christopher Clark, Kenneth Marino, and Roozbeh Mottaghi. A-okvqa: A bench- mark for visual question answering using world knowledge. In European Conference on Computer Vision, pages 146–162. Springer, 2022
2022
-
[58]
Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning
Piyush Sharma, Nan Ding, Sebastian Goodman, and Radu Soricut. Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning. In Iryna Gurevych and Yusuke Miyao, editors, Proceedings of the 56th Annual Meeting of the Association for Computati...
2018
-
[59]
Textcaps: a dataset for image captioning with reading comprehension
Oleksii Sidorov, Ronghang Hu, Marcus Rohrbach, and Aman- preet Singh. Textcaps: a dataset for image captioning with reading comprehension. In Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part II 16. Springer, 2020
2020
-
[60]
Moviechat: From dense token to sparse memory for long video understanding
Enxin Song, Wenhao Chai, Guanhong Wang, Yucheng Zhang, Haoyang Zhou, Feiyang Wu, Xun Guo, Tian Ye, Yan Lu, Jenq-Neng Hwang, et al. Moviechat: From dense token to sparse memory for long video understanding. arXiv preprint arXiv:2307.16449, 2023
2023 arXiv
-
[61]
Moviechat: From dense token to sparse memory for long video understanding
Enxin Song, Wenhao Chai, Guanhong Wang, Yucheng Zhang, Haoyang Zhou, Feiyang Wu, Xun Guo, Tianbo Ye, Yang Lu, Jenq-Neng Hwang, and Gaoang Wang. Moviechat: From dense token to sparse memory for long video understanding. ArXiv preprint, 2023
2023
-
[62]
Moviellm: Enhancing long video understanding with ai-generated movies
Zhende Song, Chenchen Wang, Jiamu Sheng, Chi Zhang, Gang Yu, Jiayuan Fan, and Tao Chen. Moviellm: Enhancing long video understanding with ai-generated movies. CoRR, abs/2403.01422, 2024
2024 arXiv
-
[63]
Llama: Open and efficient foundation language models.arXiv preprint arXiv:2302.13971, 2023
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Mar- tinet, Marie-Anne Lachaux, Timoth ´ee Lacroix, Baptiste Rozi`ere, Naman Goyal, Eric Hambro, Faisal Azhar, et al. Llama: Open and efficient foundation language models.arXiv preprint arXiv:2302.13971, 2023
2023 arXiv
-
[64]
Ocrvqa: A new dataset for optical character recognition in visual question answering
Yashaswi Verma, Akshay Krishna, Anoop Namboodiri, and C V Jawahar. Ocrvqa: A new dataset for optical character recognition in visual question answering. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2018
2018
-
[65]
Lvbench: An extreme long video understanding benchmark
Weihan Wang, Zehai He, Wenyi Hong, Yean Cheng, Xiaohan Zhang, Ji Qi, Shiyu Huang, Bin Xu, Yuxiao Dong, Ming Ding, and Jie Tang. Lvbench: An extreme long video understanding benchmark. CoRR, abs/2406.08035, 2024
2024 arXiv
-
[66]
Videohallucer: Evaluating intrinsic and ex- trinsic hallucinations in large video-language models
Yuxuan Wang, Yueqian Wang, Dongyan Zhao, Cihang Xie, and Zilong Zheng. Videohallucer: Evaluating intrinsic and ex- trinsic hallucinations in large video-language models. CoRR, abs/2406.16338, 2024
2024 arXiv
-
[67]
Videollamb: Long video understanding with recurrent mem- ory bridges
Yuxuan Wang, Cihang Xie, Yang Liu, and Zilong Zheng. Videollamb: Long video understanding with recurrent mem- ory bridges. arxiv, 2024
2024
-
[68]
Freeva: Offline MLLM as training-free video assistant
Wenhao Wu. Freeva: Offline MLLM as training-free video assistant. CoRR, abs/2405.07798, 2024
2024 arXiv
-
[69]
Davis, Kristen Grauman, and Rog´erio Schmidt Feris
Zuxuan Wu, Tushar Nagarajan, Abhishek Kumar, Steven Ren- nie, Larry S. Davis, Kristen Grauman, and Rog´erio Schmidt Feris. Blockdrop: Dynamic inference paths in residual net- works. In 2018 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2018, Salt Lake City, ...
2018
-
[70]
Deep learning for video classification and captioning
Zuxuan Wu, Ting Yao, Yanwei Fu, and Yu-Gang Jiang. Deep learning for video classification and captioning. In Frontiers of multimedia research, pages 3–29. ACM, 2017
2017
-
[71]
Msr-vtt: A large video description dataset for bridging video and language
Jun Xu, Tao Mei, Ting Yao, and Yong Rui. Msr-vtt: A large video description dataset for bridging video and language. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 5288–5296, 2016
2016
-
[72]
MSR-VTT: A large video description dataset for bridging video and language
Jun Xu, Tao Mei, Ting Yao, and Yong Rui. MSR-VTT: A large video description dataset for bridging video and language. In 2016 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2016, Las Vegas, NV , USA, June 27-30, 2016, pages 5288–5296. IEEE Computer Society, 2016
2016
-
[73]
Pllava: Parameter-free llava extension from images to videos for video dense captioning
Lin Xu, Yilin Zhao, Daquan Zhou, Zhijie Lin, See Kiong Ng, and Jiashi Feng. Pllava: Parameter-free llava extension from images to videos for video dense captioning. arXiv preprint arXiv:2404.16994, 2024
2024 arXiv
-
[74]
Clevrer: Collision events for video representation and reasoning.arXiv preprint arXiv:1910.01442, 2019
Kexin Yi, Chuang Gan, Yunzhu Li, Pushmeet Kohli, Jiajun Wu, Antonio Torralba, and Joshua B Tenenbaum. Clevrer: Collision events for video representation and reasoning.arXiv preprint arXiv:1910.01442, 2019
1910 arXiv
-
[75]
Tenenbaum
Kexin Yi, Chuang Gan, Yunzhu Li, Pushmeet Kohli, Jiajun Wu, Antonio Torralba, and Joshua B. Tenenbaum. CLEVRER: collision events for video representation and reasoning. In 8th International Conference on Learning Representations, ICLR 2020, Addis Ababa, Ethiopia, April 26-30, ...
2020
-
[76]
Video-llama: An instruction-tuned audio-visual language model for video un- derstanding
Hang Zhang, Xin Li, and Lidong Bing. Video-llama: An instruction-tuned audio-visual language model for video un- derstanding. arXiv preprint arXiv:2306.02858, 2023
2023 arXiv
-
[77]
Flash-vstream: Memory- based real-time understanding for long video streams
Haoji Zhang, Yiqin Wang, Yansong Tang, Yong Liu, Jiashi Feng, Jifeng Dai, and Xiaojie Jin. Flash-vstream: Memory- based real-time understanding for long video streams. CoRR, abs/2406.08085, 2024
2024 arXiv
-
[78]
Llama- adapter: Efficient fine-tuning of language models with zero- init attention
Renrui Zhang, Jiaming Han, Aojun Zhou, Xiangfei Hu, Shilin Yan, Pan Lu, Hongsheng Li, Peng Gao, and Yu Qiao. Llama- adapter: Efficient fine-tuning of language models with zero- init attention. In ICLR, 2023
2023
-
[79]
Please Carefully Think
Yuanhan Zhang, Bo Li, haotian Liu, Yong jae Lee, Liangke Gui, Di Fu, Jiashi Feng, Ziwei Liu, and Chunyuan Li. Llava- next: A strong zero-shot video understanding model, 2024. Abstract of Appendix This appendix provides the implementation of redundancy es- timation (Appendix A)...
2024
-
[200]
The Association for Computer Linguistics, 2011
2011
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.