REVIEW 4 major objections 5 minor 61 references
SportsGrounder: Proposal-Aided Interleaved Grounding for Dense Sports Video Reasoning
T0 review · 4 major / 5 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read A 2B model with interleaved object grounding and action supervision sets the top accuracy on dense sports video QA.
desk verdict The IGF architecture is a solid, well-specified subfield contribution, but the AAS module's claimed anti-language-bias mechanism is undercut by a likely shortcut, and the SOTA claim rests on an unpublished self-curated benchmark. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the Interleaved Grounding Fusion (IGF) mechanism, which builds the visual prefix as per-frame blocks of global grid tokens followed by a compact hybrid token per selected object: the object's normalized box coordinates tokenized through the LLM text embedding plus its OV-DINO semantic feature projected by an MLP. Because selection is restricted to the top 15 domain-relevant objects, sequence length stays bounded relative to naive concatenation of hundreds of detector queries. The Action-Aware Supervision (AAS) branch is the corrective mechanism: it maps the terminal hidden state to an action label through a linear head and back-propagates a masked cross-entropy loss only for action queries, preventing the model from treating the final state as pure language. Mixed Preference Optimization (MPO) is the decision-boundary calibrator: it optimizes preferred versus dispreferred answer sets, with DPO-style preference, quality, and generation losses, following the base model's training protocol.
What would settle it
Construct a text-only control by running the same QA pairs with the video frames removed or replaced by static gray frames while keeping the question text identical; if a text-only model approaches or matches the reported accuracy, the benchmark leaks non-visual cues. A second check: swap in a different action clip with identical player tracks and see whether answers track the visual change; if accuracy does not move, the model is still using language priors.
Extended reading notes
Core claim
On its own terms, the paper's discovery is that interleaved, frame-aligned fusion of explicit box coordinates and implicit object semantics with global ViT features lets a small multimodal model reason about dense sports scenes accurately, while injecting boxes as text does not. Action-Aware Supervision forces the end-of-sequence hidden state to carry action information, countering language-prior guessing. Mixed Preference Optimization sharpens choice among hard-negative distractors. The result is the top accuracy on the paper's SoccerNet and FineSports QA benchmarks, including the fine-grained action, spatial reasoning, and dense disambiguation subtasks, with the action and spatial gains coming from the visual mechanism rather than from model scale.
Load-bearing premise
The load-bearing premise is that the programmatically generated multiple-choice questions from SoccerNet and FineSports genuinely test fine-grained visual reasoning; if the questions leak the answer through wording, template, or distractor patterns, the reported accuracy gains would not reflect visual understanding.
Editorial extensions
If this is right
- A 2B-parameter LMM with this grounding architecture outperforms a 4B model on dense sports reasoning, suggesting parameter count is not the binding constraint for fine-grained video QA.
- Feeding detector boxes to the model as text consistently hurts action accuracy; interleaving them as visual tokens preserves temporal attention and gives the better trade-off.
- The AAS loss alone adds roughly 2.4 percentage points on action accuracy over the IGF-only baseline, so motion supervision is a separable, transferable component.
- The MPO stage contributes most to spatial reasoning and team-comparison subtasks, so preference optimization is where relational accuracy is won.
- Because IGF avoids sequence-length explosion, the architecture can be applied to longer videos and more frames per clip without quadratic attention blow-up.
Reading between the lines
- The paper does not test for answer leakage in its generated QA pairs; a text-only or frame-scrambled control would establish whether the visual components, rather than question templates, drive the reported margins.
- A natural next step the paper leaves implicit is applying IGF to other dense multi-agent scenes, such as traffic or surgical video, where small homogeneous subjects and fine actions dominate.
- Because the Jersey Number subtask remains the one place a larger model wins, combining the object branch with higher-resolution crops or an OCR prober is a plausible extension the paper does not explore.
- Deploying the method on raw broadcasts would need an action taxonomy or pseudo-labels, since AAS supervision depends on dataset action metadata.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. SportsGrounder augments InternVL3.5-2B with OV-DINO object proposals, fuses them frame-by-frame through Interleaved Grounding Fusion (IGF), adds an Action-Aware Supervision (AAS) loss on the EOS hidden state, and trains in three stages (representation alignment, LoRA-based SFT, and Mixed Preference Optimization). The authors programmatically generate 26k/24k four-option QA pairs from SoccerNet and FineSports and report 51.8% and 53.6% overall accuracy, outperforming LoRA-finetuned baselines including MiniCPM-V 4B.
Significance. If verified, this is a useful architectural contribution to dense sports video QA: the IGF mechanism is a clean solution to the sequence-length explosion and temporal misalignment problems that arise when injecting object-level groundings, and the use of a frozen open-vocabulary detector keeps the added cost modest. The paper is also clearly specified: the three-stage curriculum, losses, and fusion equations are explicit, and the ablation design is reasonable. However, the headline claim depends on a self-curated, unreleased benchmark, single-run results without uncertainty, and an AAS mechanism whose stated mechanism is not established by the experiments as written. These are load-bearing gaps, but they are fixable in revision.
major comments (4)
- [§4.1 (Datasets and QA Construction)] The central state-of-the-art claim is evaluated only on QA pairs the authors generated from SoccerNet and FineSports metadata, and the resulting benchmark is not released. Because the test set is not public and the generation protocol is described only at a high level, the reader cannot verify that the questions measure fine-grained visual reasoning rather than metadata-derived textual regularities. Please release the data and generation scripts, or additionally evaluate on existing public benchmarks such as Sports-QA and SportU, and specify the exact templates, distractor sampling, and filtering used.
- [§4.1 (Implementation Details) and Table 2] Top-K = 15, the action-supervision weight λ = 0.1, and the MPO weights are chosen 'empirically', but no validation split is described and all results appear to come from a single run. The key ablation gains (e.g., +2.4% Action for AAS and +4.4% Spatial for MPO) therefore have no stated uncertainty. Please report means and standard deviations across multiple seeds and specify the validation set used for hyperparameter selection.
- [§3.3 (AAS, Eq. 7)] The AAS loss is computed from h_EOS after the full answer sequence has been generated, and the ground-truth answer text almost always contains the action label. Since causal attention at the terminal EOS position can attend to the preceding answer tokens, the action head can minimize Eq. 7 by reading the label from the text, so the claim that AAS 'forces' visual motion aggregation is not established. The paper itself acknowledges that causal attention can heavily weigh adjacent textual inputs at terminal tokens, but does not describe a stop-gradient, a masking of answer tokens, or an alternative placement of h_EOS before answer generation. Please clarify the exact token position used, exclude answer-token gradients if necessary, and add a text-only control (e.g., answers with visual input ablated) to show that the +2.4% Action gain in Table 2 reflects visual grounding rather than text copying.
- [§3.4 and §4.1 (MPO Stage 3)] The construction of the preference pairs for Mixed Preference Optimization is not described: how are the positive responses y_c and hard-negative responses y_r generated for the four-option QA tasks, and what exactly is the frozen reference model π_ref? Without this information, the Stage 3 gain in Table 2 is not reproducible, and the claim that MPO 'distinguishes deceptive distractors' cannot be separated from the specific choice of negatives. Please specify the negative sampling procedure and its relation to the distractor options in the generated QA pairs.
minor comments (5)
- [Abstract and §1] The wording that AAS 'forces the network to learn accurate motion representations' is stronger than what a λ = 0.1 auxiliary loss can guarantee; consider softer phrasing such as 'encourages'.
- [§3.2, Eq. (5)] The claimed efficiency benefit depends on the actual sequence length, but the values of T, N, and L_box are not reported for the two datasets; please include the average token counts and compute cost.
- [Table 1] Several baseline differences are small (e.g., 45.2 vs. 43.8 on SoccerNet), and no significance tests or confidence intervals are reported; with single runs, those differences may not be meaningful.
- [Figure 4] The options for Task 3 Action Anticipation are printed in a non-alphabetical order (Fallback, Pass, Other, ToBasket), which is confusing in a multiple-choice question figure.
- [References [41] and [42]] Please confirm that the MPO protocol described as 'natively introduced by our base model, InternVL3.5' is indeed the same as the MPO in reference [41], since that reference predates InternVL3.5 and may correspond to an earlier InternVL version.
Circularity Check
AAS computes its action loss from h_EOS after the answer sequence, so the +2.4% Action gain in Table 2 may be a text-reading artifact rather than visual motion grounding; the self-generated QA benchmark makes the SOTA claim self-referential in scope.
-
self definitional
[Section 3.3 (Action-Aware Supervision), Eqs. (6)-(7); Section 3.4 Stage 2, Eqs. (9)-(10); Table 2.]
"After the causal forward pass, we obtain the sequence of hidden states E∈R^{L×d_llm}. We extract the final hidden state h_EOS corresponding to the end-of-sequence token. Given that standard attention mechanisms propagate prior context forward, h_EOS encapsulates the global semantic aggregation of the entire video-text input. ... While causal attention mechanisms can naturally heavily weigh the adjacent textual inputs at the terminal tokens, the explicit backpropagation of L_act directly through h_EOS mitigates this tendency."
Stage 2 optimizes Eq. (9) over Y_ans={y_1,...,y_{L_ans}} in the same causal forward pass that produces E and h_EOS, so as written h_EOS is the hidden state after the ground-truth answer tokens, not a pre-answer sentinel. The answer tokens identify the correct multiple-choice option (and often contain the action label), so h_EOS can carry the action category into the AAS head of Eq. (6), minimizing Eq. (7) with little or no use of H_vis. The paper's own remark that causal attention can naturally heavily weigh adjacent textual inputs at terminal tokens names the exact shortcut, yet no mask, stop-gradient, or answer-token exclusion is specified. Thus the +2.4% Action gain attributed to AAS in Table 2 is confounded as evidence for visual motion grounding by the given construction.
full rationale
The main circular reduction is in the AAS module. L_act is computed from h_EOS, which is obtained from the same causal forward pass used to score the ground-truth answer sequence in Eq. (9); nothing in the paper detaches the answer tokens from the AAS pathway. Because the answer tokens identify the correct choice and action label, the auxiliary head can satisfy Eq. (7) from the textual context, so the claimed visual-motion regularizer does not necessarily ground in video. The paper's own statement about heavy weighting of adjacent textual inputs names the exact shortcut. The IGF construction, MPO recipe, and baseline comparisons are not circular in themselves; the evaluation is on QA pairs programmatically generated by the authors from SoccerNet and FineSports annotations, which makes the 'state-of-the-art' claim self-referential in scope (no external benchmark), though within that protocol the comparisons are fair. Since the central architectural claim does not fully reduce to a fit and IGF/MPO remain independently testable, the circularity is partial rather than total.
Assumptions & free parameters
free parameters (5)
- Top-K entity selection =
15
- Action supervision weight lambda =
0.1
- Number of sampled frames T =
not specified
- MPO weights alpha and beta =
not specified
- DPO divergence margin beta =
0.1
assumptions (4)
- domain assumption Meta-annotations (action labels, bounding boxes, jersey numbers) from SoccerNet and FineSports are accurate and sufficient for generating valid QA pairs and AAS supervision.
- domain assumption The programmatic QA generation does not introduce answer-leaking textual biases.
- domain assumption OV-DINO's proposals reliably detect relevant small, homogeneous entities in broadcast sports frames.
- ad hoc to paper Supervising the EOS hidden state with action labels forces visual motion aggregation rather than reward text-based shortcuts.
Cite this review
Pith. "Pith review of SportsGrounder: Proposal-Aided Interleaved Grounding for Dense Sports Video Reasoning." pith.science (2026). https://pith.science/paper/ZVDYSQ5K
@misc{pith2026260807932,
author = {Pith},
title = {Pith review of: SportsGrounder: Proposal-Aided Interleaved Grounding for Dense Sports Video Reasoning},
year = {2026},
howpublished = {\url{https://pith.science/paper/ZVDYSQ5K}},
note = {Machine review of arXiv:2608.07932}
}
read the original abstract
Sports video analysis is crucial for athletic analytics and broadcasting enhancement. Dense sports video reasoning, however, demands a fine-grained understanding of numerous small-scale, highly interactive, and visually homogeneous entities (e.g., players sharing identical uniforms, the ball) across long temporal contexts. Current Large Multimodal Models (LMMs) inherently struggle with such dense visual complexities. Due to the lack of fine-grained visual details, these models often over-rely on textual priors to guess answers, especially when distinguishing visually similar actions and players. To address this, we propose \textbf{SportsGrounder}, a framework that leverages an open-vocabulary visual expert to aid interleaved grounding specifically for dense sports video reasoning. To achieve precise spatial localization, we extract domain-guided object proposals and introduce an Interleaved Grounding Fusion (IGF) mechanism. The IGF frame-by-frame integrates explicit bounding box coordinates and implicit visual semantics with global grid features. This design preserves strict temporal alignment and prevents sequence length explosion. Furthermore, we design an Action-Aware Supervision (AAS) module that directly regularizes the model's hidden states, forcing the network to learn accurate motion representations rather than relying on language bias. Optimized with Mixed Preference Optimization (MPO) to better distinguish deceptive distractors, our extensive experiments on newly curated dense sports VQA datasets (derived from SoccerNet and FineSports) demonstrate that SportsGrounder significantly improves fine-grained reasoning and achieves state-of-the-art accuracy.
Figures
Reference graph
Works this paper leans on
-
[1]
Shuai Bai, Yuxuan Cai, Ruizhe Chen, Keqin Chen, Xionghui Chen, Zesen Cheng, Lianghao Deng, Wei Ding, Chang Gao, Chunjiang Ge, et al . 2025. Qwen3-vl technical report.arXiv preprint arXiv:2511.21631(2025)
arXiv 2025
-
[2]
Haodong Chen, Haojian Huang, Xinxiang Yin, and Dian Shao. 2025. FineQuest: Adaptive Knowledge-Assisted Sports Video Understanding via Agent-of- Thoughts Reasoning. InProceedings of the 33rd ACM International Conference on Multimedia
work page 2025
-
[3]
Anthony Cioppa, Adrien Deliège, Silvio Giancola, Bernard Ghanem, and Marc Van Droogenbroeck. 2022. Scaling up SoccerNet with multi-view spatial localiza- tion and re-identification.Scientific Data9, 1 (2022), 355
work page 2022
-
[4]
Andong Deng, Tongjia Chen, Shoubin Yu, Taojiannan Yang, Lincoln Spencer, Yapeng Tian, Ajmal Saeed Mian, Mohit Bansal, and Chen Chen. 2025. Motion- Grounded Video Reasoning: Understanding and Perceiving Motion at Pixel Level. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recog- nition (CVPR). 8625–8636
work page 2025
-
[5]
Kaituo Feng, Kaixiong Gong, Bohao Li, Zonghao Guo, Yibing Wang, Tianshuo Peng, Junfei Wu, Xiaoying Zhang, Benyou Wang, and Xiangyu Yue. 2025. Video- R1: Reinforcing Video Reasoning in MLLMs. InAdvances in Neural Information Processing Systems (NeurIPS)
work page 2025
-
[6]
Chaoyou Fu, Yuhan Dai, Yongdong Luo, Lei Li, Shuhuai Ren, Renrui Zhang, Zihan Wang, Chenyu Zhou, Yunhang Shen, Mengdan Zhang, Peixian Chen, Yanwei Li, Shaohui Lin, Sirui Zhao, Ke Li, Tong Xu, Xiawu Zheng, Enhong Chen, Caifeng Shan, Ran He, and Xing Sun. 2025. Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysi...
work page 2025
-
[7]
Silvio Giancola, Anthony Cioppa, Marc Gutiérrez-Pérez, Jan Held, Carlos Hi- nojosa, Victor Joos, Arnaud Leduc, Floriane Magera, Karen Sanchez, Vladimir Somers, et al . 2025. SoccerNet 2025 Challenges Results.arXiv preprint arXiv:2508.19182(2025)
arXiv 2025
-
[8]
Sitong Gong, Yunzhi Zhuge, Lu Zhang, Zongxin Yang, Pingping Zhang, and Huchuan Lu. 2025. The Devil is in Temporal Token: High Quality Video Reasoning Segmentation. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 29183–29192
work page 2025
Show all 61 references
-
[9]
Taneesh Gupta, Rahul Madhavan, Xuchao Zhang, Nagarajan Natarajan, Chetan Bansal, and Saravan Rajmohan. 2026. Multi-Preference Optimization: Generaliz- ing DPO via Set-Level Contrasts. InProceedings of the International Conference on Learning Representations (ICLR)
2026
-
[10]
Jan Held, Hani Itani, Anthony Cioppa, Silvio Giancola, Bernard Ghanem, and Marc Van Droogenbroeck. 2024. X-VARS: Introducing Explainability in Foot- ball Refereeing with Multi-Modal Large Language Models. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern ...
2024
-
[11]
Wenyi Hong, Yean Cheng, Zhuoyi Yang, Weihan Wang, Lefan Wang, Xiaotao Gu, Shiyu Huang, Yuxiao Dong, and Jie Tang. 2025. MotionBench: Benchmarking and Improving Fine-grained Video Motion Understanding for Vision Language Models. InProceedings of the IEEE/CVF Conference on Compu...
2025
-
[12]
Fraser, Anahita Bhiwandiwalla, and Svetlana Kir- itchenko
Phillip Howard, Kathleen C. Fraser, Anahita Bhiwandiwalla, and Svetlana Kir- itchenko. 2025. Uncovering Bias in Large Vision-Language Models at Scale with Counterfactuals. InProceedings of the Conference of the North American Chapter of the Association for Computational Lingui...
2025
-
[13]
Zhenpeng Huang, Jiaqi Li, Zihan Jia, Xinhao Li, Desen Meng, Lingxue Song, Xi Chen, Liang Li, and Limin Wang. 2025. LongVPO: From Anchored Cues to Self-Reasoning for Long-Form Video Preference Optimization. InAdvances in Neural Information Processing Systems (NeurIPS)
2025
-
[14]
Minchan Kwon, Hyounguk Shon, and Junmo Kim. 2026. Learning Question- Aware Keyframe Selection with Synthetic Supervision for Video Question An- swering.arXiv preprint arXiv:2603.14953(2026)
2026
-
[15]
Hongyu Li, Jinyu Chen, Ziyu Wei, Shaofei Huang, Tianrui Hui, Jialin Gao, Xi- aoming Wei, and Si Liu. 2025. LLaVA-ST: A Multimodal Large Language Model for Fine-Grained Spatial-Temporal Understanding. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recog...
2025
-
[16]
Haopeng Li, Andong Deng, Qiuhong Ke, Jun Liu, Hossein Rahmani, Yulan Guo, Bernt Schiele, and Chen Chen. 2024. Sports-QA: A Large-Scale Video Question Answering Benchmark for Complex and Professional Sports.arXiv preprint arXiv:2401.01505(2024)
2024
-
[17]
Yixuan Li, Changli Tang, Jimin Zhuang, Yudong Yang, Guangzhi Sun, Wei Li, Zejun Ma, and Chao Zhang. 2025. Improving LLM Video Understanding with 16 Frames Per Second. InProceedings of the 42nd International Conference on Machine Learning (ICML), Vol. 267. 35942–35956
2025
-
[18]
Zeqian Li, Shangzhe Di, Zhonghua Zhai, Weilin Huang, Yanfeng Wang, and Weidi Xie. 2025. Universal Video Temporal Grounding with Generative Multi-modal Large Language Models. InAdvances in Neural Information Processing Systems (NeurIPS)
2025
-
[19]
Zhaohe Liao, Jiangtong Li, Siyu Sun, Qingyang Liu, Fengshun Xiao, Tianjiao Li, Qiang Zhang, Guang Chen, Li Niu, Changjun Jiang, and Liqing Zhang. 2025. Divide and Conquer: Exploring Language-centric Tree Reasoning for Video Question-Answering. InProceedings of the 42nd Interna...
2025
-
[20]
Xiongkun Linghu, Jiangyong Huang, Ziyu Zhu, Baoxiong Jia, and Siyuan Huang
-
[21]
Yudin, Maxim Monastyrny, and Aleksei Valenkov
Sergey Linok, Tatiana Zemskova, Svetlana Ladanova, Roman Titkov, Dmitry A. Yudin, Maxim Monastyrny, and Aleksei Valenkov. 2025. Beyond Bare Queries: Open-Vocabulary Object Grounding with 3D Scene Graph. InProceedings of the IEEE International Conference on Robotics and Automat...
2025
-
[22]
Zhaoyu Liu, Kan Jiang, Murong Ma, Zhe Hou, Yun Lin, and Jin Song Dong. 2025. F3Set: Towards Analyzing Fast, Frequent, and Fine-grained Events from Videos. InProceedings of the International Conference on Learning Representations (ICLR)
2025
-
[23]
Olga Loginova, Oleksandr Bezrukov, Ravi Shekhar, and Alexey Kravets. 2025. Addressing Blind Guessing: Calibration of Selection Bias in Multiple-Choice Question Answering by Video Language Models. InProceedings of the 63rd Annual Meeting of the Association for Computational Lin...
2025
-
[24]
Ziyang Luo, Haoning Wu, Dongxu Li, Jing Ma, Mohan Kankanhalli, and Junnan Li. 2025. VideoAutoArena: An Automated Arena for Evaluating Large Multimodal Models in Video Analysis through User Simulation. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Reco...
2025
-
[25]
Xing, Fahad Shah- baz Khan, and Salman Khan
Shehan Munasinghe, Hanan Gani, Wenqi Zhu, Jiale Cao, Eric P. Xing, Fahad Shah- baz Khan, and Salman Khan. 2025. VideoGLaMM: A Large Multimodal Model for Pixel-Level Visual Grounding in Videos. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition ...
2025
-
[26]
Trong-Thuan Nguyen, Pha Nguyen, Jackson Cothren, Alper Yilmaz, and Khoa Luu. 2025. HyperGLM: HyperGraph for Video Scene Graph Generation and Anticipation. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 29150–29160
2025
-
[27]
Ming Nie, Dan Ding, Chunwei Wang, Yuanfan Guo, Jianhua Han, Hang Xu, and Li Zhang. 2024. SlowFocus: Enhancing Fine-grained Temporal Understanding in Video LLM. InAdvances in Neural Information Processing Systems (NeurIPS), Vol. 37
2024
-
[28]
Liangyang Ouyang, Ruicong Liu, Yifei Huang, Ryosuke Furuta, and Yoichi Sato
-
[29]
J. Park, K. J. Jang, B. Alasaly, S. Mopidevi, A. Zolensky, E. Eaton, I. Lee, and K. Johnson. 2025. Assessing Modality Bias in Video Question Answering Bench- marks with Multimodal Large Language Models. InProceedings of the AAAI Conference on Artificial Intelligence, Vol. 39. ...
2025
-
[30]
Manning, Stefano Ermon, and Chelsea Finn
Rafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning, Stefano Ermon, and Chelsea Finn. 2023. Direct Preference Optimization: Your Language Model is Secretly a Reward Model. InAdvances in Neural Information Processing Systems (NeurIPS), Vol. 36. 53728–53741
2023
-
[31]
Jiayuan Rao, Zifeng Li, Haoning Wu, Ya Zhang, Yanfeng Wang, and Weidi Xie
-
[32]
Mohammadreza Salehi, Jae Sung Park, Tanush Yadav, Aditya Kusupati, Ranjay Krishna, Yejin Choi, Hannaneh Hajishirzi, and Ali Farhadi. 2024. ActionAtlas: A VideoQA Benchmark for Domain-specialized Action Recognition. InAdvances in Neural Information Processing Systems (NeurIPS),...
2024
-
[33]
Xi Tang, Jihao Qiu, Lingxi Xie, Yunjie Tian, Jianbin Jiao, and Qixiang Ye. 2025. Adaptive Keyframe Sampling for Long Video Understanding. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 29118–29128
2025
-
[34]
Shengbang Tong, Ellis Brown, Penghao Wu, Sanghyun Woo, Manoj Middepogu, Sai Charitha Akula, Jihan Yang, Shusheng Yang, Adithya Iyer, Xichen Pan, Austin Wang, Rob Fergus, Yann LeCun, and Saining Xie. 2024. Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs....
2024
-
[35]
2026.𝜙-DPO: Fairness Direct Preference Optimization Approach to Con- tinual Learning in Large Multimodal Models
Thanh-Dat Truong, Huu-Thien Tran, Jackson Cothren, Bhiksha Raj, and Khoa Luu. 2026.𝜙-DPO: Fairness Direct Preference Optimization Approach to Con- tinual Learning in Large Multimodal Models. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
2026
-
[36]
Sethuraman T V, Savya Khosla, et al. 2026. Stress Tests REVEAL Fragile Temporal and Visual Grounding in Video-Language Models.arXiv preprint arXiv:2602.11244 MM ’26, November 10–14, 2026, Rio de Janeiro, Brazil Yizhi Li, Jiawei Jiang, Guanhong Wang, Yingcai Wu, and Gaoang Wang (2026)
2026
-
[37]
Haochen Wang, Qirui Chen, Cilin Yan, Jiayin Cai, Xiaolong Jiang, Yao Hu, Weidi Xie, and Stratis Gavves. 2025. Object-centric Video Question Answering with Visual Grounding and Referring. InProceedings of the IEEE/CVF International Conference on Computer Vision (ICCV). 22274–22284
2025
-
[38]
Hao Wang, Pengzhen Ren, Zequn Jie, Xiao Dong, Chengjian Feng, Yinlong Qian, Lin Ma, Dongmei Jiang, Yaowei Wang, Xiangyuan Lan, and Xiaodan Liang. 2024. OV-DINO: Unified Open-Vocabulary Detection with Language-Aware Selective Fusion.arXiv preprint arXiv:2407.07844(2024)
2024 arXiv
-
[39]
Haibo Wang, Zhiyang Xu, Yu Cheng, Shizhe Diao, Yufan Zhou, Yixin Cao, Qifan Wang, Weifeng Ge, and Lifu Huang. 2025. Grounded-VideoLLM: Sharpening Fine-grained Temporal Grounding in Video Large Language Models. InFindings of the Association for Computational Linguistics: EMNLP
2025
-
[40]
Junjie Wang, Zhihao Yuan, Shuyi Jiang, Jiaqi Mao, Chun-Mei Feng, Shuguang Cui, Na Zhao, and Zhen Li. 2026. Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations. InProceedings of the International Conference on Learning Representations (ICLR)
2026
-
[41]
Weiyun Wang, Zhe Chen, Wenhai Wang, Yue Cao, Yangzhou Liu, Zhangwei Gao, Jinguo Zhu, Xizhou Zhu, Lewei Lu, Yu Qiao, and Jifeng Dai. 2024. Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization.arXiv preprint arXiv:2411.10442(2024)
2024 arXiv
-
[42]
Weiyun Wang, Lixin Gu, Hengjun Pu, Long Cui, Xingguang Wei, Zhaoyang Liu, Linglin Jing, Shenglong Ye, Jie Shao, et al . 2025. InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency.arXiv preprint arXiv:2508.18265(2025)
2025 arXiv
-
[43]
Yuxuan Wang, Yiqi Song, Cihang Xie, Yang Liu, and Zilong Zheng. 2025. VideoL- LaMB: Long Streaming Video Understanding with Recurrent Memory Bridges. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV). 24170–24181
2025
-
[44]
Syed Talal Wasim, Muzammal Naseer, Salman Khan, Ming-Hsuan Yang, and Fahad Shahbaz Khan. 2024. Video-GroundingDINO: Towards Open-Vocabulary Spatio-Temporal Video Grounding. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 18909–18918
2024
-
[45]
Haotian Xia, Haonan Ge, Junbo Zou, Hyun Woo Choi, Xuebin Zhang, Danny Suradja, Botao Rui, Ethan Tran, Wendy Jin, Zhen Ye, Xiyang Lin, Christopher Lai, Shengjie Zhang, Junwen Miao, Shichao Chen, Rhys Tracy, Vicente Ordonez, Weining Shen, and Hanjie Chen. 2026. SportR: A Benchma...
2026
-
[46]
Haotian Xia, Zhengbang Yang, Junbo Zou, Rhys Tracy, Yuqing Wang, Chi Lu, Christopher Lai, Yanjun He, Xun Shao, Zhuoqing Xie, Yuan fang Wang, Weining Shen, and Hanjie Chen. 2025. SPORTU: A Comprehensive Sports Understand- ing Benchmark for Multimodal Large Language Models. InPr...
2025
-
[47]
Jinglin Xu, Guohao Zhao, Sibo Yin, Wenhao Zhou, and Yuxin Peng. 2024. FineS- ports: A Multi-person Hierarchical Sports Video Dataset for Fine-grained Action Understanding. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 21773–21782
2024
-
[48]
Haolin Yang, Jiayuan Rao, Haoning Wu, and Weidi Xie. 2026. SoccerMaster: A Vision Foundation Model for Soccer Understanding. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
2026
-
[49]
Gupta, Rilyn Han, Li Fei-Fei, and Saining Xie
Jihan Yang, Shusheng Yang, Anjali W. Gupta, Rilyn Han, Li Fei-Fei, and Saining Xie. 2025. Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 10632–10643
2025
-
[50]
Rui Yang, Qianhui Wu, Zhaoyang Wang, Hanyang Chen, Ke Yang, Hao Cheng, Huaxiu Yao, Baolin Peng, Huan Zhang, Jianfeng Gao, and Tong Zhang. 2026. GUI-Libra: Training Native GUI Agents to Reason and Act with Action-aware Supervision and Partially Verifiable RL.arXiv preprint arXi...
2026 arXiv
-
[51]
Xianming Yang, Qi Li, Chengdong Qian, Haitao Wang, Yonghui Wu, and Wei Wang. 2026. Bias Correction and Explainability Framework for Large Language Models: A Knowledge-Driven Approach.Big Data and Cognitive Computing (2026)
2026
-
[52]
Zuhao Yang, Yingchen Yu, Yunqing Zhao, Shijian Lu, and Song Bai. 2025. TimeEx- pert: An Expert-Guided Video LLM for Video Temporal Grounding. InProceedings of the IEEE/CVF International Conference on Computer Vision (ICCV). 24286–24296
2025
-
[53]
Yuan Yao, Tianyu Yu, Ao Zhang, Chongyi Wang, Junbo Cui, Hongji Zhu, Tianchi Cai, Haoyu Li, Weilin Zhao, Zhihui He, et al. 2025. MiniCPM-V: A GPT-4V Level MLLM on Your Phone.Nat Commun 16, 5509 (2025)(2025)
2025
-
[54]
Jinhui Ye, Zihan Wang, Haosen Sun, Keshigeyan Chandrasegaran, Zane Durante, Cristobal Eyzaguirre, Yonatan Bisk, Juan Carlos Niebles, Ehsan Adeli, Li Fei-Fei, Jiajun Wu, and Manling Li. 2025. Re-thinking Temporal Search for Long-Form Video Understanding. InProceedings of the IE...
2025
-
[55]
Boqiang Zhang, Kehan Li, Zesen Cheng, Zhiqiang Hu, Yuqian Yuan, Guanzheng Chen, Sicong Leng, Yuming Jiang, Hang Zhang, Xin Li, et al. 2025. Videollama 3: Frontier multimodal foundation models for image and video understanding. arXiv preprint arXiv:2501.13106(2025)
2025 arXiv
-
[56]
Hauptmann, Yonatan Bisk, and Yiming Yang
Ruohong Zhang, Liangke Gui, Zhiqing Sun, Yihao Feng, Keyang Xu, Yuanhan Zhang, Di Fu, Chunyuan Li, Alexander G. Hauptmann, Yonatan Bisk, and Yiming Yang. 2025. Direct Preference Optimization of Video Large Multimodal Models from Language Model Reward. InProceedings of the Conf...
2025
-
[57]
Hanyu Zhou and Gim Hee Lee. 2026. LLaVA-4D: A General Large Multimodal Model for 4D Scene Understanding. InProceedings of the International Conference on Learning Representations (ICLR)
2026
-
[58]
Orr Zohar, Xiaohan Wang, Yann Dubois, Nikhil Mehta, Tong Xiao, Philippe Hansen-Estruch, Licheng Yu, Xiaofang Wang, Felix Juefei-Xu, Ning Zhang, Ser- ena Yeung-Levy, and Xide Xia. 2025. Apollo: An Exploration of Video Under- standing in Large Multimodal Models. InProceedings of...
2025
-
[2024]
InProceed- ings of the European Conference on Computer Vision (ECCV)
ActionVOS: Actions as Prompts for Video Object Segmentation. InProceed- ings of the European Conference on Computer Vision (ECCV). 216–235
-
[2025]
InProceed- ings of the ACM Multimedia Conference (ACM MM)
Multi-Agent System for Comprehensive Soccer Understanding. InProceed- ings of the ACM Multimedia Conference (ACM MM)
-
[2026]
InProceedings of the International Conference on Learning Representations (ICLR)
SceneCOT: Eliciting Grounded Chain-of-Thought Reasoning in 3D Scenes. InProceedings of the International Conference on Learning Representations (ICLR)
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.