REVIEW 3 major objections 6 minor 1 cited by
A recurrent latent visual cache in the decoder keeps video models grounded during reasoning and shortens their answers.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-12 09:08 UTC pith:63TQLDMJ
load-bearing objection Solid systems paper: recurrent decoder-side visual cache with matched SFT+GRPO gains and shorter answers; the soft spot is annotation-to-inference transfer, not the core idea. the 3 major comments →
Latent Visual Cache for Video Reasoning
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
Inserting a recurrent latent visual cache into the decoder, and training it with supervised contrastive alignment to key frames plus a vision-grounded RL reward that scores key-frame coverage over completion-token hidden states, counters Visual Anchoring Decay. The result is higher accuracy on diverse video reasoning benchmarks—especially grounding-intensive and long-video tasks—while producing substantially shorter answers, because the model keeps visual evidence available rather than relying on longer text.
What carries the argument
Latent Video Cache (Latent-VC): a fixed run of special latent slots whose decoder hidden states form a recurrent chain after the prompt and before the answer. Stage I projects those states into visual space and aligns them to frozen key-frame embeddings with a contrastive loss; Stage II adds a latent grounding reward under GRPO that measures how well free-generation hidden states cover the same key-frame targets, with strict train–inference consistency on native decoder states.
Load-bearing premise
That teaching the cache with annotated key frames and block assignments at training time still leaves a useful visual memory when those annotations disappear at inference.
What would settle it
Ablate Stage-I key-frame contrastive alignment and the Stage-II latent grounding reward, keep the same recurrent slots and remaining training, and measure accuracy plus mid-to-late generation visual-attention mass on long-video and spatial-grounding benchmarks; if performance and attention retention fall to the plain SFT+RL baseline, the visual-supervision path—not the cache slots alone—is carrying the claim.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes Latent-VC, a recurrent latent visual cache inserted into a video LMM decoder to mitigate Visual Anchoring Decay under the read-once, generate-many paradigm. Special <|lvc|> tokens form a recurrent hidden-state chain before answer decoding; Stage I aligns pooled cache states to annotated key-frame visual features via contrastive loss (Eqs. 4–6), and Stage II applies GRPO with answer, format, temporal, and latent-grounding rewards, the last measuring key-frame coverage over completion-token hidden states (Eqs. 7–9, 17–20). Instantiated on Qwen3.5-9B and evaluated on six public video benchmarks at 16/32/64 frames, Latent-VC outperforms CoT and SFT+GRPO baselines, with larger gains on grounding-intensive and long-video settings, while producing substantially shorter responses. Supporting analyses include generation-time visual attention, latent-step ablations, duration/difficulty splits, a 4B-scale transfer experiment, and a qualitative case study.
Significance. If the results hold under tighter controls, the work offers a practical architectural response to a widely observed failure mode in video LMMs: progressive loss of visual grounding during long generation. The combination of a recurrent native-hidden-state cache, train–inference consistency without auxiliary modules at test time, and the accuracy–efficiency finding (higher accuracy with ~50–60% shorter outputs) is of clear interest to multimodal reasoning and long-video understanding. Strengths include multi-benchmark evaluation, attention-based evidence for reduced anchoring decay (Fig. 4a), latent-depth and scale ablations, and a stated public release of code, models, and data. The contribution is primarily empirical/systems rather than theoretical, but the problem framing and efficiency results would be useful to the community if the mechanism is cleanly isolated.
major comments (3)
- [Secs. 2.3–2.4; Eqs. 4–9, 11, 17–20; Table 1] The central inference claim (Secs. 2.1, 2.4) is that the recurrent cache preserves visual evidence without key frames, block assignments, or temporal annotations. Yet Stage I partitions <|lvc|> slots into annotation-dependent blocks I_{i,m} (Eq. 11, Alg. 2) and aligns them to frozen key-frame features (Eqs. 4–6), while Stage II’s r_lat scores coverage of those same annotated targets over completion hidden states (Eqs. 17–20) with w_lat=1.0, and r_tmp also uses key-frame times. Table 1’s SFT+GRPO control shares data/GRPO style but does not isolate (i) recurrent slots without key-frame alignment/latent reward, or (ii) key-frame rewards without the cache. Without at least one of these ablations, it remains unclear whether gains come from a general visual memory versus extra annotation-derived supervision available only in training. This is load-bearing for the transfer premise and should be
- [Table 1; Sec. 3.2] Table 1 states that Latent-VC-9B results are averages over three runs, but no standard deviations or confidence intervals are reported, and the CoT / SFT+GRPO rows do not appear to use the same multi-run protocol. Given that several claimed margins over SFT+GRPO are modest (often ~1–3 points on individual benchmarks), variance is needed to assess whether those gains are stable. Please report run-level statistics for Latent-VC and, ideally, match the evaluation protocol for the main baselines on at least the primary setting (e.g., 64 frames).
- [Sec. 3.1; Appendix C; Table 1] Appendix C specifies Stage-I data as STGR with key frames/boxes for Latent-VC, while Stage II mixes several Open-o3-Video sources. The manuscript does not state with equal precision which SFT corpus, reward channels (especially r_tmp / r_lat), and hyperparameters the Qwen3.5-9B SFT+GRPO baseline uses. If the baseline omits temporal/latent rewards or uses different SFT data, Table 1 confounds architecture with training signal. Please document a fully matched training recipe for SFT+GRPO (data mixture, reward weights, frame budgets, GRPO settings) so the residual gain can be attributed to the cache.
minor comments (6)
- [Abstract; Fig. 1] Figure 1 and the abstract use both “Latent Video Cache (LATENT-VC)” and “Latent Visual Cache”; keep a single expanded name and acronym throughout.
- [Sec. 2.1, Eq. (2)] In Eq. (2), H_{1:S} is defined via Rollout_θ, but the probabilistic factorization is written only over answer tokens; a one-sentence clarification that latent slots are deterministic recurrent states (not sampled vocabulary tokens) would help readers less familiar with latent CoT.
- [Fig. 4(a); Sec. 3.3] Figure 4(a) normalizes attention by peak and aggregates over progress; state explicitly whether this is mean over layers/heads/examples and whether the same decoding temperature/length limits are used for both models.
- [Table 2; Appendix D.2] Table 2 and Appendix Table 4 report large response-length reductions; specify whether length includes special tokens / think tags and whether generation is truncated at a fixed max length for all methods.
- [Sec. 4] Related Work could more explicitly contrast Latent-VC with TVC-style take-along visual conditioning [26] and other latent CoT methods (COCONUT, SoftCoT) on the train–inference consistency axis claimed in Sec. 2.4.
- [Throughout] Minor typos/style: “Visual Anchoring Decay” is sometimes spaced inconsistently; “Prefetcher” capitalization varies; arXiv-style citation [32] dates Qwen3.5 as February 2026—ensure bibliography consistency before camera-ready.
Circularity Check
No circularity: empirical systems paper; benchmark gains are external evaluations, not identities forced by training definitions.
full rationale
Latent-VC is an empirical architecture-and-training paper. The load-bearing claims are measured accuracies on six public video benchmarks (VSI-Bench, VideoMMMU, MMVU, MVBench, TempCompass, VideoMME) against CoT and matched SFT+GRPO baselines under shared frame budgets. Stage I InfoNCE alignment (Eqs. 4–6) and Stage II latent reward (Eqs. 17–20) use key-frame annotations as training supervision only; inference and evaluation use none of that structure (Sec. 2.1, 2.4). Using ground-truth answers, formats, timestamps, and key frames inside a reward is standard supervised/RL practice and does not make reported Acc. equal to the reward by construction—the SFT+GRPO baseline shares the Open-o3-Video mixture and GRPO recipe without the cache and still underperforms. There is no self-definitional identity, no fitted parameter renamed as a prediction of a related quantity, no uniqueness theorem imported from overlapping authors, and no ansatz smuggled in via self-citation that forces the main result. Concerns about whether key-frame-only supervision transfers to free generation are transfer/correctness risks, not circular reductions. Score 0 with empty steps is the honest finding.
Axiom & Free-Parameter Ledger
free parameters (6)
- latent_step_count S
- cache_alignment_weight λ_lvc
- contrastive_temperature τ
- latent_reward_threshold δ
- reward_weights (w_acc, w_fmt, w_tmp, w_lat)
- GRPO clip/KL (ε_ℓ, ε_h, β)
axioms (5)
- domain assumption Visual Anchoring Decay: as autoregressive reasoning lengthens, attention to early video tokens dilutes and models drift toward linguistic priors.
- ad hoc to paper Native decoder hidden states of special <|lvc|> tokens can serve as non-verbal visual memory without removing the raw video prefix.
- domain assumption Key-frame annotations available in Open-o3-Video subsets are valid supervision targets for cache alignment and latent coverage rewards.
- domain assumption Group-relative policy optimization with clipped ratios and optional KL is a valid way to improve multimodal video policies from scalar rewards.
- domain assumption Frozen vision tower + merger features are adequate visual targets for contrastive alignment of decoder cache states.
invented entities (2)
-
Latent Video Cache / Latent Visual Prefetcher (<|lvc_start|>, <|lvc|>, <|lvc_end|> recurrent states)
no independent evidence
-
Latent grounding reward r_lat
no independent evidence
read the original abstract
Video reasoning requires Large Multimodal Models (LMMs) to remain grounded in dense evidence, yet existing systems largely adopt "read-once, generate-many" paradigm, in which visual grounding weakens during generation. This phenomenon has been widely observed and is known as Visual Anchoring Decay. To fill this gap, we introduce Latent Video Cache (Latent-VC), a recurrent latent visual cache inserted into the decoder to preserve compact visual memories throughout reasoning. The cache is trained with supervised contrastive cache alignment and vision-grounded GRPO with a latent grounding reward, while maintaining strict train-inference alignment through native decoder hidden states. Built on Qwen3.5-9B, Latent-VC consistently outperforms strong CoT and SFT+GRPO baselines across six video benchmarks, with especially clear gains on grounding-intensive and long-video tasks. In addition, it also achieves higher accuracy with substantially shorter responses, suggesting that latent visual caching improves video reasoning by preserving visual evidence rather than relying on longer textual chains.
Figures
Forward citations
Cited by 1 Pith paper
-
Thinking in Video: Can Video Generators Really Reason About the Real World?
Video generators show a perception-prediction gap: they can generate plausible continuations while failing explicit visual reasoning tests.
Reference graph
Works this paper leans on
-
[1]
A generalist agent.arXiv preprint arXiv:2205.06175, 2022
Scott Reed, Konrad Zolna, Emilio Parisotto, Sergio Gomez Colmenarejo, Alexander Novikov, Gabriel Barth-Maron, Mai Gimenez, Yury Sulsky, Jackie Kay, Jost Tobias Springenberg, et al. A generalist agent.arXiv preprint arXiv:2205.06175, 2022
Pith/arXiv arXiv 2022
-
[2]
Vision language models in autonomous driving: A survey and outlook.IEEE Transactions on Intelligent Vehicles, 2024
Xingcheng Zhou, Mingyu Liu, Ekim Yurtsever, Bare Luka Zagar, Walter Zimmer, Hu Cao, and Alois C Knoll. Vision language models in autonomous driving: A survey and outlook.IEEE Transactions on Intelligent Vehicles, 2024
2024
-
[3]
Vad: Vectorized scene representation for efficient autonomous driving
Bo Jiang, Shaoyu Chen, Qing Xu, Bencheng Liao, Jiajie Chen, Helong Zhou, Qian Zhang, Wenyu Liu, Chang Huang, and Xinggang Wang. Vad: Vectorized scene representation for efficient autonomous driving. InProceedings of the IEEE/CVF International Conference on Computer Vision, pages 8340–8350, 2023
2023
-
[4]
Video generation models as world simulators
Tim Brooks, Bill Peebles, Connor Holmes, Will DePue, Yufei Guo, Leo Jing, David Schnurr, Joe Taylor, Troy Luhman, Eric Luhman, et al. Video generation models as world simulators. OpenAI Blog, 1(8):1, 2024
2024
-
[5]
Wan: Open and advanced large-scale video generative models, 2025
Team Wan, Ang Wang, Baole Ai, Bin Wen, Chaojie Mao, Chen-Wei Xie, Di Chen, Feiwu Yu, Haiming Zhao, Jianxiao Yang, Jianyuan Zeng, Jiayu Wang, Jingfeng Zhang, Jingren Zhou, Jinkai Wang, Jixuan Chen, Kai Zhu, Kang Zhao, Keyu Yan, Lianghua Huang, Mengyang Feng, Ningyi Zhang, Pandeng Li, Pingyu Wu, Ruihang Chu, Ruili Feng, Shiwei Zhang, Siyang Sun, Tao Fang, T...
2025
-
[6]
The past mistake is the future wisdom: Error-driven contrastive probability optimization for chinese spell checking
Yinghui Li, Qingyu Zhou, Yangning Li, Zhongli Li, Ruiyang Liu, Rongyi Sun, Zizhen Wang, Chao Li, Yunbo Cao, and Hai-Tao Zheng. The past mistake is the future wisdom: Error-driven contrastive probability optimization for chinese spell checking. InFindings of the Association for Computational Linguistics: ACL 2022, pages 3202–3213, 2022
2022
-
[7]
Let’s think with images efficiently! an interleaved-modal chain-of-thought reasoning framework with dynamic and precise visual thoughts
Xu Liu, Yongheng Zhang, Qiguang Chen, Yao Li, Sheng Wang, and Libo Qin. Let’s think with images efficiently! an interleaved-modal chain-of-thought reasoning framework with dynamic and precise visual thoughts. InProceedings of the AAAI Conference on Artificial Intelligence, volume 40, pages 32213–32221, 2026
2026
-
[8]
Youtu-llm: Unlocking the native agentic potential for lightweight large language models, 2026
Junru Lu, Jiarui Qin, Lingfeng Qiao, Yinghui Li, Xinyi Dai, Bo Ke, Jianfeng He, Ruizhi Qiao, Di Yin, Xing Sun, Yunsheng Wu, Yinsong Liu, Shuangyin Liu, Mingkong Tang, Haodong Lin, Jiayi Kuang, Fanxu Meng, Xiaojuan Tang, Yunjia Xi, Junjie Huang, Haotong Yang, Zhenyi Shen, Yangning Li, Qianwen Zhang, Yifei Yu, Siyu An, Junnan Dong, Qiufeng Wang, Jie Wang, K...
2026
-
[9]
Kimi-vl technical report.arXiv preprint arXiv:2504.07491, 2025
Kimi Team, Angang Du, Bohong Yin, Bowei Xing, Bowen Qu, Bowen Wang, Cheng Chen, Chenlin Zhang, Chenzhuang Du, Chu Wei, et al. Kimi-vl technical report.arXiv preprint arXiv:2504.07491, 2025
Pith/arXiv arXiv 2025
-
[10]
Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks
Zhe Chen, Jiannan Wu, Wenhai Wang, Weijie Su, Guo Chen, Sen Xing, Muyan Zhong, Qinglong Zhang, Xizhou Zhu, Lewei Lu, et al. Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 24185–24198, 2024
2024
-
[11]
Openai gpt-5 system card, 2025
Aaditya Singh, Adam Fry, Adam Perelman, Adam Tart, Adi Ganesh, Ahmed El-Kishky, Aidan McLaughlin, Aiden Low, AJ Ostrow, Akhila Ananthram, Akshay Nathan, Alan Luo, Alec Helyar, Aleksander Madry, Aleksandr Efremov, Aleksandra Spyra, Alex Baker-Whitcomb, Alex Beutel, Alex Karpenko, Alex Makelov, Alex Neitz, Alex Wei, Alexandra Barr, Alexandre Kirchmeyer, Ale...
2025
-
[12]
Gheorghe Comanici, Eric Bieber, Mike Schaekermann, Ice Pasupat, Noveen Sachdeva, Inderjit Dhillon, Marcel Blistein, Ori Ram, Dan Zhang, Evan Rosen, et al. Gemini 2.5: Pushing the frontier with advanced reasoning, multimodality, long context, and next generation agentic capabilities.arXiv preprint arXiv:2507.06261, 2025
Pith/arXiv arXiv 2025
-
[13]
Yongheng Zhang, Ziang Liu, Jiaxuan Zhu, Shuai Wang, Xiangqi Chen, Haojing Huang, Jiayi Kuang, Siyu Chen, Ao Shen, Hao Wu, Qiufeng Wang, Qian-Wen Zhang, Junnan Dong, Wenhao Jiang, Ying Shen, Hai-Tao Zheng, Yinghui Li, Di Yin, Xing Sun, and Philip S. Yu. From chatbot to digital colleague: The paradigm shift toward persistent autonomous ai, 2026
2026
-
[14]
Yinghui Li, Jiayi Kuang, Peng Xing, Daixian Liu, Yongheng Zhang, Junnan Dong, Shu-Yu Guo, Yangning Li, Qingyu Zhou, Wenhao Jiang, et al. Cognitive mismatch in multimodal large language models for discrete symbol understanding.arXiv preprint arXiv:2603.18472, 2026
Pith/arXiv arXiv 2026
-
[15]
Video-chatgpt: Towards detailed video understanding via large vision and language models
Muhammad Maaz, Hanoona Rasheed, Salman Khan, and Fahad Khan. Video-chatgpt: Towards detailed video understanding via large vision and language models. InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 12585–12602, 2024
2024
-
[16]
Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis
Chaoyou Fu, Yuhan Dai, Yongdong Luo, Lei Li, Shuhuai Ren, Renrui Zhang, Zihan Wang, Chenyu Zhou, Yunhang Shen, Mengdan Zhang, et al. Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 24108–24118, 2025
2025
-
[17]
Boqiang Zhang, Kehan Li, Zesen Cheng, Zhiqiang Hu, Yuqian Yuan, Guanzheng Chen, Sicong Leng, Yuming Jiang, Hang Zhang, Xin Li, et al. Videollama 3: Frontier multimodal foundation models for image and video understanding.arXiv preprint arXiv:2501.13106, 2025
Pith/arXiv arXiv 2025
-
[18]
Internvl3.5: Advancing open-source multimodal models in versatility, reasoning, and efficiency, 2025
Weiyun Wang, Zhangwei Gao, Lixin Gu, Hengjun Pu, Long Cui, Xingguang Wei, Zhaoyang Liu, Linglin Jing, Shenglong Ye, Jie Shao, Zhaokai Wang, Zhe Chen, Hongjie Zhang, Ganlin Yang, Haomin Wang, Qi Wei, Jinhui Yin, Wenhao Li, Erfei Cui, Guanzhou Chen, Zichen Ding, Changyao Tian, Zhenyu Wu, Jingjing Xie, Zehao Li, Bowen Yang, Yuchen Duan, Xuehui Wang, Zhi Hou,...
2025
-
[19]
AutoCAP: Towards automatic cross-lingual alignment planning for zero-shot chain-of-thought
Yongheng Zhang, Qiguang Chen, Min Li, Wanxiang Che, and Libo Qin. AutoCAP: Towards automatic cross-lingual alignment planning for zero-shot chain-of-thought. In Lun-Wei Ku, Andre Martins, and Vivek Srikumar, editors,Findings of the Association for Computational Linguistics: ACL 2024, pages 9191–9200, Bangkok, Thailand, August 2024. Association for Computa...
2024
-
[20]
Qwen2-vl: Enhancing vision- language model’s perception of the world at any resolution, 2024
Peng Wang, Shuai Bai, Sinan Tan, Shijie Wang, Zhihao Fan, Jinze Bai, Keqin Chen, Xuejing Liu, Jialin Wang, Wenbin Ge, Yang Fan, Kai Dang, Mengfei Du, Xuancheng Ren, Rui Men, Dayiheng Liu, Chang Zhou, Jingren Zhou, and Junyang Lin. Qwen2-vl: Enhancing vision- language model’s perception of the world at any resolution, 2024
2024
-
[21]
Moviechat: From dense token to 11 sparse memory for long video understanding
Enxin Song, Wenhao Chai, Guanhong Wang, Yucheng Zhang, Haoyang Zhou, Feiyang Wu, Haozhe Chi, Xun Guo, Tian Ye, Yanting Zhang, et al. Moviechat: From dense token to 11 sparse memory for long video understanding. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 18221–18232, 2024
2024
-
[22]
Wrong-of-thought: An integrated reasoning framework with multi-perspective verification and wrong information
Yongheng Zhang, Qiguang Chen, Jingxuan Zhou, Peng Wang, Jiasheng Si, Jin Wang, Wenpeng Lu, and Libo Qin. Wrong-of-thought: An integrated reasoning framework with multi-perspective verification and wrong information. In Yaser Al-Onaizan, Mohit Bansal, and Yun-Nung Chen, editors,Findings of the Association for Computational Linguistics: EMNLP 2024, pages 66...
2024
-
[23]
Qwen3-vl technical report, 2025
Shuai Bai, Yuxuan Cai, Ruizhe Chen, Keqin Chen, Xionghui Chen, Zesen Cheng, Lianghao Deng, Wei Ding, Chang Gao, Chunjiang Ge, Wenbin Ge, Zhifang Guo, Qidong Huang, Jie Huang, Fei Huang, Binyuan Hui, Shutong Jiang, Zhaohai Li, Mingsheng Li, Mei Li, Kaixin Li, Zicheng Lin, Junyang Lin, Xuejing Liu, Jiawei Liu, Chenglong Liu, Yang Liu, Dayiheng Liu, Shixuan ...
2025
-
[24]
Gemma 3 technical report, 2025
Gemma Team, Aishwarya Kamath, Johan Ferret, Shreya Pathak, Nino Vieillard, Ramona Merhej, Sarah Perrin, Tatiana Matejovicova, Alexandre Ramé, Morgane Rivière, Louis Rouillard, Thomas Mesnard, Geoffrey Cideron, Jean bastien Grill, Sabela Ramos, Edouard Yvinec, Michelle Casbon, Etienne Pot, Ivo Penchev, Gaël Liu, Francesco Visin, Kathleen Kenealy, Lucas Bey...
2025
-
[25]
CCHall: A novel benchmark for joint cross-lingual and cross-modal hallucinations detection in large language models
Yongheng Zhang, Xu Liu, Ruoxi Zhou, Qiguang Chen, Hao Fei, Wenpeng Lu, and Libo Qin. CCHall: A novel benchmark for joint cross-lingual and cross-modal hallucinations detection in large language models. In Wanxiang Che, Joyce Nabende, Ekaterina Shutova, and Mohammad Taher Pilehvar, editors,Proceedings of the 63rd Annual Meeting of the Association for Compu...
2025
-
[26]
Mitigating visual forgetting via take-along visual conditioning for multi-modal long cot reasoning
Hai-Long Sun, Zhun Sun, Houwen Peng, and Han-Jia Ye. Mitigating visual forgetting via take-along visual conditioning for multi-modal long cot reasoning. InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 5158–5171, 2025
2025
-
[27]
Hao Fang, Jinyu Li, Jiawei Kong, Tianqu Zhuang, Kuofeng Gao, Bin Chen, Shu-Tao Xia, and Yaowei Wang. Seeing through the chain: Mitigate hallucination in multimodal reasoning models via cot compression and contrastive preference optimization.arXiv preprint arXiv:2602.03380, 2026
arXiv 2026
-
[28]
Yufeng Du, Minyang Tian, Srikanth Ronanki, Subendhu Rongali, Sravan Bodapati, Aram Galstyan, Azton Wells, Roy Schwartz, Eliu A Huerta, and Hao Peng. Context length alone hurts llm performance despite perfect retrieval.arXiv preprint arXiv:2510.05381, 2025
arXiv 2025
-
[29]
Visual hallucinations of multi-modal large language models
Wen Huang, Hongbin Liu, Minxin Guo, and Neil Gong. Visual hallucinations of multi-modal large language models. InFindings of the Association for Computational Linguistics: ACL 2024, pages 9614–9631, 2024
2024
-
[30]
Vigc: Visual instruction generation and correction
Bin Wang, Fan Wu, Xiao Han, Jiahui Peng, Huaping Zhong, Pan Zhang, Xiaoyi Dong, Weijia Li, Wei Li, Jiaqi Wang, et al. Vigc: Visual instruction generation and correction. InProceedings of the AAAI Conference on Artificial Intelligence, volume 38, pages 5309–5317, 2024
2024
-
[31]
Hourvideo: 1-hour video- language understanding.Advances in Neural Information Processing Systems, 37:53168–53197, 2024
Keshigeyan Chandrasegaran, Agrim Gupta, Lea M Hadzic, Taran Kota, Jimming He, Cristóbal Eyzaguirre, Zane Durante, Manling Li, Jiajun Wu, and Li Fei-Fei. Hourvideo: 1-hour video- language understanding.Advances in Neural Information Processing Systems, 37:53168–53197, 2024
2024
-
[32]
Qwen3.5: Towards native multimodal agents, February 2026
Qwen Team. Qwen3.5: Towards native multimodal agents, February 2026. 12
2026
-
[33]
Tempcompass: Do video llms really understand videos?arXiv preprint arXiv:2403.00476, 2024
Yuanxin Liu, Shicheng Li, Yi Liu, Yuxiang Wang, Shuhuai Ren, Lei Li, Sishuo Chen, Xu Sun, and Lu Hou. Tempcompass: Do video llms really understand videos?arXiv preprint arXiv:2403.00476, 2024
Pith/arXiv arXiv 2024
-
[34]
Mvbench: A comprehensive multi-modal video understanding benchmark
Kunchang Li, Yali Wang, Yinan He, Yizhuo Li, Yi Wang, Yi Liu, Zun Wang, Jilan Xu, Guo Chen, Ping Luo, Limin Wang, and Yu Qiao. Mvbench: A comprehensive multi-modal video understanding benchmark. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024
2024
-
[35]
Yilun Zhao, Lujing Xie, Haowei Zhang, Guo Gan, Yitao Long, Zhiyuan Hu, Tongyan Hu, Weiyuan Chen, Chuhan Li, Junyang Song, Zhijian Xu, Chengye Wang, Weifeng Pan, Ziyao Shangguan, Xiangru Tang, Zhenwen Liang, Yixin Liu, Chen Zhao, and Arman Cohan. Mmvu: Measuring expert-level multi-discipline video understanding.arXiv preprint arXiv:2501.12380, 2025
Pith/arXiv arXiv 2025
-
[36]
Video-mmmu: Evaluating knowledge acquisition from multi-discipline professional videos
Kairui Hu, Penghao Wu, Fanyi Pu, Wang Xiao, Yuanhan Zhang, Xiang Yue, Bo Li, and Ziwei Liu. Video-mmmu: Evaluating knowledge acquisition from multi-discipline professional videos. arXiv preprint arXiv:2501.13826, 2025
Pith/arXiv arXiv 2025
-
[37]
Gupta, Rilyn Han, Li Fei-Fei, and Saining Xie
Jihan Yang, Shusheng Yang, Anjali W. Gupta, Rilyn Han, Li Fei-Fei, and Saining Xie. Thinking in space: How multimodal large language models see, remember, and recall spaces.arXiv preprint arXiv:2412.14171, 2024
Pith/arXiv arXiv 2024
-
[38]
Zhihong Shao, Peiyi Wang, Qihao Zhu, Runxin Xu, Junxiao Song, Xiao Bi, Haowei Zhang, Mingchuan Zhang, YK Li, Yang Wu, et al. Deepseekmath: Pushing the limits of mathematical reasoning in open language models.arXiv preprint arXiv:2402.03300, 2024
Pith/arXiv arXiv 2024
-
[39]
Jiahao Meng, Xiangtai Li, Haochen Wang, Yue Tan, Tao Zhang, Lingdong Kong, Yunhai Tong, Anran Wang, Zhiyang Teng, Yujing Wang, and Zhuochen Wang. Open-o3-video: Grounded video reasoning with explicit spatio-temporal evidence.arXiv preprint arXiv:2510.20579, 2025
arXiv 2025
-
[40]
Video-r1: Reinforcing video reasoning in mllms.arXiv preprint arXiv:2503.21776, 2025
Kaituo Feng, Kaixiong Gong, Bohao Li, Zonghao Guo, Yibing Wang, Tianshuo Peng, Junfei Wu, Xiaoying Zhang, Benyou Wang, and Xiangyu Yue. Video-r1: Reinforcing video reasoning in mllms.arXiv preprint arXiv:2503.21776, 2025
Pith/arXiv arXiv 2025
-
[41]
Llama-vid: An image is worth 2 tokens in large language models.arXiv preprint arXiv:2311.17043, 2023
Yanwei Li, Chengyao Wang, and Jiaya Jia. Llama-vid: An image is worth 2 tokens in large language models.arXiv preprint arXiv:2311.17043, 2023
Pith/arXiv arXiv 2023
-
[42]
Zesen Cheng, Sicong Leng, Hang Zhang, Yifei Xin, Xin Li, Guanzheng Chen, Yongxin Zhu, Wenqi Zhang, Ziyang Luo, Deli Zhao, and Lidong Bing. Videollama 2: Advancing spatial- temporal modeling and audio understanding in video-llms.arXiv preprint arXiv:2406.07476, 2024
Pith/arXiv arXiv 2024
-
[43]
Long context transfer from language to vision
Peiyuan Zhang, Kaichen Zhang, Bo Li, Guangtao Zeng, Jingkang Yang, Yuanhan Zhang, Ziyue Wang, Haoran Tan, Chunyuan Li, and Ziwei Liu. Long context transfer from language to vision. arXiv preprint arXiv:2406.16852, 2024
Pith/arXiv arXiv 2024
-
[44]
Vila: On pre-training for visual language models
Ji Lin, Hongxu Yin, Wei Ping, Yao Lu, Pavlo Molchanov, Andrew Tao, Huizi Mao, Jan Kautz, Mohammad Shoeybi, and Song Han. Vila: On pre-training for visual language models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024
2024
-
[45]
Unhackable temporal rewarding for scalable video mllms.arXiv preprint arXiv:2502.12081, 2025
En Yu, Kangheng Lin, Liang Zhao, Yana Wei, Zining Zhu, Haoran Wei, Jianjian Sun, Zheng Ge, Xiangyu Zhang, Jingyu Wang, and Wenbing Tao. Unhackable temporal rewarding for scalable video mllms.arXiv preprint arXiv:2502.12081, 2025
Pith/arXiv arXiv 2025
-
[46]
Llava-onevision: Easy visual task transfer
Bo Li, Yuanhan Zhang, Dong Guo, Renrui Zhang, Feng Li, Hao Zhang, Kaichen Zhang, Peiyuan Zhang, Yanwei Li, Ziwei Liu, and Chunyuan Li. Llava-onevision: Easy visual task transfer. arXiv preprint arXiv:2408.03326, 2024
Pith/arXiv arXiv 2024
-
[47]
Jiajun Liu, Yibing Wang, Hanghang Ma, Xiaoping Wu, Xiaoqi Ma, Xiaoming Wei, Jianbin Jiao, Enhua Wu, and Jie Hu. Kangaroo: A powerful video-language model supporting long-context video input.arXiv preprint arXiv:2408.15542, 2024. 13
Pith/arXiv arXiv 2024
-
[48]
Chain-of-thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al. Chain-of-thought prompting elicits reasoning in large language models. Advances in neural information processing systems, 35:24824–24837, 2022
2022
-
[49]
Training language models to follow instructions with human feedback.Advances in neural information processing systems, 35:27730–27744, 2022
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. Training language models to follow instructions with human feedback.Advances in neural information processing systems, 35:27730–27744, 2022
2022
-
[50]
Video-llava: Learning united visual representation by alignment before projection
Bin Lin, Yang Ye, Bin Zhu, Jiaxi Cui, Munan Ning, Peng Jin, and Li Yuan. Video-llava: Learning united visual representation by alignment before projection. InProceedings of the 2024 conference on empirical methods in natural language processing, pages 5971–5984, 2024
2024
-
[51]
Gemini Team, Rohan Anil, Sebastian Borgeaud, Jean-Baptiste Alayrac, Jiahui Yu, Radu Soricut, Johan Schalkwyk, Andrew M. Dai, Anja Hauth, Katie Millican, David Silver, Melvin Johnson, Ioannis Antonoglou, Julian Schrittwieser, Amelia Glaese, Jilin Chen, Emily Pitler, Timothy Lillicrap, Angeliki Lazaridou, Orhan Firat, and James Molloy et al. Gemini: A famil...
2025
-
[52]
Thinking with videos: Multimodal tool-augmented reinforcement learning for long video reasoning
Haoji Zhang, Xin Gu, Jiawen Li, Chixiang Ma, Sule Bai, Chubin Zhang, Bowen Zhang, Zhichao Zhou, Dongliang He, and Yansong Tang. Thinking with videos: Multimodal tool-augmented reinforcement learning for long video reasoning. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 32903–32914, 2026
2026
-
[53]
Video-of-thought: step-by-step video reasoning from perception to cognition
Hao Fei, Shengqiong Wu, Wei Ji, Hanwang Zhang, Meishan Zhang, Mong Li Lee, and Wynne Hsu. Video-of-thought: step-by-step video reasoning from perception to cognition. InPro- ceedings of the 41st International Conference on Machine Learning, pages 13109–13125, 2024
2024
-
[54]
Vitcot: Video-text interleaved chain-of-thought for boosting video understanding in large language models
Yongheng Zhang, Xu Liu, Ruihan Tao, Qiguang Chen, Hao Fei, Wanxiang Che, and Libo Qin. Vitcot: Video-text interleaved chain-of-thought for boosting video understanding in large language models. InProceedings of the 33rd ACM International Conference on Multimedia, pages 5267–5276, 2025
2025
-
[55]
Shibo Hao, Sainbayar Sukhbaatar, DiJia Su, Xian Li, Zhiting Hu, Jason Weston, and Yuandong Tian. Training large language models to reason in a continuous latent space.arXiv preprint arXiv:2412.06769, 2024
Pith/arXiv arXiv 2024
-
[56]
SoftCoT: Soft chain-of-thought for efficient reasoning with LLMs
Yige Xu, Xu Guo, Zhiwei Zeng, and Chunyan Miao. SoftCoT: Soft chain-of-thought for efficient reasoning with LLMs. InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 23336–23351, 2025
2025
-
[57]
Hybrid latent reasoning via reinforcement learning.Advances in Neural Information Processing Systems, 38:5501–5530, 2026
Zhenrui Yue, Bowen Jin, Huimin Zeng, Honglei Zhuang, Zhen Qin, Jinsung Yoon, Lanyu Shang, Jiawei Han, and Dong Wang. Hybrid latent reasoning via reinforcement learning.Advances in Neural Information Processing Systems, 38:5501–5530, 2026
2026
-
[58]
Xinghao Chen, Anhao Zhao, Heming Xia, Xuan Lu, Hanlin Wang, Yanjun Chen, Wei Zhang, Jian Wang, Wenjie Li, and Xiaoyu Shen. Reasoning beyond language: A comprehensive survey on latent chain-of-thought reasoning.arXiv preprint arXiv:2505.16782, 2025
arXiv 2025
-
[59]
Monet: Reasoning in latent visual space beyond image and language
Qixun Wang, Yang Shi, Yifei Wang, Yuanxing Zhang, Pengfei Wan, Kun Gai, Xianghua Ying, and Yisen Wang. Monet: Reasoning in latent visual space beyond image and language. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 12030–12040, 2026
2026
-
[60]
Scaling up test-time compute with latent reasoning: A recurrent depth approach.Advances in Neural Information Processing Systems, 38:41340–41391, 2026
Jonas Geiping, Sean McLeish, Neel Jain, John Kirchenbauer, Siddharth Singh, Brian Bartoldson, Bhavya Kailkhura, Abhinav Bhatele, and Tom Goldstein. Scaling up test-time compute with latent reasoning: A recurrent depth approach.Advances in Neural Information Processing Systems, 38:41340–41391, 2026. 14 A Appendix Overview This appendix documents implementa...
2026
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.