Pith. sign in

REVIEW 3 major objections 6 minor 1 cited by

A recurrent latent visual cache in the decoder keeps video models grounded during reasoning and shortens their answers.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-12 09:08 UTC pith:63TQLDMJ

load-bearing objection Solid systems paper: recurrent decoder-side visual cache with matched SFT+GRPO gains and shorter answers; the soft spot is annotation-to-inference transfer, not the core idea. the 3 major comments →

arxiv 2607.02607 v1 pith:63TQLDMJ submitted 2026-07-01 cs.CV cs.CL

Latent Visual Cache for Video Reasoning

classification cs.CV cs.CL
keywords video reasoninglatent visual cachevisual anchoring decaymultimodal large language modelsvisual groundinglong video understandingGRPOlatent reasoning
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Video multimodal models usually read all the frames once and then generate a long answer. As that answer grows, their attention to the original visual tokens fades—a failure the authors call Visual Anchoring Decay, which produces hallucinations and detached reasoning. This paper claims the fix is a compact latent visual cache: special non-verbal tokens whose decoder hidden states form a recurrent memory between the prompt and the answer. The cache is trained in two stages so those hidden states stay tied to key visual moments—first by contrastive alignment to key-frame embeddings, then by a grounding reward during group-relative reinforcement learning—while using the model’s native states so training and inference match. On six video benchmarks the method beats strong chain-of-thought and supervised-plus-RL baselines, with the largest gains on spatial grounding and long videos, and it does so with much shorter responses. The authors take that as evidence that preserving visual evidence in latent form improves video reasoning more than lengthening textual chains.

Core claim

Inserting a recurrent latent visual cache into the decoder, and training it with supervised contrastive alignment to key frames plus a vision-grounded RL reward that scores key-frame coverage over completion-token hidden states, counters Visual Anchoring Decay. The result is higher accuracy on diverse video reasoning benchmarks—especially grounding-intensive and long-video tasks—while producing substantially shorter answers, because the model keeps visual evidence available rather than relying on longer text.

What carries the argument

Latent Video Cache (Latent-VC): a fixed run of special latent slots whose decoder hidden states form a recurrent chain after the prompt and before the answer. Stage I projects those states into visual space and aligns them to frozen key-frame embeddings with a contrastive loss; Stage II adds a latent grounding reward under GRPO that measures how well free-generation hidden states cover the same key-frame targets, with strict train–inference consistency on native decoder states.

Load-bearing premise

That teaching the cache with annotated key frames and block assignments at training time still leaves a useful visual memory when those annotations disappear at inference.

What would settle it

Ablate Stage-I key-frame contrastive alignment and the Stage-II latent grounding reward, keep the same recurrent slots and remaining training, and measure accuracy plus mid-to-late generation visual-attention mass on long-video and spatial-grounding benchmarks; if performance and attention retention fall to the plain SFT+RL baseline, the visual-supervision path—not the cache slots alone—is carrying the claim.

Watch this falsifier — get emailed when new claim-graph text bears on it.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. The paper proposes Latent-VC, a recurrent latent visual cache inserted into a video LMM decoder to mitigate Visual Anchoring Decay under the read-once, generate-many paradigm. Special <|lvc|> tokens form a recurrent hidden-state chain before answer decoding; Stage I aligns pooled cache states to annotated key-frame visual features via contrastive loss (Eqs. 4–6), and Stage II applies GRPO with answer, format, temporal, and latent-grounding rewards, the last measuring key-frame coverage over completion-token hidden states (Eqs. 7–9, 17–20). Instantiated on Qwen3.5-9B and evaluated on six public video benchmarks at 16/32/64 frames, Latent-VC outperforms CoT and SFT+GRPO baselines, with larger gains on grounding-intensive and long-video settings, while producing substantially shorter responses. Supporting analyses include generation-time visual attention, latent-step ablations, duration/difficulty splits, a 4B-scale transfer experiment, and a qualitative case study.

Significance. If the results hold under tighter controls, the work offers a practical architectural response to a widely observed failure mode in video LMMs: progressive loss of visual grounding during long generation. The combination of a recurrent native-hidden-state cache, train–inference consistency without auxiliary modules at test time, and the accuracy–efficiency finding (higher accuracy with ~50–60% shorter outputs) is of clear interest to multimodal reasoning and long-video understanding. Strengths include multi-benchmark evaluation, attention-based evidence for reduced anchoring decay (Fig. 4a), latent-depth and scale ablations, and a stated public release of code, models, and data. The contribution is primarily empirical/systems rather than theoretical, but the problem framing and efficiency results would be useful to the community if the mechanism is cleanly isolated.

major comments (3)
  1. [Secs. 2.3–2.4; Eqs. 4–9, 11, 17–20; Table 1] The central inference claim (Secs. 2.1, 2.4) is that the recurrent cache preserves visual evidence without key frames, block assignments, or temporal annotations. Yet Stage I partitions <|lvc|> slots into annotation-dependent blocks I_{i,m} (Eq. 11, Alg. 2) and aligns them to frozen key-frame features (Eqs. 4–6), while Stage II’s r_lat scores coverage of those same annotated targets over completion hidden states (Eqs. 17–20) with w_lat=1.0, and r_tmp also uses key-frame times. Table 1’s SFT+GRPO control shares data/GRPO style but does not isolate (i) recurrent slots without key-frame alignment/latent reward, or (ii) key-frame rewards without the cache. Without at least one of these ablations, it remains unclear whether gains come from a general visual memory versus extra annotation-derived supervision available only in training. This is load-bearing for the transfer premise and should be
  2. [Table 1; Sec. 3.2] Table 1 states that Latent-VC-9B results are averages over three runs, but no standard deviations or confidence intervals are reported, and the CoT / SFT+GRPO rows do not appear to use the same multi-run protocol. Given that several claimed margins over SFT+GRPO are modest (often ~1–3 points on individual benchmarks), variance is needed to assess whether those gains are stable. Please report run-level statistics for Latent-VC and, ideally, match the evaluation protocol for the main baselines on at least the primary setting (e.g., 64 frames).
  3. [Sec. 3.1; Appendix C; Table 1] Appendix C specifies Stage-I data as STGR with key frames/boxes for Latent-VC, while Stage II mixes several Open-o3-Video sources. The manuscript does not state with equal precision which SFT corpus, reward channels (especially r_tmp / r_lat), and hyperparameters the Qwen3.5-9B SFT+GRPO baseline uses. If the baseline omits temporal/latent rewards or uses different SFT data, Table 1 confounds architecture with training signal. Please document a fully matched training recipe for SFT+GRPO (data mixture, reward weights, frame budgets, GRPO settings) so the residual gain can be attributed to the cache.
minor comments (6)
  1. [Abstract; Fig. 1] Figure 1 and the abstract use both “Latent Video Cache (LATENT-VC)” and “Latent Visual Cache”; keep a single expanded name and acronym throughout.
  2. [Sec. 2.1, Eq. (2)] In Eq. (2), H_{1:S} is defined via Rollout_θ, but the probabilistic factorization is written only over answer tokens; a one-sentence clarification that latent slots are deterministic recurrent states (not sampled vocabulary tokens) would help readers less familiar with latent CoT.
  3. [Fig. 4(a); Sec. 3.3] Figure 4(a) normalizes attention by peak and aggregates over progress; state explicitly whether this is mean over layers/heads/examples and whether the same decoding temperature/length limits are used for both models.
  4. [Table 2; Appendix D.2] Table 2 and Appendix Table 4 report large response-length reductions; specify whether length includes special tokens / think tags and whether generation is truncated at a fixed max length for all methods.
  5. [Sec. 4] Related Work could more explicitly contrast Latent-VC with TVC-style take-along visual conditioning [26] and other latent CoT methods (COCONUT, SoftCoT) on the train–inference consistency axis claimed in Sec. 2.4.
  6. [Throughout] Minor typos/style: “Visual Anchoring Decay” is sometimes spaced inconsistently; “Prefetcher” capitalization varies; arXiv-style citation [32] dates Qwen3.5 as February 2026—ensure bibliography consistency before camera-ready.

Circularity Check

0 steps flagged

No circularity: empirical systems paper; benchmark gains are external evaluations, not identities forced by training definitions.

full rationale

Latent-VC is an empirical architecture-and-training paper. The load-bearing claims are measured accuracies on six public video benchmarks (VSI-Bench, VideoMMMU, MMVU, MVBench, TempCompass, VideoMME) against CoT and matched SFT+GRPO baselines under shared frame budgets. Stage I InfoNCE alignment (Eqs. 4–6) and Stage II latent reward (Eqs. 17–20) use key-frame annotations as training supervision only; inference and evaluation use none of that structure (Sec. 2.1, 2.4). Using ground-truth answers, formats, timestamps, and key frames inside a reward is standard supervised/RL practice and does not make reported Acc. equal to the reward by construction—the SFT+GRPO baseline shares the Open-o3-Video mixture and GRPO recipe without the cache and still underperforms. There is no self-definitional identity, no fitted parameter renamed as a prediction of a related quantity, no uniqueness theorem imported from overlapping authors, and no ansatz smuggled in via self-citation that forces the main result. Concerns about whether key-frame-only supervision transfers to free generation are transfer/correctness risks, not circular reductions. Score 0 with empty steps is the honest finding.

Axiom & Free-Parameter Ledger

6 free parameters · 5 axioms · 2 invented entities

The central claim rests on standard transformer/LMM machinery, the empirical premise that visual anchoring decays under long generation, and several hand-chosen training knobs plus a new latent-cache construct whose usefulness is demonstrated only within this training recipe. No physical constants; free parameters are optimization and reward hyperparameters. Invented entities are architectural (cache tokens/states and the latent grounding reward), not new physical objects.

free parameters (6)
  • latent_step_count S
    Number of recurrent <|lvc|> slots; set to 8 at inference after a VSI-Bench sweep (0–64). Performance peaks at 8 and falls at larger S, so the reported gains depend on this choice.
  • cache_alignment_weight λ_lvc
    Balances CE and contrastive alignment in Stage I; fixed at 0.1 without a broad sensitivity study in the main text.
  • contrastive_temperature τ
    InfoNCE temperature for key-frame–cache alignment; set to 0.07.
  • latent_reward_threshold δ
    Threshold in the piecewise latent-reward shaping ψ(s;δ); set to 0.2 and directly controls positive vs negative grounding reward.
  • reward_weights (w_acc, w_fmt, w_tmp, w_lat)
    Set to 2.0, 0.5, 0.5, 1.0; these weights define the scalar GRPO objective and therefore which trajectories are preferred.
  • GRPO clip/KL (ε_ℓ, ε_h, β)
    Clipping margins 0.2 and KL weight β=0.04 control policy update size relative to the Stage-I reference.
axioms (5)
  • domain assumption Visual Anchoring Decay: as autoregressive reasoning lengthens, attention to early video tokens dilutes and models drift toward linguistic priors.
    Stated in §1 and used as the problem definition; supported by citations and the paper’s own attention plots, but treated as the primary failure mode to fix.
  • ad hoc to paper Native decoder hidden states of special <|lvc|> tokens can serve as non-verbal visual memory without removing the raw video prefix.
    Core interface in §2.1–2.2 (Eqs. 2–3); architectural postulate validated only by end-task gains and attention analysis.
  • domain assumption Key-frame annotations available in Open-o3-Video subsets are valid supervision targets for cache alignment and latent coverage rewards.
    Stage I/II depend on annotated key frames and times (Appendix C); inference claims no need for them, so transfer is assumed.
  • domain assumption Group-relative policy optimization with clipped ratios and optional KL is a valid way to improve multimodal video policies from scalar rewards.
    Stage II follows DeepSeekMath GRPO [38]; standard in recent reasoning RL, not re-derived here.
  • domain assumption Frozen vision tower + merger features are adequate visual targets for contrastive alignment of decoder cache states.
    Eq. 4 freezes Ev features as vi,m; assumes that space is the right grounding target for reasoning-time hidden states.
invented entities (2)
  • Latent Video Cache / Latent Visual Prefetcher (<|lvc_start|>, <|lvc|>, <|lvc_end|> recurrent states) no independent evidence
    purpose: Compact recurrent visual memory inside the decoder to counteract Visual Anchoring Decay during answer generation.
    New architectural object defined in §2; independent evidence is only the paper’s benchmarks and attention maps, not an external measurement of cache contents.
  • Latent grounding reward r_lat no independent evidence
    purpose: Trajectory-level RL signal measuring whether completion-token hidden states cover annotated key-frame targets after projection.
    Defined in §2.3 and Appendix B.3 (Eqs. 17–20); exists only as a training objective component in this work.

pith-pipeline@v1.1.0-grok45 · 27622 in / 3782 out tokens · 34961 ms · 2026-07-12T09:08:01.728353+00:00 · methodology

0 comments
read the original abstract

Video reasoning requires Large Multimodal Models (LMMs) to remain grounded in dense evidence, yet existing systems largely adopt "read-once, generate-many" paradigm, in which visual grounding weakens during generation. This phenomenon has been widely observed and is known as Visual Anchoring Decay. To fill this gap, we introduce Latent Video Cache (Latent-VC), a recurrent latent visual cache inserted into the decoder to preserve compact visual memories throughout reasoning. The cache is trained with supervised contrastive cache alignment and vision-grounded GRPO with a latent grounding reward, while maintaining strict train-inference alignment through native decoder hidden states. Built on Qwen3.5-9B, Latent-VC consistently outperforms strong CoT and SFT+GRPO baselines across six video benchmarks, with especially clear gains on grounding-intensive and long-video tasks. In addition, it also achieves higher accuracy with substantially shorter responses, suggesting that latent visual caching improves video reasoning by preserving visual evidence rather than relying on longer textual chains.

Figures

Figures reproduced from arXiv: 2607.02607 by Di Yin, Hao Wu, Philip S. Yu, Xing Sun, Yinghui Li, Yongheng Zhang, Zhipeng Xu.

Figure 1
Figure 1. Figure 1: Comparison of long-video reasoning paradigms. (a) In the read-once paradigm, visual [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: Overall framework of Latent Video Cache. The method consists of two main components: [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Performance comparison between Qwen3.5-4B and [PITH_FULL_IMAGE:figures/full_fig_p007_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: Experiments on (a) visual anchoring dynamics during autoregressive video reasoning and [PITH_FULL_IMAGE:figures/full_fig_p007_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: Experiments on (a) Performance across different video durations and (b) The performance [PITH_FULL_IMAGE:figures/full_fig_p008_5.png] view at source ↗
Figure 6
Figure 6. Figure 6: A case study in which Qwen3.5-9B and Qwen3.5-9B [PITH_FULL_IMAGE:figures/full_fig_p009_6.png] view at source ↗
Figure 7
Figure 7. Figure 7: Head-level visual attention maps for Qwen3.5-9B and [PITH_FULL_IMAGE:figures/full_fig_p020_7.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Thinking in Video: Can Video Generators Really Reason About the Real World?

    cs.CV 2026-07 conditional novelty 6.0

    Video generators show a perception-prediction gap: they can generate plausible continuations while failing explicit visual reasoning tests.

Reference graph

Works this paper leans on

60 extracted references · 18 linked inside Pith · cited by 1 Pith paper

  1. [1]

    A generalist agent.arXiv preprint arXiv:2205.06175, 2022

    Scott Reed, Konrad Zolna, Emilio Parisotto, Sergio Gomez Colmenarejo, Alexander Novikov, Gabriel Barth-Maron, Mai Gimenez, Yury Sulsky, Jackie Kay, Jost Tobias Springenberg, et al. A generalist agent.arXiv preprint arXiv:2205.06175, 2022

  2. [2]

    Vision language models in autonomous driving: A survey and outlook.IEEE Transactions on Intelligent Vehicles, 2024

    Xingcheng Zhou, Mingyu Liu, Ekim Yurtsever, Bare Luka Zagar, Walter Zimmer, Hu Cao, and Alois C Knoll. Vision language models in autonomous driving: A survey and outlook.IEEE Transactions on Intelligent Vehicles, 2024

  3. [3]

    Vad: Vectorized scene representation for efficient autonomous driving

    Bo Jiang, Shaoyu Chen, Qing Xu, Bencheng Liao, Jiajie Chen, Helong Zhou, Qian Zhang, Wenyu Liu, Chang Huang, and Xinggang Wang. Vad: Vectorized scene representation for efficient autonomous driving. InProceedings of the IEEE/CVF International Conference on Computer Vision, pages 8340–8350, 2023

  4. [4]

    Video generation models as world simulators

    Tim Brooks, Bill Peebles, Connor Holmes, Will DePue, Yufei Guo, Leo Jing, David Schnurr, Joe Taylor, Troy Luhman, Eric Luhman, et al. Video generation models as world simulators. OpenAI Blog, 1(8):1, 2024

  5. [5]

    Wan: Open and advanced large-scale video generative models, 2025

    Team Wan, Ang Wang, Baole Ai, Bin Wen, Chaojie Mao, Chen-Wei Xie, Di Chen, Feiwu Yu, Haiming Zhao, Jianxiao Yang, Jianyuan Zeng, Jiayu Wang, Jingfeng Zhang, Jingren Zhou, Jinkai Wang, Jixuan Chen, Kai Zhu, Kang Zhao, Keyu Yan, Lianghua Huang, Mengyang Feng, Ningyi Zhang, Pandeng Li, Pingyu Wu, Ruihang Chu, Ruili Feng, Shiwei Zhang, Siyang Sun, Tao Fang, T...

  6. [6]

    The past mistake is the future wisdom: Error-driven contrastive probability optimization for chinese spell checking

    Yinghui Li, Qingyu Zhou, Yangning Li, Zhongli Li, Ruiyang Liu, Rongyi Sun, Zizhen Wang, Chao Li, Yunbo Cao, and Hai-Tao Zheng. The past mistake is the future wisdom: Error-driven contrastive probability optimization for chinese spell checking. InFindings of the Association for Computational Linguistics: ACL 2022, pages 3202–3213, 2022

  7. [7]

    Let’s think with images efficiently! an interleaved-modal chain-of-thought reasoning framework with dynamic and precise visual thoughts

    Xu Liu, Yongheng Zhang, Qiguang Chen, Yao Li, Sheng Wang, and Libo Qin. Let’s think with images efficiently! an interleaved-modal chain-of-thought reasoning framework with dynamic and precise visual thoughts. InProceedings of the AAAI Conference on Artificial Intelligence, volume 40, pages 32213–32221, 2026

  8. [8]

    Youtu-llm: Unlocking the native agentic potential for lightweight large language models, 2026

    Junru Lu, Jiarui Qin, Lingfeng Qiao, Yinghui Li, Xinyi Dai, Bo Ke, Jianfeng He, Ruizhi Qiao, Di Yin, Xing Sun, Yunsheng Wu, Yinsong Liu, Shuangyin Liu, Mingkong Tang, Haodong Lin, Jiayi Kuang, Fanxu Meng, Xiaojuan Tang, Yunjia Xi, Junjie Huang, Haotong Yang, Zhenyi Shen, Yangning Li, Qianwen Zhang, Yifei Yu, Siyu An, Junnan Dong, Qiufeng Wang, Jie Wang, K...

  9. [9]

    Kimi-vl technical report.arXiv preprint arXiv:2504.07491, 2025

    Kimi Team, Angang Du, Bohong Yin, Bowei Xing, Bowen Qu, Bowen Wang, Cheng Chen, Chenlin Zhang, Chenzhuang Du, Chu Wei, et al. Kimi-vl technical report.arXiv preprint arXiv:2504.07491, 2025

  10. [10]

    Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks

    Zhe Chen, Jiannan Wu, Wenhai Wang, Weijie Su, Guo Chen, Sen Xing, Muyan Zhong, Qinglong Zhang, Xizhou Zhu, Lewei Lu, et al. Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 24185–24198, 2024

  11. [11]

    Openai gpt-5 system card, 2025

    Aaditya Singh, Adam Fry, Adam Perelman, Adam Tart, Adi Ganesh, Ahmed El-Kishky, Aidan McLaughlin, Aiden Low, AJ Ostrow, Akhila Ananthram, Akshay Nathan, Alan Luo, Alec Helyar, Aleksander Madry, Aleksandr Efremov, Aleksandra Spyra, Alex Baker-Whitcomb, Alex Beutel, Alex Karpenko, Alex Makelov, Alex Neitz, Alex Wei, Alexandra Barr, Alexandre Kirchmeyer, Ale...

  12. [12]

    Gemini 2.5: Pushing the frontier with advanced reasoning, multimodality, long context, and next generation agentic capabilities.arXiv preprint arXiv:2507.06261, 2025

    Gheorghe Comanici, Eric Bieber, Mike Schaekermann, Ice Pasupat, Noveen Sachdeva, Inderjit Dhillon, Marcel Blistein, Ori Ram, Dan Zhang, Evan Rosen, et al. Gemini 2.5: Pushing the frontier with advanced reasoning, multimodality, long context, and next generation agentic capabilities.arXiv preprint arXiv:2507.06261, 2025

  13. [13]

    Yongheng Zhang, Ziang Liu, Jiaxuan Zhu, Shuai Wang, Xiangqi Chen, Haojing Huang, Jiayi Kuang, Siyu Chen, Ao Shen, Hao Wu, Qiufeng Wang, Qian-Wen Zhang, Junnan Dong, Wenhao Jiang, Ying Shen, Hai-Tao Zheng, Yinghui Li, Di Yin, Xing Sun, and Philip S. Yu. From chatbot to digital colleague: The paradigm shift toward persistent autonomous ai, 2026

  14. [14]

    Cognitive mismatch in multimodal large language models for discrete symbol understanding.arXiv preprint arXiv:2603.18472, 2026

    Yinghui Li, Jiayi Kuang, Peng Xing, Daixian Liu, Yongheng Zhang, Junnan Dong, Shu-Yu Guo, Yangning Li, Qingyu Zhou, Wenhao Jiang, et al. Cognitive mismatch in multimodal large language models for discrete symbol understanding.arXiv preprint arXiv:2603.18472, 2026

  15. [15]

    Video-chatgpt: Towards detailed video understanding via large vision and language models

    Muhammad Maaz, Hanoona Rasheed, Salman Khan, and Fahad Khan. Video-chatgpt: Towards detailed video understanding via large vision and language models. InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 12585–12602, 2024

  16. [16]

    Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis

    Chaoyou Fu, Yuhan Dai, Yongdong Luo, Lei Li, Shuhuai Ren, Renrui Zhang, Zihan Wang, Chenyu Zhou, Yunhang Shen, Mengdan Zhang, et al. Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 24108–24118, 2025

  17. [17]

    Videollama 3: Frontier multimodal foundation models for image and video understanding.arXiv preprint arXiv:2501.13106, 2025

    Boqiang Zhang, Kehan Li, Zesen Cheng, Zhiqiang Hu, Yuqian Yuan, Guanzheng Chen, Sicong Leng, Yuming Jiang, Hang Zhang, Xin Li, et al. Videollama 3: Frontier multimodal foundation models for image and video understanding.arXiv preprint arXiv:2501.13106, 2025

  18. [18]

    Internvl3.5: Advancing open-source multimodal models in versatility, reasoning, and efficiency, 2025

    Weiyun Wang, Zhangwei Gao, Lixin Gu, Hengjun Pu, Long Cui, Xingguang Wei, Zhaoyang Liu, Linglin Jing, Shenglong Ye, Jie Shao, Zhaokai Wang, Zhe Chen, Hongjie Zhang, Ganlin Yang, Haomin Wang, Qi Wei, Jinhui Yin, Wenhao Li, Erfei Cui, Guanzhou Chen, Zichen Ding, Changyao Tian, Zhenyu Wu, Jingjing Xie, Zehao Li, Bowen Yang, Yuchen Duan, Xuehui Wang, Zhi Hou,...

  19. [19]

    AutoCAP: Towards automatic cross-lingual alignment planning for zero-shot chain-of-thought

    Yongheng Zhang, Qiguang Chen, Min Li, Wanxiang Che, and Libo Qin. AutoCAP: Towards automatic cross-lingual alignment planning for zero-shot chain-of-thought. In Lun-Wei Ku, Andre Martins, and Vivek Srikumar, editors,Findings of the Association for Computational Linguistics: ACL 2024, pages 9191–9200, Bangkok, Thailand, August 2024. Association for Computa...

  20. [20]

    Qwen2-vl: Enhancing vision- language model’s perception of the world at any resolution, 2024

    Peng Wang, Shuai Bai, Sinan Tan, Shijie Wang, Zhihao Fan, Jinze Bai, Keqin Chen, Xuejing Liu, Jialin Wang, Wenbin Ge, Yang Fan, Kai Dang, Mengfei Du, Xuancheng Ren, Rui Men, Dayiheng Liu, Chang Zhou, Jingren Zhou, and Junyang Lin. Qwen2-vl: Enhancing vision- language model’s perception of the world at any resolution, 2024

  21. [21]

    Moviechat: From dense token to 11 sparse memory for long video understanding

    Enxin Song, Wenhao Chai, Guanhong Wang, Yucheng Zhang, Haoyang Zhou, Feiyang Wu, Haozhe Chi, Xun Guo, Tian Ye, Yanting Zhang, et al. Moviechat: From dense token to 11 sparse memory for long video understanding. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 18221–18232, 2024

  22. [22]

    Wrong-of-thought: An integrated reasoning framework with multi-perspective verification and wrong information

    Yongheng Zhang, Qiguang Chen, Jingxuan Zhou, Peng Wang, Jiasheng Si, Jin Wang, Wenpeng Lu, and Libo Qin. Wrong-of-thought: An integrated reasoning framework with multi-perspective verification and wrong information. In Yaser Al-Onaizan, Mohit Bansal, and Yun-Nung Chen, editors,Findings of the Association for Computational Linguistics: EMNLP 2024, pages 66...

  23. [23]

    Qwen3-vl technical report, 2025

    Shuai Bai, Yuxuan Cai, Ruizhe Chen, Keqin Chen, Xionghui Chen, Zesen Cheng, Lianghao Deng, Wei Ding, Chang Gao, Chunjiang Ge, Wenbin Ge, Zhifang Guo, Qidong Huang, Jie Huang, Fei Huang, Binyuan Hui, Shutong Jiang, Zhaohai Li, Mingsheng Li, Mei Li, Kaixin Li, Zicheng Lin, Junyang Lin, Xuejing Liu, Jiawei Liu, Chenglong Liu, Yang Liu, Dayiheng Liu, Shixuan ...

  24. [24]

    Gemma 3 technical report, 2025

    Gemma Team, Aishwarya Kamath, Johan Ferret, Shreya Pathak, Nino Vieillard, Ramona Merhej, Sarah Perrin, Tatiana Matejovicova, Alexandre Ramé, Morgane Rivière, Louis Rouillard, Thomas Mesnard, Geoffrey Cideron, Jean bastien Grill, Sabela Ramos, Edouard Yvinec, Michelle Casbon, Etienne Pot, Ivo Penchev, Gaël Liu, Francesco Visin, Kathleen Kenealy, Lucas Bey...

  25. [25]

    CCHall: A novel benchmark for joint cross-lingual and cross-modal hallucinations detection in large language models

    Yongheng Zhang, Xu Liu, Ruoxi Zhou, Qiguang Chen, Hao Fei, Wenpeng Lu, and Libo Qin. CCHall: A novel benchmark for joint cross-lingual and cross-modal hallucinations detection in large language models. In Wanxiang Che, Joyce Nabende, Ekaterina Shutova, and Mohammad Taher Pilehvar, editors,Proceedings of the 63rd Annual Meeting of the Association for Compu...

  26. [26]

    Mitigating visual forgetting via take-along visual conditioning for multi-modal long cot reasoning

    Hai-Long Sun, Zhun Sun, Houwen Peng, and Han-Jia Ye. Mitigating visual forgetting via take-along visual conditioning for multi-modal long cot reasoning. InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 5158–5171, 2025

  27. [27]

    Seeing through the chain: Mitigate hallucination in multimodal reasoning models via cot compression and contrastive preference optimization.arXiv preprint arXiv:2602.03380, 2026

    Hao Fang, Jinyu Li, Jiawei Kong, Tianqu Zhuang, Kuofeng Gao, Bin Chen, Shu-Tao Xia, and Yaowei Wang. Seeing through the chain: Mitigate hallucination in multimodal reasoning models via cot compression and contrastive preference optimization.arXiv preprint arXiv:2602.03380, 2026

  28. [28]

    Context length alone hurts llm performance despite perfect retrieval.arXiv preprint arXiv:2510.05381, 2025

    Yufeng Du, Minyang Tian, Srikanth Ronanki, Subendhu Rongali, Sravan Bodapati, Aram Galstyan, Azton Wells, Roy Schwartz, Eliu A Huerta, and Hao Peng. Context length alone hurts llm performance despite perfect retrieval.arXiv preprint arXiv:2510.05381, 2025

  29. [29]

    Visual hallucinations of multi-modal large language models

    Wen Huang, Hongbin Liu, Minxin Guo, and Neil Gong. Visual hallucinations of multi-modal large language models. InFindings of the Association for Computational Linguistics: ACL 2024, pages 9614–9631, 2024

  30. [30]

    Vigc: Visual instruction generation and correction

    Bin Wang, Fan Wu, Xiao Han, Jiahui Peng, Huaping Zhong, Pan Zhang, Xiaoyi Dong, Weijia Li, Wei Li, Jiaqi Wang, et al. Vigc: Visual instruction generation and correction. InProceedings of the AAAI Conference on Artificial Intelligence, volume 38, pages 5309–5317, 2024

  31. [31]

    Hourvideo: 1-hour video- language understanding.Advances in Neural Information Processing Systems, 37:53168–53197, 2024

    Keshigeyan Chandrasegaran, Agrim Gupta, Lea M Hadzic, Taran Kota, Jimming He, Cristóbal Eyzaguirre, Zane Durante, Manling Li, Jiajun Wu, and Li Fei-Fei. Hourvideo: 1-hour video- language understanding.Advances in Neural Information Processing Systems, 37:53168–53197, 2024

  32. [32]

    Qwen3.5: Towards native multimodal agents, February 2026

    Qwen Team. Qwen3.5: Towards native multimodal agents, February 2026. 12

  33. [33]

    Tempcompass: Do video llms really understand videos?arXiv preprint arXiv:2403.00476, 2024

    Yuanxin Liu, Shicheng Li, Yi Liu, Yuxiang Wang, Shuhuai Ren, Lei Li, Sishuo Chen, Xu Sun, and Lu Hou. Tempcompass: Do video llms really understand videos?arXiv preprint arXiv:2403.00476, 2024

  34. [34]

    Mvbench: A comprehensive multi-modal video understanding benchmark

    Kunchang Li, Yali Wang, Yinan He, Yizhuo Li, Yi Wang, Yi Liu, Zun Wang, Jilan Xu, Guo Chen, Ping Luo, Limin Wang, and Yu Qiao. Mvbench: A comprehensive multi-modal video understanding benchmark. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024

  35. [35]

    Mmvu: Measuring expert-level multi-discipline video understanding.arXiv preprint arXiv:2501.12380, 2025

    Yilun Zhao, Lujing Xie, Haowei Zhang, Guo Gan, Yitao Long, Zhiyuan Hu, Tongyan Hu, Weiyuan Chen, Chuhan Li, Junyang Song, Zhijian Xu, Chengye Wang, Weifeng Pan, Ziyao Shangguan, Xiangru Tang, Zhenwen Liang, Yixin Liu, Chen Zhao, and Arman Cohan. Mmvu: Measuring expert-level multi-discipline video understanding.arXiv preprint arXiv:2501.12380, 2025

  36. [36]

    Video-mmmu: Evaluating knowledge acquisition from multi-discipline professional videos

    Kairui Hu, Penghao Wu, Fanyi Pu, Wang Xiao, Yuanhan Zhang, Xiang Yue, Bo Li, and Ziwei Liu. Video-mmmu: Evaluating knowledge acquisition from multi-discipline professional videos. arXiv preprint arXiv:2501.13826, 2025

  37. [37]

    Gupta, Rilyn Han, Li Fei-Fei, and Saining Xie

    Jihan Yang, Shusheng Yang, Anjali W. Gupta, Rilyn Han, Li Fei-Fei, and Saining Xie. Thinking in space: How multimodal large language models see, remember, and recall spaces.arXiv preprint arXiv:2412.14171, 2024

  38. [38]

    Deepseekmath: Pushing the limits of mathematical reasoning in open language models.arXiv preprint arXiv:2402.03300, 2024

    Zhihong Shao, Peiyi Wang, Qihao Zhu, Runxin Xu, Junxiao Song, Xiao Bi, Haowei Zhang, Mingchuan Zhang, YK Li, Yang Wu, et al. Deepseekmath: Pushing the limits of mathematical reasoning in open language models.arXiv preprint arXiv:2402.03300, 2024

  39. [39]

    Open-o3-video: Grounded video reasoning with explicit spatio-temporal evidence.arXiv preprint arXiv:2510.20579, 2025

    Jiahao Meng, Xiangtai Li, Haochen Wang, Yue Tan, Tao Zhang, Lingdong Kong, Yunhai Tong, Anran Wang, Zhiyang Teng, Yujing Wang, and Zhuochen Wang. Open-o3-video: Grounded video reasoning with explicit spatio-temporal evidence.arXiv preprint arXiv:2510.20579, 2025

  40. [40]

    Video-r1: Reinforcing video reasoning in mllms.arXiv preprint arXiv:2503.21776, 2025

    Kaituo Feng, Kaixiong Gong, Bohao Li, Zonghao Guo, Yibing Wang, Tianshuo Peng, Junfei Wu, Xiaoying Zhang, Benyou Wang, and Xiangyu Yue. Video-r1: Reinforcing video reasoning in mllms.arXiv preprint arXiv:2503.21776, 2025

  41. [41]

    Llama-vid: An image is worth 2 tokens in large language models.arXiv preprint arXiv:2311.17043, 2023

    Yanwei Li, Chengyao Wang, and Jiaya Jia. Llama-vid: An image is worth 2 tokens in large language models.arXiv preprint arXiv:2311.17043, 2023

  42. [42]

    Videollama 2: Advancing spatial- temporal modeling and audio understanding in video-llms.arXiv preprint arXiv:2406.07476, 2024

    Zesen Cheng, Sicong Leng, Hang Zhang, Yifei Xin, Xin Li, Guanzheng Chen, Yongxin Zhu, Wenqi Zhang, Ziyang Luo, Deli Zhao, and Lidong Bing. Videollama 2: Advancing spatial- temporal modeling and audio understanding in video-llms.arXiv preprint arXiv:2406.07476, 2024

  43. [43]

    Long context transfer from language to vision

    Peiyuan Zhang, Kaichen Zhang, Bo Li, Guangtao Zeng, Jingkang Yang, Yuanhan Zhang, Ziyue Wang, Haoran Tan, Chunyuan Li, and Ziwei Liu. Long context transfer from language to vision. arXiv preprint arXiv:2406.16852, 2024

  44. [44]

    Vila: On pre-training for visual language models

    Ji Lin, Hongxu Yin, Wei Ping, Yao Lu, Pavlo Molchanov, Andrew Tao, Huizi Mao, Jan Kautz, Mohammad Shoeybi, and Song Han. Vila: On pre-training for visual language models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024

  45. [45]

    Unhackable temporal rewarding for scalable video mllms.arXiv preprint arXiv:2502.12081, 2025

    En Yu, Kangheng Lin, Liang Zhao, Yana Wei, Zining Zhu, Haoran Wei, Jianjian Sun, Zheng Ge, Xiangyu Zhang, Jingyu Wang, and Wenbing Tao. Unhackable temporal rewarding for scalable video mllms.arXiv preprint arXiv:2502.12081, 2025

  46. [46]

    Llava-onevision: Easy visual task transfer

    Bo Li, Yuanhan Zhang, Dong Guo, Renrui Zhang, Feng Li, Hao Zhang, Kaichen Zhang, Peiyuan Zhang, Yanwei Li, Ziwei Liu, and Chunyuan Li. Llava-onevision: Easy visual task transfer. arXiv preprint arXiv:2408.03326, 2024

  47. [47]

    Kangaroo: A powerful video-language model supporting long-context video input.arXiv preprint arXiv:2408.15542, 2024

    Jiajun Liu, Yibing Wang, Hanghang Ma, Xiaoping Wu, Xiaoqi Ma, Xiaoming Wei, Jianbin Jiao, Enhua Wu, and Jie Hu. Kangaroo: A powerful video-language model supporting long-context video input.arXiv preprint arXiv:2408.15542, 2024. 13

  48. [48]

    Chain-of-thought prompting elicits reasoning in large language models

    Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al. Chain-of-thought prompting elicits reasoning in large language models. Advances in neural information processing systems, 35:24824–24837, 2022

  49. [49]

    Training language models to follow instructions with human feedback.Advances in neural information processing systems, 35:27730–27744, 2022

    Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. Training language models to follow instructions with human feedback.Advances in neural information processing systems, 35:27730–27744, 2022

  50. [50]

    Video-llava: Learning united visual representation by alignment before projection

    Bin Lin, Yang Ye, Bin Zhu, Jiaxi Cui, Munan Ning, Peng Jin, and Li Yuan. Video-llava: Learning united visual representation by alignment before projection. InProceedings of the 2024 conference on empirical methods in natural language processing, pages 5971–5984, 2024

  51. [51]

    Gemini Team, Rohan Anil, Sebastian Borgeaud, Jean-Baptiste Alayrac, Jiahui Yu, Radu Soricut, Johan Schalkwyk, Andrew M. Dai, Anja Hauth, Katie Millican, David Silver, Melvin Johnson, Ioannis Antonoglou, Julian Schrittwieser, Amelia Glaese, Jilin Chen, Emily Pitler, Timothy Lillicrap, Angeliki Lazaridou, Orhan Firat, and James Molloy et al. Gemini: A famil...

  52. [52]

    Thinking with videos: Multimodal tool-augmented reinforcement learning for long video reasoning

    Haoji Zhang, Xin Gu, Jiawen Li, Chixiang Ma, Sule Bai, Chubin Zhang, Bowen Zhang, Zhichao Zhou, Dongliang He, and Yansong Tang. Thinking with videos: Multimodal tool-augmented reinforcement learning for long video reasoning. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 32903–32914, 2026

  53. [53]

    Video-of-thought: step-by-step video reasoning from perception to cognition

    Hao Fei, Shengqiong Wu, Wei Ji, Hanwang Zhang, Meishan Zhang, Mong Li Lee, and Wynne Hsu. Video-of-thought: step-by-step video reasoning from perception to cognition. InPro- ceedings of the 41st International Conference on Machine Learning, pages 13109–13125, 2024

  54. [54]

    Vitcot: Video-text interleaved chain-of-thought for boosting video understanding in large language models

    Yongheng Zhang, Xu Liu, Ruihan Tao, Qiguang Chen, Hao Fei, Wanxiang Che, and Libo Qin. Vitcot: Video-text interleaved chain-of-thought for boosting video understanding in large language models. InProceedings of the 33rd ACM International Conference on Multimedia, pages 5267–5276, 2025

  55. [55]

    Training large language models to reason in a continuous latent space.arXiv preprint arXiv:2412.06769, 2024

    Shibo Hao, Sainbayar Sukhbaatar, DiJia Su, Xian Li, Zhiting Hu, Jason Weston, and Yuandong Tian. Training large language models to reason in a continuous latent space.arXiv preprint arXiv:2412.06769, 2024

  56. [56]

    SoftCoT: Soft chain-of-thought for efficient reasoning with LLMs

    Yige Xu, Xu Guo, Zhiwei Zeng, and Chunyan Miao. SoftCoT: Soft chain-of-thought for efficient reasoning with LLMs. InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 23336–23351, 2025

  57. [57]

    Hybrid latent reasoning via reinforcement learning.Advances in Neural Information Processing Systems, 38:5501–5530, 2026

    Zhenrui Yue, Bowen Jin, Huimin Zeng, Honglei Zhuang, Zhen Qin, Jinsung Yoon, Lanyu Shang, Jiawei Han, and Dong Wang. Hybrid latent reasoning via reinforcement learning.Advances in Neural Information Processing Systems, 38:5501–5530, 2026

  58. [58]

    Reasoning beyond language: A comprehensive survey on latent chain-of-thought reasoning.arXiv preprint arXiv:2505.16782, 2025

    Xinghao Chen, Anhao Zhao, Heming Xia, Xuan Lu, Hanlin Wang, Yanjun Chen, Wei Zhang, Jian Wang, Wenjie Li, and Xiaoyu Shen. Reasoning beyond language: A comprehensive survey on latent chain-of-thought reasoning.arXiv preprint arXiv:2505.16782, 2025

  59. [59]

    Monet: Reasoning in latent visual space beyond image and language

    Qixun Wang, Yang Shi, Yifei Wang, Yuanxing Zhang, Pengfei Wan, Kun Gai, Xianghua Ying, and Yisen Wang. Monet: Reasoning in latent visual space beyond image and language. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 12030–12040, 2026

  60. [60]

    Scaling up test-time compute with latent reasoning: A recurrent depth approach.Advances in Neural Information Processing Systems, 38:41340–41391, 2026

    Jonas Geiping, Sean McLeish, Neel Jain, John Kirchenbauer, Siddharth Singh, Brian Bartoldson, Bhavya Kailkhura, Abhinav Bhatele, and Tom Goldstein. Scaling up test-time compute with latent reasoning: A recurrent depth approach.Advances in Neural Information Processing Systems, 38:41340–41391, 2026. 14 A Appendix Overview This appendix documents implementa...