Pith. sign in

REVIEW 2 major objections 1 minor 132 references

VideoLatent: Video-Language Learning via Latent Self-Forcing

T0 review · 2 major / 1 minor · reviewed 2026-06-26 · grok-4.3

Pith's one-line read VideoLatent learns video reasoning in latent space from standard QA triplets alone.

desk verdict VideoLatent shows a workable way to train video latents on plain QA data but the evidence that alignment plus diversity alone produces real reasoning structure is still thin. read the letter →

arxiv 2606.22870 v1 pith:ACIX7BNA submitted 2026-06-22 cs.CV cs.AI

classification cs.CVcs.AI
keywords VideounderstandingMultimodalLLMsLatentreasoningSelf-forcingtrainingChain-of-thoughtQAModelefficiency
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper presents VideoLatent as a way to add visual latent reasoning to multimodal models for videos. The key is a latent self-forcing training method that uses alignment and diversity goals on regular video question-answer data, avoiding the need for chain-of-thought labels or other extra signals. If successful, this would make high-performance video understanding more accessible by lowering both data preparation and compute costs. Experiments show gains over prior methods on many benchmarks with much less overhead.

What carries the argument

Latent self-forcing training paradigm consisting of latent alignment and latent diversity objectives that guide the generation of useful visual latents for reasoning.

What would settle it

Demonstrating on a held-out video reasoning task that performance does not exceed that of a standard MLLM baseline when no CoT or extra annotations are used.

Watch

Extended reading notes

Core claim

The authors claim that their VideoLatent model, equipped with a latent injection module, can perform visual latent reasoning for video tasks by training with a latent self-forcing paradigm that includes latent alignment and latent diversity objectives. These objectives are applied using only standard video-question-answer triplets, without reliance on CoT traces, auxiliary images, or fine-grained annotations. This results in consistent outperformance on general video understanding and complex reasoning across 14 benchmarks, along with major efficiency improvements.

Load-bearing premise

The latent alignment and diversity objectives trained solely on standard video-QA triplets are enough to create visual latents that enable effective reasoning without additional supervision.

Editorial extensions

If this is right

  • Outperforms standard and latent MLLMs on 14 video benchmarks for understanding and reasoning.
  • Reduces training overhead by approximately 6 times and inference overhead by approximately 68 times relative to Video-R1.
  • Generalizes effectively across different MLLM backbones and model scales.
  • Supports video-language learning without labor-intensive CoT annotations or auxiliary supervision.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The same objectives could potentially be applied to other video-related tasks such as captioning or action recognition.
  • Efficiency improvements may enable training on much larger video datasets than previously feasible.
  • Latent reasoning might transfer to real-time applications where CoT methods are too slow.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 1 minor

Summary. The paper introduces VideoLatent, an MLLM for video understanding and reasoning equipped with a latent injection module. It proposes a latent self-forcing training paradigm consisting of latent alignment and latent diversity objectives that are optimized solely on standard video-QA triplets (no CoT traces or auxiliary signals). The central claims are consistent outperformance versus standard and latent MLLMs on 14 benchmarks plus large efficiency gains versus Video-R1 (∼6× training, ∼68× inference) and generalizability across backbones and scales.

Significance. If the central sufficiency claim holds, the work would be significant: it offers a scalable route to visual latent reasoning for video without labor-intensive CoT annotations, directly addressing a scalability bottleneck in prior latent-reasoning methods. The reported efficiency numbers, if reproducible and fairly matched, would constitute a practical advance for deployment.

major comments (2)
  1. [Experiments / Method] The central claim that latent alignment + diversity objectives (trained only on video-QA triplets) suffice to induce visual latents supporting complex reasoning is load-bearing yet unsupported by visible evidence. No ablation isolates the contribution of these objectives versus the injection module or training recipe; gains on reasoning benchmarks could therefore be explained by other factors.
  2. [Experiments] Efficiency comparison to Video-R1 (∼6×/∼68×) is presented without explicit statement of how baselines were matched for model size, data, or optimization; this is required to substantiate the claim and is absent from the reported results.
minor comments (1)
  1. [Abstract] Notation for the efficiency ratios uses approximate symbols without defining the exact measurement protocol (wall-clock, FLOPs, or tokens).

Simulated Author's Rebuttal

2 responses · 0 unresolved

We thank the referee for the constructive feedback. We address the two major comments below with clarifications and commitments to revisions that strengthen the experimental evidence without altering the core claims.

read point-by-point responses
  1. Referee: [Experiments / Method] The central claim that latent alignment + diversity objectives (trained only on video-QA triplets) suffice to induce visual latents supporting complex reasoning is load-bearing yet unsupported by visible evidence. No ablation isolates the contribution of these objectives versus the injection module or training recipe; gains on reasoning benchmarks could therefore be explained by other factors.

    Authors: We agree that isolating the objectives is important for substantiating the central claim. The manuscript reports overall gains from the full VideoLatent pipeline but does not include dedicated ablations separating latent alignment, latent diversity, and the injection module. In the revision we will add these ablations (full model vs. module-only vs. objectives-ablated variants) on the same video-QA triplets to demonstrate that the self-forcing objectives are responsible for the reasoning improvements beyond the injection module alone. revision: yes

  2. Referee: [Experiments] Efficiency comparison to Video-R1 (∼6×/∼68×) is presented without explicit statement of how baselines were matched for model size, data, or optimization; this is required to substantiate the claim and is absent from the reported results.

    Authors: We acknowledge that the efficiency section would benefit from explicit matching details. The reported ∼6× training and ∼68× inference gains versus Video-R1 were obtained using the same backbone scale and comparable volumes of standard video-QA triplets under matched optimization settings. In the revision we will expand the experimental protocol subsection to state the exact model sizes, data quantities, and hyperparameter matching used for the Video-R1 baseline, ensuring the comparison is fully reproducible and fair. revision: yes

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: new objectives presented as empirical additions without definitional reduction

full rationale

The provided abstract and description introduce VideoLatent via a latent self-forcing paradigm consisting of alignment and diversity objectives trained exclusively on standard video-QA triplets. No equations, parameter-fitting steps, or self-citations are shown that would make any claimed prediction or uniqueness result equivalent to its own inputs by construction. Performance and efficiency claims are framed as outcomes of experiments across benchmarks rather than derivations that loop back to fitted values or prior author work. The central sufficiency assumption is an empirical hypothesis, not a self-referential definition, so the derivation chain remains independent of the target results.

Assumptions & free parameters 0 free parameters · 0 assumptions · 0 invented entities

Abstract-only; no equations or implementation details supplied, so no free parameters, axioms, or invented entities can be identified.

how reviews work

0 comments
Cite this review

Pith. "Pith review of VideoLatent: Video-Language Learning via Latent Self-Forcing." pith.science (2026). https://pith.science/paper/ACIX7BNA

@misc{pith2026260622870,
  author       = {Pith},
  title        = {Pith review of: VideoLatent: Video-Language Learning via Latent Self-Forcing},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/ACIX7BNA}},
  note         = {Machine review of arXiv:2606.22870}
}
abstract

Recent advancements in chain-of-thought (CoT) reasoning have shown promise in enhancing video understanding and reasoning capabilities of multimodal large language models (MLLMs). However, existing CoT-based MLLMs require labor-intensive CoT annotations and incur substantial training and inference overhead. While visual latent reasoning has emerged as a more efficient alternative, existing methods primarily focus on image tasks and heavily rely on additional supervision signals for visual latent generation (e.g., CoT traces, auxiliary images, or fine-grained annotations), limiting their scalability and transferability to video tasks. To bridge this gap, we introduce VideoLatent, a novel MLLM equipped with a latent injection module tailored for video understanding and reasoning. Specifically, VideoLatent learns to perform visual latent reasoning using a new latent self-forcing training paradigm, which comprises latent alignment and latent diversity objectives, and relies solely on standard video-question-answer triplets. Extensive experiments across 14 benchmarks demonstrate that our model consistently outperforms existing standard and latent MLLMs on general video understanding and complex video reasoning. Compared with Video-R1, our VideoLatent achieves superior computational efficiency, reducing training/inference overhead by $\sim$6$\times$/$\sim$68$\times$. Moreover, experiments demonstrate that our method has strong generalizability to different MLLM backbones and different model scales.

Figures

Figures reproduced from arXiv: 2606.22870 by the authors.

Figure 1
Figure 1. Our VideoLatent-7B consistently outperforms [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. Our VideoLatent achieves stronger or com [PITH_FULL_IMAGE:figures/full_fig_p002_2.png] view at source ↗
Figure 3
Figure 3. Overview of VideoLatent. Given an input video and a text question, our VideoLatent learns to perform visual latent reasoning (see Sec. 3.2) using our proposed latent self-forcing training paradigm (see Sec. 3.3). Specifically, we introduce a latent injection module to prevent self-generated latent thoughts from drifting away from the video and question context. Furthermore, our latent self-forcing covers both latent… view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

132 extracted references · 2 canonical work pages

  1. [1]

    CoRR , volume =

    OpenAI , title =. CoRR , volume =

  2. [2]

    LLaMA: Open and Efficient Foundation Language Models , journal =

    Hugo Touvron and Thibaut Lavril and Gautier Izacard and Xavier Martinet and Marie. LLaMA: Open and Efficient Foundation Language Models , journal =

  3. [3]

    CoRR , volume =

    Jinze Bai and Shuai Bai and Yunfei Chu and Zeyu Cui and Kai Dang and Xiaodong Deng and Yang Fan and Wenbin Ge and Yu Han and Fei Huang and Binyuan Hui and Luo Ji and Mei Li and Junyang Lin and Runji Lin and Dayiheng Liu and Gao Liu and Chengqiang Lu and Keming Lu and Jianxin Ma and Rui Men and Xingzhang Ren and Xuancheng Ren and Chuanqi Tan and Sinan Tan ...

  4. [4]

    Tom B. Brown and Benjamin Mann and Nick Ryder and Melanie Subbiah and Jared Kaplan and Prafulla Dhariwal and Arvind Neelakantan and Pranav Shyam and Girish Sastry and Amanda Askell and Sandhini Agarwal and Ariel Herbert. Language Models are Few-Shot Learners , booktitle =

  5. [5]

    CoRR , volume =

    An Yang and Baosong Yang and Beichen Zhang and Binyuan Hui and Bo Zheng and Bowen Yu and Chengyuan Li and Dayiheng Liu and Fei Huang and Haoran Wei and Huan Lin and Jian Yang and Jianhong Tu and Jianwei Zhang and Jianxin Yang and Jiaxi Yang and Jingren Zhou and Junyang Lin and Kai Dang and Keming Lu and Keqin Bao and Kexin Yang and Le Yu and Mei Li and Mi...

  6. [6]

    CoRR , volume =

    Jinze Bai and Shuai Bai and Shusheng Yang and Shijie Wang and Sinan Tan and Peng Wang and Junyang Lin and Chang Zhou and Jingren Zhou , title =. CoRR , volume =

  7. [7]

    Junnan Li and Dongxu Li and Silvio Savarese and Steven C. H. Hoi , title =

  8. [8]

    Yanwei Li and Yuechen Zhang and Chengyao Wang and Zhisheng Zhong and Yixin Chen and Ruihang Chu and Shaoteng Liu and Jiaya Jia , title =

Show all 132 references
  1. [9]

    CoRR , volume =

    Feng Li and Renrui Zhang and Hao Zhang and Yuanhan Zhang and Bo Li and Wei Li and Zejun Ma and Chunyuan Li , title =. CoRR , volume =

  2. [10]

    Bin Lin and Yang Ye and Bin Zhu and Jiaxi Cui and Munan Ning and Peng Jin and Li Yuan , title =

  3. [11]

    NeurIPS , year =

    Haotian Liu and Chunyuan Li and Qingyang Wu and Yong Jae Lee , title =. NeurIPS , year =

  4. [12]

    Beyond Embeddings: The Promise of Visual Table in Visual Reasoning , booktitle =

    Yiwu Zhong and Zi. Beyond Embeddings: The Promise of Visual Table in Visual Reasoning , booktitle =

  5. [13]

    2024 , note =

    OpenAI , title =. 2024 , note =

  6. [14]

    2025 , note =

    OpenAI , title =. 2025 , note =

  7. [15]

    CoRR , volume =

    Gemini Team , title =. CoRR , volume =

  8. [16]

    CoRR , volume =

    ByteDance Seed , title =. CoRR , volume =

  9. [17]

    Lillicrap and Jean

    Machel Reid and Nikolay Savinov and Denis Teplyashin and Dmitry Lepikhin and Timothy P. Lillicrap and Jean. Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context , journal =

  10. [18]

    Enhancing Temporal Modeling of Video LLMs via Time Gating , booktitle =

    Zi. Enhancing Temporal Modeling of Video LLMs via Time Gating , booktitle =

  11. [19]

    Ji Lin and Hongxu Yin and Wei Ping and Pavlo Molchanov and Mohammad Shoeybi and Song Han , title =

  12. [20]

    Yuanhan Zhang and Jinming Wu and Wei Li and Bo Li and Zejun Ma and Ziwei Liu and Chunyuan Li , title =. Trans. Mach. Learn. Res. , volume =

  13. [21]

    Bo Li and Yuanhan Zhang and Dong Guo and Renrui Zhang and Feng Li and Hao Zhang and Kaichen Zhang and Peiyuan Zhang and Yanwei Li and Ziwei Liu and Chunyuan Li , title =. Trans. Mach. Learn. Res. , volume =

  14. [22]

    Jiabo Ye and Haiyang Xu and Haowei Liu and Anwen Hu and Ming Yan and Qi Qian and Ji Zhang and Fei Huang and Jingren Zhou , title =

  15. [23]

    Muhammad Maaz and Hanoona Abdul Rasheed and Salman Khan and Fahad Khan , title =

  16. [24]

    Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models , booktitle =

    Matt Deitke and Christopher Clark and Sangho Lee and Rohun Tripathi and Yue Yang and Jae Sung Park and Mohammadreza Salehi and Niklas Muennighoff and Kyle Lo and Luca Soldaini and Jiasen Lu and Taira Anderson and Erin Bransom and Kiana Ehsani and Huong Ngo and Yen. Molmo and P...

  17. [25]

    CoRR , volume =

    Weiyun Wang and Zhangwei Gao and Lixin Gu and Hengjun Pu and Long Cui and Xingguang Wei and Zhaoyang Liu and Linglin Jing and Shenglong Ye and Jie Shao and Zhaokai Wang and Zhe Chen and Hongjie Zhang and Ganlin Yang and Haomin Wang and Qi Wei and Jinhui Yin and Wenhao Li and E...

  18. [26]

    Zhijian Liu and Ligeng Zhu and Baifeng Shi and Zhuoyang Zhang and Yuming Lou and Shang Yang and Haocheng Xi and Shiyi Cao and Yuxian Gu and Dacheng Li and Xiuyu Li and Haotian Tang and Yunhao Fang and Yukang Chen and Cheng

  19. [27]

    Wenliang Dai and Junnan Li and Dongxu Li and Anthony Meng Huat Tiong and Junqi Zhao and Weisheng Wang and Boyang Li and Pascale Fung and Steven C. H. Hoi , title =. NeurIPS , year =

  20. [28]

    CoRR , volume =

    Shuai Bai and Keqin Chen and Xuejing Liu and Jialin Wang and Wenbin Ge and Sibo Song and Kai Dang and Peng Wang and Shijie Wang and Jun Tang and Humen Zhong and Yuanzhi Zhu and Ming. CoRR , volume =

  21. [29]

    CoRR , volume =

    Qwen Team , title =. CoRR , volume =

  22. [30]

    CoRR , volume =

    Kaituo Feng and Kaixiong Gong and Bohao Li and Zonghao Guo and Yibing Wang and Tianshuo Peng and Benyou Wang and Xiangyu Yue , title =. CoRR , volume =

  23. [31]

    CoRR , volume =

    Xingjian Zhang and Siwei Wen and Wenjun Wu and Lei Huang , title =. CoRR , volume =

  24. [32]

    CoRR , volume =

    Qi Wang and Yanrui Yu and Ye Yuan and Rui Mao and Tianfei Zhou , title =. CoRR , volume =

  25. [33]

    CoRR , volume =

    Kaituo Feng and Manyuan Zhang and Hongyu Li and Kaixuan Fan and Shuang Chen and Yilei Jiang and Dian Zheng and Peiwen Sun and Yiyuan Zhang and Haoze Sun and Yan Feng and Peng Pei and Xunliang Cai and Xiangyu Yue , title =. CoRR , volume =

  26. [34]

    CoRR , volume =

    Shuming Liu and Mingchen Zhuge and Changsheng Zhao and Jun Chen and Lemeng Wu and Zechun Liu and Chenchen Zhu and Zhipeng Cai and Chong Zhou and Haozhe Liu and Ernie Chang and Saksham Suri and Hongyu Xu and Qi Qian and Wei Wen and Balakrishnan Varadarajan and Zhuang Liu and Hu...

  27. [35]

    CoRR , volume =

    Ziang Yan and Xinhao Li and Yinan He and Zhengrong Yue and Xiangyu Zeng and Yali Wang and Yu Qiao and Limin Wang and Yi Wang , title =. CoRR , volume =

  28. [36]

    CoRR , volume =

    Zefeng He and Xiaoye Qu and Yafu Li and Siyuan Huang and Daizong Liu and Yu Cheng , title =. CoRR , volume =

  29. [37]

    Haotian Liu and Chunyuan Li and Yuheng Li and Yong Jae Lee , title =

  30. [38]

    Advances in neural information processing systems , volume=

    Attention is all you need , author=. Advances in neural information processing systems , volume=

  31. [39]

    arXiv preprint arXiv:2010.11929 , year=

    An image is worth 16x16 words: Transformers for image recognition at scale , author=. arXiv preprint arXiv:2010.11929 , year=

  32. [40]

    International conference on machine learning , pages=

    Learning transferable visual models from natural language supervision , author=. International conference on machine learning , pages=. 2021 , organization=

  33. [41]

    Kunchang Li and Yali Wang and Yinan He and Yizhuo Li and Yi Wang and Yi Liu and Zun Wang and Jilan Xu and Guo Chen and Ping Lou and Limin Wang and Yu Qiao , title =

  34. [42]

    Yuanxin Liu and Shicheng Li and Yi Liu and Yuxiang Wang and Shuhuai Ren and Lei Li and Sishuo Chen and Xu Sun and Lu Hou , title =

  35. [43]

    Chaoyou Fu and Yuhan Dai and Yongdong Luo and Lei Li and Shuhuai Ren and Renrui Zhang and Zihan Wang and Chenyu Zhou and Yunhang Shen and Mengdan Zhang and Peixian Chen and Yanwei Li and Shaohui Lin and Sirui Zhao and Ke Li and Tong Xu and Xiawu Zheng and Enhong Chen and Caife...

  36. [44]

    NExT-QA: Next Phase of Question-Answering to Explaining Temporal Actions , booktitle =

    Junbin Xiao and Xindi Shang and Angela Yao and Tat. NExT-QA: Next Phase of Question-Answering to Explaining Temporal Actions , booktitle =

  37. [45]

    Junjie Zhou and Yan Shu and Bo Zhao and Boya Wu and Zhengyang Liang and Shitao Xiao and Minghao Qin and Xi Yang and Yongping Xiong and Bo Zhang and Tiejun Huang and Zheng Liu , title =

  38. [46]

    NeurIPS , year =

    Haoning Wu and Dongxu Li and Bei Chen and Junnan Li , title =. NeurIPS , year =

  39. [47]

    Weihan Wang and Zehai He and Wenyi Hong and Yean Cheng and Xiaohan Zhang and Ji Qi and Ming Ding and Xiaotao Gu and Shiyu Huang and Bin Xu and Yuxiao Dong and Jie Tang , title =

  40. [48]

    CoRR , volume =

    Yukun Qi and Yiming Zhao and Yu Zeng and Xikun Bao and Wenxuan Huang and Lin Chen and Zehui Chen and Jie Zhao and Zhongang Qi and Feng Zhao , title =. CoRR , volume =

  41. [49]

    Shaker and Anqi Tang and Muhammad Maaz and Ming

    Hanoona Abdul Rasheed and Abdelrahman M. Shaker and Anqi Tang and Muhammad Maaz and Ming. VideoMathQA: Benchmarking Mathematical Reasoning via Multimodal Understanding in Videos , journal =

  42. [50]

    CoRR , volume =

    Kairui Hu and Penghao Wu and Fanyi Pu and Wang Xiao and Yuanhan Zhang and Xiang Yue and Bo Li and Ziwei Liu , title =. CoRR , volume =

  43. [51]

    Yuanhan Zhang and Yunice Chew and Yuhao Dong and Aria Leo and Bo Hu and Ziwei Liu , title =

  44. [52]

    Yilun Zhao and Haowei Zhang and Lujing Xie and Tongyan Hu and Guo Gan and Yitao Long and Zhiyuan Hu and Weiyuan Chen and Chuhan Li and Zhijian Xu and Chengye Wang and Ziyao Shangguan and Zhenwen Liang and Yixin Liu and Chen Zhao and Arman Cohan , title =

  45. [53]

    Benno Krojer and Mojtaba Komeili and Candace Ross and Quentin Garrido and Koustuv Sinha and Nicolas Ballas and Mido Assran , title =. Trans. Mach. Learn. Res. , volume =

  46. [54]

    Gupta and Rilyn Han and Li Fei

    Jihan Yang and Shusheng Yang and Anjali W. Gupta and Rilyn Han and Li Fei. Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces , booktitle =

  47. [55]

    Cambrian-S: Towards Spatial Supersensing in Video , journal =

    Shusheng Yang and Jihan Yang and Pinzhi Huang and Ellis Brown and Zihao Yang and Yue Yu and Shengbang Tong and Zihan Zheng and Yifan Xu and Muhan Wang and Daohan Lu and Rob Fergus and Yann LeCun and Li Fei. Cambrian-S: Towards Spatial Supersensing in Video , journal =

  48. [56]

    Perception Test:

    Viorica Patraucean and Lucas Smaira and Ankush Gupta and Adri. Perception Test:. NeurIPS , year =

  49. [57]

    CoRR , volume =

    Junhao Cheng and Yuying Ge and Teng Wang and Yixiao Ge and Jing Liao and Ying Shan , title =. CoRR , volume =

  50. [58]

    CoRR , volume =

    Yukang Chen and Wei Huang and Baifeng Shi and Qinghao Hu and Hanrong Ye and Ligeng Zhu and Zhijian Liu and Pavlo Molchanov and Jan Kautz and Xiaojuan Qi and Sifei Liu and Hongxu Yin and Yao Lu and Song Han , title =. CoRR , volume =

  51. [59]

    CoRR , volume =

    Zihui Xue and Mi Luo and Kristen Grauman , title =. CoRR , volume =

  52. [60]

    CoRR , volume =

    Jianrui Zhang and Mu Cai and Yong Jae Lee , title =. CoRR , volume =

  53. [61]

    Guo Chen and Yicheng Liu and Yifei Huang and Baoqi Pei and Jilan Xu and Yuping He and Tong Lu and Yali Wang and Limin Wang , title =

  54. [62]

    Wenyi Hong and Yean Cheng and Zhuoyi Yang and Weihan Wang and Lefan Wang and Xiaotao Gu and Shiyu Huang and Yuxiao Dong and Jie Tang , title =

  55. [63]

    Ziyao Shangguan and Chuhan Li and Yuxuan Ding and Yanan Zheng and Yilun Zhao and Tesca Fitzgerald and Arman Cohan , title =

  56. [64]

    Daniel Cores and Michael Dorkenwald and Manuel Mucientes and Cees G. M. Snoek and Yuki M. Asano , title =. CoRR , volume =

  57. [65]

    CoRR , volume =

    Mu Cai and Reuben Tan and Jianrui Zhang and Bocheng Zou and Kai Zhang and Feng Yao and Fangrui Zhu and Jing Gu and Yiwu Zhong and Yuzhang Shang and Yao Dou and Jaden Park and Jianfeng Gao and Yong Jae Lee and Jianwei Yang , title =. CoRR , volume =

  58. [66]

    CoRR , volume =

    Kejian Zhu and Zhuoran Jin and Hongbang Yuan and Jiachun Li and Shangqing Tu and Pengfei Cao and Yubo Chen and Kang Liu and Jun Zhao , title =. CoRR , volume =

  59. [67]

    Jiashuo Yu and Yue Wu and Meng Chu and Zhifei Ren and Zizheng Huang and Pei Chu and Ruijie Zhang and Yinan He and Qirui Li and Songze Li and Zhenxiang Li and Zhongying Tu and Conghui He and Yu Qiao and Yali Wang and Yi Wang and Limin Wang , title =

  60. [68]

    Jeff Rasley and Samyam Rajbhandari and Olatunji Ruwase and Yuxiong He , title =

  61. [69]

    Girshick , title =

    Kaiming He and Haoqi Fan and Yuxin Wu and Saining Xie and Ross B. Girshick , title =

  62. [70]

    Chi and Quoc V

    Jason Wei and Xuezhi Wang and Dale Schuurmans and Maarten Bosma and Brian Ichter and Fei Xia and Ed H. Chi and Quoc V. Le and Denny Zhou , title =. NeurIPS , year =

  63. [71]

    Zhuosheng Zhang and Aston Zhang and Mu Li and Alex Smola , title =

  64. [72]

    Zhuosheng Zhang and Aston Zhang and Mu Li and Hai Zhao and George Karypis and Alex Smola , title =. Trans. Mach. Learn. Res. , volume =

  65. [73]

    Chancharik Mitra and Brandon Huang and Trevor Darrell and Roei Herzig , title =

  66. [74]

    DDCoT: Duty-Distinct Chain-of-Thought Prompting for Multimodal Reasoning in Language Models , booktitle =

    Ge Zheng and Bin Yang and Jiajin Tang and Hong. DDCoT: Duty-Distinct Chain-of-Thought Prompting for Multimodal Reasoning in Language Models , booktitle =

  67. [75]

    DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning , journal =

    DeepSeek. DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning , journal =

  68. [76]

    CoRR , volume =

    Shulin Tian and Ruiqi Wang and Hongming Guo and Penghao Wu and Yuhao Dong and Xiuying Wang and Jingkang Yang and Hao Zhang and Hongyuan Zhu and Ziwei Liu , title =. CoRR , volume =

  69. [77]

    Zhihong Shao and Peiyi Wang and Qihao Zhu and Runxin Xu and Junxiao Song and Mingchuan Zhang and Y. K. Li and Y. Wu and Daya Guo , title =. CoRR , volume =

  70. [78]

    CoRR , volume =

    Jiahao Meng and Xiangtai Li and Haochen Wang and Yue Tan and Tao Zhang and Lingdong Kong and Yunhai Tong and Anran Wang and Zhiyang Teng and Yujing Wang and Zhuochen Wang , title =. CoRR , volume =

  71. [79]

    Ziyang Wang and Jaehong Yoon and Shoubin Yu and Md Mohaiminul Islam and Gedas Bertasius and Mohit Bansal , title =

  72. [80]

    arXiv preprint arXiv:2503.13377 , year=

    Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding , author=. arXiv preprint arXiv:2503.13377 , year=

  73. [81]

    CoRR , volume =

    Shibo Hao and Sainbayar Sukhbaatar and DiJia Su and Xian Li and Zhiting Hu and Jason Weston and Yuandong Tian , title =. CoRR , volume =

  74. [82]

    Sachin Goyal and Ziwei Ji and Ankit Singh Rawat and Aditya Krishna Menon and Sanjiv Kumar and Vaishnavh Nagarajan , title =

  75. [83]

    Shieber , title =

    Yuntian Deng and Yejin Choi and Stuart M. Shieber , title =. CoRR , volume =

  76. [84]

    Zhenyi Shen and Hanqi Yan and Linhai Zhang and Zhanghao Hu and Yali Du and Yulan He , title =

  77. [85]

    Yige Xu and Xu Guo and Zhiwei Zeng and Chunyan Miao , title =

  78. [86]

    CoRR , volume =

    Yige Xu and Xu Guo and Zhiwei Zeng and Chunyan Miao , title =. CoRR , volume =

  79. [87]

    CoRR , volume =

    Jeffrey Cheng and Benjamin Van Durme , title =. CoRR , volume =

  80. [88]

    Perception Tokens Enhance Visual Reasoning in Multimodal Language Models , booktitle =

    Mahtab Bigverdi and Zelun Luo and Cheng. Perception Tokens Enhance Visual Reasoning in Multimodal Language Models , booktitle =

  81. [89]

    CoRR , volume =

    Xuan Shen and Yizhou Wang and Xiangxi Shi and Yanzhi Wang and Pu Zhao and Jiuxiang Gu , title =. CoRR , volume =

  82. [90]

    CoRR , volume =

    Zeyuan Yang and Xueyang Yu and Delin Chen and Maohao Shen and Chuang Gan , title =. CoRR , volume =

  83. [91]

    CoRR , volume =

    Yiming Qin and Bomin Wei and Jiaxin Ge and Konstantinos Kallidromitis and Stephanie Fu and Trevor Darrell and Xudong Wang , title =. CoRR , volume =

  84. [92]

    Segment Anything , booktitle =

    Alexander Kirillov and Eric Mintun and Nikhila Ravi and Hanzi Mao and Chlo. Segment Anything , booktitle =

  85. [93]

    NeurIPS , year =

    Lihe Yang and Bingyi Kang and Zilong Huang and Zhen Zhao and Xiaogang Xu and Jiashi Feng and Hengshuang Zhao , title =. NeurIPS , year =

  86. [94]

    Pixel Difference Networks for Efficient Edge Detection , booktitle =

    Zhuo Su and Wenzhe Liu and Zitong Yu and Dewen Hu and Qing Liao and Qi Tian and Matti Pietik. Pixel Difference Networks for Efficient Edge Detection , booktitle =

  87. [95]

    DINOv2: Learning Robust Visual Features without Supervision , journal =

    Maxime Oquab and Timoth. DINOv2: Learning Robust Visual Features without Supervision , journal =

  88. [96]

    CoRR , volume =

    Chengzhi Liu and Yuzhe Yang and Yue Fan and Qingyue Wei and Sheng Liu and Xin Eric Wang , title =. CoRR , volume =

  89. [97]

    CoRR , volume =

    Bangzheng Li and Ximeng Sun and Jiang Liu and Ze Wang and Jialian Wu and Xiaodong Yu and Hao Chen and Emad Barsoum and Muhao Chen and Zicheng Liu , title =. CoRR , volume =

  90. [98]

    CoRR , volume =

    Qixun Wang and Yang Shi and Yifei Wang and Yuanxing Zhang and Pengfei Wan and Kun Gai and Xianghua Ying and Yisen Wang , title =. CoRR , volume =

  91. [99]

    CoRR , volume =

    Yubo Wang and Juntian Zhang and Yichen Wu and Yankai Lin and Nils Lukas and Yuhan Liu , title =. CoRR , volume =

  92. [100]

    CoRR , volume =

    Zhangquan Chen and Manyuan Zhang and Xinlei Yu and Xufang Luo and Mingze Sun and Zihao Pan and Yan Feng and Peng Pei and Xunliang Cai and Ruqi Huang , title =. CoRR , volume =

  93. [101]

    Zhuoyang Liu and Jiaming Liu and Hao Chen and Jiale Yu and Ziyu Guo and Chengkai Hou and Chenyang Gu and Xiangju Mi and Renrui Zhang and Kun Wu and Zhengping Che and Jian Tang and Pheng. LaST\(. CoRR , volume =

  94. [102]

    CoRR , volume =

    Linquan Wu and Tianxiang Jiang and Yifei Dong and Haoyu Yang and Fengji Zhang and Shichaang Meng and Ai Xuan and Linqi Song and Jacky Keung , title =. CoRR , volume =

  95. [103]

    CoRR , volume =

    Jintao Tong and Jiaqi Gu and Yujing Lou and Lubin Fan and Yixiong Zou and Yue Wu and Jieping Ye and Ruixuan Li , title =. CoRR , volume =

  96. [104]

    CoRR , volume =

    Huanyu Zhang and Wenshan Wu and Chengzu Li and Ning Shang and Yan Xia and Yangyu Huang and Yifan Zhang and Li Dong and Zhang Zhang and Liang Wang and Tieniu Tan and Furu Wei , title =. CoRR , volume =

  97. [105]

    CoRR , volume =

    Byungwoo Jeon and Yoonwoo Jeong and Hyunseok Lee and Minsu Cho and Jinwoo Shin , title =. CoRR , volume =

  98. [106]

    CoRR , volume =

    Yudong Han and Yong Wang and Zaiquan Yang and Zhen Qu and Liyuan Pan and Xiangxiang Chu , title =. CoRR , volume =

  99. [107]

    Latent Implicit Visual Reasoning , journal =

    Kelvin Li and Chuyi Shang and Leonid Karlinsky and Rog. Latent Implicit Visual Reasoning , journal =

  100. [108]

    CoRR , volume =

    Shuai Dong and Siyuan Wang and Xingyu Liu and Zhongyu Wei , title =. CoRR , volume =

  101. [109]

    Show, Don't Tell: Morphing Latent Reasoning into Image Generation , journal =

    Harold Haodong Chen and Xinxiang Yin and Wen. Show, Don't Tell: Morphing Latent Reasoning into Image Generation , journal =

  102. [110]

    Plummer and Kate Saenko and Ranjay Krishna and Leonidas J

    Arijit Ray and Ahmed Abdelkader and Chengzhi Mao and Bryan A. Plummer and Kate Saenko and Ranjay Krishna and Leonidas J. Guibas and Wen. Mull-Tokens: Modality-Agnostic Latent Thinking , journal =

  103. [111]

    CoRR , volume =

    Jizheng Ma and Xiaofei Zhou and Yanlong Song and Han Yan , title =. CoRR , volume =

  104. [112]

    CoRR , volume =

    Yifei Shao and Kun Zhou and Ziming Xu and Mohammad Atif Quamar and Shibo Hao and Zhen Wang and Zhiting Hu and Biwei Huang , title =. CoRR , volume =

  105. [113]

    The Latent Space: Foundation, Evolution, Mechanism, Ability, and Outlook , journal =

    Xinlei Yu and Zhangquan Chen and Yongbo He and Tianyu Fu and Cheng Yang and Chengming Xu and Yue Ma and Xiaobin Hu and Zhe Cao and Jie Xu and Guibin Zhang and Jiale Tao and Jiayi Zhang and Siyuan Ma and Kaituo Feng and Haojie Huang and Youxing Li and Ronghao Chen and Huacan Wa...

  106. [114]

    Rethinking Chain-of-Thought Reasoning for Videos , journal =

    Yiwu Zhong and Zi. Rethinking Chain-of-Thought Reasoning for Videos , journal =

  107. [115]

    LaST-VLA: Thinking in Latent Spatio-Temporal Space for Vision-Language-Action in Autonomous Driving , journal =

    Yuechen Luo and Fang Li and Shaoqing Xu and Yang Ji and Zehan Zhang and Bing Wang and Yuannan Shen and Jianwei Cui and Long Chen and Guang Chen and Hangjun Ye and Zhi. LaST-VLA: Thinking in Latent Spatio-Temporal Space for Vision-Language-Action in Autonomous Driving , journal =

  108. [116]

    CoRR , volume =

    Jiaru Zou and Xiyuan Yang and Ruizhong Qiu and Gaotang Li and Katherine Tieu and Pan Lu and Ke Shen and Hanghang Tong and Yejin Choi and Jingrui He and James Zou and Mengdi Wang and Ling Yang , title =. CoRR , volume =

  109. [117]

    CoRR , volume =

    Wenxuan Huang and Bohan Jia and Zijie Zhai and Shaosheng Cao and Zheyu Ye and Fei Zhao and Zhe Xu and Yao Hu and Shaohui Lin , title =. CoRR , volume =

  110. [118]

    Mido Assran and Adrien Bardes and David Fan and Quentin Garrido and Russell Howes and Mojtaba Komeili and Matthew J. Muckley and Ammar Rizvi and Claire Roberts and Koustuv Sinha and Artem Zholus and Sergio Arnaud and Abha Gejji and Ada Martin and Francois Robert Hogan and Dani...

  111. [119]

    CoRR , volume =

    Delong Chen and Mustafa Shukor and Th. CoRR , volume =

  112. [120]

    Tenenbaum , title =

    Kexin Yi and Chuang Gan and Yunzhu Li and Pushmeet Kohli and Jiajun Wu and Antonio Torralba and Joshua B. Tenenbaum , title =

  113. [121]

    Tenenbaum and Chuang Gan , title =

    Bo Wu and Shoubin Yu and Zhenfang Chen and Joshua B. Tenenbaum and Chuang Gan , title =. CoRR , volume =

  114. [122]

    Chengzu Li and Wenshan Wu and Huanyu Zhang and Yan Xia and Shaoguang Mao and Li Dong and Ivan Vulic and Furu Wei , title =

  115. [123]

    CoRR , volume =

    Yi Xu and Chengzu Li and Han Zhou and Xingchen Wan and Caiqi Zhang and Anna Korhonen and Ivan Vulic , title =. CoRR , volume =

  116. [124]

    Video-of-Thought: Step-by-Step Video Reasoning from Perception to Cognition , booktitle =

    Hao Fei and Shengqiong Wu and Wei Ji and Hanwang Zhang and Meishan Zhang and Mong. Video-of-Thought: Step-by-Step Video Reasoning from Perception to Cognition , booktitle =

  117. [125]

    CoRR , volume =

    Sara Ghazanfari and Francesco Croce and Nicolas Flammarion and Prashanth Krishnamurthy and Farshad Khorrami and Siddharth Garg , title =. CoRR , volume =

  118. [126]

    CoRR , volume =

    Ye Liu and Kevin Qinghong Lin and Chang Wen Chen and Mike Zheng Shou , title =. CoRR , volume =

  119. [127]

    Songhao Han and Wei Huang and Hairong Shi and Le Zhuo and Xiu Su and Shifeng Zhang and Xu Zhou and Xiaojuan Qi and Yue Liao and Si Liu , title =

  120. [128]

    CoRR , volume =

    Yifan Wang and Shiyu Li and Peiming Li and Xiaochen Yang and Yang Tang and Zheng Wei , title =. CoRR , volume =

  121. [129]

    L2V-CoT: Cross-Modal Transfer of Chain-of-Thought Reasoning via Latent Intervention , booktitle =

    Yu. L2V-CoT: Cross-Modal Transfer of Chain-of-Thought Reasoning via Latent Intervention , booktitle =

  122. [130]

    Jun Gao and Yongqi Li and Ziqiang Cao and Wenjie Li , title =

  123. [131]

    Liqi He and Zuchao Li and Xiantao Cai and Ping Wang , title =

  124. [132]

    CoRR , volume =

    Xiangkai Ma and Lekai Xing and Han Zhang and Wenzhong Li and Sanglu Lu , title =. CoRR , volume =

Pith tools

Reviewed June 26, 2026 · model on record in the stance chip above.