REVIEW 3 major objections 2 minor 1 cited by
Empowering Nanoscale Connectivity through Molecular Communication: A Case Study of Virus Infection
T0 review · 3 major / 2 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read Molecular communication can model viral spread, detect infected hosts, and flag mutations at two scales.
desk verdict The submitted manuscript is internally inconsistent: the abstract describes a molecular communication study with an ORF3a validation, but the full text is an unrelated video reasoning paper, so the central claims are unverifiable and the paper should be desk rejected. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the molecular communication channel model, in which a virus source acts as a transmitter releasing molecules that propagate by diffusion and absorption to a receiver (a cell, a sensor, or a body region). At macroscale and microscale, different channel parameters capture the physics of airborne versus tissue-borne spread. The second load-bearing mechanism is the mutation-identification strategy, which treats mutations as changes in the received molecular signal and benchmarks detection on the ORF3a protein's signal signature.
What would settle it
Measure airborne virus concentration decay as a function of distance and time in a real indoor environment and compare it with the attenuation and delay predicted by the macroscale MC channel model; if turbulence, air currents, or virion inactivation produce deviations beyond the model's error tolerance, the claimed match at macroscale fails.
Extended reading notes
Core claim
The paper's central claim is that viral transmission can be analyzed through molecular communication channels at two distinct scales: macroscale MC channels for airborne virus spread and microscale MC channels for spread within tissue. On both scales it proposes detection methods for the virus or infected individuals, a localization mechanism to find their positions, and an identification strategy to flag potential virus mutations. The mutation-identification strategy is validated by simulation using the ORF3a protein as a benchmark, illustrating that a molecular signal signature can distinguish a mutant from the wild type. Taken together, the paper positions epidemic prevention as an engine
Load-bearing premise
Real viral transmission, in air and in tissue, behaves like an engineered molecular communication channel governed by diffusion, absorption, and receiver detection, so that MC channel mathematics transfers to epidemic modeling.
Editorial extensions
If this is right
- If MC channel models match viral transmission, then measured virus concentration data can be fitted with MC channel parameters, turning epidemic spread into a parameter-estimation problem.
- Detection of infected individuals becomes a receiver-side signal-detection task, allowing the same detectors used in communication systems to flag the presence of a virus.
- Localization of infected individuals becomes a source-localization problem, for which MC frameworks already provide distance and direction estimators.
- The mutation-identification strategy implies that a library of known protein signal signatures could be screened automatically for novel variants, reducing reliance on manual genomic analysis.
- A successful IoBNT built on these ideas would use nanoscale devices as transmitters and receivers in a network that reports epidemic-relevant measurements in real time.
Reading between the lines
- The ORF3a benchmark validates the mutation-identification idea for one protein; an immediate testable extension is whether the same signal-signature approach distinguishes mutations in other SARS-CoV-2 proteins or in other respiratory viruses.
- The full text supplied with this submission is an unrelated manuscript on video reasoning; the only recoverable claims are those in the abstract. The asserted simulation validation and the macroscale/microscale channel match are therefore not inspectable here, so the paper's evidence base is thinner than its conclusions suggest.
- If the MC channel match holds only for passive diffusion, the framework may not transfer to real infections where immune clearance, active cellular transport, and host heterogeneity alter virion movement; validating the match on real indoor aerosol data would settle this.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The submission is nominally a paper on molecular communication (MC) for epidemic control in the Internet of Bio-Nano Things. The abstract claims that macro- and microscale MC channel models match viral transmission, that detection and localization methods are developed for both scales, and that a mutation identification strategy is validated by simulation using the ORF3a protein as a benchmark. However, the supplied full text is an entirely different paper, 'Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video Reasoning' (header arXiv:2508.04416v2, cs.CV). The full text contains no molecular communication content, no virus transmission modeling, no ORF3a simulation, and no discussion of epidemic IoBNT. All claims in the abstract are therefore unsupported by the submitted manuscript, and the technical content cannot be assessed.
Significance. If the abstract's claims were supported, the paper would offer a potentially significant cross-disciplinary contribution: a unified MC-based modeling and signal-processing framework for viral spread, detection, localization, and mutation identification, with a concrete ORF3a benchmark. However, none of this evidence is present in the submitted full text. There are no equations, no channel derivations, no error bars, no evaluation protocol, and no simulation details. The paper as supplied cannot be verified, and its contribution remains entirely at the level of an abstract.
major comments (3)
- [Full text, Sec. 1 (title and abstract)] The submitted full text is a different paper. Its title, abstract, and Section 1 describe a video-reasoning framework (VITAL) for multimodal large language models, with arXiv header arXiv:2508.04416v2 and subject cs.CV. It contains no mention of molecular communication, virus transmission, ORF3a, or epidemic IoBNT. This is a load-bearing mismatch: the abstract's central claims—MC channel models matching viral transmission at two scales, detection/localization methods, and a validated mutation identification strategy—require actual model derivations and simulations that are entirely absent. The scientific content of the claimed paper cannot be reviewed.
- [Abstract, first sentence and 'validated through simulation'] The abstract asserts that MC channels in macroscale and microscale scenarios 'match viral transmission in both scales' and that a mutation identification strategy 'is validated through simulation using the ORF3a protein as a benchmark.' No equations, parameter values, simulation setup, dataset, or quantitative results are given anywhere in the full text. The claimed ORF3a validation appears nowhere. Because the central claim hinges on this missing evidence, the paper does not currently support its own abstract.
- [Full text, Secs. 3 and 4 (unrelated content)] The body of the manuscript addresses a completely unrelated problem. Section 3 describes tool-augmented reinforcement learning for video reasoning, and Section 4 reports benchmark comparisons on video QA and temporal grounding. None of the tables or equations (e.g., Eq. (1)-(3), Tables 1-6) bear on the abstract's MC claims. This is not a matter of a flawed derivation that could be repaired locally; the submitted manuscript does not contain the claimed work.
minor comments (2)
- [Metadata] The arXiv identifier, title, authors, and subject classification of the full text do not match the abstract. The editor should verify the submission metadata and that the correct full text was attached.
- [References] All references cited in the full text pertain to video understanding, multimodal LLMs, and reinforcement learning; there are no references to molecular communication, bio-nano networks, or virology. If the intended paper exists, its reference list is entirely missing.
Circularity Check
No circularity can be established from the supplied evidence: the full text is an unrelated video-reasoning paper containing none of the MC channel, detection, localization, or ORF3a simulation content promised by the abstract.
full rationale
The abstract of arXiv:2508.04415 claims that MC channel models match viral transmission at macroscale and microscale and that a mutation-identification strategy is validated using an ORF3a simulation benchmark. However, the supplied full text is arXiv:2508.04416v2, 'Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video Reasoning'; it contains no molecular communication equations, no virus-transmission channel derivation, no detection/localization methods, and no ORF3a simulation. Under the hard rules, circularity may only be claimed when a specific reduction can be quoted and exhibited (e.g., Eq. X = Eq. Y by construction or a fitted parameter renamed as a prediction). No such reduction is available here. The concern that the ORF3a validation might benchmark the mutation-identification strategy against data generated under the same MC assumptions is a plausible risk about missing verification, but it is not a demonstrated circular step. An honest non-finding is therefore returned: no circularity can be identified in the available evidence. This score should not be read as endorsing the abstract's scientific claims; it reflects that the derivation chain needed for a circularity audit is absent from the provided document.
Assumptions & free parameters
assumptions (2)
- domain assumption Viral transmission at macro and micro scales can be faithfully represented by engineered molecular communication channel models (diffusion-based propagation and absorption equations).
- domain assumption Virus mutations leave a detectable signature in the molecular communication channel, inferable with signal processing methods.
Cite this review
Pith. "Pith review of Empowering Nanoscale Connectivity through Molecular Communication: A Case Study of Virus Infection." pith.science (2026). https://pith.science/paper/37NWCZX3
@misc{pith2026250804415,
author = {Pith},
title = {Pith review of: Empowering Nanoscale Connectivity through Molecular Communication: A Case Study of Virus Infection},
year = {2026},
howpublished = {\url{https://pith.science/paper/37NWCZX3}},
note = {Machine review of arXiv:2508.04415}
}
read the original abstract
The Internet of Bio-Nano Things (IoBNT), envisioned as a revolutionary healthcare paradigm, shows promise for epidemic control. This paper explores the potential of using molecular communication (MC) to address the challenges in constructing IoBNT for epidemic prevention, specifically focusing on modeling viral transmission, detecting the virus/infected individuals, and identifying virus mutations. First, the MC channels in macroscale and microscale scenarios are discussed to match viral transmission in both scales separately. Besides, the detection methods for these two scales are also studied, along with the localization mechanism designed for the virus/infected individuals. Moreover, an identification strategy is proposed to determine potential virus mutations, which is validated through simulation using the ORF3a protein as a benchmark. Finally, open research issues are discussed. In summary, this paper aims to analyze viral transmission through MC and combat viral spread using signal processing techniques within MC.
Forward citations
Cited by 1 Pith paper
-
Identifying high-impact consumers' behavioural changes for flexibility and demand reduction in a net-zero energy system
In a modeled net-zero Europe, shifting demand by 2 hours cuts system costs by 0.4% and curtailing 3.7% of peak electricity demand saves 0.9%, the two largest gains among four demand-side strategies.
Reference graph
Works this paper leans on
-
[1]
Shuai Bai, Keqin Chen, Xuejing Liu, Jialin Wang, Wenbin Ge, Sibo Song, Kai Dang, Peng Wang, Shijie Wang, Jun Tang, et al. Qwen2.5-vl technical report. arXiv preprint arXiv:2502.13923, 2025. 1, 4, 5, 12, 15
arXiv 2025
-
[2]
Univg-r1: Rea- soning guided universal visual grounding with reinforcement learning
Sule Bai, Mingxing Li, Yong Liu, Jing Tang, Haoji Zhang, Lei Sun, Xiangxiang Chu, and Yansong Tang. Univg-r1: Rea- soning guided universal visual grounding with reinforcement learning. arXiv preprint arXiv:2505.14231, 2025. 1, 2, 5
arXiv 2025
-
[3]
Auroracap: Efficient, performant video detailed captioning and a new benchmark
Wenhao Chai, Enxin Song, Yilun Du, Chenlin Meng, Vashisht Madhavan, Omer Bar-Tal, Jenq-Neng Hwang, Sain- ing Xie, and Christopher D Manning. Auroracap: Efficient, performant video detailed captioning and a new benchmark. arXiv preprint arXiv:2410.03051, 2024. 1
arXiv 2024
-
[4]
Distributed deep learning model for intelli- gent video surveillance systems with edge computing
Jianguo Chen, Kenli Li, Qingying Deng, Keqin Li, and Philip S Yu. Distributed deep learning model for intelli- gent video surveillance systems with edge computing. IEEE Transactions on Industrial Informatics, 2019. 1
2019
-
[5]
Rextime: A benchmark suite for reasoning-across-time in videos
Jr-Jen Chen, Yu-Chien Liao, Hsi-Che Lin, Yu-Chu Yu, Yen- Chun Chen, and Frank Wang. Rextime: A benchmark suite for reasoning-across-time in videos. Advances in Neural In- formation Processing Systems, 37:28662–28673, 2024. 4, 6, 14
2024
-
[6]
Yukang Chen, Wei Huang, Baifeng Shi, Qinghao Hu, Han- rong Ye, Ligeng Zhu, Zhijian Liu, Pavlo Molchanov, Jan Kautz, Xiaojuan Qi, et al. Scaling rl to long videos. arXiv preprint arXiv:2507.07966, 2025. 4, 5, 14
arXiv 2025
-
[7]
Longvila: Scaling long-context visual lan- guage models for long videos
Yukang Chen, Fuzhao Xue, Dacheng Li, Qinghao Hu, Ligeng Zhu, Xiuyu Li, Yunhao Fang, Haotian Tang, Shang Yang, Zhijian Liu, et al. Longvila: Scaling long-context visual lan- guage models for long videos. In The Thirteenth Interna- tional Conference on Learning Representations, 2025. 3
2025
-
[8]
Junhao Cheng, Yuying Ge, Teng Wang, Yixiao Ge, Jing Liao, and Ying Shan. Video-holmes: Can mllm think like holmes for complex video reasoning? arXiv preprint arXiv:2505.21374, 2025. 2
arXiv 2025
Show all 103 references
-
[9]
Gemini 2.5: Pushing the frontier with advanced reasoning, multimodality, long context, and next generation agentic capabilities
Gheorghe Comanici, Eric Bieber, Mike Schaekermann, Ice Pasupat, Noveen Sachdeva, Inderjit Dhillon, Marcel Blis- tein, Ori Ram, Dan Zhang, Evan Rosen, et al. Gemini 2.5: Pushing the frontier with advanced reasoning, multimodality, long context, and next generation agentic capab...
2025 arXiv
-
[10]
Motionlcm: Real-time controllable motion generation via latent consistency model
Wenxun Dai, Ling-Hao Chen, Jingbo Wang, Jinpeng Liu, Bo Dai, and Yansong Tang. Motionlcm: Real-time controllable motion generation via latent consistency model. In European Conference on Computer Vision , pages 390–408. Springer,
-
[11]
Video-of-thought: Step-by-step video reasoning from perception to cognition
Hao Fei, Shengqiong Wu, Wei Ji, Hanwang Zhang, Meishan Zhang, Mong-Li Lee, and Wynne Hsu. Video-of-thought: Step-by-step video reasoning from perception to cognition. In International Conference on Machine Learning , pages 13109–13125. PMLR, 2024. 1
2024
-
[12]
Retool: Reinforcement learning for strategic tool use in llms
Jiazhan Feng, Shijue Huang, Xingwei Qu, Ge Zhang, Yujia Qin, Baoquan Zhong, Chengquan Jiang, Jinxin Chi, and Wan- jun Zhong. Retool: Reinforcement learning for strategic tool use in llms. arXiv preprint arXiv:2504.11536, 2025. 3
2025 arXiv
-
[13]
Video-r1: Reinforcing video reasoning in mllms
Kaituo Feng, Kaixiong Gong, Bohao Li, Zonghao Guo, Yibing Wang, Tianshuo Peng, Junfei Wu, Xiaoying Zhang, Benyou Wang, and Xiangyu Yue. Video-r1: Reinforcing video reasoning in mllms. arXiv preprint arXiv:2503.21776,
-
[14]
Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis
Chaoyou Fu, Yuhan Dai, Yongdong Luo, Lei Li, Shuhuai Ren, Renrui Zhang, Zihan Wang, Chenyu Zhou, Yunhang Shen, Mengdan Zhang, et al. Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis. In Proceedings of the Computer Vision and Pa...
2025
-
[15]
Refocus: Visual editing as a chain of thought for structured image understanding
Xingyu Fu, Minqian Liu, Zhengyuan Yang, John Cor- ring, Yijuan Lu, Jianwei Yang, Dan Roth, Dinei Floren- cio, and Cha Zhang. Refocus: Visual editing as a chain of thought for structured image understanding. arXiv preprint arXiv:2501.05452, 2025. 3
2025 arXiv
-
[16]
Mist: Multi-modal iterative spatial- temporal transformer for long-form video question answer- ing
Difei Gao, Luowei Zhou, Lei Ji, Linchao Zhu, Yi Yang, and Mike Zheng Shou. Mist: Multi-modal iterative spatial- temporal transformer for long-form video question answer- ing. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 14773–1478...
2023
-
[17]
Tall: Temporal activity localization via language query
Jiyang Gao, Chen Sun, Zhenheng Yang, and Ram Nevatia. Tall: Temporal activity localization via language query. In Proceedings of the IEEE international conference on com- puter vision, pages 5267–5275, 2017. 1, 4, 6, 14
2017
-
[18]
Arc-hunyuan-video-7b: Structured video comprehension of real-world shorts
Yuying Ge, Yixiao Ge, Chen Li, Teng Wang, Junfu Pu, Yizhuo Li, Lu Qiu, Jin Ma, Lisheng Duan, Xinyu Zuo, et al. Arc-hunyuan-video-7b: Structured video comprehension of real-world shorts. arXiv preprint arXiv:2507.20939 , 2025. 2
2025 arXiv
-
[19]
Context-guided spatio-temporal video grounding
Xin Gu, Heng Fan, Yan Huang, Tiejian Luo, and Libo Zhang. Context-guided spatio-temporal video grounding. In Pro- ceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 18330–18339, 2024. 1
2024
-
[20]
Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning
Daya Guo, Dejian Yang, Haowei Zhang, Junxiao Song, Ruoyu Zhang, Runxin Xu, Qihao Zhu, Shirong Ma, Peiyi Wang, Xiao Bi, et al. Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning. arXiv preprint arXiv:2501.12948, 2025. 1, 2, 3
2025 arXiv
-
[21]
Roomtour3d: Geometry-aware video- instruction tuning for embodied navigation
Mingfei Han, Liang Ma, Kamila Zhumakhanova, Ekaterina Radionova, Jingyi Zhang, Xiaojun Chang, Xiaodan Liang, and Ivan Laptev. Roomtour3d: Geometry-aware video- instruction tuning for embodied navigation. In Proceedings of the Computer Vision and Pattern Recognition Conference,...
2025
-
[22]
Revisionllm: Recursive vision-language model for temporal grounding in hour-long videos
Tanveer Hannan, Md Mohaiminul Islam, Jindong Gu, Thomas Seidl, and Gedas Bertasius. Revisionllm: Recursive vision-language model for temporal grounding in hour-long videos. In Proceedings of the Computer Vision and Pattern Recognition Conference, pages 19012–19022, 2025. 3
2025
-
[23]
Vlab: Enhancing video language pre-training by fea- ture adapting and blending
Xingjian He, Sihan Chen, Fan Ma, Zhicheng Huang, Xiaojie Jin, Zikang Liu, Dongmei Fu, Yi Yang, Jing Liu, and Jiashi Feng. Vlab: Enhancing video language pre-training by fea- ture adapting and blending. IEEE Transactions on Multime- dia, 2024. 3
2024
-
[24]
Video-mmmu: Evaluating knowledge acquisition from multi-discipline pro- fessional videos
Kairui Hu, Penghao Wu, Fanyi Pu, Wang Xiao, Yuanhan Zhang, Xiang Yue, Bo Li, and Ziwei Liu. Video-mmmu: Evaluating knowledge acquisition from multi-discipline pro- fessional videos. arXiv preprint arXiv:2501.13826, 2025. 6, 14
2025 arXiv
-
[25]
Vi- sual sketchpad: Sketching as a visual chain of thought for multimodal language models
Yushi Hu, Weijia Shi, Xingyu Fu, Dan Roth, Mari Ostendorf, Luke Zettlemoyer, Noah A Smith, and Ranjay Krishna. Vi- sual sketchpad: Sketching as a visual chain of thought for multimodal language models. Advances in Neural Informa- tion Processing Systems, 37:139348–139379, 2024. 3
2024
-
[26]
Vision-r1: Incentivizing reasoning capability in multimodal large language models
Wenxuan Huang, Bohan Jia, Zijie Zhai, Shaosheng Cao, Zheyu Ye, Fei Zhao, Zhe Xu, Yao Hu, and Shaohui Lin. Vision-r1: Incentivizing reasoning capability in multimodal large language models. arXiv preprint arXiv:2503.06749 ,
-
[27]
Real-time video recommendation ex- ploration
Yanxiang Huang, Bin Cui, Jie Jiang, Kunqian Hong, Wenyu Zhang, and Yiran Xie. Real-time video recommendation ex- ploration. In Proceedings of the 2016 international confer- ence on management of data, pages 35–46, 2016. 1
2016
-
[28]
Gpt-4o system card
Aaron Hurst, Adam Lerer, Adam P Goucher, Adam Perel- man, Aditya Ramesh, Aidan Clark, AJ Ostrow, Akila Weli- hinda, Alan Hayes, Alec Radford, et al. Gpt-4o system card. arXiv preprint arXiv:2410.21276, 2024. 1
2024 arXiv
-
[29]
Openai o1 system card
Aaron Jaech, Adam Kalai, Adam Lerer, Adam Richard- son, Ahmed El-Kishky, Aiden Low, Alec Helyar, Aleksander Madry, Alex Beutel, Alex Carney, et al. Openai o1 system card. arXiv preprint arXiv:2412.16720, 2024. 2
2024 arXiv
-
[30]
Search- r1: Training llms to reason and leverage search engines with reinforcement learning
Bowen Jin, Hansi Zeng, Zhenrui Yue, Jinsung Yoon, Sercan Arik, Dong Wang, Hamed Zamani, and Jiawei Han. Search- r1: Training llms to reason and leverage search engines with reinforcement learning. arXiv preprint arXiv:2503.09516 ,
-
[31]
Dense-captioning events in videos
Ranjay Krishna, Kenji Hata, Frederic Ren, Li Fei-Fei, and Juan Carlos Niebles. Dense-captioning events in videos. In Proceedings of the IEEE international conference on com- puter vision, pages 706–715, 2017. 4, 6, 14
2017
-
[32]
Gonzalez, Hao Zhang, and Ion Stoica
Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph E. Gonzalez, Hao Zhang, and Ion Stoica. Efficient memory management for large language model serving with pagedattention. In Pro- ceedings of the ACM SIGOPS 29th Symposium on Operating Syste...
2023
-
[33]
Large-scale content- only video recommendation
Joonseok Lee and Sami Abu-El-Haija. Large-scale content- only video recommendation. In Proceedings of the IEEE International Conference on Computer Vision Workshops , pages 987–995, 2017. 1
2017
-
[34]
Llava-onevision: Easy visual task transfer
Bo Li, Yuanhan Zhang, Dong Guo, Renrui Zhang, Feng Li, Hao Zhang, Kaichen Zhang, Peiyuan Zhang, Yanwei Li, Zi- wei Liu, et al. Llava-onevision: Easy visual task transfer. arXiv preprint arXiv:2408.03326, 2024. 1
2024 arXiv
-
[35]
Imag- ine while reasoning in space: Multimodal visualization-of- thought
Chengzu Li, Wenshan Wu, Huanyu Zhang, Yan Xia, Shaoguang Mao, Li Dong, Ivan Vuli ´c, and Furu Wei. Imag- ine while reasoning in space: Multimodal visualization-of- thought. arXiv preprint arXiv:2501.07542, 2025. 3
2025 arXiv
-
[36]
Start: Self-taught reasoner with tools
Chengpeng Li, Mingfeng Xue, Zhenru Zhang, Jiaxi Yang, Beichen Zhang, Xiang Wang, Bowen Yu, Binyuan Hui, Jun- yang Lin, and Dayiheng Liu. Start: Self-taught reasoner with tools. arXiv preprint arXiv:2503.04625, 2025. 3
2025 arXiv
-
[37]
Dyfo: A training-free dynamic focus visual search for enhancing lmms in fine-grained visual understanding
Geng Li, Jinglin Xu, Yunzhen Zhao, and Yuxin Peng. Dyfo: A training-free dynamic focus visual search for enhancing lmms in fine-grained visual understanding. In Proceedings of the Computer Vision and Pattern Recognition Conference, pages 9098–9108, 2025. 3
2025
-
[38]
Reinforcement learning tun- ing for videollms: Reward design and data efficiency
Hongyu Li, Songhao Han, Yue Liao, Junfeng Luo, Jialin Gao, Shuicheng Yan, and Si Liu. Reinforcement learning tun- ing for videollms: Reward design and data efficiency. arXiv preprint arXiv:2506.01908, 2025. 2, 6
2025 arXiv
-
[39]
Videochat-r1: Enhancing spatio-temporal perception via re- inforcement fine-tuning
Xinhao Li, Ziang Yan, Desen Meng, Lu Dong, Xiangyu Zeng, Yinan He, Yali Wang, Yu Qiao, Yi Wang, and Limin Wang. Videochat-r1: Enhancing spatio-temporal perception via re- inforcement fine-tuning. arXiv preprint arXiv:2504.06958 ,
-
[40]
Torl: Scaling tool-integrated rl
Xuefeng Li, Haoyang Zou, and Pengfei Liu. Torl: Scaling tool-integrated rl. arXiv preprint arXiv:2503.23383, 2025. 3
2025 arXiv
-
[41]
Llama-vid: An image is worth 2 tokens in large language models
Yanwei Li, Chengyao Wang, and Jiaya Jia. Llama-vid: An image is worth 2 tokens in large language models. In European Conference on Computer Vision , pages 323–340. Springer, 2024. 3
2024
-
[42]
Evolvenav: Self-improving embodied reasoning for llm-based vision-language navigation
Bingqian Lin, Yunshuang Nie, Khun Loun Zai, Ziming Wei, Mingfei Han, Rongtao Xu, Minzhe Niu, Jianhua Han, Liang Lin, Cewu Lu, et al. Evolvenav: Self-improving embodied reasoning for llm-based vision-language navigation. arXiv preprint arXiv:2506.01551, 2025. 1
2025
-
[43]
Rouge: A package for automatic evaluation of summaries
Chin-Yew Lin. Rouge: A package for automatic evaluation of summaries. In Text summarization branches out , pages 74–81, 2004. 13
2004
-
[44]
Plan, posture and go: Towards open-vocabulary text-to-motion generation
Jinpeng Liu, Wenxun Dai, Chunyu Wang, Yiji Cheng, Yan- song Tang, and Xin Tong. Plan, posture and go: Towards open-vocabulary text-to-motion generation. In European Conference on Computer Vision , pages 445–463. Springer,
-
[45]
Llava-plus: Learning to use tools for creating multi- modal agents
Shilong Liu, Hao Cheng, Haotian Liu, Hao Zhang, Feng Li, Tianhe Ren, Xueyan Zou, Jianwei Yang, Hang Su, Jun Zhu, et al. Llava-plus: Learning to use tools for creating multi- modal agents. In European conference on computer vision , pages 126–142. Springer, 2024. 3
2024
-
[46]
Seg-zero: Reasoning-chain guided segmentation via cognitive reinforcement
Yuqi Liu, Bohao Peng, Zhisheng Zhong, Zihao Yue, Fanbin Lu, Bei Yu, and Jiaya Jia. Seg-zero: Reasoning-chain guided segmentation via cognitive reinforcement. arXiv preprint arXiv:2503.06520, 2025. 2
2025 arXiv
-
[47]
Visual- rft: Visual reinforcement fine-tuning
Ziyu Liu, Zeyi Sun, Yuhang Zang, Xiaoyi Dong, Yuhang Cao, Haodong Duan, Dahua Lin, and Jiaqi Wang. Visual- rft: Visual reinforcement fine-tuning. arXiv preprint arXiv:2503.01785, 2025. 1, 2
2025 arXiv
-
[48]
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter. Decoupled weight decay regularization. arXiv preprint arXiv:1711.05101, 2017. 5
2017 arXiv
-
[49]
Vista-llama: Reducing hallucination in video language models via equal distance to visual tokens
Fan Ma, Xiaojie Jin, Heng Wang, Yuchen Xian, Jiashi Feng, and Yi Yang. Vista-llama: Reducing hallucination in video language models via equal distance to visual tokens. In Pro- ceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 13151–13160, 2024. 3
2024
-
[50]
Slowfocus: Enhancing fine-grained temporal understanding in video llm
Ming Nie, Dan Ding, Chunwei Wang, Yuanfan Guo, Jian- hua Han, Hang Xu, and Li Zhang. Slowfocus: Enhancing fine-grained temporal understanding in video llm. Advances in Neural Information Processing Systems, 37:81808–81835,
-
[51]
Thinking with images
OpenAI. Thinking with images. https://openai.com/index/ thinking-with-images, 2025. 3
2025
-
[52]
Deepvideo-r1: Video reinforcement fine- tuning via difficulty-aware regressive grpo
Jinyoung Park, Jeehye Na, Jinyoung Kim, and Hyun- woo J Kim. Deepvideo-r1: Video reinforcement fine- tuning via difficulty-aware regressive grpo. arXiv preprint arXiv:2506.07464, 2025. 2
2025
-
[53]
Lmm-r1: Empowering 3b lmms with strong reasoning abilities through two-stage rule-based rl
Yingzhe Peng, Gongrui Zhang, Miaosen Zhang, Zhiyuan You, Jie Liu, Qipeng Zhu, Kai Yang, Xingzhong Xu, Xin Geng, and Xu Yang. Lmm-r1: Empowering 3b lmms with strong reasoning abilities through two-stage rule-based rl. arXiv preprint arXiv:2503.07536, 2025. 2
2025 arXiv
-
[54]
Vlm-r1: A stable and generaliz- able r1-style large vision-language model
Haozhan Shen, Peng Liu, Jingcheng Li, Chunxin Fang, Yibo Ma, Jiajia Liao, Qiaoli Shen, Zilun Zhang, Kangjia Zhao, Qianqian Zhang, et al. Vlm-r1: A stable and generaliz- able r1-style large vision-language model. arXiv preprint arXiv:2504.07615, 2025. 1, 2
2025 arXiv
-
[55]
Accurate and fast compressed video captioning
Yaojie Shen, Xin Gu, Kai Xu, Heng Fan, Longyin Wen, and Libo Zhang. Accurate and fast compressed video captioning. In Proceedings of the IEEE/CVF international conference on computer vision, pages 15558–15567, 2023. 1
2023
-
[56]
Hybridflow: A flexible and efficient rlhf frame- work
Guangming Sheng, Chi Zhang, Zilingfeng Ye, Xibin Wu, Wang Zhang, Ru Zhang, Yanghua Peng, Haibin Lin, and Chuan Wu. Hybridflow: A flexible and efficient rlhf frame- work. In Proceedings of the Twentieth European Conference on Computer Systems, pages 1279–1297, 2025. 5
2025
-
[57]
Enhanc- ing video-llm reasoning via agent-of-thoughts distillation
Yudi Shi, Shangzhe Di, Qirui Chen, and Weidi Xie. Enhanc- ing video-llm reasoning via agent-of-thoughts distillation. In Proceedings of the Computer Vision and Pattern Recognition Conference, pages 8523–8533, 2025. 1
2025
-
[58]
Moviechat: From dense token to sparse memory for long video understanding
Enxin Song, Wenhao Chai, Guanhong Wang, Yucheng Zhang, Haoyang Zhou, Feiyang Wu, Haozhe Chi, Xun Guo, Tian Ye, Yanting Zhang, et al. Moviechat: From dense token to sparse memory for long video understanding. In Proceed- ings of the IEEE/CVF Conference on Computer Vision and Pa...
2024
-
[59]
Openthinkimg: Learning to think with images via visual tool reinforcement learning
Zhaochen Su, Linjie Li, Mingyang Song, Yunzhuo Hao, Zhengyuan Yang, Jun Zhang, Guanjie Chen, Jiawei Gu, Jun- tao Li, Xiaoye Qu, et al. Openthinkimg: Learning to think with images via visual tool reinforcement learning. arXiv preprint arXiv:2505.08617, 2025. 3
2025 arXiv
-
[60]
Visual agents as fast and slow thinkers
Guangyan Sun, Mingyu Jin, Zhenting Wang, Cheng-Long Wang, Siqi Ma, Qifan Wang, Tong Geng, Ying Nian Wu, Yongfeng Zhang, and Dongfang Liu. Visual agents as fast and slow thinkers. In The Thirteenth International Confer- ence on Learning Representations, 2025. 3
2025
-
[61]
Kimi k2: Open agen- tic intelligence, 2025
Kimi Team, Yifan Bai, Yiping Bao, Guanduo Chen, Jiahao Chen, Ningxin Chen, Ruijue Chen, Yanru Chen, Yuankun Chen, Yutian Chen, Zhuofu Chen, et al. Kimi k2: Open agen- tic intelligence, 2025. 3
2025
-
[62]
Kimi-vl technical report
Kimi Team, Angang Du, Bohong Yin, Bowei Xing, Bowen Qu, Bowen Wang, Cheng Chen, Chenlin Zhang, Chenzhuang Du, Chu Wei, et al. Kimi-vl technical report. arXiv preprint arXiv:2504.07491, 2025. 1
2025 arXiv
-
[63]
Vidi: Large multimodal models for video understanding and editing
Vidi Team, Celong Liu, Chia-Wen Kuo, Dawei Du, Fan Chen, Guang Chen, Jiamin Yuan, Lingxi Zhang, Lu Guo, Lusha Li, et al. Vidi: Large multimodal models for video understanding and editing. arXiv preprint arXiv:2504.15681, 2025. 5, 14
2025 arXiv
-
[64]
Hermes 3 technical report
Ryan Teknium, Jeffrey Quesnelle, and Chen Guang. Hermes 3 technical report. arXiv preprint arXiv:2408.11857 , 2024. 12
2024 arXiv
-
[65]
Video surveil- lance systems-current status and future trends
Vassilios Tsakanikas and Tasos Dagiuklas. Video surveil- lance systems-current status and future trends. Computers & Electrical Engineering, 70:736–753, 2018. 1
2018
-
[66]
Traceable evidence enhanced vi- sual grounded reasoning: Evaluation and methodology.arXiv preprint arXiv:2507.07999, 2025
Haochen Wang, Xiangtai Li, Zilong Huang, Anran Wang, Jiacong Wang, Tao Zhang, Jiani Zheng, Sule Bai, Zijian Kang, Jiashi Feng, et al. Traceable evidence enhanced vi- sual grounded reasoning: Evaluation and methodology.arXiv preprint arXiv:2507.07999, 2025. 2
2025
-
[67]
Vgr: Visual grounded reasoning
Jiacong Wang, Zijian Kang, Haochen Wang, Haiyong Jiang, Jiawen Li, Bohong Wu, Ya Wang, Jiao Ran, Xiao Liang, Chao Feng, et al. Vgr: Visual grounded reasoning. arXiv preprint arXiv:2506.11991, 2025. 2
2025 arXiv
-
[68]
Videorft: Incentivizing video reasoning capability in mllms via reinforced fine-tuning
Qi Wang, Yanrui Yu, Ye Yuan, Rui Mao, and Tianfei Zhou. Videorft: Incentivizing video reasoning capability in mllms via reinforced fine-tuning. arXiv preprint arXiv:2505.12434,
-
[69]
Executable code actions elicit better llm agents
Xingyao Wang, Yangyi Chen, Lifan Yuan, Yizhe Zhang, Yun- zhu Li, Hao Peng, and Heng Ji. Executable code actions elicit better llm agents. In Forty-first International Conference on Machine Learning, 2024. 3
2024
-
[70]
Driving into the future: Multiview visual forecasting and planning with world model for au- tonomous driving
Yuqi Wang, Jiawei He, Lue Fan, Hongxin Li, Yuntao Chen, and Zhaoxiang Zhang. Driving into the future: Multiview visual forecasting and planning with world model for au- tonomous driving. In Proceedings of the IEEE/CVF Con- ference on Computer Vision and Pattern Recognition , p...
2024
-
[71]
Uni-adafocus: spatial-temporal dynamic computation for video recognition
Yulin Wang, Haoji Zhang, Yang Yue, Shiji Song, Chao Deng, Junlan Feng, and Gao Huang. Uni-adafocus: spatial-temporal dynamic computation for video recognition. IEEE Transac- tions on Pattern Analysis and Machine Intelligence, 2024. 3
2024
-
[72]
Time-r1: Post-training large vision lan- guage model for temporal video grounding
Ye Wang, Ziheng Wang, Boshen Xu, Yang Du, Kejun Lin, Zihan Xiao, Zihao Yue, Jianzhong Ju, Liang Zhang, Dingyi Yang, et al. Time-r1: Post-training large vision lan- guage model for temporal video grounding. arXiv preprint arXiv:2503.13377, 2025. 2
2025 arXiv
-
[73]
Perception in reflection
Yana Wei, Liang Zhao, Kangheng Lin, En Yu, Yuang Peng, Runpei Dong, Jianjian Sun, Haoran Wei, Zheng Ge, Xi- angyu Zhang, et al. Perception in reflection. arXiv preprint arXiv:2504.07165, 2025. 3
2025 arXiv
-
[74]
Longvlm: Efficient long video understand- ing via large language models
Yuetian Weng, Mingfei Han, Haoyu He, Xiaojun Chang, and Bohan Zhuang. Longvlm: Efficient long video understand- ing via large language models. In European Conference on Computer Vision, pages 453–470. Springer, 2024. 3
2024
-
[75]
Visual chatgpt: Talking, draw- ing and editing with visual foundation models.arXiv preprint arXiv:2303.04671, 2023
Chenfei Wu, Shengming Yin, Weizhen Qi, Xiaodong Wang, Zecheng Tang, and Nan Duan. Visual chatgpt: Talking, draw- ing and editing with visual foundation models.arXiv preprint arXiv:2303.04671, 2023. 3
2023 arXiv
-
[76]
Towards long-form video understanding
Chao-Yuan Wu and Philipp Krahenbuhl. Towards long-form video understanding. In Proceedings of the IEEE/CVF Con- ference on Computer Vision and Pattern Recognition , pages 1884–1894, 2021. 3
2021
-
[77]
V?: Guided visual search as a core mechanism in multimodal llms
Penghao Wu and Saining Xie. V?: Guided visual search as a core mechanism in multimodal llms. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 13084–13094, 2024. 3
2024
-
[78]
Number it: Temporal grounding videos like flipping manga
Yongliang Wu, Xinting Hu, Yuyang Sun, Yizhou Zhou, Wenbo Zhu, Fengyun Rao, Bernt Schiele, and Xu Yang. Number it: Temporal grounding videos like flipping manga. In Proceedings of the Computer Vision and Pattern Recogni- tion Conference, pages 13754–13765, 2025. 14
2025
-
[79]
Can i trust your answer? visually grounded video question answering
Junbin Xiao, Angela Yao, Yicong Li, and Tat-Seng Chua. Can i trust your answer? visually grounded video question answering. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 13204– 13214, 2024. 4, 6, 14
2024
-
[80]
Vidchapters-7m: Video chapters at scale
Antoine Yang, Arsha Nagrani, Ivan Laptev, Josef Sivic, and Cordelia Schmid. Vidchapters-7m: Video chapters at scale. Advances in Neural Information Processing Systems , 36: 49428–49444, 2023. 4, 5, 14
2023
-
[81]
Qwen3 technical report
An Yang, Anfeng Li, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Gao, Chengen Huang, Chenxu Lv, et al. Qwen3 technical report. arXiv preprint arXiv:2505.09388, 2025. 3
2025 arXiv
-
[82]
Generalized predictive model for autonomous driving
Jiazhi Yang, Shenyuan Gao, Yihang Qiu, Li Chen, Tianyu Li, Bo Dai, Kashyap Chitta, Penghao Wu, Jia Zeng, Ping Luo, et al. Generalized predictive model for autonomous driving. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 14662–1467...
2024
-
[83]
Thinking in space: How multimodal large language models see, remember, and recall spaces
Jihan Yang, Shusheng Yang, Anjali W Gupta, Rilyn Han, Li Fei-Fei, and Saining Xie. Thinking in space: How multimodal large language models see, remember, and recall spaces. In Proceedings of the Computer Vision and Pattern Recognition Conference, pages 10632–10643, 2025. 6, 14
2025
-
[84]
Gpt4tools: Teaching large language model to use tools via self-instruction
Rui Yang, Lin Song, Yanwei Li, Sijie Zhao, Yixiao Ge, Xiu Li, and Ying Shan. Gpt4tools: Teaching large language model to use tools via self-instruction. Advances in Neural Informa- tion Processing Systems, 36:71995–72007, 2023. 3
2023
-
[85]
Language-aware vi- sion transformer for referring segmentation
Zhao Yang, Jiaqi Wang, Xubing Ye, Yansong Tang, Kai Chen, Hengshuang Zhao, and Philip HS Torr. Language-aware vi- sion transformer for referring segmentation. IEEE Transac- tions on Pattern Analysis and Machine Intelligence, 2024. 2
2024
-
[86]
CLEVRER: collision events for video representation and rea- soning
Kexin Yi, Chuang Gan, Yunzhu Li, Pushmeet Kohli, Ji- ajun Wu, Antonio Torralba, and Joshua B Tenenbaum. CLEVRER: collision events for video representation and rea- soning. In ICLR, 2020. 1
2020
-
[87]
Self-chained image-language model for video localization and question answering
Shoubin Yu, Jaemin Cho, Prateek Yadav, and Mohit Bansal. Self-chained image-language model for video localization and question answering. Advances in Neural Information Processing Systems, 36:76749–76771, 2023. 3
2023
-
[88]
A framework for training large language models for code generation via proximal policy optimization
Chi Zhang, Guangming Sheng, Siyao Liu, Jiahao Li, Ziyuan Feng, Zherui Liu, Xin Liu, Xiaoying Jia, Yanghua Peng, Haibin Lin, et al. A framework for training large language models for code generation via proximal policy optimization. In NL2Code Workshop of ACM KDD, 2024. 5
2024
-
[89]
Flash-vstream: Efficient real- time understanding for long video streams
Haoji Zhang, Yiqin Wang, Yansong Tang, Yong Liu, Ji- ashi Feng, and Xiaojie Jin. Flash-vstream: Efficient real- time understanding for long video streams. arXiv preprint arXiv:2506.23825, 2025. 3
2025 arXiv
-
[90]
Long context transfer from lan- guage to vision
Peiyuan Zhang, Kaichen Zhang, Bo Li, Guangtao Zeng, Jingkang Yang, Yuanhan Zhang, Ziyue Wang, Haoran Tan, Chunyuan Li, and Ziwei Liu. Long context transfer from lan- guage to vision. arXiv preprint arXiv:2406.16852, 2024. 3
2024 arXiv
-
[91]
Logo: A long-form video dataset for group action quality assessment
Shiyi Zhang, Wenxun Dai, Sujia Wang, Xiangwei Shen, Ji- wen Lu, Jie Zhou, and Yansong Tang. Logo: A long-form video dataset for group action quality assessment. In Pro- ceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 2405–2414, 2023. 3
2023
-
[92]
Mmvu: Measuring expert-level multi- discipline video understanding
Yilun Zhao, Haowei Zhang, Lujing Xie, Tongyan Hu, Guo Gan, Yitao Long, Zhiyuan Hu, Weiyuan Chen, Chuhan Li, Zhijian Xu, et al. Mmvu: Measuring expert-level multi- discipline video understanding. In Proceedings of the Com- puter Vision and Pattern Recognition Conference , pages...
2025
-
[93]
Deepresearcher: Scaling deep research via reinforcement learning in real- world environments
Yuxiang Zheng, Dayuan Fu, Xiangkun Hu, Xiaojie Cai, Lyu- manshan Ye, Pengrui Lu, and Pengfei Liu. Deepresearcher: Scaling deep research via reinforcement learning in real- world environments. arXiv preprint arXiv:2504.03160, 2025. 3
2025 arXiv
-
[94]
Deepeyes: In- centivizing” thinking with images” via reinforcement learn- ing
Ziwei Zheng, Michael Yang, Jack Hong, Chenxiao Zhao, Guohai Xu, Le Yang, Chao Shen, and Xing Yu. Deepeyes: In- centivizing” thinking with images” via reinforcement learn- ing. arXiv preprint arXiv:2505.14362, 2025. 3
2025 arXiv
-
[95]
Temporal relational reasoning in videos
Bolei Zhou, Alex Andonian, Aude Oliva, and Antonio Tor- ralba. Temporal relational reasoning in videos. In Proceed- ings of the European conference on computer vision (ECCV), pages 803–818, 2018. 1 Supplementary Material To facilitate a deeper understanding and reproducibility...
2018
-
[96]
Video clip captioning tool : This tool takes the ����� and ��� timestamps of a video clip as input, and gener- ates a descriptive ������� for the specified segment
-
[97]
It outputs an ������ to the given question based on the visual content of the specified clip
Video clip QA tool : This tool receives the ����� and ��� timestamps of a video clip, together with a natural language �������� , as input. It outputs an ������ to the given question based on the visual content of the specified clip
-
[98]
Video clipping tool: This tool takes the����� and ��� timestamps as input and outputs the visual content (rep- resented as ������ ������ ) corresponding to the se- lected video segment. Tab. 7 summarizes the input and output formats of these vi- sual tools. Tool Name Inputs Ou...
-
[99]
B.1 Training Details The training configurations are listed in Tab
during training and evaluation to print absolute times- tamps on frames, providing additional temporal information for MLLMs for accurate temporal perception. B.1 Training Details The training configurations are listed in Tab. 10. Gener- ally, we split the training procedure i...
-
[100]
In the sec- ond round, it generates the tool call
is prompted to generate a thinking process. In the sec- ond round, it generates the tool call. In the third round, it generates the reflection thinking process and the concluded answer. Notably, in the second round, the model receives a pre- defined tool parameter suggestion f...
-
[101]
The chain-of-thought is incomplete or does not reach a final answer
-
[102]
The generated answer does not match the ground truth
-
[103]
ground truth
The sample contains irrelevant or off-topic content, e.g., direct description of “ground truth” or “suggestion”. After data post-processing, we obtain the final MTVR- CoT-72k dataset, which consists of high-quality and well- formatted samples suitable for cold-start supervised...
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.