Pith. sign in

REVIEW 3 major objections 3 minor 104 references

"Harmless to You, Hurtful to Me!": Investigating the Detection of Toxic Languages Grounded in the Perspective of Youth

T0 review · 3 major / 3 minor · reviewed 2026-08-06 · deepseek-v4-flash

Pith's one-line read Youth and adults do not agree on what counts as toxic speech online.

desk verdict The review packet contains the wrong full text; based on the abstract alone, the youth-toxicity study is plausible and worth reviewing once the actual manuscript is obtained. read the letter →

arxiv 2508.02094 v1 pith:5F3D7Z52 submitted 2025-08-04 cs.CL cs.HC

classification cs.CLcs.HC
keywords youthtoxicityperceptionChineseyouth-toxicitydatasetdetectioncontextualfactorssourceofutterancemeta-informationsocialmediaharmadultversusdivergence
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper argues that toxicity is not a fixed property of words: young social-media users experience some language as harmful that adults would call harmless. To show this, the authors built what they describe as the first Chinese youth-toxicity dataset, targeting Chinese youth, and analyzed how youth judgments relate to contextual factors. They report that youth perception is tied to the source of an utterance and to text-related features, and that feeding this meta-information into existing toxicity detectors improves overall accuracy. If the claim holds, age-aware moderation becomes a measurable and technologically actionable goal rather than a vague safety concern.

What carries the argument

The central object is the Chinese youth-toxicity dataset, a curated collection of social-media utterances labeled for how young people perceive them. The mechanism that carries the argument is the coupling of that dataset with meta-information about utterances—who or what produced them, their source, and text-related features—fed into existing toxicity detection models. These meta-features are what the accuracy comparisons manipulate: the paper attributes the reported improvement to their addition, not to any new detection architecture.

What would settle it

Take a fresh, independently recruited panel of adolescents, show them a held-out sample from the dataset, and compare their toxicity judgments with the published labels and with adult judgments; if the original labels match adult intuitions rather than teen self-reports, or if the teen panel's agreement is near chance, the reported accuracy gains from meta-information would not demonstrate youth-specific toxicity.

Watch

Extended reading notes

Core claim

Youth-specific toxicity exists, and it can be captured and exploited. The paper's central claim is that languages perceived as nontoxic by adults but toxic by youth form a distinct category, and that these judgments are linked to contextual factors such as the source of an utterance and text-related features. The authors state that they constructed the first Chinese youth-toxicity dataset, performed extensive analysis of those features, and found that incorporating this meta-information into current toxicity detection methods significantly improves accuracy overall. The result positions youth-centered toxicity as a distinct object of study: a listener-relative judgment that standard adult-oriented detectors miss.

Load-bearing premise

The load-bearing premise is that the dataset's labels genuinely record what teenagers find toxic rather than what adults assume teenagers find toxic; the abstract does not report who labeled the data, the age range, or agreement among annotators.

Editorial extensions

If this is right

  • Existing toxicity classifiers are incomplete for youth audiences, and adding utterance-source and text-related meta-features should raise their accuracy on youth-annotated content.
  • A reusable Chinese youth-toxicity resource now exists for studying language that adults overlook but young users experience as harmful.
  • Youth-centered moderation could treat toxicity as listener-relative rather than a fixed property of words alone.
  • Future detectors could be evaluated separately for youth and adult populations instead of against one global gold standard.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The same listener-relative logic may extend beyond age: other groups who experience harm differently, such as marginalized communities, might need population-specific harm models rather than one universal toxicity score.
  • A testable extension would be a longitudinal study asking whether the same utterances lose their youth-toxicity as annotators age, which would show whether the boundary is developmental or merely generational slang.
  • The source-of-utterance finding suggests power relations matter—who speaks, a peer or an authority, may change perceived harm—so future benchmarks could explicitly vary speaker identity to test this mechanism.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 3 minor

Summary. The submission, as provided, consists of a title and abstract concerning youth-specific toxicity detection in Chinese social media, followed by a full text that is the paper "VLM4D: Towards Spatiotemporal Awareness in Vision Language Models" (arXiv:2508.02095v2), which is a computer vision benchmark paper with no connection to toxic language detection. The abstract claims the construction of the first Chinese "youth-toxicity" dataset, an analysis showing that youth perceptions are associated with contextual factors, and that incorporating meta-information significantly improves toxicity detection accuracy. None of these claims are supported by the supplied full text, which instead reports on a spatiotemporal video-question-answering benchmark, evaluations of vision-language models, and exploratory fine-tuning experiments. Because the manuscript body is entirely unrelated to the abstract's claims, the central results of the submission are unverifiable from the submitted materials.

Significance. If the intended study were properly delivered, it could be a meaningful contribution: a youth-grounded toxicity dataset for Chinese social media, with evidence that utterance source and text features improve detection, would be valuable to content moderation and computational social science. The motivation—that youth risk perception differs from adult assumptions—is plausible and well worth investigating. However, as submitted, the paper provides no dataset, no experimental results, no analysis, and no toxicity-related content whatsoever; the full text is an unrelated vision-language benchmark. The claimed contributions therefore have no supporting evidence in the manuscript, and the potential significance cannot be assessed.

major comments (3)
  1. [Full text (entire manuscript)] The full text provided for review is the paper "VLM4D: Towards Spatiotemporal Awareness in Vision Language Models" (arXiv:2508.02095v2), which is entirely unrelated to the submission's title and abstract about youth toxicity detection. The abstract promises a Chinese youth-toxicity dataset, analyses of youth perceptions, and improved toxicity detection; the full text describes a spatiotemporal QA benchmark for VLMs. There is no overlap in subject matter, data, methodology, or results. This is a load-bearing failure: the central claims of the submission are completely unsupported by the manuscript body, and no local revision can fix the absence of the actual paper.
  2. [Abstract] The abstract alone, the only text relevant to the claimed topic, lacks the evidence needed to evaluate the central empirical claims. It provides no dataset statistics, no information about annotator recruitment (who labeled the data, what age range defines "youth", how participants were sourced), no inter-annotator agreement, no baseline systems, no description of a held-out evaluation protocol, and no error bars. The statement that incorporating meta-information "significantly improves accuracy overall" is an empirical assertion with no quantitative support anywhere in the manuscript.
  3. [Abstract; evaluation protocol] The abstract does not state whether the contextual meta-features (utterance source and text-related features) were selected using the same corpus on which the accuracy improvement is measured, nor whether the evaluation was performed on held-out data. As written, the reported improvement could be an in-sample artifact of feature selection, and the manuscript gives no way to test this. This concern is separate from the full-text mismatch but is also load-bearing for the claimed practical benefit.
minor comments (3)
  1. [Abstract] The terms "meta information" and "text-related features" are used without definition; the reader cannot tell what specific features were added to the toxicity detectors.
  2. [Abstract] The construct "youth-toxicity" is presented in quotes and hyphenated forms without a precise operational definition, and the age range for "youth" is never specified.
  3. [Manuscript metadata] No data availability statement, code release link, or ethical review statement is present in the provided material, which is particularly relevant for a dataset involving minors or young people.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity can be established from the supplied material, because the target paper's full text is absent and the provided 'full text' is a different manuscript.

full rationale

The target manuscript (arXiv:2508.02094) is represented in the supplied material only by its abstract; the 'full text' provided is the unrelated VLM4D benchmark paper (arXiv:2508.02095v2). The abstract reports constructing a Chinese youth-toxicity dataset, analyzing associations between youth perception and contextual factors, and improving toxicity detectors by adding meta-information, but it contains no equations, no derivation, and no description of how features were selected or whether evaluation was held out. Under the hard rule that circularity must be exhibited by quoting a specific reduction or a fitted input renamed as prediction, no such step can be identified from the available text. Concerns that the labels may encode adult assumptions or that accuracy gains may not generalize are correctness and evidence-quality risks, not demonstrated circularity. Since no load-bearing step reduces by the paper's own text to its inputs, the honest finding is no significant circularity.

Assumptions & free parameters 1 free parameters · 3 assumptions · 0 invented entities

Abstract-only ledger: no fitted values, code, or dataset statistics are reported in the abstract. The entries above record the assumptions the full manuscript must discharge: the validity of the youth-toxicity construct and its labels, corpus representativeness, and fair baseline comparison. The main auditing targets in the full manuscript are the annotation protocol (who annotated, agreement rates) and whether the meta-information features were selected on the same data used to measure the improvement. No invented entities are introduced.

free parameters (1)
  • Choice of contextual meta-information features = unspecified
    The abstract says youth perception relates to 'the source of an utterance' and 'text-related features', and that adding these improves accuracy. It does not state whether this feature set was fixed in advance or selected after inspecting the collected corpus; a post hoc feature choice tested on the same corpus would inflate the reported gain. Full experimental details were not available.
assumptions (3)
  • domain assumption Toxicity perception is age-dependent, so a stable 'youth-toxicity' construct exists that is distinct from adult-perceived toxicity.
    Conceptual foundation of the dataset and research questions. The abstract asserts the divergence between adult and youth perception rather than demonstrating it; the entire labeling exercise presumes the construct is annotable.
  • domain assumption The collected Chinese social media samples are representative of the language Chinese youth encounter online.
    The abstract reports constructing the Chinese youth-toxicity dataset but gives no sampling details such as platforms, time range, or filters. The feature analysis generalizes only if the corpus is representative.
  • domain assumption Baseline toxicity detectors were applied in their intended settings, making the comparison fair.
    The accuracy-improvement claim compares an augmented model with pre-existing detectors. Fair comparison requires matched prompts, thresholds, and hyperparameters, none of which are described in the abstract.

how reviews work

0 comments
Cite this review

Pith. "Pith review of "Harmless to You, Hurtful to Me!": Investigating the Detection of Toxic Languages Grounded in the Perspective of Youth." pith.science (2026). https://pith.science/paper/5F3D7Z52

@misc{pith2026250802094,
  author       = {Pith},
  title        = {Pith review of: "Harmless to You, Hurtful to Me!": Investigating the Detection of Toxic Languages Grounded in the Perspective of Youth},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/5F3D7Z52}},
  note         = {Machine review of arXiv:2508.02094}
}
read the original abstract

Risk perception is subjective, and youth's understanding of toxic content differs from that of adults. Although previous research has conducted extensive studies on toxicity detection in social media, the investigation of youth's unique toxicity, i.e., languages perceived as nontoxic by adults but toxic as youth, is ignored. To address this gap, we aim to explore: 1) What are the features of ``youth-toxicity'' languages in social media (RQ1); 2) Can existing toxicity detection techniques accurately detect these languages (RQ2). For these questions, we took Chinese youth as the research target, constructed the first Chinese ``youth-toxicity'' dataset, and then conducted extensive analysis. Our results suggest that youth's perception of these is associated with several contextual factors, like the source of an utterance and text-related features. Incorporating these meta information into current toxicity detection methods significantly improves accuracy overall. Finally, we propose several insights into future research on youth-centered toxicity detection.

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

104 extracted references · 30 canonical work pages

  1. [1]

    Phi-3 technical report: A highly capable language model locally on your phone

    Marah Abdin, Jyoti Aneja, Hany Awadalla, Ahmed Awadallah, Ammar Ahmad Awan, Nguyen Bach, Amit Bahree, Arash Bakhtiari, Jianmin Bao, Harkirat Behl, et al. Phi-3 technical report: A highly capable language model locally on your phone. arXiv preprint arXiv:2404.14219 ,

  2. [2]

    Phi-4 technical report

    Marah Abdin, Jyoti Aneja, Harkirat Behl, S ´ebastien Bubeck, Ronen Eldan, Suriya Gunasekar, Michael Har- rison, Russell J Hewett, Mojan Javaheripi, Piero Kauff- mann, et al. Phi-4 technical report. arXiv preprint arXiv:2412.08905, 2024. 5

  3. [3]

    Phi-4-mini technical re- port: Compact yet powerful multimodal language models via mixture-of-loras, 2025

    Abdelrahman Abouelenin, Atabak Ashfaq, Adam Atkin- son, Hany Awadalla, Nguyen Bach, Jianmin Bao, Alon Benhaim, Martin Cai, Vishrav Chaudhary, Congcong Chen, Dong Chen, Dongdong Chen, Junkun Chen, Weizhu Chen, Yen-Chun Chen, Yi ling Chen, Qi Dai, Xiyang Dai, Ruchao Fan, Mei Gao, Min Gao, Amit Garg, Abhishek Goswami, Junheng Hao, Amr Hendy, Yuxuan Hu, Xin J...

  4. [4]

    Cosmos world foun- dation model platform for physical ai

    Niket Agarwal, Arslan Ali, Maciej Bala, Yogesh Balaji, Erik Barker, Tiffany Cai, Prithvijit Chattopadhyay, Yongxin Chen, Yin Cui, Yifan Ding, et al. Cosmos world foun- dation model platform for physical ai. arXiv preprint arXiv:2501.03575, 2025. 2, 3, 4

  5. [5]

    Pixtral 12b

    Pravesh Agrawal, Szymon Antoniak, Emma Bou Hanna, Baptiste Bout, Devendra Chaplot, Jessica Chudnovsky, Diogo Costa, Baudouin De Monicault, Saurabh Garg, Theophile Gervet, et al. Pixtral 12b. arXiv preprint arXiv:2410.07073, 2024. 5

  6. [6]

    System card: Claude opus 4 & claude sonnet 4

    Anthropic. System card: Claude opus 4 & claude sonnet 4. Technical report, 2025. 5

  7. [7]

    Qwen technical report

    Jinze Bai, Shuai Bai, Yunfei Chu, Zeyu Cui, Kai Dang, Xiaodong Deng, Yang Fan, Wenbin Ge, Yu Han, Fei Huang, et al. Qwen technical report. arXiv preprint arXiv:2309.16609, 2023. 3

  8. [8]

    Video generation models as world simulators

    Tim Brooks, Bill Peebles, Connor Holmes, Will DePue, Yufei Guo, Li Jing, David Schnurr, Joe Taylor, Troy Luh- man, Eric Luhman, Clarence Ng, Ricky Wang, and Aditya Ramesh. Video generation models as world simulators

Show all 104 references
  1. [9]

    Language models are few-shot learners

    Tom Brown, Benjamin Mann, Nick Ryder, Melanie Sub- biah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakan- tan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. Language models are few-shot learners. Advances in neu- ral information processing systems , 33:1877–1901, 2020. 3

  2. [10]

    Spatial memory: how egocentric and allocen- tric combine

    Neil Burgess. Spatial memory: how egocentric and allocen- tric combine. Trends in Cognitive Sciences , 10(12):551– 557, 2006. 2

  3. [11]

    On learning mechanical laws of motion from video using neu- ral networks

    Pradyumna Chari, Yunhao Ba, Shijie Zhou, Chinmay Tale- gaonkar, Shreeram Athreya, and Achuta Kadambi. On learning mechanical laws of motion from video using neu- ral networks. IEEE Access, 11:30129–30145, 2023. 3

  4. [12]

    Spatialvlm: Endow- ing vision-language models with spatial reasoning capabili- ties

    Boyuan Chen, Zhuo Xu, Sean Kirmani, Brain Ichter, Dorsa Sadigh, Leonidas Guibas, and Fei Xia. Spatialvlm: Endow- ing vision-language models with spatial reasoning capabili- ties. In Proceedings of the IEEE/CVF Conference on Com- puter Vision and Pattern Recognition, pages 14455–14465,

  5. [13]

    Visualgpt: Data-efficient adaptation of pretrained language models for image captioning

    Jun Chen, Han Guo, Kai Yi, Boyang Li, and Mohamed El- hoseiny. Visualgpt: Data-efficient adaptation of pretrained language models for image captioning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 18030–18040, 2022. 2

  6. [14]

    Are we on the right way for eval- uating large vision-language models? arXiv preprint arXiv:2403.20330, 2024

    Lin Chen, Jinsong Li, Xiaoyi Dong, Pan Zhang, Yuhang Zang, Zehui Chen, Haodong Duan, Jiaqi Wang, Yu Qiao, Dahua Lin, et al. Are we on the right way for eval- uating large vision-language models? arXiv preprint arXiv:2403.20330, 2024. 7

  7. [15]

    Sharegpt4video: Improving video understanding and generation with better captions

    Lin Chen, Xilin Wei, Jinsong Li, Xiaoyi Dong, Pan Zhang, Yuhang Zang, Zehui Chen, Haodong Duan, Bin Lin, Zhenyu Tang, et al. Sharegpt4video: Improving video understanding and generation with better captions. arXiv preprint arXiv:2406.04325, 2024. 4, 7

  8. [16]

    Expanding performance boundaries of open-source multimodal models with model, data, and test-time scaling

    Zhe Chen, Weiyun Wang, Yue Cao, Yangzhou Liu, Zhang- wei Gao, Erfei Cui, Jinguo Zhu, Shenglong Ye, Hao Tian, Zhaoyang Liu, et al. Expanding performance boundaries of open-source multimodal models with model, data, and test-time scaling. arXiv preprint arXiv:2412.05271, 2024. 5

  9. [17]

    Spatialrgpt: Grounded spatial reasoning in vision-language models

    An-Chieh Cheng, Hongxu Yin, Yang Fu, Qiushan Guo, Ruihan Yang, Jan Kautz, Xiaolong Wang, and Sifei Liu. Spatialrgpt: Grounded spatial reasoning in vision-language models. Advances in Neural Information Processing Sys- tems, 37:135062–135093, 2024. 3

  10. [18]

    Videollama 2: Advancing spatial- temporal modeling and audio understanding in video-llms

    Zesen Cheng, Sicong Leng, Hang Zhang, Yifei Xin, Xin Li, Guanzheng Chen, Yongxin Zhu, Wenqi Zhang, Ziyang Luo, Deli Zhao, et al. Videollama 2: Advancing spatial- temporal modeling and audio understanding in video-llms. arXiv preprint arXiv:2406.07476, 2024. 3

  11. [19]

    Sharegpt-4o: Comprehensive multimodal annotations with gpt-4o, 2024

    Erfei Cui, Yinan He, Zheng Ma, Zhe Chen, Hao Tian, Weiyun Wang, Kunchang Li, Yi Wang, Wenhai Wang, Xizhou Zhu, Lewei Lu, Tong Lu, Yali Wang, Limin Wang, Yu Qiao, and Jifeng Dai. Sharegpt-4o: Comprehensive multimodal annotations with gpt-4o, 2024. 7

  12. [20]

    Instructblip: Towards general- purpose vision-language models with instruction tuning,

    Wenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong, Junqi Zhao, Weisheng Wang, Boyang Li, Pascale Fung, and Steven Hoi. Instructblip: Towards general- purpose vision-language models with instruction tuning,

  13. [21]

    Myers, and Anna C

    Julian De Freitas, Nicholas E. Myers, and Anna C. Nobre. Tracking the changing feature of a moving object. Journal of Vision, 16(3):22, 2016. 2

  14. [22]

    Bert: Pre-training of deep bidirectional trans- formers for language understanding

    Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. Bert: Pre-training of deep bidirectional trans- formers for language understanding. In Proceedings of the 2019 conference of the North American chapter of the as- sociation for computational linguistics: human l...

  15. [23]

    An im- age is worth 16x16 words: Transformers for image recog- nition at scale

    Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. An im- age is worth 16x16 words: Transformers for image recog- nitio...

  16. [24]

    Palm-e: An embodied multimodal language model

    Danny Driess, Fei Xia, Mehdi SM Sajjadi, Corey Lynch, Aakanksha Chowdhery, Ayzaan Wahid, Jonathan Tompson, Quan Vuong, Tianhe Yu, Wenlong Huang, et al. Palm-e: An embodied multimodal language model. 2023. 3

  17. [25]

    Missing premise exacerbates overthinking: Are reason- ing models losing critical thinking skill? arXiv preprint arXiv:2504.06514, 2025

    Chenrui Fan, Ming Li, Lichao Sun, and Tianyi Zhou. Missing premise exacerbates overthinking: Are reason- ing models losing critical thinking skill? arXiv preprint arXiv:2504.06514, 2025. 5

  18. [26]

    Large spatial model: End-to-end unposed images to semantic 3d

    Zhiwen Fan, Jian Zhang, Wenyan Cong, Peihao Wang, Renjie Li, Kairun Wen, Shijie Zhou, Achuta Kadambi, Zhangyang Wang, Danfei Xu, et al. Large spatial model: End-to-end unposed images to semantic 3d. Advances in neural information processing systems , 37:40212–40229,

  19. [27]

    Vlm-3r: Vision-language models aug- mented with instruction-aligned 3d reconstruction

    Zhiwen Fan, Jian Zhang, Renjie Li, Junge Zhang, Runjin Chen, Hezhen Hu, Kevin Wang, Huaizhi Qu, Dilin Wang, Zhicheng Yan, et al. Vlm-3r: Vision-language models aug- mented with instruction-aligned 3d reconstruction. arXiv preprint arXiv:2505.20279, 2025. 3

  20. [28]

    Freyd and Ronald A

    Jennifer J. Freyd and Ronald A. Finke. Representational momentum. Journal of Experimental Psychology: Learn- ing, Memory, and Cognition, 10(1):126–132, 1984. 2

  21. [29]

    Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis

    Chaoyou Fu, Yuhan Dai, Yongdong Luo, Lei Li, Shuhuai Ren, Renrui Zhang, Zihan Wang, Chenyu Zhou, Yunhang Shen, Mengdan Zhang, et al. Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis. arXiv preprint arXiv:2405.21075, 2024. 3

  22. [30]

    Gemini 2.5: Pushing the frontier with advanced reasoning, multimodality, long context, and next generation agentic capabilities

    Google Gemini Team. Gemini 2.5: Pushing the frontier with advanced reasoning, multimodality, long context, and next generation agentic capabilities. Technical report, 2025. 5

  23. [31]

    Multimodal-gpt: A vision and lan- guage model for dialogue with humans

    Tao Gong, Chengqi Lyu, Shilong Zhang, Yudong Wang, Miao Zheng, Qian Zhao, Kuikun Liu, Wenwei Zhang, Ping Luo, and Kai Chen. Multimodal-gpt: A vision and lan- guage model for dialogue with humans. arXiv preprint arXiv:2305.04790, 2023. 3

  24. [32]

    Ego4d: Around the world in 3,000 hours of egocentric video

    Kristen Grauman, Andrew Westbury, Eugene Byrne, Zachary Chavis, Antonino Furnari, Rohit Girdhar, Jackson Hamburger, Hao Jiang, Miao Liu, Xingyu Liu, et al. Ego4d: Around the world in 3,000 hours of egocentric video. In Proceedings of the IEEE/CVF conference on computer vi- sio...

  25. [33]

    Mmworld: Towards multi- discipline multi-faceted world model evaluation in videos

    Xuehai He, Weixi Feng, Kaizhi Zheng, Yujie Lu, Wan- rong Zhu, Jiachen Li, Yue Fan, Jianfeng Wang, Linjie Li, Zhengyuan Yang, et al. Mmworld: Towards multi- discipline multi-faceted world model evaluation in videos. arXiv preprint arXiv:2406.08407, 2024. 3

  26. [34]

    Mojito: Motion tra- jectory and intensity control for video generation

    Xuehai He, Shuohang Wang, Jianwei Yang, Xiaoxia Wu, Yiping Wang, Kuan Wang, Zheng Zhan, Olatunji Ruwase, Yelong Shen, and Xin Eric Wang. Mojito: Motion tra- jectory and intensity control for video generation. arXiv preprint arXiv: 2412.08948, 2024. 3, 4

  27. [35]

    Gpt-4o system card

    Aaron Hurst, Adam Lerer, Adam P Goucher, Adam Perel- man, Aditya Ramesh, Aidan Clark, AJ Ostrow, Akila Weli- hinda, Alan Hayes, Alec Radford, et al. Gpt-4o system card. arXiv preprint arXiv:2410.21276, 2024. 3

  28. [36]

    Visual perception of biological motion and a model for its analysis

    Gunnar Johansson. Visual perception of biological motion and a model for its analysis. Perception & Psychophysics, 14(2):201–211, 1973. 2

  29. [37]

    How good is my video lmm? complex video reasoning and robust- ness evaluation suite for video-lmms

    Muhammad Uzair Khattak, Muhammad Ferjad Naeem, Jameel Hassan, Muzammal Naseer, Federico Tombari, Fa- had Shahbaz Khan, and Salman Khan. How good is my video lmm? complex video reasoning and robust- ness evaluation suite for video-lmms. arXiv preprint arXiv:2405.03690, 2024. 3

  30. [38]

    Open- vla: An open-source vision-language-action model

    Moo Jin Kim, Karl Pertsch, Siddharth Karamcheti, Ted Xiao, Ashwin Balakrishna, Suraj Nair, Rafael Rafailov, Ethan Foster, Grace Lam, Pannag Sanketi, et al. Open- vla: An open-source vision-language-action model. arXiv preprint arXiv:2406.09246, 2024. 3

  31. [39]

    Decomposing nerf for editing via feature field distil- lation

    Sosuke Kobayashi, Eiichi Matsumoto, and Vincent Sitz- mann. Decomposing nerf for editing via feature field distil- lation. Advances in neural information processing systems, 35:23311–23330, 2022. 8

  32. [40]

    On space-time interest points

    Ivan Laptev. On space-time interest points. International Journal of Computer Vision, 64(2-3):107–123, 2005. 3

  33. [41]

    Denker, Donnie Henderson, Richard E

    Yann LeCun, Bernhard Boser, John S. Denker, Donnie Henderson, Richard E. Howard, Wayne Hubbard, and Lawrence D. Jackel. Backpropagation applied to handwrit- ten zip code recognition. Neural Computation, 1(4):541– 551, 1989. 2

  34. [42]

    Alan M. Leslie. Spatiotemporal continuity and the per- ception of causality in infants. Perception, 13(3):287–305,

  35. [43]

    Seed-bench: Bench- marking multimodal large language models

    Bohao Li, Yuying Ge, Yixiao Ge, Guangzhi Wang, Rui Wang, Ruimao Zhang, and Ying Shan. Seed-bench: Bench- marking multimodal large language models. In Proceed- ings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 13299–13308, 2024. 3

  36. [44]

    Llava-onevision: Easy visual task transfer

    Bo Li, Yuanhan Zhang, Dong Guo, Renrui Zhang, Feng Li, Hao Zhang, Kaichen Zhang, Peiyuan Zhang, Yanwei Li, Ziwei Liu, et al. Llava-onevision: Easy visual task transfer. arXiv preprint arXiv:2408.03326, 2024. 3, 5

  37. [45]

    Aria: An open multimodal native mixture- of-experts model

    Dongxu Li, Yudong Liu, Haoning Wu, Yue Wang, Zhiqi Shen, Bowen Qu, Xinyao Niu, Guoyin Wang, Bei Chen, 10 and Junnan Li. Aria: An open multimodal native mixture- of-experts model. arXiv preprint arXiv:2410.05993, 2024. 5

  38. [46]

    Videochat: Chat-centric video understanding

    KunChang Li, Yinan He, Yi Wang, Yizhuo Li, Wen- hai Wang, Ping Luo, Yali Wang, Limin Wang, and Yu Qiao. Videochat: Chat-centric video understanding. arXiv preprint arXiv:2305.06355, 2023. 3, 7

  39. [47]

    Mvbench: A comprehensive multi- modal video understanding benchmark, 2023

    Kunchang Li, Yali Wang, Yinan He, Yizhuo Li, Yi Wang, Yi Liu, Zun Wang, Jilan Xu, Guo Chen, Ping Luo, Limin Wang, and Yu Qiao. Mvbench: A comprehensive multi- modal video understanding benchmark, 2023. 7

  40. [48]

    Mvbench: A comprehensive multi-modal video under- standing benchmark

    Kunchang Li, Yali Wang, Yinan He, Yizhuo Li, Yi Wang, Yi Liu, Zun Wang, Jilan Xu, Guo Chen, Ping Luo, et al. Mvbench: A comprehensive multi-modal video under- standing benchmark. In Proceedings of the IEEE/CVF Con- ference on Computer Vision and Pattern Recognition, pages 2219...

  41. [49]

    4k4DGen: Panoramic 4d generation at 4k resolution

    Renjie Li, Panwang Pan, Bangbang Yang, Dejia Xu, Shi- jie Zhou, Xuanyang Zhang, Zeming Li, Achuta Kadambi, Zhangyang Wang, Zhengzhong Tu, and Zhiwen Fan. 4k4DGen: Panoramic 4d generation at 4k resolution. In The Thirteenth International Conference on Learning Rep- resentations...

  42. [50]

    Videoeval: Comprehensive benchmark suite for low-cost evaluation of video foundation model

    Xinhao Li, Zhenpeng Huang, Jing Wang, Kunchang Li, and Limin Wang. Videoeval: Comprehensive benchmark suite for low-cost evaluation of video foundation model. arXiv preprint arXiv:2407.06491, 2024. 3

  43. [51]

    Scenethesis: A language and vision agentic framework for 3d scene generation

    Lu Ling, Chen-Hsuan Lin, Tsung-Yi Lin, Yifan Ding, Yu Zeng, Yichen Sheng, Yunhao Ge, Ming-Yu Liu, Aniket Bera, and Zhaoshuo Li. Scenethesis: A language and vision agentic framework for 3d scene generation. arXiv preprint arXiv:2505.02836, 2025. 3

  44. [52]

    Visual instruction tuning

    Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. Visual instruction tuning. Advances in neural infor- mation processing systems, 36:34892–34916, 2023. 3

  45. [53]

    World model on million-length video and language with ringattention

    Hao Liu, Wilson Yan, Matei Zaharia, and Pieter Abbeel. World model on million-length video and language with ringattention. arXiv preprint, 2024. 2

  46. [54]

    World model on million-length video and language with blockwise ringattention

    Hao Liu, Wilson Yan, Matei Zaharia, and Pieter Abbeel. World model on million-length video and language with blockwise ringattention. arXiv preprint arXiv:2402.08268,

  47. [55]

    Mmbench: Is your multi- modal model an all-around player? InEuropean conference on computer vision, pages 216–233

    Yuan Liu, Haodong Duan, Yuanhan Zhang, Bo Li, Songyang Zhang, Wangbo Zhao, Yike Yuan, Jiaqi Wang, Conghui He, Ziwei Liu, et al. Mmbench: Is your multi- modal model an all-around player? InEuropean conference on computer vision, pages 216–233. Springer, 2024. 3

  48. [56]

    Deepseek-vl: towards real-world vision- language understanding

    Haoyu Lu, Wen Liu, Bo Zhang, Bingxuan Wang, Kai Dong, Bo Liu, Jingxiang Sun, Tongzheng Ren, Zhuoshu Li, Hao Yang, et al. Deepseek-vl: towards real-world vision- language understanding. arXiv preprint arXiv:2403.05525,

  49. [57]

    Vil- bert: Pretraining task-agnostic visiolinguistic representa- tions for vision-and-language tasks

    Jiasen Lu, Dhruv Batra, Devi Parikh, and Stefan Lee. Vil- bert: Pretraining task-agnostic visiolinguistic representa- tions for vision-and-language tasks. In Neural Information Processing Systems, 2019. 2

  50. [58]

    Video-chatgpt: Towards detailed video understanding via large vision and language models

    Muhammad Maaz, Hanoona Rasheed, Salman Khan, and Fahad Shahbaz Khan. Video-chatgpt: Towards detailed video understanding via large vision and language models. arXiv preprint arXiv:2306.05424, 2023. 3

  51. [59]

    Marr and S

    D. Marr and S. Ullman. Directional selectivity and its use in early visual processing. Proceedings of the Royal Society of London. Series B, Biological Sciences , 211(1183):151– 180, 1981. 2

  52. [60]

    The llama 4 herd: The beginning of a new era of natively multimodal ai innovation

    Meta. The llama 4 herd: The beginning of a new era of natively multimodal ai innovation. Technical report, 2025. 5

  53. [61]

    Video-bench: A comprehensive benchmark and toolkit for evaluat- ing video-based large language models

    Munan Ning, Bin Zhu, Yujia Xie, Bin Lin, Jiaxi Cui, Lu Yuan, Dongdong Chen, and Li Yuan. Video-bench: A comprehensive benchmark and toolkit for evaluat- ing video-based large language models. arXiv preprint arXiv:2311.16103, 2023. 3

  54. [62]

    Sceneteller: Language-to-3d scene gen- eration

    Bas ¸ak Melis ¨Ocal, Maxim Tatarchenko, Sezer Karao ˘glu, and Theo Gevers. Sceneteller: Language-to-3d scene gen- eration. In European Conference on Computer Vision , pages 362–378. Springer, 2024. 3

  55. [63]

    Hello gpt-4o

    OpenAI. Hello gpt-4o. Technical report, 2024. 5

  56. [64]

    A real-to-sim-to-real approach to robotic manipulation with vlm-generated iterative keypoint rewards

    Shivansh Patel, Xinchen Yin, Wenlong Huang, Shubham Garg, Hooshang Nayyeri, Li Fei-Fei, Svetlana Lazeb- nik, and Yunzhu Li. A real-to-sim-to-real approach to robotic manipulation with vlm-generated iterative keypoint rewards. arXiv preprint arXiv:2502.08643, 2025. 3

  57. [65]

    A benchmark dataset and evaluation methodology for video object segmentation

    Federico Perazzi, Jordi Pont-Tuset, Brian McWilliams, Luc Van Gool, Markus Gross, and Alexander Sorkine-Hornung. A benchmark dataset and evaluation methodology for video object segmentation. In Proceedings of the IEEE confer- ence on computer vision and pattern recognition , p...

  58. [66]

    The 2017 davis challenge on video object segmentation

    Jordi Pont-Tuset, Federico Perazzi, Sergi Caelles, Pablo Ar- bel´aez, Alex Sorkine-Hornung, and Luc Van Gool. The 2017 davis challenge on video object segmentation. arXiv preprint arXiv:1704.00675, 2017. 2, 4

  59. [67]

    Improving language understanding by gen- erative pre-training

    Alec Radford, Karthik Narasimhan, Tim Salimans, Ilya Sutskever, et al. Improving language understanding by gen- erative pre-training. 2018. 3

  60. [68]

    Learning transferable visual models from natural language supervision

    Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. Learning transferable visual models from natural language supervision. In Proceedings of th...

  61. [69]

    Learning to localize objects improves spatial reasoning in visual-llms

    Kanchana Ranasinghe, Satya Narayan Shukla, Omid Pour- saeed, Michael S Ryoo, and Tsung-Yu Lin. Learning to localize objects improves spatial reasoning in visual-llms. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 12977–12987, 2024. 3

  62. [70]

    Two-stream con- volutional networks for action recognition in videos

    Karen Simonyan and Andrew Zisserman. Two-stream con- volutional networks for action recognition in videos. InAd- vances in Neural Information Processing Systems (NIPS) , pages 568–576, 2014. 3

  63. [71]

    Spelke and Katherine D

    Elizabeth S. Spelke and Katherine D. Kinzler. Core knowl- edge. Developmental Science, 10(1):89–96, 2007. 2

  64. [72]

    Alanavlm: A multimodal embodied ai foundation model for egocentric video understanding

    Alessandro Suglia, Claudio Greco, Katie Baker, Jose L Part, Ioannis Papaioannou, Arash Eshghi, Ioannis Konstas, 11 and Oliver Lemon. Alanavlm: A multimodal embodied ai foundation model for egocentric video understanding. arXiv preprint arXiv:2406.13807, 2024. 3

  65. [73]

    Gemini 1.5: Unlocking multimodal understanding across millions of tokens of con- text

    Gemini Team, Petko Georgiev, Ving Ian Lei, Ryan Burnell, Libin Bai, Anmol Gulati, Garrett Tanzer, Damien Vincent, Zhufeng Pan, Shibo Wang, et al. Gemini 1.5: Unlocking multimodal understanding across millions of tokens of con- text. arXiv preprint arXiv:2403.05530, 2024. 3

  66. [74]

    Llama: Open and efficient foundation language mod- els

    Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timoth ´ee Lacroix, Bap- tiste Rozi `ere, Naman Goyal, Eric Hambro, Faisal Azhar, et al. Llama: Open and efficient foundation language mod- els. arXiv preprint arXiv:2302.13971, 2023. 3

  67. [75]

    Wan: Open and advanced large-scale video generative models

    Team Wan, Ang Wang, Baole Ai, Bin Wen, Chaojie Mao, Chen-Wei Xie, Di Chen, Feiwu Yu, Haiming Zhao, Jianx- iao Yang, et al. Wan: Open and advanced large-scale video generative models. arXiv preprint arXiv:2503.20314, 2025. 4

  68. [76]

    Vlm see, robot do: Human demo video to robot action plan via vision language model

    Beichen Wang, Juexiao Zhang, Shuwen Dong, Irving Fang, and Chen Feng. Vlm see, robot do: Human demo video to robot action plan via vision language model. arXiv preprint arXiv:2410.08792, 2024. 3

  69. [77]

    Action recognition by dense trajectories

    Heng Wang, Alexander Kl ¨aser, Cordelia Schmid, and Cheng-Lin Liu. Action recognition by dense trajectories. In 2011 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 3169–3176, 2011. 3

  70. [78]

    Grounded-videollm: Sharpening fine-grained tem- poral grounding in video large language models

    Haibo Wang, Zhiyang Xu, Yu Cheng, Shizhe Diao, Yu- fan Zhou, Yixin Cao, Qifan Wang, Weifeng Ge, and Lifu Huang. Grounded-videollm: Sharpening fine-grained tem- poral grounding in video large language models. arXiv preprint arXiv:2410.03290, 2024. 3

  71. [79]

    Qwen2-vl: Enhancing vision-language model’s perception of the world at any resolution

    Peng Wang, Shuai Bai, Sinan Tan, Shijie Wang, Zhihao Fan, Jinze Bai, Keqin Chen, Xuejing Liu, Jialin Wang, Wenbin Ge, et al. Qwen2-vl: Enhancing vision-language model’s perception of the world at any resolution. arXiv preprint arXiv:2409.12191, 2024. 3, 5

  72. [80]

    Compositional 4d dy- namic scenes understanding with physics priors for video question answering

    Xingrui Wang, Wufei Ma, Angtian Wang, Shuo Chen, Adam Kortylewski, and Alan Yuille. Compositional 4d dy- namic scenes understanding with physics priors for video question answering. arXiv preprint arXiv:2406.00622 ,

  73. [81]

    Internvideo2: Scaling foundation models for multimodal video understanding

    Yi Wang, Kunchang Li, Xinhao Li, Jiashuo Yu, Yinan He, Guo Chen, Baoqi Pei, Rongkun Zheng, Zun Wang, Yan- song Shi, et al. Internvideo2: Scaling foundation models for multimodal video understanding. In European Confer- ence on Computer Vision, pages 396–416. Springer, 2024. 5, 8

  74. [82]

    Internvideo2

    Yi Wang, Xinhao Li, Ziang Yan, Yinan He, Jiashuo Yu, Xi- angyu Zeng, Chenting Wang, Changlian Ma, Haian Huang, Jianfei Gao, et al. Internvideo2. 5: Empowering video mllms with long and rich context modeling. arXiv preprint arXiv:2501.12386, 2025. 5

  75. [83]

    Finetuned language models are zero-shot learn- ers

    Jason Wei, Maarten Bosma, Vincent Y Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, Andrew M Dai, and Quoc V Le. Finetuned language models are zero-shot learn- ers. arXiv preprint arXiv:2109.01652, 2021. 3

  76. [84]

    Chain-of-thought prompting elicits reasoning in large lan- guage models

    Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al. Chain-of-thought prompting elicits reasoning in large lan- guage models. Advances in neural information processing systems, 35:24824–24837, 2022. 6, 7

  77. [85]

    Cat4d: Create anything in 4d with multi-view video dif- fusion models

    Rundi Wu, Ruiqi Gao, Ben Poole, Alex Trevithick, Changxi Zheng, Jonathan T Barron, and Aleksander Holynski. Cat4d: Create anything in 4d with multi-view video dif- fusion models. In Proceedings of the Computer Vision and Pattern Recognition Conference, pages 26057–26068,

  78. [86]

    Videollm-mod: Efficient video-language streaming with mixture-of-depths vision computation

    Shiwei Wu, Joya Chen, Kevin Qinghong Lin, Qimeng Wang, Yan Gao, Qianli Xu, Tong Xu, Yao Hu, Enhong Chen, and Mike Zheng Shou. Videollm-mod: Efficient video-language streaming with mixture-of-depths vision computation. Advances in Neural Information Processing Systems, 37:10992...

  79. [87]

    Grok-2 beta release

    xAI. Grok-2 beta release. Technical report, 2024. 5

  80. [88]

    Youtube-vos: Sequence-to-sequence video object segmentation

    Ning Xu, Linjie Yang, Yuchen Fan, Jianchao Yang, Dingcheng Yue, Yuchen Liang, Brian Price, Scott Cohen, and Thomas Huang. Youtube-vos: Sequence-to-sequence video object segmentation. In Proceedings of the Euro- pean conference on computer vision (ECCV) , pages 585– 601, 2018. 2, 4

  81. [89]

    An Yang, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chengyuan Li, Dayiheng Liu, Fei Huang, Haoran Wei, et al. Qwen2. 5 technical report.arXiv preprint arXiv:2412.15115, 2024. 5

  82. [90]

    Thinking in space: How mul- timodal large language models see, remember, and recall spaces

    Jihan Yang, Shusheng Yang, Anjali W Gupta, Rilyn Han, Li Fei-Fei, and Saining Xie. Thinking in space: How mul- timodal large language models see, remember, and recall spaces. arXiv preprint arXiv:2412.14171, 2024. 3, 6

  83. [91]

    Videorefer suite: Advanc- ing spatial-temporal object understanding with video llm

    Yuqian Yuan, Hang Zhang, Wentong Li, Zesen Cheng, Boqiang Zhang, Long Li, Xin Li, Deli Zhao, Wenqiao Zhang, Yueting Zhuang, et al. Videorefer suite: Advanc- ing spatial-temporal object understanding with video llm. arXiv preprint arXiv:2501.00599, 2024. 3

  84. [92]

    Mmmu: A massive multi- discipline multimodal understanding and reasoning bench- mark for expert agi

    Xiang Yue, Yuansheng Ni, Kai Zhang, Tianyu Zheng, Ruoqi Liu, Ge Zhang, Samuel Stevens, Dongfu Jiang, Weiming Ren, Yuxuan Sun, et al. Mmmu: A massive multi- discipline multimodal understanding and reasoning bench- mark for expert agi. In Proceedings of the IEEE/CVF Con- ference...

  85. [93]

    Improving 2d feature representations by 3d-aware fine-tuning

    Yuanwen Yue, Anurag Das, Francis Engelmann, Siyu Tang, and Jan Eric Lenssen. Improving 2d feature representations by 3d-aware fine-tuning. In European Conference on Com- puter Vision, pages 57–74. Springer, 2024. 8

  86. [94]

    Videollama 3: Frontier multimodal foundation models for image and video under- standing

    Boqiang Zhang, Kehan Li, Zesen Cheng, Zhiqiang Hu, Yuqian Yuan, Guanzheng Chen, Sicong Leng, Yuming Jiang, Hang Zhang, Xin Li, et al. Videollama 3: Frontier multimodal foundation models for image and video under- standing. arXiv preprint arXiv:2501.13106, 2025. 3, 5

  87. [95]

    Video-llama: An instruction-tuned audio-visual language model for video understanding

    Hang Zhang, Xin Li, and Lidong Bing. Video-llama: An instruction-tuned audio-visual language model for video understanding. arXiv preprint arXiv:2306.02858, 2023. 3

  88. [96]

    Combo: compositional world models for embodied 12 multi-agent cooperation

    Hongxin Zhang, Zeyuan Wang, Qiushi Lyu, Zheyuan Zhang, Sunli Chen, Tianmin Shu, Yilun Du, and Chuang Gan. Combo: compositional world models for embodied 12 multi-agent cooperation. arXiv preprint arXiv:2404.10775,

  89. [97]

    Llava-next: A strong zero-shot video understanding model,

    Yuanhan Zhang, Bo Li, haotian Liu, Yong jae Lee, Liangke Gui, Di Fu, Jiashi Feng, Ziwei Liu, and Chunyuan Li. Llava-next: A strong zero-shot video understanding model,

  90. [98]

    Mmvu: Measuring expert- level multi-discipline video understanding

    Yilun Zhao, Lujing Xie, Haowei Zhang, Guo Gan, Yitao Long, Zhiyuan Hu, Tongyan Hu, Weiyuan Chen, Chuhan Li, Junyang Song, et al. Mmvu: Measuring expert- level multi-discipline video understanding. arXiv preprint arXiv:2501.12380, 2025. 3, 6

  91. [99]

    Lla- mafactory: Unified efficient fine-tuning of 100+ language models

    Yaowei Zheng, Richong Zhang, Junhao Zhang, Yanhan Ye, Zheyan Luo, Zhangchi Feng, and Yongqiang Ma. Lla- mafactory: Unified efficient fine-tuning of 100+ language models. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 3: Sys- ...

  92. [100]

    Feature 3dgs: Supercharging 3d gaussian splatting to enable distilled fea- ture fields

    Shijie Zhou, Haoran Chang, Sicheng Jiang, Zhiwen Fan, Zehao Zhu, Dejia Xu, Pradyumna Chari, Suya You, Zhangyang Wang, and Achuta Kadambi. Feature 3dgs: Supercharging 3d gaussian splatting to enable distilled fea- ture fields. In Proceedings of the IEEE/CVF Conference on Comput...

  93. [101]

    Dreamscene360: Uncon- strained text-to-3d scene generation with panoramic gaus- sian splatting

    Shijie Zhou, Zhiwen Fan, Dejia Xu, Haoran Chang, Pradyumna Chari, Tejas Bharadwaj, Suya You, Zhangyang Wang, and Achuta Kadambi. Dreamscene360: Uncon- strained text-to-3d scene generation with panoramic gaus- sian splatting. In European Conference on Computer Vi- sion, pages 3...

  94. [102]

    Feature4x: Bridging any monocular video to 4d agentic ai with versatile gaussian feature fields

    Shijie Zhou, Hui Ren, Yijia Weng, Shuwang Zhang, Zhen Wang, Dejia Xu, Zhiwen Fan, Suya You, Zhangyang Wang, Leonidas Guibas, et al. Feature4x: Bridging any monocular video to 4d agentic ai with versatile gaussian feature fields. In Proceedings of the Computer Vision and Patter...

  95. [103]

    Minigpt-4: Enhancing vision-language understanding with advanced large language models

    Deyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li, and Mo- hamed Elhoseiny. Minigpt-4: Enhancing vision-language understanding with advanced large language models. arXiv preprint arXiv:2304.10592, 2023. 3

  96. [104]

    left” and “right

    Orr Zohar, Xiaohan Wang, Yann Dubois, Nikhil Mehta, Tong Xiao, Philippe Hansen-Estruch, Licheng Yu, Xiao- fang Wang, Felix Juefei-Xu, Ning Zhang, et al. Apollo: An exploration of video understanding in large multimodal models. arXiv preprint arXiv:2412.10360, 2024. 7 13 VLM4D:...

Pith tools

Reviewed August 6, 2026 · model on record in the stance chip above.