REVIEW 3 major objections 5 minor 61 references
Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times
T0 review · 3 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read This paper argues that state-of-the-art video language models cannot reliably tell completed from ongoing actions in video when the distinction is carried by grammatical aspect, and supports the claim with a four-language benchmark on…
desk verdict A genuinely new quadrilingual benchmark for temporal aspect reasoning in video, with a plausible but under-stated empirical claim that current VLMs are far below human performance on it. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is the Perfect Times template system grounded in Allen's interval algebra, a formal scheme for all possible temporal relations between two intervals. The system assigns each video pair a main-clause action (mca) and a dependent-clause action (dca), groups their relation as precedence, succession, or simultaneity, and generates questions in four languages using aspect-appropriate verb forms, with every action class labeled telic or atelic. Each question carries three distractor types: Type 1 varies only the aspect or completeness of the correct answer, Type 2 uses another action from the same video with the opposite aspect, and Type 3 draws from an action absent from the video; this design lets the authors localize a model's error to failed action recognition, failed temporal ordering, or failed aspect-form matching.
What would settle it
A decisive control is to rerun the benchmark using per-video telicity labels instead of one label per verb class, or to restrict evaluation to the verb classes on which annotators fully agreed. If model accuracy stays near 26% to 47% on the unambiguous subset, the deficit is genuinely temporal; if accuracy rises substantially, the context-free labels were producing misleading questions and the lack-of-temporal-fusion claim would need to be limited to the label-noise-free subset.
Extended reading notes
Core claim
The paper's central claim, stated in the evaluation section, is that the accuracy of all tested models in all languages remains significantly below the human gold standard, and this is taken as evidence that state-of-the-art VLMs lack robust mechanisms for precise temporal fusion. The discovery is operationalized through a new semi-synthetic benchmark: 3,739 multiple-choice questions generated from 400 household activity videos, with each question built from one of twelve templates that systematically covers precedence, succession, and simultaneity, and each correct answer paired with three distractors that differ in aspectual form, in temporal order, or in scene relevance. Analysis of the error patterns shows that models rarely choose out-of-context actions, so the failure is not basic action recognition; their errors concentrate on distractors that change the verb form's aspect (Type 1) or swap the temporal order of two actions that both occur in the video (Type 2). The paper further reports that all models except GPT-4o answer more telic than atelic questions correctly, and that mistaken models prefer perfect over durative answers, which the authors read as a reliance on lexical and causal biases rather than on grammatical aspect.
Load-bearing premise
The benchmark's correctness depends on telicity labels assigned to 157 verb classes once, out of context, with only substantial annotator agreement (kappa 0.67); a wrong label for a particular video can generate an invalid question or misleading distractor, lowering model scores for reasons unrelated to temporal reasoning.
Editorial extensions
If this is right
- Current video language models cannot be trusted on video questions whose correct answer hinges on grammatical aspect, even in English.
- The template-based design means the benchmark can be extended to any language and any action-annotated video collection, enabling controlled cross-linguistic comparisons of temporal reasoning.
- Because out-of-context distractors are rarely selected, the models' deficit is specifically in temporal and aspectual integration, not in recognizing which actions occur in a scene.
- The consistent preference for telic or perfect answers indicates that models lean on lexical and causal priors, so benchmarks that do not balance aspectual form will overstate true temporal understanding.
- Cross-linguistic results suggest that languages with explicit lexical aspect markers (Russian) are somewhat easier for several models, while languages that leave more to visual disambiguation (Japanese) are harder, pointing to surface syntactic information as a partial crutch.
Reading between the lines
- The paper does not run this control, but if telicity-label noise is the real driver of low scores, then rerunning the benchmark on only the verb classes with unanimous annotator agreement should raise model accuracy more than it raises human accuracy; the paper's own kappa of 0.67 makes this a plausible confound.
- The distractor typology implies targeted training interventions: a model that systematically falls for Type 1 distractors needs better aspect-form matching, while a Type 2 pattern calls for better temporal ordering, so the benchmark could be used as a diagnostic during fine-tuning rather than only as an evaluation.
- The reported telic-choice bias may reflect an over-representation of completed, causally linked events in training data; a testable extension would be to rebalance the answer options or apply a debiasing prompt and observe whether the distractor error profile shifts.
- Because the templates are semi-synthetic, one could generate distractors at graded aspectual distance from the correct answer, producing a difficulty calibration curve for temporal reasoning that the current dataset does not provide.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces Perfect Times, a quadrilingual (English, Italian, Russian, Japanese) multiple-choice video question-answering benchmark for evaluating temporal reasoning in video-language models (VLMs). Using 400 Charades videos re-annotated with Action Genome action classes, the authors generate 3,739 QA pairs from 12 templates that systematically cover Allen's interval relations and vary the perfectivity/atelicity of main and dependent clause actions, with three distractor types. They evaluate six VLMs and a human gold standard (93.36% accuracy), reporting model accuracies between 25.92% and 46.99% across languages. The central claim is that all tested SoTA VLMs fall far below human performance and therefore lack robust mechanisms for precise temporal fusion, particularly for distinguishing completed versus ongoing actions via grammatical aspect. The dataset and code are released.
Significance. If the benchmark is valid, it fills a real gap: existing temporal video QA benchmarks are predominantly English, before/after-oriented, and do not systematically control for grammatical aspect. The cross-linguistic design is a genuine contribution, as is the grounding in Allen's interval algebra and linguistic theories of telicity. The paper ships the dataset and code, includes a human baseline, and reports distractor-type error patterns. The headline finding—VLMs massively underperform humans on aspect-sensitive temporal reasoning—is plausible and worth publishing, provided the instrument's validity and the statistical basis for the comparative claims are strengthened. The main risk is that the telicity labels (Fleiss κ=0.67) underpin every correct answer and distractor; if many items are mislabeled, the benchmark partially measures label noise rather than temporal reasoning.
major comments (3)
- [§4.1 and §4.4] The telicity labels for 157 verb classes are assigned 'irrespective of their context' with only substantial inter-annotator agreement (Fleiss' κ=0.67), and the authors themselves acknowledge inherent ambiguity. Since every correct answer and Distractor Type 1 is generated from these labels, a verb class whose telicity is context-dependent (e.g., 'read') can yield QA pairs whose gold answer is wrong or whose Type 1 distractor is not a true minimal aspect pair. The model-human gap and the aspect-blindness interpretation therefore depend on the validity of these labels. Please re-audit the evaluation on the subset of verb classes with unanimous (or near-unanimous) telicity agreement, report model and human accuracies on that subset, and discuss whether the gap persists. Also report error rates per verb class to identify label-driven failures.
- [§4.4 and Table 2] The claim that 'the accuracy of all models in all languages remains significantly below the human gold standard' uses 'significantly' without any statistical test, and Table 2 reports no confidence intervals or pairwise significance tests. While the raw gap (46.99% max vs 93.36% human) is large enough to be robust, several comparative claims—'Gemini-2.0-flash-lite consistently performed better', 'InternVL2 ... worst in Italian', 'models perform better on Russian and worse on Japanese'—rest on differences of 2–5 percentage points (e.g., Gemini English 43.41 vs Russian 46.99; GPT-4o Japanese 38.49 vs Russian 45.04). These differences may be within sampling error. Please provide confidence intervals (e.g., bootstrap) and appropriate tests (e.g., McNemar for per-item comparisons) for the headline gap and for all cross-language and cross-model rankings.
- [§4.3 and Table 2] The Distractor Rate by Type is defined in Eq. (1) as the proportion of errors attributable to each distractor type, which should sum to 100% across the three types in error cases. However, the entries in Table 2 sum to the overall error rate, not to 100 (e.g., Gemini English: 22.15+30.42+4.01=56.58 ≈ 100−43.41). The footnote (footnote 8) even acknowledges an alternative 'relative to all predictions' definition. This inconsistency makes the central error-pattern analysis ('most errors are Type 2', 'models ignore linguistic aspect') hard to interpret, because the numbers conflate accuracy with error composition. Please state clearly which definition is used in Table 2, label the columns accordingly, and ideally report both error-conditional rates (summing to 100) and overall rates.
minor comments (5)
- [§4.4] The text states that '5.67% (208) of questions were answered incorrectly by all models'; however, 208/3739 = 5.56%, so either the percentage or the count is incorrect. Please reconcile.
- [§4.1 and Table 10] The human evaluation is described as 'speakers of each language were asked to take this MCQ test', but Table 10 lists only one row per language. Please clarify how many annotators answered each language version, how the 93.36% gold standard was aggregated, and how Fleiss' κ=0.8 was computed if there was only one annotator per language.
- [Appendix G] The Russian prompt for LLaVA-NeXT-Video contains a stray quotation mark in 'Вопрос: "Пожалуйста, выберите правильный ответ', which could corrupt the prompt at inference time. Please fix the formatting.
- [§3.4] For Russian and Japanese, 'the opposite aspect word' is not always a grammatical minimal pair (e.g., lexical perfective/imperfective pairs in Russian, te-form vs i-form in Japanese). Please clarify how Distractor Type 1 is operationalized in each language and whether any templates yield distractors that are not exact aspect alternations.
- [§4.2 and Table 11] The evaluation protocol differs across models: some receive the full video, others receive frames sampled every 3 seconds. This could confound cross-model comparisons. The paper mentions this, but it should be stated more prominently as a limitation and considered in the interpretation of rankings.
Circularity Check
No circular derivation: the benchmark is externally anchored and the model–human gap is an empirical result; only a non-load-bearing self-citation raises the score to 2.
full rationale
The paper's central claim is an empirical evaluation: Perfect Times is an external benchmark built from Charades videos, Action Genome classes, and human telicity annotations, and VLM accuracies (Table 2) are compared with a human gold standard produced by independent annotators answering the same MCQ test. There is no fitted parameter that is later renamed as a prediction, and no uniqueness theorem is imported from the authors' prior work. The only self-citation is Loginova et al. (2024), used to justify balancing answer options; that citation is methodological and not load-bearing for the temporal-reasoning conclusion. The strongest fragility is a validity issue, not circularity: telicity labels were assigned 'irrespective of their context' with Fleiss' kappa 0.67 (Section 4.1), so some QA pairs may be noisy for context-dependent verbs; the authors explicitly acknowledge this ambiguity and the small 157-class inventory in Section 6. Such noise could affect the quantitative gap, but it does not mean the prediction reduces to its own inputs. The empirical ranking of models and the comparison to the human gold standard remain an independent measurement, so the circularity score is low.
Assumptions & free parameters
assumptions (4)
- standard math Allen's interval algebra exhaustively describes possible temporal relations between two actions.
- domain assumption Telicity can be assigned to verb classes in isolation from context.
- domain assumption The semi-synthetic template sentences are natural enough for native speakers to answer from the video.
- domain assumption The Charades videos paired with Action Genome labels provide sufficient visual information to determine the temporal relation between the queried actions.
Cite this review
Pith. "Pith review of Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times." pith.science (2026). https://pith.science/paper/KYV7GWTC
@misc{pith2026250600928,
author = {Pith},
title = {Pith review of: Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times},
year = {2026},
howpublished = {\url{https://pith.science/paper/KYV7GWTC}},
note = {Machine review of arXiv:2506.00928}
}
read the original abstract
Human perception of events is intrinsically tied to distinguishing between completed (perfect and telic) and ongoing (durative) actions, a process mediated by both linguistic structure and visual cues. In this work, we introduce the \textbf{Perfect Times} dataset, a novel, quadrilingual (English, Italian, Russian, and Japanese) multiple-choice question-answering benchmark designed to assess video-language models (VLMs) on temporal reasoning. By pairing everyday activity videos with event completion labels and perfectivity-tailored distractors, our dataset probes whether models truly comprehend temporal dynamics or merely latch onto superficial markers. Experimental results indicate that state-of-the-art models, despite their success on text-based tasks, struggle to mirror human-like temporal and causal reasoning grounded in video. This study underscores the necessity of integrating deep multimodal cues to capture the nuances of action duration and completion within temporal and causal video dynamics, setting a new standard for evaluating and advancing temporal reasoning in VLMs.
Figures
Figures from the paper (12 more)
Reference graph
Works this paper leans on
-
[1]
Marah Abdin, Jyoti Aneja, Hany Awadalla, Ahmed Awadallah, Ammar Ahmad Awan, Nguyen Bach, Amit Bahree, Arash Bakhtiari, Jianmin Bao, Harkirat Behl, Alon Benhaim, Misha Bilenko, Johan Bjorck, Sébastien Bubeck, Martin Cai, Qin Cai, Vishrav Chaudhary, Dong Chen, Dongdong Chen, Weizhu Chen, Yen-Chun Chen, Yi-Ling Chen, Hao Cheng, Parul Chopra, Xiyang Dai, Matt...
arXiv 2024
-
[2]
OpenAI Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, Red Avila, Igor Babuschkin, Suchir Balaji, Valerie Balcom, Paul Baltescu, Haiming Bao, Mo Bavarian, Jeff Belgum, Irwan Bello, Jake Berdine, Gabriel Bernadett-Shapiro, Christopher Berner, Lenny...
work page 2023
-
[3]
Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katie Millican, Malcolm Reynolds, Roman Ring, Eliza Rutherford, Serkan Cabi, Tengda Han, Zhitao Gong, Sina Samangooei, Marianne Monteiro, Jacob Menick, Sebastian Borgeaud, Andy Brock, Aida Nematzadeh, Sahand Sharifzadeh, Mikolaj Binkowski, R...
arXiv 2022
-
[4]
James F. Allen. 1984. Towards a general theory of action and time. Artificial Intelligence, 23(2):123--154
work page 1984
-
[5]
Jinze Bai, Shuai Bai, Yunfei Chu, Zeyu Cui, Kai Dang, Xiaodong Deng, Yang Fan, Wenbin Ge, Yu Han, Fei Huang, Binyuan Hui, Luo Ji, Mei Li, Junyang Lin, Runji Lin, Dayiheng Liu, Gao Liu, Chengqiang Lu, Keming Lu, Jianxin Ma, Rui Men, Xingzhang Ren, Xuancheng Ren, Chuanqi Tan, Sinan Tan, Jianhong Tu, Peng Wang, Shijie Wang, Wei Wang, Shengguang Wu, Benfeng X...
arXiv 2023
-
[6]
Bosch, Mathilde Chailleux, and Francesca Foppolo
Jasmijn E. Bosch, Mathilde Chailleux, and Francesca Foppolo. 2021. https://doi.org/10.3765/elm.1.4880 Incremental processing of telicity in italian children
-
[7]
Mu Cai, Reuben Tan, Jianrui Zhang, Bocheng Zou, Kai Zhang, Feng Yao, Fangrui Zhu, Jing Gu, Yiwu Zhong, Yuzhang Shang, Yao Dou, Jaden Park, Jianfeng Gao, Yong Jae Lee, and Jianwei Yang. 2024. Temporalbench: Towards fine-grained temporal understanding for multimodal video models. arXiv preprint arXiv:2410.10818
arXiv 2024
-
[8]
Franklin Chang, Tomoko Tatsumi, Yuna Hiranuma, and Colin Bannard. 2023. https://api.semanticscholar.org/CorpusID:260332667 Visual heuristics for verb production: Testing a deep-learning model with experiments in japanese . Cognitive science, 47 8:e13324
work page 2023
Show all 61 references
-
[9]
Zhe Chen, Jiannan Wu, Wenhai Wang, Weijie Su, Guo Chen, Sen Xing, Muyan Zhong, Qinglong Zhang, Xizhou Zhu, Lewei Lu, et al. 2024. Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks. In Proceedings of the IEEE/CVF Conference on Comput...
2024
-
[10]
Zesen Cheng, Sicong Leng, Hang Zhang, Yifei Xin, Xin Li, Guanzheng Chen, Yongxin Zhu, Wenqi Zhang, Ziyang Luo, Deli Zhao, and Lidong Bing. 2024. https://arxiv.org/abs/2406.07476 Videollama 2: Advancing spatial-temporal modeling and audio understanding in video-llms . arXiv pre...
2024 arXiv
-
[11]
Gonzalez, Ion Stoica, and Eric P
Wei-Lin Chiang, Zhuohan Li, Zi Lin, Ying Sheng, Zhanghao Wu, Hao Zhang, Lianmin Zheng, Siyuan Zhuang, Yonghao Zhuang, Joseph E. Gonzalez, Ion Stoica, and Eric P. Xing. 2023. https://lmsys.org/blog/2023-03-30-vicuna/ Vicuna: An open-source chatbot impressing gpt-4 with 90\
2023
-
[12]
Hyung Won Chung, Le Hou, S. Longpre, Barret Zoph, Yi Tay, William Fedus, Eric Li, Xuezhi Wang, Mostafa Dehghani, Siddhartha Brahma, Albert Webson, Shixiang Shane Gu, Zhuyun Dai, Mirac Suzgun, Xinyun Chen, Aakanksha Chowdhery, Dasha Valter, Sharan Narang, Gaurav Mishra, Adams W...
2022 arXiv
-
[13]
Daniel Cores, Michael Dorkenwald, Manuel Mucientes, Cees G. M. Snoek, and Yuki M. Asano. 2024. https://api.semanticscholar.org/CorpusID:273233203 Lost in time: A new temporal benchmark for videollms
2024
-
[14]
Dima Damen, Hazel Doughty, Giovanni Maria Farinella, Sanja Fidler, Antonino Furnari, Evangelos Kazakos, Davide Moltisanti, Jonathan Munro, Toby Perrett, Will Price, and Michael Wray. 2020. https://api.semanticscholar.org/CorpusID:218470131 The epic-kitchens dataset: Collection...
2020
-
[15]
Bosch, Ciro Greco, Maria Nella Carminati, and Francesca Panzeri
Francesca Foppolo, Jasmijn E. Bosch, Ciro Greco, Maria Nella Carminati, and Francesca Panzeri. 2021. https://api.semanticscholar.org/CorpusID:238423539 Draw a star and make it perfect: Incremental processing of telicity . Cognitive science, 45 10:e13052
2021
-
[16]
Francesca Foppolo, Francesca Panzeri, Cesare Greco, and Matteo Carminati. 2016. https://api.semanticscholar.org/CorpusID:125787569 The incremental processing of accomplishment predicates
2016
-
[17]
Chaoyou Fu, Yuhan Dai, Yongdong Luo, Lei Li, Shuhuai Ren, Renrui Zhang, Zihan Wang, Chenyu Zhou, Yunhang Shen, Mengdan Zhang, Peixian Chen, Yanwei Li, Shaohui Lin, Sirui Zhao, Ke Li, Tong Xu, Xiawu Zheng, Enhong Chen, Rongrong Ji, and Xing Sun. 2024. https://arxiv.org/abs/2405...
2024 arXiv
-
[18]
Fabian Caba Heilbron, Victor Escorcia, Bernard Ghanem, and Juan Carlos Niebles. 2015. https://api.semanticscholar.org/CorpusID:1710722 Activitynet: A large-scale video benchmark for human activity understanding . 2015 IEEE Conference on Computer Vision and Pattern Recognition ...
2015
-
[19]
Jang, Yale Song, Chris Dongjoo Kim, Youngjae Yu, Youngjin Kim, and Gunhee Kim
Y. Jang, Yale Song, Chris Dongjoo Kim, Youngjae Yu, Youngjin Kim, and Gunhee Kim. 2019. https://api.semanticscholar.org/CorpusID:190638738 Video question answering with spatio-temporal reasoning . International Journal of Computer Vision, 127:1385 -- 1412
2019
-
[20]
Jingwei Ji, Ranjay Krishna, Li Fei-Fei, and Juan Carlos Niebles. 2019. https://api.semanticscholar.org/CorpusID:209376177 Action genome: Actions as compositions of spatio-temporal scene graphs . 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages ...
2019
-
[21]
Will Kay, Jo \ a o Carreira, Karen Simonyan, Brian Zhang, Chloe Hillier, Sudheendra Vijayanarasimhan, Fabio Viola, Tim Green, Trevor Back, Apostol Natsev, Mustafa Suleyman, and Andrew Zisserman. 2017. https://api.semanticscholar.org/CorpusID:27300853 The kinetics human action ...
2017 arXiv
-
[22]
Richard Landis and Gary G
J. Richard Landis and Gary G. Koch. 1977. The measurement of observer agreement for categorical data. Biometrics, 33(1)
1977
-
[23]
Jie Lei, Licheng Yu, Mohit Bansal, and Tamara L. Berg. 2018. https://api.semanticscholar.org/CorpusID:52171684 Tvqa: Localized, compositional video question answering . In Conference on Empirical Methods in Natural Language Processing
2018
-
[24]
Junnan Li, Dongxu Li, Silvio Savarese, and Steven C. H. Hoi. 2023. https://api.semanticscholar.org/CorpusID:256390509 Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models . In International Conference on Machine Learning
2023
-
[25]
Junnan Li, Dongxu Li, Caiming Xiong, and Steven C. H. Hoi. 2022. https://api.semanticscholar.org/CorpusID:246411402 Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation . In International Conference on Machine Learning
2022
-
[26]
Bin Lin, Bin Zhu, Yang Ye, Munan Ning, Peng Jin, and Li Yuan. 2023. https://api.semanticscholar.org/CorpusID:265281544 Video-llava: Learning united visual representation by alignment before projection . ArXiv, abs/2311.10122
2023 arXiv
-
[27]
Yi Liu, Limin Wang, Xiao Ma, Yali Wang, and Y. Qiao. 2021. https://api.semanticscholar.org/CorpusID:237485612 Fineaction: A fine-grained video dataset for temporal action localization . IEEE Transactions on Image Processing, 31:6937--6950
2021
-
[28]
Yuanxin Liu, Shicheng Li, Yi Liu, Yuxiang Wang, Shuhuai Ren, Lei Li, Sishuo Chen, Xu Sun, and Lu Hou. 2024. https://arxiv.org/abs/2403.00476 Tempcompass: Do video llms really understand videos? Preprint, arXiv:2403.00476
2024 arXiv
-
[29]
Olga Loginova, Oleksandr Bezrukov, and Alexey Kravets. 2024. https://arxiv.org/abs/2410.14248 Addressing blind guessing: Calibration of selection bias in multiple-choice question answering by video language models . Preprint, arXiv:2410.14248
2024 arXiv
-
[30]
Agnese Lombardi and Alessandro Lenci. 2023. https://arxiv.org/abs/2307.02910 Agentivit\`a e telicit\`a in gilberto: implicazioni cognitive . Preprint, arXiv:2307.02910
2023 arXiv
-
[31]
Haoyu Lu, Wen Liu, Bo Zhang, Bingxuan Wang, Kai Dong, Bo Liu, Jingxiang Sun, Tongzheng Ren, Zhuoshu Li, Hao Yang, Yaofeng Sun, Chengqi Deng, Hanwei Xu, Zhenda Xie, and Chong Ruan. 2024. https://arxiv.org/abs/2403.05525 Deepseek-vl: Towards real-world vision-language understand...
2024 arXiv
-
[32]
Eleni Metheniti, Tim Van De Cruys, and Nabil Hathout. 2022. https://doi.org/10.18653/v1/2022.cmcl-1.10 About time: Do transformers learn temporal verbal aspect? In Proceedings of the Workshop on Cognitive Modeling and Computational Linguistics, pages 88--101, Dublin, Ireland. ...
2022 doi
-
[33]
Antoine Miech, Dimitri Zhukov, Jean-Baptiste Alayrac, Makarand Tapaswi, Ivan Laptev, and Josef Sivic. 2019. How T o100 M : L earning a T ext- V ideo E mbedding by W atching H undred M illion N arrated V ideo C lips. In ICCV
2019
-
[34]
Serge Minor, Natalia Mitrofanova, Gustavo Guajardo, Myrte Vos, and Gillian Catriona Ramchand. 2022. https://api.semanticscholar.org/CorpusID:255297282 Temporal information and event bounding across languages: Evidence from visual world eyetracking from . Semantics and Linguist...
2022
-
[35]
Marc Moens and Mark Steedman. 1988. Temporal ontology and temporal reference. Comput. Linguist., 14(2):15–28
1988
-
[36]
Viorica Pătrăucean, Lucas Smaira, Ankush Gupta, Adrià Recasens Continente, Larisa Markeeva, Dylan Banarse, Skanda Koppula, Joseph Heyward, Mateusz Malinowski, Yi Yang, Carl Doersch, Tatiana Matejovicova, Yury Sulsky, Antoine Miech, Alex Frechette, Hanna Klimczak, Raphael Koste...
2023
-
[37]
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. 2021. https://api.semanticscholar.org/CorpusID:231591445 Learning transferable visual model...
2021
-
[38]
Nishant Rai, Haofeng Chen, Jingwei Ji, Rishi Desai, Kazuki Kozuka, Shun Ishizaka, Ehsan Adeli, and Juan Carlos Niebles. 2021. https://api.semanticscholar.org/CorpusID:234357543 Home action genome: Cooperative compositional action understanding . 2021 IEEE/CVF Conference on Com...
2021
-
[39]
Anwer, Tim Baldwin, Michael Felsberg, and Fahad S
Hanoona Rasheed, Muhammad Maaz, Abdelrahman Shaker, Salman Khan, Hisham Cholakal, Rao M. Anwer, Tim Baldwin, Michael Felsberg, and Fahad S. Khan. 2025. Palo: A large multilingual multimodal language model. In Proceedings of the IEEE/CVF Winter Conference on Applications of Com...
2025
-
[40]
Arka Sadhu, Kan Chen, and Ram Nevatia. 2021. https://api.semanticscholar.org/CorpusID:233181708 Video question answering with phrases via semantic roles . ArXiv, abs/2104.03762
2021 arXiv
-
[41]
Sigurdsson, G \"u l Varol, X
Gunnar A. Sigurdsson, G \"u l Varol, X. Wang, Ali Farhadi, Ivan Laptev, and Abhinav Kumar Gupta. 2016. https://api.semanticscholar.org/CorpusID:18061547 Hollywood in homes: Crowdsourcing data collection for activity understanding . In European Conference on Computer Vision
2016
-
[42]
Gemini Team, Rohan Anil, Sebastian Borgeaud, Jean-Baptiste Alayrac, Jiahui Yu, Radu Soricut, Johan Schalkwyk, Andrew M. Dai, Anja Hauth, Katie Millican, David Silver, Melvin Johnson, Ioannis Antonoglou, Julian Schrittwieser, Amelia Glaese, Jilin Chen, Emily Pitler, Timothy Lil...
2024 arXiv
-
[43]
Carol Tenny. 1994. https://api.semanticscholar.org/CorpusID:62581588 Aspectual roles and the syntax-semantics interface
1994
-
[44]
Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, Daniel M
Hugo Touvron, Louis Martin, Kevin R. Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, Daniel M. Bikel, Lukas Blecher, Cristian Cant \'o n Ferrer, Moya Chen, Guillem Cucurull, David Esiobu, Jude Fernandes,...
2023 arXiv
-
[45]
Angeliek van Hout. 2008. https://www.sciencedirect.com/science/article/pii/S0024384107001520 Acquiring perfectivity and telicity in dutch, italian and polish . Lingua, 118 11:1740--1765
2008
-
[46]
Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N
Ashish Vaswani, Noam M. Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. In Neural Information Processing Systems
2017
-
[47]
Peng Wang, Shuai Bai, Sinan Tan, Shijie Wang, Zhihao Fan, Jinze Bai, Keqin Chen, Xuejing Liu, Jialin Wang, Wenbin Ge, Yang Fan, Kai Dang, Mengfei Du, Xuancheng Ren, Rui Men, Dayiheng Liu, Chang Zhou, Jingren Zhou, and Junyang Lin. 2024. Qwen2-vl: Enhancing vision-language mode...
2024 arXiv
-
[48]
Shaojie Wang, Wentian Zhao, Ziyi Kou, Jing Shi, and Chenliang Xu. 2021. https://api.semanticscholar.org/CorpusID:231813023 How to make a blt sandwich? learning vqa towards understanding web instructional videos . 2021 IEEE Winter Conference on Applications of Computer Vision (...
2021
-
[49]
Haiwan Wei, Yitian Yuan, Xiaohan Lan, Wei Ke, and Lin Ma. 2025. https://api.semanticscholar.org/CorpusID:277621021 Instructionbench: An instructional video understanding benchmark . ArXiv, abs/2504.05040
2025 arXiv
-
[50]
Bo Wu and Shoubin Yu. 2021. https://api.semanticscholar.org/CorpusID:244907062 Star: A benchmark for situated reasoning in real-world videos . In NeurIPS Datasets and Benchmarks
2021
-
[51]
Junbin Xiao, Xindi Shang, Angela Yao, and Tat seng Chua. 2021. https://api.semanticscholar.org/CorpusID:234763093 Next-qa: Next phase of question-answering to explaining temporal actions . 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 9772--9781
2021
-
[52]
Xu, Zhou Zhao, Jun Xiao, Fei Wu, Hanwang Zhang, Xiangnan He, and Yueting Zhuang
D. Xu, Zhou Zhao, Jun Xiao, Fei Wu, Hanwang Zhang, Xiangnan He, and Yueting Zhuang. 2017. https://api.semanticscholar.org/CorpusID:3864050 Video question answering via gradually refined attention over appearance and motion . Proceedings of the 25th ACM international conference...
2017
-
[53]
Yuan Yao, Tianyu Yu, Ao Zhang, Chongyi Wang, Junbo Cui, Hongji Zhu, Tianchi Cai, Haoyu Li, Weilin Zhao, Zhihui He, et al. 2024. Minicpm-v: A gpt-4v level mllm on your phone. arXiv preprint arXiv:2408.01800
2024 arXiv
-
[54]
Xu, Jun Yu, Ting Yu, Zhou Zhao, Yueting Zhuang, and Dacheng Tao
Zhou Yu, D. Xu, Jun Yu, Ting Yu, Zhou Zhao, Yueting Zhuang, and Dacheng Tao. 2019. https://api.semanticscholar.org/CorpusID:69645185 Activitynet-qa: A dataset for understanding complex web videos via question answering . ArXiv, abs/1906.02467
2019 arXiv
-
[55]
Xiang Yue, Yueqi Song, Akari Asai, Seungone Kim, Jean de Dieu Nyandwi, Simran Khanuja, Anjali Kantharuban, Lintang Sutawika, Sathyanarayanan Ramamoorthy, and Graham Neubig. 2024. https://arxiv.org/abs/2410.16153 Pangea: A fully open multilingual multimodal llm for 39 languages...
2024 arXiv
-
[56]
Xiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, and Lucas Beyer. 2023. https://arxiv.org/abs/2303.15343 Sigmoid loss for language image pre-training . Preprint, arXiv:2303.15343
2023 arXiv
-
[57]
Yuanhan Zhang, Jinming Wu, Wei Li, Bo Li, Zejun Ma, Ziwei Liu, and Chunyuan Li. 2024. https://arxiv.org/abs/2410.02713 Video instruction tuning with synthetic data . Preprint, arXiv:2410.02713
2024 arXiv
-
[58]
Yiyun Zhao, Jian Gang Ngui, Lucy Hall Hartley, and Steven Bethard. 2021. https://doi.org/10.18653/v1/2021.conll-1.6 Do pretrained transformers infer telicity like humans? In Proceedings of the 25th Conference on Computational Natural Language Learning, pages 72--81, Online. As...
2021 doi
-
[59]
Yaoyao Zhong, Wei Ji, Junbin Xiao, Yicong Li, Wei Deng, and Tat seng Chua. 2022. https://api.semanticscholar.org/CorpusID:247218478 Video question answering: Datasets, algorithms and challenges . ArXiv, abs/2203.01225
2022 arXiv
-
[60]
online" 'onlinestring :=
ENTRY address archivePrefix author booktitle chapter edition editor eid eprint eprinttype howpublished institution journal key month note number organization pages publisher school series title type volume year doi pubmed url lastchecked label extra.label sort.label short.list...
-
[61]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 gl...
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.