REVIEW 3 major objections 5 minor 91 references
InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models
T0 review · 3 major / 5 minor · reviewed 2026-08-10 · deepseek-v4-flash
Pith's one-line read Injecting a procedural graph—automatically mined from training videos—into a multimodal large language model improves its real-time understanding of instructional tasks across recognition, prediction, and error-detection benchmarks.
desk verdict Core graph-injection claim holds up across backbones, but the closed-set evaluation and graph-built error tasks mean part of the gain is lookup, not understanding; still worth a serious review. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the procedural graph, a directed graph whose nodes are action step labels and whose edges are temporally ordered transitions between consecutive steps observed in training videos. The load-bearing operation is the online search path: as the video unfolds, the model's predicted action text is mapped to a node via one-hot similarity, and the resulting prefix path is inserted into the prompt. This path carries the dependency information that the LLM would otherwise have to infer from pixels alone.
What would settle it
Take a held-out set of instructional videos whose step sequences deliberately include transitions absent from the training-derived graph, and compare the graph-augmented model against the video-plus-query-only model on action prediction and plan prediction. If the graph-augmented model does not beat the base model on these out-of-graph videos—or performs worse—then the proposed benefit depends on closed-world coverage rather than on the graph representation itself.
Extended reading notes
Core claim
On its own terms, the paper establishes that conditioning the answer distribution on a procedural graph consistently improves a multimodal LLM over conditioning on video and query alone. The graph is constructed before training by scanning training videos, collecting step annotations as nodes, and adding an edge for every temporally adjacent pair of steps. At inference, the model recognizes each action as free-form text, maps it onto the nearest graph node, and accumulates a predicted online path; that path is then verbalized and fed into the LLM along with frames and prompts. The experiments report gains on every tested task on both COIN and CrossTask, including large absolute improvements when the same graph is appended to existing online video LLMs and proprietary API-based models, and the authors take this as evidence that explicit dependency structure eases the reasoning burden on the LLM.
Load-bearing premise
The procedural graph is built only from training videos, so the entire method assumes that every action and transition the assistant will see at test time is already represented in the graph; if a test video contains an unseen step or a new ordering, the mapped graph path becomes empty or misleading.
Editorial extensions
If this is right
- If the central claim holds, adding a mined procedural graph is a drop-in improvement for any multimodal LLM in this setting; the experiments show gains even without retraining the underlying model.
- Explicit graphs reduce the LLM's need to do long-range planning from raw video, so gains should be largest on tasks with longer dependencies, such as plan prediction, and on detection of inserted wrong steps.
- The same graph can generate streaming-dialog training data, lowering the annotation effort for multi-turn assistance.
- Because the graph is built from training data, deployments whose test procedures overlap with training procedures can benefit without new annotation.
Reading between the lines
- The paper's comparison of graph sources suggests a design principle: the best source may depend on whether the task rewards focused local context or broad dependency coverage; this could be tested by varying the retrieval pool size.
- The larger gain on wrong-step detection than on order-error detection suggests the graph mainly encodes local adjacency; a testable extension is to add explicit global ordering constraints.
- A deployed assistant would need the graph to grow as new tasks appear, since the online path is built from a fixed graph; incremental graph expansion at inference is a natural next step.
- The paper's future-work admission that errors compound along the graph path implies that better node-mapping or error-recovery mechanisms, rather than larger LLMs, are the main lever for plan prediction.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper presents InsTALL, a multi-modal LLM for real-time assistance on instructional videos. It combines a VideoLLM-style architecture (CLIP encoder, temporal pooling, MLP, Mistral-7B-Instruct with LoRA) with a procedural graph G that is automatically mined from training videos. G is verbalized and injected into the LLM prompt at training and inference time; at inference, the model's recognized actions are mapped onto G to form an online search path bG_t that conditions all subsequent predictions. The model is trained on five sub-tasks—task recognition, action recognition, next-action prediction, plan prediction with and without the goal—and two auxiliary error-detection tasks introduced by corrupting video steps. Experiments on COIN and CrossTask report state-of-the-art results, including large gains when the graph is added to multiple MLLM backbones (VideoLLM-online+, GPT-4o-mini, GPT-4-turbo, GPT-4o).
Significance. If the reported effects are stable, the work makes a useful contribution: it shows a data-driven procedural graph can be injected into MLLMs as an inductive bias, consistently improving several backbones (Table 8), a finding that generalizes the idea of GraphRAG to multimodal instruction following. The algorithmic specification of graph construction (Alg. 1) and online path formation (Alg. 2) is clear and reproducible. The two error-detection tasks are a practical addition for assistive systems. However, the strength of the evidence is tempered by the absence of statistical significance reporting, by the fact that the error tasks are constructed from the same graph used for prediction, and by the lack of robustness analysis for the online path mapping.
major comments (3)
- [Section 5.3, Table 8] Table 8 reports single-run accuracies for each backbone and condition, with no error bars, confidence intervals, or significance tests. Since some gains are small (e.g., +0.6 to +3.3 for VideoLLM-online+ on several tasks) and MLLM training and few-shot prompting are known to be noisy, the claim of 'unanimous improvement' is not statistically substantiated. Please provide results over multiple seeds with a paired significance test across tasks, or at least report variance, to confirm the gains are not within run-to-run noise.
- [Section 4.4, Eqs. (8)-(9)] The two error-detection tasks are generated by substituting a step v not in N(vt) ∪ vt and by shuffling to a sequence not present in the graph G. Consequently, an oracle that only has access to G can solve these tasks without any video understanding. The VQG vs. VQ gap in Table 7C could therefore reflect graph-lookup rather than visual error detection. Please add an ablation that either withholds the graph at inference (while still training with it) or constructs errors using visually plausible out-of-vocabulary actions, to disentangle the graph's role from genuine visual error understanding.
- [Section 4.2, Eq. (7) and Future Works] The model conditions all predictions on the online search path bG_t, which is obtained by one-hot mapping the AR output to graph nodes. If the AR output is incorrect or the action is absent from G, bG_t is corrupted and the error propagates to all subsequent conditioned predictions. The Future Works paragraph admits that such errors compound, but the evaluation never isolates this failure mode. Since COIN and CrossTask test sets share the closed action taxonomy used to build G, the benefit of graph injection under recognition failures or out-of-vocabulary actions is unknown. Add a robustness experiment (e.g., injecting controlled AR errors or holding out a subset of action nodes from G) and report how AR, AP, PP, and PP+ degrade relative to the VQ baseline.
minor comments (5)
- [Abstract vs. Section 4.3] The abstract in the main text states that the graph is leveraged 'at inference time,' while the abstract in the submission header and Section 4.3 describe using the graph at both training and inference time. The text should be made consistent, since training-time graph verbalization is explicitly described.
- [Table 4] The 'Pooling' row in Table 4 lists the shape as 'T×N×D/W'; it is unclear what W denotes and how spatial pooling is implemented. Please specify the exact tensor transformation.
- [Eq. (1)] Equation (1) writes 'min EV,Y' without defining the expectation; please clarify over which random variables the expectation is taken and how it is approximated empirically in training.
- [Table 8] Table 8 omits the VideoLLM-online (VQ) row on CrossTask, which makes it difficult to compare the magnitude of graph gains across datasets. Please add the row or explain why this baseline is not evaluated on CrossTask.
- [Section 4.2, Eq. (7)] The node mapping bv_t = arg max 1_VG(at) is ambiguous when at is free-form text; specify the similarity metric used (e.g., cosine similarity between text embeddings) and the tie-breaking procedure for matching to graph nodes.
Circularity Check
Central TR/AR/AP/PP results are independent visual predictions, but the two novel error-detection tasks define ground truth as deviations from the same procedural graph injected as input, so those gains partly reduce to graph-membership lookup.
-
self definitional
[Section 4.4, Incorrect Action Detection (Eq. 8), with graph conditioning from Eq. 6]
"We create samples with incorrect actions by modifying each video in the data. Precisely, we randomly replace one step with an incorrect step v /∈N (vt)∪vt. This leads to an erroneous graph ¯G generation. The task is to identify this mistaken step within the sequence"
The ground-truth label is defined as v∉N(vt)∪vt, i.e., non-membership in the graph neighborhood. The same graph G is an input to the model at inference (Eq. 6), and the online path is mapped onto G via Eq. 7. Therefore, once the model recognizes the action text, the correct error judgment is exactly a graph-membership check. The label is constructed from the very graph that is provided as input, so the reported error-detection improvement partly measures lookup on the injected graph rather than independent visual error detection. Visual recognition of the action node is still required, so the reduction is partial rather than total.
-
self definitional
[Section 4.4, Incorrect Order Detection (Eq. 9), with graph conditioning from Eq. 6]
"By randomly shuffling the order of steps, we create a dataset for detecting mistakes in the ordering. ... we ensure the randomly shuffled order of actions is different from all action orderings present in the videos belonging to a particular task."
The 'incorrect order' label is defined as an ordering that is absent from the set of orderings encoded in G. Since G is supplied to the model at inference (Eq. 6) and the recognized action sequence is mapped to nodes of G (Eq. 7), the model can answer by checking whether the recognized sequence is a valid path in G. The ground truth is thus a property of the input graph, not an independent temporal-visual judgment. This makes the order-error sub-task self-definitional: the answer rule is contained in the graph given as input, with visual recognition again providing the action nodes.
full rationale
The paper's main tasks (TR, AR, AP, PP, PP+) are not circular in a damaging way: the procedural graph is mined from training action annotations, test videos are disjoint, and the model still has to map pixels to action nodes via visual recognition before the graph can help. Conditioning on the graph is a legitimate prior, and the online path bGt is built from the model's own outputs rather than from ground-truth labels. The Future Works admission that 'errors compound across prediction steps' is a limitation of that self-conditioning loop, not a circular definition. The two novel error-detection sub-tasks, however, are constructed by corrupting videos with actions/orders defined relative to the very graph G that is injected as input at inference (Eq. 6). For those tasks the ground truth is a graph-membership condition, so the reported gains over non-graph baselines partly reflect lookup on the input graph rather than an independent visual error-detection capability. Because these are auxiliary tasks and the central recognition/prediction claims retain independent visual content, the overall circularity is partial, not total.
Assumptions & free parameters
free parameters (4)
- Temporal aggregation operation =
Pooling (selected empirically, Table 4)
- CLIP backbone =
CLIP-H-14 (selected via retrieval F1, Table 5)
- Graph construction source =
Entire training set (Table 6)
- LoRA hyperparameters =
rank=128, scale=256
assumptions (3)
- domain assumption The procedural graph G constructed from training data (Alg. 1) is a faithful and comprehensive model of the action dependencies for test videos.
- domain assumption The one-hot similarity mapping from free-form LLM action text to graph nodes (Eq. 7) is sufficiently reliable to build the online path.
- domain assumption The LLM (Mistral-7B-Instruct) can effectively exploit the injected graph tokens alongside video tokens.
Cite this review
Pith. "Pith review of InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models." pith.science (2026). https://pith.science/paper/G5XNBDHL
@misc{pith2026250112231,
author = {Pith},
title = {Pith review of: InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models},
year = {2026},
howpublished = {\url{https://pith.science/paper/G5XNBDHL}},
note = {Machine review of arXiv:2501.12231}
}
read the original abstract
The improved competence of generative models can help building multi-modal virtual assistants that leverage modalities beyond language. By observing humans performing multi-step tasks, one can build assistants that have situational awareness of actions and tasks being performed, enabling them to cater assistance based on this understanding. In this paper, we develop a Context-aware Instructional Task Assistant with Multi-modal Large Language Models (InsTALL) that leverages an online visual stream (e.g. a user's screen share or video recording) and responds in real-time to user queries related to the task at hand. To enable useful assistance, InsTALL 1) trains a multi-modal model on task videos and paired textual data, and 2) automatically extracts task graph from video data and leverages it at training and inference time. We show InsTALL achieves state-of-the-art performance across proposed sub-tasks considered for multimodal activity understanding -- task recognition (TR), action recognition (AR), next action prediction (AP), and plan prediction (PP) -- and outperforms existing baselines on two novel sub-tasks related to automatic error identification.
Figures
Reference graph
Works this paper leans on
-
[1]
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al. Gpt-4 technical report. arXiv preprint arXiv:2303.08774, 2023. 7
arXiv 2023
-
[2]
Ht-step: Aligning instructional articles with how-to videos
Triantafyllos Afouras, Effrosyni Mavroudi, Tushar Nagara- jan, Huiyu Wang, and Lorenzo Torresani. Ht-step: Aligning instructional articles with how-to videos. Advances in Neural Information Processing Systems, 36, 2024. 2, 4
2024
-
[3]
Unsuper- vised learning from narrated instruction videos
Jean-Baptiste Alayrac, Piotr Bojanowski, Nishant Agrawal, Josef Sivic, Ivan Laptev, and Simon Lacoste-Julien. Unsuper- vised learning from narrated instruction videos. In Proceed- ings of the IEEE conference on computer vision and pattern recognition, pages 4575–4583, 2016. 1
2016
-
[4]
Flamingo: a visual language model for few-shot learning
Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katherine Millican, Malcolm Reynolds, et al. Flamingo: a visual language model for few-shot learning. Advances in neural information processing systems, 35:23716–23736,
-
[5]
Localizing moments in video with natural language
Lisa Anne Hendricks, Oliver Wang, Eli Shechtman, Josef Sivic, Trevor Darrell, and Bryan Russell. Localizing moments in video with natural language. In Proceedings of the IEEE international conference on computer vision , pages 5803– 5812, 2017. 2
2017
-
[6]
https://www.anthropic.com/news/developing- computer-use, 2024
Anthropic. https://www.anthropic.com/news/developing- computer-use, 2024. 2
2024
-
[7]
Video-mined task graphs for keystep recognition in instructional videos
Kumar Ashutosh, Santhosh Kumar Ramakrishnan, Triantafyl- los Afouras, and Kristen Grauman. Video-mined task graphs for keystep recognition in instructional videos. Advances in Neural Information Processing Systems, 36, 2024. 2, 3, 5, 8
2024
-
[8]
Detours for navigating instructional videos
Kumar Ashutosh, Zihui Xue, Tushar Nagarajan, and Kristen Grauman. Detours for navigating instructional videos. In Pro- ceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 18804–18815, 2024. 1
2024
Show all 91 references
-
[9]
Is space-time attention all you need for video understanding? In ICML, page 4, 2021
Gedas Bertasius, Heng Wang, and Lorenzo Torresani. Is space-time attention all you need for video understanding? In ICML, page 4, 2021. 2, 8
2021
-
[10]
Procedure planning in instructional videos via contextual modeling and model-based policy learning
Jing Bi, Jiebo Luo, and Chenliang Xu. Procedure planning in instructional videos via contextual modeling and model-based policy learning. In Proceedings of the IEEE/CVF Interna- tional Conference on Computer Vision, pages 15611–15620,
-
[11]
Activitynet: A large-scale video bench- mark for human activity understanding
Fabian Caba Heilbron, Victor Escorcia, Bernard Ghanem, and Juan Carlos Niebles. Activitynet: A large-scale video bench- mark for human activity understanding. In Proceedings of the ieee conference on computer vision and pattern recognition, pages 961–970, 2015. 2
2015
-
[12]
Procedure planning in instructional videos
Chien-Yi Chang, De-An Huang, Danfei Xu, Ehsan Adeli, Li Fei-Fei, and Juan Carlos Niebles. Procedure planning in instructional videos. In European Conference on Computer Vision, pages 334–350. Springer, 2020. 2
2020
-
[13]
Temporally grounding natural sentence in video
Jingyuan Chen, Xinpeng Chen, Lin Ma, Zequn Jie, and Tat- Seng Chua. Temporally grounding natural sentence in video. In Proceedings of the 2018 conference on empirical methods in natural language processing, pages 162–171, 2018. 2
2018
-
[14]
Videollm-online: Online video large language model for streaming video
Joya Chen, Zhaoyang Lv, Shiwei Wu, Kevin Qinghong Lin, Chenan Song, Difei Gao, Jia-Wei Liu, Ziteng Gao, Dongxing Mao, and Mike Zheng Shou. Videollm-online: Online video large language model for streaming video. In Proceedings of the IEEE/CVF Conference on Computer Vision and P...
2024
-
[15]
Semantic proposal for activity localization in videos via sentence query
Shaoxiang Chen and Yu-Gang Jiang. Semantic proposal for activity localization in videos via sentence query. In Proceedings of the AAAI Conference on Artificial Intelligence, pages 8199–8206, 2019. 2
2019
-
[16]
KGPT: Knowledge-grounded pre-training for data-to-text generation
Wenhu Chen, Yu Su, Xifeng Yan, and William Yang Wang. KGPT: Knowledge-grounded pre-training for data-to-text generation. In Proceedings of the 2020 Conference on Em- pirical Methods in Natural Language Processing (EMNLP), pages 8635–8648, Online, 2020. Association for Computa-...
2020
-
[17]
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Syl- vain Gelly, Jakob Uszkoreit, and Neil Houlsby. An image is worth 16x16 words: Transformers for image recognition ...
2021
-
[18]
From local to global: A graph rag approach to query-focused summarization
Darren Edge, Ha Trinh, Newman Cheng, Joshua Bradley, Alex Chao, Apurva Mody, Steven Truitt, and Jonathan Larson. From local to global: A graph rag approach to query-focused summarization. arXiv preprint arXiv:2404.16130, 2024. 2, 3
2024 arXiv
-
[19]
Masked diffusion with task-awareness for procedure planning in instructional videos
Fen Fang, Yun Liu, Ali Koksal, Qianli Xu, and Joo-Hwee Lim. Masked diffusion with task-awareness for procedure planning in instructional videos. arXiv preprint arXiv:2309.07409 ,
-
[20]
Slowfast networks for video recognition
Christoph Feichtenhofer, Haoqi Fan, Jitendra Malik, and Kaiming He. Slowfast networks for video recognition. In Proceedings of the IEEE/CVF international conference on computer vision, pages 6202–6211, 2019. 8
2019
-
[21]
Tall: Temporal activity localization via language query
Jiyang Gao, Chen Sun, Zhenheng Yang, and Ram Nevatia. Tall: Temporal activity localization via language query. In Proceedings of the IEEE international conference on com- puter vision, pages 5267–5275, 2017. 2
2017
-
[22]
Retrieval- augmented generation for large language models: A survey
Yunfan Gao, Yun Xiong, Xinyu Gao, Kangxiang Jia, Jinliu Pan, Yuxi Bi, Yi Dai, Jiawei Sun, and Haofen Wang. Retrieval- augmented generation for large language models: A survey. arXiv preprint arXiv:2312.10997, 2023. 2
2023 arXiv
-
[23]
Mac: Mining activity concepts for language-based temporal local- ization
Runzhou Ge, Jiyang Gao, Kan Chen, and Ram Nevatia. Mac: Mining activity concepts for language-based temporal local- ization. In 2019 IEEE winter conference on applications of computer vision (WACV), pages 245–253. IEEE, 2019. 2
2019
-
[24]
Imagebind: One embedding space to bind them all
Rohit Girdhar, Alaaeldin El-Nouby, Zhuang Liu, Mannat Singh, Kalyan Vasudev Alwala, Armand Joulin, and Ishan Misra. Imagebind: One embedding space to bind them all. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 15180–15190, 2023. 2
2023
-
[25]
Radar: automated task planning for proactive decision support
Sachin Grover, Sailik Sengupta, Tathagata Chakraborti, Aditya Prasad Mishra, and Subbarao Kambhampati. Radar: automated task planning for proactive decision support. Human–Computer Interaction, 35(5-6):387–412, 2020. 2
2020
-
[26]
Retrieval augmented language model pre- training
Kelvin Guu, Kenton Lee, Zora Tung, Panupong Pasupat, and Mingwei Chang. Retrieval augmented language model pre- training. In International conference on machine learning, pages 3929–3938. PMLR, 2020. 2
2020
-
[27]
Ma-lmm: Memory-augmented large multimodal model for long-term video understanding
Bo He, Hengduo Li, Young Kyun Jang, Menglin Jia, Xue- fei Cao, Ashish Shah, Abhinav Shrivastava, and Ser-Nam Lim. Ma-lmm: Memory-augmented large multimodal model for long-term video understanding. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recogni...
2024
-
[28]
Denoising diffu- sion probabilistic models
Jonathan Ho, Ajay Jain, and Pieter Abbeel. Denoising diffu- sion probabilistic models. Advances in neural information processing systems, 33:6840–6851, 2020. 2
2020
-
[29]
LoRA: Low-rank adaptation of large language models
Edward J Hu, yelong shen, Phillip Wallis, Zeyuan Allen- Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. LoRA: Low-rank adaptation of large language models. In International Conference on Learning Representations, 2022. 4, 6
2022
-
[30]
Audiogpt: Understanding and generating speech, music, sound, and talking head
Rongjie Huang, Mingze Li, Dongchao Yang, Jiatong Shi, Xuankai Chang, Zhenhui Ye, Yuning Wu, Zhiqing Hong, Jiawei Huang, Jinglin Liu, et al. Audiogpt: Understanding and generating speech, music, sound, and talking head. In Proceedings of the AAAI Conference on Artificial Intell...
2024
-
[31]
Language is not all you need: Aligning perception with language models
Shaohan Huang, Li Dong, Wenhui Wang, Yaru Hao, Saksham Singhal, Shuming Ma, Tengchao Lv, Lei Cui, Owais Khan Mohammed, Barun Patra, et al. Language is not all you need: Aligning perception with language models. Advances in Neural Information Processing Systems, 36:72096–72109,
-
[32]
Leveraging passage retrieval with generative models for open domain question answering
Gautier Izacard and Edouard Grave. Leveraging passage retrieval with generative models for open domain question answering. In Proceedings of the 16th Conference of the Eu- ropean Chapter of the Association for Computational Linguis- tics: Main Volume, pages 874–880, Online, 20...
2021
-
[33]
Mistral 7b
Albert Q Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, et al. Mistral 7b. arXiv preprint arXiv:2310.06825,
-
[34]
Gen- erating images with multimodal language models
Jing Yu Koh, Daniel Fried, and Russ R Salakhutdinov. Gen- erating images with multimodal language models. Advances in Neural Information Processing Systems, 36, 2024. 2
2024
-
[35]
Wikihow: A large scale text summarization dataset
Mahnaz Koupaee and William Yang Wang. Wikihow: A large scale text summarization dataset. arXiv preprint arXiv:1810.09305, 2018. 4, 6
2018 arXiv
-
[36]
A survey on temporal sentence grounding in videos
Xiaohan Lan, Yitian Yuan, Xin Wang, Zhi Wang, and Wenwu Zhu. A survey on temporal sentence grounding in videos. ACM Transactions on Multimedia Computing, Communica- tions and Applications, 19(2):1–33, 2023. 2
2023
-
[37]
Tvr: A large-scale dataset for video-subtitle moment retrieval
Jie Lei, Licheng Yu, Tamara L Berg, and Mohit Bansal. Tvr: A large-scale dataset for video-subtitle moment retrieval. In Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part XXI 16, pages 447–463. Springer, 2020. 2
2020
-
[38]
Less is more: Clipbert for video-and-language learning via sparse sampling
Jie Lei, Linjie Li, Luowei Zhou, Zhe Gan, Tamara L Berg, Mohit Bansal, and Jingjing Liu. Less is more: Clipbert for video-and-language learning via sparse sampling. In Proceed- ings of the IEEE/CVF conference on computer vision and pattern recognition, pages 7331–7341, 2021. 2, 8
2021
-
[39]
Chatting makes perfect: Chat-based image retrieval
Matan Levy, Rami Ben-Ari, Nir Darshan, and Dani Lischin- ski. Chatting makes perfect: Chat-based image retrieval. Ad- vances in Neural Information Processing Systems, 36, 2024. 2
2024
-
[40]
Retrieval- augmented generation for knowledge-intensive nlp tasks
Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, et al. Retrieval- augmented generation for knowledge-intensive nlp tasks. Ad- vances in Neural Information Processing Syst...
2020
-
[41]
Blip- 2: Bootstrapping language-image pre-training with frozen image encoders and large language models
Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi. Blip- 2: Bootstrapping language-image pre-training with frozen image encoders and large language models. In International conference on machine learning, pages 19730–19742. PMLR,
-
[42]
Skip-plan: Procedure planning in instructional videos via condensed action space learning
Zhiheng Li, Wenjia Geng, Muheng Li, Lei Chen, Yansong Tang, Jiwen Lu, and Jie Zhou. Skip-plan: Procedure planning in instructional videos via condensed action space learning. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 10297–10306, 2023. 2
2023
-
[43]
Cohn, and Janet B
Fangru Lin, Emanuele La Malfa, Valentin Hofmann, Elle Michelle Yang, Anthony G. Cohn, and Janet B. Pier- rehumbert. Graph-enhanced large language models in asyn- chronous plan reasoning. In Forty-first International Confer- ence on Machine Learning, 2024. 3
2024
-
[44]
Learning to recognize procedural activities with distant supervision
Xudong Lin, Fabio Petroni, Gedas Bertasius, Marcus Rohrbach, Shih-Fu Chang, and Lorenzo Torresani. Learning to recognize procedural activities with distant supervision. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 13853–13863, 2022. 2, 8
2022
-
[45]
Jointly cross-and self-modal graph attention network for query-based moment localization
Daizong Liu, Xiaoye Qu, Xiao-Yang Liu, Jianfeng Dong, Pan Zhou, and Zichuan Xu. Jointly cross-and self-modal graph attention network for query-based moment localization. In Proceedings of the 28th ACM International Conference on Multimedia, pages 4070–4078, 2020. 2, 3
2020
-
[46]
Improved baselines with visual instruction tuning
Haotian Liu, Chunyuan Li, Yuheng Li, and Yong Jae Lee. Improved baselines with visual instruction tuning. In Pro- ceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 26296–26306, 2024. 1
2024
-
[47]
Visual instruction tuning
Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. Visual instruction tuning. Advances in neural information processing systems, 36, 2024. 4
2024
-
[48]
Attentive moment retrieval in videos
Meng Liu, Xiang Wang, Liqiang Nie, Xiangnan He, Baoquan Chen, and Tat-Seng Chua. Attentive moment retrieval in videos. In The 41st international ACM SIGIR conference on research & development in information retrieval, pages 15–24, 2018. 2
2018
-
[49]
Language models of code are few-shot commonsense learners
Aman Madaan, Shuyan Zhou, Uri Alon, Yiming Yang, and Graham Neubig. Language models of code are few-shot commonsense learners. In Proceedings of the 2022 Confer- ence on Empirical Methods in Natural Language Processing, pages 1384–1403, Abu Dhabi, United Arab Emirates, 2022. A...
2022
-
[50]
What’s cookin’? interpreting cooking videos using text, speech and vision
Jonathan Malmaud, Jonathan Huang, Vivek Rathod, Nicholas Johnston, Andrew Rabinovich, and Kevin Murphy. What’s cookin’? interpreting cooking videos using text, speech and vision. In Proceedings of the 2015 Conference of the North American Chapter of the Association for Computa...
2015
-
[51]
Howto100m: Learning a text-video embedding by watching hundred mil- lion narrated video clips
Antoine Miech, Dimitri Zhukov, Jean-Baptiste Alayrac, Makarand Tapaswi, Ivan Laptev, and Josef Sivic. Howto100m: Learning a text-video embedding by watching hundred mil- lion narrated video clips. In Proceedings of the IEEE/CVF international conference on computer vision , pag...
2019
-
[52]
End-to-end learn- ing of visual representations from uncurated instructional videos
Antoine Miech, Jean-Baptiste Alayrac, Lucas Smaira, Ivan Laptev, Josef Sivic, and Andrew Zisserman. End-to-end learn- ing of visual representations from uncurated instructional videos. In Proceedings of the IEEE/CVF conference on com- puter vision and pattern recognition, page...
2020
-
[53]
Learning and verification of task structure in instructional videos
Medhini Narasimhan, Licheng Yu, Sean Bell, Ning Zhang, and Trevor Darrell. Learning and verification of task structure in instructional videos. arXiv preprint arXiv:2303.13519 ,
-
[54]
SCHEMA: State CHanges MAtter for procedure planning in instructional videos
Yulei Niu, Wenliang Guo, Long Chen, Xudong Lin, and Shih- Fu Chang. SCHEMA: State CHanges MAtter for procedure planning in instructional videos. In The Twelfth International Conference on Learning Representations, 2024. 2
2024
-
[55]
Virtualhome: Sim- ulating household activities via programs
Xavier Puig, Kevin Ra, Marko Boben, Jiaman Li, Tingwu Wang, Sanja Fidler, and Antonio Torralba. Virtualhome: Sim- ulating household activities via programs. In Proceedings of the IEEE conference on computer vision and pattern recogni- tion, pages 8494–8502, 2018. 4
2018
-
[56]
Learning transferable visual models from natural language supervi- sion
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning transferable visual models from natural language supervi- sion. In International conference on machine learning, ...
2021
-
[57]
Grounding action descriptions in videos
Michaela Regneri, Marcus Rohrbach, Dominikus Wetzel, Ste- fan Thater, Bernt Schiele, and Manfred Pinkal. Grounding action descriptions in videos. Transactions of the Association for Computational Linguistics, 1:25–36, 2013. 2
2013
-
[58]
FLAP: Flow-adhering planning with constrained decoding in LLMs
Shamik Roy, Sailik Sengupta, Daniele Bonadiman, Saab Man- sour, and Arshit Gupta. FLAP: Flow-adhering planning with constrained decoding in LLMs. In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Languag...
2024
-
[59]
proScript: Par- tially ordered scripts generation
Keisuke Sakaguchi, Chandra Bhagavatula, Ronan Le Bras, Niket Tandon, Peter Clark, and Yejin Choi. proScript: Par- tially ordered scripts generation. In Findings of the Associ- ation for Computational Linguistics: EMNLP 2021 , pages 2138–2149, Punta Cana, Dominican Republic, 20...
2021
-
[60]
As- sembly101: A large-scale multi-view video dataset for un- derstanding procedural activities
Fadime Sener, Dibyadip Chatterjee, Daniel Shelepov, Kun He, Dipika Singhania, Robert Wang, and Angela Yao. As- sembly101: A large-scale multi-view video dataset for un- derstanding procedural activities. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern ...
2022
-
[61]
Hugginggpt: Solving ai tasks with chatgpt and its friends in hugging face
Yongliang Shen, Kaitao Song, Xu Tan, Dongsheng Li, Weim- ing Lu, and Yueting Zhuang. Hugginggpt: Solving ai tasks with chatgpt and its friends in hugging face. Advances in Neural Information Processing Systems, 36, 2024. 2
2024
-
[62]
Moviechat: From dense token to sparse memory for long video understanding
Enxin Song, Wenhao Chai, Guanhong Wang, Yucheng Zhang, Haoyang Zhou, Feiyang Wu, Haozhe Chi, Xun Guo, Tian Ye, Yanting Zhang, et al. Moviechat: From dense token to sparse memory for long video understanding. In Proceedings of the IEEE/CVF Conference on Computer Vision and Patt...
2024
-
[63]
Mpnet: Masked and permuted pre-training for language understanding
Kaitao Song, Xu Tan, Tao Qin, Jianfeng Lu, and Tie-Yan Liu. Mpnet: Masked and permuted pre-training for language understanding. Advances in neural information processing systems, 33:16857–16867, 2020. 8
2020
-
[64]
Language models can see: Plugging visual controls in text generation
Yixuan Su, Tian Lan, Yahui Liu, Fangyu Liu, Dani Yogatama, Yan Wang, Lingpeng Kong, and Nigel Collier. Language models can see: Plugging visual controls in text generation. arXiv preprint arXiv:2205.02655, 2022. 2
2022 arXiv
-
[65]
PandaGPT: One model to instruction-follow them all
Yixuan Su, Tian Lan, Huayang Li, Jialu Xu, Yan Wang, and Deng Cai. PandaGPT: One model to instruction-follow them all. In Proceedings of the 1st Workshop on Taming Large Language Models: Controllability in the era of Interactive Assistants!, pages 11–23, Prague, Czech Republic...
2023
-
[66]
Plate: Visually-grounded planning with transformers in procedural tasks
Jiankai Sun, De-An Huang, Bo Lu, Yun-Hui Liu, Bolei Zhou, and Animesh Garg. Plate: Visually-grounded planning with transformers in procedural tasks. IEEE Robotics and Automa- tion Letters, 7(2):4924–4930, 2022. 2
2022
-
[67]
Coin: A large-scale dataset for comprehensive instructional video analysis
Yansong Tang, Dajun Ding, Yongming Rao, Yu Zheng, Danyang Zhang, Lili Zhao, Jiwen Lu, and Jie Zhou. Coin: A large-scale dataset for comprehensive instructional video analysis. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 1207–1216,
-
[68]
On the planning abilities of large language models-a critical investigation
Karthik Valmeekam, Matthew Marquez, Sarath Sreedharan, and Subbarao Kambhampati. On the planning abilities of large language models-a critical investigation. Advances in Neural Information Processing Systems , 36:75993–76005,
-
[69]
Gomez, Lukasz Kaiser, and Illia Polosukhin
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkor- eit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. Attention is all you need. Advances in Neural Information Processing Systems, 2017. 2
2017
-
[70]
Event-guided procedure planning from in- structional videos with text supervision
An-Lan Wang, Kun-Yu Lin, Jia-Run Du, Jingke Meng, and Wei-Shi Zheng. Event-guided procedure planning from in- structional videos with text supervision. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 13565–13575, 2023. 2
2023
-
[71]
Pdpp: Projected diffusion for procedure planning in instructional videos
Hanlin Wang, Yilu Wu, Sheng Guo, and Limin Wang. Pdpp: Projected diffusion for procedure planning in instructional videos. In Proceedings of the IEEE/CVF Conference on Com- puter Vision and Pattern Recognition, pages 14836–14845,
-
[72]
Temporal segment networks: Towards good practices for deep action recognition
Limin Wang, Yuanjun Xiong, Zhe Wang, Yu Qiao, Dahua Lin, Xiaoou Tang, and Luc Van Gool. Temporal segment networks: Towards good practices for deep action recognition. In European conference on computer vision , pages 20–36. Springer, 2016. 8
2016
-
[73]
Visual chatgpt: Talking, draw- ing and editing with visual foundation models
Chenfei Wu, Shengming Yin, Weizhen Qi, Xiaodong Wang, Zecheng Tang, and Nan Duan. Visual chatgpt: Talking, draw- ing and editing with visual foundation models. arXiv preprint arXiv:2303.04671, 2023. 2
2023 arXiv
-
[74]
NExt-GPT: Any-to-any multimodal LLM
Shengqiong Wu, Hao Fei, Leigang Qu, Wei Ji, and Tat-Seng Chua. NExt-GPT: Any-to-any multimodal LLM. In Forty- first International Conference on Machine Learning , 2024. 1
2024
-
[75]
Rethinking spatiotemporal feature learning: Speed-accuracy trade-offs in video classification
Saining Xie, Chen Sun, Jonathan Huang, Zhuowen Tu, and Kevin Murphy. Rethinking spatiotemporal feature learning: Speed-accuracy trade-offs in video classification. In Proceed- ings of the European conference on computer vision (ECCV), pages 305–321, 2018. 8
2018
-
[76]
Translating natural language to plan- ning goals with large-language models
Yaqi Xie, Chen Yu, Tongyao Zhu, Jinbin Bai, Ze Gong, and Harold Soh. Translating natural language to plan- ning goals with large-language models. arXiv preprint arXiv:2302.05128, 2023. 3
2023 arXiv
-
[77]
Multilevel language and vision integration for text-to-clip retrieval
Huijuan Xu, Kun He, Bryan A Plummer, Leonid Sigal, Stan Sclaroff, and Kate Saenko. Multilevel language and vision integration for text-to-clip retrieval. In Proceedings of the AAAI Conference on Artificial Intelligence, pages 9062–9069,
-
[78]
VideoCLIP: Contrastive pre- training for zero-shot video-text understanding
Hu Xu, Gargi Ghosh, Po-Yao Huang, Dmytro Okhonko, Ar- men Aghajanyan, Florian Metze, Luke Zettlemoyer, and Christoph Feichtenhofer. VideoCLIP: Contrastive pre- training for zero-shot video-text understanding. In Proceed- ings of the 2021 Conference on Empirical Methods in Natu...
2021
-
[79]
Retrieval- augmented generation with knowledge graphs for customer service question answering
Zhentao Xu, Mark Jerome Cruz, Matthew Guevara, Tie Wang, Manasi Deshpande, Xiaofeng Wang, and Zheng Li. Retrieval- augmented generation with knowledge graphs for customer service question answering. In Proceedings of the 47th In- ternational ACM SIGIR Conference on Research an...
2024
-
[80]
Activitynet-qa: A dataset for understanding complex web videos via question answering
Zhou Yu, Dejing Xu, Jun Yu, Ting Yu, Zhou Zhao, Yueting Zhuang, and Dacheng Tao. Activitynet-qa: A dataset for understanding complex web videos via question answering. In Proceedings of the AAAI Conference on Artificial Intelligence, pages 9127–9134, 2019. 3, 5
2019
-
[81]
Distilling script knowledge from large language models for constrained language planning
Siyu Yuan, Jiangjie Chen, Ziquan Fu, Xuyang Ge, Soham Shah, Charles Jankowski, Yanghua Xiao, and Deqing Yang. Distilling script knowledge from large language models for constrained language planning. In Proceedings of the 61st Annual Meeting of the Association for Computationa...
2023
-
[82]
Man: Moment alignment network for natural language moment retrieval via iterative graph adjustment
Da Zhang, Xiyang Dai, Xin Wang, Yuan-Fang Wang, and Larry S Davis. Man: Moment alignment network for natural language moment retrieval via iterative graph adjustment. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 1247–1257, 2019. 2
2019
-
[83]
SpeechGPT: Empowering large language models with intrinsic cross-modal conversa- tional abilities
Dong Zhang, Shimin Li, Xin Zhang, Jun Zhan, Pengyu Wang, Yaqian Zhou, and Xipeng Qiu. SpeechGPT: Empowering large language models with intrinsic cross-modal conversa- tional abilities. In Findings of the Association for Computa- tional Linguistics: EMNLP 2023, pages 15757–1577...
2023
-
[84]
Video-LLaMA: An instruction-tuned audio-visual language model for video un- derstanding
Hang Zhang, Xin Li, and Lidong Bing. Video-LLaMA: An instruction-tuned audio-visual language model for video un- derstanding. In Proceedings of the 2023 Conference on Em- pirical Methods in Natural Language Processing: System Demonstrations, pages 543–553, Singapore, 2023. Ass...
2023
-
[85]
Temporal sentence grounding in videos: A survey and fu- ture directions
Hao Zhang, Aixin Sun, Wei Jing, and Joey Tianyi Zhou. Temporal sentence grounding in videos: A survey and fu- ture directions. IEEE Transactions on Pattern Analysis and Machine Intelligence, 45(8):10443–10465, 2023. 2
2023
-
[86]
P3iv: Probabilistic procedure planning from instructional videos with weak su- pervision
He Zhao, Isma Hadji, Nikita Dvornik, Konstantinos G Derpa- nis, Richard P Wildes, and Allan D Jepson. P3iv: Probabilistic procedure planning from instructional videos with weak su- pervision. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition,...
-
[87]
Learning procedure-aware video repre- sentation from instructional videos and their narrations
Yiwu Zhong, Licheng Yu, Yang Bai, Shangwen Li, Xueting Yan, and Yin Li. Learning procedure-aware video repre- sentation from instructional videos and their narrations. In Proceedings of the IEEE/CVF Conference on Computer Vi- sion and Pattern Recognition, pages 14825–14835, 2023. 2, 8
2023
-
[88]
Procedure-aware pretraining for instructional video understanding
Honglu Zhou, Roberto Martín-Martín, Mubbasir Kapadia, Silvio Savarese, and Juan Carlos Niebles. Procedure-aware pretraining for instructional video understanding. In Proceed- ings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 10727–10738, 2023. 2, 3, 8
2023
-
[89]
Towards auto- matic learning of procedures from web instructional videos
Luowei Zhou, Chenliang Xu, and Jason Corso. Towards auto- matic learning of procedures from web instructional videos. In Proceedings of the AAAI Conference on Artificial Intelli- gence, 2018. 1
2018
-
[90]
MiniGPT-4: Enhancing vision-language understanding with advanced large language models
Deyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li, and Mo- hamed Elhoseiny. MiniGPT-4: Enhancing vision-language understanding with advanced large language models. In The Twelfth International Conference on Learning Representa- tions, 2024. 2
2024
-
[91]
Cross- task weakly supervised learning from instructional videos
Dimitri Zhukov, Jean-Baptiste Alayrac, Ramazan Gokberk Cinbis, David Fouhey, Ivan Laptev, and Josef Sivic. Cross- task weakly supervised learning from instructional videos. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 3537–3545, 2...
2019
Reviewed August 10, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.