REVIEW 4 major objections 4 minor 48 references
VG-TVP: Multimodal Procedural Planning via Visually Grounded Text-Video Prompting
T0 review · 4 major / 4 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read The paper claims that grounding LLM plans in video captions and regenerating them as video yields multimodal how-to instructions preferred over text-only baselines.
desk verdict Useful dataset and pipeline, but the headline 'outperforms' claim lacks statistical support. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing component is the Fusion of Captioning (FoC) step, which takes video captions from several instructional videos of the same task and uses an LLM to reorder, merge, and align them into a single ranked caption list that matches the procedural steps. Around it sit two bridges: Video-to-Text (V2T-B), which turns instructional videos into captions, and Text-to-Video (T2V-B), which turns the revised visual descriptions into short clips. FoC is what converts scattered, timestamped captions into a visually grounded text plan. The paper claims that without FoC, the generated videos lose plan accuracy and visual informativeness.
What would settle it
Show the model a task whose instructional videos contain an object the captioner consistently mislabels (the paper's own example is an apple called an orange), and check whether the generated text and video plans misname or misorder that step; if the plans stay correct despite the bad caption, the claimed grounding is not doing the work.
Extended reading notes
Core claim
The central discovery is that video knowledge enters an LLM most usefully as fused caption text rather than as raw pixels, and that the key difficulty is temporal alignment between what the video captions show and what the text plan says. The proposed VG-TVP pipeline first asks an LLM for a vanilla text plan, then captions multiple instructional videos of the same task, uses a Fusion of Captioning step to reorder and merge those captions into a coherent procedure, and then prompts the LLM to rewrite the original plan as text/context/visual triplets. The visual sentence in each triplet is fed to a text-to-video model, producing a short video per step. Human preference experiments on the Daily-PP dataset show the resulting multimodal plans winning against unimodal baselines on most comparisons, with the largest margins in visual informativeness.
Load-bearing premise
The load-bearing premise is that the video captioning model describes the real objects and actions accurately enough for the LLM to trust; the appendix notes that when VLog mislabels an object, FoC can propagate the mistake and misorder the plan.
Editorial extensions
If this is right
- Procedural planning can be improved without fine-tuning or training new models: composing a zero-shot LLM, a video captioning model, and a text-to-video model is enough.
- For tasks that have no instructional videos, the model can still produce multimodal plans by borrowing and fusing captions from related seen tasks.
- Generated step videos are short and human-centred, which the paper argues can lower a learner's cognitive load compared with watching long instructional videos.
- Human preference, not BLEU or METEOR, is the appropriate yardstick for daily-life procedural plans, because such tasks have no single ground-truth sequence.
Reading between the lines
- The quality ceiling of the whole pipeline is set by the captioning step, so a better video-understanding model should directly raise both text and video plan quality without any redesign.
- The same caption-fusion recipe could be transferred to other instruction formats, such as diagrams or audio narration, whenever a model can turn those modalities into text.
- A testable extension would measure whether users actually complete tasks faster or more successfully with text-video plans than with text-only plans, since the paper evaluates preference, not performance.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes VG-TVP, a zero-shot multimodal procedural planning framework that combines an LLM-generated vanilla text plan, VLog video captions fused through a Fusion of Captioning (FoC) module, and ModelScope text-to-video generation to produce paired textual and video step plans. The authors introduce a new dataset, Daily-PP, with 50 seen and 15 unseen daily-life tasks, and evaluate their method with human win/tie/lose comparisons across textual informativeness, visual informativeness, temporal coherence, and plan accuracy, supplemented by BLEU/METEOR scores and an LLM-based evaluation protocol. The central claim is that VG-TVP outperforms unimodal baselines on Daily-PP.
Significance. If the human-preference results held up statistically, the paper would offer a practical zero-shot recipe for multimodal procedural planning, and the released Daily-PP dataset and code would be reusable assets for the community. The FoC idea of aligning and fusing unordered video captions with vanilla text plans is plausible, and the V2T-B/T2V-B pipeline is a reasonable way to couple video understanding with text-to-video generation. However, the evaluation as presented does not yet establish the central claim: the human evaluation tables lack uncertainty quantification, several comparisons favor the baselines, and the acknowledged caption-error/misordering limitation in FoC is never quantified. The contribution is potentially useful but not yet demonstrated at the bar of the stated claim.
major comments (4)
- [Human Evaluation Metric; Tables 1 and 2] The paper's headline claim rests entirely on the win/tie/lose percentages in Tables 1 and 2, but the manuscript reports no significance test, confidence interval, per-cell rating count, or inter-rater agreement measure. With 28 subjects and 50 seen/15 unseen tasks under the stated 'no subject sees the same task twice' protocol, a margin such as 40% vs. 36% could easily be sampling noise; the authors must state how many ratings contribute to each cell and report, for example, exact binomial confidence intervals and a test that the VG-TVP win proportion exceeds the baseline win proportion (or a tie-aware model). Without this, 'outperforms' is not statistically established.
- [Abstract; Results, Tables 1 and 2] The abstract's claim that VG-TVP 'outperforms unimodal baselines' overstates what the tables show. In Table 2 (unseen tasks), VG-TVP loses to GPT-3.5 on temporal coherence (20.00 win vs. 40.00 lose) and on plan accuracy (26.67 vs. 40.00), and ties on textual informativeness (80.00). In Table 1, it also loses textual informativeness to Llama2-13B-q4 (40.00 win vs. 36.00 lose). The claim should be scoped to the aspects and baselines where the data support it, or be backed by the statistical analysis requested above.
- [Appendix, The Impact of V2T-B and FoC] The appendix states that FoC 'may also misorder information' when VLog mislabels objects, giving the example of an apple mislabeled as an orange or lemon based on its color. This directly affects temporal coherence and plan accuracy, which are exactly the aspects where VG-TVP already loses to GPT-3.5 on unseen tasks. Because FoC is the mechanism that grounds text plans in visual content, the paper needs at least a quantitative check of caption and step-order accuracy (e.g., agreement with human-annotated step order on a subset) or an ablation with an alternative captioner; otherwise the visual-grounding link in the pipeline remains unverified.
- [Experiments, Human Evaluation Protocol] The protocol description says 'No subject was shown the same task twice,' but it is not explained how the 28 subjects were assigned across the 50 seen and 15 unseen tasks and across the eight baselines in Table 1. As a result, the reader cannot tell whether the comparisons in each cell are based on a single subject, a few subjects, or many, and whether the percentages are comparable across rows. The authors should describe the assignment procedure and report the number of comparisons behind each percentage.
minor comments (4)
- [Ablations] The capitalization of FoC is inconsistent: the Ablations paragraph uses 'FOC' instead of 'FoC'; please unify the terminology.
- [Table 3] Table 3 is hard to interpret because the header does not clearly separate the BLEU and METEOR columns for the baseline and VG-TVP conditions, and the BLEU values are near zero; please clarify the layout and consider reporting confidence intervals or omitting the metric if it is not intended to support the main claim.
- [Tables 4 and 5 and surrounding text] The appendix text says VG-TVP 'consistently outperforms baselines and TIP' on CLIP-based MSS, but Tables 4 and 5 do not include a TIP row and report no error bars or significance information; please add the TIP comparison and uncertainty estimates or temper the claim.
- [Figures 1, 5, 18, 24] Several task names appear as concatenated strings such as 'How toBakeKofta/MeatballsandPotatoes?' and 'How toCookSpaghetti'; please use spaced and readable task names in figures and text.
Circularity Check
No significant circularity: VG-TVP is an empirical pipeline whose central claim rests on external human evaluation, not on a derivation from its inputs.
full rationale
This paper is an empirical systems comparison rather than a mathematical derivation. There are no fitted parameters, no equations that reduce a predicted quantity to an input, and no uniqueness theorem invoked from prior work. Each stage of the pipeline receives distinct inputs: the LLM generates a vanilla text plan from the task description, VLog produces captions from instructional videos, FoC fuses those captions, and the alignment prompt combines the vanilla plan with FoC to produce the final text and video plans. The central claim, that VG-TVP outperforms unimodal baselines on Daily-PP, is supported by human preference judgments and additional automatic metrics comparing generated outputs with reference plans. None of these evaluation quantities is an input to the generation pipeline by construction. The only authorship overlap with a cited component is VLog, whose author is also a co-author of this paper, but VLog is an off-the-shelf external video captioning tool used as a bridge rather than as the source of the claimed result; the claim is not that VLog works but that the full VG-TVP pipeline is preferred by human raters. The appendix even acknowledges VLog's captioning errors and states that FoC may misorder information, which is a limitation of the input component, not a circular justification. Concerns about missing significance tests or the post hoc explanation of GPT-3.5's lower unseen-task performance relate to statistical validity, not to circularity, so they do not affect this score.
Assumptions & free parameters
assumptions (5)
- domain assumption LLMs can produce coherent zero-shot procedural text plans and can revise them from captions.
- domain assumption VLog captions are accurate and complete enough to support procedural planning.
- domain assumption ModelScope text-to-video generates clips that follow the visual prompt semantics.
- domain assumption Human judgments from 28 subjects are a valid measure of plan quality.
- domain assumption Daily-PP splitting into seen and unseen tasks, with captions from related seen tasks, provides sufficient visual knowledge for unseen planning.
Cite this review
Pith. "Pith review of VG-TVP: Multimodal Procedural Planning via Visually Grounded Text-Video Prompting." pith.science (2026). https://pith.science/paper/UKCQ2WEC
@misc{pith2026241211621,
author = {Pith},
title = {Pith review of: VG-TVP: Multimodal Procedural Planning via Visually Grounded Text-Video Prompting},
year = {2026},
howpublished = {\url{https://pith.science/paper/UKCQ2WEC}},
note = {Machine review of arXiv:2412.11621}
}
read the original abstract
Large Language Model (LLM)-based agents have shown promise in procedural tasks, but the potential of multimodal instructions augmented by texts and videos to assist users remains under-explored. To address this gap, we propose the Visually Grounded Text-Video Prompting (VG-TVP) method which is a novel LLM-empowered Multimodal Procedural Planning (MPP) framework. It generates cohesive text and video procedural plans given a specified high-level objective. The main challenges are achieving textual and visual informativeness, temporal coherence, and accuracy in procedural plans. VG-TVP leverages the zero-shot reasoning capability of LLMs, the video-to-text generation ability of the video captioning models, and the text-to-video generation ability of diffusion models. VG-TVP improves the interaction between modalities by proposing a novel Fusion of Captioning (FoC) method and using Text-to-Video Bridge (T2V-B) and Video-to-Text Bridge (V2T-B). They allow LLMs to guide the generation of visually-grounded text plans and textual-grounded video plans. To address the scarcity of datasets suitable for MPP, we have curated a new dataset called Daily-Life Task Procedural Plans (Daily-PP). We conduct comprehensive experiments and benchmarks to evaluate human preferences (regarding textual and visual informativeness, temporal coherence, and plan accuracy). Our VG-TVP method outperforms unimodal baselines on the Daily-PP dataset.
Figures
Figures from the paper (18 more)
Reference graph
Works this paper leans on
-
[1]
Alayrac, J.; Bojanowski, P.; Agrawal, N.; Sivic, J.; Laptev, I.; and Lacoste - Julien, S. 2016. Unsupervised Learning from Narrated Instruction Videos. In Conference on Computer Vision and Pattern Recognition, CVPR , Las Vegas, NV, USA , 4575--4583. IEEE
work page 2016
-
[2]
Banerjee, S.; and Lavie, A. 2005. METEOR : An Automatic Metric for MT Evaluation with Improved Correlation with Human Judgments. In Proceedings of the ACL Workshop on Intrinsic and Extrinsic Evaluation Measures for Machine Translation and/or Summarization , 65--72. Association for Computational Linguistics
2005
-
[3]
Blattmann, A.; Rombach, R.; Ling, H.; Dockhorn, T.; Kim, S. W.; Fidler, S.; and Kreis, K. 2023. Align Your Latents: High-Resolution Video Synthesis with Latent Diffusion Models. In Conference on Computer Vision and Pattern Recognition, Vancouver, BC, Canada, 22563--22575. IEEE
work page 2023
-
[4]
Brown, T. B.; Mann, B.; Ryder, N.; Subbiah, M.; Kaplan, J.; Dhariwal, P.; Neelakantan, A.; Shyam, P.; Sastry, G.; Askell, A.; Agarwal, S.; Herbert - Voss, A.; Krueger, G.; Henighan, T.; Child, R.; Ramesh, A.; Ziegler, D. M.; Wu, J.; Winter, C.; Hesse, C.; Chen, M.; ...; and Amodei, D. 2020. Language Models are Few-Shot Learners. In Annual Conference on Ne...
work page 2020
-
[5]
Chang, C.; Huang, D.; Xu, D.; Adeli, E.; Fei - Fei, L.; and Niebles, J. C. 2020. Procedure Planning in Instructional Videos. In 16th European Conference Computer Vision, ECCV , Glasgow,UK , volume 12356 of Lecture Notes in Computer Science, 334--350. Springer
work page 2020
-
[6]
Chang, E. Y. 2023. Prompting Large Language Models With the Socratic Method. arXiv:2303.08769
arXiv 2023
-
[7]
Du, X.; Rush, A. M.; and Cardie, C. 2021. GRIT: Generative Role-filler Transformers for Document-level Event Entity Extraction. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume, EACL , 634--644. Association for Computational Linguistics
work page 2021
-
[8]
Dvornik, N.; Hadji, I.; Zhang, R.; Derpanis, K. G.; Wildes, R. P.; and Jepson, A. D. 2023. StepFormer: Self-Supervised Step Discovery and Localization in Instructional Videos. In Conference on Computer Vision and Pattern Recognition, CVPR , Vancouver, BC, Canada , 18952--18961. IEEE
work page 2023
Show all 48 references
-
[9]
Elhamifar, E.; and Naing, Z. 2019. Unsupervised Procedure Learning via Joint Dynamic Summarization. In 2019 International Conference on Computer Vision, ICCV , Seoul, Korea (South) , 6340--6349. IEEE
2019
-
[10]
Fang, F.; Liu, Y.; Koksal, A.; Xu, Q.; and Lim, J. 2023. Masked Diffusion with Task-awareness for Procedure Planning in Instructional Videos. CoRR, abs/2309.07409
2023 arXiv
-
[11]
Huang, W.; Abbeel, P.; Pathak, D.; and Mordatch, I. 2022. Language Models as Zero-Shot Planners: Extracting Actionable Knowledge for Embodied Agents. In International Conference on Machine Learning, ICML , Baltimore, Maryland, USA , volume 162, 9118--9147. PMLR
2022
-
[12]
F.; Song, C.; Chen, J.; Gao, D.; Lei, W.; Xu, Q.; Lim, J.; and Shou, M
Ilaslan, M. F.; Song, C.; Chen, J.; Gao, D.; Lei, W.; Xu, Q.; Lim, J.; and Shou, M. 2023. G aze VQA : A Video Question Answering Dataset for Multiview Eye-Gaze Task-Oriented Collaborations. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processi...
2023
-
[13]
Q.; Sablayrolles, A.; Mensch, A.; Bamford, C.; Chaplot, D
Jiang, A. Q.; Sablayrolles, A.; Mensch, A.; Bamford, C.; Chaplot, D. S.; de Las Casas, D.; Bressand, F.; Lengyel, G.; Lample, G.; Saulnier, L.; Lavaud, L. R.; Lachaux, M.; Stock, P.; Scao, T. L.; Lavril, T.; Wang, T.; Lacroix, T.; and Sayed, W. E. 2023. Mistral 7B. CoRR, abs/2...
2023 arXiv
-
[14]
Khachatryan, L.; Movsisyan, A.; Tadevosyan, V.; Henschel, R.; Wang, Z.; Navasardyan, S.; and Shi, H. 2023. Text2Video-Zero: Text-to-Image Diffusion Models are Zero-Shot Video Generators. In International Conference on Computer Vision, Paris, France, 15908--15918. IEEE
2023
-
[15]
S.; Reid, M.; Matsuo, Y.; and Iwasawa, Y
Kojima, T.; Gu, S. S.; Reid, M.; Matsuo, Y.; and Iwasawa, Y. 2022. Large Language Models are Zero-Shot Reasoners. In Annual Conference on Neural Information Processing Systems, NeurIPS, New Orleans, USA. Curran Associates Inc
2022
-
[16]
B.; and Serre, T
Kuehne, H.; Arslan, A. B.; and Serre, T. 2014. The Language of Actions: Recovering the Syntax and Semantics of Goal-Directed Human Activities. In Conference on Computer Vision and Pattern Recognition, CVPR , Columbus, OH, USA , 780--787. IEEE
2014
-
[17]
Li, J.; Li, D.; Savarese, S.; and Hoi, S. C. H. 2023. BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models. In International Conference on Machine Learning, ICML , Honolulu, Hawaii, USA , volume 202 of Proceedings of Machine Le...
2023
-
[18]
Lian, L.; Shi, B.; Yala, A.; Darrell, T.; and Li, B. 2023. LLM-grounded Video Diffusion Models. CoRR, abs/2309.17444
2023 arXiv
-
[19]
Lin, H.; Zala, A.; Cho, J.; and Bansal, M. 2023. VideoDirectorGPT: Consistent Multi-scene Video Generation via LLM-Guided Planning. CoRR, abs/2309.15091
2023 arXiv
-
[20]
Q.; and Lei, S
Lin, K. Q.; and Lei, S. W. 2023. VLog: Video as a Long Document. https://github.com/showlab/VLog. Accessed:2024-12-15
2023
-
[21]
E.; Eckstein, M
Lu, Y.; Feng, W.; Zhu, W.; Xu, W.; Wang, X. E.; Eckstein, M. P.; and Wang, W. Y. 2023 a . Neuro-Symbolic Procedural Planning with Commonsense Prompting. In The Eleventh International Conference on Learning Representations, ICLR , Kigali, Rwanda
2023
-
[22]
E.; and Wang, W
Lu, Y.; Lu, P.; Chen, Z.; Zhu, W.; Wang, X. E.; and Wang, W. Y. 2023 b . Multimodal Procedural Planning via Dual Text-Image Prompting. CoRR, abs/2305.01795
2023 arXiv
-
[23]
Miech, A.; Zhukov, D.; Alayrac, J.; Tapaswi, M.; Laptev, I.; and Sivic, J. 2019. HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video Clips. In International Conference on Computer Vision, ICCV , Seoul, Korea (South) , 2630--2640. IEEE
2019
-
[24]
Niu, Y.; Guo, W.; Chen, L.; Lin, X.; and Chang, S.-F. 2024. SCHEMA : State Changes Matter for Procedure Planning in Instructional Videos. In The Twelfth International Conference on Learning Representations
2024
-
[25]
OpenAI. 2023. GPT-4 Technical Report. CoRR, abs/2303.08774
2023 arXiv
-
[26]
Papineni, K.; Roukos, S.; Ward, T.; and Zhu, W.-J. 2002. BLEU: a method for automatic evaluation of machine translation. In Proceedings of the 40th Annual Meeting on Association for Computational Linguistics, ACL '02, 311–318. USA: Association for Computational Linguistics
2002
-
[27]
W.; Xu, T.; Brockman, G.; McLeavey, C.; and Sutskever, I
Radford, A.; Kim, J. W.; Xu, T.; Brockman, G.; McLeavey, C.; and Sutskever, I. 2023. Robust Speech Recognition via Large-Scale Weak Supervision. In International Conference on Machine Learning, ICML , Honolulu, Hawaii, USA , volume 202 of Proceedings of Machine Learning Resear...
2023
-
[28]
Ramesh, A.; Pavlov, M.; Goh, G.; Gray, S.; Voss, C.; Radford, A.; Chen, M.; and Sutskever, I. 2021. Zero-Shot Text-to-Image Generation. In Proceedings of the 38th International Conference on Machine Learning, ICML , volume 139, 8821--8831. PMLR
2021
-
[29]
Shen, Y.; Wang, L.; and Elhamifar, E. 2021. Learning To Segment Actions From Visual and Language Instructions via Differentiable Weak Sequence Alignment. In Conference on Computer Vision and Pattern Recognition, CVPR , 10156--10165. IEEE
2021
-
[30]
Singer, U.; Polyak, A.; Hayes, T.; Yin, X.; An, J.; Zhang, S.; Hu, Q.; Yang, H.; Ashual, O.; Gafni, O.; Parikh, D.; Gupta, S.; and Taigman, Y. 2023. Make-A-Video: Text-to-Video Generation without Text-Video Data. In The Eleventh International Conference on Learning Representat...
2023
-
[31]
H.; Sadler, B
Song, C. H.; Sadler, B. M.; Wu, J.; Chao, W.; Washington, C.; and Su, Y. 2023. LLM-Planner: Few-Shot Grounded Planning for Embodied Agents with Large Language Models. In International Conference on Computer Vision, ICCV , Paris, France , 2986--2997. IEEE
2023
-
[32]
Soucek, T.; Damen, D.; Wray, M.; Laptev, I.; and Sivic, J. 2024. GenHowTo: Learning to Generate Actions and State Transformations from Instructional Videos. In Conference on Computer Vision and Pattern Recognition, CVPR , Seattle, WA, USA , 6561--6571. IEEE
2024
-
[33]
Tang, Y.; Ding, D.; Rao, Y.; Zheng, Y.; Zhang, D.; Zhao, L.; Lu, J.; and Zhou, J. 2019. COIN: A Large-Scale Dataset for Comprehensive Instructional Video Analysis. In Conference on Computer Vision and Pattern Recognition, CVPR , Long Beach, CA, USA , 1207--1216. IEEE
2019
-
[34]
Touvron, H.; Martin, L.; Stone, K.; Albert, P.; Almahairi, A.; Babaei, Y.; Bashlykov, N.; Batra, S.; Bhargava, P.; Bhosale, S.; Bikel, D.; Blecher, L.; Canton - Ferrer, C.; Chen, M.; Cucurull, G.; Esiobu, D.; Fernandes, J.; Fu, J.; Fu, W.; Fuller, B.; Gao, C.; ...; and Scialom...
2023 arXiv
-
[35]
Wang, H.; Wu, Y.; Guo, S.; and Wang, L. 2023 a . PDPP: Projected Diffusion for Procedure Planning in Instructional Videos. In Conference on Computer Vision and Pattern Recognition, CVPR , Vancouver, BC, Canada , 14836--14845. IEEE
2023
-
[36]
Wang, J.; Yuan, H.; Chen, D.; Zhang, Y.; Wang, X.; and Zhang, S. 2023 b . ModelScope Text-to-Video Technical Report. CoRR, abs/2308.06571
2023 arXiv
-
[37]
Z.; Ge, Y.; Wang, X.; Lei, S
Wu, J. Z.; Ge, Y.; Wang, X.; Lei, S. W.; Gu, Y.; Shi, Y.; Hsu, W.; Shan, Y.; Qie, X.; and Shou, M. Z. 2023. Tune-A-Video: One-Shot Tuning of Image Diffusion Models for Text-to-Video Generation. In International Conference on Computer Vision, ICCV , Paris, France , 7589--7599. IEEE
2023
-
[38]
H.; Miech, A.; Pont - Tuset, J.; Laptev, I.; Sivic, J.; and Schmid, C
Yang, A.; Nagrani, A.; Seo, P. H.; Miech, A.; Pont - Tuset, J.; Laptev, I.; Sivic, J.; and Schmid, C. 2023. Vid2Seq: Large-Scale Pretraining of a Visual Language Model for Dense Video Captioning. In Conference on Computer Vision and Pattern Recognition, CVPR , Vancouver, BC, C...
2023
-
[39]
Yang, Y.; Yao, W.; Zhang, H.; Wang, X.; Yu, D.; and Chen, J. 2022. Z-LaVI: Zero-Shot Language Solver Fueled by Visual Imagination. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, EMNLP , Abu Dhabi, UAE , 1186--1203. Association for Co...
2022
-
[40]
B.; Versari, L.; Sohn, K.; Minnen, D.; Cheng, Y.; Gupta, A.; Gu, X.; Hauptmann, A
Yu, L.; Lezama, J.; Gundavarapu, N. B.; Versari, L.; Sohn, K.; Minnen, D.; Cheng, Y.; Gupta, A.; Gu, X.; Hauptmann, A. G.; Gong, B.; Yang, M.; Essa, I.; Ross, D. A.; and Jiang, L. 2023. Language Model Beats Diffusion - Tokenizer is Key to Visual Generation. CoRR, abs/2310.05737
2023 arXiv
-
[41]
G.; Wildes, R
Zhao, H.; Hadji, I.; Dvornik, N.; Derpanis, K. G.; Wildes, R. P.; and Jepson, A. D. 2022. P\( ^ 3 \)IV: Probabilistic Procedure Planning from Instructional Videos with Weak Supervision. In Conference on Computer Vision and Pattern Recognition, CVPR , New Orleans, USA , 2928--2...
2022
-
[42]
Zhou, D.; Wang, W.; Yan, H.; Lv, W.; Zhu, Y.; and Feng, J. 2022. MagicVideo: Efficient Video Generation With Latent Diffusion Models. CoRR, abs/2211.11018
2022 arXiv
-
[43]
Zhou, H.; Mart \' n - Mart \' n, R.; Kapadia, M.; Savarese, S.; and Niebles, J. C. 2023. Procedure-Aware Pretraining for Instructional Video Understanding. In Conference on Computer Vision and Pattern Recognition, CVPR , Vancouver, BC, Canada , 10727--10738. IEEE
2023
-
[44]
Zhou, L.; Xu, C.; and Corso, J. J. 2018. Towards Automatic Learning of Procedures From Web Instructional Videos. In Proceedings of the Thirty-Second AAAI Conference on Artificial Intelligence, (AAAI-18), New Orleans, Louisiana, USA , 7590--7598. AAAI Press
2018
-
[45]
Zhu, L.; and Yang, Y. 2020. ActBERT: Learning Global-Local Video-Text Representations. In Conference on Computer Vision and Pattern Recognition, CVPR , Seattle, WA, USA , 8743--8752. IEEE
2020
-
[46]
G.; Fouhey, D
Zhukov, D.; Alayrac, J.; Cinbis, R. G.; Fouhey, D. F.; Laptev, I.; and Sivic, J. 2019. Cross-Task Weakly Supervised Learning From Instructional Videos. In Conference on Computer Vision and Pattern Recognition, CVPR , Long Beach, CA, USA , 3537--3545. IEEE
2019
-
[47]
, " * write output.state after.block = add.period write newline
ENTRY address archivePrefix author booktitle chapter edition editor eid eprint howpublished institution isbn journal key month note number organization pages publisher school series title type volume year label extra.label sort.label short.list INTEGERS output.state before.all...
-
[48]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 gl...
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.