VG-TVP enriches LLM-generated text plans with captions from instructional videos and produces a short video per step; human raters prefer it over text-only baselines.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CV 1years
2024 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
VG-TVP: Multimodal Procedural Planning via Visually Grounded Text-Video Prompting
VG-TVP enriches LLM-generated text plans with captions from instructional videos and produces a short video per step; human raters prefer it over text-only baselines.