Fine-tuning video-language models on educational content helps only some models, while a transcript-only baseline produces more relevant and answerable questions.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CV 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Enhancing the Learning Experience: Using Vision-Language Models to Generate Questions for Educational Videos
Fine-tuning video-language models on educational content helps only some models, while a transcript-only baseline produces more relevant and answerable questions.