V-CASS combines a vision model, a knowledge-infused language model, and expressive text-to-speech to generate video commentary speech aligned with on-screen mood, and users prefer it over neutral speech.
Does speech rate influence intertemporal decisions? an experimental investigation,
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
background 1
citation-polarity summary
fields
cs.HC 1years
2025 1verdicts
CONDITIONAL 1roles
background 1polarities
background 1representative citing papers
citing papers explorer
-
V-CASS: Vision-context-aware Expressive Speech Synthesis for Enhancing User Understanding of Videos
V-CASS combines a vision model, a knowledge-infused language model, and expressive text-to-speech to generate video commentary speech aligned with on-screen mood, and users prefer it over neutral speech.