SimTube generates pre-publication video comments from multimodal video understanding and sampled user personas, and its evaluations suggest these simulated comments are often rated as helpful as real ones.
Video In Sentences Out
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
We present a system that produces sentential descriptions of video: who did what to whom, and where and how they did it. Action class is rendered as a verb, participant objects as noun phrases, properties of those objects as adjectival modifiers in those noun phrases,spatial relations between those participants as prepositional phrases, and characteristics of the event as prepositional-phrase adjuncts and adverbial modifiers. Extracting the information needed to render these linguistic entities requires an approach to event recognition that recovers object tracks, the track-to-role assignments, and changing body posture.
fields
cs.HC 1years
2024 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
SimTube: Generating Simulated Video Comments through Multimodal AI and User Personas
SimTube generates pre-publication video comments from multimodal video understanding and sampled user personas, and its evaluations suggest these simulated comments are often rated as helpful as real ones.