REVIEW 1 cited by
Video In Sentences Out
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
We present a system that produces sentential descriptions of video: who did what to whom, and where and how they did it. Action class is rendered as a verb, participant objects as noun phrases, properties of those objects as adjectival modifiers in those noun phrases,spatial relations between those participants as prepositional phrases, and characteristics of the event as prepositional-phrase adjuncts and adverbial modifiers. Extracting the information needed to render these linguistic entities requires an approach to event recognition that recovers object tracks, the track-to-role assignments, and changing body posture.
Forward citations
Cited by 1 Pith paper
-
SimTube: Generating Simulated Video Comments through Multimodal AI and User Personas
SimTube generates pre-publication video comments from multimodal video understanding and sampled user personas, and its evaluations suggest these simulated comments are often rated as helpful as real ones.
Discussion (0). Continue with ORCID to comment.