Across three short traffic video sequences (two real, one synthetic), VideoLLaMA-2 outperformed GPT-4o, Gemini 1.5 Pro, InternVL, and LLaVA-NeXT-Video with 57 percent average accuracy, while all models exhibited clear gaps in multi-object tracking and temporal reasoning.
Align and aggregate: Compositional reasoning with video alignment and answer aggregation for video question-answering, 2024
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
method 1
citation-polarity summary
fields
cs.CV 1years
2024 1verdicts
CONDITIONAL 1roles
method 1polarities
use method 1representative citing papers
citing papers explorer
-
Eyes on the Road: State-of-the-Art Video Question Answering Models Assessment for Traffic Monitoring Tasks
Across three short traffic video sequences (two real, one synthetic), VideoLLaMA-2 outperformed GPT-4o, Gemini 1.5 Pro, InternVL, and LLaVA-NeXT-Video with 57 percent average accuracy, while all models exhibited clear gaps in multi-object tracking and temporal reasoning.