CrashSight is a new infrastructure-focused benchmark showing that state-of-the-art vision-language models can describe crash scenes but fail at temporal and causal reasoning.
Surveillancevqa-589k: A benchmark for comprehensive surveillance video-language understanding with large models
4 Pith papers cite this work. Polarity classification is still indexing.
fields
cs.CV 4representative citing papers
ProcObject-10K is the first benchmark for object-centric procedural reasoning in videos that exposes a large gap where models answer questions plausibly but fail to ground their answers in the correct video segments.
Introduces the first benchmark for metaphorical video understanding, identifies MLLM weaknesses in cross-domain mapping, and proposes an inference-time enhancement using a knowledge graph.
An agentic three-stage video annotation pipeline with an MSTED intermediate and top-down domain adaptation produces CoT training data that lifts Cosmos-Reason2 past Gemini on traffic event reasoning.
citing papers explorer
-
CrashSight: A Phase-Aware, Infrastructure-Centric Video Benchmark for Traffic Crash Scene Understanding and Reasoning
CrashSight is a new infrastructure-focused benchmark showing that state-of-the-art vision-language models can describe crash scenes but fail at temporal and causal reasoning.
-
ProcObject-10K: Benchmarking Object-Centric Procedural Understanding in Instructional Videos
ProcObject-10K is the first benchmark for object-centric procedural reasoning in videos that exposes a large gap where models answer questions plausibly but fail to ground their answers in the correct video segments.
-
MetaphorVU: Towards Metaphorical Video Understanding
Introduces the first benchmark for metaphorical video understanding, identifies MLLM weaknesses in cross-domain mapping, and proposes an inference-time enhancement using a knowledge graph.
-
MAVEN: A Multi-stage Agentic Annotation Pipeline for Video Reasoning Tasks
An agentic three-stage video annotation pipeline with an MSTED intermediate and top-down domain adaptation produces CoT training data that lifts Cosmos-Reason2 past Gemini on traffic event reasoning.