State-of-the-art pose-based video anomaly detection models achieve over 52% frame-level AUC-ROC but drop below 10% event-level precision and 0.11 average F1 when evaluated with temporal action localization metrics on standard benchmarks.
Actionformer: Lo- calizing moments of actions with transformers
2 Pith papers cite this work. Polarity classification is still indexing.
fields
cs.CV 2verdicts
CONDITIONAL 2representative citing papers
ProcObject-10K is the first benchmark for object-centric procedural reasoning in videos that exposes a large gap where models answer questions plausibly but fail to ground their answers in the correct video segments.
citing papers explorer
-
From Frames to Events: Rethinking Evaluation in Human-Centric Video Anomaly Detection
State-of-the-art pose-based video anomaly detection models achieve over 52% frame-level AUC-ROC but drop below 10% event-level precision and 0.11 average F1 when evaluated with temporal action localization metrics on standard benchmarks.
-
ProcObject-10K: Benchmarking Object-Centric Procedural Understanding in Instructional Videos
ProcObject-10K is the first benchmark for object-centric procedural reasoning in videos that exposes a large gap where models answer questions plausibly but fail to ground their answers in the correct video segments.