REVIEW 4 major objections 8 minor 88 references
Event Stream based Multi-Modal Video Anomaly Detection: A Benchmark Dataset and Algorithms
T0 review · 4 major / 8 minor · reviewed 2026-07-13 · grok-4.5
Pith's one-line read Real event cameras plus video and text alignment make surveillance anomaly detection far more reliable under bad light and fast motion.
desk verdict Real visible–event VAD dataset is the real contribution; the pipeline works on it, but “events are essential” is oversold and public-benchmark events are mostly non-physical. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
E-VAD's two-stage pipeline: contrastive multi-modal pretraining that projects event frames into a shared CLIP-anchored space with video and ConceptNet-expanded text, followed by adaptive weighting fusion (a learned per-feature gate α that mixes video and event features) that supplies the anomaly classifier.
What would settle it
Run the identical E-VAD pipeline on a held-out real hybrid-camera capture set where events are never accumulated into frames (or where event frames are deliberately degraded to match the noise profile of monitor-re-recorded events) and check whether the large AUC gap over video-only and re-trained multi-modal baselines disappears.
Extended reading notes
Core claim
Jointly using real event streams with visible video, via contrastive multi-modal pretraining that aligns event-video-text embeddings and an adaptive weighting fusion that balances temporal event cues against spatial video features, yields consistently higher weakly-supervised anomaly detection performance than video-only or naively fused baselines, establishing that authentic event sensing is essential for robust real-world VAD.
Load-bearing premise
That turning raw asynchronous events into ordinary image-like frames and projecting them into a vision-language space trained on everyday photos still keeps enough of the original high-speed, high-dynamic-range motion signal for the claimed complementarity to hold.
Editorial extensions
If this is right
- Surveillance systems that add a real event camera can keep detecting anomalies under low light, glare and rapid motion that currently force video-only detectors to fail.
- New VAD benchmarks and methods must treat genuine event streams as a first-class modality rather than as a derived or simulated add-on.
- The same contrastive alignment-plus-adaptive-fusion recipe can be reused for other safety-critical perception tasks that already have hybrid event-video sensors.
- Industrial settings with subtle, sparse anomalies (labs, production lines) become practical targets for weakly supervised multi-modal detectors once real event data are available.
Reading between the lines
- If the event-frame + CLIP-projection step is the weakest link, later work that keeps events in native sparse or spiking form should widen the gap still further on real industrial data.
- The large performance jump on TJUTCM Pha versus the more modest gains on simulated or re-recorded public sets already hints that dataset realism, not just model architecture, is the main driver of the reported advance.
- The adaptive gate itself could become a diagnostic: high event weight regions may serve as automatic pointers to lighting or blur failures in deployed video systems.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes E-VAD, a weakly supervised multi-modal video anomaly detection framework that fuses conventional RGB video with asynchronous event streams. It introduces TJUTCM Pha, claimed as the first real-scene visible–event VAD benchmark (6.3B events, ~376k frames) collected with a DAVIS346 in pharmaceutical lab/production settings, and a two-stage method: (i) contrastive multi-modal pretraining that aligns event, video, and ConceptNet-expanded text embeddings in a CLIP-anchored space, and (ii) an adaptive weighting fusion (AWF) module that gates video vs. event features before temporal modeling and MIL-style snippet scoring. Experiments on ShanghaiTech-Event (simulated events), UCF-Crime-DVS (monitor re-recorded events), and TJUTCM Pha report strong gains, especially 82.25% AUC / 0.00% FAR on TJUTCM Pha versus re-trained multi-modal baselines around 72–74%, and the authors conclude that event sensing is complementary and essentially necessary for robust real-world VAD.
Significance. If the sensing-complementarity claim holds under fair controls, the work is significant for VAD: it reframes robustness as a sensing problem rather than purely architectural, and TJUTCM Pha would be a useful community resource because it is real hybrid capture rather than screen-replay or pure simulation. Strengths that should be credited include the scale and industrial realism of TJUTCM Pha, the explicit acknowledgment of limitations of UCF-Crime-DVS and ShanghaiTech-Event, thorough ablations of fusion variants / temporal modules / event representations on TJUTCM Pha, and qualitative score-curve and t-SNE analyses that support the design narrative. The free parameters (contrastive weights, μ, λ, top-k rule, etc.) are standard for this literature and do not by themselves invalidate the results. The main open question is attribution: how much of the large TJUTCM gap is physical event sensing versus the new CLIP+TM+AWF+text pipeline.
major comments (4)
- Central attribution claim vs. missing matched control (Abstract; §I contributions; Table V; §V.C–D). The paper asserts that event sensing is not merely beneficial but “essential,” with the 82.25% AUC on TJUTCM Pha (vs. ~73.8% best re-trained baseline) as primary evidence. Table V shows video-only 75.30, event-only 78.15, video+event 80.91, and full model 82.25, but there is no control that runs the identical CLIP encoder + temporal module + AWF-style head + MIL/KL training on video alone (or with a non-event second stream) on TJUTCM Pha. Without that matched architecture control, the jump from ~73–75% external video baselines to 82.25% cannot be cleanly attributed to physical event complementarity rather than the new pretraining/fusion pipeline. Please add this control (and ideally a non-event auxiliary stream) and revise the “essential” language to match what the controls support.
- Overstated “consistently outperforms” relative to public benchmarks (Abstract; §I; Tables II–III). On ShanghaiTech-Event, E-VAD reaches 98.67% AUC and is best among reported methods; on UCF-Crime-DVS, E-VAD AUC is 88.75%, below ITC’s 89.04% and only modestly above VADCLIP (88.02). Several re-trained video+event baselines also drop or barely move (e.g., VADCLIP* 87.94, UML* 85.95), which the paper itself links to non-physical event generation. The abstract and contribution bullets should be revised to state dataset-specific outcomes accurately rather than a blanket “consistently outperforms,” and to separate gains on real hybrid capture from gains on simulated/re-recorded events.
- Public-benchmark event validity and self-constructed evaluation risk (§IV Table I; §V.A; ShanghaiTech-Event construction). ShanghaiTech-Event is the authors’ DVS-Voltmeter simulation extension, and UCF-Crime-DVS is monitor re-recording; both inherit video temporal/photometric limits, as the manuscript notes. Using a self-built simulation extension as a primary SOTA table (Table II) while claiming event-driven superiority creates a mild circularity risk. Please (i) move ShanghaiTech-Event results to a clearly labeled “controlled / simulated” analysis, (ii) report any available statistics on event noise/sparsity differences vs. TJUTCM Pha, and (iii) avoid treating simulated/re-recorded gains as equivalent evidence for “real event sensing” in the conclusion.
- Event-frame representation vs. claimed microsecond / HDR advantages (§III.B.1; Table VIII; weakest assumption in the design). Asynchronous events are accumulated into CLIP-compatible event frames and projected into a natural-image CLIP space. Table VIII shows event frames beat voxel grids and time surfaces under the authors’ protocol, but this does not demonstrate that microsecond timing and high dynamic range survive the accumulation+CLIP pathway—especially when public events are already non-physical. Please quantify temporal binning (interval length, polarity handling) and discuss, with evidence, which claimed event advantages remain after framing; if the benefit is mainly motion-salient 2D patterns rather than true asynchronous sensing, the introduction and conclusion should say so.
minor comments (8)
- Abstract and elsewhere: “E VAD” / “EVAD” / “E-V AD” spacing is inconsistent; standardize to E-VAD.
- Table I: “#Frame” for TJUTCM Pha is listed as 377k while the abstract/text use 376,368; reconcile the count.
- Eq. (5)–(6): notation mixes L_ρq / L_et / L_ve / L_vt and “ρ! =q”; clarify that the three pairwise losses are the only terms and give default θ, β, γ values used in experiments.
- Eq. (12): L_kd is written with p_v2t / q_v2t but the surrounding text does not fully specify how semantic consistency labels are built for multi-class anomaly categories under weak labels.
- Fig. 2 caption and §III: “CLIP is not used as a pretrained event encoder” is important; consider elevating this clarification earlier to avoid reader confusion with EventCLIP/EventBind.
- §V.B: batch size 512 for 1000 epochs with max 200-frame sequences is heavy; a short note on wall-clock cost or hardware would aid reproducibility.
- Project link for TJUTCM Pha is deferred (“will be release later”); for a dataset paper this should be concrete before acceptance, including split files and event format.
- Minor typos: “outperforms methods” (abstract), “dimen-sion” (Eq. 9 context), “A WF” spacing in Table VI, “V oxel” in Table VIII.
Circularity Check
No derivation circularity; standard contrastive/MIL training and empirical SOTA claims on a new real event dataset plus public benchmarks (with simulated/re-recorded events noted as limited).
full rationale
This is an empirical multi-modal VAD systems paper. The load-bearing claims (E-VAD outperforms via event–video–text contrastive pretraining + adaptive fusion; event sensing is essential) rest on standard losses (pairwise InfoNCE-style L_et/L_ve/L_vt, MIL top-k classification L_cs, and KL L_kd) and measured AUC/FAR on held-out test splits. No equation defines a quantity in terms of the target it then “predicts,” no parameter is fitted to a subset and then reported as an independent prediction of a near-identical quantity, and no uniqueness theorem or ansatz is imported via self-citation to force the architecture. ShanghaiTech-Event is an author-constructed simulation extension (via DVS-Voltmeter) used as one controlled benchmark, and TJUTCM Pha is their new real capture set; both are normal dataset contributions, not algebraic self-definitions of the reported scores. Public UCF-Crime-DVS results and ablations (video-only, event-only, fusion variants) supply independent grounding. Mild self-construction risk exists only in the sense that the largest absolute gains appear on the authors’ own industrial set, but that does not reduce any claimed result to its inputs by construction. Score 1 reflects that residual dataset-construction proximity without elevating it to circular derivation.
Assumptions & free parameters
free parameters (7)
- contrastive loss weights θ, β, γ (L_totalpre)
- detection loss weight λ (L = L_ce + λ L_kd)
- temporal mix μ (F_o = μ X_g + (1-μ) X_l)
- contrastive temperature τ
- fusion gate initialization α≈0.5 and linear gate parameters W_α, b_α
- top-k snippet rule (k=⌊T/16+1⌋ anomalous, k=1 normal)
- training hyperparameters (lr 1e-3, batch 512, 1000 epochs, max 200 frames, local interval 5)
assumptions (6)
- domain assumption Frame-like accumulation of events over intervals preserves enough temporal dynamics for VAD (inter-frame evolution dominates fine intra-frame timing).
- domain assumption CLIP image/text embedding space is a valid semantic anchor for projecting event features via a learnable projector.
- domain assumption Video-level binary labels plus multiple-instance learning (top-k mean) suffice to localize snippet anomalies under weak supervision.
- domain assumption ConceptNet-expanded category text provides useful supervisory semantics for anomaly subclasses.
- ad hoc to paper Simulated (ShanghaiTech-Event) and monitor-rerecorded (UCF-Crime-DVS) events are informative enough for multi-modal comparison, even if not fully realistic.
- standard math Scaled dot-product attention and causal convolution are appropriate temporal models for snippet scoring.
invented entities (3)
-
TJUTCM Pha dataset
-
E-VAD adaptive weighting fusion (AWF) module
-
ShanghaiTech-Event extension
Cite this review
Pith. "Pith review of Event Stream based Multi-Modal Video Anomaly Detection: A Benchmark Dataset and Algorithms." pith.science (2026). https://pith.science/paper/FZ5C3BWR
@misc{pith2026260709114,
author = {Pith},
title = {Pith review of: Event Stream based Multi-Modal Video Anomaly Detection: A Benchmark Dataset and Algorithms},
year = {2026},
howpublished = {\url{https://pith.science/paper/FZ5C3BWR}},
note = {Machine review of arXiv:2607.09114}
}
read the original abstract
Video anomaly detection (VAD) is critical for automated surveillance but remains fragile under challenging conditions such as illumination variations, fast motion, and complex backgrounds when relying solely on visible light videos. To address these limitations, we propose EVAD, an event enhanced VAD framework that jointly exploits conventional video and event streams captured by bio inspired event cameras. Event sensors asynchronously capture brightness changes with high temporal resolution, offering robustness to motion blur and extreme lighting, and providing motion salient cues complementary to video based visual information. To support multi modal VAD research, we construct a large scale visible event benchmark comprising 6.3 billion events and 376,368 video frames collected under diverse illumination levels, motion patterns, and background complexities, filling the gap of realistic and scalable datasets for event based anomaly detection. Building upon this dataset, we design a contrastive multi modal pretraining framework to learn discriminative event representations by aligning semantic embeddings across event streams, visible videos, and textual descriptions. An adaptive fusion module then dynamically integrates event based temporal cues with video based spatial semantics, improving robustness to environmental disturbances. Experiments on benchmarks and the proposed TJUTCM Pha dataset demonstrate that E VAD consistently outperforms methods, validating the effectiveness of event-based sensing for VAD in real world scenarios.
Figures
Figures from the paper (5 more)
Reference graph
Works this paper leans on
-
[1]
Narwade, R
P. Narwade, R. Kawamura, G. Gajbhiye, K. Niinuma, Synthetic video generation for weakly supervised cross-domain video anomaly detection, in: International Conference on Pattern Recognition, Springer, 2024, pp. 375–391
2024
-
[2]
Zhang, J
M. Zhang, J. Wang, Q. Qi, H. Sun, Z. Zhuang, P. Ren, R. Ma, J. Liao, Multi-scale video anomaly detection by multi-grained spatio-temporal representation learning, in: 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), IEEE, 2024, pp. 17385–17394
2024
-
[3]
J. Ahn, J. Park, S. S. Lee, K.-H. Lee, H. Do, J. Ko, Safefac: Video-based smart safety monitoring for preventing industrial work accidents, Expert Systems with Applications 215 (2023) 119397
2023
-
[4]
Z. Chen, H. Wang, C. Li, C. Liu, F. Yang, D. Zhang, A. J. Fauci, J. Zhang, Large language models in traditional chinese medicine: a systematic review, Acupuncture and Herbal Medicine 5 (1) (2025) 57– 67
2025
-
[5]
X. Liu, T. Gong, Artificial intelligence and evidence-based research will promote the development of traditional medicine, Acupuncture and Herbal Medicine 4 (1) (2024) 134–135
2024
-
[6]
Bogdoll, M
D. Bogdoll, M. Nitsche, J. M. Z ¨ollner, Anomaly detection in autonomous driving: A survey, in: 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), IEEE, 2022, pp. 4487– 4498
2022
-
[7]
F. U. M. Ullah, M. S. Obaidat, A. Ullah, K. Muhammad, M. Hijji, S. W. Baik, A comprehensive review on vision-based violence detection in surveillance videos, ACM Computing Surveys 55 (10) (2023) 1–44
2023
-
[8]
P. Zhu, X. Wang, Y . Luo, Z. Sun, W.-S. Zheng, Y . Wang, C. Chen, Unpaired image captioning by image-level weakly-supervised visual concept recognition, IEEE Transactions on Multimedia 25 (2022) 6702– 6716
2022
Show all 88 references
-
[9]
Z. Ren, L. He, P. Zhu, Super-resolution learning strategy based on expert knowledge supervision, Remote Sensing 16 (16) (2024) 2888
2024
-
[10]
F. Liu, Y . Wen, J. Sun, P. Zhu, L. Mao, G. Niu, J. Li, Iterative mamba diffusion change-detection model for remote sensing, Remote Sensing 16 (19) (2024) 3651
2024
-
[11]
Y . Bao, H. Ding, Z. Zhang, K. Yang, Q. Tran, Q. Sun, T. Xu, Intelligent acupuncture: data-driven revolution of traditional chinese medicine, Acupuncture and Herbal Medicine 3 (4) (2023) 271–284
2023
-
[12]
Zhong, G
Y . Zhong, G. Yan, Y . Hu, D. Zhu, R. Zhu, A two-stage framework with memory for anomaly detection via video decomposition and bidirectional consistency, IEEE Transactions on Circuits and Systems for Video Technology (2025)
2025
-
[13]
Xie, Y .-W
X.-D. Xie, Y .-W. Zhan, Z.-X. Ma, H.-M. Liu, Z.-D. Chen, X. Luo, X.-S. Xu, Distributed learning for privacy-preserving semi-supervised video anomaly detection, IEEE Transactions on Circuits and Systems for Video Technology (2025)
2025
-
[14]
Zhong, R
Y . Zhong, R. Zhu, G. Yan, P. Gan, X. Shen, D. Zhu, Inter-clip feature similarity based weakly supervised video anomaly detection via multi- scale temporal mlp, IEEE Transactions on Circuits and Systems for Video Technology (2024)
2024
-
[15]
Gehrig, D
D. Gehrig, D. Scaramuzza, Low-latency automotive vision with event cameras, Nature 629 (8014) (2024) 1034–1040
2024
-
[16]
Z. Wu, M. Gehrig, Q. Lyu, X. Liu, I. Gilitschenski, Leod: Label-efficient object detection for event cameras, in: 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), IEEE, 2024, pp. 16933–16942
2024
-
[17]
X. Wang, S. Wang, C. Tang, L. Zhu, B. Jiang, Y . Tian, J. Tang, Event stream-based visual object tracking: A high-resolution benchmark dataset and a novel baseline, in: Proceedings of the IEEE/CVF Confer- ence on Computer Vision and Pattern Recognition, 2024, pp. 19248– 19257
2024
-
[18]
Gallego, T
G. Gallego, T. Delbr ¨uck, G. Orchard, C. Bartolozzi, B. Taba, A. Censi, S. Leutenegger, A. J. Davison, J. Conradt, K. Daniilidis, et al., Event- based vision: A survey, IEEE transactions on pattern analysis and machine intelligence 44 (1) (2020) 154–180
2020
-
[19]
Y . Qian, S. Ye, C. Wang, X. Cai, J. Qian, J. Wu, Ucf-crime-dvs: A novel event-based dataset for video anomaly detection with spiking neural networks, in: Proceedings of the AAAI Conference on Artificial Intelligence, V ol. 39, 2025, pp. 6577–6585
2025
-
[20]
Flaborea, L
A. Flaborea, L. Collorone, G. M. D. Di Melendugno, S. D’Arrigo, B. Prenkaj, F. Galasso, Multimodal motion conditioned diffusion model for skeleton-based video anomaly detection, in: Proceedings of the IEEE/CVF international conference on computer vision, 2023, pp. 10318–10329
2023
-
[21]
G. Wang, Y . Wang, J. Qin, D. Zhang, X. Bao, D. Huang, Video anomaly detection by solving decoupled spatio-temporal jigsaw puzzles, in: European Conference on Computer Vision, Springer, 2022, pp. 494– 511
2022
-
[22]
H. Liu, L. He, M. Zhang, F. Li, Vadiffusion: Compressed domain information guided conditional diffusion for video anomaly detection, IEEE Transactions on Circuits and Systems for Video Technology 34 (9) (2024) 8398–8411
2024
-
[23]
C. Guo, L. Li, Y . Ren, X. Zhang, G. Feng, Aligning normal representa- tions in diffusion model for video anomaly detection, IEEE Transactions on Circuits and Systems for Video Technology (2025)
2025
-
[24]
C. Guo, H. Wang, Y . Xia, G. Feng, Learning appearance-motion synergy via memory-guided event prediction for video anomaly detection, IEEE Transactions on Circuits and Systems for Video Technology 34 (3) (2023) 1519–1531
2023
-
[25]
Z. Yang, J. Liu, P. Wu, Text prompt with normality guidance for weakly supervised video anomaly detection, in: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2024, pp. 18899– 18908
2024
-
[26]
M. I. Georgescu, R. T. Ionescu, F. S. Khan, M. Popescu, M. Shah, A background-agnostic framework with adversarial training for abnormal 13 event detection in video, IEEE transactions on pattern analysis and machine intelligence 44 (9) (2021) 4505–4523
2021
-
[27]
Georgescu, A
M.-I. Georgescu, A. Barbalau, R. T. Ionescu, F. S. Khan, M. Popescu, M. Shah, Anomaly detection in video via self-supervised and multi- task learning, in: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2021, pp. 12742–12752
2021
-
[28]
P. Wu, X. Zhou, G. Pang, Y . Sun, J. Liu, P. Wang, Y . Zhang, Open- vocabulary video anomaly detection, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 18297–18307
2024
-
[29]
X. Wang, Y . Jin, W. Wu, W. Zhang, L. Zhu, B. Jiang, Y . Tian, Object detection using event camera: A moe heat conduction based detector and a new benchmark dataset, in: Proceedings of the Computer Vision and Pattern Recognition Conference, 2025, pp. 29321–29330
2025
-
[30]
Chakravarthi, A
B. Chakravarthi, A. A. Verma, K. Daniilidis, C. Fermuller, Y . Yang, Recent event camera innovations: A survey, in: European Conference on Computer Vision, Springer, 2024, pp. 342–376
2024
-
[31]
C. Feng, W. Yu, X. Cheng, Z. Tang, J. Zhang, L. Yuan, Y . Tian, Ae-nerf: Augmenting event-based neural radiance fields for non-ideal conditions and larger scenes, in: Proceedings of the AAAI Conference on Artificial Intelligence, V ol. 39, 2025, pp. 2924–2932
2025
-
[32]
Y . Yang, L. Pan, L. Liu, Event camera data pre-training, in: Proceedings of the IEEE/CVF international conference on computer vision, 2023, pp. 10699–10709
2023
-
[33]
Zubic, M
N. Zubic, M. Gehrig, D. Scaramuzza, State space models for event cameras, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 5819–5828
2024
-
[34]
P. Zhu, X. Wang, L. Zhu, Z. Sun, W.-S. Zheng, Y . Wang, C. Chen, Prompt-based learning for unpaired image captioning, IEEE Transac- tions on Multimedia 26 (2023) 379–393
2023
-
[35]
Radford, J
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, et al., Learning transferable visual models from natural language supervision, in: International con- ference on machine learning, PmLR, 2021, pp. 8748–8763
2021
-
[36]
Chen, D.-Z
F.-L. Chen, D.-Z. Zhang, M.-L. Han, X.-Y . Chen, J. Shi, S. Xu, B. Xu, Vlp: A survey on vision-language pre-training, Machine Intelligence Research 20 (1) (2023) 38–56
2023
-
[37]
Zhang, Z
R. Zhang, Z. Guo, W. Zhang, K. Li, X. Miao, B. Cui, Y . Qiao, P. Gao, H. Li, Pointclip: Point cloud understanding by clip, in: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2022, pp. 8552–8562
2022
-
[38]
S. T. Wasim, M. Naseer, S. Khan, F. S. Khan, M. Shah, Vita-clip: Video and text adaptive clip via multimodal prompting, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023, pp. 23034–23044
2023
-
[39]
Zhang, Z
R. Zhang, Z. Zeng, Z. Guo, Y . Li, Can language understand depth?, in: Proceedings of the 30th ACM International Conference on Multimedia, 2022, pp. 6868–6874
2022
-
[40]
Z. Wu, X. Liu, I. Gilitschenski, Eventclip: Adapting clip for event-based object recognition, arXiv preprint arXiv:2306.06354 (2023)
2023 arXiv
-
[41]
J. Zhou, X. Zheng, Y . Lyu, L. Wang, Eventbind: Learning a unified representation to bind them all for event-based open-world understand- ing, in: European Conference on Computer Vision, Springer, 2024, pp. 477–494
2024
-
[42]
Speer, C
R. Speer, C. Havasi, Conceptnet 5: A large semantic network for relational knowledge, in: The People’s Web Meets NLP: Collaboratively Constructed Language Resources, Springer, 2013, pp. 161–176
2013
-
[43]
Y . Fan, Y . Yu, W. Lu, Y . Han, Weakly-supervised video anomaly de- tection with snippet anomalous attention, IEEE Transactions on Circuits and Systems for Video Technology 34 (7) (2024) 5480–5492
2024
-
[44]
S. Yan, R. Zhang, Z. Guo, W. Chen, W. Zhang, H. Li, Y . Qiao, H. Dong, Z. He, P. Gao, Referred by multi-modality: A unified temporal transformer for video object segmentation, in: Proceedings of the AAAI Conference on Artificial Intelligence, V ol. 38, 2024, pp. 6449–6457
2024
-
[45]
Y . Pu, X. Wu, L. Yang, S. Wang, Learning prompt-enhanced context features for weakly-supervised video anomaly detection, IEEE Transac- tions on Image Processing (2024)
2024
-
[46]
Z.-Y . Hu, Y . Zhong, S. Huang, M. Lyu, L. Wang, Enhancing temporal modeling of video llms via time gating, in: Findings of the Association for Computational Linguistics: EMNLP 2024, 2024, pp. 2845–2856
2024
-
[47]
P. Wu, J. Liu, Learning causal temporal relation and feature discrimina- tion for anomaly detection, IEEE Transactions on Image Processing 30 (2021) 3513–3527
2021
-
[48]
Sultani, C
W. Sultani, C. Chen, M. Shah, Real-world anomaly detection in surveil- lance videos, in: Proceedings of the IEEE conference on computer vision and pattern recognition, 2018, pp. 6549–6558
2018
-
[49]
A. Adam, E. Rivlin, I. Shimshoni, R. Reinitz, Robust real-time unusual event detection using multiple fixed-location monitors, IEEE Transac- tions on Pattern Analysis and Machine Intelligence 30 (3) (2008) 555– 560
2008
-
[50]
C. Lu, J. Shi, J. Jia, Abnormal event detection at 150 fps in matlab, in: Proceedings of the IEEE International Conference on Computer Vision, 2013, pp. 2720–2727
2013
-
[51]
Mahadevan, W
V . Mahadevan, W. Li, V . Bhalodia, N. Vasconcelos, Anomaly detection in crowded scenes, in: 2010 IEEE Computer Society Conference on Computer Vision and Pattern Recognition, IEEE, 2010, pp. 1976–1983
2010
-
[52]
P. Wu, J. Liu, Y . Shi, Y . Sun, F. Shao, Z. Wu, Z. Yang, Not only look, but also listen: Learning multimodal violence detection under weak supervision, in: European conference on computer vision, Springer, 2020, pp. 322–339
2020
-
[53]
A. Amir, B. Taba, D. Berg, T. Melano, J. McKinstry, C. Di Nolfo, T. Nayak, A. Andreopoulos, G. Garreau, M. Mendoza, et al., A low power, fully event-based gesture recognition system, in: Proceedings of the IEEE conference on computer vision and pattern recognition, 2017, pp. 7243–7252
2017
-
[54]
Y . Bi, A. Chadha, A. Abbas, E. Bourtsoulatze, Y . Andreopoulos, Graph-based object classification for neuromorphic vision sensing, in: Proceedings of the IEEE/CVF international conference on computer vision, 2019, pp. 491–501
2019
-
[55]
Y . Dong, Y . Li, D. Zhao, G. Shen, Y . Zeng, Bullying10k: a large-scale neuromorphic dataset towards privacy-preserving bullying recognition, Advances in Neural Information Processing Systems 36 (2023) 1923– 1937
2023
-
[56]
W. Liu, W. Luo, D. Lian, S. Gao, Future frame prediction for anomaly detection–a new baseline, in: Proceedings of the IEEE conference on computer vision and pattern recognition, 2018, pp. 6536–6545
2018
-
[57]
S. Lin, Y . Ma, Z. Guo, B. Wen, Dvs-voltmeter: Stochastic process-based event simulator for dynamic vision sensors, in: European Conference on Computer Vision, Springer, 2022, pp. 578–593
2022
-
[58]
Zhang, L
J. Zhang, L. Qing, J. Miao, Temporal convolutional network with complementary inner bag loss for weakly supervised anomaly detection, in: 2019 IEEE International Conference on Image Processing (ICIP), IEEE, 2019, pp. 4030–4034
2019
-
[59]
Zhong, N
J.-X. Zhong, N. Li, W. Kong, S. Liu, T. H. Li, G. Li, Graph convolutional label noise cleaner: Train a plug-and-play action classifier for anomaly detection, in: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2019, pp. 1237–1246
2019
-
[60]
M. Z. Zaheer, A. Mahmood, M. Astrid, S.-I. Lee, Claws: Clustering assisted weakly supervised learning with normalcy suppression for anomalous event detection, in: European Conference on Computer Vision, Springer, 2020, pp. 358–376
2020
-
[61]
Feng, F.-T
J.-C. Feng, F.-T. Hong, W.-S. Zheng, Mist: Multiple instance self- training framework for video anomaly detection, in: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2021, pp. 14009–14018
2021
-
[62]
Y . Tian, G. Pang, Y . Chen, R. Singh, J. W. Verjans, G. Carneiro, Weakly-supervised video anomaly detection with robust temporal feature magnitude learning, in: Proceedings of the IEEE/CVF international conference on computer vision, 2021, pp. 4975–4986
2021
-
[63]
S. Li, F. Liu, L. Jiao, Self-training multi-sequence learning with trans- former for weakly supervised video anomaly detection, in: Proceedings of the AAAI Conference on Artificial Intelligence, V ol. 36, 2022, pp. 1395–1403
2022
-
[64]
S. Park, H. Kim, M. Kim, D. Kim, K. Sohn, Normality guided multiple instance learning for weakly supervised video anomaly detection, in: Proceedings of the IEEE/CVF winter conference on applications of computer vision, 2023, pp. 2665–2674
2023
-
[65]
Wu, H.-Y
J.-C. Wu, H.-Y . Hsieh, D.-J. Chen, C.-S. Fuh, T.-L. Liu, Self-supervised sparse representation for video anomaly detection, in: European Confer- ence on Computer Vision, Springer, 2022, pp. 729–745
2022
-
[66]
M. Cho, M. Kim, S. Hwang, C. Park, K. Lee, S. Lee, Look around for anomalies: Weakly-supervised anomaly detection via context-motion relational learning, in: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2023, pp. 12137–12146
2023
-
[67]
B. Wan, Y . Fang, X. Xia, J. Mei, Weakly supervised video anomaly detection via center-guided discriminative learning, in: 2020 IEEE International Conference on Multimedia and Expo (ICME), IEEE, 2020, pp. 1–6
2020
-
[68]
H. Lv, Z. Yue, Q. Sun, B. Luo, Z. Cui, H. Zhang, Unbiased multiple instance learning for weakly supervised video anomaly detection, in: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2023, pp. 8022–8031. 14
2023
-
[69]
T. Zhu, Q. Yu, X. Dong, S. Li, Y . Liu, J. Jiang, L. Shu, Prodisc-vad: An efficient system for weakly-supervised anomaly detection in video surveillance applications, arXiv e-prints (2025) arXiv–2505
2025
-
[70]
Y . Zhu, S. Newsam, Motion-aware feature for improved video anomaly detection, arXiv preprint arXiv:1907.10211 (2019)
1907 arXiv
-
[71]
Y . Zhen, Y . Guo, J. Wei, X. Bao, D. Huang, Multi-scale background suppression anomaly detection in surveillance videos, in: 2021 IEEE International Conference on Image Processing (ICIP), IEEE, 2021, pp. 1114–1118
2021
-
[72]
Y . Pu, X. Wu, Locality-aware attention network with discriminative dynamics learning for weakly supervised anomaly detection, in: 2022 IEEE International Conference on Multimedia and Expo (ICME), IEEE, 2022, pp. 1–6
2022
-
[73]
Zhang, G
C. Zhang, G. Li, Q. Xu, X. Zhang, L. Su, Q. Huang, Weakly supervised anomaly detection in videos considering the openness of events, IEEE transactions on intelligent transportation systems 23 (11) (2022) 21687– 21699
2022
-
[74]
Zhang, G
C. Zhang, G. Li, Y . Qi, S. Wang, L. Qing, Q. Huang, M.-H. Yang, Exploiting completeness and uncertainty of pseudo labels for weakly supervised video anomaly detection, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023, pp. 16271–16280
2023
-
[75]
P. Wu, X. Zhou, G. Pang, L. Zhou, Q. Yan, P. Wang, Y . Zhang, Vadclip: Adapting vision-language models for weakly supervised video anomaly detection (2024)
2024
-
[76]
Liu, K.-M
T. Liu, K.-M. Lam, B.-K. Bao, Injecting text clues for improving anoma- lous event detection from weakly labeled videos, IEEE Transactions on Image Processing (2024)
2024
-
[77]
Y . Chen, Z. Liu, B. Zhang, W. Fok, X. Qi, Y .-C. Wu, Mgfn: Magnitude-contrastive glance-and-focus network for weakly-supervised video anomaly detection, in: Proceedings of the AAAI conference on artificial intelligence, V ol. 37, 2023, pp. 387–395
2023
-
[78]
H. K. Joo, K. V o, K. Yamazaki, N. Le, Clip-tsa: Clip-assisted temporal self-attention for weakly-supervised video anomaly detection, in: 2023 IEEE International Conference on Image Processing (ICIP), IEEE, 2023, pp. 3230–3234
2023
-
[79]
Y . Zhou, Y . Qu, X. Xu, F. Shen, J. Song, H. T. Shen, Batchnorm- based weakly supervised video anomaly detection, IEEE Transactions on Circuits and Systems for Video Technology (2024)
2024
-
[80]
H. Zhou, J. Yu, W. Yang, Dual memory units with uncertainty regulation for weakly supervised video anomaly detection, in: Proceedings of the AAAI Conference on Artificial Intelligence, V ol. 37, 2023, pp. 3769– 3777
2023
-
[81]
S. Tu, Q. Dai, Z. Wu, Z.-Q. Cheng, H. Hu, Y .-G. Jiang, Implicit temporal modeling with learnable alignment for video recognition, in: Proceedings of the ieee/cvf international conference on computer vision, 2023, pp. 19936–19947
2023
-
[82]
Y . Tang, L. Qi, F. Xie, X. Li, C. Ma, M.-H. Yang, Video predic- tion transformers without recurrence or convolution, in: arXiv preprint arxiv:2410.04733, 2024
2024
-
[83]
H. Lu, G. Yang, N. Fei, Y . Huo, Z. Lu, P. Luo, M. Ding, Vdt: General- purpose video diffusion transformers via mask modeling, arXiv preprint arXiv:2305.13311 (2023)
2023 arXiv
-
[84]
S. T. Wasim, M. U. Khattak, M. Naseer, S. Khan, M. Shah, F. S. Khan, Video-focalnets: Spatio-temporal focal modulation for video action recognition, in: Proceedings of the IEEE/CVF International Conference on Computer Vision, 2023, pp. 13778–13789
2023
-
[85]
S. Yang, Y . Zhou, Z. Liu, C. C. Loy, Fresco: Spatial-temporal correspon- dence for zero-shot video translation, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 8703–8712
2024
-
[86]
L. Qian, J. Li, Y . Wu, Y . Ye, H. Fei, T.-S. Chua, Y . Zhuang, S. Tang, Momentor: advancing video large language model with fine-grained temporal reasoning, in: Proceedings of the 41st International Conference on Machine Learning, 2024, pp. 41340–41356
2024
-
[87]
L. v. d. Maaten, G. Hinton, Visualizing data using t-sne, Journal of machine learning research 9 (Nov) (2008) 2579–2605. Peipei Zhureceived the B.E. degree from Hebei University of Technology, Tianjin, China, in 2014, the M.E. degree from Beijing University of Posts and Teleco...
2008
-
[88]
He is currently an Associate Professor at Guangzhou Institute of Technology, Xi- dian University
From January to March 2020, he was an Intern at Bell Labs France, where he worked on indoor localization systems. He is currently an Associate Professor at Guangzhou Institute of Technology, Xi- dian University. His research interests include UA V– UGV cooperative systems, mil...
2020
Reviewed July 13, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.