Pith. sign in

REVIEW 1 cited by

JOSENet: A Joint Stream Embedding Network for Violence Detection in Surveillance Videos

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2405.02961 v2 pith:RYVXQLYZ submitted 2024-05-05 cs.CV eess.IV

classification cs.CVeess.IV
keywords detectionsurveillancevideosviolencejosenetvideoactionrecognition
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The increasing proliferation of video surveillance cameras and the escalating demand for crime prevention have intensified interest in the task of violence detection within the research community. Compared to other action recognition tasks, violence detection in surveillance videos presents additional issues, such as the wide variety of real fight scenes. Unfortunately, existing datasets for violence detection are relatively small in comparison to those for other action recognition tasks. Moreover, surveillance footage often features different individuals in each video and varying backgrounds for each camera. In addition, fast detection of violent actions in real-life surveillance videos is crucial to prevent adverse outcomes, thus necessitating models that are optimized for reduced memory usage and computational costs. These challenges complicate the application of traditional action recognition methods. To tackle all these issues, we introduce JOSENet, a novel self-supervised framework that provides outstanding performance for violence detection in surveillance videos. The proposed model processes two spatiotemporal video streams, namely RGB frames and optical flows, and incorporates a new regularized self-supervised learning approach for videos. JOSENet demonstrates improved performance compared to state-of-the-art methods, while utilizing only one-fourth of the frames per video segment and operating at a reduced frame rate. The source code is available at https://github.com/ispamm/JOSENet.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. SHARDeg: A Benchmark for Skeletal Human Action Recognition in Degraded Scenarios

    cs.CV 2025-05 conditional novelty 5.0 of 10

    A new benchmark degrades NTU-120 skeleton data three ways, shows degradation type strongly affects accuracy, and finds LogSigRNN overtakes DeGCN at 3 FPS once missing frames are interpolated.

Pith tools