IET Computer Vision (Apr 2023)

Violence 4D: Violence detection in surveillance using 4D convolutional neural networks

  • Mai Magdy,
  • Mohamed Waleed Fakhr,
  • Fahima A. Maghraby

DOI
https://doi.org/10.1049/cvi2.12162
Journal volume & issue
Vol. 17, no. 3
pp. 282 – 294

Abstract

Read online

Abstract As violence has increased around the world, surveillance cameras are everywhere, and they are only going to get more ubiquitous. Due to the massive volume of video footage, automatic activity detection systems must be used to create an online warning in the event of aberrant activity. A deep learning architecture is presented in this study using four‐dimensional video‐level convolution neural networks. The proposed architecture includes residual blocks that are used with three‐Dimensional Convolution Neural Networks 3D (CNNs) to learn long‐term and short‐term spatiotemporal representation from the video as well as record inter‐clip interaction. ResNet50 is used as the backbone for three‐dimensional convolution networks and dense optical flow for the region of interest. The proposed architecture is applied on four benchmarks for violence and non‐violence videos, which are commonly used for violent detection. It obtained test accuracies of 94.67% on RWF2000, 97.29% on Crowd violence, 100% on Movie fight and 100% on the Hockey Fight dataset. These results outperform the previous methods used on RWF2000 datasets.

Keywords