Mathematics (Jan 2023)

An Effective Fuzzy Clustering of Crime Reports Embedded by a Universal Sentence Encoder Model

  • Aparna Pramanik,
  • Asit Kumar Das,
  • Danilo Pelusi,
  • Janmenjoy Nayak

DOI
https://doi.org/10.3390/math11030611
Journal volume & issue
Vol. 11, no. 3
p. 611

Abstract

Read online

Crime reports clustering is crucial for identifying and preventing criminal activities that frequently happened in society. In the proposed work, named entities in a report are recognized to extract the crime-related phrases and subsequently, the phrases are preprocessed by applying stopword removal and lemmatization operations. Next, the module of the universal encoder model, called the transformer, is applied to extract phrases of the report to get a sentence embedding for each associated sentence, aggregation of which finally provides the vector representation of that report. An innovative and efficient graph-based clustering algorithm consisting of splitting and merging operations has been proposed to get the cluster of crime reports. The proposed clustering algorithm generates overlapping clusters, which indicates the existence of reports of multiple crime types. The fuzzy theory has been used to provide a score to the report for expressing its membership into different clusters, and accordingly, the reports are labelled by multiple categories. The efficiency of the proposed method has been assessed by taking into account different datasets and comparing them with other state-of-the-art approaches with the help of various performance measure metrics.

Keywords