Textual outlier detection with an unsupervised method using text similarity and density peak

Sereshki Mahnaz Taleb; Zanjireh Morteza Mohammadi; Bahaghighat Mahdi

doi:10.2478/ausi-2023-0008

Acta Universitatis Sapientiae: Informatica (Aug 2023)

Textual outlier detection with an unsupervised method using text similarity and density peak

Sereshki Mahnaz Taleb,
Zanjireh Morteza Mohammadi,
Bahaghighat Mahdi

Affiliations

Sereshki Mahnaz Taleb: 1Computer Engineering Department, Imam, Khomeini International University, Qazvin, Iran
Zanjireh Morteza Mohammadi: 2Computer Engineering Department, Imam Khomeini International University, Qazvin, Iran
Bahaghighat Mahdi: 3Computer Engineering Department, Imam Khomeini International, University, Qazvin, Iran

DOI: https://doi.org/10.2478/ausi-2023-0008
Journal volume & issue: Vol. 15, no. 1
pp. 91 – 110

Abstract

Read online

Text mining is an intriguing area of research, considering there is an abundance of text across the Internet and in social medias. Nevertheless outliers pose a challenge for textual data processing. The ability to identify this sort of irrelevant input is consequently crucial in developing high-performance models. In this paper, a novel unsupervised method for identifying outliers in text data is proposed. In order to spot outliers, we concentrate on the degree of similarity between any two documents and the density of related documents that might support integrated clustering throughout processing. To compare the e ectiveness of our proposed approach with alternative classification techniques, we performed a number of experiments on a real dataset. Experimental findings demonstrate that the suggested model can obtain accuracy greater than 98% and performs better than the other existing algorithms.

Published in Acta Universitatis Sapientiae: Informatica

ISSN: 2066-7760 (Online)
Publisher: Scientia Publishing House
Country of publisher: Romania
LCC subjects: Science: Mathematics: Instruments and machines: Electronic computers. Computer science
Website: https://acta.sapientia.ro/en/series/informatica

About the journal

Abstract

Keywords