Indonesian Online News Extraction and Clustering Using Evolving Clustering

Muhammad Alfian; Ali Ridho Barakbah; Idris Winarno

doi:10.30630/joiv.5.3.537

JOIV: International Journal on Informatics Visualization (Sep 2021)

Indonesian Online News Extraction and Clustering Using Evolving Clustering

Muhammad Alfian,
Ali Ridho Barakbah,
Idris Winarno

Affiliations

Muhammad Alfian: Informatics and Computer Engineering Department, Politeknik Elektronika Negeri Surabaya, Indonesia
Ali Ridho Barakbah: Informatics and Computer Engineering Department, Politeknik Elektronika Negeri Surabaya, Indonesia
Idris Winarno: Informatics and Computer Engineering Department, Politeknik Elektronika Negeri Surabaya, Indonesia

DOI: https://doi.org/10.30630/joiv.5.3.537
Journal volume & issue: Vol. 5, no. 3
pp. 280 – 290

Abstract

Read online

43,000 online media outlets in Indonesia publish at least one to two stories every hour. The amount of information exceeds human processing capacity, resulting in several impacts for humans, such as confusion and psychological pressure. This study proposes the Evolving Clustering method that continually adapts existing model knowledge in the real, ever-evolving environment without re-clustering the data. This study also proposes feature extraction with vector space-based stemming features to improve Indonesian language stemming. The application of the system consists of seven stages, (1) Data Acquisition, (2) Data Pipeline, (3) Keyword Feature Extraction, (4) Data Aggregation, (5) Predefined Cluster using Automatic Clustering algorithm, (6) Evolving Clustering, and (7) News Clustering Result. The experimental results show that Automatic Clustering generated 388 clusters as predefined clusters from 3.000 news. One of them is the unknown cluster. Evolving clustering runs for two days to cluster the news by streaming, resulting in a total of 611 clusters. Evolving clustering goes well, both updating models and adding models. The performance of the Evolving Clustering algorithm is quite good, as evidenced by the cluster accuracy value of 88%. However, some clusters are not right. It should be re-evaluated in the keyword feature extraction process to extract the appropriate features for grouping. In the future, this method can be developed further by adding other functions, updating and adding to the model, and evaluating.

Published in JOIV: International Journal on Informatics Visualization

ISSN: 2549-9610 (Print); 2549-9904 (Online)
Publisher: Politeknik Negeri Padang
Country of publisher: Indonesia
LCC subjects: Science: Mathematics: Instruments and machines: Electronic computers. Computer science: Computer software
Website: http://joiv.org

About the journal

Abstract

Keywords