Applied Sciences (Jul 2020)

Attributes Reduction in Big Data

  • Waleed Albattah,
  • Rehan Ullah Khan,
  • Khalil Khan

DOI
https://doi.org/10.3390/app10144901
Journal volume & issue
Vol. 10, no. 14
p. 4901

Abstract

Read online

Processing big data requires serious computing resources. Because of this challenge, big data processing is an issue not only for algorithms but also for computing resources. This article analyzes a large amount of data from different points of view. One perspective is the processing of reduced collections of big data with less computing resources. Therefore, the study analyzed 40 GB data to test various strategies to reduce data processing. Thus, the goal is to reduce this data, but not to compromise on the detection and model learning in machine learning. Several alternatives were analyzed, and it is found that in many cases and types of settings, data can be reduced to some extent without compromising detection efficiency. Tests of 200 attributes showed that with a performance loss of only 4%, more than 80% of the data could be ignored. The results found in the study, thus provide useful insights into large data analytics.

Keywords