Cybernetics and Information Technologies (Mar 2022)

Combination of Resnet and Spatial Pyramid Pooling for Musical Instrument Identification

  • Dewi Christine,
  • Chen Rung-Ching

DOI
https://doi.org/10.2478/cait-2022-0007
Journal volume & issue
Vol. 22, no. 1
pp. 104 – 116

Abstract

Read online

Identifying similar objects is one of the most challenging tasks in computer vision image recognition. The following musical instruments will be recognized in this study: French horn, harp, recorder, bassoon, cello, clarinet, erhu, guitar saxophone, trumpet, and violin. Numerous musical instruments are identical in size, form, and sound. Further, our works combine Resnet 50 with Spatial Pyramid Pooling (SPP) to identify musical instruments that are similar to one another. Next, the Resnet 50 and Resnet 50 SPP model evaluation performance includes the Floating-Point Operations (FLOPS), detection time, mAP, and IoU. Our work can increase the detection performance of musical instruments similar to one another. The method we propose, Resnet 50 SPP, shows the highest average accuracy of 84.64% compared to the results of previous studies.

Keywords