Uncertainty-aware music genre classification using evidential deep learning
Music is known to play a primary role in stress reduction, thereby enhancing our mental well-being. Retrieval of musical category tailored to specific needs of an individual is gaining traction but remains challenging and quite cumbersome. This is because of hazy and ambiguous nature of music classification, due to human subjectivity and disagreement which necessitates effective methods of classification. In the proposed work, the open source GTZAN dataset for musical genre classification from Music Analysis, Retrieval and Synthesis for Audio Signals (MARSYAS) collection has been utilized. Models despite achieving higher accuracy may provide incorrect predictions when genres overlap. In order to overcome this issue, this study proposes an uncertainty quantification by Dirichlet based evidence modeling with hybrid convolutional neural network-long short term memory network (CNN-LSTM), where the incorrect predictions have been penalized via Kullback-Leibler (KL) divergence. The initial expected calibration error (ECE) of 0.1401 and the corresponding reliability diagram suggest the overconfidence of the model in incorrect predictions. The ECE of 0.0791 after temperature scaling suggests the alignment of predicted confidence and the accuracy. From selective prediction plots, it can be observed that top 20% confident samples before calibration achieve an accuracy of 93–94%. Upon calibration, a similar accuracy is achieved by top 50% samples, which is a significant improvement. The study underscores the need for insights of quantitative trust upon the model, which is a crucial need for deploying music recommendation systems.