Skip to content
Open access

Neural network monophonic melody generation with RNN, LSTM and GAN architectures

Sep 2026 · Scientific Reports · 0 citations

Abstract

Music is an expression of emotion and a universal language that connects people across various cultures. From timeless classical pieces to modern electronic beats, artificial intelligence is capable of composing, arranging, and even crafting refined musical arrangements in a wide variety of styles. This study, trains and develops an automated model to create melodies, on the monophonic genre specific datasets that serve as inputs to the Melody Generative System (MGS) to generate and compose melody automatically. To accomplish this, Recurrent Neural Network (RNN), Long Short-Term Memory (LSTM) network and Generative Adversarial Network (GAN) was used. To train the model, we used musical datasets, in Kern and MIDI formats primarily featuring solo piano compositions. Here, the Kern dataset uses specific text based notations to represent musical elements, whereas MIDI is a format containing information about musical performance data like notes and duration. This article also compares the results obtained from three different architectures using different activation functions, distinctly tanh and softmax. The study found that, tanh activation function led to higher accuracy and lower loss compared to softmax function in generating melodies. The comparative analysis of the 3 networks prove that, RNN-LSTM has the best accuracy of 98.39% for 80 − 20 train-test-split. Whereas, the performance of GAN network fell short compared to Simple RNN and RNN-LSTM when trained with both Kern and MIDI datasets. This study showcases the capabilities of different melody generation algorithms for generating monophonic melodies, and also demonstrates how deep learning can effectively recognize and replicate musical patterns.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.