Jul 2026· International Conference on Image, Video and Signal Processing· Vol 14268, pp. 1426811 - 1426811-13· 0 citations· 22 references
Engineering
TL;DR
A style-aware data augmentation framework that combines rule-based design with statistical constraints, applied to small-scale, highly constrained Jiangnan symbolic music generation tasks is proposed, demonstrating the need for carefully designed augmentation strategies in highly constrained symbolic music generation tasks.
Abstract
This study proposes a style-aware data augmentation framework that combines rule-based design with statistical constraints, applied to small-scale, highly constrained Jiangnan symbolic music generation tasks. By leveraging YNote representation, fixed rhythmic frameworks, and Markov-style local transition statistics, we systematically expand the training data while maintaining musical structural plausibility. Using the augmented data, we fine-tune a GPT-2 model to analyze how different training data scales affect generation behavior. Experimental results show that Bilingual Evaluation Understudy (BLEU) -based reference-overlap metrics exhibit only minor fluctuations across different training scales and are insufficient to directly reflect style improvement. In contrast, Kullback-Leibler (KL) divergence and bigram statistics effectively characterize the style consistency of the generated set in terms of overall distribution proximity and local transition plausibility. Further analysis indicates that a medium-scale training set (approximately 3,000–12,000 samples) achieves the best balance between distribution consistency and transition coverage, whereas excessive augmentation may lead to distribution calibration drift. Overall, the study demonstrates that data augmentation has a positive but non-monotonic effect on style consistency, highlighting the need for carefully designed augmentation strategies in highly constrained symbolic music generation tasks.
Experimental results indicate that fully automated data curation combined with imbalance-aware training yields non-trivial improvements, but models still struggle to capture fine-grained acoustic features, indicating a gap between surface-level alignment and deep musical comprehension.
Ziya Zhou, Shangda Wu, Shenyang Xu et al.· 0 citations
What is music style? Though often described using text labels such as"swing,""classical,"or"emotional,"the real style remains implicit and hidden in concrete music examples. In this paper, we introduce a cross-modal framework that learns implicit music styles from raw audio and applies them to symbolic music generation. Inspired by BLIP-2, our model leverages a Querying Transformer (Q-Former) to extract style representations from a large, pre-trained audio language model (LM), and further applies them to condition a symbolic LM for generating piano arrangements. We adopt a two-stage training strategy: contrastive learning to align auditory style with symbolic expression, followed by generative modeling for music arrangement. Our model generates piano performances jointly conditioned on a lead sheet (content) and a reference audio example (style), enabling controllable and stylistically faithful arrangement. Experiments demonstrate the effectiveness of our approach in piano cover generation, style transfer, and audio-to-MIDI retrieval, achieving substantial improvements in style-aware alignment and music quality.
Jing-Wei Zhao, Gus G. Xia, Ziyu Wang et al.· 0 citations
Objective evaluation of expressive MIDI piano performances typically relies on attribute statistics such as timing, velocity, and duration of individual notes. However, these methods often disregard dependencies between notes, which poses a potential limitation in assessing the similarity between two sets of performances. In generative applications, the wide variety of expressive attributes makes it difficult to aggregate them into a single scalar metric for model selection. In this work, we reexamine attribute-scoped metrics and explore the perceptual properties of contextual embeddings from self-supervised symbolic music models, Aria and CLaMP3. Results from our listening study indicate that these models can be used as perceptual proxies, showing agreement with per-sample human ratings on par with traditional metrics. To measure conditional distributional similarity, we adapt Kernel Audio Distance to the symbolic music domain. Unlike Pearson correlation and reconstruction error, kernel-based methods on contextual embeddings do not require note alignment and are sensitive to contextual perturbations. To facilitate reproducibility, we release Pereval, an open-source library that integrates performance evaluation utilities, including both attribute-scoped and deep feature metrics.
D. Gavrilev, Ilya Borovik, Vladimir Viro· arXiv.org· 0 citations
Music is known to play a primary role in stress reduction, thereby enhancing our mental well-being. Retrieval of musical category tailored to specific needs of an individual is gaining traction but remains challenging and quite cumbersome. This is because of hazy and ambiguous nature of music classification, due to human subjectivity and disagreement which necessitates effective methods of classification. In the proposed work, the open source GTZAN dataset for musical genre classification from Music Analysis, Retrieval and Synthesis for Audio Signals (MARSYAS) collection has been utilized. Models despite achieving higher accuracy may provide incorrect predictions when genres overlap. In order to overcome this issue, this study proposes an uncertainty quantification by Dirichlet based evidence modeling with hybrid convolutional neural network-long short term memory network (CNN-LSTM), where the incorrect predictions have been penalized
via
Kullback-Leibler (KL) divergence. The initial expected calibration error (ECE) of 0.1401 and the corresponding reliability diagram suggest the overconfidence of the model in incorrect predictions. The ECE of 0.0791 after temperature scaling suggests the alignment of predicted confidence and the accuracy. From selective prediction plots, it can be observed that top 20% confident samples before calibration achieve an accuracy of 93–94%. Upon calibration, a similar accuracy is achieved by top 50% samples, which is a significant improvement. The study underscores the need for insights of quantitative trust upon the model, which is a crucial need for deploying music recommendation systems.
Long-form song generation models continue to improve in duration, structural integrity, and acoustic complexity, making reliable aesthetic rewards increasingly important for aligning these models with human preferences. However, reward models for complete songs remain limited, and existing evaluators typically predict scores in a single forward pass without providing readable explanations. We introduce MUSECRITIC, a semi-scalar reward model that generates a natural-language critique covering five aesthetic dimensions and uses it as an intermediate representation to predict continuous reward scores. MUSECRITIC follows a two-stage training pipeline: a teacher model first provides high-quality critiques for supervised fine-tuning, after which the fine-tuned model generates its own critiques for reward learning, mitigating distribution shift between training and inference. On an in-domain test set of 200 SongEval songs, MUSECRITIC reduces macro-averaged mean squared error from 0.2875 to 0.2316 and improves macro-averaged LCC, SRCC, and Kendall's tau to 0.9068, 0.8838, and 0.7178, respectively. On the out-of-domain Music Arena benchmark with 733 preference pairs, it achieves the highest accuracy of 71.35%. Moreover, using MUSECRITIC with GRPO improves Muse-0.6B on all nine aesthetic metrics from SongEval and Audiobox Aesthetics. These results demonstrate that critique-conditioned reward modeling reduces scoring error and provides an effective optimization signal for song generation. The project repository is available at https://github.com/WuqnEl/MuseCritic.
Jiabao Zhuang, Changhao Jiang, Hanchen Wang et al.· 0 citations
In this paper, we study the task of automatic music upmixing, wherein a system predicts spatial mixing parameters from a multi-stem recording. Different from existing methods that rely on task-specific music encoders, we approach this task via audio language model (ALM) post-training, leveraging rich representations from existing ALMs, which encode both music semantics and mixing knowledge. Specifically, we propose a post-training recipe that first employs rejection sampling SFT, followed by reinforcement learning (RL) with verifiable rewards (RLVR) via GRPO. We propose Sphere (Spatial Heuristic Rewards), a deterministic reward suite inspired by music mixing conventions, to guide our post-training. It consists of 6 perceptually-motivated sub-rewards and encourages the output mix to be centered, balanced and spacious. More broadly, our results suggest that expert domain knowledge can be encoded as verifiable rewards and distilled into language models, without task-specific architectures.
Zi-Xun Guo, Calvin Murdock, Sanjeel Parekh et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.