Skip to content

MIDI-RAE-JEPA: Hierarchical Representation Learning and Generation for Symbolic Music

Jul 2026 · arXiv.org · Vol abs/2607.14537 · 1 citation · 28 references
Computer Science Engineering

TL;DR

The results suggest that equivariance-based SSL objectives, combined with sufficient fine-level encoder capacity, provide a viable path toward semantically rich, generatively useful representations of symbolic music.

Abstract

Rich internal representations of musical structure are essential for music understanding tasks such as machine-assisted music co-writing, yet self-supervised approaches for symbolic music representation remain underexplored, particularly those that encode the hierarchical multiscale nature of musical structures. We present MIDI-RAE-JEPA, combining a pitch- and time-shift equivariance objective with LeJEPA and a Swin Transformer V2 encoder to learn such hierarchical representations of symbolic music encoded as piano roll images. The time-shift equivariance objective encourages the model to internalize temporal musical relationships. The encoder is trained purely on self-supervised objectives -- including a masked embedding predictor (MEP) -- with collapse prevented via SIGReg. A separate decoder trained on the frozen encoder embeddings achieves reconstruction F1 of 0.995, and a flow matching generative model conditioned on those embeddings produces generations that closely match the pitch register and rhythmic density of the conditioning excerpt, while mismatched conditioning yields unrelated but musically plausible output. Learned representations outperform a Haar scattering transform baseline on a downstream emotion classification task, and embedding distances increase monotonically with pitch and time shift magnitude, confirming measurable equivariance. These results suggest that equivariance-based SSL objectives, combined with sufficient fine-level encoder capacity, provide a viable path toward semantically rich, generatively useful representations of symbolic music.

View source

Similar papers

Preprint Aug 2026

Helping Music Co-Creation Agents'Listen'Well: Hierarchical Self-Supervised World Models for Understanding and Generation

Collaborative music agents need internal representations rich enough to support both understanding and generation, yet flexible enough for a workflow where the human retains agency. We present a hierarchical self-supervised ``world model''for symbolic music: a 2.55M-parameter Swin V2 encoder trained on MIDI piano-roll...

Scott H. Hawley · 0 citations
Preprint Aug 2026

Equivariant Music Transformer

Humans recognize a musical passage even when it is shifted in time or transposed in pitch, indicating a notion of equivariance in the representation space. Our analysis, however, shows that standard music transformers map such time-shifted or pitch-transposed inputs onto uncorrelated representations: these models becom...

Zi-Xun Guo, Simon Dixon · 0 citations
Preprint Aug 2026

SPHERE: Automatic Music Upmixing via Audio Language Model Post-Training with Spatial Heuristic Rewards

In this paper, we study the task of automatic music upmixing, wherein a system predicts spatial mixing parameters from a multi-stem recording. Different from existing methods that rely on task-specific music encoders, we approach this task via audio language model (ALM) post-training, leveraging rich representations fr...

Zi-Xun Guo, Calvin Murdock, Sanjeel Parekh et al. · 0 citations
Preprint Aug 2026

Learning Music Style for Piano Arrangement Through Cross-Modal Bootstrapping

A cross-modal framework that learns implicit music styles from raw audio and applies them to symbolic music generation and generates piano performances jointly conditioned on a lead sheet and a reference audio example, enabling controllable and stylistically faithful arrangement.

Jing-Wei Zhao, Gus G. Xia, Ziyu Wang et al. · 0 citations
#artificial intelligence Preprint Sep 2026

CMA-OT: Hierarchical Expert Supervision for Dance-to-Music Generation

Dance-to-music (D2M) generation aims to synthesize music that is rhythmically and stylistically aligned with dance videos. A key challenge arises from the semantic mismatch between sparse dance cues, such as rhythm and style, and the dense information required for music composition, including structure, instrumentation...

Jin-Ting Wang, Chen-Xing Li, Dong Yu et al. · 0 citations
Jul 2026

Music-JEPA: Learning a World Model of Sound from Action

This work proposes to learn a world model of piano sound using JEPA by framing music as an action-conditioned system, and shows that the learned model captures the relationships between musical actions and their resulting sound.

Ziyu Wang, Kun Fang, Yann LeCun · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.