MusicLayout, an explicit intermediate representation for controlling musical structure in text-to-music generation, is introduced, providing evidence that explicit layout planning can improve long-range structural organization and support layout-level control.
Abstract
Text-to-music generation has advanced rapidly, but current systems still rely primarily on global text prompts, leaving the structural organization of generated music implicit and difficult to inspect, control, or revise before audio generation. To address this issue, we introduce MusicLayout, an explicit intermediate representation for controlling musical structure in text-to-music generation. MusicLayout describes a musical piece as a time-aligned layout of sections, textures, repetitions, variations, and instrument-level arrangements, serving as an interpretable planning layer between textual intent and the generated music. We integrate MusicLayout into a text-to-music framework built on a unified autoregressive formulation, where the model first generates a MusicLayout representation and subsequently predicts audio tokens conditioned on this representation within a single sequence. The resulting MusicLayout can be inspected and modified prior to audio generation, providing a mechanism for layout-level structural control. We evaluate MusicLayout through layout-conditioned generation, layout manipulation experiments, and matched-data ablations, providing evidence that explicit layout planning can improve long-range structural organization and support layout-level control. We have released the implementation as open source on GitHub at https://github.com/XaryLee/MusicLayout.
We introduce StepAudio 3 Music, a large-scale, long-form music generation model that supports explicit musical planning and open-domain text-controlled generation. The StepAudio Music Tokenizer represents audio as a 50-Hz stream from a 65536-entry single codebook, using semantically informed self-supervised and multi-t...
Chengli Feng, Zhi-Yue Wu, Jia-Hao Song et al.· 0 citations
Most music-generation systems are still framed and evaluated primarily as producers of complete outputs, whereas composition often proceeds through successive revisions to a shared musical artifact. This paper studies a different use of a general-purpose instruction-following large language model: not as a one-shot mus...
A. Fernandes, Jean-Pierre Briot, S. Barbosa et al.· 0 citations
Dance-to-music (D2M) generation aims to synthesize music that is rhythmically and stylistically aligned with dance videos. A key challenge arises from the semantic mismatch between sparse dance cues, such as rhythm and style, and the dense information required for music composition, including structure, instrumentation...
Jin-Ting Wang, Chen-Xing Li, Dong Yu et al.· 0 citations
A cross-modal framework that learns implicit music styles from raw audio and applies them to symbolic music generation and generates piano performances jointly conditioned on a lead sheet and a reference audio example, enabling controllable and stylistically faithful arrangement.
Jing-Wei Zhao, Gus G. Xia, Ziyu Wang et al.· 0 citations
Text-to-music language models begin with a choice usually made by default: how to tokenize music. Normally entangled with backbone, data, and recipe, its effect has never been measured in isolation. We fix pretrained Qwen3.5 (0.8B-27B), data, budget, and decoding, and swap only the representation across seven tokenizat...
Junhao Chen, Ming-Jin Chen, Jingjia Mao et al.· 1 citation
Large audio-language models can produce fluent and musically plausible responses, yet it often remains unclear whether those responses are grounded in the audio input. We introduce temporal music grounding, a task in which a model returns one or more time spans corresponding to a queried musical note, event, or pattern...
Kun Fang, Ziyu Wang, Ichiro Fujinaga· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.