Skip to content

Author

Shu-Han Yang

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Conference Open access 2026

Interpretable Vision-to-Music Mapping from Low-Level Visual Statistics to Mode, Rhythm, and Harmonic Structure

Vision-to-music generation transforms visual input into structured musical output, but many recent systems rely on end-to-end neural models whose internal cross-modal decisions are difficult to explain. This paper studies an interpretable alternative based on explicit visual analysis and rule-based symbolic generation. The method extracts low-level HSV statistics from video frames, constructs a smoothed visual activity signal S(t), derives non-uniform temporal segments, and maps global, segment-level, and local cues to tempo, mode, harmony, bass, pad, and melody. The framework outputs four-track symbolic music together with MIDI, stems, and metadata for direct inspection. Case studies on four clips show that the method can generate musically coherent results while keeping the relationship between visual structure and musical organization transparent. The analysis also reveals a current limitation, namely that abrupt visual change is often expressed through segmentation and accompaniment updates more strongly than through dense melodic motion. Future work may further improve melodic flexibility, introduce stronger user control, and evaluate the framework on broader video collections and listener studies.

Shu-Han Yang · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.