SonicWeave: Chunk-Routed Mixture-of-Experts for Unified Audio Scene Generation
Text-conditioned general audio generation is moving beyond isolated speech, music, and sound-effect synthesis toward a single model that can compose them into controllable, coherent audio scenes. This unified setting is particularly challenging: heterogeneous components impose conflicting structural requirements on a s...