Preprint
Aug 2026
AudioLens: Multi-Perspective Speech Clustering with Reasoning Audio-Language Models
This work introduces audio multi-perspective clustering, where a model directly partitions speech recordings according to a natural-language perspective while inferring both the number of clusters and their assignments.
Wen-Jun Huang, Qiao-Song Chu, Tiger Shao et al.
· 0 citations