Comparing human and nonhuman judgments of perceived similarity of gender diverse talkers
Listeners conduct fine-grained analyses of gender from speech, and their perceptual organization of voices is presumed to be grounded in structured acoustic-phonetic information. Here, we compare human perceived similarity to automated self-supervised speech representations from a deep-learning model (i.e., HuBERT), wh...