Skip to content

Probing Low-Level Acoustic Attribute Encoding in CLAP Audio Embeddings

Jul 2026 · arXiv.org · Vol abs/2607.03806 · 1 citation · 38 references
Engineering Computer Science

TL;DR

This work analyzes CLAP audio embeddings through a probing framework, studying the encoding of three fundamental perceptual dimensions: reverberation, loudness, and spectral content, measured via spectral centroid (SC) and relative pitch (RP).

Abstract

Audio foundation models are widely adopted as general-purpose feature extractors, yet the internal structure of their learned representations remains insufficiently understood. In this work, we analyze CLAP audio embeddings through a probing framework, studying the encoding of three fundamental perceptual dimensions: reverberation (RT60), loudness (LUFS), and spectral content, measured via spectral centroid (SC) and relative pitch (RP). Probes of increasing complexity are trained to predict each attribute from frozen embeddings across five datasets spanning noise, speech, monophonic musical notes, and music mixtures. Our primary finding is that all of these attributes are reliably recoverable from the CLAP embedding space across the examined datasets. Within this global picture, two encoding regimes emerge: RT60, LUFS, and RP are approximately linearly encoded, while SC requires non-linear probes. Both regimes generalize across eight additional audio foundation models, with the notable exception that amplitude-invariant architectures discard loudness entirely by construction. The identified linear feature directions are geometrically consistent across datasets for RT60 and LUFS, while highly domain-specific for RP. Finally, we provide a qualitative demonstration of cross-modal consistency, showing that text embeddings of acoustic descriptors align geometrically with the identified RT60 feature direction.

View source

Similar papers

Preprint Aug 2026

Listening Forward: Next Patch Embedding Prediction Enables Scalable Audio Learners

NAPE (Next-Audio-Patch-Embedding prediction), a self-supervised framework in which a causal Transformer predicts each next patch embedding of a log-mel spectrogram from the previous ones, using causal masking and stop-gradient as its sole training signal is introduced.

Umberto Cappellazzo, Xu-Bo Liu, Stavros Petridis et al. · 0 citations
#artificial intelligence Preprint Sep 2026

MADS: A Multiview Acoustic Descriptor Set Beyond Standard Spectral Summaries

Results establish MADS (Multi-view Acoustic Descriptor Set) not merely as a competitive standalone descriptor set, but as the foundational descriptor layer of a broader acoustically grounded representation program for future frame-level and deep-learning-compatible audio modeling.

Utsab Ghosh, Roshni Chakraborty · 0 citations
Preprint Sep 2026

Semantic Refinement of Universal Audio Representations through Audio-Description Alignment

Universal audio representations must preserve acoustic detail while making high-level concepts accessible across speech, music, environmental sound, and downstream models of different capacities. We study semantic refinement of an acoustically pretrained encoder by adding audio-description alignment to a foundation of...

Le-Jun Min, Jun-Yu Dai, Rui-Chen Zheng et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Heard but Not Heeded: Paralinguistic Information Encoding and Loss in Audio-Language Models

A mechanistic analysis of paralinguistic information in four open source models using the Expresso dataset with controlled speaking styles identifies a gap between what models encode and what they use, highlighting a key limitation in current audio language models.

Bhuvan Koduru, Dareen Alharthi, Rita Singh et al. · 2 citations · ⚡1
Preprint Aug 2026

Textual Acoustic Grounding for Generalizable LLM-Based Deepfake Voice Detection

Deepfake voice detection suffers from poor generalization across unseen domains. While Audio Large Language Models (ALLMs) show promise, the modality gap between continuous audio embeddings which capture the subtle acoustic details necessary for deepfake detection and the semantic space of LLMs remains a critical, unde...

Yassine El Kheir, Xin Wang, Wanying Ge et al. · 1 citation
Jul 2026

Content is What Remains: Invariant Speech Tokenization from Parallel Utterances

PINT (Parallel INvariant Tokenization), which fine-tunes an SSL encoder with alignment losses across parallel utterances and augmentations to distill this shared residual of linguistic content, and shows a 98.7% relative reduction in speaker probe accuracy.

Laurin Wagner, Bernhard Thallinger, Miroslav Stankovič et al. · 2 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.