Results show that the original full-resolution LAM is the strongest, that separately trained lightweight models are the most competitive learned approaches, and that representation alignment between the upsampler and LAM matters more than model complexity.
Abstract
Latent Acoustic Mapping (LAM) is a self-supervised learning method that generates high-resolution spherical acoustic maps from multichannel recordings without labelled data, matching supervised baselines on direction-of-arrival benchmarks. However, LAM degrades significantly with sparse 4-channel arrays, as the low-resolution cross-spectral matrix captures far less spatial information than the 32-channel inputs LAM was designed for. We benchmark a diverse set of upsampling architectures, spanning lightweight convolutional networks, iterative back-projection models, physics-informed networks, and generative adversarial approaches. We also study whether aligning these upsamplers with LAM by training them jointly or in different stages helps preserve the spatial structure that LAM depends on. Results show that the original full-resolution LAM is the strongest, that separately trained lightweight models are the most competitive learned approaches, and that representation alignment between the upsampler and LAM matters more than model complexity.
Audio-Text Foundation Models (ATMs) fail catastrophically under severe acoustic noise, yet existing adaptation strategies either rely on gradient-based Test-Time Adaptation (TTA), which reinforces noise rather than signal, or on prompt tuning that requires privileged noise annotations unavailable at inference. We address these failures with PRISM (Prototype-Rectified Iterative Self-supervised Manifold Denoising), a training-free, source-free TTA framework grounded in the Affine Noise Hypothesis: severe acoustic noise induces a low-rank affine shift in the multimodal latent space, with more than 90% of distortion energy confined to the leading 60 principal components. PRISM estimates and reverses this distortion from an unlabeled target batch using frozen text prototypes as geometric anchors via three closed-form geometric corrections compiled into a single static projection matrix by Affine Bias Regression. At inference, adaptation reduces to one matrix-vector multiplication in 0.0009 ms, making it substantially faster than gradient-based TTA while requiring no additional training. On UrbanSound8K, PRISM improves over the zero-shot baseline by 12.94 percentage points and surpasses an oracle-assisted TTA baseline by 9.41 percentage points, despite never observing its privileged augmented noise prompts. We further identify the Polyphonic Trap, a principled failure mode of subspace deflation for broadband classes, and resolve it via Confidence-Aware Regression (CAR), recovering up to 8.16 percentage points for the worst-affected class.
A. Shukla, R. Thakur, Aryan Das et al.· 0 citations
We present an amortised neural framework for three-dimensional acoustic diffraction tomography that reconstructs scene geometry as signed distance functions from boundary element method (BEM) pressure observations. Our central finding is that domain-specific encoder design, rather than physics-informed loss terms, drives reconstruction accuracy. Building on a transfer function decomposition that separates scattering from free-space propagation, we introduce an encoder that maps complex pressure measurements to latent geometry codes through magnitude–phase input representation and frequency-selective attention pooling. Across 100 synthetic scenes spanning 12 shape categories (3 random seeds), this encoder attains 0.644±0.015 mean intersection over union, a 51.0% improvement over a domain-agnostic baseline ( p=0.0015), while the physics losses evaluated in this amortised setting provide zero or negative benefit. Inference is a single forward pass ( ∼127 ms, a ∼950× speedup over per-scene optimisation). A systematic Eikonal regularisation ablation reveals a topology- and architecture-dependent asymmetry: any benefit is confined to specific convex shapes with reconstruction headroom in the per-scene auto-decoder and vanishes under encoder amortisation, whereas the harm—a flattening of topological features—is universal across architectures. Combined with Helmholtz partial differential equation failure and Laplacian supervision dead-ends, these findings challenge the assumption that physics losses universally benefit neural acoustic surrogates and suggest that architectural physics integration is more effective than loss-based approaches. We will release a 100-scene 3D BEM benchmark dataset, deterministically regenerable from the accompanying code, for reproducibility.
Ju O Kim, Deokwoo Lee· Measurement science and tech...· 0 citations
This work proposes Normalizing Autoencoder (NAE), which employs a novel conditional loss that aligns the surrogate loss gradient with that of reconstruction loss, directly improving upon the current standard.
The proliferation of multimicrophone IoT devices—such as smartphones, smart speakers, and wearables—has enabled diverse acoustic sensing applications. However, scarcely labeled data and poor cross-task transferability continue to limit traditional supervised methods. To address this, we propose acoustic-URL, an unsupervised representation learning framework using cross-domain fusion (cdf) and cross-channel contrastive learning (CCCL) to extract transferable features from unlabeled multichannel audio. Our approach integrates time-domain waveforms and spatial-domain channel impulse responses (CIRs) to capture both temporal and spatial information. By aligning features from different microphones that capture the same acoustic event, the model learns spatially consistent representations. Evaluations on four tasks—digit recognition, writer identification, gesture recognition, and footstep-based identification—under varying supervision and task-domain shifts show that acoustic-URL consistently outperforms supervised baselines, especially in low-label settings, and demonstrates strong cross-task generalization.
Bingzhi Wang, Yongzhao Zhang, Jiajun Yu et al.· IEEE Internet of Things Jour...· 1 citation
We investigate how the realism of synthetic room impulse response (RIR) datasets affects the training of DeepFilterNet3 for single-channel speech enhancement. We compare a DNS4 image-source-method (ISM) RIR dataset with a higher-acoustic-fidelity dataset generated using hybrid wave-based and geometrical acoustics simulation. Rather than isolating individual simulation factors, we compare complete RIR generation pipelines while keeping the enhancement model unchanged. Models are evaluated on unseen measured RIRs using objective speech enhancement metrics and downstream automatic speech recognition (ASR). Training with the higher-fidelity dataset consistently yields modest improvements in objective metrics and substantially lower ASR word error rates than the ISM dataset. Although the experiments do not attribute these gains to individual modelling components, they show that increasing the overall realism of synthetic acoustic training data improves the generalization of DeepFilterNet3 to unseen measured environments.
A. Milo, Georg Götz, Steinar Guðjónsson et al.· 0 citations
Pre-trained Vision Foundation Models (VFMs) provide strong visual representations for diverse downstream tasks. The key challenge of VFM adaptation stems from the prohibitive costs of full fine-tuning and catastrophic forgetting. To address this, Low-Rank Adaptation (LoRA) has emerged as the prevailing paradigm for Parameter-Efficient Fine-Tuning (PEFT). However, LoRA is typically designed for transformer self-attention layers parameterized by 2D matrices. Since convolutional kernels inherently couple spatial and channel information within a 4D tensor, forcing them into a monolithic 2D matrix disrupts the inherent spatial topology. In this paper, we propose Low-Rank Convolutional Adaptation (LoCA), a convolution-aware PEFT framework that addresses spatial-channel entanglement by decoupling channel and spatial adaptation. LoCA introduces a low-rank channel adaptation for dense cross-channel mixing and refines spatial bases extracted from pre-trained kernels via Singular Value Decomposition (SVD). Experimental results show that LoCA preserves pre-trained spatial priors and achieves competitive or state-of-the-art performance across fine-grained classification, domain-generalized semantic segmentation, and generative benchmarks.
Sojung An, Junha Lee, Sujeong You et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.