MBCE: Multi-Band Contrastive Encoding for Frequency-Aware Ultrasound Representation Learning
Abstract
Ultrasound imaging relies on high-frequency sound waves to visualize internal anatomy in a non-invasive manner, yet automated analysis remains difficult due to low signal-to-noise ratios, speckle artifacts, and inconsistencies across scanners. These challenges are compounded by the scarcity of annotated datasets, limiting feature extraction and generalization across anatomical regions. Conventional self-supervised learning (SSL) models, mostly developed for natural images, often overlook spectral characteristics that are important in ultrasound and may therefore learn representations that are sensitive to scanner- and acquisition-dependent variations. To address this gap, we propose Multi-Band Contrastive Encoding (MBCE), a spectral contrastive learning framework designed to learn frequency-aware ultrasound representations. MBCE extracts spatial features while incorporating log-magnitude Fourier sub-band features through a shared ResNet-50 backbone enhanced with lightweight cross-attention, enabling joint modeling of spatial and spectral information without separate encoders for each view. We further introduce Spectral-Band Contrastive Loss (SBCL), which aligns low-, mid-, and high-frequency embeddings from the same image while contrasting them against embeddings from other images. Experiments on breast, liver, and pancreas ultrasound datasets show consistent improvements over representative SSL baselines, including cross-organ evaluation on an unseen liver dataset. By explicitly leveraging spectral structure, our approach offers a label-efficient direction for ultrasound representation learning and motivates future evaluation on tasks such as segmentation and detection with larger data corpora. The source code and weights for MBCE are available at: https://github.com/dsatyam09/Multi-Band-Contrastive-Encoder.git