Skip to content

Map as a Prompt: Learning Multi-Modal Spatial-Signal Foundation Models for Cross-scenario Wireless Localization

Jul 2026 · arXiv.org · Vol abs/2607.15713 · 1 citation · 19 references
Computer Science Engineering

TL;DR

SigMap is a multimodal foundation model that introduces two key innovations: a cycle-adaptive masking strategy that dynamically adjusts masking patterns based on channel periodicity characteristics to learn robust wireless representations and a novel"map-as-prompt"framework that integrates 3D geographic information through lightweight soft prompts for effective cross-scenario adaptation.

Abstract

Accurate and robust wireless localization is a critical enabler for emerging 5G/6G applications, including autonomous driving, extended reality, and smart manufacturing. Despite its importance, achieving precise localization across diverse environments remains challenging due to the complex nature of wireless signals and their sensitivity to environmental changes. Existing data-driven approaches often suffer from limited generalization capability, requiring extensive labeled data and struggling to adapt to new scenarios. To address these limitations, we propose SigMap, a multimodal foundation model that introduces two key innovations: (1) A cycle-adaptive masking strategy that dynamically adjusts masking patterns based on channel periodicity characteristics to learn robust wireless representations; (2) A novel"map-as-prompt"framework that integrates 3D geographic information through lightweight soft prompts for effective cross-scenario adaptation. Extensive experiments demonstrate that our model achieves state-of-the-art performance across multiple localization tasks while exhibiting strong zero-shot generalization in unseen environments, significantly outperforming both supervised and self-supervised baselines by considerable margins.

View source

Similar papers

Review Aug 2026

Foundation Models for Wireless Localization: Pretraining, Adaptation, and Utilization

Accurate wireless localization is a key enabler for 6G networks, yet remains challenging under diverse and rapidly changing propagation conditions. Model-based methods degrade when multipath channels are non-resolvable and model mismatches occur, while supervised deep learning demands large labeled datasets and generalizes poorly to new deployments. Inspired by foundation models (FMs) in language and vision, this article presents a unified framework for FM-based wireless localization that learns transferable channel representations from large-scale unlabeled channel state information and adapts to new environments with minimal or even no supervision. We review the fundamentals of FMs, compare the FM paradigm with existing localization approaches, and introduce a three-stage framework spanning large-scale pretraining, localization-oriented fine-tuning, and context-augmented inference, together with the location-aware applications it enables. Ray-tracing-based case studies show improved positioning accuracy and cross-environment generalization. Finally, we present an outlook on key research directions toward AI-native networks for wireless localization.

Guangjin Pan, Jiajia Guo, Zheng Xing et al. · 0 citations
#artificial intelligence Review Sep 2026

Wireless Foundation Models: State-of-the-Art and Open Challenges

Wireless foundation models (WFMs) have emerged as a promising approach for learning reusable representations from large-scale wireless data and adapting them to downstream tasks. However, the rapidly growing literature remains fragmented across modalities, pretraining objectives, architectures, adaptation strategies, and evaluation protocols, making it difficult to assess progress toward broadly transferable models. This survey provides a systematic analysis of WFMs for physical-layer applications. We first introduce the main WFM design components, including pretraining, backbone architectures, and downstream adaptation. We then organize the literature into five physical-layer task families: signal recognition and demodulation, channel representation learning, RF sensing and localization, beam management, and spectrum sensing and monitoring, while separately examining multi-task PHY models. Across these categories, we analyze how existing models are pretrained, adapted, and evaluated, with particular attention to downstream task diversity and the distinction between in-distribution, partial-shift, and out-of-distribution transfer. Our analysis shows that current WFMs provide increasing evidence of reusable wireless representations, but this evidence varies considerably across task families and evaluation settings. Differences in datasets, modalities, architectures, pretraining objectives, adaptation protocols, and distribution shifts make it difficult to determine which design choices drive transfer and generalization. We conclude by identifying open directions for improving data availability, evaluation rigor, generalization, efficient adaptation, and real-world deployment, providing a unified framework for understanding the current WFM landscape and the requirements for developing more reusable foundation models for future physical-layer wireless systems.

Alonso M. Pacheco Huachaca, J. J. Rodríguez Rodríguez, Ahmed Aboulfotouh et al. · 0 citations
Sep 2026

Environment-Aware Generalized Wireless Localization via Transformer

Wireless localization is expected to play a key role in future communication systems by providing location-aware services and supporting efficient network operation. However, existing deep learning (DL)-based localization methods often suffer from limited generalization when the deployment environment changes, since they tend to learn environment-specific propagation patterns. To address this issue, this article proposes an environment-aware generalized wireless localization framework that jointly exploits wireless channel, base station (BS) geometric information, and environmental information. Irregular city structures are represented by voxel-based occupancy maps, enabling explicit modeling of environmental factors that affect radio propagation. A transformer-based architecture is developed to comprehensively process wireless channel, geometric information of network nodes, and environmental information, thereby capturing the interaction between channel observations and surrounding urban structures. In addition, the proposed framework estimates a confidence map instead of directly regressing user equipment (UE) coordinates, which improves robustness under ambiguous propagation conditions. To support training and evaluation, we also develop an urban environment generator and a ray tracing-based channel simulator that produce large-scale datasets with physically consistent alignment between channels and 3-D environments. This framework enables systematic evaluation and robust localization in previously unseen urban environments.

Y. Noh, Kae Won Choi · 0 citations
Preprint Aug 2026

Electromagnetic World Model for 6G: A Unified Framework for Joint Environment Reconstruction and Channel Prediction

The integration of sensing, communication, and intelligence is becoming a key enabler for sixth generation (6G) wireless systems, where intelligent terminals are expected to simultaneously support efficient link establishment and reliable environmental sensing. However, existing studies mainly exploit sensing information or communication information to address a single task, such as channel prediction or environment reconstruction. Motivated by the shared dependence of optical and radio-frequency signals on the surrounding environment, we propose the electromagnetic world model (EMWM), the first unified framework for joint environment reconstruction and channel prediction. EMWM learns a common electromagnetic representation with the potential to provide a modeling foundation for 6G tasks. Specifically, partial channel state information (CSI) and multi-view red-green-blue (RGB) images are encoded into CSI and visual tokens and jointly processed by a hierarchical world-model backbone with local and global aggregation. Based on the learned representation, a mixture-of-experts (MoE)-based CSI prediction head reconstructs the complete CSI, while a depth prediction head estimates multi-view depth maps that are further converted into three-dimensional (3D) point clouds. Moreover, a large-scale multi-modal dataset is constructed based on a campus digital twin. Experimental results show that EMWM outperforms conventional neural network and large language model (LLM) baselines in both CSI prediction and environment reconstruction, achieving a squared generalized cosine similarity (SGCS) of 0.9699 for CSI prediction while demonstrating robustness across different signal-to-noise ratio (SNR) conditions and zero-shot generalization at 28 GHz.

Yizhu Zhao, Li Yu, Jian-Hua Zhang et al. · 0 citations
#artificial intelligence Preprint Sep 2026

IIns-VAE+: A Robust Transfer Learning Framework for Environmental Identification in Wireless Sensing

Environmental identification in wireless sensing is essential for 6G integrated sensing and communication (ISAC) systems to achieve reliable situational awareness. However, deep learning (DL) models for this task often fail to generalize under domain shift across diverse environments. While the Inter-Instance Variational Auto-encoder (IIns-VAE) learns features of rich representation, its neural classifier remains vulnerable to these distribution changes. In this paper, we propose IIns-VAE+, a hybrid model that combines the IIns-VAE framework with Minimax Risk Classifiers (MRC) to improve adaptability in transfer learning scenarios. We use real-world datasets to evaluate our framework across three transfer learning scenarios, including general to specific room environments, high to low label resolutions, and mixed to specific environments. The experimental results indicate that IIns-VAE+ significantly outperforms baselines, demonstrating its critical value in building adaptable and robust perceptive networks in future 6G systems.

Yu-Xiao Li, K. Hu, Bo-Bai Zhao et al. · 0 citations
#small language model Open access Sep 2026

Reconstructing wireless signals for low altitude networks using small language models

Wireless signal reconstruction is essential for RF-based positioning in GPS-denied environments. However, multipath propagation, shadowing, and non-Gaussian noise complicate this, and traditional methods require extensive site-specific calibration that precludes rapid deployment. We present In-Context Signal Completion (ICSC), demonstrating that small language models fine-tuned with Group Relative Policy Optimization and physics-informed rewards can reconstruct RSSI across sequential extrapolation and spatial interpolation tasks. Our 0.5B-parameter model attains 55% recall within 2 dB and a 2.85 dB mean absolute error on sequential prediction. This achieves a 49% error reduction over the untrained baseline, performing on par with GPT-4o (51%) with fewer parameters. Successful zero-shot transfer to spatial interpolation indicates the model acquires transferable physical reasoning rather than task-specific memorization. Operating at 3 ms latency for real-time edge inference, ICSC reduces deployment from weeks of per-site data collection to immediate inference using sequential context.

Xin Li, Ran Liu, Chau Yuen · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.