Skip to content

RadioDiff-v2: Generative Angular Radio Maps for Multi-Beam Selection and Localization

Jul 2026 · arXiv.org · Vol abs/2607.08045 · 0 citations · 46 references
Computer Science Engineering Mathematics

TL;DR

RadioDiff-v2, a dual-branch one-dimensional diffusion transformer trained with flow matching, is proposed, a dual-branch one-dimensional diffusion transformer trained with flow matching that leads every baseline on every metric.

Abstract

Angular radio maps describe the received-power distribution over the angle of arrival and underpin beam selection and receiver localization in sixth-generation (6G) networks. Predicting the angular power spectrum (APS) from geometry is difficult, because the mapping is ill-posed in non-line-of-sight (NLOS) conditions and must generalize to unseen environments. Distortion-minimizing regressors return the conditional mean, which over-smooths the spectrum and erases the multipath structure that downstream tasks need. We cast the task as a perception-distortion problem and propose RadioDiff-v2, a dual-branch one-dimensional diffusion transformer trained with flow matching. It couples periodic angular encoding, adaptive layer-normalization conditioning, a Fourier angular mixer, and joint velocity and clean-signal heads. A per-metric estimator portfolio reads every deployment quantity from this single model, so that samples carry the distribution, the clean-signal head supplies a regression-grade point estimate, Bayes-optimal rules select beams, and the conditional likelihood localizes the receiver. We prove that a concentrated conditional yields a straight probability-flow trajectory that one step integrates exactly, identifying deterministic transport as the correct inductive bias. On a zero-shot test of 99 environments and one million links, RadioDiff-v2 leads every baseline on every metric, with a 0.39 dB Wasserstein-1 distance, per-bin error below the regression baseline, a 2.43 dB eight-beam NLOS sweep loss, and a 20.6-pixel localization error with four base stations. Code is available at https://github.com/UNIC-Lab/RadioDiff-v2.

View source

Similar papers

#small language model Open access Sep 2026

Reconstructing wireless signals for low altitude networks using small language models

Wireless signal reconstruction is essential for RF-based positioning in GPS-denied environments. However, multipath propagation, shadowing, and non-Gaussian noise complicate this, and traditional methods require extensive site-specific calibration that precludes rapid deployment. We present In-Context Signal Completion (ICSC), demonstrating that small language models fine-tuned with Group Relative Policy Optimization and physics-informed rewards can reconstruct RSSI across sequential extrapolation and spatial interpolation tasks. Our 0.5B-parameter model attains 55% recall within 2 dB and a 2.85 dB mean absolute error on sequential prediction. This achieves a 49% error reduction over the untrained baseline, performing on par with GPT-4o (51%) with fewer parameters. Successful zero-shot transfer to spatial interpolation indicates the model acquires transferable physical reasoning rather than task-specific memorization. Operating at 3 ms latency for real-time edge inference, ICSC reduces deployment from weeks of per-site data collection to immediate inference using sequential context.

Xin Li, Ran Liu, Chau Yuen · 0 citations
Preprint Aug 2026

Physics-informed VAE-EVT for Tail Aware Radio Map Prediction

This work introduces a physics- and tail-informed VAE-EVT (variational autoencoder-extreme value theory) framework that distinctly models both the bulk and tail distribution of SNR, and significantly outperforms the state-of-the-art GAN-based model.

A. Gamage, Niloofar Mehrnia, James Gross · 0 citations
Preprint Aug 2026

Physics-Guided Neural Airy Beamforming for Near-Field Blockage Mitigation

High-frequency communication systems heavily rely on line-of-sight(LoS) paths, so blockage of the LoS path can cause severe performance loss. Near-field Airy beams with curved trajectories can steer energy around obstacles, offering a promising solution for blockage mitigation. However, existing methods for selecting a near-optimal Airy beam trajectory either rely on high-overhead beam training, or employ data-driven learning without a clear, physically interpretable rule. To address this problem, we propose a physics-guided neural Airy beamforming framework that selects a near-optimal trajectory in one shot with clear physical interpretability. Specifically, we first formulate a single-edge representation of the blocker model in 3GPP TR 38.901 and reveal the trajectory--edge coupling mechanism. This analysis yields a trajectory-selection optimality condition that defines the candidate trajectories. Although these trajectories generally cannot be expressed in closed form, we show that they form a continuous structure. This continuous structure is then exploited to construct a compact physics-defined region that captures near-optimal trajectories. Guided by this region, a lightweight neural predictor is finally designed to directly select a near-optimal trajectory without beam training. Simulations show that the compact physics-defined region effectively captures near-optimal trajectories, while occupying only about 6% of the candidate-space area on average. The proposed framework retains 99.7% of the reference rate obtained through numerical optimization, and nearly matches the rate of the data-driven method despite using approximately 112\times fewer neural-network parameters.

Yi Wang, Linglong Dai · 1 citation
Aug 2026

A dual-branch fusion network for footstep sound source localization in non-line-of-sight corridors.

Non-line-of-sight (NLOS) acoustic source localization using microphone arrays has attracted increasing attention in robotics due to its non-intrusive nature, low cost, and ease of integration. However, existing methods typically rely on accurate environmental models and ideal propagation assumptions, resulting in limited robustness in complex multipath environments. To address these issues, this paper proposes CorridorLocNet, a dual-branch fusion network for footstep sound source localization in NLOS corridors. Specifically, Mel-spectrogram and generalized cross correlation with phase transform features are first concatenated into a joint input. Subsequently, a dual-branch architecture is proposed, comprising a residual convolutional branch to extract local time-frequency patterns and a lightweight Conformer branch to capture global temporal dependencies. Furthermore, a cross-attention module adaptively fuses high-dimensional representations from both branches. Finally, a multi-layer perceptron outputs the estimated position. By learning the complex mapping from spatial sound field variations to source positions, the proposed method provides a robust solution in NLOS corridors. Additionally, a real-world dataset of footsteps occurring behind a corridor corner is constructed for evaluation. Experimental results demonstrate that CorridorLocNet achieves 98.83% classification accuracy and reduces the average error by 2.56 m compared with the reflection-aware localization approach, validating its feasibility in NLOS localization scenarios.

Xiaonan Wang, Zhe Chen, F. Yin · 0 citations
Jul 2026

Task-guided line-of-sight power-ratio calibration for bias-limited visible light positioning.

Channel impulse response (CIR)-assisted visible light positioning (VLP) improves indoor ranging by separating the line-of-sight (LoS) contribution from multipath-contaminated observations. Practical CIR processing, however, does not provide an ideal LoS measurement: delay-bin aggregation, path-support errors, tap-estimation noise, and residual non-line-of-sight leakage introduce a structured bias in the estimated LoS power ratio (LPR). This Letter formulates bias-limited VLP as a task-guided measurement-calibration problem. Instead of replacing the localization backend with coordinate regression, the proposed task-guided LoS power-ratio calibration (T-LPRCal) framework learns a bounded correction of the CIR-derived LPR before range estimation. A first-order propagation model relates LPR bias to range and position errors, while a compact recurrent classifier selects the correction through position-supervised labels generated offline by the physical localization pipeline. Simulations in a representative indoor optical channel show that T-LPRCal reduces the bias floor in the high signal-to-noise ratio (SNR) regime, improves median and 90th-percentile position errors, and preserves explicit Lambertian ranging and a least-squares (LS) localization backend.

Pablo Palacios Játiva, Iván Sánchez Salazar, Ismael Soto et al. · 0 citations
Preprint Aug 2026

ControlRadio: Prompt-Driven Controllable Diffusion for Cross-Modal Radio Map Generation

ControlRadio is presented, a controllable generative framework that produces radio maps from natural-language descriptions and environmental layouts, including building structures and transmitter locations, while reducing computation time by more than four orders of magnitude compared with conventional simulation-based methods.

Kangjun Liu, Xiying Pan, Shuhang Zhang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.