This work proposes Physics-Unrolled Hybrid Neural Operator (PU-HNO), a three-stage cascade that predicts high-fidelity indoor radio maps from low-fidelity ray-tracing outputs and scene priors by progressively capturing reflection, diffraction, and scattering effects, rather than treating radio maps as generic images.
Abstract
Radio maps are essential for wireless decision-making tasks such as access-point placement, coverage planning, and localization, but their fine spatial details are governed by complex propagation effects and are costly to simulate accurately. Machine learning offers a path to high-fidelity radio-map prediction without running expensive high-fidelity simulations for every scene. However, generating high-quality training labels at scale is also difficult: the affordable labels come from finite-ray simulations, which are richer than low-fidelity inputs but carry residual Monte Carlo noise. We address this challenge with Physics-Unrolled Hybrid Neural Operator (PU-HNO), a three-stage cascade that predicts high-fidelity indoor radio maps from low-fidelity ray-tracing outputs and scene priors by progressively capturing reflection, diffraction, and scattering effects, rather than treating radio maps as generic images. We prove that, under conditionally unbiased label noise, the model can learn stable propagation structure and outperform its own training labels. Experiments across diverse floorplans show that PU-HNO outperforms image-to-image baselines, wireless learning models, and monolithic neural operators across both image-quality and wireless deployment metrics.
Radio maps describe how wireless signals propagate across space and are essential for wireless communication, sensing, and network planning. However, constructing accurate radio maps traditionally requires either dense measurements or computationally expensive physical simulations, which limits scalability and real-time deployment. Recent advances in generative artificial intelligence offer a promising alternative, but existing approaches lack fine-grained control and physical consistency when applied to real-world wireless environments. Here we present \textbf{ControlRadio}, a controllable generative framework that produces radio maps from natural-language descriptions and environmental layouts, including building structures and transmitter locations. Joint semantic and spatial conditioning enables interpretable, propagation-plausible generation, while a controlled latent prior and layout-aware conditioning improve stability and structural consistency. Extensive experiments demonstrate that ControlRadio achieves state-of-the-art accuracy and strong generalization across diverse urban scenarios, while reducing computation time by more than four orders of magnitude compared with conventional simulation-based methods. Such results suggest a new paradigm for scalable and controllable wireless environment modeling, with broad implications for next-generation communication systems and data-driven radio sensing.
Kangjun Liu, Xiying Pan, Shuhang Zhang et al.· 0 citations
Coverage map estimation is fundamental in planning wireless communication systems to enhance service quality and minimize operational costs. Although traditional deterministic propagation models offer high precision, their computational costs and processing times increase exponentially, especially in urban areas. In this study, a ResNet-based Conditional Variational Autoencoder (ResNet-CVAE) architecture is proposed for rapid and accurate coverage map generation in scenarios involving multiple diffractions. In order to generate dataset, a dynamic ray-tracing code integrating Geometric Optics (GO) and the Uniform Theory of Diffraction (UTD) was developed. By processing obstacle geometry and transmitter locations as both numerical and spatial condition information, the proposed deep learning model successfully captures abrupt signal level drops and physical shadowing effects behind obstacles. Experimental results demonstrate that the ResNet-CVAE model produces high accurate results with significantly lower computational overhead compared to traditional methods and adapts effectively to complex obstacle configurations. This approach offers significant potential for real-time analysis in network planning and base station placement processes.
Uğur Erbaş, Bahtiyar Bayram, Adem Avcı et al.· European Conference on Artif...· 0 citations
Radio map construction aims to infer dense received-power fields from environmental layouts, sparse observations, and physical priors. While crucial for environment-aware wireless systems, it remains challenging in complex urban scenes with building-induced non-line-of-sight (NLOS) shadows and dynamic blockages. Non-iterative methods, such as interpolation techniques, RadioUNet, and RME-GAN, often struggle to accurately model these complex obstruction effects. Conversely, iterative generative methods tend to produce physically implausible hallucinations in shadowed or strongly obstructed areas. To address these issues, we propose LSK-RM, a physics-guided Large Selective Kernel U-Net for one-stage dense radio map reconstruction. Specifically, an LSK-based encoder-decoder is introduced to adaptively aggregate local shadow-boundary details and long-range attenuation context within a single forward pass. Furthermore, we develop a multi-source physical prior representation that fuses environmental geometry, sparse measurements, and fast ray-tracing visibility cues. To suppress physically implausible energy leakage, we design a novel logarithmic physics-guided objective combining pixel-wise supervision with Laplacian and obstacle-boundary consistency. Experiments on the RadioMapSeer dynamic blockage dataset demonstrate that LSK-RM outperforms representative baselines, including RadioUNet, RME-GAN, RMDM, and RadioFlow. Notably, it achieves higher accuracy across quantitative metrics such as NMSE, and exhibits significantly better modeling performance in diffraction transitions and shadow regions.
Zhengyan Liao, Weidong Zou, Chunlei Wang et al.· 2026 8th International Confe...· 0 citations
High-fidelity simulation of mmWave radar signals for dynamic human motion is valuable for developing radar-based human sensing models; yet collecting accurately labeled measurements for a specific deployment site remains expensive. We present HybridSim, a physics-learning hybrid simulator that synthesizes mmWave radar signals from dynamic human meshes under a fixed indoor room configuration, explicitly decoupling propagation into two components. To parameterize the human subject, we use a tri-plane representation to extract human features and a Graph Convolutional Network to stabilize optimization and mitigate gradient instability. The direct signal path is modeled via an inverse-rendering formulation with a microfacet BRDF to capture primary surface reflections. In parallel, the indirect path is approximated by combining 3D Gaussian Splatting with a virtual-receiver geometry to fit and reproduce site-specific multipath interference patterns, achieving substantially lower computational cost than explicit full ray tracing. Experiments in a fixed-room setting show improved agreement with a physically based reference and consistent gains on downstream radar-based human sensing tasks when using HybridSim for site-specific data augmentation.
Weitao Xiong, Tianyu Liu, Peng Li et al.· 0 citations
Ultra-reliable low-latency communication (URLLC) requires precise identification of spatial regions where the signal-to-noise ratio (SNR) falls below an outage threshold. In this context, an outage refers to instances in which SNR falls below a specified threshold, which, for URLLC, can be as stringent as the 0.1% quantile of the SNR distribution. Traditional generative radio map models tend to focus on reconstructing average signal levels, often overlooking the low SNR that is crucial for accurate outage prediction. To address this limitation, we introduce a physics- and tail-informed VAE-EVT (variational autoencoder-extreme value theory) framework that distinctly models both the bulk and tail distribution of SNR. Our approach begins with a physics-informed preprocessing stage that extracts deterministic features, including line-of-sight, shadowing, and distance, from the scene geometry. A dual-latent encoder then captures the bulk SNR using a Gaussian mixture and the tail using a generalized Pareto distribution (GPD). By employing a modified variational objective, the model is trained to jointly supervise both regimes, ensuring focused attention on extreme fading events. Evaluated on the RadioMapSeer dataset, our method achieves an SNR RMSE of 4.83 dB in the outage region defined by the low threshold of 0.1% SNR quantile. This significantly outperforms the state-of-the-art GAN-based model, which records an SNR RMSE of 21.90 dB, with the performance gap widening as the outage threshold becomes more stringent.
A. Gamage, Niloofar Mehrnia, James Gross· 0 citations
Millimeter-wave (mmWave) radar enables privacy-preserving and illumination-robust human motion reconstruction, but training generalizable models typically requires costly paired radar-motion recordings. Simulation can scale such supervision, yet even physics-based simulators cannot fully reproduce real-world multipath, clutter, hardware-specific response statistics, or distance-dependent resolution degradation, leaving a sim-to-real gap. We present mmSimPrior, a simulation-pretrained framework that factorizes transferable knowledge into signal, motion, and radar-to-motion mapping priors. To learn transferable signal and motion priors, we pretrain a multimodal radar encoder with a physics-informed domain-randomization curriculum designed to mitigate the sim-to-real gap by approximating real-world propagation- and acquisition-level variations, while a joint-temporal tokenizer learns a discrete prior over plausible human motion. A dual-mode mapping module predicts either motion-code distributions for structurally constrained zero-shot reconstruction or continuous motion parameters for flexible adaptation from limited real data. We further construct a 4.2M-frame, 31K-sequence dataset suite and introduce a No-Overlap Setting that prevents any exact subject-environment-location-motion tuple from appearing in both the adaptation and test sets. Experiments on mmSimPrior-Real and RT-Pose demonstrate consistent gains: with only 24 paired real sequences, mmSimPrior-Reg reduces MPJPE by 24.7-39.0% over the strongest baseline across the three environments, while mmSimPrior-Cls reduces zero-shot MPJPE by 8.5% without fine-tuning.
Cheng Guo, Qiming Cao, Shengkai Xu et al.· 0 citations