Robust Joint Sensing and Task Inference for Resource-Constrained IoT Agents
Reliable intelligent sensing over bandwidthlimited and noise-impaired wireless links is a key challenge for resource-constrained Internet-of-Things (IoT) agents. Existing semantic communication methods mainly optimize either image reconstruction or task inference, which limits their use in joint sensing scenarios under severe compression. This paper proposes R-SemCom, a robust joint sensing and task inference framework for resourceconstrained IoT agents. R-SemCom employs a heterogeneous dual-stream encoder with a Convolutional Neural Network (CNN) branch for local structural modeling and a Vision Transformer (ViT) branch for global semantic representation. An orthogonality-constrained decoupling mechanism is introduced to improve latent-space efficiency, while a serial reconstruction-guided inference strategy is designed to enhance task robustness under noisy channels. Experimental results on the CIFAR-10 dataset show that, at a compression ratio of 1/12 and under lowSNR conditions ranging from -5 dB to 5 dB, R-SemCom achieves the best overall performance among the compared methods and provides a more favorable trade-off between reconstruction quality and recognition accuracy.