Aug 2026· Electronics· Vol 15, pp. 3415· 0 citations· 11 references
TL;DR
A novel steganographic framework based on generative artificial intelligence and predefined semantic mapping that yields graceful degradation under common channel distortions together with high visual camouflage, while also revealing that recovery reliability decreases as more semantic dimensions are activated simultaneously.
Abstract
With the advancement of deep learning-based steganalysis and prevalence of lossy compression mechanisms in social network transmission, traditional steganography based on cover modification (such as LSB substitution) faces dual challenges of security and robustness. This study proposes a novel steganographic framework based on generative artificial intelligence and predefined semantic mapping. Unlike embedding ciphertext in pixel noise, this method utilizes a shared mapping protocol (Codebook) to transform abstract information into concrete visual elements (such as characters, actions, scenes, and styles), and constructs stego-images through generative models. Experimental results show that when both parties share the same key, the system achieves full semantic recovery. Using an explicitly specified pipeline (Gemini 2.5 Flash Image for synthesis and Gemini 2.5 Flash for parsing), we evaluate the scheme on an enlarged, randomly sampled scenario set spanning three to six active semantic dimensions. Across these scenarios, we report per-dimension accuracy, the end-to-end full-recovery rate with 95% confidence intervals, and the partial-recovery rate under controlled JPEG compression, re-scaling, Gaussian noise, and cropping, rather than a single aggregate figure. The results indicate that carrying information at the semantic level yields graceful degradation under common channel distortions together with high visual camouflage, while also revealing that recovery reliability decreases as more semantic dimensions are activated simultaneously. We therefore present these findings as a proof of concept and explicitly separate demonstrated results from hypotheses left to future work. We therefore present these findings as a proof of concept and explicitly separate demonstrated results from hypotheses left to future work.
Deep learning-based video steganography has made significant strides, yet conventional explicit methods often suffer from cover distortion and reduced extraction accuracy at high capacities. In this paper, we propose an implicit video steganography framework that treats video hiding and recovery as a dual-stream generation process leveraging implicit neural representations. Instead of altering existing carriers, secret information is encoded within the neural network’s weights, making it an inherent part of the generation process. We introduce a dual-stream input encoding mechanism that decouples the input space into temporal and cryptographic encodings to ensure covert transmission, allowing only authorized receivers to recover hidden content. Furthermore, a multi-scale generation network, incorporating frequency-aware upscaling and statistical distribution loss, is presented to achieve high-quality reconstruction. Extensive experiments demonstrate that our approach achieves state-of-the-art results, minimizing detectable discrepancies while concealing up to seven secret videos within a single carrier. Our method significantly outperforms existing benchmarks by a margin of over 10 dB in peak signal-to-noise ratio, highlighting its superior imperceptibility, accuracy, and security.
Yifei Wang, Gaozhi Liu, Sheng Li et al.· Computer/law journal· 0 citations
Diffusion-based generative image steganography enables covert communication by synthesizing stego images without relying on cover images. However, existing latent-space methods still struggle to balance robustness, steganographic security, and visual fidelity, especially under practical channel distortions such as compression, blur, resizing, and noise. To address these challenges, we propose RIS-MoE, a robust and secure latent-space image steganography framework that integrates distortion-tolerant message representation with receiver-side adaptive latent restoration. At the sender side, a learnable orthogonal transformation converts the secret message into a distributed representation, which is embedded into the diffusion latent through a residual-guided Hide Network. At the receiver side, a plug-and-play Mixture-of-Experts (MoE) denoising module estimates the distortion composition and adaptively fuses specialized restoration experts before message extraction. Extensive experiments show that RIS-MoE achieves strong robustness under single, mixed, and real-world distortions. It maintains extraction accuracy above 90% under all evaluated simulated combined distortions and achieves 94.62% and 95.29% extraction accuracy after real-world Weibo and Instagram transmission, respectively. RIS-MoE also achieves competitive empirical resistance against spatial-domain, latent-domain, and diffusion-aware steganalyzers, while maintaining favorable visual quality with an FID of 7.35 and an LPIPS of 0.21 on Flickr8K. In addition, the proposed MoE module consistently improves representative latent-space steganography pipelines as a plug-and-play restoration component, demonstrating its transferability. The source code is publicly available at: https://github.com/angle-cell/RIS_MOE.
Gen-Fan Yang, Rong-Chang Duan, Hong Zhang et al.· Cybersecurity· 0 citations
This survey presents a structured review of deep learning-based techniques for image data hiding, proposing a three-paradigm taxonomy organized by the method’s operational relationship to the carrier image. We classify existing methods into modification-based, synthesis-based, and logic-based approaches. In the modification-based tier, we trace the architectural progression from foundational Convolutional Neural Networks and Generative Adversarial Networks to high-capacity Invertible Neural Networks and Transformers, analyzing their distinct trade-offs between embedding capacity, imperceptibility, and robustness. In the synthesis-based tier, we examine how Diffusion Probabilistic Models and generative adversarial frameworks reframe data hiding as a carrier generation problem rather than a pixel editing task. This paradigm encompasses both generative steganography (where carriers are synthesized from scratch) and proactive watermarking (where provenance is embedded during AI content generation). In the logic-based tier, we review zero-watermarking and coverless steganography, where ownership is established through feature extraction and semantic mapping without modifying any image, a critical property for sensitive domains such as medical imaging. Finally, we identify four persistent infrastructure gaps: benchmarking fragmentation, narrow robustness evaluation, domain generalization failures, and computational infeasibility that prevent real-world deployment despite architectural progress, and we propose concrete research directions.
Matúš Janok, Radoslav Forgáč, Ladislav Hluchý· Journal of Imaging· 0 citations
Digital image steganography aims to imperceptibly embed secret information into a cover image to enable covert communication. This paper focuses on image-level imperceptibility and recovery quality, and proposes a cost-guided joint mask–perturbation optimization with attentive decoding for image steganography method (CMAD). In an end-to-end differentiable framework, CMAD jointly optimizes the embedding mask, perturbation magnitude, and decoding-network parameters, thereby improving recovery accuracy while preserving imperceptibility. During optimization, the proposed AniCost cost-map guidance mechanism computes pixel-level embedding costs through wavelet-based anisotropy analysis, and uses a probability-map guidance loss to directly encourage the mask to activate in complex-texture regions and deactivate in smooth regions. The channel-attention-based decoding network is fine-tuned for each image pair during optimization to adapt to the current pair. Experimental results show that the stego images generated by CMAD achieve PSNR values of 55–59 dB, while the recovered secret images achieve PSNR values of 35–40 dB.
VisionStego is proposed, a duallayer security architecture that pairs symmetric-key encryption with an artificial-intelligence-guided steganographic embedding stage, so that cloud-hosted data is protected in both substance and appearance.
R. Saxena, Priti Maheshwary· International Journal for Re...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.