Jul 2026· International Conference Computing Methodologies and Communication· pp. 1163-1168· 0 citations· 19 references
Abstract
Image synthesis has become a central problem in generative AI, with applications spanning virtual reality, medical imaging, autonomous systems, and creative content generation. Diffusion-based generative models have substantially advanced the field by producing high-fidelity, visually consistent outputs, yet a fundamental tension remains: hybrid synthesis tasks require photorealism and geometric consistency to be achieved together, so that perceptual quality does not come at the cost of structural accuracy. This review surveys recent progress in conditional diffusion frameworks, with particular attention to geometry-aware conditioning mechanisms that connect appearance and structure. Our contributions are a taxonomy of 18 conditioning frameworks, a six-parameter controllability analysis (Table II), formal definitions of the evaluation metrics most commonly reported in the literature (FID, SSIM, LPIPS, IS, Precision, Recall, Dice) with a consolidated reference table, an explicit architectural comparison of DDPM, DDIM, and LDM together with transformer-based generative models (DiT, VQGAN+Transformer), a structured comparison of diffusion models and GANs across deployment-relevant criteria, and a discussion of open challenges that includes the limitations of current experimental validation practice.
Generative diffusion models have emerged as a class of powerful techniques for various imaging applications, including but not limited to synthesis, reconstruction, and segmen- tation. Despite their success, current generative models pose two key limitations. First, they primarily rely on image intensity and texture in...
Nian Wu, Nivetha Jayakumar, Jiarui Xing et al.· Machine Learning for Biomedi...· 0 citations
Generative diffusion models have emerged as a class of powerful techniques for various imaging applications, including but not limited to synthesis, reconstruction, and segmentation. Despite their success, current generative models pose two key limitations. First, they primarily rely on image intensity and texture info...
Nian Wu, Nivetha Jayakumar, Jiarui Xing et al.· 0 citations
SpatialCrafter is presented, a novel two-stage framework that addresses explorable image-to-scene generation issues by introducing a global 3D proxy for high-fidelity image-to-scene generation and appearance refinement and introduces Parallel Geometry Injection and Proxy-Aware Corruption training strategies.
Chuan Fang, Lingteng Qiu, Yixun Liang et al.· 1 citation
DiGS-Avatar is proposed, which reformulates this task as an efficient, diffusion-based UV-latent completion task, ensuring 3D consistency by design, and introduces a teacher-student framework where a multi-view teacher provides geometrically aligned pseudo-ground-truth latents to supervise a single-view diffusion stude...
Axolotl3D is presented, a multi-modal and occlusion-aware 3D generation model that jointly conditions on images, visibility masks, camera parameters, and a partial point cloud that synthesizes diverse conditioning regimes from large-scale 3D data, enabling robust cross-modal reasoning.
A. Hu, Maria Shugrina· arXiv.org· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.