AtlasCT: Report-Conditioned 3D CT Synthesis with a Learnable Population Atlas Prior
Abstract
Medical image synthesis can reduce data scarcity, but volumetric generation must preserve anatomy across planes. In chest computed tomography (CT), report conditioning specifies pathology but gives little spatial guidance, while mask-guided methods require a case-specific segmentation at inference. AtlasCT removes that requirement: a population occupancy atlas, built once from the training set, is compressed into an embedding and injected through a text-conditioned gate and an atlas-affinity attention bias, so each report sets how strongly the prior applies. The contribution is the mechanism that makes a static population prior case-adaptive rather than the choice of prior source. A compact convolution-transformer backbone with bidirectional text-volume fusion and single-stage super-resolution diffusion generates volumes at 512 by 512 by 512 resolution. On CT-RATE, AtlasCT achieves the best MedicalNet-based Frechet distance and maximum mean discrepancy among the evaluated report-conditioned methods. Removing the atlas reduces the seven-region mean Dice similarity coefficient from 0.759 to 0.704, a mean-anatomy-bias analysis shows that adaptive gating avoids contraction toward homogeneous anatomy, and synthetic augmentation improves the classifier macro-average area under the receiver operating characteristic curve by 0.009, with a 95 percent confidence interval from 0.004 to 0.014.