Dual-domain transformer-based learned primal-dual reconstruction for PET imaging
Background and Objective Positron Emission Tomography (PET) is widely used for assessing metabolic activity and diagnosing cancer. Due to the inherently high noise levels in PET data, advanced reconstruction algorithms are essential for accurate imaging. Convolutional Neural Networks (CNNs) within the Learned Primal-Dual (LPD) framework have shown good performance for PET reconstruction; however, the locality assumption imposed by CNNs can limit their ability to capture long-range contextual dependencies in sinogram data. Motivated by our previous work on transformer-based architecture for low-dose and low-count sinogram denoising, we comprehensively investigate the integration of attention mechanisms, both self-attention and cross-attention, within the LPD framework. Methods We propose and evaluate three novel transformer-based LPD architectures: the Dual-Domain Stacked Transformer-based LPD, which employs sinusoidal patches to model long-range dependencies in the dual domain (sinogram); the Dual-Domain Restormer-based LPD, a hybrid design that combines CNNs for fine-grained local feature extraction with transformers for global information exchange; and the Dual-Domain UNet-based LPD incorporating Cross-Attention mechanism, which augments the U-Net LPD with cross-attention to capture complementary information across image and sinogram domains. To enable robust transfer learning, we further introduce a system-aligned synthetic data generation process that replicates the MiniPET-3 system, our in-house preclinical PET scanner dedicated to small-animal imaging, incorporating realistic acquisition geometry, resolution blurring, positron range effects, and Poisson noise. Results Evaluations on synthetic data demonstrate that both the Stacked Transformer-based LPD and Dual-Domain Restormer-based LPD outperform the UNet-based LPD baseline in MSE and PSNR across all noise levels, with the Restormer additionally achieving higher SSIM at low noise. The UNet-based LPD incorporating a cross-attention mechanism also improves MSE and PSNR over the baseline, while exhibiting a slight reduction in SSIM, indicating a trade-off between pixel-wise accuracy and structural similarity. Trained exclusively on synthetic data, the proposed models generalize well to experimental measurements, producing high-quality reconstructions. Conclusions Overall, this work highlights the promise of transformer-integrated LPD frameworks for enhancing PET reconstruction quality and robustness in low-count imaging settings.