2 papers indexed here
We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.
Not the right person? Other researchers publish under this name.
Dual-View X-ray to 3D CT Reconstruction via Discrete Latent Translation with VQ-VAE and Transformer
Three-dimensional (3D) CT reconstruction from dual-view X-ray data is a severely ill-posed inverse problem, due to the substantial loss of structural information caused by the sparsity of the projection data. To address this challenge, we propose a dual-view X-ray-to-CT reconstruction framework based on vector quantized variational autoencoders (VQ-VAEs) and Transformer-based latent code translation. Specifically, a 3D VQ-VAE model is pre-trained to learn discrete latent representations of CT volumes, while a 2D VQ-VAE encodes dual-view X-ray projections into discrete tokens. Two separate Transformers are then employed to transform the 2D discrete representations into top-level and bottom- level 3D latent codes, which are subsequently decoded by the pretrained 3D VAE decoder, yielding an end-to-end mapping from dual-view X-rays to 3D CT volumes. Furthermore, a genetic algorithm (GA) is introduced to optimize the predicted bottom-level 3D latent codes. The proposed method is evaluated on a synthetic paired dataset constructed from the LIDC-IDRI CT database and digital reconstructed radiographs (DRRs). The framework successfully synthesizes structurally plausible 3D CT volumes from only two X-ray views, achieving a structural similarity (SSIM) of 0.7352 and a mean squared error (MSE) of 0.0087. Latent code optimization with the GA further improves structural similarity and reduces reconstruction error. The results demonstrate the potential of the proposed approach for recovering bone structures such as ribs, scapulae, and the sternum.