LiftXR is proposed, an interleaved, geometry-guided framework that explicitly incorporates spatial layout recovery into CT reconstruction, and consistently outperforms recent X-ray-to-CT reconstruction methods, establishing a new state of the art.
Abstract
X-ray imaging can be approximately modeled as the projection of an underlying volumetric attenuation field, with each measurement recording the accumulated attenuation along a corresponding ray path. Reconstructing a CT volume from only a few X-ray views is therefore severely ill-posed, as the projections collapse depth information and leave 3D locations of anatomical regions and their corresponding intensity distributions highly entangled and ambiguous. We observe that once the spatial organization of anatomical regions is established, estimating their CT intensities becomes substantially more tractable. Motivated by this, we propose LiftXR, an interleaved, geometry-guided framework that explicitly incorporates spatial layout recovery into CT reconstruction. Specifically, a layout lifter first generates a 3D anatomical layout from bi-planar X-rays, providing spatial guidance for an intensity renderer to reconstruct a CT volume. An anatomical parser then performs volumetric perception on the reconstruction, exploiting its spatially resolved boundary and intensity cues to recover a refined anatomical layout. This transition from projection-conditioned layout generation to reconstruction-conditioned anatomical perception allows the parsed layout to provide feedback for region-specific intensity calibration. Extensive experiments on two public datasets demonstrate that LiftXR consistently outperforms recent X-ray-to-CT reconstruction methods, establishing a new state of the art. Moreover, the reconstructed CT achieves superior performance in external downstream segmentation, indicating improved anatomical fidelity. Code will be released.
This work investigates whether a single diffusion model trained across several imaging domains can instead serve as a prior for many CT problems simultaneously, and proposes a frozen model that out-performs analytic reconstructions in all three cases.
Haley Duba-Sullivan, Patxi Fernandez-Zelaia, Obaidullah Rahman et al.· 0 citations
HiGDiff is proposed, a feed-forward hierarchical Gaussian diffusion framework that decomposes reconstruction both spatially and from structure to detail in three distinct CT benchmark datasets.
This paper proposes K-NeAS, a unified and scalable architecture for automated, multi-material surface reconstruction that replaces independent material networks with a shared latent backbone and introduces a fully differentiable $K$-material sequential soft selector to model an arbitrary number of overlapping tissues.
Daksh K. Shah, Emmanouil Nikolakakis, Razvan V. Marinescu· arXiv.org· 0 citations
The reconstruction of volumetric computed tomography (CT) from two orthogonal X-ray projections is an inverse problem that is difficult to solve as only limited depth information can be obtained from the two-dimensional (2D) projections of three dimensional (3D) anatomies. The majority of previous work has concentrated on improving the reconstruction model, whereas the goal of this work is to improve the input representation before the 3D reconstruction stage, we introduce a diagnostic benchmarking framework to obtain intermediate pseudo-volumetric diagnostic inputs before CT prediction from the digitally reconstructed radiographs (DRRs) in both anterior-posterior (AP) and lateral (LAT) views. This validates the pipeline by experimentation: CT-to-CT; simplified DRR-to-CT and Plastimatch-based DRR-to-CT. Then, it tests one-channel, three-channel and five-channel pseudo-volumetric representations on top of a 3D U-Net-based reconstruction backbone. The peak signal-to-noise ratio (PSNR) and three-dimensional structural similarity index measure (SSIM3D) were used to quantify the evaluation. For multi-patient experiments, enriching the pseudo-volumetric representation was found to yield better validation PSNR from 17.24 dB with pseudol to 18.55 dB with pseudo5, and a mean PSNR of 19.27 dB during validation inference. The results demonstrate that input representation quality and projection consistency are two important factors for biplanar X-ray-to-CT reconstruction. The proposed framework is thus designed as a way to analyze input representations and is not a clinically deployable CT replacement system.
Sahlah Abd Ali, Ielaf O. Abdul Majjed Dahl· IEEE Jordan Conference on Ap...· 0 citations
EpiC-NeRF is proposed, a CT-specific closed-loop framework that actively feeds estimated epistemic uncertainty back into sparse-view reconstruction and achieves improved reconstruction fidelity over existing analytic, iterative, and neural implicit reconstruction methods.
Donghyuk Choo, Haill An, Younhyun Jung· Mathematics· 0 citations
Recovering the 6-DoF pose of the knee bones from a plain radiograph, given the patient's segmented pre-operative CT, turns a routine low-dose image into a quantitative measurement of joint geometry, without the added dose of a repeat CT or a fixed biplanar rig. Classic solutions align a rendered bone silhouette to image edges; recent alternatives refine pose by backpropagating an image-similarity loss through a differentiable X-ray renderer. Both operate one patient at a time and are fragile under a single view. Silhouettes are depth-ambiguous, and differentiable-rendering refinement has a narrow capture range at substantial per-iteration cost. We instead learn an amortized, subject-agnostic dense 2D-3D correspondence, supervised solely by projection geometry. One shared-weight model per bone, trained across 758 patients, registers patients unseen during training. The pose then follows in closed form from a global, initialization-free, render-free PnP+RANSAC solve. Because X-ray formation is transmissive, our correspondence target is transmission-aware rather than tied to a single surface. Though trained only to register, the representation is anatomically semantic: a simple classifier reads a landmark's anatomical region from its embedding across held-out patients, and the same features separate the knee's bones into a 2D-3D-consistent identity learned without any bone label. On a large single-institution cohort the model generalizes well to held-out patients.
Rembert Daems, Jonas Grammens, Caro Roten et al.· arXiv.org· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.