Aug 2026· Frontiers of Physics· Vol 14· 0 citations· 29 references
TL;DR
TLCE-morph, a Tri-Path Lie Convolution Encoder-based learning framework for deformable respiratory motion correction in thoracic PET, achieves the most favorable overall quantitative performance and shows more consistent local structural recovery in representative motion-sensitive regions.
Abstract
Respiratory motion in thoracic positron emission tomography (PET) introduces spatially heterogeneous non-rigid deformation that can blur lesions, weaken local boundary definition, and reduce structural fidelity. To address this problem, we developed TLCE-morph, a Tri-Path Lie Convolution Encoder-based learning framework for deformable respiratory motion correction in thoracic PET. The framework combines an SO(3)-based group-aware convolution module with a Tri-Path Fusion Encoder to couple orientation-aware geometric modeling with structurally guided feature encoding at local, global, and cross-scale levels. TLCE-morph was evaluated on simulated respiratory motion datasets and a two-center clinical gated PET cohort using Dice, correlation coefficient, and 95th percentile Hausdorff distance. Additional analyses included lesion-level normalized PET uptake consistency, local line-profile and full width at half maximum measurements in motion-sensitive regions, architectural ablation, group-representation comparison, and computational profiling. Across the simulated datasets, TLCE-morph remained comparatively stable as deformation increased from relatively regular displacement to more heterogeneous and coupled motion. In the clinical gated PET cohort, it achieved the most favorable overall quantitative performance among the evaluated methods and showed more consistent local structural recovery in representative motion-sensitive regions. Additional comparisons of group representations and architectural ablation indicated that the observed advantage was associated with the joint contribution of 3D orientation-aware feature modeling and complementary structural constraints rather than with any single component alone. These findings suggest that stable respiratory motion correction in thoracic PET may benefit from coupling geometric sensitivity with structurally guided feature encoding under heterogeneous deformation, rather than relying on appearance matching alone.
Purpose: Respiratory motion remains a major source of quantitative bias in PET and becomes increasingly relevant for high-sensitivity long axial field-of-view (LAFOV) PET/CT. Although numerous respiratory motion correction (MoCo) methods have been proposed, their quantitative accuracy cannot be established clinically because a patient-specific motion-free reference is fundamentally unavailable in vivo. This study combined clinical PET imaging with a digital twin, a realistic representation of both the PET/CT system and the patient, to objectively validate respiratory MoCo against a corresponding motion-free reference. Methods: Twenty patients (10 [18F]FDG with predominantly pulmonary lesions and 10 [18F]SiFAlin-TATE with predominantly hepatic lesions; total 135 lesions) were analyzed. The digital twin combined a validated LAFOV PET/CT simulation model with an anatomically realistic phantom containing 14 lung and liver lesions, two patient-derived respiratory patterns, and respiratory motion amplitudes of 2 and 3 cm, generating patient-like datasets with corresponding motion-free references. Data-driven and image-based MoCo were evaluated using lesion morphology, SUVmean, SUVmax, and metabolic tumor volume (MTV). Results: In patients, data-driven MoCo produced larger SUVmean increases than image-based MoCo for liver (48.1{+/-}18.9% vs. 17.0 {+/-} 12.0%; p<0.01), lower-lung (32.5{+/-}21.2% vs. 16.3{+/-}15.6%, p=0.06), and upper-lung lesions (28.4{+/-}32.0% vs. 10.4 {+/-} 17.2%; p<0.01), with similar findings for SUVmax and larger MTV reductions. Simulation revealed marked motion-induced SUVmean underestimation before correction, particularly in liver (-31.2{+/-}6.8%) and lower lung (-15.5{+/-}13.9%). Relative to the motion-free reference, data-driven MoCo most accurately recovered hepatic uptake (4.3{+/-}11.7% vs. -10.0 {+/-} 9.2%; p=0.01) but overestimated pulmonary uptake (lower lung: 19.8{+/-}16.3% vs. -1.6 {+/-} 10.2%; p=0.02). SUVmax showed the same regional behavior, whereas image-based MoCo yielded MTV estimates closer to the reference. Quantitative recovery was largely independent of respiratory pattern, while larger motion amplitudes mainly affected image-based MoCo. Conclusion: Combining clinical PET with a realistic digital twin and corresponding motion-free ground truth enabled objective validation of respiratory MoCo beyond conventional clinical evaluation. Larger correction-induced quantitative changes should not be equated with greater quantitative accuracy. Instead, MoCo performance was region- and metric-dependent, highlighting the value of ground-truth-based validation for developing and benchmarking respiratory motion correction and quantitative PET on LAFOV PET/CT systems.
W. Lan, S. Weigel, E. Calderón et al.· medRxiv· 0 citations
Four-dimensional cone beam CT (4D CBCT) is important for image-guided radiation therapy of thoracic cancers, but its use is limited by long scan times, causing high patient dose and motion/sparse-sampling artifacts. We propose a deep learning method for motion-resolved 4D CBCT reconstruction from conventional free-breathing scans, without a respiratory signal or explicit projection binning. Our CNN takes free-breathing 3D CBCT projections as input and predicts a static volume at maximum inhalation plus ten displacement vector fields (DVFs) spanning a breathing cycle. The network extends U-Net: the encoder acts on filtered projection stacks, the decoder acts in the volume domain, and skip connections are replaced with non-trainable back-projection functions at multiple resolutions to transfer features between domains. The model is trained on simulated CBCT scans and evaluated on 11 unseen simulated patients and 13 clinical free-breathing scans. Two additional models (60 s and 6 s scans) were evaluated by clinical experts on three and two scans, comparing single phases of our 4D reconstruction to reference 3D SART-TV images for tumor and esophagus visibility. Experts preferred our method for tumor visibility (59% vs. 36% no preference, 5% reference) and esophagus visibility (47% vs. 42%, 11%). On simulated data, image quality matched SART-TV (mean RMSE: -1.19 HU, PSNR: +0.09 dB, SSIM: -0.009) while enabling 4D reconstruction. On clinical scans, our method showed sharper dynamic structures (e.g., diaphragm) and fewer motion streak artifacts than traditional reconstruction. This non-patient-specific CNN predicts static volumes and full 4D respiratory motion models from a single free-breathing scan, without a respiratory surrogate or projection binning, reducing motion artifacts while adding motion-modeling capability.
Ivo Herzig, P. Paysan, Daniel Barco et al.· 0 citations
Computed tomography (CT) is widely accessible for abdominal imaging but provides limited soft-tissue contrast compared with magnetic resonance (MR) imaging. Cross-modality synthesis offers a potential means to virtually complete multimodal protocols, yet current CT-to-MR translation methods often prioritize perceptual realism over anatomical reliability, resulting in geometric distortion and limited downstream utility. To address this gap, we introduce Reg-APGAN, a registration-guided, anatomy-preserving framework that synthesizes MR-like contrast from CT while constraining organ geometry. Using abdominal CT and T1-weighted MR from 114 patients, we convert both modalities into a unified coronal space and perform hierarchical rigid registration using skeletal and multi-organ labels. Reg-APGAN then performs 2D slice-wise translation on rigidly aligned 3D CT-MR volumes under structure-aware supervision to preserve anatomical topology. Evaluation is performed on the full abdominal cavity, a highly deformable and challenging setting for multimodal synthesis. Under identical weak-alignment conditions, Reg-APGAN yields a 0.51 dB PSNR improvement, the highest MS-SSIM, a 6-7% reduction in MAE relative to CycleGAN, and a 13-15% reduction in ROI-based intensity and distributional errors. Moreover, pseudo-MR images generated by Reg-APGAN enabled more anatomically coherent downstream segmentation than CT alone and pseudo-MR images generated by baseline translation methods. By coupling anatomical registration with contrast-level translation, Reg-APGAN enables structurally consistent MR-like visualization from CT and supports multimodal AI development. However, the method is not intended to replace diagnostic MR and requires external validation and radiologist reader studies and external datasets is required.
This work provides an openly released resource consisting of tracked real motion and pregenerated synthetic motion, along with a pretrained variational autoencoder (VAE) to generate larger ground-truth datasets, intended to support reproducible development, benchmarking, and comparison of head motion estimation methods in medical imaging modalities.
M. Goldmann, Felix Damm, F. Goldmann et al.· Journal of Medical Imaging· 0 citations
Whole-heart segmentation from CT and MRI is essential for quantitative cardiac image analysis, but remains challenging under multi-center and multi-modality distribution shift. In the CARE whole-heart segmentation task, models must generalize from limited labeled sites to unseen acquisition distributions, where variation in spacing, intensity, reconstruction texture, and anatomy can degrade out-of-distribution performance. We propose a modality-routed 3D cardiac segmentation pipeline that combines TotalSegmentator-initialized nnU-Netv2 models with site-characterized, label-preserving appearance augmentation. We first characterize the available sites using measurable image properties and use this analysis to motivate candidate data-space generalization routes. The final retained recipe applies Bias Field + Bezier appearance augmentation, combining smooth spatial intensity perturbation with nonlinear intensity remapping, followed by lightweight class-wise largest-connected-component cleanup. On the primary held-out-site validation splits, the final configuration improves CT mean Dice from 0.8350 to 0.9135 and MRI mean Dice from 0.7695 to 0.7830, while also reducing HD95. These results suggest that site-motivated appearance augmentation is a practical strategy for improving cross-site robustness in limited-data whole-heart segmentation. Our code can be found in https://github.com/Purdue-M2/Improving-Cross-Site-Whole-Heart-Segmentation
Tanishqua H Mudaliar, Justin Li, Daniel Lin et al.· 0 citations
Accurate 3D segmentation is central to quantitative lesion assessment and anatomy mapping for clinical planning and follow-up. Thin, elongated, and fine anatomical/pathological structures (e.g., vessels) are a particularly challenging case: a one-voxel boundary error can disconnect a branch and change clinically relevant topology. In encoder-decoder networks (e.g., U-Net), repeated downsampling and fixed-grid convolution blur or alias fine structures and weaken orientation cues, so early mistakes propagate across scales. We propose a geometry-guided local operator that steers where features are sampled, rather than deforming convolutional kernels, under a single formulation for both feature refinement (stride 1) and resolution reduction (stride>1). At each voxel, it predicts a local orientation and bounded step sizes, samples symmetrically along these directions, and transforms paired samples into compact geometric and boundary cues with lightweight mixing; a cross-scale consensus aligns encoder and decoder features at skip connections to reduce geometric mismatch. Replacing all stride 1 and stride 2 operators in a 3D U-Net yields consistent improvements on BraTS, MSD Hepatic Vessel, and TDSC-ABUS, with notably better boundary metrics (e.g., BraTS Dice 86.1 to 88.9, HD95 7.1 to 6.2; TDSC-ABUS HD95 39.1 to 27.8) while reducing parameters from 2.3M to 0.8M. We further demonstrate that the operator can be integrated into other backbones (e.g., nnU-Net, Swin-UNETR, and MedNeXt) without changing their macro-architectures while providing consistent performance gains.