This work introduces a sparse feature-matching model driven by a universal retinal vessel segmentation to achieve robust coarse global alignment across modalities and develops a modality-invariant optical flow estimation network, termed MI-RAFT, to refine the alignment through dense local registration.
Abstract
Retinal image registration is essential for ophthalmic diagnosis, longitudinal disease monitoring, and multimodal retinal image analysis. Existing retinal registration methods are typically modality-dependent: they are designed or optimized either for a single imaging modality in mono-modal registration or for a fixed pair of modalities in cross-modal registration. This limits their flexibility and applicability in practical scenarios involving diverse retinal imaging modalities and different combinations of them. In this work, we propose a generalizable two-stage, modality-invariant framework for retinal image registration. First, we introduce a sparse feature-matching model driven by a universal retinal vessel segmentation to achieve robust coarse global alignment across modalities. Second, we develop a modality-invariant optical flow estimation network, termed MI-RAFT, to refine the alignment through dense local registration. Extensive experiments demonstrate that the proposed method can handle diverse combinations of commonly used retinal imaging modalities, exhibiting strong modality invariance while outperforming state-of-the-art modality-dependent registration methods.
A large-scale pre-training framework that utilizes synthetic cross-modal training signals, incorporating diverse data from various sources, to teach models to recognize and match fundamental structures across images, which generalizes effectively across more than eight unseen cross-modality registration tasks.
Xingyi He He, Hao Yu, Sida Peng et al.· IEEE Transactions on Pattern...· 0 citations
Accurate segmentation of ocular regions is a fundamental prerequisite for robust ocular biometric recognition, as it affects feature extraction, multimodal representation, and subsequent matching under unconstrained conditions. This review focuses on ocular region segmentation as an essential component of biometric recognition rather than an isolated semantic segmentation task. We survey major segmentation paradigms, including traditional image-processing methods, deep learning architectures, and emerging foundation-model-based approaches, and analyze their applications in delineating the iris, sclera, pupil, and periocular structures under challenges such as occlusion, illumination variation, specular reflection, and motion blur. We further categorize multimodal ocular fusion strategies into three levels—data-, feature-, and decision-level fusion—and compare their information preservation, alignment requirements, computational complexity, and deployment suitability. To facilitate reproducible evaluation, we review representative public datasets and benchmark protocols, summarizing their acquisition characteristics, annotation properties, and evaluation metrics. Finally, we discuss practical considerations for method selection and highlight persistent challenges, including image-quality degradation, cross-spectral and cross-sensor domain shifts, limited dataset diversity, and cross-modal misalignment. Future research directions are outlined toward standardized, quality-aware, efficient, and deployable ocular biometric recognition systems.
Received: 5 April 2026 | Revised: 30 June 2026 | Accepted: 27 July 2026
Conflicts of Interest
The authors declare that they have no conflicts of interest to this work.
Data Availability Statement
No new data were generated or collected in this review article. This study is based exclusively on previously published literature and publicly available ocular biometric datasets. The datasets mentioned in this review, including iris, sclera, pupil, and periocular image databases, can be accessed through the original publications or repositories maintained by their respective dataset providers. The authors did not create, modify, or redistribute any dataset during this study.
Author Contribution Statement
Yansuo Yu: Conceptualization, Methodology, Software, Validation, Formal analysis, Investigation, Resources, Data curation, Writing – review & editing, Supervision, Project administration, Funding acquisition. Yongbin Qi: Conceptualization, Methodology, Software, Validation, Formal analysis, Investigation, Resources, Data curation, Writing – original draft, Writing – review & editing, Visualization. Shilin Zhao: Validation, Formal analysis, Investigation, Resources, Data curation, Visualization. Huanzhang Qi: Validation, Formal analysis, Investigation, Resources, Data curation, Visualization. Wendao Li: Resources, Visualization. Haoqi Zhang: Resources, Visualization. Da Teng: Validation, Formal analysis, Resources, Supervision, Project administration. Qiang Liu: Supervision, Project administration, Funding acquisition.
Longitudinal brain computed tomography (CT) is routinely used to monitor disease progression in patients with focal lesions and other neuropathologies. Accurate registration across time points is essential for reliable quantitative analysis. However, this task remains challenging due to the intrinsically low soft-tissue contrast of CT, particularly between gray and white matter, and the presence of space-occupying lesions, which induce substantial anatomical deformation and violate the intensity-consistency assumptions underlying conventional registration methods. To address these challenges, we propose a lesion-robust framework for longitudinal brain CT registration based on MRguided image translation. Specifically, an Image-to-Image Schrödinger Bridge (I2SB) network is employed to translate CT images into pseudo-MR images with enhanced anatomical contrast. We then perform joint deformable registration on both pseudo-MR and original CT images, where pseudo-MR provides structurally informative guidance while CT enforces modality-consistent data fidelity. This joint formulation enables more reliable correspondence estimation in the presence of lesion-induced deformation and intensity inconsistencies. The resulting deformation field is subsequently applied to the original CT images to achieve accurate longitudinal alignment. We evaluate the proposed method on an in-house longitudinal CT dataset of patients with intracerebral hemorrhage. Experimental results demonstrate that our approach consistently outperforms conventional intensity-based registration methods as well as existing image-translation–assisted strategies in both quantitative metrics and visual alignment quality. By jointly leveraging contrast-enhanced structural cues and modality-consistent constraints, the proposed framework provides a robust solution for longitudinal brain CT registration in the presence of large lesions.
Zijun Cheng, Huixiang Zhuang, Yue Guan et al.· International Conference on...· 0 citations
Qualitative evaluations on heterogeneous medical spectral acquisition systems demonstrate the practical relevance of the proposed training data augmentation protocol as an enabler for spatially coherent spectral fusion in HSI workflows.
E. Wisotzky, Jost Triller, Simon W. Härtl et al.· 0 citations
Most advances in keypoint descriptions address monomodal settings, where image variations arise from viewpoint, illumination, or contrast changes. Multimodal scenarios involve images produced by fundamentally different sensing processes, such as multispectral imaging, RGB-depth, satellite imagery, or medical imaging, causing the same structures to appear differently. A common solution to cross-modal description is to train descriptors for each modality pair, which requires retraining whenever the modalities change, or to train large models, which incur a significant increase in runtime. Instead, we propose CrossFeat, a framework that enables an existing monomodal descriptor to operate across modalities. Our method learns a crossing function in descriptor space that maps features from one modality to a representation compatible with another. To preserve the structural information captured by the original descriptor, CrossFeat introduces a geometry-appearance disentanglement such that only appearance is altered while the geometric properties are preserved. Experiments across multiple domains and datasets demonstrate improved performance in multimodal matching.
Paul-Werner Schneider, Nazim Haouchine· 0 citations
Retinal fundus photography is widely used for screening and monitoring ocular diseases, but many modern classification pipelines rely on deep latent representations and provide limited interpretability. This study develops an interpretable fundus image classification framework based on a ring-structured representation of the retinal vasculature centered on the optic disc. The method quantifies vessel geometry, color appearance, oxygenation-related vascular appearance, and vessel--background entropy within concentric retinal regions. These physiologically motivated descriptors are derived from vessel masks, image intensities, and optical-density measurements and aggregated across rings to capture spatial variation in vascular properties. Using only quantitative vascular descriptors, the proposed method achieved strong classification performance across three public fundus datasets. On HRF, it achieved 91.1\% accuracy using automatically generated vessel masks, matching RETFound, a vision transformer pretrained on large-scale retinal fundus image data, under the same evaluation setting. Additional analyses suggest that pretrained image models are sensitive to acquisition-related spatial cues, including fundus scale and retinal position within the field of view, as well as broader non-vessel image characteristics. This framework may support interpretable disease classification, quantitative retinal phenotyping, and retinal biomarker discovery without requiring large task-specific training datasets.
Xiaoyan Li, Shi-qian Xu, Arvind Gupta et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.