From Perspective to Fisheye Depth Estimation and Open-Vocabulary Segmentation
The crux of the method lies in a set of learnable parameters, termed Distortion Extenders (DEX), that model the fisheye distortion coefficients and the distributional shift between fisheye and perspective images encoded in the latent space, which transforms the latent embeddings of fisheye images to resemble those of perspective images to recover high-fidelity estimates.