Skip to content

Author

Harryson Yumnam

2 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#edge computing Open access Sep 2026

Dioptra: An Ultra-Lightweight Geometry-Aware Architecture for Monocular Metric Depth on Edge Devices

Dense metric depth estimation from a single image is crucial for autonomous robotic navigation and spatial computing. However, recent foundation models increasingly rely on heavy vision transformers (>300 M parameters) and massive pre-training, making edge deployment on memory-constrained devices challenging. We present Dioptra, a compact, geometry-aware neural architecture (8.1 M parameters, 32.4 MB weights) designed to learn metric depth directly from scratch. Dioptra integrates two camera-grounded mechanisms: (i) Trivision Ray Positional Encoding, which unprojects optical ray triplets per patch token using camera intrinsics with reflection symmetry tracking, and (ii) Angular Residual Attention (ARA), an intrinsic geometric attention bias that penalizes off-axis angular distortion to sharpen surface discontinuities. In addition, we decouple global scale supervision from relative shape regularization via a log-median scale consistency loss. On a strict cross-environment split of the TartanAir dataset (15 training environments, 8,186 held-out test frames across two unseen 3D maps, one relit at night), Dioptra Stage 1 baseline achieves an Aligned AbsRel of 0.3073 (δ₁ = 51.74%) and an unaligned Metric AbsRel of 0.9293 (Metric RMSE 6.79 m). Furthermore, to prevent ray unprojection from degenerating into a static spatial embedding on fixed-sensor benchmarks, we introduce dynamic pinhole crop augmentation; Stage 2 fine-tuning achieves our headline performance of 0.2901 Aligned AbsRel and 0.6043 Metric AbsRel (35.0% error reduction), while multi-camera field-of-view sweeps (50° to 100°) demonstrate camera-intrinsic equivariance with up to -55.9% error reduction (>2.2× relative accuracy improvement) at telephoto angles. Local evaluation on 40 uncompressed ground-truth test arrays across all three held-out test environments demonstrates strong zero-shot transfer on daylight industrial geometry (0.2290 AbsRel, 86.6% δ₂), resilient low-light transfer (0.3291–0.5195 AbsRel), and graceful outdoor theme-park degradation (0.6239 domain mean, best-frame 0.2663), yielding an overall held-out mean of 0.5545 AbsRel (0.4577 robust mean) for the Stage 1 baseline, with Stage 2 fine-tuning achieving 0.7420 Metric AbsRel alongside sub-meter fidelity on in-domain clinical corridors (best-frame 0.64 m RMSE, domain mean 3.22 m). Operating at under 180 MB peak RAM with 52.8 ms steady-state FP32 forward inference via layer-shared geometric caching (≈ 18.9 FPS; 98.4 ms unoptimized reference; 28.2 ms on NVIDIA T4; 334.9 ms full 3D mesh pipeline), Dioptra demonstrates that explicit geometric inductive biases enable viable metric depth estimation within a lightweight embedded footprint. Code is available at https://github.com/SeranomTheGreat/dioptra

Harryson Yumnam · 0 citations
#edge computing Open access Sep 2026

Dioptra: An Ultra-Lightweight Geometry-Aware Architecture for Monocular Metric Depth on Edge Devices

Dense metric depth estimation from a single image is crucial for autonomous robotic navigation and spatial computing. However, recent foundation models increasingly rely on heavy vision transformers (>300 M parameters) and massive pre-training, making edge deployment on memory-constrained devices challenging. We present Dioptra, a compact, geometry-aware neural architecture (8.1 M parameters, 32.4 MB weights) designed to learn metric depth directly from scratch. Dioptra integrates two camera-grounded mechanisms: (i) Trivision Ray Positional Encoding, which unprojects optical ray triplets per patch token using camera intrinsics with reflection symmetry tracking, and (ii) Angular Residual Attention (ARA), an intrinsic geometric attention bias that penalizes off-axis angular distortion to sharpen surface discontinuities. In addition, we decouple global scale supervision from relative shape regularization via a log-median scale consistency loss. On a strict cross-environment split of the TartanAir dataset (15 training environments, 8,186 held-out test frames across two unseen 3D maps, one relit at night), Dioptra Stage 1 baseline achieves an Aligned AbsRel of 0.3073 (δ₁ = 51.74%) and an unaligned Metric AbsRel of 0.9293 (Metric RMSE 6.79 m). Furthermore, to prevent ray unprojection from degenerating into a static spatial embedding on fixed-sensor benchmarks, we introduce dynamic pinhole crop augmentation; Stage 2 fine-tuning achieves our headline performance of 0.2901 Aligned AbsRel and 0.6043 Metric AbsRel (35.0% error reduction), while multi-camera field-of-view sweeps (50° to 100°) demonstrate camera-intrinsic equivariance with up to -55.9% error reduction (>2.2× relative accuracy improvement) at telephoto angles. Local evaluation on 40 uncompressed ground-truth test arrays across all three held-out test environments demonstrates strong zero-shot transfer on daylight industrial geometry (0.2290 AbsRel, 86.6% δ₂), resilient low-light transfer (0.3291–0.5195 AbsRel), and graceful outdoor theme-park degradation (0.6239 domain mean, best-frame 0.2663), yielding an overall held-out mean of 0.5545 AbsRel (0.4577 robust mean) for the Stage 1 baseline, with Stage 2 fine-tuning achieving 0.7420 Metric AbsRel alongside sub-meter fidelity on in-domain clinical corridors (best-frame 0.64 m RMSE, domain mean 3.22 m). Operating at under 180 MB peak RAM with 52.8 ms steady-state FP32 forward inference via layer-shared geometric caching (≈ 18.9 FPS; 98.4 ms unoptimized reference; 28.2 ms on NVIDIA T4; 334.9 ms full 3D mesh pipeline), Dioptra demonstrates that explicit geometric inductive biases enable viable metric depth estimation within a lightweight embedded footprint. Code is available at https://github.com/SeranomTheGreat/dioptra

Harryson Yumnam · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.