Skip to content
Conference

Cross-Domain Adaptive Transformer Framework for Robust Object Tracking in Unconstrained Multi-Domain Environments

Sep 2026 · Automation, Control, and Information Technology · pp. 1569-1600 · 0 citations · 30 references

Abstract

Visual object tracking (VOT) in real-world scenarios necessitates a high degree of adaptability to overcome the distributions discrepancies inherent in multi-domain environments. While standard deep learning models excel in well-conditioned laboratory datasets, their performance typically degrades when encountering adverse weather, fluctuating illumination, or sensor-specific noise. This research proposes a novel Context-Aware Multi-Domain Adaptation (CAMA) framework that integrates frequency-aware visual adapters and long-term memory modules into a frozen foundation transformer backbone. By leveraging a textconditioned scenario generator for synthetic domain expansion and an optimal transport-based confidence alignment mechanism, the system bridges the gap between source and unlabeled target domains without requiring redundant model updates. Evaluation across multiple benchmarks, including trimodal (RGB, Depth, Thermal) and adverse-weather datasets, indicates that the proposed framework significantly mitigates tracking drift and identity switches. The findings suggest a robust paradigm for real-time tracking in high-stakes applications such as autonomous navigation, aerial surveillance, and intelligent manufacturing.

View source

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.