A cross-modal alignment IP for accelerating open-vocabulary object detection inference on heterogeneous edge platforms
The cross-modal similarity calculation between large-scale visual features and text dictionaries in the Open-Vocabulary Object Detection (OVOD) inference stage exerts a significant degree of memory access pressure and computational overhead, which consequently becomes a primary bottleneck that limits the real-time perf...