Skip to content
Preprint

Cloud, Edge, or Split? Profiling Onboard and Split Vision-Language Model Deployment for Drone AI

Sep 2026 · 0 citations · 32 references
Computer Science

TL;DR

This paper benchmarks the performance trade-offs among fully onboard, cloud-based, and split-computing architectures for lightweight VLMs using SmolVLM-256M as a representative lightweight VLM and shows that no deployment strategy is universally optimal.

Abstract

Vision-Language Models (VLMs) enable edge devices like unmanned aerial vehicles (UAVs) to interpret visual observations and reason about complex environments using natural-language instructions. However, their practical deployment remains challenging as onboard inference is constrained by limited computational, memory, and energy resources, whereas cloud-based inference introduces communication latency, bandwidth overhead, and dependence on network connectivity. To address these limitations, split computing offers a promising alternative by partitioning VLM inference between the resource-constrained UAVs and more capable remote servers. However, the performance trade-offs among fully onboard, cloud-based, and split-computing architectures for lightweight VLMs have not yet been systematically profiled. This paper benchmarks these three deployment paradigms using SmolVLM-256M as a representative lightweight VLM. We quantify their inference latency, computational resource utilization, communication overhead, and energy consumption across varying image resolutions and network conditions. Our results show that no deployment strategy is universally optimal; instead, the preferred strategy depends on the interaction between network conditions and input image resolution.

View source

Similar papers

Preprint Sep 2026

SCORAS-MoE: Joint Compression and Resource-Adaptive Deployment of MoE-VLMs in LEO Satellite Networks

Deploying large vision-language models (VLMs) onboard satellites enables onboard data processing and reduces raw data downlink. However, onboard inference faces two resource challenges. Limited onboard memory and energy require model compression and distributed deployment. Dynamic resource availability requires fast de...

Tong Quan, Yuan-Long Wan, Hua-Sen He et al. · 0 citations
Preprint Sep 2026

EdgeVLN: Runtime-Aware Deployment Ready Quantized Vision Language Navigation Model

Vision-language navigation (VLN) models perform well but target compute-rich platforms, limiting deployment on memory- and power-constrained robotic edge devices. Compression alone does not establish whether a VLN model fits the memory, latency, and energy budgets of an edge platform while preserving navigation behavio...

Rithvik Jonna, Man Namgung, Aakash Gurram et al. · 0 citations
Preprint Aug 2026

Risk-Adaptive Edge--Cloud Visual Reasoning for Communication-Efficient Autonomous Driving

A risk-adaptive edge-cloud architecture in which onboard traffic assessment determines when cloud reasoning is requested is presented, in which onboard traffic assessment served as a practical trigger for selective VLM inference in these experiments.

Meng Ma, Shu-Yang Li, Nai-Gang Wang et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Distilling Vision-Language Models for On-Device Fire Understanding

A teacher-student knowledge distillation framework in which large VLMs fine-tuned for fire understanding can be distilled into lightweight students is developed, which provides broader guidance for deploying domain-specialized VLMs in resource-constrained, safety-critical settings.

Mohammad Kazzazi, Zi-Xuan Liu, S. Khajavi · 0 citations
#edge computing Preprint Aug 2026

AI Infrastructure in Space: How Far Can We Go?

A systems vision for AI infrastructure in space is developed as the systems layer that manages AI capabilities across spacecraft, orbital networks, ground stations, and cloud backends, while treating orbital and physical state as part of the resource model.

Qing Li, Qi-Yang Zhang, Da-Liang Xu et al. · 0 citations
#edge computing Preprint Sep 2026

AceSpec: An Asymmetric Edge-Cloud Collaborative Framework for Communication-Efficient LLM Inference

AceSpec, an asymmetric edge-cloud collaborative framework that employs an asymmetric communication protocol that transmits minimal main-chain indices uplink and compact sparse distributions downlink and introduces a network-aware, Lagrangian-optimized resource allocation strategy that dynamically maximizes the local ca...

Yi-Da Zhang, Zhi-Yong Gao, Shuai-Bing Yue et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.