Efficient VLM Inference System With LoRA Adapters at the Edge
Vision Language Models (VLMs) extend large language models with visual perception, enabling complex vision tasks at the edge. Low-rank adaptation (LoRA) adapters offer a lightweight method to inject domain-specific knowledge into a shared base VLM, making them attractive for edge serving where concurrent mobile clients...