Small Language Models Fine-Tuning to Enable Intent-Driven Management of Kubernetes Resources
Abstract
Cloud-native infrastructure management increasingly demands intent-driven interfaces that translate natural language into executable commands, yet large language models (LLMs) introduce privacy risks, operational costs, and deployment constraints. This paper investigates whether small language models (SLMs) fine-tuned for domain-specific tasks can achieve comparable reliability for Kubernetes command generation. We evaluate six SLM architectures (220M–7B parameters) across multiple parameter-efficient fine-tuning methods (LoRA, QLoRA, Sparse LoRA, IA3) and compare them against inference-only LLMs using both similarity metrics (BLEU, ROUGE, METEOR, BERTScore, LLM-as-a-Judge) and live cluster execution through the Model Context Protocol. Results reveal a clear capacity threshold at 0.6B–1B parameters: models below this range fail to internalize kubectl syntax, while those above achieve strong performance with diminishing returns beyond 1B parameters. Parameter-efficient methods approach full fine-tuning quality while enabling on-premise deployment. Fine-tuned Qwen3-0.6B achieves 87% task completion compared to 94% for a commercial LLM, demonstrating that compact, domain-adapted models provide operational reliability sufficient for production intent-driven automation. These findings establish empirically validated strategies for deploying resource-efficient, privacy-preserving infrastructure automation without dependence on external APIs.