QuerySmith: Fine-Tuning Phi-4-mini for Text-to-SQL Generation with CPU-Efficient Partial-Residual Quantization
Overview Large language models fine-tuned for text-to-SQL generation are typically evaluated and deployed assuming GPU inference, which limits their use in resource-constrained or on-premise settings where only CPU hardware is available. This work makes two contributions: Fine-tuning Phi-4-mini-instruct (3.8B params, d...