Back to feed
Open access

Enhancing Small Language Models via Evolutionary Data Generation and Context-Aware DPO for Low-Resource Legal Question Answering

2026 · IEEE Access · Vol 14, pp. 113150-113173 · 0 citations · 70 references

Abstract

Deploying pre-trained language models for question answering in specialized low-resource legal domains remains challenging due to the scarcity of expert-annotated data, the complexity of legal reasoning, and limited computational resources. Existing retrieval-augmented generation approaches improve factual grounding but often fail to preserve logically consistent legal reasoning when retrieved evidence is noisy, while reinforcement learning-based alignment methods incur substantial computational costs. This study proposes an automated end-to-end framework that integrates contrastive retrieval with context-aware preference alignment to enhance the legal reasoning performance of small language models. The framework employs an evolutionary instruction-tuning methodology to systematically mutate seed statutes, generating a diverse synthetic corpus conditioned on structural legal reasoning to address data scarcity. We introduce a multi-objective rule-based reward model that functions as a virtual legal assistant by jointly evaluating logical entailment and statutory faithfulness. The resulting reward scores are used to automatically construct preference pairs for Context Direct Preference Optimization, enabling efficient policy alignment while substantially reducing the need for costly human annotation. The proposed framework is evaluated on a comprehensive benchmark comprising expert-validated queries derived from the Vietnamese Civil Code. Experimental results demonstrate consistent improvements across all evaluated baselines, achieving a 13.5% relative improvement in citation error rates over the strongest standard DPO baseline, while also demonstrating a 66% reduction compared to zero-shot prompting. The proposed framework further achieves superior performance in logical faithfulness (0.79), answer relevance (0.85), semantic similarity (0.855), and expert human evaluation (4.5/5). These findings demonstrate that integrating a rule-based virtual legal assistant with contrastive retrieval and preference optimization provides an effective and computationally efficient strategy for aligning the legal reasoning capabilities of small language models, offering a scalable solution for legal question answering in specialized low-resource environments.

Read PDF