SSPA: Enhancing Pseudo-Corpus Quality on Tibetan Machine Translation via Semantic-Syntax Prealignment
Tibetan-to-English machine translation (MT) models frequently falter under extreme domain data scarcity, often producing translations that violate the distinctive agglutinative rules of Tibetan and suffer from domain-specific stylistic mismatches. To overcome these limitations, we propose Semantic-Syntax Prealignment (SSPA), an innovative corpus generation framework. SSPA constructs high-quality pseudo-parallel pairs by explicitly minimizing the deviation between the syntactic-semantic profiles of generated samples and professional reference texts. Specifically, source-target structural representations are standardized through length-unified truncation and terminology normalization, followed by a dual-domain alignment process that maximizes syntactic cosine similarity under rigorous structural constraints. We further augment these aligned frames via a cross-length dynamic filling mechanism, which is integrated with an Expectation-over-Transformation (EOT)-based style regularization mechanism specifically adapted for stylistic perturbations, to simulate authentic linguistic variations. Extensive evaluations on our newly constructed Tibetan Medicine-Tibetan English (TM-TE) dataset demonstrate that SSPA significantly outperforms existing competitive baselines. Notably, SSPA achieves a BLEU-4 score of 36.2 and improves long-sentence BLEU-4 by 16.8 points, with a parser-verified grammatical compliance rate of 96.2%. The framework exhibits remarkable cross-domain adaptability and stylistic consistency, offering a robust, versatile solution for low-resource Tibetan professional domain MT.