A Factorial Study of Synthetic Data Generation for Low-Resource Machine Translation using Grammar Books
A pipeline that uses large language models to extract grammatical rules, example sentences, and lexicons from grammar books and generate synthetic parallel corpora for fine-tuning-rather than feeding grammar content into prompts at inference time, as in prior work is introduced.
V. Ravikumar, Sina Ahmadi, L. Jäger et al.
· 0 citations