Skip to content
Book Open access

Evolutionary Optimization of Instruction-Tuned Large Language Model Inference Parameters for Spanish Text Simplification

Jul 2026 · Proceedings of the Genetic and Evolutionary Computation Conference Companion · pp. 1465-1473 · 0 citations · 31 references

TL;DR

The framework offers a replicable methodology for inference parameter optimization applicable to Spanish and other under-resourced languages in Latin American Natural Language Processing (NLP) contexts and provides empirical support for the single-objective approach within the studied setting.

Abstract

Large Language Models (LLM) have become valuable tools for automatic text simplification, yet their output quality is highly sensitive to inference-time parameters such as temperature, repetition penalty, top-p, and top-k. These parameters are typically set heuristically rather than systematically optimized. In this work, we apply Covariance Matrix Adaptation Evolution Strategy (CMA-ES) to identify high-performing inference configurations for an instruction-tuned Large Language Model (LLM) fine-tuned for Spanish Text Simplification (TS) on the Financial Education Corpus IN SpAnish (FEINA) benchmark. Using System output Against References and against the Input sentence (SARI) as the optimization objective and Sentence Bidirectional Encoder Representations from Transformers (SBERT) as a monitoring metric, CMA-ES converges in 9 of 25 generations, improving test-set SARI from 36.57 (default) to 38.87 (+6.3%) while increasing SBERT from 0.85 to 0.88. Grid search analysis reveals that repetition penalty and temperature are the dominant factors influencing simplification quality, while top-k and maximum tokens have negligible effects. The strong positive correlation between SARI and SBERT across all 2,500 evaluated configurations provides empirical support for the single-objective approach within the studied setting. Our framework offers a replicable methodology for inference parameter optimization applicable to Spanish and other under-resourced languages in Latin American Natural Language Processing (NLP) contexts.

Read PDF

Similar papers

2026

Optimizing Large Language Models for Robust Domain-Specific Text-to-SQL: From Prompting to Preference Alignment

This work compares Proximal Policy Optimization (PPO), Direct Preference Optimization (DPO), and Odds Ratio Preference Optimization (ORPO) using a novel reward modeling approach based on execution and semantic principles, revealing that while standard PPO suffers from reward sparsity and catastrophic collapse on 7B models, monolithic alignment via ORPO scales efficiently to 20B parameter models.

Noah Hampp, Katya Mirylenka, Michael R. Glass · 1 citation
#natural language process... Preprint Jul 2026

A Factorial Study of Synthetic Data Generation for Low-Resource Machine Translation using Grammar Books

A pipeline that uses large language models to extract grammatical rules, example sentences, and lexicons from grammar books and generate synthetic parallel corpora for fine-tuning-rather than feeding grammar content into prompts at inference time, as in prior work is introduced.

V. Ravikumar, Sina Ahmadi, L. Jäger et al. · 0 citations
Conference Jul 2026

Instruction-Tuned Translation Language Model Focused On Turkish

Large language models have high computation and inference costs. Recently, Small Language Models (SLMs) have become more important because they require fewer resources and offer high efficiency. Different training approaches can be used for SLMs to achieve high performance, even on resource-constrained hardware.In this study, we present our MT-270M translation model. It was trained using instruction fine-tuning to provide high efficiency and success for bidirectional translation between Turkish and English. We observe how we selected the datasets for the training phase and how data selection affects translation quality. Then, we explain how we prepared our high-quality training data. Finally, we examine the effects of data quality changes and including different tasks in the training process on the success of the small language model.

Ali Efe Çoban, Oguz Dikenelli · 0 citations
Book Open access Jul 2026

Audience-Customized Translation of Rule-Based Evidence with Large Language Models Across Multiplexer Benchmarks

Evolutionary rule-based machine learning (ERBML) algorithms can capture complex relationships while still yielding highly interpretable models comprised of IF:THEN rules. During prediction, 'matching' rules contribute to, and form the explanation for, the model's prediction. However, rules and their associated parameters (in their raw form) are likely too technical for their intended users. This study examines the feasibility of using a large language model (LLM) to translate the prediction evidence from matching rules into natural language text for different audiences, e.g. layman, clinician, expert. Using models trained by the 'HEROS' ERBML on MUX benchmarks, we evaluate LLM text quality metrics under different scenarios (i.e. 1,800 prediction explanations). We observe that (1), intuitively, LLM quality performance improves on HEROS models that have been more ideally trained, (2) making a glossary available to the LLM to define feature names generally raises explanation audience-fit scores and sometimes lowers overstatement (beyond rule-evidence), but it also lengthens explanations and often increases hallucination rate, and (3) audience customization creates some LLM performance trade-offs. These results suggest that constrained LLM translation of rules for natural language prediction explanations is feasible, while highlighting the importance of carefully designing the LLM prompts and evidence input from the ERBML.

H. Bandhey, Gabriel Lipschutz-Villa, Khoi Dinh et al. · 0 citations
Preprint Aug 2026

Efficient Multilingual Neural Machine Translation via Corpus-Driven Vocabulary Pruning: An English-Arabic Case Study

This paper proposes a general optimization framework that combines a vocabulary pruning method with a targeted fine-tuning protocol for MNMT models, and reduces the vocabulary size from over 128,000 to approximately 10,000 tokens, enabling a 60% memory saving without any loss in performance.

Ahmed Amine Aliane, N. Semmar, H. Aliane · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.