Skip to content

From Normative Frameworks to Alignment Data: Constructing and Evaluating SFT and Preference Data

Sep 2026 · 0 citations · 34 references
Computer Science

TL;DR

An expert-driven methodology for constructing such alignment data and applying it to a normative framework grounded in Islamic ethical, theological, and jurisprudential traditions is presented.

Abstract

Aligning language models with a specified normative framework requires translating abstract principles into concrete examples and preference signals from which models can learn. We present an expert-driven methodology for constructing such alignment data and apply it to a normative framework grounded in Islamic ethical, theological, and jurisprudential traditions. Over approximately one year, seven domain experts systematically probed language models to identify alignment deficiencies, curated desired responses, and constructed preference pairs from model outputs and expert judgments. The resulting Arabic-English datasets contain approximately 2.8K supervised fine-tuning (SFT) examples and 5.4K preference pairs spanning a broad range of normative domains. We evaluate the datasets through controlled post-training experiments comparing a Baseline model with models incorporating the curated SFT data alone and both the SFT and preference data. In blind expert evaluation on 150 separately constructed prompts, the model trained with the curated SFT data was preferred over the Baseline in 51.3% of assessor judgments, compared with 14.4% in the opposite direction (p<.001 at the prompt level). Adding the preference data resulted in a smaller difference, with the model trained with both datasets preferred over the SFT model in 28.0% of judgments versus 20.9% in the opposite direction; this difference was not statistically significant at the prompt level (p = .166). Standard Arabic and English benchmarks show no broad degradation in general-purpose capabilities. These results demonstrate how expert-defined normative principles can be systematically operationalized into alignment data and evaluated through controlled model training.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

AlignDiff: Exploiting Model-Intrinsic Information for Better Preference Data Selection

AlignDiff, a preference data filtering framework driven by intrinsic model signals, first identifies samples with clear preferences using both positive and inverse signals, then prioritizes the more challenging samples based on the average negative log-likelihood gap, encouraging the model to learn richer information f...

Peng Lai, He Zhu, Zhiwen Ruan et al. · 1 citation
#machine learning Preprint Sep 2026

From Constitutions to Control: Interpretable Rewards for Aligning Language Models

Current approaches to aligning language models often make it hard to know what behavior is being rewarded or to change that reward in a targeted way. In particular, standard preference-based methods collapse multiple considerations into aggregate human judgments, obscuring what drives the resulting reward, while princi...

Johann D. Gaebler, C. Isley, Max Lamparth et al. · 0 citations
#artificial intelligence Review Sep 2026

Do LLMs Have Values? A Quantitative Analysis and Alignment Framework for Values in Large Language Models

This study empirically confirms the presence of LLM values, accurately quantifies their shifts, and achieves more efficient and precise steering than conventional blind training, all without degrading general capabilities.

Kelvin Zhang, Jing-Yu-Gin Chen, Yu-Fan Liu et al. · 0 citations
#natural language process... Preprint Sep 2026

Chronologic: Measuring Language Models'Ability to Represent the Past

Language models are appealing tools for research on the past. But to trust the evidence a model provides, researchers need to know whether its responses fit the period represented. Validation is challenging, because this is not a task living people ordinarily perform, and because many questions have multiple correct an...

Ted Underwood, Zi-Liang Qiu, Sarah Griebel et al. · 0 citations
Review Open access Aug 2026

Artificial Minds, Cultural Shadows: Cultural Alignment, Identity, and Voice Across Multiple Large Language Models

Comparison of five widely used large language models suggests that AI-generated language may shape how culturally situated perspectives are expressed, with differences across models indicating that AI-generated language may shape how culturally situated perspectives are expressed.

Ashkan Goudarzi, Aylar Naderi Zonouz · 0 citations
2026

Mind the Language Gap: Assessing LLM Safety in Italian

This paper presents a methodology for building safety evaluation datasets that comprehensively cover the full spectrum of sensitive topics relevant to LLM safety, and releases a public repository containing the list of categorized Italian Wikipedia pages, the automatically generated prompts, and the standard prompt tem...

Elena Marafatto, Roberto Navigli · 0 citations

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.