Skip to content
Conference

A Turkish Biology Dataset and LLM Model

Jul 2026 · Signal Processing and Communications Applications Conference · pp. 1-4 · 0 citations · 12 references

Abstract

General-purpose, low-parameterized language models often produce imprecise or shallow explanations when handling specialized scientific subjects, particularly in non-English contexts like Turkish, where training data is limited. In this paper, an original Turkish biology dataset within the scope of the high school curriculum has been developed; and subsequently, large natural language model studies conducted on this dataset are presented. To synthesize high-quality question-answer pairs, video transcripts and written content were first scraped from public Khan Academy resources. In the two-stage synthesis strategy, by using the GPT-4 API, firstly questions depending on the content and then answers were generated. Data diversity was ensured by producing more than one answer for a single question. The Gemma-3 1B model was fine-tuned on this specialized dataset using the Low-Rank Adaptation (LoRA) method. To measure model performance, LLM-as-a-judge, standard n-gram metrics such as ROUGE, and human evaluation were used. The best-performing configuration, Gemma-3 1B trained on standard-length answers, achieved superior results across all dimensions, reaching M-Prometheus scores of 3.8 for coherence and 3.0 for both correctness and completeness. Additionally, it has been shown that LLM-as-a-judge metrics are closer to human evaluation compared to the ROUGE metric.

View source

Similar papers

Open access Aug 2026

ADAPTIVE TURKISH FEDERATED RAG ARCHITECTURE FOR LOW-RESOURCE SYSTEMS

Today, Large Language Models (LLMs) perform many tasks in the field of natural language processing with high success, from text generation to translation, semantic analysis to code writing. However, these models have some fundamental limitations that make their reliable use challenging. They can produce factual errors known in the literature as hallucinations and cannot directly access developments after their training period. They also sometimes reproduce biases present in the training data. The Re-trieval-Augmented Generation (RAG) approach aims to produce more up-to-date and verifiable outputs by dynamically feeding the model with external information sources, thereby reducing hallucination and temporal limitations. This study examines distributed learning approaches aimed at protecting data privacy and QLoRA-based fine-tuning strategies within a mathematical framework. It also addresses 4-bit NormalFloat (NF4) quantization techniques used to improve system efficiency. Furthermore, it de-tails morphology-aware hybrid access methods and adaptive routing mechanisms that can make deci-sions based on query complexity to achieve results more suitable for Turkish. The study demonstrates that combining RAG-based systems with federated learning and homomorphic encryption creates a secure, decentralized architecture, and that LLM can be efficiently run on low-resource systems using the NF4-QLoRA combination. However, limitations such as encryption latency, erroneous information from untrusted sources, and the inadequacy of standard evaluation methods for Turkish also exist.

M. Toy, Ahmet Ali Süzen · 0 citations
Conference Aug 2026

Data-Centric Evaluation of Arabic Abstractive Summarization Using a Large-scale Curated News Corpus

This paper introduces MAAD, a high-quality, carefully constructed and curated by the authors large-scale Arabic dataset for abstractive news summarisation. The authors selected a high-quality subset of 50,000 articles from the dataset Original, which contains 602,792 articles. To maintain the quality, diversity, and training suitability of the subset, the subset underwent a multi-stage preprocessing pipeline involving noise removal, duplicate filtering, linguistic normalisation, and expert validation. The experimental evaluation was executed in two phases. In the first phase, three transformer-based models (ArabicT5, AraBART, and mT5) were evaluated on a controlled subset of 1,110 articles to establish fair baseline comparisons among models, where ArabicT5 achieved the best performance (ROUGE-1: 23.64, ROUGE-2: 11.82, ROUGE-L: 22.10). In the second phase, ArabicT5-base was trained on all 50,000 articles to evaluate scalability, achieving substantially improved results of 68.4, 52.3, and 64.1, respectively, with a BLEU score of 58.7. The findings emphasise the significance of scale, effective preprocessing, and the benefits of Arabic-specific pretraining on the quality of summarisation. Moreover, a human evaluation on 500 randomly sampled instances verified fluency and adequacy scores of 4.86 and 4.35, respectively, with a strong inter-annotator agreement (Cohen's Kappa: 0.78 and 0.74). Overall, the findings indicate that MAAD is a reliable and scalable dataset with strong potential to serve as a benchmark for Arabic abstractive summarisation and to support the development of robust transformer-based models.

M. Al-Nahari, Ayedh Abdulaziz Mohsen, Nada Abdu Al-Humidi et al. · 0 citations
Aug 2026

STAR: instruction tuning for Arabic across tasks, datasets, and models

An in-depth evaluation of instruction tuning for Arabic NLP tasks using three prominent LLMs: LLaMA 3.1-8B, AceGPT-v2-8B, and Qwen3-8B shows that instruction tuning consistently improves performance across most tasks, with notable variations in effectiveness across different tasks and prompts.

Maged Saeed Al-shaibani, Zaid Alyafeai, Irfan Ahmad · 0 citations
#natural language process... Preprint Aug 2026

MameLoshnLM: Yiddish Language Model and Evaluation Benchmark

MameLoshnLM, the first open-source 8B-parameter language model built specifically for Yiddish, is presented, providing both a foundation for Yiddish NLP and a practical template for language model development in historically rich but digitally underrepresented languages.

Uri Katz, Omer Goldman, Tomasz Limisiewicz et al. · 0 citations
#artificial intelligence Preprint Aug 2026

Reading the News: Adapting Large Language Models to Swedish Journalism Through Continued Pre-Training

This work investigates continued pre-training for adapting large language models to Swedish journalism, using a high-quality dataset that is curate from millions of news articles and demonstrates the importance of targeted evaluation in the adaptation process.

Lukas Borggren, Jenny Kunz, Marco Kuhlmann · 0 citations
#natural language process... Preprint Sep 2026

Domain-specific Pretraining Profile and Transformer Performance: Evidence from Modeling Digital Pragmatics in Arabic-English Code-switching

This study highlights the role of domain-specific pretraining profile (DSPP) in Transformer performance for modeling digital pragmatics in Arabic-English code-switched discourse. It evaluates MARBERT and XLM-R(oBERTa), with BERT serving as a general-purpose baseline. The models were evaluated on their ability to classify context-sensitive pragmatic functions in code-switched social-media discourse. 11695 unique X posts were collected via Python and utilized for the study. The study employs a quantitative and qualitative NLP approach, following a supervised pipeline. Findings unveil that MARBERT consistently surpasses XLM-R with validation Macro F1 increasing from 0.39 to 0.84 and validation loss decreasing from 0.55 to 0.19. On an independent test set, it achieved 0.96 accuracy, 0.83 macro precision, 0.87 macro recall, and 0.85 Macro F1, while XLM-R achieved 0.92 test accuracy but a substantially lower Macro F1 of 0.52. This was also supported by class-level performance where MARBERT outperforms XLM-R considerably with F1 improvements ranging from +0.33 to +0.60, demonstrating a clear advantage in modeling Arabic digital pragmatics. The study concludes that Transformer performance depends more on DSPP than multilingual coverage alone, as the latter does not guarantee optimal performance on a highly specialized pragmatic classification task.

Fahad Saud Al Hussen, King Saud University, Riyadh et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.