Skip to content
#small language model Dataset Open access

A Dataset for Multidimensional Evaluation of French Synthetic Texts Generated by SLMs

Oct 2026 · Zenodo (CERN European Organization for Nuclear Research)
Topic Modeling

Abstract

This dataset comprises synthetic news descriptions in French, generated from real newspaper headlines using a controlled Small Language Models (SLM) pipeline. The generation process follows three distinct configurations: Retrieval-Augmented Generation (RAG) and generation without external context (NO RAG) and QLoRA-based fine-tuning (FT). In the zero-shot configuration, the base model generates each description using only the corresponding headline and generation instruction. In the RAG configuration, the same base model additionally receives contextual information retrieved from a historical French news corpus. In the fine-tuning configuration, the models are adapted using headline-description pairs from the historical corpus through QLoRA and subsequently used to generate descriptions without external retrieval. For each configuration, three temperature settings (0.9, 0.6, and 0.3) were applied to control the variability and determinism of the generated outputs. The generation process relied on a corpus of real-world French news articles collected between February 16 and March 26, 2026. To prevent information leakage, the source data was chronologically strictly partitioned: Temporal Window Timeframe News Items Description Knowledge Window 16 Feb 2026 – 18 Mar 2026 7,767 Served as the RAG knowledge base and the QLORA fine-tuning dataset. Generation Window 18 Mar 2026 – 26 Mar 2026 2,071 Reserved exclusively for synthetic text generation and evaluation. The experimental pipeline originally generated 55,917 synthetic descriptions across 27 configurations, derived from three Small Language Models (Ministral-3B, TinyLlama-1.1B, and Qwen2.5-0.5B), three generation strategies, and three temperature values. To ensure copyright compliance, exact verbatim reproductions of the input headlines were removed, resulting in a final publicly released dataset of 54,751 synthetic descriptions. The dataset is distributed in the fr_slm_multieval.zip. Inside the archive, the data is organized into newline-delimited JSON (.ndjson) files. Each file corresponds to a specific experimental configuration, allowing researchers to isolate model behaviors. The files follow the naming convention: [strategy]_[RAG-status]-[model]-[temperature].ndjson where [strategy] specifies the model adaptation state: zs: base model, used either in zero-shot or RAG generation. ft: QLoRA fine-tuned model. Examples of the configuration files included in the file: ● ft_NO-RAG-ministral-0.3.ndjson ● zs_RAG-qwen-0.9.ndjson

View source

Similar papers

#computer vision Open access Jun 2016

Software Development in Startup Companies: The Greenfield Startup Model

The results are packaged in the Greenfield Startup Model (GSM), which explains the priority of startups to release the product as quickly as possible, and the need to shorten time-to-market, by speeding up the development through low-precision engineering activities.

Carmine Giardino, Nicolò Paternoster, M. Unterkalmsteiner et al. · 178 citations · ⚡14
#computer vision Open access Oct 2016

Software Startups - A Research Agenda

Software startup companies develop innovative, software-intensive products within limited timeframes and with few resources, searching for sustainable and scalable business models.

M. Unterkalmsteiner, P. Abrahamsson, Xiaofeng Wang et al. · 157 citations · ⚡17
#machine learning Review Open access Oct 2016

“Failures” to be celebrated: an analysis of major pivots of software startups

This study conducts a case survey study based on the secondary data of the major pivots happened in 49 software startups, and demonstrates that customer need pivot is the most common among all pivot types.

Sohaib Shahid Bajwa, Xiaofeng Wang, Anh Nguyen-Duc et al. · 127 citations · ⚡15
#computer vision Review Open access May 2015

A survey study on major technical barriers affecting the decision to adopt cloud services

The comparison of adopter and non-adopter sample reveals three potential adoption inhibitor, security, data privacy, and portability, which underlines the importance of the technical and security perspectives for research investigating the adoption of technology.

Nattakarn Phaphoom, Xiaofeng Wang, S. Samuel et al. · 111 citations · ⚡8
#computer vision Open access Feb 2018

Lean Internal Startups for Software Product Innovation in Large Companies: Enablers and Inhibitors

This study investigates how Lean internal startup facilitates software product innovation in large companies and identifies its enablers and inhibitors, and shows the potential of the method-in-action framework to investigate the Lean startup approach in non-startup context.

Henry Edison, Nina M. Smørsgård, Xiaofeng Wang et al. · 78 citations · ⚡6
#computer vision Conference Sep 2010

Exploring the Sources of Waste in Kanban Software Development Projects

The application of agile software methods and more recently the integration of Lean practices contribute to the trend of continuous improvement in the software industry. One such area warranting proper empirical evidence is a project’s operational efficiency when using the Kanban method. This short paper takes a new an...

Marko Ikonen, Petri Kettunen, Nilay V. Oza et al. · 67 citations · ⚡9

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.