Skip to content
Open access

Evaluating the Effectiveness of Open‐Source LLMs for Automated Analysis of Multilingual Consultation Feedback: A Swiss Case Study

Sep 2026 · Policy Studies Journal · 0 citations · 24 references

TL;DR

Assessment of an open‐source and proprietary large language models for analyzing stakeholder inputs from three Swiss pre‐parliamentary consultations highlights both the potential and limits of algorithmic support for participatory governance in linguistically and technically diverse environments.

Abstract

Public consultations are a cornerstone of democratic policymaking, yet the sheer volume and heterogeneity of submissions often overwhelm traditional analysis. Emerging large language models (LLMs) offer new opportunities for making sense of such data, but their performance in contexts with high linguistic and technical diversity remains underexplored. In this research note, we present a case‐based demonstration rather than a generalizable evaluation, assessing an open‐source LLM (Gemma3:27b) and a proprietary LLM (GPT‐5‐mini) for analyzing stakeholder inputs from three Swiss pre‐parliamentary consultations. The coexistence of multiple languages and varying technical complexity in submissions creates uniquely challenging datasets, providing a rigorous test for automated tools. We benchmark the model against official administrative reports, assessing accuracy and completeness. Results show the open‐source model performs excellently in straightforward consultations, achieving a mean accuracy of 90%, but encounters difficulties in retrieving the substantive content (overall completeness: 70%). Comparatively, the proprietary model obtains better results than the open‐source one in these two settings. Overall, LLMs can streamline analysis and enhance interpretive capacity, complementing (but not replacing) human expertise. The study highlights both the potential and limits of algorithmic support for participatory governance in linguistically and technically diverse environments.

Read PDF

Similar papers

2026

Mind the Language Gap: Assessing LLM Safety in Italian

This paper presents a methodology for building safety evaluation datasets that comprehensively cover the full spectrum of sensitive topics relevant to LLM safety, and releases a public repository containing the list of categorized Italian Wikipedia pages, the automatically generated prompts, and the standard prompt tem...

Elena Marafatto, Roberto Navigli · 0 citations
Review Aug 2026

From Values to Benchmarks: Evaluating Large Language Models for Governmental Use in Dutch

Large language models are increasingly being deployed in governmental settings, yet few existing evaluation frameworks jointly reflect the values of public administration and the linguistic requirements of non-English contexts. We present the"Grip on LLMs"framework, a systematic evaluation suite for Dutch governmental...

Laurens Samson, Iva Gornishka, Gossa Lô et al. · 0 citations

BankGPT: the use of large language models in official communications 1

No overall preference is indicated between human and machine-generated summaries; however, experts with greater familiarity with the Bulletin and higher educational attainment showed a marked preference for the official summaries.

Claudia Biancotti, C. Camassa, Marco Fruzzetti et al. · 0 citations

Beyond Good Intentions: When Does the Framing of Multilingual and Low-Resource NLP Research Become a Caricature?

Building language technologies and conducting NLP research for low-resource languages---particularly when led by native speakers or involving participatory research practices---are often framed as means of addressing inequality, serving local communities, and, at times, contributing to *decolonisation*. In this paper,...

Nedjma Djouhra Ousidhoum, Noopur Zambare, Mohamed Abdalla · 0 citations
#artificial intelligence Preprint Sep 2026

What Do We Expect from LLMs? Mapping the Design of LLM Benchmarks

This work systematically map 14,767 papers introducing or updating evaluation resources from arXiv submissions between January 2022 and August 2026, using staged screening and automated full-text coding to examine changes in target systems and domains, evaluation materials and conditions, and scoring mechanisms.

Chao Wang · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.