Sep 2026· Utrecht University Repository (Utrecht University)
Abstract
Large Language Models (LLMs) have transformed natural language processing (NLP), but their billions of parameters make them opaque. This lack of transparency is especially problematic in high-risk areas such as healthcare, finance, and content moderation, where understanding model decisions is essential for responsible use. This dissertation focuses on Explainable NLP (XNLP), making NLP models transparent, as a fundamental requirement for trustworthy AI systems. The dissertation addresses four research questions. RQ1 asks how XNLP methods can be designed and applied to meet the unique demands of high-risk domains such as healthcare, finance, and social media moderation. RQ2 investigates how token-level explanation methods can provide transparency in text classification systems and reveal vulnerabilities to adversarial manipulation. RQ3 examines the extent to which annotator demographics influence labeling decisions and how content-driven XAI techniques compare to demographic persona prompting for LLM-based annotation. RQ4 explores how moral alignment in LLMs can be evaluated across cultures in a systematic and transparent manner. To address these questions, the dissertation employs token-level explanation methods such as SHAP and LIME across multiple tasks. It surveys XNLP applications across domains, identifying gaps between methodological research and practical deployment; develops a transparent sexism-detection pipeline that combines BERT-based classification with SHAP explanations so that content moderators can verify decisions at the word level; and applies explainability methods to AI-generated text detection, revealing that detectors often rely on superficial features and showing that token-replacement experiments can expose these vulnerabilities and improve robustness. The dissertation then shifts to human-centered evaluation and moral alignment. It finds that text content is the dominant factor in annotation decisions, far outweighing annotator demographics, and that content-focused SHAP explanations are more effective than demographic persona prompting for guiding LLM annotations. It evaluates how well LLMs capture moral attitudes across cultures, finding that instruction-tuned models achieve moderate alignment with human survey data but show a persistent Western-centric bias. Finally, it introduces the EvalMORAAL framework, which combines Chain-of-Thought (CoT) reasoning with LLM-as-judge peer review; explicit reasoning consistently improved alignment compared with implicit scoring, though a significant gap between Western and non-Western regions remains. Together, these contributions show that explainability methods can improve both the reliability and the transparency of NLP systems. The dissertation also acknowledges limitations, including a focus on classification tasks and mainly English text, and limited human evaluation. It outlines a vision where explanations are actively integrated into model training, creating a feedback loop between human evaluation and model improvement, and the persistent regional difference in moral alignment underscores that making AI systems transparent and fair is an ongoing effort that requires continued attention to cultural diversity.
The results are packaged in the Greenfield Startup Model (GSM), which explains the priority of startups to release the product as quickly as possible, and the need to shorten time-to-market, by speeding up the development through low-precision engineering activities.
Carmine Giardino, Nicolò Paternoster, M. Unterkalmsteiner et al.· IEEE Transactions on Softwar...· 178 citations· ⚡14
Software startup companies develop innovative, software-intensive products within limited timeframes and with few resources, searching for sustainable and scalable business models.
M. Unterkalmsteiner, P. Abrahamsson, Xiaofeng Wang et al.· e-Informatica Software Engin...· 157 citations· ⚡17
This study conducts a case survey study based on the secondary data of the major pivots happened in 49 software startups, and demonstrates that customer need pivot is the most common among all pivot types.
Sohaib Shahid Bajwa, Xiaofeng Wang, Anh Nguyen-Duc et al.· Empirical Software Engineeri...· 127 citations· ⚡15
The comparison of adopter and non-adopter sample reveals three potential adoption inhibitor, security, data privacy, and portability, which underlines the importance of the technical and security perspectives for research investigating the adoption of technology.
Nattakarn Phaphoom, Xiaofeng Wang, S. Samuel et al.· Journal of Systems and Softw...· 111 citations· ⚡8
The ongoing work building a Raspberry Pi cluster consisting of 300 nodes is presented, with potential use cases being an inexpensive and green test bed for cloud computing research and a robust and mobile data center for operating in adverse environments.
P. Abrahamsson, S. Helmer, Nattakarn Phaphoom et al.· IEEE International Conferenc...· 110 citations· ⚡7
The results indicate that software developers are a slightly happy population, but the need for limiting the unhappiness of developers remains, and 219 factors representing causes of unhappiness while developing software are identified.
D. Graziotin, Fabian Fagerholm, Xiaofeng Wang et al.· International Conference on...· 84 citations· ⚡6
Related blog posts
MIT News · Artificial Intelligence· news.mit.eduSep 14, 2026
The “HardFlow” algorithm could help generative AI models produce high-quality outputs that obey strict requirements when “pretty close” doesn’t cut it.
AI may appear weightless, but every model depends on physical infrastructure. To understand responsible AI, we need to look beyond algorithms and consider the entire lifecycle of the hardware behind them. The post Responsible AI Must Consider Its Afterlife appeared first on GPT-Lab.
Adaptive AI agents can help make BIM data more machine-readable by navigating IFC models, interpreting inconsistent information, and mapping it to defined standards. In this blog, Alok Rawat shares findings from a real-world pilot in construction workflows. The post Adaptive AI Agents in Construction Workflows appeared first on GPT-Lab.
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.