Skip to content
Conference Open access

Implementation of BERTopic for Topic Modelling on KNKT’s Aviation Accident Investigation Reports

Aug 2026 · ICEETE Conference Series · Vol 4, pp. 944-950 · 0 citations

Abstract

Aircraft accident investigation reports contain important information regarding the chronology of events, findings, contributing factors, and safety recommendations that can be used to understand accident patterns. However, this information is generally presented in the form of unstructured text, making manual analysis less efficient. This research aims to apply BERTopic to identify latent themes in aviation accident investigation reports published by the National Transportation Safety Committee (KNKT). A total of 182 investigation reports classified as Final Reports were processed through text extraction, corpus formation, and preprocessing, resulting in 172 documents used for topic modelling. BERTopic was implemented using sentence embedding, UMAP dimensionality reduction, HDBSCAN clustering, and c-TF-IDF-based topic representation. The modelling results yielded ten main topics reflecting various aspects of aviation safety, including technical, operational, and human factors. Evaluation showed that BERTopic achieved a topic coherence (Cv) value of 0.6111 and generated more specific keywords compared to Latent Dirichlet Allocation (LDA). The research results indicate that BERTopic is capable of effectively extracting latent themes from KNKT investigation reports and has the potential to support the analysis of aviation accident patterns and serve as a foundation for the development of domain knowledge-based research in the field of aviation safety.

Read PDF

Similar papers

Open access Aug 2026

TMCAS: Efficient Large Language Model-Assisted Topic Modeling for Civil Aviation Safety Reports

This paper proposes TMCAS, an efficient large language model-assisted topic modeling framework for civil aviation safety reports that achieves superior clustering and interpretability while substantially reducing inference cost compared with document-wise LLM baselines.

Xiangge Li, Haofeng Wang, Xiuting Zhou et al. · 0 citations
#large language models Open access Sep 2026

Data-Driven Initialization for Topic Modeling of Financial Reports: Evidence from Borsa Istanbul Using MATLAB

This study analyzes annual reports of firms listed on Borsa I​stanbul (BIST) for the period from 1990 to 2026, and demonstrates that data-driven initia’sation significantly improves topi’sc interpre⁠tability, reduces optimizat​ion iterations, decreases model unce​rtainty, and e​n’hances reproducibility compared with tr​aditional initializa’on st​rategies.

Unknown authors · 0 citations
Review Open access Aug 2026

Decision-oriented explainable artificial intelligence: a PDR-based review of methods, applications, and emerging frontiers

Explainable artificial intelligence (XAI) is increasingly used in consequential decisions, but method selection still requires evidence about predictive reliability, explanatory fidelity, and stakeholder usefulness. This latent Dirichlet allocation (LDA)-assisted structured review used the OpenAlex Works metadata application programming interface (API) to retrieve 4,200 English-language records published from 2018 to 13 July 2026. Sequential DOI, exact-title, and fuzzy-title deduplication retained 3,166 records; title, abstract, and keyword screening retained 694; eligibility assessment retained 666; and 666 records entered the final document-term matrix. Candidate LDA models with k = 5 − 16 were compared using c_v and u_mass coherence, conventional perplexity, matched-topic stability, and Jensen–Shannon separation. The selected k = 8 solution achieved c_v = 0.5226, stability = 0.6809, and mean topic separation = 0.5824. Manual inspection of the top terms and documents identified themes concerning robust counterfactual generation, responsible and regulated decision support, interpretable clinical risk prediction, industrial, cybersecurity, and infrastructure XAI, causal and actionable algorithmic recourse, visual and multimodal medical explanation, clinical XAI adoption, trust, and workflow, and human-centered XAI evaluation and taxonomy. These evidence-derived themes are integrated with the adopted Predictive-Descriptive-Relevance (PDR) framework as a decision-oriented synthesis lens. The review contributes a reproducible literature map, a critical method comparison, a stakeholder-centered cross-domain matrix, and a research roadmap for causal, robust, uncertainty-aware, multimodal, generative, foundation-model XAI.

Jiangshan Zhu · 0 citations

Automatic Classification of Industrial Accident Causes using NLP: A Case Study in eMARS

This study proposes a lightweight Natural Language Processing (NLP) pipeline to automatically classify the primary cause of major accidents using the European eMARS database and shows that Word2Vec+SVM provides the strongest and most stable baseline on the full labelled set, while SBERT performance improves markedly under higher label fidelity.

Valerio Cozzani, B. Fabiano, G. Reniers et al. · 0 citations
Jul 2026

Integrating FRACAS and FMECA with Natural Language Processing (NLP): An AI-Assisted Approach to Reliability Analysis

An algorithm is developed that automates and streamlines the analysis of equipment field-failure reports and other unstructured maintenance records and reduces the resource-intensive manual work required to prepare, interpret, and process FRACAS reports, thus enabling timelier, data-driven equipment reliability analysis.

Esther Yu, Guangjiang Cao, Y. Khalil et al. · 0 citations
Open access Jul 2026

RAFE-XAI: A Retrieval-Augmented Feature Engineering and Explainable NLP Framework for Urban Infrastructure Risk Classification

This study introduces RAFE-XAI, a retrieval-augmented feature engineering and explainable natural language processing framework for urban infrastructure risk classification that incorporates semantic sentence embeddings, retrieval-based evidence, neighborhood-derived label distributions, domain-specific risk indicators, infrastructure asset cues, location indicators, and evidence-based explainability.

Abdulaziz Almaleh, Abdullah M. Alqahtani · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.