2026· Schriften aus der Fakultät Wirtschaftsinformatik und Angewandte Informatik der Otto-Friedrich-Universität Bamberg· 0 citations
TL;DR
The role of traditional models in current NLP research and practice is discussed, especially in contrast and comparison to modern neural network-based approaches including LLMs.
Abstract
The research field of Natural Language Processing (NLP) has experienced a major shift since the introduction of Large Language Models (LLMs). All facets and application scenarios within NLP have been impacted by the use of LLMs. Current research as well as practice of text processing tools is focused mainly on the application and development of LLMs. Major investments, not only by LLM providers but also other companies applying LLMs in their workflows, have only solidified the role of LLMs in NLP - and in other research and application areas - as part of the artificial intelligence boom in recent years.
However, limitations and downsides of the application of LLMs have also emerged. Problems regarding the generated texts as well as the environmental impact of the large-scale use of LLMs are just two of many factors that should be critically analyzed, despite the hype and the prevalence of LLMs for NLP tasks. These restrictions provide the main motivation for this thesis. Traditional models as alternatives to LLMs will be discussed from different perspectives. The characterization of traditional models will be progressively developed as features of alternatives to LLMs will emerge during the course of this thesis. This process will be grounded in experiments, observations and evaluations. Several NLP applications will be presented by surveying the state of the art with neural network-based models such as LLMs as well as the current usage of traditional models. The concrete NLP applications comprise information and relation extraction, text classification, text segmentation, text simplification and text summarization.
The first half of this thesis will present the emergence of LLMs contextualized along previous developments within NLP. Characteristics of the selected NLP applications will be collected before a structured literature review will display the prevalence of LLMs regarding each application and will discuss if traditional models are still actively researched. Lessons from domains with long-standing development procedures and processes will also be taken into account to provide a purposeful and structured manner of approaching NLP tasks. A collection of challenges within current NLP will conclude the first half of the thesis, which will serve as motivation for the analysis of experiments and applications of the latter half.
The second half of this thesis will present observations and evaluations from use cases, aligned towards the challenges recognized in the first half. Through the analysis of these use cases, benefits of applying traditional models will be collected and supported, in particular through the analysis of a text segmentation use case that is purposefully applied with the lessons drawn from the first half of the thesis in mind. The interpretation of these results will conclude in a discussion on the applicability of traditional models in contrast to LLMs and also give recommendations of both model types for different use cases.
Concrete use cases for information extraction, entity matching, text classification and text segmentation will be presented, in which traditional models match or surpass the performance of modern methods. Through improved efficiency as well as enhanced explainability and reproducibility in comparison with neural network-based techniques, these showcases demonstrate the continued relevancy of traditional techniques in today's NLP landscape.
Overall, this thesis discusses the role of traditional models in current NLP research and practice, especially in contrast and comparison to modern neural network-based approaches including LLMs. The applicability of modern and less modern techniques is analyzed through a case-based analysis of NLP tasks in a structured and purposeful manner.
A comprehensive review of the evolution of NLP from traditional rule-based approaches to modern transformer models including BERT and GPT demonstrates that NLP continues to transform intelligent systems and is expected to play an increasingly significant role in the development of next-generation AI technologies.
P. Kalaiselvi· International Journal of Eme...· 0 citations
Whether contemporary LLMs can reproduce the research outcomes of a fully documented human study: a 1991 article that identified dermatophytosis (ringworm) in historical fine art was evaluated.
Large Language Models (LLMs) are being increasingly used in everyday applications. A major challenge in the context of LLMs or Artificial Intelligence (AI) in general is to ensure privacy when using them, meaning that personally identifiable information (PII) is removed from any text that enters an LLM. These challenges have become more urgent with novel EU legislation. Uncertainty around LLM usage with respect to privacy concerns in EU countries can be a major blocker for the speed of innovation and transfer from research to applications. Here we present \textbf{Redakto}, a tool that can be used for anonymizing text prior to feeding it to an LLM or other downstream text processing. We provide state-of-the-art functionalities for both redaction of PII but also when used for pseudonymization. These functionalities are exposed such that they can easily be used by end-users, through the Redakto web application, and by developers and researchers, via REST APIs and model context protocol (MCP) hooks. The implementation is fully open source, requires modest compute resources, and can be readily deployed on local hardware. In contrast to prior work and in order to better assess the quality of the anonymized texts, we conduct extensive empirical evaluations on textual data from legal and medical domain with respect to both privacy and utility of the redacted texts. Our empirical results demonstrate that the texts anonymized with different redaction strategies achieve utility scores on par with the original texts, suggesting that anonymization with Redakto can be used for LLM tasks without substantial negative impact for the tasks we explored.
This survey examines how these methods integrate graphs into various stages of the LLM pipeline, including the input, model, and output phases, and outlines the challenges and future research directions for developing more efficient and interpretable solutions.
Xinyan Zhu, Cheng Yang, Qiuyue Wang et al.· 0 citations
It is argued that the ability of long context should not only come from increasing the context window, but also from the ability of the model to locate, integrate and reason about important information in long text.
Jun-Hao Wu· Applied and Computational En...· 0 citations
The recent meteoric rise of LLMs (Large Language Models) and associated tools was largely unexpected and surprising to most. The rapid ascent of this technology has caught many software developers unawares, leaving them suddenly somewhat ignorant, and arguably under-skilled.
LLMs, whilst still advancing, have recently demonstrated impressive capabilities in their ability to assist software developers in their day-to-day tasks (e.g., coding new features, and locating and fixing issues). However, the use and adoption of LLMs presents many larger challenges for society as a whole; many of which are not in themselves technical concerns.
This paper examines the current and perceived impact of this technology in the context of Open Source. We identify several social, economic, environmental, political, legal, and technical concerns regarding the use of LLMs in Open Source projects.
We contribute guidance around defining an AI Policy for Open Source projects. We further offer an AI Policy Score Card to assist projects in clearly defining and declaring how they wish to work with AI or not.
Adam Retter· Balisage Series on Markup Te...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.