Skip to content
Review Open access

Optimizing NLP-Text Classification in Knowledge Management Systems: A Literature Review

Aug 2026 · Engineering and Technology Journal · Vol 11 · 0 citations

TL;DR

An outline of the evolution of NLP-based text classification methods from initial machine learning methods such as Naïve Bayes and Support Vector Machines to current sophisticated deep learning algorithms such as Convolutional Neural Networks, Recurrent Neural Networks, and Transformers is offered.

Abstract

Knowledge Management Systems (KMS) are required to organize and assign meaning to huge amounts of organizational knowledge that are largely in the form of unstructured text. Natural Language Processing (NLP), and more immediately methods of text categorization, has been one of the principal enabler technologies to enable KMS to be simpler by helping to automatically categorize documents, enhance searching for information, and assist in decision-making. This paper offers an outline of the evolution of NLP-based text classification methods from initial machine learning methods such as Naïve Bayes and Support Vector Machines to current sophisticated deep learning algorithms such as Convolutional Neural Networks, Recurrent Neural Networks, and Transformers. We offer real-world industry use cases, issues of scalability, explainability, and ethics and encapsulate research areas of existing gaps. The findings underscore the enormous potential of NLP text classification to assist the effectiveness and efficiency of knowledge management (KM) activities.

Read PDF

Similar papers

Review Open access Aug 2026

Intelligent Business Document Processing Using AI- and NLP-Based Techniques: A Systematic Literature Review

The findings show that AI- and NLP-based methods have significantly improved the automation, retrieval, interpretation, and structuring of business documents, and large language models (LLMs), particularly when combined with prompt engineering, retrieval-augmented generation, knowledge graphs, and agent-based architectures, offer promising opportunities to address these gaps.

Naif N. Alotaibi, Morteza Saberi, M. Bandara et al. · 0 citations
Open access 2019

Intelligent Document Classification Using Neural Networks

Subsequent to smart document categorization has become a core operation in contemporary information discovery, online libraries, business information management and big-data analytics. As the amount of unstructured textual information in terms of academic repositories, social media sites, corporate archives and government databases grows exponentially, automated and intelligent classification methods are critical in terms of efficient data organization and knowledge discovery. Although effective in limited situations, traditional rule-based and statistical machine learning methods fail to scale and generalize when faced with semantic ambiguity, contextual differences and high-dimensional feature spaces. Neural networks and especially deep learning models have proven themselves able to extract semantic representations and contextual relationships in text data remarkably. In this paper, the intelligent document classification using neural networks will be studied in a very comprehensive manner in terms of the architecture, features representation, training techniques, and evaluation procedures. The framework proposed combines text preprocessing, embedding, and neural classification models to obtain robust and scalable document classification. An overall experimental study is performed based on benchmark datasets in an attempt to determine the accuracy of classification, the precision, recall, and the computational efficiency. The findings show that neural network methods are very effective compared to traditional methods particularly when dealing with large and complicated document collections. The paper concludes by mentioning practical implications, challenges, and future research directions in the area of neural document classification.

Kwame Nkosi · 0 citations
#large language models Open access Sep 2026

Research on Text Information Extraction and Imbalanced Classification Methods for Enterprise Profiling

This research focuses on enterprise profiling in scenarios where large volumes of diverse texts—such as registration records, annual reports, news articles, and bidding notices—are continuously generated. Instead of relying solely on a single data representation or classification model, we developed a comprehensive natural language processing (NLP) pipeline for extracting key information and identifying industries. The pipeline consists of several steps. First, we use a BERT-BiLSTM-CRF model to identify core e/nterprise entities. Then, we combine TF-IDF with BERT embeddings to create a hybrid feature scheme that captures both lexical cues and contextual semantics. To address the challenge of imbalanced industry labels, we apply SMOTE in the dense semantic space and pair it with Focal Loss to enhance learning for minority classes. Additionally, we introduce a Stacking strategy to integrate outputs from different models, making predictions more stable. Tests on a self-compiled dataset covering ten national economic sectors and about 50,000 enterprises show that our method achieves a macro-F1 score of 95.4%. It outperforms traditional machine learning baselines and single deep learning models, offering more reliable recognition for minority classes. These results suggest that our framework is well-suited for applications such as supply chain partner discovery, industrial mapping, and targeted investment promotion.

Xin-Yi Xu · 0 citations
Review Open access Aug 2026

Applications of Natural Language Processing: A Comprehensive Study

A comprehensive review of the evolution of NLP from traditional rule-based approaches to modern transformer models including BERT and GPT demonstrates that NLP continues to transform intelligent systems and is expected to play an increasingly significant role in the development of next-generation AI technologies.

P. Kalaiselvi · 0 citations
Open access Aug 2026

Large language model-based automated knowledge extraction and prediction system using Artificial Intelligence

This study presents an automated knowledge extraction and prediction system using the advancements in Artificial Intelligence (AI) tools, referred to as APEX-LLM, which is a scalable, domain-independent system which can be customized and applied to health, financial and business sectors, and education.

Jun Yin · 0 citations
Open access 2021

Deep Learning Models for Document Classification

Document classification is an essential process of natural language processing (NLP) which presupposes assigning textual documents to predefined categories depending on their contents. As the amount of digital text that needs to be classified has grown exponentially through the sources of social media, scholarly repositories, law archives, news portals and enterprise document management systems, effective and correct document classification has become more important. Most of the common machine learning models such as Naïve Bayes, Support Vector Machines, and k-Nearest Neighbors have proven to be of acceptable performance but they heavily depend on manually crafted features and are not as accurate in detecting semantic and contextual information in text. The latest technology in deep learning greatly altered the methods of document classification as it allowed extracting features and learning representations based on the context. Convolutional neural networks (CNNs) models, recurrent neural networks (RNNs), Long Short Memory networks (LSTMs), Gated Recurrent Units (GRUs) and transformer-based have been used to set the state of the art on benchmark datasets. Such models employ dense word encodings, attention, and hierarchical models to represent document-level semantic structures on both local and global levels. In the present paper, the systematic investigation of the document classification frameworks using deep learning models is offered. It analyzes background information, architectural design, learning process and optimization plans. Moreover, it evaluates the advantages and weaknesses of the different deep learning methods in processing long texts, multi-label classification, domain adaptation, and scalability. The socio-economic metrics are also standard performance metrics, and a single approach to the methodology is suggested that incorporates preprocessing, embedding learning, model training, and evaluation using these metrics. The results of the experiment when exploring representative datasets are addressed to emphasize the trends of the comparative performance. The paper will end by presenting some of the current challenges and the direction of future research and stressing the aspects of explainability, efficiency, and domain robustness.

S. Rahman · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.