Skip to content
Review Open access

Intelligent Business Document Processing Using AI- and NLP-Based Techniques: A Systematic Literature Review

Aug 2026 · Analytics · 0 citations · 41 references

TL;DR

The findings show that AI- and NLP-based methods have significantly improved the automation, retrieval, interpretation, and structuring of business documents, and large language models (LLMs), particularly when combined with prompt engineering, retrieval-augmented generation, knowledge graphs, and agent-based architectures, offer promising opportunities to address these gaps.

Abstract

This systematic literature review examines the application of artificial intelligence (AI) and natural language processing (NLP) techniques in intelligent business document processing. The study systematically analyses 46 peer-reviewed articles published between 2014 and 2025 and indexed in the Scopus database. The reviewed literature was grouped into six core NLP-based analytical tasks: semantic search, question answering, summarisation, text data integration and matching, event extraction, and business process management. The findings show that AI- and NLP-based methods have significantly improved the automation, retrieval, interpretation, and structuring of business documents. Semantic search methods enhance information retrieval by moving beyond keyword matching, while question-answering systems and summarisation techniques support automated knowledge discovery and content reduction. Deep learning and transformer-based models have also improved entity matching, event extraction, and predictive business process monitoring. However, the review identifies several persistent limitations, including the continued dominance of extractive approaches, limited adoption of abstractive summarisation, insufficient integration of knowledge graphs, fragmented system development, limited enterprise-scale validation, and a lack of reusable code and shared resources. The findings further indicate that large language models (LLMs), particularly when combined with prompt engineering, retrieval-augmented generation, knowledge graphs, and agent-based architectures, offer promising opportunities to address these gaps. Overall, this review highlights both the progress and remaining challenges in developing scalable, explainable, and domain-adaptable AI-driven systems for intelligent business document processing.

Read PDF

Similar papers

Review Open access Aug 2026

Optimizing NLP-Text Classification in Knowledge Management Systems: A Literature Review

An outline of the evolution of NLP-based text classification methods from initial machine learning methods such as Naïve Bayes and Support Vector Machines to current sophisticated deep learning algorithms such as Convolutional Neural Networks, Recurrent Neural Networks, and Transformers is offered.

Jenifer Mchory, Kelvin Kabeti Omieno, Collins Odoyo et al. · 0 citations
Open access Aug 2026

Large language model-based automated knowledge extraction and prediction system using Artificial Intelligence

This study presents an automated knowledge extraction and prediction system using the advancements in Artificial Intelligence (AI) tools, referred to as APEX-LLM, which is a scalable, domain-independent system which can be customized and applied to health, financial and business sectors, and education.

Jun Yin · 0 citations
Review Open access 2025

Large Language Models for Intelligent Research Knowledge Discovery and Automation

This study presents a conceptual framework that combines semantic retrieval, intelligent reasoning, automated literature analysis, and workflow orchestration, demonstrating how LLM-powered systems can transform scientific research into scalable, accurate, ethical, and collaborative knowledge discovery processes.

Narendra Karmarkar, Iyengar P. K. · 0 citations
Review Open access Aug 2026

Applications of Natural Language Processing: A Comprehensive Study

A comprehensive review of the evolution of NLP from traditional rule-based approaches to modern transformer models including BERT and GPT demonstrates that NLP continues to transform intelligent systems and is expected to play an increasingly significant role in the development of next-generation AI technologies.

P. Kalaiselvi · 0 citations
#large language models Open access Sep 2026

Research on Text Information Extraction and Imbalanced Classification Methods for Enterprise Profiling

This research focuses on enterprise profiling in scenarios where large volumes of diverse texts—such as registration records, annual reports, news articles, and bidding notices—are continuously generated. Instead of relying solely on a single data representation or classification model, we developed a comprehensive natural language processing (NLP) pipeline for extracting key information and identifying industries. The pipeline consists of several steps. First, we use a BERT-BiLSTM-CRF model to identify core e/nterprise entities. Then, we combine TF-IDF with BERT embeddings to create a hybrid feature scheme that captures both lexical cues and contextual semantics. To address the challenge of imbalanced industry labels, we apply SMOTE in the dense semantic space and pair it with Focal Loss to enhance learning for minority classes. Additionally, we introduce a Stacking strategy to integrate outputs from different models, making predictions more stable. Tests on a self-compiled dataset covering ten national economic sectors and about 50,000 enterprises show that our method achieves a macro-F1 score of 95.4%. It outperforms traditional machine learning baselines and single deep learning models, offering more reliable recognition for minority classes. These results suggest that our framework is well-suited for applications such as supply chain partner discovery, industrial mapping, and targeted investment promotion.

Xin-Yi Xu · 0 citations
Open access Aug 2026

Automated Classification of SAP Literature: Predicting Impact and Trends

The rapid evolution of business systems such as SAP (Systems, Applications, and Products in Data Processing) has generated a growing body of research on implementations, innovations, and business impacts. Determining high-impact papers and detecting emerging trends remains challenging due to the volume of literature. This study presents a machine learning powered pipeline for collecting, pre-processing, and classifying SAP-related research articles retrieved from Semantic Scholar. The pipeline employs natural language processing techniques, including text cleaning, lemmatization, and SciBERT embeddings, to generate richer feature representations, along with metadata features such as vocabulary diversity, novelty score, paper age, and citation velocity. To analyse research impact and trends, we trained a set of classical machine learning models, Random Forest, XGBoost, and LightGBM, and a set of large language models (LLMs), BERT, RoBERTa, and ELECTRA, fine-tuned for classification. The LLMs demonstrated superior performance compared to classical models, achieving accuracies of approximately 93% to 97% for impact classification and 96% to 97% for trend categorization. The models classify papers along two key dimensions: impact classification (High Impact, Niche, Low Impact) and trend categorization (Hot Trend, Recent Classic, Established, Historic), defined using proxy-based bibliometric indicators. This contribution provides an automated framework for literature analysis in the SAP context, enabling researchers and practitioners to identify high-impact studies and support the identification of emerging research directions.

Kawkab Bouressace, Tamás Orosz · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.