Jul 2026· Dandao Xuebao/Journal of Ballistics· Vol 38, pp. 306-314· 0 citations
TL;DR
An automated research gap detection system that integrates natural language processing, citation network analysis, and ensemble machine learning to identify research gaps across scientific literature systematically is presented.
Abstract
The exponential growth of scientific publications creates a critical challenge for researchers attempting to navigate their fields. Manual literature reviews, once sufficient for identifying research opportunities, now consume disproportionate time and often lack comprehensiveness. This paper presents an automated research gap detection system that integrates natural language processing, citation network analysis, and ensemble machine learning to identify research gaps across scientific literature systematically. The proposed system uses transformer-based models (SciBERT, BioBERT) for semantic understanding, graph neural networks for citation structure analysis, and Support Vector Machines, Random Forests, and Gradient Boosting for gap classification. We implement a complete pipeline processing documents at scale, extracting semantic content, analyzing citation relationships, and identifying knowledge gaps through multiple complementary techniques. The system was designed to detect three categories of gaps: knowledge discrepancies (conflicting information), knowledge voids (completely missing information), and methodological limitations (inadequate research methods). Experimental evaluation across multiple scientific domains demonstrates that the system identifies research opportunities with accuracy comparable to expert assessments while processing millions of documents efficiently. The results show 78% accuracy in predicting emerging research areas up to two years in advance and 88% validation rate for identified gaps when reviewed by domain experts.
The rapid evolution of business systems such as SAP (Systems, Applications, and Products in Data Processing) has generated a growing body of research on implementations, innovations, and business impacts. Determining high-impact papers and detecting emerging trends remains challenging due to the volume of literature. This study presents a machine learning powered pipeline for collecting, pre-processing, and classifying SAP-related research articles retrieved from Semantic Scholar. The pipeline employs natural language processing techniques, including text cleaning, lemmatization, and SciBERT embeddings, to generate richer feature representations, along with metadata features such as vocabulary diversity, novelty score, paper age, and citation velocity. To analyse research impact and trends, we trained a set of classical machine learning models, Random Forest, XGBoost, and LightGBM, and a set of large language models (LLMs), BERT, RoBERTa, and ELECTRA, fine-tuned for classification. The LLMs demonstrated superior performance compared to classical models, achieving accuracies of approximately 93% to 97% for impact classification and 96% to 97% for trend categorization. The models classify papers along two key dimensions: impact classification (High Impact, Niche, Low Impact) and trend categorization (Hot Trend, Recent Classic, Established, Historic), defined using proxy-based bibliometric indicators. This contribution provides an automated framework for literature analysis in the SAP context, enabling researchers and practitioners to identify high-impact studies and support the identification of emerging research directions.
Kawkab Bouressace, Tamás Orosz· Acta Technica Jaurinensis· 0 citations
The lack of high-quality ground truth datasets to train machine learning (ML) models impedes the potential of artificial intelligence for science research. Scientific information extraction (SIE) from the literature using LLMs is emerging as a powerful approach to automate the creation of these datasets. However, existing LLM-based approaches and benchmarking studies for SIE focus on broad topics such as biomedicine and chemistry, are limited to choice-based tasks, and focus on extracting information from short and well-formatted text. The potential of SIE methods in complex, open-ended tasks is considerably under-explored. In this study, we use a domain that has been virtually ignored in SIE, namely virology, to address these research gaps. We design a unique, open-ended SIE task of extracting mutations in a given virus that modify its interaction with the host. We develop a new, multi-step retrieval augmented generation (RAG) framework called VILLA for SIE. In parallel, we curate a novel dataset of 629 mutations in ten influenza A virus proteins obtained from 293 scientific publications to serve as ground truth for the mutation extraction task. We demonstrate VILLA's superior performance using a novel and comprehensive evaluation and comparison with vanilla RAG and other state-of-the art RAG- and agent-based tools for SIE. Finally, we evaluate the generalizability of our proposed method using another unique dataset of mutations in hepatitis E virus.
Blessy Antony, Amartya Dutta, Sneha Aggarwal et al.· Proceedings of the 32nd ACM...· 0 citations
This research focuses on enterprise profiling in scenarios where large volumes of diverse texts—such as registration records, annual reports, news articles, and bidding notices—are continuously generated. Instead of relying solely on a single data representation or classification model, we developed a comprehensive natural language processing (NLP) pipeline for extracting key information and identifying industries. The pipeline consists of several steps. First, we use a BERT-BiLSTM-CRF model to identify core e/nterprise entities. Then, we combine TF-IDF with BERT embeddings to create a hybrid feature scheme that captures both lexical cues and contextual semantics. To address the challenge of imbalanced industry labels, we apply SMOTE in the dense semantic space and pair it with Focal Loss to enhance learning for minority classes. Additionally, we introduce a Stacking strategy to integrate outputs from different models, making predictions more stable. Tests on a self-compiled dataset covering ten national economic sectors and about 50,000 enterprises show that our method achieves a macro-F1 score of 95.4%. It outperforms traditional machine learning baselines and single deep learning models, offering more reliable recognition for minority classes. These results suggest that our framework is well-suited for applications such as supply chain partner discovery, industrial mapping, and targeted investment promotion.
Xin-Yi Xu· Applied and Computational En...· 0 citations
The fast development of Large Language Models is a problem for keeping academic integrity in scientific publishing. The usual tools that detect this kind of thing use statistics like perplexity and linguistic features. These tools demonstrate limited effectiveness against sophisticated domain-specific AIgenerated text. This paper presents ResearchNet, a hybrid detection framework for identifying whether scientific text is human-authored or LLM-generated. ResearchNet uses a kind of encoder called Frozen SciBERT and a Graph Convolutional Network, which looks at text as a graph where the sentences are connected by logical transitions. It also uses something called DeepScientificAttention to combine information about terminology and citations. The structure of the text to make a strong classification. ResearchNet was evaluated on the AIGTxt dataset across ten scientific fields such as Astrophysics, Medicine and Social Sciences ResearchNet achieves a ROC AUC of 90.02 percent and the highest Mixed-class F1 of 0.69 which is better than all the other models we compared it to and it was 2.7 percent points better, than the next best model, which was SciBERT+GCN.
Rola Islait, M. Alhawamdeh· IEEE Jordan Conference on Ap...· 0 citations
This study presents a conceptual framework that combines semantic retrieval, intelligent reasoning, automated literature analysis, and workflow orchestration, demonstrating how LLM-powered systems can transform scientific research into scalable, accurate, ethical, and collaborative knowledge discovery processes.
Narendra Karmarkar, Iyengar P. K.· International Journal of Eme...· 0 citations
This work presents a scalable, reproducible framework for evaluating, optimizing, and interpreting LLMs for biomedical knowledge extraction, with a focus on gene–gene regulatory relation prediction, pathway component recognition, multimodal pathway figure understanding, and automated prompt optimization.
Muhammad Azam· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.