2026· International journal of engineering and technology· 0 citations· 15 references
TL;DR
This study clarifies the evolution of Chinese word segmentation technology, providing a reference for the selection, engineering implementation, and optimization of word segmentation algorithms in the large model era, and is highly valuable for advancing the high-quality development of Chinese natural language processing.
Abstract
— With the rapid development of large language models and the deep integration of artificial intelligence into various industries the requirements for accuracy and generalization ability in Natural Language Processing (NLP) are constantly rising. Large models centered on Transformers are propelling NLP into a new stage. Chinese lacks natural word boundaries, making word segmentation a fundamental step in Chinese NLP, and its accuracy directly determines the effectiveness of subsequent tasks. Due to its linguistic characteristics, Chinese word segmentation has long faced three core challenges: inconsistencies between general vocabularies and segmentation standards, difficulties in handling ambiguous segments, and poor performance in out-of-vocabulary word identification. The paper highlights the advantages of deep learning for word segmentation, detailing classic neural networks such as CNN, RNN, LSTM, and BiLSTM-CRF, as well as the application of pre-trained models including BERT, RoBERTa, and lightweight real-time models. The paper emphasizes the advantages of deep learning in word segmentation, detailing classic neural network models such as CNN, RNN, LSTM, and BiLSTM CRF, as well as the application of BERT, RoBERTa pre-trained models, and lightweight real-time models in word segmentation. Research shows that deep learning-based word segmentation methods offer the best overall performance, effectively solving challenges in ambiguous segmentation and out-of-vocabulary word recognition. Different algorithms and systems can meet the diverse needs of scientific research, industry, and vertical fields. This study clarifies the evolution of Chinese word segmentation technology, providing a reference for the selection, engineering implementation, and optimization of word segmentation algorithms in the large model era, and is highly valuable for advancing the high-quality development of Chinese natural language processing
This study integrates parallel multi-kernel word-level convolutional features into conventional and hybrid deep learning models for Arabic text analysis tasks, providing a systematic within-study assessment of model sensitivity to architecture, preprocessing, and learning-rate selection.
Ahmed I.Taloba, George Samy Rady, Khaled F. Hussain· International Journal of Adv...· 0 citations
This paper addresses the task of automatic identification of German word forms and constructs an end-to-end model framework based on deep neural networks that employs character-level and subword-level dual-channel feature representations, and combines encoder-decoder architecture, scaled dot-product attention, and positional encoding.
Bo Wang· International Conference on...· 0 citations
This study validates the effectiveness of the BERT model in semantic similarity calculation, providing more accurate technical support for related application scenarios, and laying the foundation for subsequent model optimization and lightweighting research.
Jiachen Gao· International Conference on...· 0 citations
Current advances in neural network models have improved state-of-the-art performance in natural language processing tasks such as named-entity recognition, sentiment analysis, and machine translation. In particular, neural language models are applied to encode information in word embeddings. These approaches are generally trained on large corpora using semi-supervised learning. Word embeddings encode the syntactic and semantic properties of words as dense vectors. In agglutinative languages such as Turkish, Finnish, and Hungarian, word-embedding construction is challenging because extensive suffixation and polysemy can cause information loss. To overcome these limitations, character n-grams are often preferred for embedding representations. Nevertheless, character n-grams do not guarantee the capture of information in long word sequences. In this study, a method that partitions word sequences according to frequent patterns within a given context is proposed for training a neural language model. In this respect, likelihood- and ranking-based inference are combined with n-gram and syllable partitioning for word-embedding generation from a text corpus. The proposed approach provides a language-agnostic, context-sensitive segmentation mechanism that can complement language processing methods such as lemmatization, morphological analysis, and stemming. For embedding generation, the SkipGram and FastText models are used, and the effects of word partitioning are evaluated using analogy, named-entity recognition, POS tagging, sentiment analysis, and morphological disambiguation datasets for Turkish. The results indicate task-dependent and generally limited improvements over traditional token-based word-embedding extraction. In particular, skip n-gram partitioning produces a substantial improvement over partitioning based on frequent-ngrams, sentencepiece-bpe, sentence-unigram and morfessor. No consistent relationship was observed across tasks between performance and either graph density or the average number of distinct n-grams per sentence.
With the rapid development of social media, a large number of users' comments are constantly generated in the network, and some harmful content containing sensitive information seriously affects the network environment. The development of deep learning has promoted research and governance in this field. This article reviews deep learning-based sensitive word recognition methods from three categories: Convolutional Neural Network (CNN), Recurrent Neural Network (RNN), and Graph Neural Network (GNN), and summarizes the classic public data sets in this field. In addition, it also points out the characteristics of each method and the problems to be solved in this field, such as the lack of public data resources, the weakness of multimodal comprehensive analysis, and the dialect bias of the model. This article aims to systematically review the research progress in the field of sensitive word recognition, summarize the applicable scenarios and limitations of different deep learning methods, and provide useful references for subsequent research.
A comprehensive review of the evolution of NLP from traditional rule-based approaches to modern transformer models including BERT and GPT demonstrates that NLP continues to transform intelligent systems and is expected to play an increasingly significant role in the development of next-generation AI technologies.
P. Kalaiselvi· International Journal of Eme...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.