This paper constructs TransBERT, which uses channel attention and bidirectional encoding representation of character-level features to explore the information in text, and shows that the performance of TransBERT is better than the current most advanced baseline model.
In order to solve the problems of domain entity lack and high noise, little human supervision in low resource specialized corpora, this work adopts the method of multi-feature fusion of dynamic domain dictionary expansion. The method includes dictionary matching, semantic confidence filtering and confidence-weighted na...
A comprehensive natural language processing (NLP) pipeline for extracting key information and identifying industries, which outperforms traditional machine learning baselines and single deep learning models, offering more reliable recognition for minority classes.
Xin-Yi Xu· Applied and Computational En...· 0 citations
This study validates the effectiveness of the BERT model in semantic similarity calculation, providing more accurate technical support for related application scenarios, and laying the foundation for subsequent model optimization and lightweighting research.
Jia-Cheng Gao· International Conference on...· 0 citations
To address the challenges of diverse domain-specific terminology, highly colloquial expressions, and limited annotated samples in sentiment analysis of stock forum texts, this study proposes an ERNIE-Transformer sentiment classification model that integrates ERNIE and Transformer architectures. First, a systematic data...
Xiu-Mei Li, Fei Chen, Wen-Chao Ling et al.· Journal of Electrical System...· 0 citations
Unified modeling poses significant challenges for named entity recognition (NER), where effectively modeling contextual semantics and achieving efficient feature fusion remain critical challenges. The Word-Word Relation Classification for Named Entity Recognition (W
2
NER) framework, based on word-word relationship...
Financial statement information extraction and technical document data management both face significant challenges due to complex formats, scattered layouts, and intricate proximity-based semantic relations within semi-structured documents. These challenges are particularly evident in engineering enterprises involving...