KSEC (Knowledge-enhanced Splitting Error Corrector), a novel framework tailored for variable-length corrections, is proposed that achieves state-of-the-art performance among lightweight models of similar size and outperforms existing methods across multiple evaluation metrics.
Abstract
Chinese Spelling Correction (CSC) is a fundamental task in Natural Language Processing (NLP) aimed at identifying and correcting character errors in Chinese texts. It significantly enhances text readability and semantic accuracy. Most deep learning-based CSC methods focus on isometric correction, ensuring identical lengths for input and output sequences. However, they struggle with variable-length errors like splitting errors—where a single character is incorrectly divided into two (e.g., splitting “明” into “日” and “月”). These errors are challenging because they disrupt token alignment, preventing standard sequence-labeling models from mapping inputs to outputs effectively. To overcome this limitation, we propose KSEC (Knowledge-enhanced Splitting Error Corrector), a novel framework tailored for variable-length corrections. KSEC automatically constructs a splitting character knowledge base from public corpora to provide factual validation for correction outcomes. Furthermore, we design a variable-length architecture integrating an attention mechanism and introduce an alignment-aware loss function that optimizes sequence-to-sequence token mapping. Extensive experiments on standard CSC and CSEC benchmarks demonstrate that KSEC achieves state-of-the-art performance among lightweight models of similar size and outperforms existing methods across multiple evaluation metrics.
Improved robustness and explainability for automatic Chinese learner error correction is demonstrated by an alignment consistency loss to ensure character-level consistency, and the combination of word-segmentation augmentation and multi-reference soft-label training to reduce conflicts caused by segmentation differenc...
Y.-L. Diao, W. Gao· Advanced Electromagnetics· 0 citations
Tokenization forms the foundation of modern Natural Language Processing (NLP) systems by transforming raw text into discrete units that neural language models can process. The effectiveness of this process directly influences vocabulary efficiency, sequence length, computational cost, and downstream model performance....
V. HariKrishnanK, Sudarsun Santhiappan· 0 citations
This paper proposes an alignment-aware multi-granularity tagging framework. First, this method uses a cross-lingual pre-trained model to encode source and target language contexts jointly while explicitly modeling-level bilingual correspondences via a learnable soft alignment layer. Second, a gated local enhancement mo...
Bi-Juan Wang, Lingli Zhu, Hongli Wen· International Journal of Inf...· 0 citations
Advanced large language models (LLMs) with long context windows can substantially reduce input truncation in document-level machine translation (DocMT). However, direct Doc2Doc translation remains prone to n-gram repetition and progressive quality degradation. A common remedy is to segment the document into finer-grain...
Xiao-Tian Wang, You-Yuan Lin, Zhan Shen et al.· 0 citations
Vocabulary integration is an effective approach for improving named entity recognition performance. However, existing methods exhibit insufficient decoupling capability in modeling the position and content between characters and words, and under static matching mechanisms, models tend to memorize fixed character–word c...
Mingjian Li, Fei Gao, Nuo Qun et al.· Applied Sciences· 0 citations
Before using GenAI models as EdTech tools, their pedagogical suitability should be corroborated. In this paper, we present ShAnEL-2 , a novel multilingual dataset comprising 1,185 student responses to short-answer language learning exercises corrected by teachers. We use ShAnEL-2 to establish an initial benchmark of (1...
Jasper Degraeuwe, Thomas Moerman· International Conference on...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.