Skip to content
Open access

KSEC: A Knowledge-Enhanced Approach for Variable-Length Chinese Spelling Correction

Jul 2026 · Electronics · Vol 15, pp. 3238 · 0 citations · 33 references

TL;DR

KSEC (Knowledge-enhanced Splitting Error Corrector), a novel framework tailored for variable-length corrections, is proposed that achieves state-of-the-art performance among lightweight models of similar size and outperforms existing methods across multiple evaluation metrics.

Abstract

Chinese Spelling Correction (CSC) is a fundamental task in Natural Language Processing (NLP) aimed at identifying and correcting character errors in Chinese texts. It significantly enhances text readability and semantic accuracy. Most deep learning-based CSC methods focus on isometric correction, ensuring identical lengths for input and output sequences. However, they struggle with variable-length errors like splitting errors—where a single character is incorrectly divided into two (e.g., splitting “明” into “日” and “月”). These errors are challenging because they disrupt token alignment, preventing standard sequence-labeling models from mapping inputs to outputs effectively. To overcome this limitation, we propose KSEC (Knowledge-enhanced Splitting Error Corrector), a novel framework tailored for variable-length corrections. KSEC automatically constructs a splitting character knowledge base from public corpora to provide factual validation for correction outcomes. Furthermore, we design a variable-length architecture integrating an attention mechanism and introduce an alignment-aware loss function that optimizes sequence-to-sequence token mapping. Extensive experiments on standard CSC and CSEC benchmarks demonstrate that KSEC achieves state-of-the-art performance among lightweight models of similar size and outperforms existing methods across multiple evaluation metrics.

Read PDF

Similar papers

Open access Aug 2026

BERT-based Automatic Error Correction System for Chinese Learners

Improved robustness and explainability for automatic Chinese learner error correction is demonstrated by an alignment consistency loss to ensure character-level consistency, and the combination of word-segmentation augmentation and multi-reference soft-label training to reduce conflicts caused by segmentation differenc...

Y.-L. Diao, W. Gao · 0 citations
#natural language process... Preprint Sep 2026

A Grapheme-Aware Indic Tokenizer for Tamil: Large-Scale Training and Intrinsic Evaluation

Tokenization forms the foundation of modern Natural Language Processing (NLP) systems by transforming raw text into discrete units that neural language models can process. The effectiveness of this process directly influences vocabulary efficiency, sequence length, computational cost, and downstream model performance....

V. HariKrishnanK, Sudarsun Santhiappan · 0 citations
Open access Aug 2026

Improved Sequence Labeling Algorithms and Their Applications for Translation Error Detection

This paper proposes an alignment-aware multi-granularity tagging framework. First, this method uses a cross-lingual pre-trained model to encode source and target language contexts jointly while explicitly modeling-level bilingual correspondences via a learnable soft alignment layer. Second, a gated local enhancement mo...

Bi-Juan Wang, Lingli Zhu, Hongli Wen · 0 citations
#natural language process... Preprint Sep 2026

Doc2FRC: Length-Consistent Document-Level Machine Translation via Fixed-Range Chunking

Advanced large language models (LLMs) with long context windows can substantially reduce input truncation in document-level machine translation (DocMT). However, direct Doc2Doc translation remains prone to n-gram repetition and progressive quality degradation. A common remedy is to segment the document into finer-grain...

Xiao-Tian Wang, You-Yuan Lin, Zhan Shen et al. · 0 citations
Open access Aug 2026

Decoupled Attention and Character–Word Mask for Chinese Nested Named Entity Recognition

Vocabulary integration is an effective approach for improving named entity recognition performance. However, existing methods exhibit insufficient decoupling capability in modeling the position and content between characters and words, and under static matching mechanisms, models tend to memorize fixed character–word c...

Mingjian Li, Fei Gao, Nuo Qun et al. · 0 citations
Open access 2026

ShAnEL-2: A Multilingual Benchmarking Dataset for Short-Answer Language Learning Exercises

Before using GenAI models as EdTech tools, their pedagogical suitability should be corroborated. In this paper, we present ShAnEL-2 , a novel multilingual dataset comprising 1,185 student responses to short-answer language learning exercises corrected by teachers. We use ShAnEL-2 to establish an initial benchmark of (1...

Jasper Degraeuwe, Thomas Moerman · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.