Enhancing Consistency in Academic English Writing Feedback Generation with a MacBERT-large Model Combining Adversarial Training and Contrastive Learning
Aug 2026· Advanced Electromagnetics· Vol 15, pp. 8858-8862· 0 citations· 10 references
TL;DR
An enhanced MacBERT-large encoder–decoder model integrating Fast Gradient Method adversarial training and supervised contrastive learning is proposed, providing a semantic consistency modeling framework for intelligent text generation and academic writing assistance systems.
Abstract
Academic English writing feedback generation requires robust semantic understanding, stable feedback output, and accurate discrimination among similar error types. Existing feedback generation systems often produce inconsistent suggestions for semantically equivalent inputs and show limited generalization to complex academic expressions. To improve feedback consistency, this study proposes an enhanced MacBERT-large encoder–decoder model integrating Fast Gradient Method adversarial training and supervised contrastive learning. The MacBERT-large encoder extracts contextual semantic representations of academic text, while a Transformer decoder generates feedback sequences using a dedicated academic vocabulary. FGM adversarial training introduces controlled perturbations into the embedding layer, enabling the model to maintain stable predictions under paraphrased or slightly varied inputs. A supervised contrastive learning module maps text samples into a representation space where feedback cases with the same error type are pulled closer and different error types are separated through NT-Xent loss. A multi-task learning framework jointly optimizes cross-entropy loss, adversarial loss, and contrastive loss to balance generation quality, robustness, and category discrimination. Experiments on the AEW-Feedback dataset containing 15,000 academic papers show that the proposed model achieves 67.3% BLEU-4, 71.2% ROUGE-L, 74.8% METEOR, and a feedback consistency score of 0.891, outperforming MacBERT-large and single-enhancement variants. The method provides a semantic consistency modeling framework for intelligent text generation and academic writing assistance systems.
An Efficient Dual-BERT Adversarial Network (DBAN) is proposed to improve the translation of noisy UGC by integrating contextual representation learning with adversarial training and significantly improves contextual understanding and cross-lingual semantic alignment while maintaining computational efficiency.
A. A. Aliero, Nasiru Muhammad Dankolo· International Journal Of Eng...· 0 citations
As machine translation (MT) systems continue to improve, standard benchmarks become less informative for exposing remaining weaknesses. Traditional methods for creating challenging test sets rely on expensive manual creation or curation, while automated approaches struggle to produce sets with the necessary translation difficulty and linguistic diversity. We propose a scalable reinforcement-learning-based approach for rewriting existing source texts into instances that are more difficult to translate for MT systems. We fine-tune a large language model with Group Relative Policy Optimization (GRPO), using reward signals based on translation difficulty together with constraints for semantic similarity, grammaticality, and approximate length preservation. On WMT25, our approach substantially reduces average COMET translation quality from 0.63 to 0.48, while preserving grammaticality and readability, whereas the base model remains at 0.64. Evaluations on the unseen WMT19-WMT24 benchmarks confirm that this behavior generalizes beyond the training data, and human evaluation further shows that the rewrites substantially lower translation quality while incurring a moderate drop in naturalness and only a small change in grammaticality. We release our code to support reproducibility.
Florian Zogaj, Jakob H\"utteneder, Giovanni De Muri et al.· 0 citations
This work proposes augmenting existing benchmarks to increase translation difficulty by combining adversarial optimization with a differentiable translation difficulty estimator, and uses gradients from a combined difficulty and fluency objective to iteratively replace tokens in Adversarial Translation Optimization (ATO).
William Kalikman, Šimon Sukup, Michal Tesnar et al.· European Association for Mac...· 2 citations
Assessing the quality of scientific literature translation remains challenging because of strong subjectivity, dense domain-specific terminology, and the limited availability of standardized reference translations. These issues are particularly relevant for the international dissemination of research in advanced electromagnetic engineering, where precise multilingual communication supports the reliable exchange of knowledge on electromagnetic waves, antennas, and propagation technologies. This paper proposes a Contrastive Learning-based Chinese-English Scientific Translation Quality Evaluation model (C-TQE). By constructing multi-level positive and negative sample pairs, the model learns the relative ordinal relationships of translation quality within a shared representation space. A dual-encoder architecture encodes source sentences and candidate translations through a shared pre-trained language model, while a contrastive loss function draws high-quality translations closer to the source representation and separates low-quality ones. To address the characteristics of scientific texts, a term-aware negative sampling strategy exploits domain dictionaries and syntactic structures to generate semantically similar but terminologically incorrect examples. Experiments on 11, 238 human-annotated instances from the WMT20–22 Chinese-English scientific translation tasks show that C-TQE achieves a Kendall’s tau correlation coefficient of 0.564 with human judgments, outperforming COMET (0.512) and BLEURT (0.497). Ablation studies confirm the effectiveness of term-aware negative sampling and the contrastive learning objective, while diagnostic analysis demonstrates high consistency in evaluating terminological accuracy and syntactic structures. The proposed framework provides an effective solution for large-scale scientific translation quality assessment and facilitates the accurate international communication of multidisciplinary engineering research, including electromagnetic and antenna-related studies.
Large language models continue to face challenges in translating low-resource languages with scarce parallel data. This study investigates how to fine-tune them effectively using target-side monolingual data. Existing approaches—dominated by back-translation and recent LLM-based rewriting—remain limited by noisy synthetic sources, unguided simplification, and the absence of a principled mechanism for integrating monolingual sentences into the training objective. To address this, we developed a semi-supervised framework that integrates marginal distribution estimation and curriculum-guided rewriting to exploit monolingual data for low-resource translation. Experiments in four low-resource directions demonstrated substantial gains, averaging +8 spBLEU and +10 COMET over strong baselines, while three additional mid-resource directions showed stable improvements and consistent trends. Reference-free metrics further validated robust gains in fluency and adequacy. The findings establish a scalable paradigm for low-resource translation, revealing that the principled integration of marginal likelihood estimation and generative rewriting enables large language models to achieve superior performance under extreme data scarcity.
Wenjie Yu, Zhiqiang Yu, Zuo Jiang et al.· ACM Transactions on Asian an...· 0 citations
This paper proposes a general optimization framework that combines a vocabulary pruning method with a targeted fine-tuning protocol for MNMT models, and reduces the vocabulary size from over 128,000 to approximately 10,000 tokens, enabling a 60% memory saving without any loss in performance.
Ahmed Amine Aliane, N. Semmar, H. Aliane· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.