Skip to content
Conference

ViePAWS: A Vietnamese Adversarial Dataset for Paraphrase Identification under LLM-based Word Scrambling

Aug 2026 · International Conference on Multimedia Analysis and Pattern Recognition · pp. 43-48 · 0 citations · 21 references

Abstract

Paraphrase identification remains challenging when sentence pairs exhibit high lexical overlap but subtle semantic differences, as models often rely on surface similarity rather than true meaning. Existing benchmarks such as PAWS highlight this issue, but comparable resources for Vietnamese are still lacking. In this paper, we introduce ViePAWS1, a Vietnamese benchmark for paraphrase identification under adversarial high-overlap conditions. The dataset is constructed to contain sentence pairs that are lexically and semantically similar while differing in meaning, making them difficult to distinguish using surface cues alone. We evaluate both pretrained models and large language models and observe consistent performance degradation compared to standard settings. In particular, models struggle to correctly classify non-paraphrase pairs with high similarity, revealing a strong bias toward lexical overlap. Our analysis further shows that different types of meaning variation lead to different failure patterns, indicating that current models lack robust mechanisms for capturing fine-grained semantic distinctions. These results highlight the need for more challenging evaluation benchmarks that better reflect true semantic understanding in paraphrase identification.

View source

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.