ViePAWS: A Vietnamese Adversarial Dataset for Paraphrase Identification under LLM-based Word Scrambling
Paraphrase identification remains challenging when sentence pairs exhibit high lexical overlap but subtle semantic differences, as models often rely on surface similarity rather than true meaning. Existing benchmarks such as PAWS highlight this issue, but comparable resources for Vietnamese are still lacking. In this p...