Skip to content
Conference

Intent Classification Under Label Noise: A Comparative Analysis of Fine-Tuned Transformers and Large Language Models

Jul 2026 · Signal Processing and Communications Applications Conference · pp. 1-4 · 0 citations · 12 references

Abstract

Label noise, which is frequently encountered in real-world data, is a critical problem that can directly degrade model performance. In this study, we systematically investigated the impact of label noise on intent classification by comparing the in-context learning (ICL) approach of large language models (LLMs) with fine-tuned transformer models. In the experiments conducted on the Banking77 dataset, we evaluated four LLMs and two transformer models under three different noise types and three different noise levels. We also tested the LLMs with different prompting strategies to examine the effect of taking precautions against potential noise on different LLMs. Our findings show that strong LLMs experience less than 2 percent loss in accuracy and F1 score even under the heaviest noise conditions, whereas in fine-tuned transformer models and relatively weaker LLMs, the drop can reach the 15-20 percent range.

View source

Similar papers

Preprint Aug 2026

From Specialization to Generalization: Instruction-tuned LLMs for Robust Harmful Content Mitigation

By thoroughly unifying 36 English hate speech datasets spanning multiple labeling schemes, this work fine-tune a generalist LLM, based on Qwen3 (Qwen Team, 2025), specifically for hate speech mitigation, demonstrating not only state-of-the-art performance on in-domain benchmarks but also substantial improvements in cross-domain and cross-lingual generalization--areas where encoder-based specialist classifiers often struggle.

Lukas Edman, Daryna Dementieva, Alexander Fraser · 0 citations
Open access 2026

Advancing Machine-generated Text Detection: A Comprehensive Evaluation of Transformer-based Models

Test set results show that Decoding-Enhanced Bert with Disentangled Attention (DeBERTa) achieves the highest macro F1 − Score of 85.48%, surpassing the previously top-ranked Multi-Task Learning (MTL) system, which attains a macro F1 of 83.07%.

Batyr Sharimbayev, S. Kadyrov · 0 citations
Preprint Aug 2026

When Is Noise Response Universal? Tokenization as the Hidden Variable in Language Models

The degradation rate across neural models, both sentence embeddings and decoder-only LLMs, is studied, and how consistent it is depends on the scale of the noise: under word-level noise, models with very different architectures decline along nearly the same curve, while under character-level noise they separate.

Yefan Tao, Gerald Friedland, Luyang Kong · 0 citations
#artificial intelligence Preprint Sep 2026

Prompt-Robust Language Models: Which Training Strategies Work?

Despite their strong performance, large language models remain highly sensitive to prompt formulation. Prior work addresses this through refined data construction or through dedicated robustness objectives. We reproduce and compare these strategies under controlled conditions, and measure how effective they are in addressing models'prompt sensitivity. We find the current robustness fine-tuning methods improve over standard fine-tuning and in-context learning, but the best-to-worst prompt gap remains as high as 40-57% of performance. Moreover, the recent robustness-enhancing methods we test - CoIN for contrastive alignment and PPCL for consistency regularization - often fail to outperform the simplest data construction strategy: training on one template per batch. Our diagnostics explain these results. The auxiliary objectives move the quantity they penalize, but do not generalize beyond it. Additionally, data construction strategies differ due to the conflicting signs of per-template gradients on 57-64% of parameters. Thus, batches that mix formulations force the optimizer to reconcile competing updates instead of finding a shared, prompt-agnostic one.

F. Sadrieh, Michal Štefánik · 0 citations
Open access Jul 2026

Beyond The Surface: Characterizing Adversarial Boundaries in Synthetic Text Attribution Across Heterogeneous Domains

A hybrid detection framework which combines semantically deep embeddings from the RoBERTa transformer with a set of carefully designed language statistics and linguistic statistics and shows excellent resistance to the surface-level adversarial paraphrasing strategy.

Anita Rani, Suman · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.