Skip to content

Category

natural language processing

6,613 papers

#artificial intelligence Preprint Oct 2026

Mind the Accent Gap: British Accent Robustness in Speech-Driven Financial Voice Assistants

AI voice assistants often use Automatic Speech Recognition (ASR) with LLM-based reasoning, yet existing systems struggle with regional British accents, including Scottish, Irish, and Welsh accents, since most ASR models are trained predominantly on American English voice data. Consequently, errors can carry through to...

Aadam Haq, Oggi Rudovic, Malcolm Chadwick et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Anatomy of LLM Sycophancy: What a Flip Rate Hides

A model under pushback can correct itself, capitulate, or hold, and one flip rate counts a correction and a capitulation alike. Using SycoLens, a modular replay protocol, we test how user pressure and evaluation settings shape measured flip rates. Each measurement is one stateless replay of an item, a committed answer,...

Haonan Huang · 0 citations
#artificial intelligence Preprint Oct 2026

The Assistance Dilemma: Learning to Teach via Multi-Turn Reinforcement Learning

Large language models (LLMs) trained to answer questions are natively poor at teaching. Reinforcement Learning (RL) against a simulated student is a promising approach to improve their pedagogy, but existing RL-trained tutors reward the student's success on the tutored problem with the tutor's words still in context. T...

Jakub Macina, Manu Kapur, Mrinmaya Sachan · 0 citations
#artificial intelligence Preprint Oct 2026

Better Call Reward: Reward Hacking as Strategic Abstention in Legal Reasoning Models

What happens when a legal AI model learns to look like a lawyer instead of reasoning like one? We fine tune Qwen3-8B with Group Relative Policy Optimisation (GRPO) against a proxy built from three surface features: citation count, legalese density, and response length. The model does not learn to reason more effectivel...

Subramanyam Sahoo, Justin C. Shenk · 0 citations
#artificial intelligence Preprint Open access Oct 2026

SpatialChain: A Benchmark for Auditing Spatial Reasoning Faithfulness in VLMs

Thinking-enabled vision-language models (VLMs) report ever-higher accuracy on spatial benchmarks, yet final-answer scores cannot reveal whether a correct prediction reflects faithful spatial reasoning or a linguistic shortcut. We introduce SpatialChain, a dataset of 28,350 training and 899 test examples pairing spatial...

Rafael Teixeira Sousa, Vin\'icius Paulo Lopes de Oliveira, Elisa Ayumi Masasi de Oliveira et al. · 0 citations
#artificial intelligence Preprint Oct 2026

RAISED: Self-Distillation for Robustness to Prompt Injection in LLM Agents

Tool-using language-model agents are vulnerable to indirect prompt injection because they must act on untrusted external content. Existing training-time defenses can reduce attack success rates, but often at the cost of general capabilities. We show that training-based defenses induce substantial drift in the model's o...

Mohamed Dhouib, Clément Elliker, Alexi Canesse et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Steering by Influence: Curvature Aware Data Weighting for Activation Steering

Inference-time steering offers cheap, fine-grained control over a language model's outputs by estimating a concept's representation in activation space and shifting activations towards it. Existing methods build these representations from activation averages over contrastive datasets. These averages incorporate unrelat...

James A. E. Dixon, Stephen J. Roberts, Francesco Quinzan · 0 citations
#artificial intelligence Preprint Oct 2026

Agentic schema-guided extraction of materials process knowledge from scientific literature

Materials literature contains detailed experimental knowledge, but procedures, chemical entities and measurements remain difficult to aggregate because they are reported in heterogeneous forms and depend on process-specific context. We present SciKGExtract, a schema-guided framework that combines large-language-model e...

Sameer Sadruddin, Jennifer D'Souza · 0 citations
#artificial intelligence Review Oct 2026

DialectSentEval 2026: Arabic Dialect Sentiment Analysis and Swapping Shared Task

Sentiment analysis is a fundamental problem in Natural Language Processing (NLP). Standard sentiment classification for the Arabic language remains challenging due to the high volume of dialectal Arabic. To advance research in this area, this paper proposes the Shared Task on Sentiment Analysis and Swapping in Arabic D...

Saad Ezzini, Shadi Abudalfa, Maram Alharbi et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Cross-lingual Calibration of Pre-Generation Success Probes for Multilingual LLM Routing

Pre-generation success probes estimate response correctness from a language model's hidden activations before decoding, enabling cost-aware routing. While prior work has demonstrated their utility primarily on English inputs, we study their reliability across languages along three dimensions: (1) whether they preserve...

Andrea Paganelli, Stefano Civelli, Pietro Bernardelle et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Anosognosia in LLMs: Probing Self-Awareness of Quantized Computational Substrate

Can LLMs recognize degradation in their own computational substrate? Inspired by anosognosia, a neurological condition in which patients fail to recognize impairments in their own abilities, we investigate whether LLMs can recognize degradation in their computational substrate induced by quantization. We first show tha...

Yoshihiro Izawa, Gouki Minegishi, Yoko Yamakata · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Cross-Lingual Transferability of Training Data Extraction Attacks to Recover Memorized PII

The robustness of Personally Identifiable Information (PII) protection in Large Language Models (LLMs) is a critical concern, yet the risks associated with cross-lingual data extraction remain under-explored. This study evaluates the vulnerability of English-centric and multilingual models to Training Data Extraction (...

Alexandru Nazare, Agnese Profico, Nicol\`o Vania et al. · 0 citations

From tech blogs

See all →
MIT News · Artificial Intelligence Sep 24, 2026

Estimating suicide risk from text

A new language-processing tool could help identify the highest-risk individuals from natural language, enabling swifter interventions.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.