Skip to content

Category

natural language processing

6,429 papers

#artificial intelligence Preprint Oct 2026

Which Language Should a Skeleton Speak? Language Choices in Multilingual Reasoning

Skeleton-based reasoning prompting is a promising training-free approach for structuring LLM reasoning, but prior work largely assumes an English-centric setting. We propose the Language-Aware Skeleton Exploration Framework (LASEF) to study skeleton-language choice in multilingual mathematical reasoning. Across math be...

HyeonSeok Lim, Seung-Woo Song, Inho Won et al. · 0 citations
#artificial intelligence Preprint Oct 2026

Collaborative Reasoning Distillation via Cross-Feedback and Coherent Curation

Reasoning capabilities are critical for advancing Large Language Models, yet current approaches either require massive computational budgets or struggle to effectively distill reasoning to smaller models. Standard distillation methods rely on outcome-based rewards, failing to distinguish between sound reasoning and luc...

Tae-hong Kim, Seunggeun Cho, Dong-Su Han · 0 citations
#artificial intelligence Preprint Open access Oct 2026

How Do LLMs Change Predictions Under Negation?

Negation is an essential feature of human language, yet large language models (LLMs) remain unreliable in processing it. We evaluate recent open-source and closed-source LLMs on our negation benchmark and find that, in 37-71% of cases, they repeat the same answer under negation (e.g., "Madrid" for "What is not the capi...

Jongwook Yoon, Jongwon Lim, Sungjib Lim et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Arctic Questions, Missing Answers: A Dataset and Benchmark for LLM Abstention in Arctic Science

Large language models (LLMs) should abstain from scientific multiple-choice questions when no option is valid, but frequent abstention alone does not demonstrate sensitivity to answer availability. We introduce ArcticQA, a dataset of 194 questions derived from primary Arctic research, with automated checks of answer su...

Benjamin Wilcox, Dawei Gao, Pradeeban Kathiravelu et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

The Confidence Game: Strategic Miscalibration in Human-AI Delegation

Calibrated uncertainty quantification is essential to ensuring AI agents are trustworthy and reliable. However, when agents seek to maximize user engagement or revenue, confidence reports may be strategically distorted, detracting from their informativeness. We formalize this problem in the Confidence Game: a repeated...

Raghu Arghal, Saswati Sarkar, Shirin Saeedi Bidokhti · 0 citations
#artificial intelligence Preprint Open access Oct 2026

OnlineQAT: On-Policy Distillation for Ultra-Low-Bit Large Language Models

Quantization-aware training (QAT) can recover much of the accuracy lost when large language models are compressed below four bits. Existing re- covery stages, however, are commonly optimized on fixed completions or teacher-generated answers, whereas the deployed quantized model condi- tions on prefixes generated by its...

Wenjun Wang, Heng Li, Yanggan Gu et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Adversarial Images Hijack Web Agents from Visual Grounding to Browser Execution

Modern web agents built on large vision-language models process webpages, select relevant UI elements, and translate model outputs into browser actions. Existing visual red-teaming approaches use adversarial visual content to manipulate this process. However, they primarily target model inference and do not explicitly...

Wanjing Han, Levi Taiji Li, Mu Zhang et al. · 0 citations
#artificial intelligence Preprint Oct 2026

Multi-Objective Aligned Small Language Model Framework for SUD Patient Dialogue Generation

Substance Use Disorder (SUD) counseling requires patient responses that reflect underlying cognitive states such as beliefs, coping strategies, and readiness for change. Although large language models (LLMs) can generate fluent text, they often fail to produce cognitively coherent and clinically realistic patient behav...

Thushara Manjari Naduvilakandy, Hyeju Jang, M. Al Hasan · 0 citations
#artificial intelligence Preprint Open access Oct 2026

LayerRoPE: Dynamic Depth-wise Magnitude & Angular Superposition

As data propagates through a Transformer, the norm of its hidden states grows by orders of magnitude with depth, a phenomenon framed as 'curse of depth' and nearly universally treated as a pathology to be suppressed. We take the opposite view. Across 16 pre-trained LLMs from 9 families, spanning dense, mixture-of-exper...

Shikhar Srivastava, Christopher Kanan · 0 citations
#artificial intelligence Preprint Open access Oct 2026

ToolRACER: A Robust Agentic Conversation Emulation Resource for Agent Training and Evaluation

Task-oriented conversational agents remain fragile under real world conversation scenarios as they rarely follow a predictable script, especially when users exhibit non-cooperative behavior. Existing function-calling benchmarks often emphasize successful, cooperative interactions and underrepresent adversarial conversa...

Arkajyoti Chakraborty, Aryan Tayal, Ishika Agarwal et al. · 0 citations
#artificial intelligence Preprint Oct 2026

sk-bench: A Native-First Benchmark for Evaluating Large Language Models in Slovak

Multilingual LLM benchmarks omit Slovak, a morphologically rich West Slavic language of five million speakers, or cover it only by machine translation. We present sk-bench, a native-first Slovak benchmark with 30 datasets (33 scored task variants) across ten skill categories. Eleven resources are introduced or first pa...

Marek Suppa, Ivan Vykopal, Andrej Ridzik et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

U-Space: Uncovering When and Why Uncertainty Arises in Language Models

Large language models are informing decisions with ever-higher stakes. As the consequences of their errors grow, a central question becomes harder to ignore: how much can we trust an individual answer? Yet recognizing when to defer remains difficult because language models can present incorrect conclusions with fluent...

Tobias Braun, Nils Loose, Alexander Herzog et al. · 0 citations

From tech blogs

See all →
MIT News · Artificial Intelligence Sep 24, 2026

Estimating suicide risk from text

A new language-processing tool could help identify the highest-risk individuals from natural language, enabling swifter interventions.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.