Skip to content

Category

natural language processing

6,613 papers

#artificial intelligence Preprint Oct 2026

Playing social deduction games with reinforcement fine-tuned large language models

Reinforcement fine-tuning (RFT) is increasingly used in applications where large language models (LLMs) interact with humans and other agents. Here we use social deduction games to study how RFT changes LLMs'social behaviour. We let fine-tuned and base LLM agents play hidden-role games that require hidden-state inferen...

Ling-Zhe Zhang, Yun-Peng Zhai, Tong Jia et al. · 0 citations
#artificial intelligence Preprint Oct 2026

Clean: Second-order LLM Training at Linear Memory Cost via Nystr\"om Sketching

Training large language models (LLMs) entails a fundamental trade-off: memory-efficient optimizers such as Adam discard cross-parameter curvature, whereas full-curvature methods such as SOAP can accelerate convergence at prohibitive memory costs. We introduce Clean, a memory-efficient and full-curvature optimizer desig...

Beheshteh T. Rakhshan, S. Rajabi, Maziar Sargordi Shikai Fang et al. · 0 citations
#artificial intelligence Preprint Oct 2026

Trajectory-Derived Confidence for Reliable, Resource-Aware Clinical Text-to-SQL Agents

LLM agents for clinical text-to-SQL applications reason autonomously over multiple steps but cannot assess whether their own reasoning or outputs can be trusted. In high leverage applications such as healthcare, this presents a critical risk where system mistakes can be costly. These reliability failures are also resou...

Mincheol Daniel Song, Joshua Ward, Jake Jung et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

PB-GRPO: Learning Socially Adaptive LLM Agents from Persona-Driven Simulation with Preference-Batched GRPO

Building LLMs that behave well socially, not merely correctly, requires Building LLMs that behave well socially, not merely correctly, requires more than producing locally helpful responses. A socially competent agent must infer users' unstated goals, respect their preferences, and adapt as the conversation unfolds. Th...

Jingquan Wang, Jun Yin, Xu Han et al. · 0 citations
#artificial intelligence Review Oct 2026

Representational Control over Self-Report&Behavior Coherence in LLM Risk-Taking

Self-report is an appealing low-cost probe of an LLM's dispositions, but recent work finds only selective agreement between what models report and how they behave. Prior accounts establish these patterns by prompting black-box LLMs, leaving open whether the gap is a prompting artefact or a fact about how the underlying...

R. Kocielnik, Pei-Yang Song, Peng-Rui Han et al. · 0 citations
#artificial intelligence Preprint Oct 2026

SEER: Self-Evolving Event Reasoning and Retrieval for Time Series Forecasting

Real-world time series are frequently driven by exogenous events and structural shifts, rendering conventional forecasting based solely on historical numerical observations insufficient. While language models can retrieve external news, standard retrieval-augmented approaches struggle with high noise, missing signals,...

Ming-Tian Tan, Palash Goyal, Mihir Parmar et al. · 0 citations
#artificial intelligence Preprint Oct 2026

Slaying the Hydra: Interaction-Aware Circuit Discovery in Language Models

Localizing behavior to individual components of a language model is a central goal of mechanistic interpretability. However, scoring components one at a time misses context-dependent effects: a primary component can inhibit the activation of a backup, leading to issues with ranking components. Actual causality studies...

Sankaran Vaidyanathan, R. Urbaniak, Emily Bunnapradist et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

A Step Towards Forgetting: Optimiser History and the Loss of Answer Mass

During fine-tuning, a language model can assign less probability to previously learned answers even when the current gradient acts to preserve that probability. With momentum, each update also carries gradients computed at earlier model states, and these stored contributions can push the model in the opposite direction...

Valeria Ruscio, Seth Nabarro, Keiran Thompson · 0 citations
#artificial intelligence Preprint Oct 2026

SYNLAT: Syntax-Aligned Text-Latent Compression for Chain-of-Thought Reasoning

Long chain-of-thought (CoT) traces impose substantial output-token costs. Under constrained budgets, compression must preserve answer-critical information, making boundary placement central. Token-level and fixed-length boundaries can fragment coherent spans such as phrases, formulas, and local derivations, whereas ste...

Yi-Feng Zhao, Hong-Jun Yu, Shi-Bo Wang et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

The Score Is Not the Structure: Brain Alignment and Cross-Lingual Transfer

Researchers often support the claim that a model shares structure with the brain, or across languages, by reporting a similarity score. We ask what such a score reads when the shared structure is absent, or when the tool that measures it does not work. We check two settings, and in both the score is not what it appears...

Saman Rahbar · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Do Motion Tokenizers for Co-Speech Gesture Generation Encode Gesture Semantics?

Discrete motion tokenizers encode motion as atomic units and are widely used for co-speech gesture generation. It remains unclear which motion properties, especially those relevant to gesture semantics, are recoverable from these codebooks. We probe a reconstruction-trained codebook using 19 co-speech gesture descripto...

Varsha Suresh, Divij Jain, Jia Liu et al. · 0 citations
#artificial intelligence Preprint Oct 2026

The Pushback Paradox: A Two-Probe Diagnostic for Language Model Compliance

Are language models compliant with user instructions? A model that always complies can be stopped but also exploited, while one that always resists can be neither exploited nor stopped. We contribute an open two-probe benchmark that can place any language model on this spectrum. In the active probe, a user instructs th...

Stefan Bühler, David Exler, M. Reischl et al. · 0 citations

From tech blogs

See all →
MIT News · Artificial Intelligence Sep 24, 2026

Estimating suicide risk from text

A new language-processing tool could help identify the highest-risk individuals from natural language, enabling swifter interventions.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.