Skip to content

Category

natural language processing

6,613 papers

#natural language process... Preprint Open access Oct 2026

Quantifying the Gap between Understanding and Generation within Unified Multimodal Models

Recent advances in unified multimodal models (UMM) have demonstrated remarkable progress in both understanding and generation tasks. However, whether these two capabilities are genuinely aligned and integrated within a single model remains unclear. To investigate this question, we introduce GapEval, a bidirectional ben...

Chenlong Wang, Yuhang Chen, Zhihan Hu et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Cross-Lingual Activation Steering for Multilingual Language Models

Large language models exhibit strong multilingual capabilities, yet significant performance gaps persist between dominant and non-dominant languages. Prior work attributes this gap to imbalances between shared and language-specific neurons in multilingual representations. We propose Cross-Lingual Activation Steering (C...

Rhitabrat Pokharel, Ameeta Agrawal, Tanay Nagar · 0 citations
#natural language process... Preprint Open access Oct 2026

TabiBERT: A Large-Scale ModernBERT Foundation Model and A Unified Benchmark for Turkish

The introduction of BERT established encoder-only transformer models as a foundational paradigm in natural language processing. Encoder-only models remain the standard tool for classification, tagging and retrieval, where contextual representations and low inference cost matter more than text generation, yet Turkish la...

Melik\c{s}ah T\"urker, A. Ebrar K{\i}z{\i}lo\u{g}lu, Onur G\"ung\"or et al. · 0 citations
#natural language process... Preprint Open access Oct 2026

Artificial Hivemind: The Open-Ended Homogeneity of Language Models (and Beyond)

Language models (LMs) often struggle to generate diverse, human-like creative content, raising concerns about the long-term homogenization of human thought through repeated exposure to similar outputs. Yet scalable methods for evaluating LM output diversity remain limited, especially beyond narrow tasks such as random...

Liwei Jiang, Yuanjun Chai, Margaret Li et al. · 0 citations
#natural language process... Preprint Open access Oct 2026

VietBinoculars: A Zero-Shot Approach for Detecting Vietnamese LLM-Generated Text

The rapid proliferation of Large Language Models has intensified the challenge of distinguishing LLM-generated text from human writing in non-English languages. This study introduces VietBinoculars, a zero-shot detection framework coupling PhoGPT-4B observer and performer models with calibrated global decision threshol...

Trieu Hai Nguyen, Sivaswamy Akilesh · 0 citations
#artificial intelligence Preprint Open access Oct 2026

FedCoT: Communication-Efficient Federated Reasoning Enhancement for Large Language Models

Enhancing LLM reasoning in federated settings is nontrivial due to stringent computational, communication, and privacy constraints, especially in healthcare, where clinically consequential decisions require not only accuracy but also interpretable, auditable rationales to meet safety, accountability, and regulatory req...

Chuan Li, Qianyi Zhao, Fengran Mo et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Too Categorical to be Human: Emotion Concepts in LLMs and Humans

Understanding human emotions is central to user-facing AI applications, safety alignment, and the simulation of human behavior. As emotional stimuli shape high-stakes behavior in Large Language Models (LLMs), there is increasing interest in how models represent emotion concepts internally. Mechanistic accounts of these...

Sree Bhattacharyya, Evgenii Kuriabov, Lucas Craig et al. · 0 citations
#natural language process... Preprint Open access Oct 2026

Common Corpus: The Largest Collection of Ethical Data for LLM Pre-Training

Large Language Models (LLMs) are pre-trained on large amounts of data from different sources and domains. Such datasets often contain trillions of tokens, including large portions of copyrighted or proprietary content, which raises questions about the legal use of such models. This underscores the need for truly open p...

Pierre-Carl Langlais, Pavel Chizhov, Catherine Arnett et al. · 0 citations
#natural language process... Preprint Open access Oct 2026

Even Small Reasoners Should Quote Their Sources: Introducing the Pleias-RAG Model Family

We introduce a new generation of small reasoning models for RAG, search, and source summarization. Pleias-RAG-350m and Pleias-RAG-1B are mid-trained on a large synthetic dataset emulating the retrieval of a wide variety of multilingual open sources from the Common Corpus. They provide native support for citation and gr...

Pierre-Carl Langlais, Pavel Chizhov, Mattia Nee et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Sherpa: Teaching LLMs to Teach Adaptively

Large language models (LLMs) have become increasingly capable problem solvers, but being able to solve a problem is not the same as being able to teach it. Existing approaches to training LLMs as teachers rely on demonstrations, preference data, or predefined pedagogical criteria that specify what good teaching looks l...

Weixian Xu, Yanzhe Zhang, Zora Zhiruo Wang et al. · 0 citations
#natural language process... Preprint Open access Oct 2026

A Systematic Study of Semantic ID Spaces for Generative Information Retrieval

Generative Information Retrieval (GIR) has emerged as a transformative paradigm, shifting document retrieval from a traditional "retrieve-and-rank" workflow to sequence-to-sequence generation, where a model directly predicts document identifiers (DocIDs). While the semantic design of these DocIDs is known to be critica...

Alexia Allal, Hicham Randrianarivo, Sylvain Lamprier · 0 citations
#natural language process... Preprint Open access Oct 2026

Disentangling Paradigm, Identifier, and Decoding in Generative Retrieval

Generative retrieval trains a language model to generate the identifier of a relevant document. Recent work replaces the autoregressive decoder with diffusion, but changes identifiers, training recipe and decoding at once, so differences cannot be credited to the paradigm. On NQ320K and MS300K, we train autoregressive,...

Hicham Randrianarivo, Logan Renaud, Alexia Allal · 0 citations

From tech blogs

See all →
MIT News · Artificial Intelligence Sep 24, 2026

Estimating suicide risk from text

A new language-processing tool could help identify the highest-risk individuals from natural language, enabling swifter interventions.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.