Skip to content

Category

natural language processing

6,613 papers

#machine learning Preprint Open access Oct 2026

Studying the Soupability of Documents in State Space Models

We investigate whether hidden states from Structured State Space Models (SSMs) can be merged post hoc to support downstream reasoning. Inspired by model souping, we study document souping, a strategy where documents are encoded independently, and their representations are pooled, via simple operations like averaging, i...

Yasaman Jafari, Zixian Wang, Leon Bergen et al. · 0 citations
#machine learning Preprint Open access Oct 2026

Two Calls, Two Moments, and the Vote-Accuracy Curve of Repeated LLM Inference

Repeated sampling can improve LLM accuracy, and quantifying the gains from additional calls is essential for allocating test-time compute. We study binary decisions, a fundamental setting where repeated answers to the same question are aggregated by majority vote. We show that two independently sampled responses per ex...

Yi Liu · 0 citations
#machine learning Preprint Open access Oct 2026

Aligning the Query Space: Greedy Information Projection for Language Model Data Selection

Data selection for language models is often framed as balancing example quality and diversity. We argue that both are consequences of a more fundamental principle: selected examples should preserve the downstream query space induced by instructions, task signals, or retrieval needs. We present Greedy Information Projec...

Victor Ye Dong, Kuan-Yun Lee, Jiamei Shuai et al. · 0 citations
#machine learning Preprint Open access Oct 2026

Systematic Hazard Sampling: Minimal-Variance Inference for Discrete Diffusion and Flow Models

Uniform-noise discrete diffusion and flow models generate sequences non-autoregressively through iterative, context-dependent token replacements. However, these models are typically formulated as time-inhomogeneous continuous- or discrete-time Markov chains (CTMC/DTMC), sampled using independent Bernoulli change decisi...

Seunghwan Jang, Wonje Jeung, SooJean Han · 0 citations
#machine learning Preprint Open access Oct 2026

EfficientXpert: Efficient Domain Adaptation for Large Language Models via Propagation-Aware Pruning

Deploying domain-specialized large language models on resource-constrained hardware motivates reducing both adaptation cost and the number of retained weights. LoRA lowers adaptation cost but leaves a dense backbone, while separate pruning can discard weights that become important during adaptation. We propose Efficien...

Songlin Zhao, Michael Pitts, Zhuwei Qin · 0 citations
#machine learning Preprint Open access Oct 2026

ufakzeka-karar: An Open Turkish Typed-Decision Model with Order-Invariant Option Scoring

ufakzeka-karar is an open Turkish decision model with 182,494,466 parameters. Given a Turkish text and questions of a fixed answer type (a choice, a level on an ordered scale, or yes or no), it returns a temperature-scaled probability for every option and an expected error that serves as a "not sure" signal, without ge...

Sait Furkan Teke (ufak AI) · 0 citations
#machine learning Preprint Open access Oct 2026

Reading the Mood: Emotion-Guided Book-to-Music Recommendation via CGANs and LLMs

Background music that matches the mood of a text has been shown to make readers feel more immersed and improve their reading experience, motivating recommender systems that pair books with mood-matched music. In this direction, we present Sentiment Aware Generative Adversarial Network for Cross Domain Recommendation (S...

Manousos Linardakis, Georgios Alexandridis · 0 citations
#machine learning Preprint Oct 2026

Representation-Space MMD for Diffusion Language Models

We introduce a post-training method for diffusion language models (DLMs) that minimizes Maximum Mean Discrepancy (MMD) between generated and reference distributions in the feature space of a frozen pretrained DLM. To estimate MMD, we retain contextual features at individual token positions, obtaining multiple observati...

Ilya Drobyshevskiy, I. Sudakov, Maksim Semenov et al. · 0 citations
#machine learning Preprint Oct 2026

SOL: Measuring Gaps between Text Distributions by Double Sliced Wasserstein Metrics

Evaluating text generation requires measuring how well the generated distribution matches the data distribution. For autoregressive models, this is done by the perplexity. Diffusion and flow-based language models can only provide a likelihood bound, whose tightness differs between model families. Sample-based substitut...

Gregor Kornhardt, Moritz Piening, J. Chemseddine et al. · 0 citations
#machine learning Preprint Oct 2026

Ontology Concept Overlap as a Training Signal: Knowledge-Grounded Reinforcement Learning for Clinical Question Answering

Reinforcement learning post-training for language models relies on two reward designs: human preferences (RLHF, DPO) and binary verifiers (RLVR). Clinical question answering fits neither. Near-correct answers differ by a single substituted entity, and no executable check decides clinical correctness. We instantiate a s...

Aditya Tanna, Abhishek Jindal · 0 citations
#machine learning Preprint Open access Oct 2026

From Abusive Language Classification to Sequence Labeling Identification

Industrial content moderation must process massive message streams under tight latency constraints, yet most abusive language (AL) detection systems rely on sentence-level classification (ALC), which neither localizes abusive spans nor identifies who is targeted. We define Abusive Language Identification (ALI) as a seq...

Nicolas Zampieri, Ignacio Lopez, Manon Girard et al. · 0 citations
#machine learning Preprint Open access Oct 2026

What Does It Cost to Simulate a Quantum Sentence Classifier? An Energy and Compute Perspective on Near-Term QNLP

Near-term quantum natural language processing (QNLP) experiments often run on classical simulators, so simulator cost is part of the field's practical compute burden, yet accuracy tables do not show it. We measure that cost for a variational quantum classifier (VQC) on binary SST-2 sentiment classification, using Penny...

Kishlay Kashyap, Sandipan Ganguly · 0 citations

From tech blogs

See all →
MIT News · Artificial Intelligence Sep 24, 2026

Estimating suicide risk from text

A new language-processing tool could help identify the highest-risk individuals from natural language, enabling swifter interventions.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.