Skip to content

Category

natural language processing

6,613 papers

#machine learning Preprint Open access Oct 2026

StagQ: Constraint-Driven Multi-Precision Weight Quantization for LLMs

Serving a large language model (LLM) across a fleet of deployments requires several weight-precision operating points. Multi-precision formats serve them all from one stream whose prefixes are valid lower-precision codes, instead of storing multiple copies. We present StagQ, a multi-precision weight format whose main s...

Zhe Wei, Mengqi Guo, Yuan Yuan et al. · 0 citations
#machine learning Preprint Open access Oct 2026

AdaSpark: Adaptive DSpark with Online Learning for Tree Verification and N-gram Fill

Block drafters such as DSpark propose ranked candidates for several positions in one forward pass, and a tree verifier checks them in one pass of the target. The number of rows to verify trades the tokens a wider tree is expected to accept against the time a wider verify takes. Most schedulers that choose this number t...

Liquan Liu, Yifan Zhang, Bowei Xu · 0 citations
#machine learning Preprint Open access Oct 2026

Voltic: Distinguishing Volatility from Stochasticity in Recurrent Memory

Recurrent sequence models must decide how strongly to overwrite their memory at each token. Read as Bayesian filtering, this write is the gain of a Kalman update, set by uncertainty from two sources that pull it in opposite directions: volatility, how quickly the underlying associations change, and stochasticity, how n...

Parsa Hejabi, Morteza Dehghani, Payam Piray · 0 citations
#machine learning Preprint Oct 2026

ColdDDI: Evaluating Knowledge Utilization in Cold-Start Drug-Drug Interaction Prediction

Cold-start drug-drug interaction (DDI) prediction tests whether models can identify clinically significant interactions for drugs without training-time interaction history. Existing benchmarks mostly report aggregate edge-prediction scores, leaving a key evaluation question unanswered: when models receive molecular, te...

Ji-Heng Liang, Chen Zhao, Di Wu et al. · 0 citations
#machine learning Preprint Oct 2026

Universal Test-Time Training

Recent Test-Time Training (TTT) architectures compress context into fast weights that are updated online and queried as memory. Existing TTT designs keep this memory private to each layer: it recurs only over time, and depth merely indexes L separate memories. We argue that memory ownership need not be tied to depth, a...

Ze-Fan Cai, Qin-Zhe Hu, Ziqiao Ma et al. · 1 citation
#machine learning Preprint Oct 2026

Task Vector Descent: Learning from Non-IID Batches

A central challenge in continual learning is to acquire new knowledge without forgetting what the model has already learned. This challenge appears in language model training when training data comes from various domain-, user-, or task-specific distributions that are encountered unevenly over time. In such settings, s...

Anton Baumann, Jonas Hübotter, Zeynep Akata et al. · 0 citations
#machine learning Preprint Oct 2026

When Does Longer Reasoning Help? Predicting Mathematical Reasoning Through Discovery and Execution

Test-time compute can improve mathematical reasoning, but can short-budget runs predict how mathematical reasoning scales with additional compute? We introduce a Discovery--Execution (DE) framework that predicts the aggregate held-out scaling curves through a convolution of strategy discovery and conditional execution....

Adib Hasan, Lay Jain, Thanic Nur Samin · 0 citations
#machine learning Preprint Oct 2026

One Token Can Be Enough: Bridging Prompting and Activation Steering with Prefix Steering

Prompting guides language model behavior through the initial context, whereas activation steering often intervenes throughout generation. A natural question is whether steering can produce effects on subsequent computation similar to those of prompting. Under fixed-state attention assumptions, we establish sufficient c...

Xudong Zhu, Zhi-Hui Zhu · 0 citations
#machine learning Preprint Open access Oct 2026

Residual Visual Credit Optimization: Conserved Evidence Routing for Multimodal Reinforcement Learning

Reinforcement learning with verifiable rewards scales multimodal reasoning, but an outcome reward says how much a trajectory is worth, not how that value should be spread over the decisions that produced it. We introduce Residual Visual Credit Optimization (RVCO), which treats token credit as a conserved routing proble...

Lin Qiu, Yao Liu, Diyi Hu et al. · 0 citations
#machine learning Preprint Open access Oct 2026

CURIO: Curiosity-Driven Test-Time Learning for Open-Ended Discovery

Open-ended discovery requires learning from repeated attempts while continuing to explore directions whose value is not yet apparent. Search with a frozen large language model (LLM) can reuse previous solutions in context, but cannot update the model from its successes and failures on the test problem. Reinforcement le...

Tao Feng, Fangxu Yu, Zijie Lei et al. · 0 citations
#machine learning Preprint Oct 2026

Hierarchical Credit Assignment for RLVR on Fused Gromov-Wasserstein Geometry

Reinforcement learning with verifiable rewards (RLVR) has been shown to improve the reasoning capability of large language models (LLMs) across diverse reasoning tasks. However, group-based RLVR methods, such as GRPO, assign a uniform advantage to all tokens within rollouts of the same outcome. While existing works ref...

Qi Yu, Rui-Zhong Qiu, Zhichen Zeng et al. · 0 citations
#machine learning Preprint Oct 2026

Principled Top-$k$ Selection for Language Models with Hybrid Gradients

Selecting the best $k$ items out of $m$ candidates is a critical component of modern large language model systems, such as document selection in Retrieval-Augmented Generation (RAG) and expert routing in Mixture-of-Experts (MoEs). However, training these selection modules remains challenging due to weak gradient signal...

Xu-Chen Gong, Ju Sun, Tian Li · 0 citations

From tech blogs

See all →
MIT News · Artificial Intelligence Sep 24, 2026

Estimating suicide risk from text

A new language-processing tool could help identify the highest-risk individuals from natural language, enabling swifter interventions.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.