Skip to content

Category

natural language processing

6,613 papers

#machine learning Preprint Open access Oct 2026

Lost in the bf16 Cast: Exporting Ternary Language Models Can Revert Most Low-Learning-Rate Code Changes

Ternary language models such as BitNet b1.58, Falcon-E and BitCPM are fine-tuned with higher-precision latent weights and deployed as ternary codes produced by an export step that, in the labs' documented pipelines, first casts the latents to bf16. We audit those pipelines across three labs. In released checkpoints, fp...

Avichal Sahai (Ofbusiness), Nishant Raj (Ofbusiness), Animesh Srivastava (Ofbusiness) · 0 citations
#machine learning Preprint Oct 2026

$\alpha$Transfer: Coefficient Transfer for Efficient Model Merging

Model merging offers a promising solution for combining multiple fine-tuned checkpoints into a single model through parameter arithmetic. However, finding optimal merging coefficients requires an extensive search that becomes prohibitively expensive as models scale in both size and number, due to high memory requiremen...

Shih-Cheng Huang, Zhi Rui Tam, Chieh-Yen Lin et al. · 0 citations
#machine learning Preprint Open access Oct 2026

TRACE: Rollout-Guided Quantization-Aware Training for FP4 Reinforcement Learning of MoE Language Models

Reinforcement learning (RL) for post-training large language models (LLMs) incurs substantial computation and memory overhead during rollout generation, which motivates low-precision rollout for efficient RL training. However, existing FP4 RL methods suffer from a key limitation: they primarily optimize quantization ac...

Xin Wang, Hao Yu, Zhengyang Zhuge et al. · 0 citations
#machine learning Preprint Open access Oct 2026

Improving Synthetic Data Generation for Argument Mining via Adversarial Reinforcement Learning

Argument Mining (AM) is fundamentally constrained by the scarcity of high-quality structure-annotated datasets. While LLMs have shown promise in synthetic data generation, producing synthetic AM data that is both structurally accurate and sufficiently diverse remains a challenging problem. To address this problem, we r...

Zhijun Zhang, Qianlong Wang, Keyang Ding et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Stateless Language Agents: Scaling Long-Horizon Automated Research

Automated research systems increasingly run LLM agents over long horizons, but more inference does not by itself produce more progress: agents replay growing histories, duplicate one another's work, or stop experimenting while token consumption continues. Yet most evaluations use short budgets or benchmarks that satura...

Qizheng Zhang, Changxiu Ji, Isaac Sun et al. · 0 citations
#artificial intelligence Preprint Oct 2026

AlignQuant: Tile-Aligned Mixed-Precision Quantization for Efficient LLM Generation

Fine-grained mixed-precision quantization promises efficient large language model inference, but local precision choices can conflict with regular GPU storage and computation units. This precision-boundary mismatch limits the translation of compression into practical acceleration. We introduce AlignQuant, a post-traini...

Han-Zhi Zhang, Qiao Zhang, Qing-Lei Cao et al. · 0 citations
#artificial intelligence Preprint Oct 2026

Stepped MoE: Segment-Level Routing with Configurable Inference Complexity

Training large language models (LLMs) is resource-intensive, and adapting them for diverse deployment scenarios with varying computational constraints remains challenging. While elastic architectures enable flexible model deployment and sparsely activated models allow input-adaptive computation, existing approaches tre...

Arnav Kundu, Zhao-Yang Xu, Bai-Ru Hou et al. · 0 citations
#machine learning Preprint Open access Oct 2026

Structuring MoE Expert Selection for Agentic Reinforcement Learning

Long-horizon LLM agents are frequently implemented using sparse mixture-of-experts (MoE) models, yet the co-design of agentic behavior and MoE structures remains underexplored. In this work, we comprehensively study the connections between agentic post-training and MoE expert selection. In off-the-shelf MoE models, we...

Bolian Li, Ting-Yao Hu, Cheng-Yu Hsieh et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Minimal Witness Reinforcement Learning

``What are the irreducible conditions that are sufficient to produce an outcome?'' is one of the most common questions that recur across computation and science. Its answers, the minimal sufficient witnesses, are what we mean by explanations, mechanisms and reasons. These problems usually ask for multiple minimal witne...

T. Y. Tsui, Zihao Ye, Pengxiang Cai et al. · 0 citations
#machine learning Preprint Open access Oct 2026

Learning Scientific Exploration from Human Research Decision Trajectories

A key challenge in building AI systems for scientific research is enabling $\textit{scientific exploration}$: the systematic process of investigating unknown phenomena or ideas to gain new knowledge through sequences of research decisions and actions. Yet this process is largely missing from existing scientific corpora...

Xuchen Gong, Shane Gu, Haokun Liu et al. · 0 citations
#artificial intelligence Preprint Oct 2026

CLM-as-a-Judge: Evaluating an Open Contrastive Decision Model on Public Judge Benchmarks

An open contrastive decision model is near chance as a judge on the hard public benchmarks: Contrastive-LM/CLM-v0.1-8B scores between 0.351 (best- of-four, chance 0.250) and 0.593 (pairwise, chance 0.500), is statistically indistinguishable from coin flipping on RM-Bench and JudgeBench, and answers every HaluEval item...

Gowthamkumar Nandakishore · 0 citations
#machine learning Preprint Open access Oct 2026

A theory of platonic representations in language models

Representations of translated sentences are similar in the inner layers of multilingual language models -- an observation connected to the platonic representation hypothesis, yet unexplained theoretically. We provide an explanation based on the assumption that data have a hidden hierarchical structure whose abstract le...

Darshil Doshi, Wenjie Zhou, Corinna Elena Wegner et al. · 0 citations

From tech blogs

See all →
MIT News · Artificial Intelligence Sep 24, 2026

Estimating suicide risk from text

A new language-processing tool could help identify the highest-risk individuals from natural language, enabling swifter interventions.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.