Skip to content

Category

natural language processing

6,613 papers

#machine learning Preprint Open access Oct 2026

ExpertMuon-Compass: Alignment-Guided Step Sizes for Mixture-of-Experts Training

Mixture-of-experts (MoE) language models send each token to a few experts, so each expert is trained on a different part of the data, and this part changes during training. With a shared learning rate, Muon applies updates of roughly the same size to expert matrices of the same shape, even when an expert's update is po...

Omatharv Bharat Vaidya, Ashwin Vinod, Pedram Akbarian et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

MemArena: An Ego-Centric Benchmark for On-Device Agentic Personal Memory Assistants at Scale

Edge-deployed personal memory assistants must handle private interpersonal conversations on-device with open-weight models. Yet, existing memory benchmarks often under-test the combination of activity-dense interaction, ego-centric perspective, and coherent multi-session worlds. MemArena fills these gaps with a single-...

Jiadong Zhang, Xiaosong Ma · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Cost-Effective Automated Judging of Natural-Language Mathematical Proofs

Grading natural-language mathematical proofs is a recurring cost in evaluating math-reasoning systems, and frontier LLM judges are expensive. We ask whether cheap open-weight models can serve as reliable judges given a candidate proof, a ground-truth proof, and a human-grading rubric. On a 200-instance validation sampl...

Benjamin Grayzel · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Safe Inference-Time Alignment via Lagrangian Reward Augmentation

Inference-time alignment steers a frozen language model during decoding using auxiliary reward signals, avoiding the cost of repeated weight updates. However, existing inference-time alignment methods typically optimize a single scalar score, so explicit safety constraints must either be ignored or encoded through manu...

Yaswanth Chittepu, Ativ Joshi, Sohini Chintala et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Office Comprehension Benchmark

We introduce Office Comprehension Bench (OCB), the first public benchmark to jointly evaluate LLM systems on Word, Excel, and PowerPoint comprehension over native file formats (.docx, .xlsx, .pptx) and their variants. OCB consists of two tracks. File Fidelity Q&A tests structural and visual perception of office artifac...

Firoz Shaik, Mateus Pican\c{c}o Lima Gomes, Tanvir Aumi et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Narrative-UFET: Narrative Generation for Ultra-Fine Entity Typing

Ultra-fine entity typing (UFET) assigns highly specific types to entity mentions, but current approaches struggle with types in the long tail. We hypothesize that a key limitation is the reliance on sentence-level context, since disambiguating evidence is often spread across multiple sentences. Testing this has been di...

Mreedul Gupta, Advait Deshmukh, Ashwin Umadi et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Gefen: Optimized Stochastic Optimizer

AdamW is a default optimizer for deep learning, but its moment states add two parameter-sized buffers to training memory, increasing the cost of large-scale pretraining. We propose Gefen, a memory-efficient optimizer that automatically shares second-moment estimates across parameter blocks and quantizes the first momen...

Nadav Benedek, Tomer Koren, Ohad Fried · 0 citations
#artificial intelligence Preprint Open access Oct 2026

LEAF: Growing Trees Without Branching for Speech-Aware Large Language Model Post-Training

State-of-the-art GRPO-style methods for speech-aware large language model post-training suffer from coarse credit assignment, broadcasting the same terminal-reward advantage to every token in a response. This ignores useful structure within rollout batches, where speech-conditioned completions often share prefixes befo...

Argyrios Gerogiannis, Yekaterina Yegorova, Mark Hasegawa-Johnson et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Low-Resource Safety Failures Are Action Failures, Not Representation Failures

Language models often answer harmful requests in low-resource languages (LRLs) that they refuse in high-resource languages (HRLs). Across three instruction-tuned models and 23 languages, harmful refusal falls from 87.9% in HRLs to 43.9% in LRLs, while harmless refusal remains low. A common explanation is that models re...

Rashad Aziz, Ikhlasul Akmal Hanif, Fajri Koto · 0 citations
#artificial intelligence Preprint Open access Oct 2026

EPIC: Efficient and Parallel Inference under CFG Constraints for Diffusion Language Models

Controlling language model outputs is essential for ensuring structural validity, reliability, and downstream usability, and diffusion language models are no exception. Recent advances in diffusion language model decoding have extended output control beyond regular constraints to context-free grammar (CFG) constraints....

Hyundong Jin, Yo-Sub Han · 0 citations
#artificial intelligence Preprint Open access Oct 2026

EUDAIMONIA: Evaluating Undesirable Dynamics in AI

Large language models (LLMs) are increasingly used as conversational partners for companionship, emotional disclosure, and interpersonal advice, but the social dynamics of these interactions can create harms that are not captured by current evaluations in real-world settings. We introduce the Social AI Design Code, a s...

Jun Rui Huang, Wang Bill Zhu, Ziyi Liu et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

A Controlled Synthetic Benchmark for Educational Aspect-Based Sentiment Analysis

Educational aspect-based sentiment analysis (ABSA) can support course improvement, but public aspect-labeled student feedback remains scarce because educational reviews are private, institution-specific, and expensive to annotate. This study introduces a controlled synthetic benchmark for educational ABSA built from 10...

Yehudit Aperstein, Alexander Apartsin · 0 citations

From tech blogs

See all →
MIT News · Artificial Intelligence Sep 24, 2026

Estimating suicide risk from text

A new language-processing tool could help identify the highest-risk individuals from natural language, enabling swifter interventions.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.