Skip to content

Category

natural language processing

6,613 papers

#natural language process... Preprint Open access Oct 2026

A generative-informed neuro-symbolic framework for syntactic ambiguity resolution: Evidence from Arabic DPs

Syntactic ambiguity poses a persistent challenge for Arabic NLP, particularly in morphologically rich nominal constructions where multiple structu6ral interpretations may be compatible with the same surface sequence. This study proposes a generatively informed neuro-symbolic framework for resolving structural ambiguity...

Mohammed Damom, Muneef Y. Alshawsh, Ashraf A. Naji et al. · 0 citations
#natural language process... Preprint Oct 2026

CUEing User Simulators: Calibrated User Embeddings for Multi-Turn Benchmarking

Recent benchmarks rely on user simulators to evaluate AI agents in multi-turn interaction. While existing simulation techniques demonstrate surface fidelity to human style and behavior, ecologically valid interactive benchmarking also requires alignment in when and how agents fail across simulated and real user populat...

Anjali Kantharuban, Jonas Mueller · 0 citations
#natural language process... Preprint Oct 2026

Finding the Move Is Not Winning the Game: XiangqiBench for Closed-Loop Evaluation of LLM Agents

Static evaluations credit a language model for naming the right move, but an agent must carry a plan through to a verified outcome while an opponent responds. We introduce XiangqiBench, an executable benchmark that measures this difference in Chinese chess: starting from 119 tactical endgames with forced mates supporte...

Yekun Chai, Qiwei Peng, Hao-Yi Xiong · 0 citations
#natural language process... Preprint Oct 2026

HakemBench: A Turkish Benchmark of Typed Decisions

HakemBench is a Turkish benchmark of typed decisions, in which the model under test reads a text, a question and a fixed set of options and returns a probability for every option. Version 1.0 is released fully open under CC BY 4.0, with 2,346 items and 4,275 choice, yes/no and score questions in seven tracks (fact-chec...

Sait Furkan Teke · 1 citation
#machine learning Preprint Open access Oct 2026

Useful Features, Backward Scores: OOD in Language-Model Trajectories

Out-of-distribution (OOD) detectors prioritize inputs for closer inspection. Yet features that distinguish input groups need not yield a useful anomaly ranking. We analyze this gap in language-model trajectories under text-length control and fixed score directions. On Spam development data, an input adaptation of D^2HS...

Hamidreza Saghir · 0 citations
#machine learning Preprint Open access Oct 2026

Denser $\neq$ Better: Limits of On-Policy Self-Distillation for Continual Post-Training

Continual post-training enables foundation models to acquire new knowledge while preserving existing capabilities. Recent work suggests that on-policy learning can mitigate forgetting, with self-distillation as a particularly attractive approach. We revisit this optimistic claim through self-distillation policy optimiz...

Meng Wang, Haohan Zhao, Wenzhuo Liu et al. · 0 citations
#machine learning Preprint Oct 2026

FALCON: A Model and Dataset Agnostic Framework for Synthetic Data Generation for NL2SQL Pairs

Relational databases are among the most widely deployed forms of structured knowledge, and natural language access to them requires grounding language onto schema entities and relations while handling the ambiguity inherent in how people phrase requests. Existing synthetic NL-to-SQL data generation methods largely igno...

Darian Lee, Shannon Rumsey, Jack St. Clair et al. · 0 citations
#machine learning Preprint Open access Oct 2026

MintEval: Do LLMs Implement the Trading Strategy You Asked For? A Behavioural-Equivalence Benchmark for Natural-Language-to-Strategy Code

Large language models are moving from producing trading signals to writing the code that executes them. The failure mode of the second role is silent: generated code runs, a backtest plots, yet the risk logic that the trader described is not the logic being executed. Existing code benchmarks test functional correctness...

Siyu Wang, Yifan Wang, Yuecheng He · 0 citations
#machine learning Preprint Oct 2026

HyperThink: Text-to-Parameter Hypernetworks for Efficient Reasoning

Long-form thinking traces can substantially improve the multi-step reasoning performance of large language models (LLMs), but they introduce high inference-time overhead, with latency dominated by sequential decoding. We propose HyperThink, a text-to-parameter approach that amortizes this reasoning computation into a s...

Donggyun Kim, Jack Lu, Chanwoo Kim et al. · 0 citations
#machine learning Preprint Oct 2026

Clinical Concept Centers in LLMs

Large language models are increasingly used in clinical settings. However, research into the reliability and performance of these models has focused almost entirely on the language substrate, scoring what the model says. Mechanistic interpretability has found that the latent space carries a higher fidelity of represent...

Aishik Nagar, A. Vaidyanathan, Arun-Kumar Kaliya-Perumal et al. · 0 citations
#machine learning Preprint Open access Oct 2026

RMCW: A Deletion-Robust Watermark Based on Reed--Muller Codes for Language Models

Large Language Model (LLM) watermarking provides a lightweight mechanism for identifying text generated by a specific model, but its robustness remains fragile under post-processing attacks. Deletion attacks are particularly challenging because they shift token positions and break the alignment between observed tokens...

Yi Wang, Baicheng Chen, Yu Wang et al. · 0 citations
#machine learning Preprint Open access Oct 2026

WakeKV: Reactive, Reversible KV Residency for Heads That Change Their Minds

Most KV-cache compression methods classify attention heads once, either offline or during prefill, and keep this classification fixed throughout generation. Across three models (1.5B-8B) and three regimes (needle retrieval, long chain-of-thought, and multi-turn recall), we measure head behavior on four model-regime com...

Utkarsh Ranjan · 0 citations

From tech blogs

See all →
MIT News · Artificial Intelligence Sep 24, 2026

Estimating suicide risk from text

A new language-processing tool could help identify the highest-risk individuals from natural language, enabling swifter interventions.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.