Skip to content

Category

natural language processing

6,613 papers

#machine learning Preprint Oct 2026

Trained Agentic Context Management

We study long context language models. Instead of training long context natively, or designing a long context harness, we train a model over the simplest possible harness: a tool to call itself with any specified prompt and a tool to read tokens in a range from the input context. We finetune Qwen3.6-35B-A3B on a divers...

Bryce Sandlund · 0 citations
#machine learning Preprint Oct 2026

Hesitation Has a Geometry: Entropy-Trained Hyperbolic Probes for Sparse Activation Steering

When a large language model solves a mathematical problem, its reasoning is largely hierarchical, and the solution often branches at a few tokens where the next-token entropy is high. Such tree-like structure embeds in hyperbolic space with far lower distortion than in Euclidean space. Activation steering, however, usu...

Ze-Yong Zhang, Tung Sum Thomas Kwok, Teng-Fei Ma et al. · 0 citations
#machine learning Preprint Open access Oct 2026

Does Every User Need a Private LoRA? Decoupling Personalization from Per-User Adaptation

Personalized large language models often require a complete adaptation state for each user. However, this paradigm scales poorly as the user population grows. We revisit this design through the lens of personalization capacity allocation: how much adaptation capacity can be shared across users, how the shared capacity...

Songyuan Sui, Srikanth Malla, Chiho Choi et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Assessing Rule Adherence of LLM Adjudicators in Call of Cthulhu TRPG

As LLMs are increasingly deployed as autonomous adjudicators in games such as Call of Cthulhu (CoC), robust rule adherence becomes critical when user intent conflicts with system rules. However, as these models are trained to be helpful and compliant, they may be vulnerable to a class of manipulations we term Rhetorica...

Weiying Chen, Junlong Shen, Zhanyuan Guo et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

How Far Can You Get Without a GPU? A Systematic Benchmark of Lightweight Hallucination Detection Across Question Answering, Dialogue, and Summarisation

Hallucination detection has become a pressing requirement for trustworthy AI deployment at scale. The most accurate detection methods depend on GPU-intensive inference, proprietary API calls, or white-box access to the generating model, putting them out of reach for resource-constrained researchers and practitioners. W...

Kriti Faujdar, Smit Kadvani · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Morpheus: A Morphology-Aware Neural Tokenizer and Word Embedder for Turkish

Turkish is agglutinative: meaning is carried by morphemes, yet the subword tokenizers that drive modern language models split words by corpus statistics, fragmenting semantically loaded suffixes and -- in the case of WordPiece and rule-based analyzers -- failing to decode their output back to the original text. This pa...

Tolga \c{S}akar · 0 citations
#artificial intelligence Preprint Open access Oct 2026

A Language Model from 1913: Pretraining on Historical Text

While modern language models increasingly rely on ever-larger web corpora, we show that pretraining on historical text (e.g., pre-1913 text) in a data-constrained setting can produce a temporally grounded language model that still shows reasonable performance on language understanding. However, developing History LMs r...

Xiaoxi Luo, Zachary Shinnick, Niclas Griesshaber et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

On the Limits of LLM Adaptability: Impact of Model-Internalized Priors on Annotation Task Performance

Large Language Models (LLMs) are increasingly used for zero-shot annotation and LLM-as-a-judge tasks, yet their reliability hinges on how model-internalized priors interact with user-provided instructions. We investigate three dimensions of this interaction: (1) how an LLM's familiarity with data and task definitions r...

Etienne Casanova, Rafal Kocielnik, R. Michael Alvarez · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Where Do Apparent LLM Clinical Triage Failures Arise? Localizing the Multiple-Choice Format Effect

LLM evaluations using clinician-authored triage vignettes have reported substantial under-triage under constrained multiple-choice testing. Yet model performance on the same clinical cases can change when responses are generated in free text. We test whether this format effect appears while the case is processed or whe...

David Fraile Navarro, Berardino Como, Jialei Sheng et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

EchoDistill: Robust Large Audio Language Models via Noisy-to-Clean Self-Distillation

Large Audio Language Models (LALMs) remain vulnerable to acoustic noise, which can obscure task-relevant evidence and produce unreliable responses. We propose EchoDistill, a noisy-to-clean self-distillation framework that uses clean audio as privileged information during post-training. A noisy-input student samples can...

Kaiwen Luo, Chunxi Luo, Liang Lin et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

FastKernels: Benchmarking GPU Kernel Generation in Production

LLM-based agents for GPU kernel generation are advancing rapidly, but the benchmarks they optimize against evaluate kernels in isolation, with synthetic inputs and weak baselines, rewarding sandbox speedups that break or vanish in real inference systems. We introduce FastKernels, a benchmark of 384 tasks drawn from 47...

Gabriele Oliaro, Jaeseong Lee, Yichao Fu et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

HINT-SD: Targeted Hindsight Self-Distillation for Long-Horizon Agents

Training long-horizon LLM agents with reinforcement learning is challenging because sparse outcome rewards reveal whether a task succeeds, but not which intermediate actions caused the outcome or how they should be corrected. Recent methods alleviate this issue by generating rewards or textual hints from turn-level act...

Woongyeong Yeo, Yumin Choi, Taekyung Ki et al. · 0 citations

From tech blogs

See all →
MIT News · Artificial Intelligence Sep 24, 2026

Estimating suicide risk from text

A new language-processing tool could help identify the highest-risk individuals from natural language, enabling swifter interventions.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.