Skip to content

Category

natural language processing

6,613 papers

#artificial intelligence Preprint Open access Oct 2026

GISTBench: Evaluating LLM User Understanding via Evidence-Based Interest Verification

We introduce GISTBench, a benchmark for evaluating Large Language Models' (LLMs) ability to understand users from their interaction histories in recommendation systems. Unlike traditional RecSys benchmarks that focus on item prediction accuracy, our benchmark evaluates how well LLMs can extract and verify user interest...

Iordanis Fostiropoulos, Muhammad Rafay Azhar, Abdalaziz Sawwan et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

On the Tip of the Tongue: Why LLMs Hallucinate Answers They Can Decode

A language model can give the wrong answer even when the correct answer is decodable from its intermediate states. To study this gap between decodability and selection, we distinguish \textit{read} from \textit{write} at the first answer token. Read asks whether the gold token can be decoded from intermediate residual...

Valeria Ruscio, Keiran Thompson · 0 citations
#artificial intelligence Preprint Oct 2026

FrugalEvo: Towards Cost-Aware LLM-Guided Program Evolution

LLM-guided evolutionary methods, such as AlphaEvolve, have emerged as powerful approaches for challenging computational optimization problems, such as circle packing. However, prior work typically optimizes performance gain over a fixed number of iterations. We argue that practical optimization should maximize gain per...

Hui Chen, Xuan Qi, J. Zhao et al. · 0 citations
#artificial intelligence Preprint Oct 2026

SyntaxBench: A Statistical Diagnostic Framework for Character-Level Reasoning in Large Language Models

Large language models are increasingly used where small syntactic errors matter, yet character-level reasoning is still evaluated mostly through isolated probes and aggregate accuracy. We introduce SyntaxBench, a diagnostic benchmark and statistical evaluation framework for character-level reasoning. It contains five c...

Mohsen Larni, Sobhan Ebrahimi Azar, Pouyan Nahed et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Not Until the Evidence Says So: Teaching LLM Investigators When to Close a Case

Accident, defect and outage investigations end with a decision that ordinary question answering never faces: whether the evidence gathered so far is enough to close the case. We study this decision for LLM investigators, which request evidence from a case file, revise their hypotheses, and either close the case with a...

Tingzhu Bi, Ping Wang, Meng Ma · 0 citations
#artificial intelligence Preprint Oct 2026

Predicting Steering Vectors and Adapter Weights for Few-Shot Author-Style Transfer

Adapting large language models to an individual author's style from a few examples is challenging, and scientific writing sharpens the difficulty: formal conventions leave little surface variation, and authors write about their own topics, so extracted ``style''easily entangles with content. We study style-conditioned...

L. Popp, Danni Liu, Supriti Sinhamahapatra et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

The Fragility of Trigger-Tag Mechanisms for Misuse Detection in Open-Weight LLMs

Open-weight language models can be downloaded, modified, and deployed beyond their developers' control, limiting the effectiveness of centrally enforced safeguards. Recent work has therefore proposed \emph{trigger-tag} mechanisms that produce a detectable signal when a model is used under a target condition, such as ge...

Toluwani Aremu, Manit Baser, Mohan Gurusamy et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Ask, Relax, or Act? Evaluating Actionable Indeterminacy in LLM Preference Reasoning

An LLM agent can recognize uncertainty yet still choose the wrong next step: asking when action is already justified, or seeking clarification when the constraints must change. We formalize actionable indeterminacy: act when an accepted action is shared across all admissible preferences or objectives, clarify when each...

Ang Li, Yue Lin, Feifei Kou et al. · 0 citations
#artificial intelligence Preprint Oct 2026

Tailoring the Quantization Space for 1-Bit KV Cache Compression

The key-value (KV) cache becomes a major memory bottleneck in long-context LLM inference, placing substantial pressure on memory capacity and bandwidth. To mitigate this bottleneck, vector quantization (VQ) has emerged as a promising approach for aggressive KV cache compression. However, existing VQ methods degrade sub...

Minsoo Cheong, Donghyun Son, Sungjoo Yoo · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Sentry: Learning to Recover from LLM Agent Failures at Test Time

LLM agents often fail mid-task due to invalid tool calls, repeated actions, or poorly grounded reasoning, and learning from these failures is a path to reliability. We find that how failure knowledge reaches the agent matters as much as what it contains. Failure lessons are conditional: kept in the agent's context, the...

Changxiu Ji, Amy Lu, Qizheng Zhang et al. · 0 citations
#artificial intelligence Preprint Oct 2026

Misinformation Without Triggers: From Factual Answers to Downstream Decisions

Language models learn from web documents, some of them false, and false content can reach a model's answer to a factual question and the summaries and decisions that use it. Most data-poisoning studies add a trigger to the training data and activate it in the prompt. False documents can also change factual responses wi...

Lin Tian, M. Rizoiu · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Query-aware routing for Cross-lingual performance gains in Encoders

Multilingual encoders can exhibit reduced retrieval effectiveness when queries and relevant documents differ in language, despite strong same-language performance. We investigate whether Finnish and Swedish cross-lingual retrieval can improve while preserving an encoder's existing same-language performance and document...

Akshay Jain, Edward Kim · 0 citations

From tech blogs

See all →
MIT News · Artificial Intelligence Sep 24, 2026

Estimating suicide risk from text

A new language-processing tool could help identify the highest-risk individuals from natural language, enabling swifter interventions.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.