Skip to content

Category

natural language processing

6,613 papers

#machine learning Preprint Open access Oct 2026

Rethinking Faithfulness in LLMs: A Pairwise Context-Sensitive Perspective

Large language models (LLMs) are expected to answer questions faithfully based on the provided context, abstaining when the context information is insufficient to answer the questions. Existing faithfulness evaluations typically assess each question-context instance in isolation; however, such instance-level evaluation...

Zizhuo Zhang, Xiong Peng, Jingwei Sun et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

ReFold: Training-Free Reversible Inter-Turn Context Folding for Long-Horizon Agents

Long-horizon LLM agents act on an append-only interaction history that is re-sent to the model at every step, so the context and its cost grow with steps until the sessions exceed the context window. Existing methods manage the context through context requirement prediction, relying on additional model calls, heuristic...

Yupeng Su, Jiayi Tian, Zheng Zhang et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Persistent Memory in Multi-Agent LLM Inference: What It Costs, What It Buys, and When You Can Tell

Decomposing long-context inference across cooperating agents bounds the active KV cache per call rather than total evidence, which matters when KV-cache memory binds. Many such systems add a persistent tier storing and recalling reasoning traces, usually validated by an ablation reporting an accuracy gain. We measure b...

Hochan Son, Kyungdoe Han, Jaehan Koh et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Quantization Effects on Tool-Failure Recovery Vary Across Prompts and Evaluation Designs

Post-training quantization reduces the cost of deploying language-model agents, but its effect on recovery from temporary tool failures can depend on how recovery is evaluated. We compare 8-bit and 4-bit variants of Llama-3.1-8B-Instruct and Qwen2.5-7B-Instruct on twenty deterministic tool-use tasks and five prompts. T...

Yuhe Hu · 0 citations
#machine learning Preprint Open access Oct 2026

APEX: Speculate smarter, not deeper

Speculative decoding reduces large language model inference latency by drafting multiple tokens before target-model verification, but its effectiveness depends on both the proposal mechanism and draft depth. Fixed configurations cannot respond to changes in predictability, repetition, and acceptance during generation,...

Manvi Jha, Zach Zhang, Zhichao Xu et al. · 0 citations
#artificial intelligence Preprint Oct 2026

Learning to Retrieve via Reinforcement Learning in Embedding Space

Dense retrieval models are typically trained with contrastive objectives that learn effective representations but do not directly optimize retrieval metrics or downstream task performance. To address this problem, we introduce RELER (REinforcement LEarning for Retrieval), a reinforcement learning framework that enables...

Qi Liu, Feng-Ming Liang, Yi-Qun Chen et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

SanSi: A Looped Typed Decision Model for System 1.5 Thinking

Typed decision models answer a declared question without generating text: a decision head returns a probability for each of the declared options in a single forward pass. A single pass is fast, intuitive System 1 thinking. We study what lies between one pass and generated reasoning: looping, in which the same layers ar...

Shuyu Gan, Young-Jun Lee, Dongyeop Kang · 0 citations
#artificial intelligence Preprint Oct 2026

Recurrent Looped Transformer

State tracking requires an update at every input, but the depth a Transformer applies to each token is fixed regardless of sequence length. We introduce the Recurrent Looped Transformer (RLT), which splits its layers between a parallel causal encoder and a recurrent decoder. At each token, the decoder merges the encode...

Yi-Fan Zhang, Ji-Chen Feng, Shi-Han Qin · 0 citations
#machine learning Preprint Open access Oct 2026

Two Vectors Replace In-Context Demos: Structured Task Adaptation via Embeddings

In-context learning (ICL) adapts frozen large multimodal models (LMMs) to new tasks from a few demonstrations (demos), but re-encodes them at every query, where each demo image adds up to hundreds of visual tokens. Demo-free methods remove this cost with a compact task state. However, they add it at locations searched...

Xi Ding, Naichen Shi, Jiawei Zhang · 0 citations
#machine learning Preprint Open access Oct 2026

AccentCL: Robust Accent Classification with Incremental Expansion

Accent classifiers are typically trained with a fixed label inventory and cannot accommodate new accent categories as new data becomes available. Moreover, accented speech corpora often exhibit substantial class imbalance and/or domain shift due to differences in recording conditions across corpora. We present AccentCL...

Mu-Ruei Tseng, Waris Quamer, Ghady Nasrallah et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Logbook: Extremely Long-form Audio Event Understanding

Audio benchmarks are built around short, pre-segmented clips, limiting model design to brief inputs or fixed vocabularies. To close this gap, we introduce Logbook, a benchmark for hour-scale audio understanding, with recordings ranging from ten minutes to six days. Given a continuous audio recording and an event label...

Kwanghee Choi, Suwon Shon, Dmitriy Serdyuk et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Lineage-Aware Memory Governance: A Derivation-Gated Framework for Privacy-Preserving Column-Level Access Control in Enterprise AI Agents

Enterprise AI agents that share a memory store face two unaddressed risks: sensitive data can leak through legitimately computed results the requester could not derive, and departments can silently compute a same-named key performance indicator (KPI) through conflicting logic. Existing agent-memory systems (e.g., MemGP...

Venkata M Sangaraju, Sudhir Vissa · 0 citations

From tech blogs

See all →
MIT News · Artificial Intelligence Sep 24, 2026

Estimating suicide risk from text

A new language-processing tool could help identify the highest-risk individuals from natural language, enabling swifter interventions.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.