Skip to content

Category

natural language processing

6,613 papers

#artificial intelligence Preprint Open access Oct 2026

SWE-Journey: Towards More Realistic Evaluation of Coding Assistants through Long-Horizon, Multi-Turn Interaction

Coding assistants such as Claude Code and Codex have become a major application of LLM agents, yet existing benchmarks remain far from real-world use, particularly in task horizon and interaction length. Code assistants require completing long chains of development work in continuously evolving repositories, while repe...

Hexuan Deng, Yue Wang, Wenyu Jiang et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Constitutional Gating and Deterministic Recovery for Multi-Agent LLM Negotiation: Ablations Against a Stateful Adversarial Gatekeeper

Multi-agent LLM systems negotiating with a stateful counterpart waste model calls in three ways: polite loops that never meet the counterpart's hidden acceptance condition, malformed outputs that trigger retries, and compliance deadlocks in which the counterpart demands something the agent must refuse. We study a three...

Masaaki Nakatsu (AO, Inc. / OrbLabs AG), Reno Wang (AO et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

MetaEncoder: Exploring the Limit of Bi-Encoders for Multimodal System One Decision Making with Natural Language Interface

System One models output constrained decisions and probability distributions rather than free-form text generation. While prevailing paradigms rely on structured schema objects to encode state, intent, and candidate choices, we revisit a fully natural language-based System One interface. In this framework, both the use...

Jianpeng Cheng, Guangyu Sun, Aashu Singh et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

REMORY: Learning Residual Memory for Context Compaction

Long-horizon agents compact their history to continue within a finite context window, but a textual summary alone may not support every subsequent decision. We introduce REMORY, a neural memory network that supplements the summary with a bounded sequence of soft memory tokens. Given the history and summary, the network...

Hanchen Xia, Baoyou Chen, Yutang Ge et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Gated Memory: Admission-Controlled Memory Formation for Conversational AI

Personalized conversational AI relies on long-term memory systems that extract facts from user utterances and store them in persistent vector stores. Despite progress in retrieval, deduplication, and lifecycle management, the formation stage, the moment a fact is first written to storage has received almost no principl...

Preeti Saraswat, Divya Neelagiri, Ajay Manoj · 0 citations
#artificial intelligence Preprint Open access Oct 2026

SafeInferCom: Safe Inference-Time Compute via Verifier-Guided Mid-Generation Intervention for Robotic Task Planning

Large Reasoning Language Models (LRLMs) enable multi-step reasoning for robotic task planning, but continued reasoning can overwrite valid intermediate plans or leave constraint violations unresolved, reducing planning reliability and wasting inference-time computation. We develop an inference-time monitor that exposes...

Weizhe Xu, Jialiang Fan, Mengyu Liu et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

RAG-Stress: Probing the Limits of Evidence Reliance in Retrieval-Augmented Generation

Following retrieved evidence does not guarantee factual correctness: misleading evidence can induce a model to replace an answer it previously gave correctly. Standard accuracy measures obscure this behavior by combining answer replacement with preexisting errors. We introduce RAG-Stress, a controlled diagnostic protoc...

Shunyuan Zhou, Hao Chen, Tianyu Wang et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

ActiveMedAgent: Cost-Aware Trajectory Learning for Multimodal Medical Diagnosis

Clinical diagnosis is inherently sequential: clinicians escalate from cheap to costly tests only when additional evidence is expected to resolve diagnostic uncertainty. We present ActiveMedAgent, a framework that brings this cost-aware sequential logic to multimodal medical AI. Given a frozen, API-accessed vision-langu...

Weiwei Ma, Xiaobing Yu, Peijie Qiu et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

SFT-as-Context Mitigates Forgetting in Supervised Fine-Tuning

Supervised fine-tuning (SFT) equips large language models (LLMs) with specialized capabilities, but often comes at the cost of forgetting the general capabilities of their parent models (i.e., the pretrained models before fine-tuning). This trade-off is especially limiting for queries that require both specialized and...

Kenan Tang, Andong Hua, Chengxuan Qian et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

When Citations Mislead? A Claim-Level Benchmark for Legal Hallucination Detection

Large language models are increasingly used in legal research and drafting, but they can still produce claims that sound convincing without being supported by the cited source. We introduce PARCEL, a benchmark for checking whether a legal claim is supported by the underlying authority. Using recent New York State Court...

M. Mikail Demir, M. Abdullah Canbaz · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Stochastic Teacher Intervention for Agentic On-Policy Distillation

On-policy distillation (OPD) efficiently transfers capabilities from a stronger teacher to a student language model through dense token-level supervision on student-generated rollouts and has shown promise on complex tasks such as mathematical reasoning. However, in multi-turn agentic tasks, student decisions shape sub...

Junnan Liu, Linhao Luo, Zhijun Chen et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Real Long-Term Memory for AI: A 50-Million-Token Window That Is Faster and Cheaper Than Recompute

A large language model can only use the text that fits in its context window, and it recomputes its internal key-value (KV) state for a prompt every time the prompt is sent. We test a memory layer, the public package galahad-kv, that saves the KV state of each block of about 16,000 tokens to encrypted local NVMe disk a...

Sietse Schelpe · 0 citations

From tech blogs

See all →
MIT News · Artificial Intelligence Sep 24, 2026

Estimating suicide risk from text

A new language-processing tool could help identify the highest-risk individuals from natural language, enabling swifter interventions.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.