Skip to content

Category

natural language processing

6,613 papers

#natural language process... Preprint Open access Oct 2026

Can Language Models Actually Retrieve In-Context? Drowning in Documents at Million Token Scale

Language models (LMs) raise an intriguing alternative to vector-based retrieval: conditioning on an in-context corpus and directly generating a relevant answer. However, prior work has largely focused on proprietary systems or the smaller-scale reranking task, leaving corpus-scale in-context retrieval largely unexplore...

Siddharth Gollapudi, Nilesh Gupta, Prasann Singhal et al. · 0 citations
#computer vision Preprint Open access Oct 2026

AVOC: Enhancing Hour-Level Audio-Video Understanding in Omni-Modal LLMs via Retrieval-Inspired Token Compression

Multimodal Large Language Models have achieved remarkable progress in short-form audio-video understanding, yet long-form audio-video comprehension remains challenged by limited context windows and severe information redundancy. To address these bottlenecks, we propose AVOC, a framework for long-form audio-video unders...

Yijing Chen, Wenhui Tan, Xiaoyi Yu et al. · 0 citations
#natural language process... Preprint Open access Oct 2026

Deeper is Not Always Better: Mitigating the Alignment Tax via Confident Layer Decoding

Autoregressive generation in large language models (LLMs) conventionally decodes from the final layer, assuming that deeper representations yield more reliable next-token predictions. We revisit this assumption by revealing a recurring Guess-Refine-Perturb dynamic: early layers form coarse guesses, intermediate layers...

Xuanming Zhang, Sining Zhoubian, Yuxuan Chen et al. · 0 citations
#natural language process... Preprint Open access Oct 2026

Compression Is Not Evaluation-Neutral: Fixed RAG Compression Can Distort Reader Comparisons

Retrieval-augmented generation (RAG) evaluations often compare readers after a compressor has changed their evidence. This mixes two questions: which complete pipeline works best, and how much of a reader upgrade survives compression. We show that fixed compression can raise average pipeline accuracy while hiding most...

Sugam Panthi, Rabab Abdelfattah · 0 citations
#natural language process... Preprint Open access Oct 2026

Who Checks the Citations? Benchmarking Legal Hallucination Detection

Attorneys, judges, and pro se filers increasingly use AI to draft legal documents, yet these tools frequently fabricate citations. Despite predictions that newer models would hallucinate less or that court sanctions would deter negligent filers, we found over 1,000 filings containing fabricated citations---with this nu...

Patty Liu, Dominik Stammbach, Peter Henderson · 0 citations
#natural language process... Preprint Open access Oct 2026

Validation-Gated Causal Interventions for Interpreting High-Stakes Large Language Model Behavior: A Case Study in Suicidality Detection

Large language models are proposed to flag suicidal content, but representing a concept and acting on it are distinct. A suicidality feature is decodable from 0.5B parameters upward, while the behavior that uses it emerges only in the low billions. Where it emerges, it rests on a compact mid-network feature: in Llama-3...

Nafiz Ahmed, Sarah Sharif, Dingjing Shi et al. · 0 citations
#natural language process... Preprint Open access Oct 2026

Retrospective Progress-Aware Self-Refinement for LLM Agent Training

Long-horizon LLM-based agents receive rich environmental observations during interaction, yet outcome rewards provide limited explicit supervision about how individual actions advance task completion. We investigate whether agents can turn this interaction evidence into useful training signals through retrospective pro...

Xinbei Ma, Congmin Zheng, Jiyang Qiu et al. · 0 citations
#natural language process... Preprint Open access Oct 2026

Agent Planning Benchmark: A Diagnostic Framework for Planning Capabilities in LLM Agents

Planning is central to LLM agents: before acting, an agent must decompose goals, select tools, reason over constraints, and decide when a task is infeasible. Yet existing agent evaluations often report only end-to-end success, making it difficult to determine whether failures stem from planning or execution. We introdu...

Haoyu Sun, Wenxuan Wang, Mingyang Song et al. · 0 citations
#natural language process... Preprint Open access Oct 2026

SMADE-IE: Sparse Multi-Agent Framework with Evidence-Driven Debate for Zero-Shot Information Extraction

Zero-shot information extraction (IE) with large language models (LLMs) enables adaptation to new schemas and domains without task-specific training. Existing methods mainly follow three paradigms. Monolithic prompting is efficient but prone to missed mentions, boundary errors, and type confusion. Each-type prompting i...

Kenfeng Huang, Yi Cai, Xin Wu et al. · 0 citations
#natural language process... Preprint Open access Oct 2026

ExpHarness: Model-Agnostic Experience Learning through a Trainable Harness

Large language model (LLM) agents increasingly operate within a harness, the scaffolding that determines what enters the executor's context, yet the experience they accumulate across tasks rarely flows back into this harness. Existing approaches include executor fine-tuning and external memory retrieval, but combining...

Tao Feng, Chongrui Ye, Fangxu Yu et al. · 0 citations
#natural language process... Preprint Open access Oct 2026

ElasticMem: Latent Memory as a Learnable Resource for LLM Agents

Long-term memory is essential for LLM agents to reason coherently across extended interactions, personalize responses, and reuse past experience. However, existing memory-augmented methods typically treat memory as a fixed resource: text-space approaches concatenate retrieved memories into the context window, causing s...

Tao Feng, Chongrui Ye, Fangxu Yu et al. · 0 citations
#natural language process... Preprint Open access Oct 2026

Beyond Phones: Structured Phonemic Modeling for Vietnamese Automatic Speech Recognition

Phone-based representations provide a compact and acoustically grounded alternative to conventional orthographic modeling for automatic speech recognition (ASR). However, phones describe surface pronunciations and may lose lexical distinctions under dialect-dependent sound mergers, making their conversion back to ortho...

Nghia Hieu Nguyen, Quan Ngoc Hoang, Long Hoang Huu Nguyen et al. · 0 citations

From tech blogs

See all →
MIT News · Artificial Intelligence Sep 24, 2026

Estimating suicide risk from text

A new language-processing tool could help identify the highest-risk individuals from natural language, enabling swifter interventions.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.