Skip to content

Category

natural language processing

6,613 papers

#natural language process... Preprint Open access Oct 2026

Evidence-Bound Reasoning: Neuro-Semantic Verification of Biomedical AI in Glioblastoma Radiogenomics

Background: Biomedical AI can generate plausible explanations without reliably verifying whether each statement is supported by patient-specific evidence. We developed a neuro-semantic verification framework that converts radiomic measurements into addressable evidence records and machine-checkable claims. Methods: UPe...

Mariya Miteva, Maria Nisheva-Pavlova · 0 citations

Towards In-Parameter Memory Augmentation for Large Language Models

Recently Large Language Models (LLMs) and LLM-based agents increasingly need to incorporate knowledge acquired after pretraining, e.g., domain facts, user preferences, documents, and interaction experience. In-context learning (ICL) and ICL-based agent harness remain flexible, but they consume context capacity and incu...

Haoyu Huang, Zhong-Wei Xie, Jiaxin Bai et al. · 0 citations
#natural language process... Preprint Open access Oct 2026

InterCorrect: Intersection-Aware Correction of Demographic Model Merging for Fair ASR

Automatic Speech Recognition (ASR) systems often show uneven performance across demographic groups, and errors can be especially difficult to address for speakers belonging to multiple demographic groups. This work studies demographic-aware model merging for fair Speech-LLM-based ASR. Starting from a SLAM-ASR-based mod...

Ashley E. Bravo-Bravo, Yuchen Zhang, Haralambos Mouratidis et al. · 0 citations
#natural language process... Preprint Oct 2026

Generative AI translations in high-stakes emergency messaging

Emergency messaging such as extreme-weather reports and earthquake instructions can involve high stakes, to the extent that translation errors can lead to tragic consequences. The use of machine translation or generative artificial intelligence might therefore not be recommended. On the other hand, time savings in the...

Nune Ayvazan, Anthony Pym, Yu Hao · 0 citations
#natural language process... Preprint Oct 2026

Incidental information contaminates patient notes and disrupts clinical reasoning in large language models

Large language models (LLMs) are increasingly relied upon to support ambient documentation and clinical reasoning. Here we examine the impact of a failure mode shared between these two applications by assessing their sensitivity to information incidental to the patient encounter. In 576 patient-clinician dialogues, we...

K. Vishwanath, Brandon Ye, A. Alyakin et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

How High Is 0.6? Floors, Ceilings, and Headroom in Interpretability Probing

Probes are the workhorse of interpretability. If a model's hidden states predict a variable, the model is said to represent it. But a probe score has no fixed meaning. An $R^2$ of 0.6 may only reflect what the input already gives away, and the same score can mean different things on different data. We propose reading e...

Pranjal Garg · 0 citations
#artificial intelligence Preprint Oct 2026

Wiki-Talkie: Multilingual Benchmarking of Persona-Based Agents on Real-World Discussions

LLMs are increasingly deployed as autonomous agents in social environments, making it critical to study their ability to faithfully simulate human interactions. Central to this is grounding agents in realistic user personas, yet existing datasets rely on fictional personas and are limited to a handful of languages, lac...

Dennis Fucci, Andrea Bacciu, Dong Liu et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Language-model ratings of depression reflect the rater more than the patient

Depression has no diagnostic blood test. Language models promise tireless, consistent assessment, but can accurate raters disagree about individuals? We pre-registered 880 language-model raters, crossing 11 open models with prompting and scoring choices, and applied them to 189 interviews against the eight-item Patient...

Baihan Lin · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Rethinking Cross-Tokenizer On-Policy Distillation: From Alignment Coverage to Supervision Reliability

On-Policy Distillation (OPD) trains a student on its own generations using teacher feedback. With different tokenizers, comparing teacher and student predictions requires alignment at both sequence and vocabulary levels. In this paper, we examine whether expanding this alignment coverage improves learning. Across three...

Bingxi Hou, Guochao Jiang, Guofeng Quan et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Foresight-over-Graph: Reasoning Beyond Local Horizons for Knowledge Base Question Answering

Large language models (LLMs) have demonstrated strong capabilities in question answering, yet they still frequently suffer from hallucinations on knowledge-intensive tasks. Knowledge graphs (KGs) provide LLMs with structured, interpretable, and updatable factual grounding, making them a promising external knowledge sou...

Yang Hong, Yajun Yang, Xin Wang et al. · 0 citations
#natural language process... Preprint Oct 2026

Language Unalignability: Why Some Concepts Resist Cross-Cultural Benchmark Evaluation

Current evaluation of multilingual Large Language Models (LLMs) rests on an implicit Translation-Isomorphism Assumption (TIA): that semantic structures across languages are congruent and mutually mappable without loss of information. We argue that this assumption is not merely violated in practice, but ill-posed in pri...

Shu-Kai Hsieh, Da-Chen Lian · 0 citations
#natural language process... Preprint Open access Oct 2026

Memory Depth and Reconstructed Context Width: A Controlled Evaluation of Hierarchical Retrieval

Long-term conversational memory is becoming an integral component of modern LLM systems. Proposed architectures group records by topics and events, construct hierarchies and graphs, and connect facts through causal and temporal relations. We experimentally study the interaction between two memory parameters: structural...

Michael Andreev · 0 citations

From tech blogs

See all →
MIT News · Artificial Intelligence Sep 24, 2026

Estimating suicide risk from text

A new language-processing tool could help identify the highest-risk individuals from natural language, enabling swifter interventions.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.