Skip to content

Category

natural language processing

6,613 papers

#computer vision Preprint Oct 2026

The Failure Is in the Readout: Fine-Grained Emotion Recognition Benchmarks Measure Elicitation, Not Perception

Fine-grained emotion recognition supports therapy tools and social robots, but it needs facial data, which raises privacy and data-protection concerns. EmoNet-Face-HQ answers that with generated portraits, expert-rated over a $40$-category taxonomy far finer than the usual six to eight basic emotions. Under the protoco...

T. Hallmen, Fabian Deuser, Robin-Nico Kampa et al. · 0 citations
#natural language process... Preprint Oct 2026

Conversation Is a Two-Body Problem: Dyadic Evaluation of Full-Duplex Dialogue Models

Full-duplex spoken dialogue models listen and speak at the same time, enabling voice agents to have natural, low-latency interactions that turn-based systems cannot offer. However, they are commonly evaluated against single-sided interlocutors: pre-recorded audio that cannot react, or an automated examiner that reacts...

Sungnyun Kim, Sungwoo Cho, Ji-Yun Oh et al. · 0 citations
#artificial intelligence Preprint Oct 2026

Natural Language Questions as an Interface for Knowledge Graphs: QRAKEN Graph Distillation and Semantic Self-Healing

Natural-language access to RDF knowledge graphs is a core Semantic Web ambition. Large language models (LLMs) have advanced Text-to-SPARQL, yet on unfamiliar graphs they often generate valid queries that misrepresent the populated data model. QRAKEN is a training-free, ontology-agnostic neurosymbolic pipeline grounding...

Remo Grillo, L. Klic, Giovanni Colavizza · 0 citations
#artificial intelligence Preprint Oct 2026

POLAR: Ontology-Guided Risk Prevention for Tool-Calling LLM Agents

LLM tool-use agents operate in dynamic environments where many actions carry operational risk. However, most safety mechanisms react only after errors manifest. Existing pre-emptive approaches either fine-tune the agent on chain-of-thought deliberation or compile natural-language guardrails into runtime checks, but the...

Yun-Ju Kang, Seonghyeon Cho, Irene Li et al. · 0 citations
#natural language process... Preprint Open access Oct 2026

HINTT Submission to the 2nd MLC-SLM Challenge: Comparing Cascaded and Unified Approaches to Diarization and ASR

This paper presents the HINTT system submitted to the 2nd Challenge and Workshop on Multilingual Conversational Speech Language Model (MLC-SLM). We address multilingual speaker-attributed ASR, where systems must determine who spoke when and what was spoken. We investigate two modeling strategies for this problem: a cas...

Takanori Ashihara, Kohei Matsuura, Masato Mimura · 0 citations
#artificial intelligence Preprint Oct 2026

VisionWeave: Weaving Elastic Visual Representations as a Native Capability of MLLMs

Multimodal large language models have become the dominant paradigm for visual understanding, but incur substantial costs by encoding inputs into dense, fixed-size patch tokens. However, visual information is unevenly distributed: some regions require fine-grained detail, while others admit compact representations. Down...

Yuan Feng, Qi-Ze Yang, Rui-Zhe Chen et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Confidence Reasoning Graphs: Structured Confidence Estimation for LLM Agents

When using an LLM agent in a consequential domain, making an informed decision about whether to trust its output or intervene requires calibrated confidence in the agent's success. Confidence estimation for agents is difficult because evidence about success is distributed across heterogeneous, interdependent steps of a...

Brendan King, Farima Fatahi Bayat, Jean-Flavien Bussotti et al. · 0 citations
#artificial intelligence Preprint Oct 2026

Harness Engineering for Software Engineering via Modular Executable Dev-Primitives

Large language models (LLMs) equipped with terminal access have demonstrated strong capabilities in automating software engineering tasks. However, existing agents remain brittle on long-horizon workflows, where they must repeatedly reconstruct program state scattered across source files, configurations, tests, depende...

Hai-Bo Jin, Xin-Jie Li, Peng Kuang et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

ThinkFuse: Trajectory-Aware Test-Time Fusion for Small Reasoning Models

Small reasoning models (SRMs) have shown strong performance on complex reasoning tasks by generating extended chain-of-thought trajectories, but they often fail to recover once their reasoning enters an erroneous path. Existing test-time fusion methods rely on local fusion signals to determine when to trigger fusion, w...

Myunghoon Kang, Jungseob Lee, Jaehyung Seo et al. · 0 citations
#natural language process... Preprint Open access Oct 2026

When Old Facts Return: Re-Reads, Reverts, and the Limits of Temporal Memory

A memory system can retire an obsolete value and later restore it merely because the same old statement appears again. A re-read of an old source and a genuine revert can produce the same observed sequence of values while requiring opposite current answers. We study this ambiguity on 130 extractor-selected atomic trans...

Neeraj Yadav (Called It Inc.) · 0 citations
#artificial intelligence Preprint Oct 2026

Where Rules End and Judges Begin: Measuring the Judgment Boundary in Multi-Agent Systems Security

LLM-based multi-agent systems (MAS) engage tools, share memory, and delegate tasks, often encountering adversarial content. Current defenses for MAS are typically evaluated in isolation, focusing on one attack type at a time, which can lead to costly and hard-to-audit outcomes. This study organizes defenses into five p...

Shaswata Mitra, Raj Patel, Subash Neupane et al. · 0 citations
#natural language process... Preprint Open access Oct 2026

Loud and Clear: Dynamic Activation Steering for Improving Speech Intelligibility in Noisy Environments

Speech becomes less intelligible in noisy environments, and humans naturally adapt their voice to compensate. Inspired by this behavior, we investigate whether a text-to-speech (TTS) model can be guided to produce more intelligible speech using activation steering, without retraining. We focus on two characteristics of...

Seymanur Akti, Alexander Waibel · 0 citations

From tech blogs

See all →
MIT News · Artificial Intelligence Sep 24, 2026

Estimating suicide risk from text

A new language-processing tool could help identify the highest-risk individuals from natural language, enabling swifter interventions.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.