Skip to content

Category

natural language processing

6,613 papers

#artificial intelligence Preprint Open access Oct 2026

ATOD: An Evaluation Framework and Benchmark for Agentic Task-Oriented Dialogue Systems

Agentic task-oriented dialogue (TOD) requires systems to track concurrent goals, dependencies, and long-horizon state. We examine goal-lifecycle recovery from fixed dialogue trajectories. ATOD contains 1,000 synthetic dialogues annotated for six advanced-TOD properties, and ATOD-Eval defines metrics for dependency-sens...

Yifei Zhang, Hooshang Nayyeri, Rinat Khaziev et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Symphonym: Universal Phonetic Embeddings for Cross-Script Toponym Matching

Matching place names across writing systems is a persistent obstacle to integrating multilingual geographic sources, from modern gazetteers to medieval itineraries and colonial-era surveys. Existing approaches rely on language-specific phonetic algorithms or on romanisation that discards phonetic information, and none...

Stephen Gadd · 0 citations
#artificial intelligence Preprint Open access Oct 2026

EulerESG: Automating ESG Disclosure Analysis with LLMs

Environmental, Social, and Governance (ESG) reports have become central to how companies communicate climate risk, social impact, and governance practices, yet they are still published primarily as long, heterogeneous PDF documents. This makes it difficult to systematically answer seemingly simple questions. Existing t...

Yi Ding, Xushuo Tang, Zhengyi Yang et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Oolong: Evaluating Long Context Reasoning and Aggregation Capabilities

As model context lengths continue to grow, concerns about whether models effectively use the full context length have persisted. While several carefully designed long-context evaluations have recently been released, these evaluations tend to rely on retrieval from one or more sections of the context, which allows nearl...

Amanda Bertsch, Adithya Pratapa, Teruko Mitamura et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

MCA: Modality Composition Awareness for Robust Composed Multimodal Retrieval

Multimodal retrieval, which seeks to retrieve relevant content across modalities such as text or image, supports applications from AI search to contents production. Despite the success of separate-encoder approaches like CLIP aligning modality-specific embeddings with contrastive learning, recent multimodal large langu...

Qiyu Wu, Shuyang Cui, Satoshi Hayakawa et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Demystifying Hybrid Thinking: Can LLMs Truly Switch Between Think and No-Think?

Hybrid thinking enables LLMs to switch between reasoning and direct answering, offering a balance between efficiency and reasoning capability. Yet our experiments reveal that current hybrid thinking LLMs only achieve partial mode separation: reasoning behaviors often leak into the no-think mode. To understand and mitig...

Shouren Wang, Wang Yang, Xianxuan Long et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

COCORELI: Enforcing Execution Preconditions for Reliable Collaborative Instruction Following

Autonomous agents executing human instructions must operate reliably even when instructions are incomplete. While recent approaches improve detection of missing information, detection alone is insufficient: agents often proceed to execution even after recognizing underspecification, leading to incorrect or unsafe actio...

Swarnadeep Bhar, Omar Naim, Eleni Metheniti et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Resource-Limited Joint Multimodal Sentiment Reasoning and Classification via Chain-of-Thought Enhancement and Distillation

Current approaches for Multimodal Sentiment Analysis (MSA) primarily leverage the knowledge and reasoning capabilities of parameter-heavy (Multimodal) LLMs for classification, overlooking autonomous multimodal sentiment reasoning generation in resource-constrained environments. In this paper, we focus on the Resource-L...

Haonan Shangguan, Xiaocui Yang, Shi Feng et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Dimensionality and Measurement Precision in HLE's Multiple-Choice Subset

Humanity's Last Exam (HLE) is widely used to evaluate frontier language models. HLE organizes its questions into eight subject-domain categories, whose subscores are often interpreted as evidence of distinct capabilities. However, no study has assessed whether these labels correspond to empirically separable latent con...

Mayank Sharma, Savira Nadela, Tyler Matteson · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Subliminal Clocks: Latent Time Modelling in Diffusion Language Models

Diffusion Language Models (DLMs) have recently emerged as a promising alternative to autoregressive models. Unlike standard diffusion-based approaches, DLMs are not explicitly conditioned on a timestep, raising a natural question: do these models internally represent denoising progress, and how is such information used...

Maximo Eduardo Rulli, Thomas Vaitses Fontanari, Simone Petruzzi et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Scaling Participation in Modular AI Systems

Humanity is a mosaic of multifaceted talents and needs, and any truly intelligent AI must reflect that richness. Yet the LLMs used by all are built by the few -- a centralized market of monolithic AI models structurally ill-suited to capture the diversity of human knowledge, reasoning, and values. Here we introduce sca...

Shangbin Feng, Yike Wang, Weijia Shi et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Nudging Beyond the Comfort Zone: Efficient Strategy-Guided Exploration for RLVR

Reinforcement learning with verifiable rewards (RLVR) is a scalable paradigm for improving the mathematical reasoning of large language models, but it is fundamentally limited by exploration: the policy can only improve on trajectories it has already sampled. Sampling more rollouts alleviates this at prohibitive comput...

Chanuk Lee, Sangwoo Park, Minki Kang et al. · 0 citations

From tech blogs

See all →
MIT News · Artificial Intelligence Sep 24, 2026

Estimating suicide risk from text

A new language-processing tool could help identify the highest-risk individuals from natural language, enabling swifter interventions.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.