Agentic task-oriented dialogue (TOD) requires systems to track concurrent goals, dependencies, and long-horizon state. We examine goal-lifecycle recovery from fixed dialogue trajectories. ATOD contains 1,000 synthetic dialogues annotated for six advanced-TOD properties, and ATOD-Eval defines metrics for dependency-sens...
Yifei Zhang, Hooshang Nayyeri, Rinat Khaziev et al.· 0 citations
Matching place names across writing systems is a persistent obstacle to integrating multilingual geographic sources, from modern gazetteers to medieval itineraries and colonial-era surveys. Existing approaches rely on language-specific phonetic algorithms or on romanisation that discards phonetic information, and none...
Environmental, Social, and Governance (ESG) reports have become central to how companies communicate climate risk, social impact, and governance practices, yet they are still published primarily as long, heterogeneous PDF documents. This makes it difficult to systematically answer seemingly simple questions. Existing t...
Yi Ding, Xushuo Tang, Zhengyi Yang et al.· 0 citations
As model context lengths continue to grow, concerns about whether models effectively use the full context length have persisted. While several carefully designed long-context evaluations have recently been released, these evaluations tend to rely on retrieval from one or more sections of the context, which allows nearl...
Amanda Bertsch, Adithya Pratapa, Teruko Mitamura et al.· 0 citations
Reach audiences
Advertise in front of researchers, engineers, and readers.
Multimodal retrieval, which seeks to retrieve relevant content across modalities such as text or image, supports applications from AI search to contents production. Despite the success of separate-encoder approaches like CLIP aligning modality-specific embeddings with contrastive learning, recent multimodal large langu...
Qiyu Wu, Shuyang Cui, Satoshi Hayakawa et al.· 0 citations
Hybrid thinking enables LLMs to switch between reasoning and direct answering, offering a balance between efficiency and reasoning capability. Yet our experiments reveal that current hybrid thinking LLMs only achieve partial mode separation: reasoning behaviors often leak into the no-think mode. To understand and mitig...
Shouren Wang, Wang Yang, Xianxuan Long et al.· 0 citations
Autonomous agents executing human instructions must operate reliably even when instructions are incomplete. While recent approaches improve detection of missing information, detection alone is insufficient: agents often proceed to execution even after recognizing underspecification, leading to incorrect or unsafe actio...
Swarnadeep Bhar, Omar Naim, Eleni Metheniti et al.· 0 citations
Current approaches for Multimodal Sentiment Analysis (MSA) primarily leverage the knowledge and reasoning capabilities of parameter-heavy (Multimodal) LLMs for classification, overlooking autonomous multimodal sentiment reasoning generation in resource-constrained environments. In this paper, we focus on the Resource-L...
Haonan Shangguan, Xiaocui Yang, Shi Feng et al.· 0 citations
Humanity's Last Exam (HLE) is widely used to evaluate frontier language models. HLE organizes its questions into eight subject-domain categories, whose subscores are often interpreted as evidence of distinct capabilities. However, no study has assessed whether these labels correspond to empirically separable latent con...
Diffusion Language Models (DLMs) have recently emerged as a promising alternative to autoregressive models. Unlike standard diffusion-based approaches, DLMs are not explicitly conditioned on a timestep, raising a natural question: do these models internally represent denoising progress, and how is such information used...
Maximo Eduardo Rulli, Thomas Vaitses Fontanari, Simone Petruzzi et al.· 0 citations
Humanity is a mosaic of multifaceted talents and needs, and any truly intelligent AI must reflect that richness. Yet the LLMs used by all are built by the few -- a centralized market of monolithic AI models structurally ill-suited to capture the diversity of human knowledge, reasoning, and values. Here we introduce sca...
Shangbin Feng, Yike Wang, Weijia Shi et al.· 0 citations
Reinforcement learning with verifiable rewards (RLVR) is a scalable paradigm for improving the mathematical reasoning of large language models, but it is fundamentally limited by exploration: the policy can only improve on trajectories it has already sampled. Sampling more rollouts alleviates this at prohibitive comput...
Chanuk Lee, Sangwoo Park, Minki Kang et al.· 0 citations
A new method, called CW-Net, translates the reasoning process of an autonomous vehicle’s AI system into understandable concepts that explain its behavior.
MIT News · Artificial Intelligence· news.mit.eduAug 31, 2026
With millions of users across the world, Julia has been used to conduct cutting-edge research and to design new drugs, jet engines, heat pumps, and more.