Language models (LMs) raise an intriguing alternative to vector-based retrieval: conditioning on an in-context corpus and directly generating a relevant answer. However, prior work has largely focused on proprietary systems or the smaller-scale reranking task, leaving corpus-scale in-context retrieval largely unexplore...
Siddharth Gollapudi, Nilesh Gupta, Prasann Singhal et al.· 0 citations
Multimodal Large Language Models have achieved remarkable progress in short-form audio-video understanding, yet long-form audio-video comprehension remains challenged by limited context windows and severe information redundancy. To address these bottlenecks, we propose AVOC, a framework for long-form audio-video unders...
Yijing Chen, Wenhui Tan, Xiaoyi Yu et al.· 0 citations
Autoregressive generation in large language models (LLMs) conventionally decodes from the final layer, assuming that deeper representations yield more reliable next-token predictions. We revisit this assumption by revealing a recurring Guess-Refine-Perturb dynamic: early layers form coarse guesses, intermediate layers...
Xuanming Zhang, Sining Zhoubian, Yuxuan Chen et al.· 0 citations
Retrieval-augmented generation (RAG) evaluations often compare readers after a compressor has changed their evidence. This mixes two questions: which complete pipeline works best, and how much of a reader upgrade survives compression. We show that fixed compression can raise average pipeline accuracy while hiding most...
Sugam Panthi, Rabab Abdelfattah· 0 citations
Reach audiences
Advertise in front of researchers, engineers, and readers.
Attorneys, judges, and pro se filers increasingly use AI to draft legal documents, yet these tools frequently fabricate citations. Despite predictions that newer models would hallucinate less or that court sanctions would deter negligent filers, we found over 1,000 filings containing fabricated citations---with this nu...
Patty Liu, Dominik Stammbach, Peter Henderson· 0 citations
Large language models are proposed to flag suicidal content, but representing a concept and acting on it are distinct. A suicidality feature is decodable from 0.5B parameters upward, while the behavior that uses it emerges only in the low billions. Where it emerges, it rests on a compact mid-network feature: in Llama-3...
Nafiz Ahmed, Sarah Sharif, Dingjing Shi et al.· 0 citations
Long-horizon LLM-based agents receive rich environmental observations during interaction, yet outcome rewards provide limited explicit supervision about how individual actions advance task completion. We investigate whether agents can turn this interaction evidence into useful training signals through retrospective pro...
Xinbei Ma, Congmin Zheng, Jiyang Qiu et al.· 0 citations
Planning is central to LLM agents: before acting, an agent must decompose goals, select tools, reason over constraints, and decide when a task is infeasible. Yet existing agent evaluations often report only end-to-end success, making it difficult to determine whether failures stem from planning or execution. We introdu...
Haoyu Sun, Wenxuan Wang, Mingyang Song et al.· 0 citations
Zero-shot information extraction (IE) with large language models (LLMs) enables adaptation to new schemas and domains without task-specific training. Existing methods mainly follow three paradigms. Monolithic prompting is efficient but prone to missed mentions, boundary errors, and type confusion. Each-type prompting i...
Large language model (LLM) agents increasingly operate within a harness, the scaffolding that determines what enters the executor's context, yet the experience they accumulate across tasks rarely flows back into this harness. Existing approaches include executor fine-tuning and external memory retrieval, but combining...
Tao Feng, Chongrui Ye, Fangxu Yu et al.· 0 citations
Long-term memory is essential for LLM agents to reason coherently across extended interactions, personalize responses, and reuse past experience. However, existing memory-augmented methods typically treat memory as a fixed resource: text-space approaches concatenate retrieved memories into the context window, causing s...
Tao Feng, Chongrui Ye, Fangxu Yu et al.· 0 citations
Phone-based representations provide a compact and acoustically grounded alternative to conventional orthographic modeling for automatic speech recognition (ASR). However, phones describe surface pronunciations and may lose lexical distinctions under dialect-dependent sound mergers, making their conversion back to ortho...
Nghia Hieu Nguyen, Quan Ngoc Hoang, Long Hoang Huu Nguyen et al.· 0 citations
A new method, called CW-Net, translates the reasoning process of an autonomous vehicle’s AI system into understandable concepts that explain its behavior.
MIT News · Artificial Intelligence· news.mit.eduAug 31, 2026
With millions of users across the world, Julia has been used to conduct cutting-edge research and to design new drugs, jet engines, heat pumps, and more.