Large language models are increasingly used to simulate how individuals respond to new situations, yet the behavioral reasoning behind these responses is either inherited from pretraining or learned from individual-level annotations, which offer limited behavioral diversity and little supervision of the reasoning itsel...
Yining Zhao, Bushi Liu, Haofei Yu et al.· 0 citations
Block diffusion language models keep a large key-value (KV) cache throughout generation and attend to it at every denoising step, limiting both memory capacity and generation speed. Reducing these costs requires deciding which past tokens to use for denoising the current block (selection) and which to keep in memory fo...
Gleb Molodtsov, Ekaterina Alimaskina, Evgeny Uskov et al.· 0 citations
Reinforcement learning with verifiable rewards (RLVR) has become the dominant paradigm for eliciting multi-step reasoning in large language models, and a recent wave of methods (LUFFY, ExPO, PAPO, TAPO) further augments RL with \emph{external guidance} - expert traces, self-explanations, or retrieved thought patterns....
Sofia Torres, Gabriel Almeida, Carter Adams et al.· 0 citations
Data integration is the process of combining data from multiple, heterogeneous sources into a consistent, unified representation. Data integration involves a sequence of interdependent tasks including schema matching, value normalization, blocking, entity matching, and data fusion. Existing table-based benchmarks eithe...
Aaron Steiner, Ralph Peeters, Christian Bizer· 0 citations
Reach audiences
Advertise in front of researchers, engineers, and readers.
Polarization in online communities is often studied through either language or interaction structure, but the two views are rarely connected within a unified framework. Prior work has linked them by constructing interaction graphs from human judgements of agreement and disagreement, leaving a gap between language as ob...
Zhijin Guo, Li Zhang, Tyler Bonnet et al.· 0 citations
Unlike code completion, debugging requires localizing faults and applying targeted edits. We observe that frontier LLMs often regenerate correct but over-edited solutions during debugging. To evaluate how far LLMs are from precise debugging, we introduce the Precise Debugging Benchmark (PDB) framework, which automatica...
Miaosen Chai, Wang Bill Zhu, Shangshang Wang et al.· 0 citations
High-fidelity diagram creation requires the complex orchestration of semantic topology, visual styling, and spatial layout, posing a significant challenge for automated systems. Existing methods also suffer from a representation gap: pixel-based models often lack precise control, while code-based synthesis limits intui...
Tianfu Wang, Leilei Ding, Ziyang Tao et al.· 0 citations
Turning a pretrained language model (LM) into a vision-language model (VLM) through multimodal fine-tuning often erodes its native language ability, a form of catastrophic forgetting that shows up even on text-only tasks. This loss is hard to undo with further fine-tuning, and existing remedies add adapters or alignmen...
Patrick Amadeus Irawan, Erland Hilman Fuadi, Shanu Kumar et al.· 0 citations
Current Large Language Models have achieved Olympiad-level logic, yet Vision-Language Models paradoxically falter on elementary spatial tasks like block counting. This capability mismatch reveals a critical ``spatial intelligence gap,'' where models fail to construct coherent 3D mental representations from 2D observati...
Shaoxiong Zhan, Yanlin Lai, Zheng Liu et al.· 0 citations
The rapid advancement of long-context vision language models (LCVLMs) has led to a significant expansion of their context windows. However, an extended context window does not guarantee the effective utilization of the context, posing a critical challenge for real-world applications. Current evaluations of such long-co...
Keyan Zhou, Zecheng Tang, Lingfeng Ming et al.· 0 citations
The proliferation of disinformation demands reliable and scalable fact-checking solutions. We present Dynamic Evidence-based FAct-checking with Multimodal Experts (DEFAME), a modular, zero-shot MLLM pipeline for open-domain, text-image claim verification. DEFAME operates in a six-stage process, dynamically selecting th...
Tobias Braun, Mark Rothermel, Marcus Rohrbach et al.· 0 citations
Whether idiosyncratic, item-specific knowledge is learned before abstract class-level generalizations, or vice versa, is a central question in language learning, with exemplar and abstraction-based theories making opposite predictions. Recent methods have claimed to show that, at least for large language models, abstra...
Zachary Nicholas Houghton, Vsevolod Kapatsinski· 0 citations
A new method, called CW-Net, translates the reasoning process of an autonomous vehicle’s AI system into understandable concepts that explain its behavior.
MIT News · Artificial Intelligence· news.mit.eduAug 31, 2026
With millions of users across the world, Julia has been used to conduct cutting-edge research and to design new drugs, jet engines, heat pumps, and more.