We study long context language models. Instead of training long context natively, or designing a long context harness, we train a model over the simplest possible harness: a tool to call itself with any specified prompt and a tool to read tokens in a range from the input context. We finetune Qwen3.6-35B-A3B on a divers...
When a large language model solves a mathematical problem, its reasoning is largely hierarchical, and the solution often branches at a few tokens where the next-token entropy is high. Such tree-like structure embeds in hyperbolic space with far lower distortion than in Euclidean space. Activation steering, however, usu...
Ze-Yong Zhang, Tung Sum Thomas Kwok, Teng-Fei Ma et al.· 0 citations
Personalized large language models often require a complete adaptation state for each user. However, this paradigm scales poorly as the user population grows. We revisit this design through the lens of personalization capacity allocation: how much adaptation capacity can be shared across users, how the shared capacity...
Songyuan Sui, Srikanth Malla, Chiho Choi et al.· 0 citations
As LLMs are increasingly deployed as autonomous adjudicators in games such as Call of Cthulhu (CoC), robust rule adherence becomes critical when user intent conflicts with system rules. However, as these models are trained to be helpful and compliant, they may be vulnerable to a class of manipulations we term Rhetorica...
Weiying Chen, Junlong Shen, Zhanyuan Guo et al.· 0 citations
Reach audiences
Advertise in front of researchers, engineers, and readers.
Hallucination detection has become a pressing requirement for trustworthy AI deployment at scale. The most accurate detection methods depend on GPU-intensive inference, proprietary API calls, or white-box access to the generating model, putting them out of reach for resource-constrained researchers and practitioners. W...
Turkish is agglutinative: meaning is carried by morphemes, yet the subword tokenizers that drive modern language models split words by corpus statistics, fragmenting semantically loaded suffixes and -- in the case of WordPiece and rule-based analyzers -- failing to decode their output back to the original text. This pa...
While modern language models increasingly rely on ever-larger web corpora, we show that pretraining on historical text (e.g., pre-1913 text) in a data-constrained setting can produce a temporally grounded language model that still shows reasonable performance on language understanding. However, developing History LMs r...
Xiaoxi Luo, Zachary Shinnick, Niclas Griesshaber et al.· 0 citations
Large Language Models (LLMs) are increasingly used for zero-shot annotation and LLM-as-a-judge tasks, yet their reliability hinges on how model-internalized priors interact with user-provided instructions. We investigate three dimensions of this interaction: (1) how an LLM's familiarity with data and task definitions r...
Etienne Casanova, Rafal Kocielnik, R. Michael Alvarez· 0 citations
LLM evaluations using clinician-authored triage vignettes have reported substantial under-triage under constrained multiple-choice testing. Yet model performance on the same clinical cases can change when responses are generated in free text. We test whether this format effect appears while the case is processed or whe...
David Fraile Navarro, Berardino Como, Jialei Sheng et al.· 0 citations
Large Audio Language Models (LALMs) remain vulnerable to acoustic noise, which can obscure task-relevant evidence and produce unreliable responses. We propose EchoDistill, a noisy-to-clean self-distillation framework that uses clean audio as privileged information during post-training. A noisy-input student samples can...
Kaiwen Luo, Chunxi Luo, Liang Lin et al.· 0 citations
LLM-based agents for GPU kernel generation are advancing rapidly, but the benchmarks they optimize against evaluate kernels in isolation, with synthetic inputs and weak baselines, rewarding sandbox speedups that break or vanish in real inference systems. We introduce FastKernels, a benchmark of 384 tasks drawn from 47...
Gabriele Oliaro, Jaeseong Lee, Yichao Fu et al.· 0 citations
Training long-horizon LLM agents with reinforcement learning is challenging because sparse outcome rewards reveal whether a task succeeds, but not which intermediate actions caused the outcome or how they should be corrected. Recent methods alleviate this issue by generating rewards or textual hints from turn-level act...
Woongyeong Yeo, Yumin Choi, Taekyung Ki et al.· 0 citations
A new method, called CW-Net, translates the reasoning process of an autonomous vehicle’s AI system into understandable concepts that explain its behavior.
MIT News · Artificial Intelligence· news.mit.eduAug 31, 2026
With millions of users across the world, Julia has been used to conduct cutting-edge research and to design new drugs, jet engines, heat pumps, and more.