The increasing use of automated translation quality estimation (QE) systems calls for practical, decision-oriented methods for evaluating their performance. We propose that Receiver Operating Characteristic (ROC) analysis is a useful approach for this purpose. Our study shows that ROC analysis not only produces results...
Evelyn Y. Garland (Acta Language Services, LLC), Carola F. Berger (CFB Scientific Translations LLC)· 0 citations
We preregistered a comparison of two ways to help an LLM answer questions over a small research corpus: single-round Vector RAG and an LLM-compiled markdown wiki browsed by a tool-using agent. Both answered the same 13 questions over 24 papers with the same answer model, scored by two blinded LLM judges. The three prer...
Large language model (LLM) unlearning aims to remove specific data influences from pre-trained model without costly retraining, addressing privacy, copyright, and safety concerns. However, recent studies reveal a critical vulnerability: unlearned models rapidly recover "forgotten" knowledge through relearning attacks....
Zeguan Xiao, Xuanzhe Xu, Yong Wang et al.· 0 citations
We study whether a decoder-only language model requires an independently trainable input vector for every token. For a vocabulary of size $V$, an injective fixed-length binary identifier requires $K=\lceil\log_2 V\rceil$ bits. We replace the usual trainable $V\times d_{\mathrm{model}}$ input table with fixed minimal bi...
A. Bochkov· 0 citations
Reach audiences
Advertise in front of researchers, engineers, and readers.
Large language models are increasingly used in K-12 education, but existing benchmarks mainly test exam question answering rather than understanding how curriculum knowledge is structured and visually presented. We call this capability curriculum cognition. It covers prerequisite chains, concept taxonomies, experiment-...
Hao Liang, Qihan Lin, Mingrui Chen et al.· 0 citations
Hardware infrastructure is a critical bottleneck for LLM-driven circuit design, limiting what agents can express, compile, and iteratively refine within an agentic loop. To address this bottleneck, we introduce CKTLEAN, a typed hardware infrastructure embedded in Lean. It supports hardware description, compilation to S...
Jing Xiong, Qi Han, Chenchen Ding et al.· 0 citations
Distillation and reinforcement learning through verifiable rewards (RLVR) have achieved progress in enhancing the reasoning ability of large language models (LLMs). However, we note that negative rollouts may admit no gradation of failure severity, and the combinatorial vastness makes penalizing a few sampled negatives...
LLMs have so far failed both to generate consistently compelling stories and to recognize this failure--on the leading creative-writing benchmark (EQ-Bench), LLM judges rank zero-shot AI stories above New Yorker short stories, a gold standard for literary fiction. We argue that existing rubrics overlook a key dimension...
Peiqi Sui, Yutong Zhu, Tianyi Cheng et al.· 0 citations
GraphRAG is increasingly adopted for converting unstructured corpora into graph structures to enable multi-hop reasoning. However, standard graph algorithms rely heavily on static connectivity and explicit edges, often failing in real-world scenarios where Knowledge Graphs (KGs) are noisy, sparse, or incomplete. To add...
The increasing adoption of Large Language Models (LLMs) has enabled AI scientists to perform complex end-to-end scientific discovery tasks requiring coordination of specialized roles, including idea generation and experimental execution. However, most state-of-the-art AI scientist systems rely on static, hand-designed...
Yougang Lyu, Xi Zhang, Yuyue Zhao et al.· 0 citations
We argue that uncertainty is a key and understudied limitation of LLMs' performance in creative writing, which is often characterized as trite and clich\'e-ridden. Literary theory identifies uncertainty as a necessary condition for creative expression, while current alignment strategies steer models away from uncertain...
Large language models (LLMs) are increasingly deployed worldwide, yet their safety alignment remains predominantly English-centric. This allows for vulnerabilities in non-English contexts, especially with low-resource languages. We introduce a novel application of knowledge distillation (KD) in the context of multiling...
Max Zhang, Derek Liu, Kai Zhang et al.· 0 citations
A new method, called CW-Net, translates the reasoning process of an autonomous vehicle’s AI system into understandable concepts that explain its behavior.
MIT News · Artificial Intelligence· news.mit.eduAug 31, 2026
With millions of users across the world, Julia has been used to conduct cutting-edge research and to design new drugs, jet engines, heat pumps, and more.