A common assumption in language model development is that cognitive abilities are organized around a general, domain-free intelligence factor, like fluid intelligence in humans. This assumption is rarely tested directly, and prior attempts have done so only at a much smaller scale. We take a latent variable approach to...
Cross-modal associations are systematic pairings of features across modalities, such as the association of'bouba'with round shapes and'kiki'with sharp shapes. Prior work has compared humans and vision-language models (VLMs) on such associations, but often using different stimuli or tasks between humans and models. Here...
Su-Min Hong, Katsumi Ibaraki, Renee Shi et al.· 0 citations
Decision-only language models return a probability for every answer option instead of generating text, which makes them attractive as survey respondents and as judges. We audit the cultural values of one such model, TypeSafe's JEV, with the Values Survey Module 2013. We asked it the 24 items as 12 matched Saudi and 12...
A large-scale benchmark suite for open-world aerial object-goal search, with 3 times as many scenes and 18.7 times as many task instances as the largest existing benchmark for this task, and a unified evaluation framework with a unified evaluation framework.
Tong-Tong Feng, Xin Wang, Hao-Ran Hou et al.· 0 citations
Reach audiences
Advertise in front of researchers, engineers, and readers.
A small trainable advisor can steer a frozen language-model executor using natural-language advice. In addition to learning from task rewards, the advisor can use feedback from completed interactions to improve its advice. However, a plausible correction need not change execution, yet learning from such corrections can...
Rishabh Agrawal, He-Jie Cui, Sha-Sha Li et al.· 0 citations
Running a language model locally offers advantages in privacy, latency, and cost, but local hardware fits only small models, which are less capable than frontier models. The usual remedy for a hard query, escalating it to a cloud model, gives up the privacy and cost advantages of running locally. A deployment that stay...
Kenan Alkiek, Moontae Lee, David Jurgens et al.· 0 citations
Sparse mixture-of-experts (MoE) large language models scale model capacity by routing each token to a small subset of experts. Their routers are regularized with load balancing terms and learn affinity scores through the language-model objective. However, these objectives do not provide direct alignment between routing...
Yury Nahshan, Nati Daniel, Jacob Goldberger et al.· 0 citations
Despite the recent success of coding agents built on large language models, it remains challenging to run them over long horizons, since every observation is appended to the context and the context grows with each one. History-based maintenance is a common remedy, which masks or summarizes old observations, or prunes w...
Zi-Yang Yu, Liang Zhao, Bo-Wen Zhu et al.· 0 citations
Long-context sequence models face a fundamental tradeoff: softmax attention uses flexible token-level interactions at quadratic cost, whereas linear attention obtains linear-time training and constant-time decoding by compressing history into a fixed-size state. In this work, we ask whether we can connect these regimes...
E. Anand, Abdullah Ateyeh, Archer Wang et al.· 2 citations
This work proposes a large-scale multimodal large language model-based supervision generation framework that improves supervision diversity, fidelity, and semantic coverage, and introduces a granularity-aware text representation with separate CLS tokens for summary and detailed views.
F. M. Thoker, Renaud Vandeghen, Karen Sanchez et al.· 0 citations
Backdoor as Probe is proposed, a test-time adversarial defense for CLIP that improves average robust accuracy from 1.0\% to 52.3\% while retaining clean accuracy, achieving performance comparable to state-of-the-art methods with up to a \(5.7\times\) inference speedup.
Zhong Ling Wang, Jie Zhang, Sen Nie et al.· 0 citations
A novel end-to-end framework integrating a self-attention mechanism to address limitations in effectively detecting small animals in low-contrast trap images and small animals while also demonstrating zero-shot detection capability leveraging the MLLM.
Nowshin Amin, Nafisa Tabassum Oyshi, Tahmid Abrar Zidan et al.· 0 citations