Frozen language models (LMs) are increasingly used as fixed feature extractors for downstream reranking, scoring, and preference modeling, raising a practical question: how should a compact module represent interactions among features in a fixed low-dimensional bottleneck? Common linear and low-rank adapters remain lin...
E. Roh, Hyojun Ahn, Hoyeong Lee et al.· 0 citations
Recent approaches to reinforcement learning (RL) post-training for large language models increasingly remove the critic to reduce training instability and memory overhead. Even where a critic is trained, it is discarded once training ends, although it has learned to predict outcomes. We revisit this trend and show that...
Hong-Yang Li, Xiao Li, Caesar Wu et al.· 0 citations
This work introduces \method{}, a framework for jailbreaking through text-only continuation interfaces that permit repeated sampling and assistant-prefix continuation, and achieves the highest mean score most comparisons against baselines.
Jesson Wang, Shawn Li, Wei Yang et al.· 0 citations
Fine-tuning-as-a-service lets users adapt a safety-aligned language model to their own data, but it also creates a harmful fine-tuning attack surface: a small amount of harmful data mixed into an otherwise benign fine-tuning set can degrade the model's alignment. Two recent alignment-stage defenses address this problem...
Muhammad Zeeshan Akram, Mufid Kamel Marican, Anvesh Reddy Yenugu et al.· 0 citations
Reach audiences
Advertise in front of researchers, engineers, and readers.
Spotter is proposed, which reverses the roles: the embodied model leads and executes continuously, while the VLM runs in parallel, monitors through a lightweight local screener, intervenes only when an error is detected, reflects on and corrects it, and returns control.
Long Li, Qi-Chao Zhao, Yue Yang et al.· 0 citations
A common assumption in language model development is that cognitive abilities are organized around a general, domain-free intelligence factor, like fluid intelligence in humans. This assumption is rarely tested directly, and prior attempts have done so only at a much smaller scale. We take a latent variable approach to...
Cross-modal associations are systematic pairings of features across modalities, such as the association of'bouba'with round shapes and'kiki'with sharp shapes. Prior work has compared humans and vision-language models (VLMs) on such associations, but often using different stimuli or tasks between humans and models. Here...
Su-Min Hong, Katsumi Ibaraki, Renee Shi et al.· 0 citations
Decision-only language models return a probability for every answer option instead of generating text, which makes them attractive as survey respondents and as judges. We audit the cultural values of one such model, TypeSafe's JEV, with the Values Survey Module 2013. We asked it the 24 items as 12 matched Saudi and 12...
A large-scale benchmark suite for open-world aerial object-goal search, with 3 times as many scenes and 18.7 times as many task instances as the largest existing benchmark for this task, and a unified evaluation framework with a unified evaluation framework.
Tong-Tong Feng, Xin Wang, Hao-Ran Hou et al.· 0 citations
A small trainable advisor can steer a frozen language-model executor using natural-language advice. In addition to learning from task rewards, the advisor can use feedback from completed interactions to improve its advice. However, a plausible correction need not change execution, yet learning from such corrections can...
Rishabh Agrawal, He-Jie Cui, Sha-Sha Li et al.· 0 citations
Running a language model locally offers advantages in privacy, latency, and cost, but local hardware fits only small models, which are less capable than frontier models. The usual remedy for a hard query, escalating it to a cloud model, gives up the privacy and cost advantages of running locally. A deployment that stay...
Kenan Alkiek, Moontae Lee, David Jurgens et al.· 0 citations
Sparse mixture-of-experts (MoE) large language models scale model capacity by routing each token to a small subset of experts. Their routers are regularized with load balancing terms and learn affinity scores through the language-model objective. However, these objectives do not provide direct alignment between routing...
Yury Nahshan, Nati Daniel, Jacob Goldberger et al.· 0 citations