Understanding gender biases in large language models (LLMs) is increasingly important as these systems become embedded in decision-support tools with real consequences. Prior research has focused only on a small set of models, leaving open the extent to which gender biases are common and heterogeneous across LLMs. We a...
On-policy distillation (OPD) trains a student on its own generated responses using dense, token-level supervision from a stronger teacher. Vanilla OPD treats all teacher signals equally, assuming that the teacher's supervision is equally important for every token. However, teacher signals at different tokens may have v...
Zhen-Yu Wang, Tian-Ze Wang, Lin-Jun Zhang et al.· 0 citations
MIST, the Misleading-Image Stress Test: 200 English sentences, each built around a phrase readable either figuratively or literally and shown with an aligned image depicting its reading, a misleading image depicting the opposite, or no image at all is introduced.
Nagham Omar, Mahmoud Jabarin, Kinan Ibraheem et al.· 0 citations
This paper revisits the standard state-only formulation of value estimation and proposes PPO, a self-privileged actor-critic framework that consistently improves value-estimation quality by a substantial margin and outperforms representative actor-critic and critic-free RLVR baselines on challenging mathematical reason...
Kun Liang, Chenming Tang, Clive Bai et al.· 0 citations
Reach audiences
Advertise in front of researchers, engineers, and readers.
Jev answers MMLU's calculation-heavy questions more accurately than other MMLU questions (94% vs. 91%), whereas both open models, and all three on C-Eval, find them harder.
Tobias Deußer, L. Sparrenberg, R. Sifa· 4 citations
Speculative decoding accelerates autoregressive generation by using a smaller drafter to propose tokens for batched verification by a larger target. However, conventional speculative decoding couples drafting to the target's evolving verified prefix, serializing drafting and verification. We ask whether this dependency...
Yun-Zhe Li, Kyoungjun Park, Hong-Zi Zhu et al.· 0 citations
Test-time compute scaling has emerged as a cornerstone of advanced machine reasoning, yet performing iterative deliberation directly within continuous latent representation spaces reveals a catastrophic pathology: the Deliberation Drift Cliff. While unconstrained recurrent latent models achieve initial reasoning gains...
Frozen language models (LMs) are increasingly used as fixed feature extractors for downstream reranking, scoring, and preference modeling, raising a practical question: how should a compact module represent interactions among features in a fixed low-dimensional bottleneck? Common linear and low-rank adapters remain lin...
E. Roh, Hyojun Ahn, Hoyeong Lee et al.· 0 citations
Recent approaches to reinforcement learning (RL) post-training for large language models increasingly remove the critic to reduce training instability and memory overhead. Even where a critic is trained, it is discarded once training ends, although it has learned to predict outcomes. We revisit this trend and show that...
Hong-Yang Li, Xiao Li, Caesar Wu et al.· 0 citations
This work introduces \method{}, a framework for jailbreaking through text-only continuation interfaces that permit repeated sampling and assistant-prefix continuation, and achieves the highest mean score most comparisons against baselines.
Jesson Wang, Shawn Li, Wei Yang et al.· 0 citations
Fine-tuning-as-a-service lets users adapt a safety-aligned language model to their own data, but it also creates a harmful fine-tuning attack surface: a small amount of harmful data mixed into an otherwise benign fine-tuning set can degrade the model's alignment. Two recent alignment-stage defenses address this problem...
Muhammad Zeeshan Akram, Mufid Kamel Marican, Anvesh Reddy Yenugu et al.· 0 citations
Spotter is proposed, which reverses the roles: the embodied model leads and executes continuously, while the VLM runs in parallel, monitors through a lightweight local screener, intervenes only when an error is detected, reflects on and corrects it, and returns control.
Long Li, Qi-Chao Zhao, Yue Yang et al.· 0 citations