Mamba-style and hybrid language models compress their past into a fixed-size recurrent state that is rewritten at every generated token. Storing this state in low precision saves memory bandwidth, but every rounding error is fed back into the next update and can accumulate over long generations. Production systems roun...
Scaling Large Language Models (LLMs) via Mixture-of-Experts (MoE) enables massive parameter growth with nearly constant per-token computation. However, further scaling the parameter count requires increasingly sparse routing, where expert load imbalance becomes more severe. This imbalance reduces parameter utilization...
Peng Jin, Zi-Han Qiu, Ze-Kun Wang et al.· 0 citations
Radar hardware faults threaten automated perception, motivating accurate, compact diagnosis and understandable maintenance guidance. We introduce SCORE-LM, which couples a small scatterer-conditioned operator-response encoder (SCORE) to an adapted local language model. SCORE combines self-referenced complex trajectorie...
Language models are typically pretrained from random initialization. Recent work challenges this convention, showing that a brief warm-up on abstract, algorithmically generated data can provide a better starting point for subsequent learning of natural language. In this paper, we show that in small language models, suc...
Zachary Shinnick, Hemanth Saratchandran, Damien Teney et al.· 0 citations
Reach audiences
Advertise in front of researchers, engineers, and readers.
The scale and expense of pre-training language models make efficient hyperparameter tuning essential, yet a principled guidance is still missing. In this work, we analyze language model pre-training dynamics from a local landscape geometry perspective. Our study reveals two distinct phases. In Phase I, sharpness of the...
Zhan-Peng Zhou, Yu-Han Sun, Bing-Rui Li et al.· 1 citation
On-policy distillation (OPD) has demonstrated two important capabilities in language models: compressing large teachers into smaller students and merging expert models into a single model. Existing diffusion OPD, however, mostly focus on the latter, with teachers and students sharing the same backbone and scale. We inv...
Zhen-Xing Zhang, Jia-Yan Teng, Wen-Xu Wu et al.· 0 citations
As Large Language Model (LLM) agents are increasingly deployed in complex environments, multi-turn interaction attacks have become a significant security challenge. Existing detection methods typically rely on historical context. However, this retrospective logic struggles to identify deep malicious intents that are sp...
Ze-Zhong Wang, Xue-Yang Tang, Rui Lian et al.· 1 citation
Reinforcement learning with a verifiable reward (RLVR) offers a scalable approach to training language-model agents, yet sparse outcome rewards can leave early training with little signal for policy improvement. We identify an On-Policy Acceleration Phenomenon: in our main comparisons, RLVR initialized with on-policy d...
Yi-Tong Qiao, Tian-Tian He, Lei Liu et al.· 0 citations
Large language model (LLM)-based automatic heuristic design (AHD) iteratively proposes and refines heuristics, pairing design rationales with executable code. Task-specific evaluators assess programs; execution outcomes and performance scores guide search. Many AHD systems keep the generator frozen; EvoTune and Co-Evol...
Group Relative Policy Optimization (GRPO) has emerged as a memory-efficient reinforcement fine-tuning (RFT) technique for reasoning-intensive tasks. How- ever, GRPO training dynamics on small language models (SLMs) remain poorly understood, limiting its reliable adoption and reproducibility in open and resource- constr...
Rajat Ghosh, Vaishnavi Bhargava, Henry Wong et al.· 0 citations
Emotion recognition in conversation has been widely studied, but applying Large Language Models (LLMs) to continuous dimensional emotion evaluation in multimodal dialogue remains largely unexplored. We propose an LLM-based framework that performs discrete emotion recognition and Valence-Arousal-Dominance (VAD) dimensio...
Large language models rely heavily on human text, which often conveys surface answers rather than the spatial and structural logic behind them. Protein folding is a natural testbed, because one solved structure yields thousands of exactly checkable spatial and topological statements. We ask: can learning to fold protei...
Yong Liu, Zhan-Peng Shi, Yi-Zhou Dang et al.· 0 citations