Skip to content

Category

small language model

2,884 papers

#machine learning Preprint Sep 2026

Quality Determines Direction, Length Shapes Magnitude: Length Control for Open-Ended Reinforcement Learning

Quality-Gated Length Advantage Shaping (QGLAS), which first computes advantages from quality rewards alone, then adds bounded bonuses only to shorter positive-advantage responses, leaving all other advantages unchanged, consistently achieves a stronger quality--length trade-off than representative baselines.

Zi-Jun Weng, Zhong-An Bi, Xuan-Ang Gao et al. · 0 citations
#machine learning Preprint Sep 2026

The Low-Rank Structure of VLA Reinforcement Learning

Reinforcement learning (RL) is increasingly used to post-train vision-language-action (VLA) models, yet how RL reshapes these policies remains poorly understood. We find that RL across widely used flow-based VLA models, including $\pi_{0.5}$ and GR00T~N1.5/N1.6, on LIBERO, ManiSkill, MetaWorld, and CALVIN induces subst...

Minjae Oh, Yoonah Park, Jongwon Lim et al. · 0 citations
#machine learning Preprint Sep 2026

Routing Without Embeddings: Fast And Interpretable Routing With Regular Expressions

Large Language Model (LLM) routers commonly rely on neural query embeddings, with larger encoders expected to better capture query intent and difficulty. Yet scaling Qwen2.5 encoders from 0.5B to 72B parameters brings little improvement in routing accuracy (Figure 1b), suggesting that small encoders may already capture...

Yi-Fan Lu, Qi-Yue Zhang, Haotian Shan et al. · 0 citations
#machine learning Preprint Sep 2026

PQ-HSA: Reusing Product-Quantized Scores for Hybrid Sparse-Approximate Attention

PQ-HSA (hybrid sparse-approximate attention) attends the selected tokens with their original keys and values, and the unselected tokens, the background, enter the same softmax through those scores, summed per inverted list and multiplied by the list's mean value.

Kun-Ming Shao, Jie-Run Chen, Yan-Li Wang et al. · 0 citations
#machine learning Preprint Sep 2026

You Only Edit Once: Incentivizing In-Context Capability of LLMs via Local Demonstration Refinement

In-context learning (ICL) is crucial for boosting the inference performance of large language models (LLMs). However, the effectiveness of ICL in LLMs is greatly influenced by the choice of demonstration sets. Exhaustive searches over these sets are combinatorial, and existing selectors often rely on relevance or likel...

Jia-Rong Wen, Qi Wang, Yun Qu et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Large Language Models for Automated Cross-Domain Machine Learning Task Type Identification: A Benchmark Dataset and Evaluation

Machine learning task type identification is essential for constructing valid ML pipelines, yet in practice it is typically specified manually. We investigate whether large language models (LLMs) can infer both the data domain and the downstream prediction task directly from dataset-level information when only the targ...

Petros Tsialis, Steffen Limmer, Tobias Rodemann et al. · 0 citations
#artificial intelligence Preprint Sep 2026

A mechanistic study of language model introspection

Large language models (LLMs) can sometimes report perturbations to their internal activations---even when the input provides no evidence that an intervention occurred. How do models detect and localize such internal changes? We study this question using a controlled task that keeps the input text fixed. We either injec...

Jia-Hong Zou, Xiang-Kun Sun, Ling-Kai Kong et al. · 0 citations
#artificial intelligence Preprint Sep 2026

From Preference to Reciprocity: Decentralized Matching with Empirically Grounded LLM-agent Based Modeling

Bipartite matching is a fundamental problem in game theory and market design. Classical approaches such as Gale--Shapley assume complete preferences and centralized computation, whereas many real-world matching processes are decentralized, asynchronous, and shaped by sequential interaction under limited information. We...

Wang-Xuan Fan, Xiao-Yu Nie, Zhou-Tian Shi et al. · 0 citations
#artificial intelligence Preprint Sep 2026

ABC-Align: Prediction-Powered Alignment with Adaptive Bias Control

This work proposes ABC-Align, leveraging abundant pseudo label signal to minimize variance and applying a lightweight, adaptive correction grounded in the human-labeled subset, and empirically demonstrates that ABC-Align achieves superior performance over prior semi-supervised baselines in a series of experiments on an...

Eric Frankel, Bang-Hua Zhu, Sewoong Oh et al. · 0 citations

From tech blogs

See all →

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.