Skip to content

Category

small language model

2,863 papers

#small language model Preprint Sep 2026

Selective Elicitation as a Commercial Influence Channel: A Reproducible Synthetic Shopping-Agent Stress Test

This work contrasts a neutral question, a soft commercial instruction, and an explicitly adversarial instruction to ask about the sponsor's advantage while omitting the rival's advantage to establish neither typical behavior under advertising incentives nor effects on actual consumers.

Jia-Peng Li · 0 citations
#natural language process... Preprint Sep 2026

Video2Skill: From Streaming Experience to Reusable Embodied Skills

This work introduces Video2Skill, a benchmark that covers robot tabletop manipulation and human kitchen activity and tests three core capabilities: locating manipulation events in time, grouping events of the same transformation, and deciding when to reuse an existing skill or create a new one.

Jian-Shu Zhang, Ce Zhang, Xi-Yuan Yang et al. · 0 citations
#natural language process... Preprint Sep 2026

When Updating Stops Being Learning: Rethinking LLM Self-Evolution via learnable information gain

A holistic framework based on learnable information gain, which measures how much novel, parameterizable information a round provides relative to the previous round, and proposes ATRI (Adaptive Training Regulation via Information-gain), which reweights samples within a round and halts training across rounds when inform...

Chen-Xu Wang, Chao-Zhuo Li, Xin-Ze Shi et al. · 0 citations
#small language model Preprint Sep 2026

The Editor Has Read-Only Access: Correctness Signals in Diffusion Language Models

Across six diffusion models, linear probes distinguish passing from failing attempts, with the strongest reads generally appearing beyond the early layers, and probe point estimates offer no consistent advantage in comparisons with model confidence.

A. Miglani, Samrath Singh Chadha, Kevin Li et al. · 0 citations
#machine learning Preprint Sep 2026

Fine-Tuning on Self-Generated and Reward-Weighted Data: Learning Dynamics, Convergence Rates, and Benefits of Off-Policyness

A unified theory for RE(S) is developed that covers the full spectrum of S, and can be interpreted as a stage-wise optimization process, where each stage takes $S$ gradient steps for minimizing the Kullback-Leibler distance to a fixed reward-weighted rollout distribution.

Zhi-Wei Wang, Yan-Xi Chen, Ya-Liang Li et al. · 0 citations
#natural language process... Preprint Sep 2026

Distilling What Matters: Confidence-Aware Selective Distillation for Large Language Models

CaRE-KD is proposed, a confidence-gated distillation framework that replaces static objectives with uncertainty-adaptive optimization and provides a gradient-level analysis showing how this dual-granularity design induces a conditional calibration mechanism that prior static divergences cannot reproduce.

Ayan Sengupta, Vaibhav Seth, Tanmoy Chakraborty · 0 citations
#machine learning Preprint Sep 2026

Invariant Atoms: Sparse Coordinates of Local Semantic Geometry in Language Model Representations

This work learns a shared semantic frame and sparse coordinates that reconstruct semantic displacements while suppressing nuisance variation, with anchor-dependent diagonal modulation adjusting atom strengths without sample-specific rotations to support reusable invariant directions as a sparse coordinate system for lo...

Muhammad Ahtesham, Xin Zhong · 0 citations
#machine learning Preprint Sep 2026

EasyPPO: Stabilizing the Critic Is Key

A key strength of Proximal Policy Optimization (PPO) is its learned critic, which uses historical trajectories collected during reinforcement learning to estimate expected returns and reduce policy-gradient variance. However, we find that the critic is also a major source of instability in reinforcement learning for la...

Xuan-Yi Zhou, Qiu-Yang Mang, Huan-Zhi Mao et al. · 0 citations
#machine learning Preprint Sep 2026

CorrGRPO: Correlation-Normalized GRPO for Multi-Reward Learning

Group Relative Policy Optimization (GRPO) is widely used to train reasoning language models, where it computes advantages by centering and normalizing rewards across rollouts of the same prompt. For multiple rewards, GRPO sums the reward components and normalizes the total reward by its within-group standard deviation....

Wen-Bin Hu, Hui-Hao Jing, Hao-Chen Shi et al. · 0 citations
#small language model Open access Sep 2026

Effect of an Electronic Health Record-Integrated Clinical Dashboard for Radiologists on STAT Priority Chest Radiograph Reporting and Downstream Care: A Stepped-Wedge Cluster Randomized Trial

Background. Chest radiographs (CXRs) are the most common imaging exam performed, but CXR reports can be nonspecific due to incomplete history, relying on vague terms like opacities to convey diagnostic uncertainty. We prospectively evaluated whether introducing a clinical dashboard with relevant patient information aff...

J. Balkman, S. Chimmula, Alysha Lam et al. · 0 citations

From tech blogs

See all →

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.