Skip to content

Category

small language model

2,764 papers

#machine learning Preprint Oct 2026

Static Bootstrap Placement for Encrypted Language Model Decoding

Language models increasingly serve prompts that carry private data, and secure inference under homomorphic encryption lets a client outsource the computation without revealing the prompt. Existing secure inference systems run a forward pass without consuming a token under encryption, and generating text with them requi...

Halil Ibrahim Kanpak, Didem Unat · 0 citations
#machine learning Preprint Oct 2026

Saying, Not Knowing: Aggressively GGUF-Quantized Small Language Models Still Write Rare Words They Can No Longer Define

Post-training quantization to the GGUF format's mixed-precision K-quants is commonly how open-weight language models reach consumer hardware, yet its effect on fine-grained lexical competence is uncharacterized. We audit 27 quantized artifacts across 13 families and four architecture backbones, 0.35B-14B parameters, ev...

Saurabh Kumar Singh, Yogeshwar Singh Dadwhal, Malhar Vedak · 0 citations
#machine learning Preprint Oct 2026

Towards Looped Models Done Right, Part II: Rethinking at Fixed Points

Every recurrence of a looped language model adds cost in training, decoding, prefill, and reinforcement learning (RL). The closer recurrent states get to fixed points, the less the path to them matters. This enables truncated backpropagation in training; terminal key-value (KV) sharing for decoding with almost no loss...

Benhao Huang, Chu-Fan Shi, Jun-Lin Chen et al. · 0 citations
#machine learning Preprint Oct 2026

Learning What to Imitate: Entropy-Aware Distribution Mixing

Small language models are often post-trained as students on reasoning traces from stronger teacher models to efficiently learn new skills. However, token-level imitation on traces that lie far outside the student's expected distribution often produces \textit{confident conflicts}, whereby the student is required to imi...

Juan Garcia Giraldo, Matteo Santelmo, E. Durech et al. · 0 citations
#machine learning Preprint Oct 2026

Separators Make Carry Propagation Learnable:The Geometry of Latent Carry in a Multiplication Transformer

Transformers asked to multiply multi-digit numbers in a single forward pass often fail, and interpretability studies of pretrained language models find arithmetic solved by input-range heuristics rather than by an explicit carry. We train small Llama-style transformers from scratch on 4x4 multiplication without chain o...

Sama Satariyan, Raphael Cousin, Gérard Biau · 0 citations
#machine learning Preprint Oct 2026

A Fine-Grained Analysis of the LoRA Fine-Tuning Landscape with Implications for Data Selection

Low-Rank Adaptation (LoRA) has become a standard approach for parameter-efficient fine-tuning, yet a fundamental practical question remains unresolved: how should the adapter rank be chosen? An overly small rank may lead to a poorly conditioned optimization landscape, whereas an unnecessarily large rank sacrifices the...

Bo-Wen Zhang, Chang-Rui Fang, Xin-Song Ma et al. · 0 citations
#machine learning Preprint Oct 2026

Robust Parameter-Efficient LLM Adaptation on Analog Hardware

Analog in-memory computing is a promising platform for on-device execution of large language models because it performs matrix--vector multiplications (MVMs) in memory and in parallel, reducing data movement. However, limited digital-to-analog converter precision, input noise, and finite conductance states can degrade...

Jin-Dan Li, Zhao-Xian Wu, Tian-Yi Chen · 0 citations
#machine learning Preprint Oct 2026

How Should Teachers Be Prepared? RL on Student-Induced States for On-Policy Distillation

On-policy distillation (OPD) improves the reasoning capabilities of small language models through token-level teacher supervision on student-generated trajectories. Yet can teachers that excel at solving problems independently also guide student reasoning effectively? Prior work shows that when student prefixes follow...

Xiao-Yu Ma, Hao-Yue Liu, Zhi-Chao Wang et al. · 0 citations
#artificial intelligence Preprint Oct 2026

Choosing an energy-efficient software architecture for building system diagnostic support

Around 30\% of global energy expenditure can be attributed to the building sector, where a large portion of energy-consumption could be avoided by repairing existing faults. Fault detection and diagnosis (FDD) software addresses this issue; however, its creation and operation also have an environmental impact. The magn...

Roxane Koitz-Hristov, Franz Wotawa · 0 citations
#artificial intelligence Review Oct 2026

Scaling Down the Scaling Laws: Parameter Efficiency and Compute-Optimal Training in Resource-Constrained Large Language Models

Large language models (LLMs) have achieved substantial performance gains through increases in model size, training data, and computational resources. However, traditional scaling approaches produce diminishing returns, rising financial and environmental costs, and barriers to participation for researchers operating out...

J. Dwyer · 0 citations
#artificial intelligence Preprint Oct 2026

Backdooring Sparse Autoencoders

Sparse autoencoders (SAEs) are increasingly used not only to interpret language models but also to intervene on their internal representations. We show that this creates a supply-chain attack surface: a maliciously modified SAE can induce attacker-chosen behavior when inserted into the forward pass of an otherwise unch...

E. Ahlers, Daniel Passon, Tobias Kiecker et al. · 0 citations
#artificial intelligence Preprint Oct 2026

Nexus: An Execution Fabric for AI Agents Across Cloud, Edge, and Devices

Language-model agents are evolving into long-running services that interact with models, tools, computers, mobile devices, and distributed environments. Existing agent frameworks simplify reasoning and tool invocation. However, cloud-centric designs face three limitations: centralized execution increases failure impact...

C. Chang, Jia-Lin Zhou · 0 citations

From tech blogs

See all →

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.