Skip to content

Category

small language model

2,807 papers

#natural language process... Preprint Oct 2026

Towards Unbiased On-Policy Distillation for Block Diffusion Language Models

On-policy distillation (OPD) has emerged as an effective post-training paradigm for language models, with recent efforts extending it to block diffusion language models (BDLMs). However, existing studies focus almost exclusively on small block sizes, leaving distillation into student models with larger blocks underexpl...

Zai-Quan Yang, Fei Wei, Yong Wang et al. · 0 citations
#natural language process... Preprint Oct 2026

When Evidence Changes: Evaluating Memory Repair and Re-reading in Language-Model Agents

When documents supporting an agent's derived facts are revoked or replaced, should it repair memory or re-read current evidence? We introduce an evidence-revision evaluation on medication- and problem-list tasks from public ICU records. Under revocation, replacement and control events, we compare full and source-filter...

Wen-Hui Chu · 0 citations
#machine learning Preprint Oct 2026

Ontology Concept Overlap as a Training Signal: Knowledge-Grounded Reinforcement Learning for Clinical Question Answering

Reinforcement learning post-training for language models relies on two reward designs: human preferences (RLHF, DPO) and binary verifiers (RLVR). Clinical question answering fits neither. Near-correct answers differ by a single substituted entity, and no executable check decides clinical correctness. We instantiate a s...

Aditya Tanna, Abhishek Jindal · 0 citations
#machine learning Preprint Oct 2026

Readout Blindness: VLM Scores Miss the Spatial Direction Their Frozen Encoders Retain

CLIP-like vision-language models remain a cornerstone of multimodal systems, yet their scores stay near chance on directed spatial relations, such as whether one object is left of another. We call this failure readout blindness and analyze, theoretically and empirically, why deployed scores miss the direction: when sco...

Guang-Yuan Li, Tian-Ming Du, Yan Jiang et al. · 0 citations
#machine learning Preprint Oct 2026

Fitting Vision Adapters at Frontier Scales

Training a small projector between a frozen vision encoder and language model is an established approach to multimodal learning. As the parameter count of language models scales dramatically, we revisit which vision capabilities this approach can add while keeping their pretrained weights fixed. Here we train a 50M par...

Jaehoon Lee, Harry B. Partridge, M. Jayasekara et al. · 0 citations
#machine learning Preprint Oct 2026

FORGE: Verification-Gated Behavioral Repair for Generative Language Models

Generative large language models (LLMs) inherit undesirable behaviors from pre-training, including demographic bias and toxic generation, that often emerge only after deployment and affect a small subset of inputs. A repair should eliminate the identified defect, preserve the model's overall functionality and, ideally,...

Hsin-Ling Hsu, Min-Yue Chen, Nai-Chia Chen et al. · 0 citations
#machine learning Preprint Oct 2026

Static Bootstrap Placement for Encrypted Language Model Decoding

Language models increasingly serve prompts that carry private data, and secure inference under homomorphic encryption lets a client outsource the computation without revealing the prompt. Existing secure inference systems run a forward pass without consuming a token under encryption, and generating text with them requi...

Halil Ibrahim Kanpak, Didem Unat · 0 citations
#machine learning Preprint Oct 2026

Saying, Not Knowing: Aggressively GGUF-Quantized Small Language Models Still Write Rare Words They Can No Longer Define

Post-training quantization to the GGUF format's mixed-precision K-quants is commonly how open-weight language models reach consumer hardware, yet its effect on fine-grained lexical competence is uncharacterized. We audit 27 quantized artifacts across 13 families and four architecture backbones, 0.35B-14B parameters, ev...

Saurabh Kumar Singh, Yogeshwar Singh Dadwhal, Malhar Vedak · 0 citations
#machine learning Preprint Oct 2026

Towards Looped Models Done Right, Part II: Rethinking at Fixed Points

Every recurrence of a looped language model adds cost in training, decoding, prefill, and reinforcement learning (RL). The closer recurrent states get to fixed points, the less the path to them matters. This enables truncated backpropagation in training; terminal key-value (KV) sharing for decoding with almost no loss...

Benhao Huang, Chu-Fan Shi, Jun-Lin Chen et al. · 0 citations
#machine learning Preprint Oct 2026

Learning What to Imitate: Entropy-Aware Distribution Mixing

Small language models are often post-trained as students on reasoning traces from stronger teacher models to efficiently learn new skills. However, token-level imitation on traces that lie far outside the student's expected distribution often produces \textit{confident conflicts}, whereby the student is required to imi...

Juan Garcia Giraldo, Matteo Santelmo, E. Durech et al. · 0 citations
#machine learning Preprint Oct 2026

Separators Make Carry Propagation Learnable:The Geometry of Latent Carry in a Multiplication Transformer

Transformers asked to multiply multi-digit numbers in a single forward pass often fail, and interpretability studies of pretrained language models find arithmetic solved by input-range heuristics rather than by an explicit carry. We train small Llama-style transformers from scratch on 4x4 multiplication without chain o...

Sama Satariyan, Raphael Cousin, Gérard Biau · 0 citations
#machine learning Preprint Oct 2026

A Fine-Grained Analysis of the LoRA Fine-Tuning Landscape with Implications for Data Selection

Low-Rank Adaptation (LoRA) has become a standard approach for parameter-efficient fine-tuning, yet a fundamental practical question remains unresolved: how should the adapter rank be chosen? An overly small rank may lead to a poorly conditioned optimization landscape, whereas an unnecessarily large rank sacrifices the...

Bo-Wen Zhang, Chang-Rui Fang, Xin-Song Ma et al. · 0 citations

From tech blogs

See all →

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.