Skip to content

Category

small language model

2,824 papers

#machine learning Preprint Sep 2026

ScaffoldM3C: A Multimodal Sequential Monte Carlo Framework for Generative Stable Construction Planning

Autonomously constructing physically realizable 3D structures remains a significant challenge due to combinatorial action spaces, interchangeable components, equifinal assembly sequences, and strict stability requirements during construction. State-of-the-art methods fine-tune large language models for text-based gener...

Gadiel Sznaier Camps, Cheng-Yang He, G. Sartoretti et al. · 0 citations
#machine learning Preprint Oct 2026

TACO: Ternary Absolute-max Column-wise One-sparse Optimizer for LLM Fine-Tuning

Full-parameter fine-tuning of large language models (LLMs) incurs substantial optimizer state memory overhead, limiting the model sizes that fit on modern GPUs. Existing approaches either compress optimizer state, abandon first-order gradients, or change the update geometry while retaining dense state. The recently int...

Ji-Chao Jiang, Cristian McGee, El Houcine Bergou et al. · 0 citations
#machine learning Preprint Oct 2026

Every Ablation Is a Dose: Counterweights and the Semblance of Self-Repair

Ablate a component of a language model, and other components often appear to adjust and compensate. This phenomenon, termed self-repair, has been observed repeatedly, but its mechanism remains unclear. The most systematic study to date concluded that self-repair is noisy and unlikely to have a single explanation. We ar...

Areeb Ahmad, Pratinav Seth, Vinay Kumar Sankarapu · 0 citations
#machine learning Preprint Oct 2026

Stochastic Rounding in Low-Precision Transformer Inference: A Variable-Precision Emulation Study of a Small GPT-2

Should low-precision transformer inference use stochastic rounding (SR) or round-to-nearest (RN)? The answer depends on where in the network you look. We isolate this effect by holding the numerical format fixed and varying only the rounding rule at individual operation sites. To enable experiments at freely chosen pre...

Yohan Chatelain, Pablo de Oliveira Castro Krembil Centre for Neuroinformatics, Camh et al. · 0 citations
#machine learning Preprint Sep 2026

Analysis of Quantized and Efficiently Adapted Protein Language Models

Background: Protein language models (PLMs) are increasingly used for sequence generation and property prediction, but their size makes fine-tuning and deployment expensive. The effects of quantization and parameter efficient fine-tuning on performance, representations and generation remain insufficiently characterized....

I. Zeisler, Sebastian Clancy, P. Bayat et al. · 0 citations
#machine learning Preprint Sep 2026

MoRA: MoE Pruning via Router Bias Learning and Expert Approximation

Mixture-of-Experts (MoE) models enable parameter scaling with limited per-token computation by activating only a small subset of experts for each token, but deploying them still requires loading the complete expert pool into memory. Structured expert pruning can effectively reduce the memory usage by removing experts....

Yu-Shuai Sun, Zi-Kun Zhou, Lin Gao et al. · 0 citations
#artificial intelligence Preprint Oct 2026

Scalable, Transferable Meta-network for Data Selection Requires a Different Loss (and Why the Obvious Choice is Problematic)

Data selection is critical for training large language models on massive and heterogeneous corpora. Meta-learning for Training-data Selection offers a principled alternative to heuristic scoring by learning data weights from a target validation objective, but existing methods face a trade-off between fine-grained valua...

Zi-Lin Du, Bo-Wen Yang, Bo-Yang Li · 0 citations
#artificial intelligence Preprint Oct 2026

Selection-Based Structured Reasoning: Toward Efficient Multimodal Search Agents

Multimodal agents commonly generate free-form reasoning before each action. For small models, limited model capacity can result in lengthy reasoning that provides little useful guidance for action generation while incurring substantial inference cost. To address this challenge, we introduce Selection-based Structured R...

Feiyu Zhu, Xiao-Yu Zhu, Ji-Qian Yang et al. · 0 citations
#artificial intelligence Preprint Oct 2026

Managing Context and Communication in Distributed Agentic UAV Swarms

Unmanned aerial vehicle (UAV) swarms increasingly rely on language-model agents to provide adaptive mission-level reasoning in uncertain environments. Fully distributed control, in which each UAV hosts an independent Small Language Model (SLM), removes reliance on a centralized coordinator but introduces an information...

Andrea Iannoli, Ivan D. Zyrianoff, A. Trotta et al. · 0 citations
#artificial intelligence Preprint Oct 2026

It Takes Workflows to Evolve Better Workflows

Tackling complex real-world tasks can exceed the capabilities of a single large language model (LLM), motivating the use of multi-agent workflows that coordinate specialized agents to work together on these tasks. Recent methods train LLMs to construct better workflows from execution outcomes, but they optimize only th...

Xue-Hang Guo, Hao-Yu Wang, Hai-Feng Chen et al. · 0 citations
#artificial intelligence Preprint Oct 2026

Distilling Directional Verification

Knowledge distillation aims to transfer the factual knowledge of large language models to smaller models for efficient deployment. Yet a teacher may recall a relation in one direction while failing to generate the answer in the reverse direction. Distillation from its generated answers can therefore propagate this dire...

Jungseob Lee, Sugyeong Eo, Seongtae Hong et al. · 0 citations
#artificial intelligence Preprint Sep 2026

How Divergence Becomes Decision Flips in Compressed Language Models

Compression reports summarize how far a compressed language model moved from the dense one, usually by a KL divergence; a deployment that relies on the dense model's outputs needs to know how many of its decisions changed. We show that total variation, not KL, answers this directly. Across 802 compressed and perturbed...

Beatriz Almeida Felício · 0 citations

From tech blogs

See all →

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.