Skip to content

Category

small language model

2,885 papers

#machine learning Preprint Sep 2026

Nonparametric In-Context Learning under Growing Geometric Complexity: Minimax Optimality and Local Geometry-Adaptivity of Transformers

These results identify conditions under which the resulting predictor exploits local geometry and attains the aggregate minimax rate, and derive an in-context generalization bound for near empirical risk minimizers over this class.

J. Seo, Jisu Kim · 0 citations
#machine learning Preprint Sep 2026

Benchmarking Attention for Tabular Foundation Models

Tabular in-context learners such as TabPFN, Mitra, or ConTextTab rely on alternating row and column attention over 2D sequences of latent embeddings. These attention patterns differ markedly from the one-dimensional case in language models: row attention involves longer sequences while column attention operates on much...

Maximilian Schambach, Clemens Biehl, Sam Thelin · 0 citations
#machine learning Conference Open access Sep 2026

Audio emotion recognition for atypical hearing

This research focuses on auditory hypersensitivity in people with autism, a phenomenon that is often difficult to evaluate and unique to each individual.

Ulysse Roussel · 0 citations
#machine learning Preprint Sep 2026

Reinforcement Learning of Communication in a Mesh of Small Language Models

TalkMesh, a decentralized mesh of small language model agents that learns when and what to communicate is presented, a decentralized mesh of small language model agents that reaches the accuracy of majority voting over 32 samples with each of three models.

Mehmet Kerem Türkcan · 0 citations
#machine learning Review Sep 2026

Fake News Theories: Harnessing Disciplinary Insights for Computational Modeling, Detection, and Explanation

Disinformation research has produced increasingly accurate automated fake-news detectors, but many systems remain difficult to interpret and are weakly connected to established theories of persuasion, credibility, and human judgment. In this paper, we develop a theory-informed computational framework that translates cr...

Zhao-Yang Cao, Miriam J. Metzger, Reza Zafarani · 0 citations
#machine learning Preprint Sep 2026

Adaptive Multi-Value Control in LLMs via Causal Activation Steering

Large language models (LLMs) are increasingly deployed in settings where responses must reflect multiple, potentially interacting social norms and human values. Activation steering offers a lightweight alternative to training-based alignment by modifying internal activations at inference time. However, prior human-valu...

Payel Bhattacharjee, Ravi Tandon · 0 citations
#artificial intelligence Preprint Sep 2026

Can You Check That? The Checkability Boundary for Local LLM Network Automation

This work introduces checkability as a criterion for determining which tasks are suitable for local inference, in Touchstone, a local-first pipeline that uses seven off-the-shelf SLMs (1-8B parameters) to generate candidates, uses task-specific intrinsic checks to reject responses, and escalates unresolved inputs to a...

Maleeha Masood, Momina Nofal · 0 citations
#artificial intelligence Preprint Sep 2026

Intent2Tc: Automated Intent-to-Traffic Control Translation with Language Models

Intent2Tc is presented, a closed-loop language-model-driven framework that translates business-level traffic-shaping intents into declarative sub-intents and subsequently into validated, executable Linux traffic control (tc) configurations and demonstrates the practical applicability of the proposed framework.

Andrea Masini, Sudipta Acharya, P. Bellavista et al. · 1 citation
#artificial intelligence Preprint Sep 2026

Softmax Reparameterization for Output-Head Quantization

Large vocabularies make output heads a substantial inference cost in small language models. We introduce softmax reparameterization, a post-training method that searches over functionally equivalent output heads before quantization. The method subtracts a scalar multiple of the vocabulary-row mean from every output row...

Asim Kadav, Christian Flores, C. Arora et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Beyond the Last Truffula Tree: SustainAI - A Water-Aware, Closed-Loop Framework for Environmentally Accountable AI

SustainAI provides a practical foundation for integrating ethical care and environmental responsibility into AI infrastructure design and lifecycle management, framing AI sustainability around relational ethics, regional equity, and ecological stewardship.

Farnaz Farid, Tashfia Towkee, S. Nasreen et al. · 0 citations

From tech blogs

See all →

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.