Skip to content

Category

small language model

2,746 papers

#artificial intelligence Open access Oct 2026

Works for 10/6/2026 - Sergei Pavlovich Korolev

Listen and ask questions on Gemini Notebook: https://notebook.google.com/notebook/31300077-8066-4826-bca1-faf9fe8dbd2e?authuser=1 The Little Twist in the Middle: History, Accessibility, Logos, and the Problem of Making the Other Reachable develops a unified account of how accumulated history changes what becomes access...

Ryan MacLean · 0 citations
#artificial intelligence Open access Oct 2026

Using Artificial Intelligence to estimate how common alcohol references and alcohol-related brand references are among influencers on Instagram

Background: Young people spend hours on social media every day, where following influencers is the norm. Influencers post frequently about alcohol and alcohol brands and there has been a rise in influencer marketing. However, prevalence estimates rely on manual analyses of only a small number of posts. We used a large...

Jack Delmenico, Dan Anderson‐Luxford, Emmanuel N. Kuntsche et al. · 0 citations
#small language model Preprint Oct 2026

COMPASS: Finding Where Reasoning Lives in Language Models

Comparing three model families and multiple math benchmarks, COMPASS outperforms the activation-steering baselines the authors compare against, improves GSM8K accuracy by 16 percentage points on average, and approaches CoT accuracy with 20-70\% fewer generated tokens.

Pratyay Dutta, Kowshik Thopalli, V. Narayanaswamy · 0 citations
#small language model Preprint Oct 2026

Agent in a Bottle: Can LLM Agents Turn Their Capabilities Into Cheap, Scalable Artifacts?

This work introduces BOTTLED, a benchmark in which agents receive an entire unlabelled workload and must complete it under fixed time, compute and LLM API budgets, and finds that strong zero-shot task performance does not reliably translate into strong bottling capabilities.

Ankit Sonthalia, Haritz Puerto, Alexander Rubinstein et al. · 0 citations
#small language model Review Oct 2026

Small Language Models for Smart Data Model Classification at the Edge: A Cost-Aware Hybrid Approach

This study evaluates the performance of lightweight open-source language models to resolve an input data entity against its corresponding best fitting SDM representation under resource-constrained conditions and provides significant and valuable insights into model selection, task formulation, and deployment strategies...

Cristian Martella, A. Martella, Antonella Longo et al. · 0 citations
#small language model Preprint Oct 2026

A Shape-Adaptive Architecture with Disaggregated Quantization for Efficient LLM Serving

DynaCore is presented, a unified architecture for efficient LLM serving via system-architecture co-design that substantially reduces service-level latency over quantization and reconfigurable accelerators, and proposes disaggregated quantization, applying dual-side quantization to prefill and weight-only quantization t...

Cong Guo, Chi-Yue Wei, Bo-Wen Duan et al. · 0 citations
#small language model Preprint Oct 2026

HarnessSecurity-Bench: Do Security Mechanisms Really Protect Coding Agent Harnesses?

This work presents HarnessSecurity, the first systematic empirical study and benchmark of open- and closed-source coding agent harnesses, and derives a ten-mechanism taxonomy and assesses 400 harness-mechanism cells using independent ratings by researchers and large language model judges.

Zheng-Yang Zhu, Li-Ming Huang, Run-Min Ji et al. · 0 citations
#small language model Preprint Oct 2026

Massive Activation Gating Channel in Large Language Models

This paper finds that the emergence of massive activations is controlled by a single channel in the input embedding to a spike feed-forward network (FFN), and names this channel the massive activation gating channel (MAGC).

Min-Jia Mao, Shi Chen, Bo-Wen Yin et al. · 0 citations
#small language model Preprint Oct 2026

Preparing an AI-Augmented SIEM for the EU Cyber Resilience Act: A Practitioner Case Study

This case study documents a CRA preparedness pilot for one such product, SEUXDR, an AI-augmented security monitoring product with a large-language-model active-response component on the open-source CYBERFORT platform, and offers practitioners a replicable starting point for translating CRA legal text into operational p...

Georgios Koutidis, Nikolaos Kekatos, Marina Korgiala-Karyda et al. · 0 citations
#small language model Preprint Oct 2026

SCSM: A Traffic-Native Foundation Model for Transferable Website Fingerprinting

SCSM constructs pairs of pretraining views from the same group of unlabeled traces through Segmentation, Combination, Scaling, and Masking, which produces diverse observable patterns while preserving the underlying packet events and local traffic dynamics of real trace fragments.

Xian-Wen Deng, Rui-Jie Zhao, Ming-Wei Zhan et al. · 0 citations

From tech blogs

See all →

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.