Skip to content

Category

small language model

2,883 papers

#artificial intelligence Preprint Sep 2026

Large-scale factor analysis shows machine intelligence is only partially interpretable

A common assumption in language model development is that cognitive abilities are organized around a general, domain-free intelligence factor, like fluid intelligence in humans. This assumption is rarely tested directly, and prior attempts have done so only at a much smaller scale. We take a latent variable approach to...

Faiz Ghifari Haznitrama, Afrizal Hasbi Azizy, Faeyza Rishad Ardi · 0 citations
#artificial intelligence Preprint Sep 2026

Similar Choices, Different Attention: Cross-Modal Associations in Humans and Vision-Language Models

Cross-modal associations are systematic pairings of features across modalities, such as the association of'bouba'with round shapes and'kiki'with sharp shapes. Prior work has compared humans and vision-language models (VLMs) on such associations, but often using different stimuli or tasks between humans and models. Here...

Su-Min Hong, Katsumi Ibaraki, Renee Shi et al. · 0 citations
#artificial intelligence Review Sep 2026

Calibrated to Whom? Persona and Language Effects on Cultural Values in JEV

Decision-only language models return a probability for every answer option instead of generating text, which makes them attractive as survey respondents and as judges. We audit the cultural values of one such model, TypeSafe's JEV, with the Values Survey Module 2013. We asked it the 24 items as 12 matched Saudi and 12...

Bushra Asseri, Abdulaziz M. Asseri · 0 citations
#artificial intelligence Preprint Sep 2026

AerialDojo-200K: A Large-Scale Benchmark Suite for Open-World Aerial Object-Goal Search

A large-scale benchmark suite for open-world aerial object-goal search, with 3 times as many scenes and 18.7 times as many task instances as the largest existing benchmark for this task, and a unified evaluation framework with a unified evaluation framework.

Tong-Tong Feng, Xin Wang, Hao-Ran Hou et al. · 0 citations
#artificial intelligence Preprint Sep 2026

AdviSD: Learning to Advise Frontier LLMs via Targeted Multi-Turn Self-Distillation

A small trainable advisor can steer a frozen language-model executor using natural-language advice. In addition to learning from task rewards, the advisor can use feedback from completed interactions to improve its advice. However, a plausible correction need not change execution, yet learning from such corrections can...

Rishabh Agrawal, He-Jie Cui, Sha-Sha Li et al. · 0 citations
#artificial intelligence Preprint Sep 2026

HARISSA: Inference-Time Self-Checks for Efficient and Safe Local Language Model Deployment

Running a language model locally offers advantages in privacy, latency, and cost, but local hardware fits only small models, which are less capable than frontier models. The usual remedy for a hard query, escalating it to a cloud model, gives up the privacy and cost advantages of running locally. A deployment that stay...

Kenan Alkiek, Moontae Lee, David Jurgens et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Cross-Entropy Guided Routing in Mixture-of-Experts Large Language Models

Sparse mixture-of-experts (MoE) large language models scale model capacity by routing each token to a small subset of experts. Their routers are regularized with load balancing terms and learn affinity scores through the language-model objective. However, these objectives do not provide direct alignment between routing...

Yury Nahshan, Nati Daniel, Jacob Goldberger et al. · 0 citations
#artificial intelligence Preprint Sep 2026

StateTape: Action-Conditioned Evidence Lifecycle Modeling for Long-Horizon Coding Agents

Despite the recent success of coding agents built on large language models, it remains challenging to run them over long horizons, since every observation is appended to the context and the context grows with each one. History-based maintenance is a common remedy, which masks or summarizes old observations, or prunes w...

Zi-Yang Yu, Liang Zhao, Bo-Wen Zhu et al. · 0 citations
#artificial intelligence Preprint Sep 2026

SMat-Attention: Structured Long-Context Sequence Modeling

Long-context sequence models face a fundamental tradeoff: softmax attention uses flexible token-level interactions at quadratic cost, whereas linear attention obtains linear-time training and constant-time decoding by compressing history into a fixed-size state. In this work, we ask whether we can connect these regimes...

E. Anand, Abdullah Ateyeh, Archer Wang et al. · 2 citations
#small language model Preprint Sep 2026

Advancing Video-Text Pretraining with Multi-View Captions

This work proposes a large-scale multimodal large language model-based supervision generation framework that improves supervision diversity, fidelity, and semantic coverage, and introduces a granularity-aware text representation with separate CLS tokens for summary and detailed views.

F. M. Thoker, Renaud Vandeghen, Karen Sanchez et al. · 0 citations
#small language model Preprint Sep 2026

Backdoor as Probe: Test-Time Adversarial Defense for CLIP

Backdoor as Probe is proposed, a test-time adversarial defense for CLIP that improves average robust accuracy from 1.0\% to 52.3\% while retaining clean accuracy, achieving performance comparable to state-of-the-art methods with up to a \(5.7\times\) inference speedup.

Zhong Ling Wang, Jie Zhang, Sen Nie et al. · 0 citations
#small language model Preprint Sep 2026

Automated Species Identification in Camera Trap Images for Wildlife Conservation

A novel end-to-end framework integrating a self-attention mechanism to address limitations in effectively detecting small animals in low-contrast trap images and small animals while also demonstrating zero-shot detection capability leveraging the MLLM.

Nowshin Amin, Nafisa Tabassum Oyshi, Tahmid Abrar Zidan et al. · 0 citations

From tech blogs

See all →

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.