Skip to content

Category

small language model

2,864 papers

#artificial intelligence Preprint Sep 2026

SQUARE: Structured Quantum Representation Adapters as Compact Quadratic Feature Maps for Frozen Language Models

Frozen language models (LMs) are increasingly used as fixed feature extractors for downstream reranking, scoring, and preference modeling, raising a practical question: how should a compact module represent interactions among features in a fixed low-dimensional bottleneck? Common linear and low-rank adapters remain lin...

E. Roh, Hyojun Ahn, Hoyeong Lee et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Unlocking the Critic: Reward-Free Policy Optimization for LLM Post-Training

Recent approaches to reinforcement learning (RL) post-training for large language models increasingly remove the critic to reduce training instability and memory overhead. Even where a critic is trained, it is discarded once training ends, although it has learned to predict outcomes. We revisit this trend and show that...

Hong-Yang Li, Xiao Li, Caesar Wu et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Controlled Decoding Attacks on Black-Box LLMs

This work introduces \method{}, a framework for jailbreaking through text-only continuation interfaces that permit repeated sampling and assistant-prefix continuation, and achieves the highest mean score most comparisons against baselines.

Jesson Wang, Shawn Li, Wei Yang et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Safer Content or Firmer Refusals? A Hybrid Perturbation Defense for Alignment under Harmful Fine-tuning

Fine-tuning-as-a-service lets users adapt a safety-aligned language model to their own data, but it also creates a harmful fine-tuning attack surface: a small amount of harmful data mixed into an otherwise benign fine-tuning set can degrade the model's alignment. Two recent alignment-stage defenses address this problem...

Muhammad Zeeshan Akram, Mufid Kamel Marican, Anvesh Reddy Yenugu et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Spotter: Let the Embodied Model Lead, and the VLM Reflect for It

Spotter is proposed, which reverses the roles: the embodied model leads and executes continuously, while the VLM runs in parallel, monitors through a lightweight local screener, intervenes only when an error is detected, reflects on and corrects it, and returns control.

Long Li, Qi-Chao Zhao, Yue Yang et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Large-scale factor analysis shows machine intelligence is only partially interpretable

A common assumption in language model development is that cognitive abilities are organized around a general, domain-free intelligence factor, like fluid intelligence in humans. This assumption is rarely tested directly, and prior attempts have done so only at a much smaller scale. We take a latent variable approach to...

Faiz Ghifari Haznitrama, Afrizal Hasbi Azizy, Faeyza Rishad Ardi · 0 citations
#artificial intelligence Preprint Sep 2026

Similar Choices, Different Attention: Cross-Modal Associations in Humans and Vision-Language Models

Cross-modal associations are systematic pairings of features across modalities, such as the association of'bouba'with round shapes and'kiki'with sharp shapes. Prior work has compared humans and vision-language models (VLMs) on such associations, but often using different stimuli or tasks between humans and models. Here...

Su-Min Hong, Katsumi Ibaraki, Renee Shi et al. · 0 citations
#artificial intelligence Review Sep 2026

Calibrated to Whom? Persona and Language Effects on Cultural Values in JEV

Decision-only language models return a probability for every answer option instead of generating text, which makes them attractive as survey respondents and as judges. We audit the cultural values of one such model, TypeSafe's JEV, with the Values Survey Module 2013. We asked it the 24 items as 12 matched Saudi and 12...

Bushra Asseri, Abdulaziz M. Asseri · 0 citations
#artificial intelligence Preprint Sep 2026

AerialDojo-200K: A Large-Scale Benchmark Suite for Open-World Aerial Object-Goal Search

A large-scale benchmark suite for open-world aerial object-goal search, with 3 times as many scenes and 18.7 times as many task instances as the largest existing benchmark for this task, and a unified evaluation framework with a unified evaluation framework.

Tong-Tong Feng, Xin Wang, Hao-Ran Hou et al. · 0 citations
#artificial intelligence Preprint Sep 2026

AdviSD: Learning to Advise Frontier LLMs via Targeted Multi-Turn Self-Distillation

A small trainable advisor can steer a frozen language-model executor using natural-language advice. In addition to learning from task rewards, the advisor can use feedback from completed interactions to improve its advice. However, a plausible correction need not change execution, yet learning from such corrections can...

Rishabh Agrawal, He-Jie Cui, Sha-Sha Li et al. · 0 citations
#artificial intelligence Preprint Sep 2026

HARISSA: Inference-Time Self-Checks for Efficient and Safe Local Language Model Deployment

Running a language model locally offers advantages in privacy, latency, and cost, but local hardware fits only small models, which are less capable than frontier models. The usual remedy for a hard query, escalating it to a cloud model, gives up the privacy and cost advantages of running locally. A deployment that stay...

Kenan Alkiek, Moontae Lee, David Jurgens et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Cross-Entropy Guided Routing in Mixture-of-Experts Large Language Models

Sparse mixture-of-experts (MoE) large language models scale model capacity by routing each token to a small subset of experts. Their routers are regularized with load balancing terms and learn affinity scores through the language-model objective. However, these objectives do not provide direct alignment between routing...

Yury Nahshan, Nati Daniel, Jacob Goldberger et al. · 0 citations

From tech blogs

See all →

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.