Skip to content

Category

small language model

2,822 papers

#natural language process... Preprint Oct 2026

Learning When to Commit from Partial Speech for End-to-End Simultaneous Speech Translation

Simultaneous speech translation must emit useful target text before the source is complete while preserving every committed token. We adapt a full-utterance speech language model using prefix supervision derived from its own complete- and partial-waveform translations, requiring neither transcripts nor human translatio...

Hieu Hoang, Amittai Axelrod · 0 citations
#machine learning Preprint Oct 2026

FALCON: A Model and Dataset Agnostic Framework for Synthetic Data Generation for NL2SQL Pairs

Relational databases are among the most widely deployed forms of structured knowledge, and natural language access to them requires grounding language onto schema entities and relations while handling the ambiguity inherent in how people phrase requests. Existing synthetic NL-to-SQL data generation methods largely igno...

Darian Lee, Shannon Rumsey, Jack St. Clair et al. · 0 citations
#machine learning Preprint Oct 2026

HyperThink: Text-to-Parameter Hypernetworks for Efficient Reasoning

Long-form thinking traces can substantially improve the multi-step reasoning performance of large language models (LLMs), but they introduce high inference-time overhead, with latency dominated by sequential decoding. We propose HyperThink, a text-to-parameter approach that amortizes this reasoning computation into a s...

Donggyun Kim, Jack Lu, Chanwoo Kim et al. · 0 citations
#machine learning Preprint Oct 2026

LESSER: Post-Training Data Selection with Output-Layer Gradients

The choice of post-training data for large language models substantially affects downstream performance. Gradient-based data selection is a popular approach that ranks training data by how well their gradients align with those of a small validation set. However, ranking with full-parameter gradients requires an expensi...

Lyuxin David Zhang, Eric Wong, Surbhi Goel et al. · 0 citations
#machine learning Preprint Oct 2026

Exact Memory-Time Optimization for Prefix-Cached Language Model Serving

Retaining language-model prefix states trades recomputation against storage time. Optimizing each cached block independently can overcount savings: a resident block is usable only when the required preceding prefix is also available. We introduce Prefix-Certificate Retention (PCR), an exact finite-trace formulation for...

Shivam Gupta · 0 citations
#machine learning Preprint Oct 2026

Online Verification of Language Model Responses Under Cost Constraints

As large language models are increasingly deployed for multi-step reasoning, verifying the correctness of their outputs has become essential for maintaining reliability at scale. Verifying the correctness of large language model outputs is often done by querying a costly ground-truth oracle, which is impractical to inv...

Erfan Hajihashemi, Yan-Ning Shen · 0 citations
#machine learning Preprint Oct 2026

Trained Agentic Context Management

We study long context language models. Instead of training long context natively, or designing a long context harness, we train a model over the simplest possible harness: a tool to call itself with any specified prompt and a tool to read tokens in a range from the input context. We finetune Qwen3.6-35B-A3B on a divers...

Bryce Sandlund · 0 citations
#artificial intelligence Preprint Oct 2026

SyntaxBench: A Statistical Diagnostic Framework for Character-Level Reasoning in Large Language Models

Large language models are increasingly used where small syntactic errors matter, yet character-level reasoning is still evaluated mostly through isolated probes and aggregate accuracy. We introduce SyntaxBench, a diagnostic benchmark and statistical evaluation framework for character-level reasoning. It contains five c...

Mohsen Larni, Sobhan Ebrahimi Azar, Pouyan Nahed et al. · 0 citations
#artificial intelligence Preprint Oct 2026

DyRA: Dynamic Residual Approximation for Efficient Matrix Multiplication in DNNs

Large-scale foundation models achieve strong performance across diverse tasks, but their size makes inference costly, largely due to dense matrix multiplications. Prior work reduces this cost by replacing dense weight matrices with efficient structured forms such as low-rank factorizations. However, these methods appro...

Daewon Chae, Hyunwon Chung, Changwoo Lee et al. · 0 citations
#artificial intelligence Preprint Oct 2026

FastOPD: On-Policy Distillation for Lightweight VLA Deployment

Vision-Language-Action (VLA) foundation models have scaled rapidly to enhance manipulation performance and generalizability, but this scaling incurs high computational costs that render real-world deployment increasingly challenging. Existing approaches typically mitigate this issue by designing smaller architectures o...

Yoojin Oh, Jeongsol Kim, Yeonwoo Seo et al. · 0 citations
#artificial intelligence Preprint Oct 2026

Improving the Energy-Efficiency of the Code Generated by LLMs through Effective Prompting

As AI-assisted programming becomes increasingly mainstream, the environmental impact of AI-generated software has emerged as an important consideration. This motivates evaluating LLM-generated code beyond functional correctness by considering execution efficiency and energy consumption. However, despite substantial adv...

Ritika Rekhi, Bing Zhang, Md Arman Islam et al. · 0 citations
#artificial intelligence Preprint Oct 2026

Toward SLM-based agentic task-tool intent matching

Tool-equipped AI agents use tool calls to access data and act on external systems. Horizontal growth of agentic systems increases the number of these interactions, and further motivates the need for automated, per-call oversight that can operate at low latency and/or on-prem. Conventional authorization schemes can dete...

C. Troiani, Arash Salarian, Majed El Helou et al. · 0 citations

From tech blogs

See all →

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.