Skip to content

Category

small language model

2,884 papers

#artificial intelligence Preprint Sep 2026

When Harness Beats Scale, and When Reading Beats Both

We describe our system for DocSem, the document-grounded quantitative reasoning shared task at DocInsights 2026, and analyze why it succeeded on labeled data and failed on the test set. The pipeline pairs hybrid block retrieval with Program-of-Thoughts (PoT) generation executed in a sandboxed interpreter, self-consiste...

Ivan Bondarenko, Nikolay O. Nikitin · 0 citations
#artificial intelligence Preprint Sep 2026

Learning to Steer, Steering to See: Unveiling the Geometry of RLVR in Large Language Models via Trainable Vectors

The Alpha-Stabler framework is proposed, a plug-and-play framework with a Predictor that monitors principal-subspace intrusion for early collapse warnings, and a Controller that removes the principal-subspace component of activation gradients during backpropagation while preserving the orthogonal complement.

Yu-Chen Cai, Ding Cao, Qi-Xiang Yin et al. · 1 citation
#artificial intelligence Preprint Sep 2026

Bayesian Active Learning for Intent Disambiguation in Interactive Robot Planning

This work proposes a Bayesian framework that treats clarification as an active learning problem over grounded Signal Temporal Logic task specifications and uses LLMs to initialize candidate formal specifications and translate informative contrasts into natural-language clarification questions, while Bayesian optimizati...

Hu-Ao Li, Carson Sobolewski, A. Saravanos et al. · 0 citations
#artificial intelligence Preprint Sep 2026

SAGE: Symbolic Action-Gating and Editing for LLM Task Planners

SAGE (Symbolic Action-Gating and Editing), a single-LLM planner built from two lightweight mechanisms: a domain-agnostic symbolic gate that blocks precondition-violating actions with typed reasons as a runtime safety monitor, and a local edit that regenerates only the failed sub-goal's suffix, keeping completed and unt...

T. Bui, Jongsul Moon, Youngouk Kim et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Unknown is not normal: separating language-model extraction from rule-based decision logic for clinical risk scores

Large language models (LLMs) are increasingly used to compute clinical risk scores from free-text notes. Notes are often incomplete, and treating undocumented findings as normal can silently misclassify patients. We test whether separating three-state extraction (present, absent or unknown, by an LLM) from decision log...

Nicolás Vera Zúñiga · 0 citations
#artificial intelligence Preprint Sep 2026

API Secrets Should Never Become Tokens in the LLM's Vocabulary: A Threat Analysis of API Credential Handling in LLM Agent Systems and an Empirical Evaluation of a Vault-Mediated Execution Boundary

Tool-using large language model (LLM) agents turn credential hygiene from a storage problem into an execution-security problem. A key pasted into a prompt, or embedded in a system prompt or tool configuration, crosses from an authentication boundary into a data pipeline, where it may persist in conversation history, lo...

P. Kenney, Hadi Ahmadi, Denis Lusson et al. · 0 citations
#artificial intelligence Open access Sep 2026

Query, Align, and Distill: Navigation-Aware Cross-Modal Interaction for Efficient Vision-and-Language Navigation

A high-performing teacher is built that makes navigation evidence selection explicit and compressible, and a compact student is trained by transferring both where to attend and what to do, then further match action distributions during fine-tuning.

Zhihao Chen, Yi-Yuan Ge, Zi-Yang Wang et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Communication between Frozen Large Language Models via Prompt Optimization in a Referential Game

Communication between two frozen large language models from different providers, with different tokenizers, accessed through their API endpoints is studied, finding that successful place value communication in some runs is rare.

V. Anand, Muthu Kumar Chandrasekaran, Shiva Chaitanya · 0 citations
#artificial intelligence Review Sep 2026

Overview of the TREC 2025 Million Large Language Models track

The TREC Million LLM Track operationalizes a retrieval-based paradigm in which an assistant agent infers expertise dynamically by examining models'observable behavior, providing the first large-scale benchmark for expertise retrieval in agentic AI.

Evangelos Kanoulas, Panagiotis Eustratiadis, Jamie Callan et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Don't Inoculate Everything: Stratified Inoculation Prompting Narrows Backdoor Triggers and Preserves Desired Traits

Supervised fine-tuning can teach language models undesired behaviours alongside desired ones. Inoculation prompting (IP) aims to limit unwanted generalisation by requesting the undesired behaviour during training and removing the request at inference. However, undesired behaviour can still appear under unrelated prompt...

Kajetan Dymkiewicz, Tim Farrelly, Adam Práda et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Action-Space Shaping for LLM Agents: Measuring and Mitigating Tool-Schema Bias

Large Language Models (LLMs) have shown strong performance on tool-use agentic tasks when given a fixed tool schema. Yet a tool schema is not the action space of an agent; it is merely one interface representation of it. The same executable action can be exposed through many different, functionally equivalent tool defi...

Yinhong Liu, Zhi-Li Tan, Zi-Lin Wang et al. · 0 citations

From tech blogs

See all →

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.