We describe our system for DocSem, the document-grounded quantitative reasoning shared task at DocInsights 2026, and analyze why it succeeded on labeled data and failed on the test set. The pipeline pairs hybrid block retrieval with Program-of-Thoughts (PoT) generation executed in a sandboxed interpreter, self-consiste...
Commutator memory is defined per source pair, not per example, and its projection on $b_{AB}$ decays with further training, which identifies which came from which order in 92% of cases across four LLMs (chance 50%).
The Alpha-Stabler framework is proposed, a plug-and-play framework with a Predictor that monitors principal-subspace intrusion for early collapse warnings, and a Controller that removes the principal-subspace component of activation gradients during backpropagation while preserving the orthogonal complement.
Yu-Chen Cai, Ding Cao, Qi-Xiang Yin et al.· 1 citation
This work proposes a Bayesian framework that treats clarification as an active learning problem over grounded Signal Temporal Logic task specifications and uses LLMs to initialize candidate formal specifications and translate informative contrasts into natural-language clarification questions, while Bayesian optimizati...
Hu-Ao Li, Carson Sobolewski, A. Saravanos et al.· 0 citations
Reach audiences
Advertise in front of researchers, engineers, and readers.
SAGE (Symbolic Action-Gating and Editing), a single-LLM planner built from two lightweight mechanisms: a domain-agnostic symbolic gate that blocks precondition-violating actions with typed reasons as a runtime safety monitor, and a local edit that regenerates only the failed sub-goal's suffix, keeping completed and unt...
T. Bui, Jongsul Moon, Youngouk Kim et al.· 0 citations
Large language models (LLMs) are increasingly used to compute clinical risk scores from free-text notes. Notes are often incomplete, and treating undocumented findings as normal can silently misclassify patients. We test whether separating three-state extraction (present, absent or unknown, by an LLM) from decision log...
Tool-using large language model (LLM) agents turn credential hygiene from a storage problem into an execution-security problem. A key pasted into a prompt, or embedded in a system prompt or tool configuration, crosses from an authentication boundary into a data pipeline, where it may persist in conversation history, lo...
P. Kenney, Hadi Ahmadi, Denis Lusson et al.· 0 citations
A high-performing teacher is built that makes navigation evidence selection explicit and compressible, and a compact student is trained by transferring both where to attend and what to do, then further match action distributions during fine-tuning.
Zhihao Chen, Yi-Yuan Ge, Zi-Yang Wang et al.· IEEE transactions on circuit...· 0 citations
Communication between two frozen large language models from different providers, with different tokenizers, accessed through their API endpoints is studied, finding that successful place value communication in some runs is rare.
V. Anand, Muthu Kumar Chandrasekaran, Shiva Chaitanya· 0 citations
The TREC Million LLM Track operationalizes a retrieval-based paradigm in which an assistant agent infers expertise dynamically by examining models'observable behavior, providing the first large-scale benchmark for expertise retrieval in agentic AI.
Evangelos Kanoulas, Panagiotis Eustratiadis, Jamie Callan et al.· 0 citations
Supervised fine-tuning can teach language models undesired behaviours alongside desired ones. Inoculation prompting (IP) aims to limit unwanted generalisation by requesting the undesired behaviour during training and removing the request at inference. However, undesired behaviour can still appear under unrelated prompt...
Kajetan Dymkiewicz, Tim Farrelly, Adam Práda et al.· 0 citations
Large Language Models (LLMs) have shown strong performance on tool-use agentic tasks when given a fixed tool schema. Yet a tool schema is not the action space of an agent; it is merely one interface representation of it. The same executable action can be exposed through many different, functionally equivalent tool defi...
Yinhong Liu, Zhi-Li Tan, Zi-Lin Wang et al.· 0 citations