Small language model (SLM) agents need safety controls that track consequences across tool calls with little monitoring overhead. A private read, for example, becomes a leak when a later action sends that data outside the system. We introduce CARB (Conformal Agent Risk Budget), which calibrates when to stop an agent us...
Zi-Jun Yu, Yu-Tong Gu, Vahid Partovi Nia et al.· 0 citations
Infrared small-target detection plays an important role in maritime monitoring and aerial surveillance. Although multimodal large language models (MLLMs) offer promising capabilities for visual understanding, existing MLLM-based approaches struggle to precisely localize infrared small targets. In this paper, we propose...
Jia-Wen Xi, Yu Zhang, Tian-Yi Zhao et al.· 0 citations
Large language model watermarking embeds detectable statistical signals during decoding, but the resulting changes to token probabilities can degrade generation quality. This trade-off is particularly important for code, where small changes in token selection can break syntax or alter program behavior. Existing code wa...
Hyundong Jin, Hyeseon An, Soohan Lim et al.· 0 citations
Training large language models (LLMs) entails a fundamental trade-off: memory-efficient optimizers such as Adam discard cross-parameter curvature, whereas full-curvature methods such as SOAP can accelerate convergence at prohibitive memory costs. We introduce Clean, a memory-efficient and full-curvature optimizer desig...
Beheshteh T. Rakhshan, S. Rajabi, Maziar Sargordi Shikai Fang et al.· 0 citations
Reach audiences
Advertise in front of researchers, engineers, and readers.
Large language model (LLM) agents are increasingly capable of acting in complex tool-use environments, yet they often fail to recognize when tasks are infeasible and no valid solution exists. Recent work has formalized this reliability gap as the problem of agentic abstention, and existing approaches typically optimize...
Hang Luo, Bing-Bing Wen, Guang Yang et al.· 0 citations
Adapting general-purpose large language models to specific tasks requires substantial human effort in designing data and training strategies. Sustaining improvement is especially challenging because model updates change the error distribution, requiring strategies to be continually refined. We introduce ImproveAnyTask,...
Xing-Bo Yao, Xiao-Man Wang, Zheng-Wu Lei et al.· 0 citations
LLM agents are starting to own the full customer experience. Soon, LLMs may be selling and buying on behalf of companies and customers respectively. Small models are more cost-efficient at scale, but can reinforcement learning train them into competent sellers? We train four Gemma 4 checkpoints (2.3B to 31B effective p...
We develop and test a theory of language model representations in which there exist atomic features. Our main theoretical insight is that in such a model, sparse dictionaries (e.g., SAEs) of increasing size recover an increasing prefix of the most prevalent atoms in the training data. This"recovery principle"yields thr...
Kenny Peng, Jon M. Kleinberg, Nikhil Garg· 0 citations
Small language models (SLMs) offer a promising foundation for on-device agents through low-latency, resource-efficient inference, yet limited reasoning and planning capabilities constrain their performance on long-horizon tasks requiring multi-step interaction with the environment. Step-level collaboration between SLMs...
Zhe-Wei Fang, Yu-Xin Zhang, Zhen-Wei Shao et al.· 0 citations
Diffusion large language models (DLLMs) generate text through iterative block denoising, and multi-branch speculative decoding accelerates this process by verifying a main branch together with multiple draft branches in a single forward pass. While prior DLLM acceleration methods primarily exploit temporal redundancy a...
Chung-En Ho, Wei-Yu Sun, Cheng-Jhih Shih et al.· 0 citations
Large language models (LLMs) can generate clinical narratives that are insufficiently grounded in patient-specific evidence. In traditional Chinese medicine (TCM), errors can propagate from etiology and pathogenesis through syndrome diagnosis and treatment principles to prescription generation. We developed TCMClinical...
Ji Dai, Chen-Kai Zhang, Yan Jia et al.· 0 citations
An annotation framework covering key aspects of travel history, including destination, exposures, travel duration, and multiple temporal variables is developed, used to annotate a corpus of 100 clinical case reports by five clinicians, yielding over 5,000 clinician annotations.
P. Kanithi, A. Pradhan, H. Cui et al.· medRxiv· 0 citations