LLM agents are starting to own the full customer experience. Soon, LLMs may be selling and buying on behalf of companies and customers respectively. Small models are more cost-efficient at scale, but can reinforcement learning train them into competent sellers? We train four Gemma 4 checkpoints (2.3B to 31B effective p...
We develop and test a theory of language model representations in which there exist atomic features. Our main theoretical insight is that in such a model, sparse dictionaries (e.g., SAEs) of increasing size recover an increasing prefix of the most prevalent atoms in the training data. This"recovery principle"yields thr...
Kenny Peng, Jon M. Kleinberg, Nikhil Garg· 0 citations
Small language models (SLMs) offer a promising foundation for on-device agents through low-latency, resource-efficient inference, yet limited reasoning and planning capabilities constrain their performance on long-horizon tasks requiring multi-step interaction with the environment. Step-level collaboration between SLMs...
Zhe-Wei Fang, Yu-Xin Zhang, Zhen-Wei Shao et al.· 0 citations
Diffusion large language models (DLLMs) generate text through iterative block denoising, and multi-branch speculative decoding accelerates this process by verifying a main branch together with multiple draft branches in a single forward pass. While prior DLLM acceleration methods primarily exploit temporal redundancy a...
Chung-En Ho, Wei-Yu Sun, Cheng-Jhih Shih et al.· 0 citations
Reach audiences
Advertise in front of researchers, engineers, and readers.
Large language models (LLMs) can generate clinical narratives that are insufficiently grounded in patient-specific evidence. In traditional Chinese medicine (TCM), errors can propagate from etiology and pathogenesis through syndrome diagnosis and treatment principles to prescription generation. We developed TCMClinical...
Ji Dai, Chen-Kai Zhang, Yan Jia et al.· 0 citations
An annotation framework covering key aspects of travel history, including destination, exposures, travel duration, and multiple temporal variables is developed, used to annotate a corpus of 100 clinical case reports by five clinicians, yielding over 5,000 clinician annotations.
P. Kanithi, A. Pradhan, H. Cui et al.· medRxiv· 0 citations
Whether an open-weight pipeline running entirely on one consumer graphics card can accurately extract clinical information from faxed referral packets is evaluated.
Findings suggest model scaledoes not fully determine performance within a fixed pipeline,supporting compact local models as a practical, low-cost, lowlatencyalternative for enterprise deployment.
S. Shah, Aman Sheikh, S. Nath et al.· Proceedings of International...· 0 citations
Semantic scheduling is formulates as an uncertainty-aware framework that builds an admissible, quality-controlled scheduling instance from incomplete artifacts and optimizes it into feasible calendar-resource schedules.
D. Tyshchenko, V. Antypenko, A. Nenia et al.· Engineering, Technology &...· 0 citations
Yemeni students in Saudi universities occupy an unusual position. They study abroad, yet in a country that shares their language, religion, and much of their culture, while their families remain in a state that has been at war since 2014. Depression and anxiety are common among international students generally, and rec...
Abdulqader Ba Abbad, A. Alasheq, Mohammed Bokir et al.· Scientific Reports· 0 citations
The Robot Motion Framework is introduced, which combines modular robot hardware with a model-based toolchain consisting of a model database, a graphical modeling environment, and a model interaction language and is used in teaching MBSE and DSM concepts to aerospace engineering students at the University of Stuttgart.
Vanessa Tietz, Michael Wojczik, Bjoern Annighoefer· 0 citations
In an evaluation on two PREEvision Electric/Electronic-architecture models using automatically derived questions, it is found that text embedding models are capable of capturing the semantics of model elements, however, they often only find anchor points in the model and struggle to retrieve all relevant model elements...
Julian Roßkothen, David Inca Pilco, Tobias Hey et al.· 0 citations