An autonomous cyber defender trained with reinforcement learning (RL) is typically tied to the network on which it was trained, limiting its ability to generalize as network scale changes. Hierarchical RL reduces decision complexity by separating strategic targeting from tactical execution, but it does not eliminate th...
Harshith Doppalapudi, Nathaniel D. Bastian, Ankit Shah· 0 citations
As Large Language Models (LLMs) increasingly serve as foundational reasoning engines, their tendency to hallucinate remains a critical vulnerability. While recent internal state probes offer a promising alternative to slow external retrieval systems, they largely reduce hallucination detection to a token-wise binary cl...
Natural language descriptions can provide rich semantic representations of audio-visual urban scenes, yet datasets that jointly describe both auditory and visual information remain limited. In this paper, we introduce AVSD-Scenes, a paired audio-visual scene description dataset for urban environments. The dataset conta...
Dhanunjaya Varma Devalraju, Arshdeep Singh, Mark D. Plumbley· 0 citations
An emerging hypothesis about reinforcement learning (RL) post-training of large language models (LLMs) is that it merely sharpens existing behaviors of a base model, improving single-shot accuracy at the cost of solution coverage. Although this trade-off has been observed in math and coding tasks, it need not extend to...
Changdae Oh, Qi Zeng, Qi Qi et al.· 0 citations
Reach audiences
Advertise in front of researchers, engineers, and readers.
Molecular tumor boards integrate genomic findings, clinical context, and therapeutic evidence to support precision oncology. As AI enters this workflow, a key safety challenge is distinguishing truly unsupported recommendations from evidence-supported options that still require oncologist review because of incomplete i...
Negin Ashrafi, Jia Luo, Stacey M. Frumm et al.· 0 citations
Collaboration between a small language model (SLM) and a large language model (LLM) offers an opportunity to combine the efficiency of smaller models with the strong reasoning capabilities of larger ones. Existing approaches primarily frame such collaboration as a computation allocation problem, determining which model...
Concurrent actions in large language model (LLM) agent environments require arbitration even when each proposal is individually valid. We implement a typed snapshot-settlement contract and audit three distinct properties: order sensitivity, useful progress, and replay consistency. Five settlement policies are tested in...
Hao-Tian Chen, Bo-Wen Ye, Yu-Ning Zhang et al.· 0 citations
Large language models can translate business descriptions into optimization models, but executable code may misrepresent constraints or objectives. A solver can then return an optimal solution to the wrong problem. Even when the solution satisfies the intended operating rules, a better plan may exist. For organizations...
Jin-Zhi Bu, Hai-Xin Tang, Hua-Nan Zhang· 0 citations
Large language models (LLMs) increasingly underpin scientific AI applications that reason over structured knowledge, from biomedical question answering to materials informatics. However, their logical reasoning often falls short, producing factual inaccuracies unacceptable in these settings. Reliable evaluation remains...
Nishtha N. Vaidya, S. Grimm, T. Hubauer et al.· 0 citations
Reusable Latent Correction (RLC) is proposed, which converts one-off natural-language guidance from a black-box LLM into persistent corrective experiences in the hidden space of an SLM, enabling the SLM to reuse LLM-derived corrections during inference without any online LLM calls.
Bo-Han Zhang, Li-Nan Yue, Weibo Gao et al.· 0 citations
Combining two-step denoising with representative execution methods substantially reduces inference cost with a small reduction in task performance, motivating joint optimization of model-inference efficiency and robot-system timing.
Di Wu, Rong-Tian Shen, Ping Liu et al.· 0 citations
Comparing empirical cross-lingual transfer with typology-based similarity, it is found that transfer BLEU identifies closely interacting language pairs better than URIEL similarity, though neither predicts which varieties benefit from joint training.
Frank Lawrence Nii Adoquaye Acquaye, Eric George Parakal, Jesse Johnson et al.· 0 citations