Retrieval-Augmented Generation (RAG) significantly improves Large Language Models (LLMs) but introduces massive input sequences that severely bottleneck the prefill stage. While KV-cache reuse reduces redundant computation for shared document prefixes, the reusable KV working set in RAG serving can exceed GPU memory ca...
Wenfeng Wang, Xiaofeng Hou, Peng Tang et al.· ACM Transactions on Architec...· 0 citations
Abstract Evaluating the quality of search systems traditionally requires a significant number of human relevance annotations. In recent times, several systems have explored the usage of Large Language Models (LLMs) as automated judges for this task while their inherent biases prevent direct use for metric estimation. W...
Abhishek Divekar, Anirban Majumder· AI Magazine· 0 citations
Abstract Rapid advances in large language models (LLMs) have empowered autonomous agents to generate social networks, communicate, and form shared and diverging opinions on political issues. However, our understanding of their collective behaviours and underlying mechanisms remains incomplete. In this paper, we simulat...
The ability of Large Language Models (LLMs) to outperform humans on many intellectual tasks has created a difficult landscape for assessing student work. One strategy is designing assignments to be specifically difficult for AI. However, LLMs have been shown to perform well on general chemistry problems across topics,...
Emilie Ernst, Jennifer R. DeRosa, Brian G. Ernst· American Journal of STEM Edu...· 0 citations
This record provides the evaluation data accompanying the paper “A Safety-First Two-Stage Mental Health SupportChatbot: Robust Crisis Routing andRetrieval-Grounded Response Generation” (SU26 DAP Group 4, FPT University, Can Tho Campus). The data is designed to stress-test the robustness of a two-stage mental-health sup...
Vinh Hung Ho· Zenodo (CERN European Organi...· 0 citations
SACMQ — South Asian Cultural Multilingual Question Dataset contains multilingual question-answering data in five South Asian languages: Bangla, Hindi, Urdu, Tamil, and Nepali. It was developed to evaluate large language models (LLMs) and Retrieval-Augmented Generation (RAG) systems on multilingual knowledge, with topic...
Fazle Mohammad Tasfiq, Sabab Attin, Sheikh Md. Tahmid Hossain Faiyaz et al.· Zenodo (CERN European Organi...· 0 citations
Objective: To critically evaluate the current evidence regarding the application of generative artificial intelligence (GenAI) in oncology patient counselling, with particular emphasis on oncology pharmacy practice, and to distinguish direct oncology evidence from indirect evidence derived from general healthcare and o...
Hemandh S N, Subha Gayathri Munta, Jeeshitha Javvadi· Journal of Oncology Pharmacy...· 0 citations
This report constitutes Deliverable D2.5 of Work Package 2 under the Horizon Europe-funded project Longitudinal Educational Achievements: Reducing Inequalities (LEARN; Grant Agreement No. 101132531). Focusing on Task 2.5, the report systematically examines the adoption, adaptation, and structural resistance to Evidence...
Stephen P. Morris, Lee Bentley· Zenodo (CERN European Organi...· 0 citations
ABSTRACT The present monographic study constitutes a fundamental, interdisciplinary oeuvre of the V. A. Alekseev, establishing the theoretical and instrumentation foundation for the next generation of non-silicon computing paradigms. The monograph unfolds a rigorously sequential, mathematically verified formulaic frame...
Valery Alekseev· Zenodo (CERN European Organi...· 0 citations
Large Language Models (LLMs) are increasingly integrated into software systems used in healthcare, public administration, and customer service. Decisions about data collection, use, and retention can introduce privacy and security risks, reinforce biases, and affect people who interact with or are subject to these syst...
Anonymous Anonymous· Zenodo (CERN European Organi...· 0 citations
Large language models (LLMs) increasingly excel at mathematics tasks, but their unreliability limits their utility in mathematics research. A mitigation is to use LLMs to generate formal proofs in languages such as Lean, in which the compiler verifies every proof step. We present the first demonstration of this method’...
George Tsoukalas, Anton Kovsharov, Sergey Shirobokov et al.· Science· 3 citations