Performance within the foundation paradox: the case for small qualitative evaluations
This paper argues for SQEs as a complementary evaluation strategy for AI systems operating in dynamic, contested, and interdisciplinary settings and introduces small qualitative evaluations (SQEs) as a human-in-the-loop framework for assessing LLM performance in such less-bounded domains.
Michael Simeone, Jacki Hyatt, E. Bienenstock et al.
· Neural computing & applicati... · 0 citations