WildSEEK: Evaluating Language Models for Information-Seeking
This work introduces WildSEEK, a manually annotated dataset of 3k information-seeking queries from real user interactions, and an evaluation framework for LLM-generated responses, and finds that over a third of information-seeking queries are high-risk and more often analytical.
Tanise Ceron, Joachim Baumann, Elisa Bassignana et al.
· 0 citations