LLM value studies often merge questionnaire ratings, pairwise choices, and values inferred from generated text into one profile. That merge assumes that the three observations describe the same stable preference. STONIC tests this assumption on 5,144 situations from four banks and 35 fixed model configurations. It comp...
A. Chetvergov, S. Ukolov, Timofei Sivoraksha et al.· 1 citation
Aggregate factuality scores hide where a language model succeeds, which relations it confuses, and whether an answer survives innocuous changes to the question or decoder. We introduce PROOF, a profile-oriented benchmark for factual coverage in instruction-tuned language models. PROOF converts a frozen Wikidata snapsho...
A. Chetvergov, Mikhail Solovev, Timofei Sivoraksha et al.· 0 citations
VIBE is introduced, a benchmark for entity-centered affective profiling of LLM outputs in Valence-Arousal-Dominance (VAD) space and its core contribution is a measurement contract, which motivates entity-centered affective profiling as a documented practice.
A. Chetvergov, Alexander Evseev, Timofei Sivoraksha et al.· 0 citations
Large language models trained and aligned within different linguistic and regional ecosystems may frame the same political, cultural, and geopolitical entities in different ways. Such differences are often evaluated through sentiment, favorability, or stance, reducing model attitudes to a single positive-negative axis....
A. Chetvergov, Alexander Evseev, Mikhail Solovev et al.· arXiv.org· 0 citations
This work evaluates 21 instruction-tuned LLM runs under a fixed ranked-response protocol, showing that models often locate the correct motivational region while ranking close alternatives unstably, and motivates value-recognition evaluation that combines exact accuracy, ranked recovery, and directed error analysis.
A. Chetvergov, S. Ukolov, Timofei Sivoraksha et al.· arXiv.org· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.