A controlled protocol for evaluating answer stability is introduced: after a model answers a multiple-choice question correctly, it is challenged with a coherent argument for an incorrect option and measured whether the model flips, finding that self-attribution consistently increases flip rates and pooling wrong-answer arguments across models yields stronger adversarial challenges.
Nafiseh Nikeghbal, Amir Hossein Kargaran, Shaghayegh Kolli et al.· arXiv.org· 0 citations
Evaluating four state-of-the-art models finds that placing a role in a semantically unrelated context does not suppress role-linked attributes; instead, cross-role attribute concentration increases (pooled BI $+0.047$).
Shaghayegh Kolli, S. Emami, Moreno D'Incà et al.· 0 citations
Seeking to unify the evaluation of text-to-text privatization, PrivBench is introduced, a holistic and modular benchmarking platform for researchers and practitioners working on text privatization that evaluates privatization on a series of defined desiderata, which are structured into modules.
Stephen Meisenbacher, Andreea-Elena Bodea, A. Akın et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.