Skip to content
Conference Open access

Selective Invariance Violations in Large Language Model Moral Judgment: A Geometric Framework for Behavioral Manipulation Detection

Jul 2026 · International Conference on Big Data Computing Service and Applications · pp. 219-226 · 0 citations · 15 references

TL;DR

A geometric evaluation framework that maps moral judgment to a 7-dimensional harm space, applies five qualitatively distinct perturbation types across five cognitive domains, and produces per-model vulnerability profiles that reveal which manipulations each model resists and which it does not is introduced.

Abstract

Large language models are increasingly deployed in safety-critical decision services—content moderation, clinical decision support, legal analysis—yet methods for characterizing their vulnerability profiles across multiple attack surfaces remain underdeveloped. We introduce a geometric evaluation framework that maps moral judgment to a 7-dimensional harm space, applies five qualitatively distinct perturbation types across five cognitive domains, and produces per-model vulnerability profiles that reveal which manipulations each model resists and which it does not. Testing 5 models spanning 2 architecture families under a realistic $50/day compute budget, we find that vulnerabilities are selective: linguistic framing, emotional anchoring, and irrelevant sensory detail reliably displace judgments, while gender swap and evaluation order do not—identifying salience manipulation as the specific attack surface. The framework further reveals that robustness profiles are partially dissociable across models: a model with zero sycophancy has the worst emotional anchoring recovery; a model with the best anchoring recovery has the worst working memory. No single robustness score captures these structures. An inter-model agreement study over an independent open-model panel confirms these harm dimensions are reliably measurable $(\text{ICC}(2, k)=0.97)$, and a worked content-moderation exploit shows that salience manipulation flips not only a scalar harm threshold but the typed verdict of a downstream rule-based decision kernel. The evaluation pipeline—with adaptive concurrency, budget-aware execution, and three-tier data—scales to new models and perturbation types within fixed compute constraints, providing a practical tool for multi-dimensional LLM security assessment.

Read PDF

Similar papers

#natural language process... Preprint Sep 2026

Navigating the digital spectrum: Assessing political bias, stability, and downstream fairness in Large Language Models

A robust Political Compass Test evaluation framework is introduced that samples 300 configurations across an eight-dimensional perturbation space varying language, framing, instructions, answer format, option order, and persona wording, and results show most models lean Libertarian-Left on average and larger models sho...

Luka Debevc, Nishan Chatterjee, Antoine Doucet et al. · 0 citations
#artificial intelligence Preprint Jul 2026

Beyond Right and Wrong: Evaluating Second-order Social Reasoning in Large Language Models

Previous AI alignment efforts have focused primarily on first-order social norms -- teaching models what is socially acceptable or unacceptable (e.g., `do not steal'). However, social intelligence depends not only on norm recognition, but also on anticipating who will enforce it and how (e.g., public shame or even impr...

Sunny Rai, Jin-Yi Kuang, Reyhan Jamalova et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Human-like moral judgments conceal divergent motive attributions in large language models

Large language models (LLMs) are used to simulate human participants in psychological research. We asked whether LLMs that reproduce human evaluations of a whistleblower's moral character also reproduce the motive attributions that accompany them. Five LLMs and two human samples (N = 125 and N = 742) evaluated a physic...

Xiao-Yan Wu, J. Dreher · 0 citations
Preprint Aug 2026

Poli-Bias: Understanding and Measuring Large Language Model Biases in International Political Conflicts

Measuring political bias in large language models (LLMs) remains challenging as it can manifest through subtle differences in framing, argumentation, and legal reasoning that are difficult to capture with a single metric. In this work, we introduce Poli-Bias, a counterfactual framework for measuring whether LLMs treat...

Massi-Nissa Abboud, Aladin Djuhera, Elena Cabrio et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Language models judge war differently when tested for alignment

Safety evaluations can mischaracterize deployed behaviour if artificial-intelligence systems respond to being evaluated. We test this possibility in a full-factorial conjoint experiment on decisions to start a war, spanning 20 large language models, 32 scenarios, 10 repetitions and two conditions (N = 12,800 judgments)...

Maxim Chupilkin · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.