This work trains six independent linear probes on open-weight language models, one per Moral Foundations Theory (MFT) category, and examines how the resulting directions relate to each other in representation space, finding the directions neither collapse into a single moral detector nor isolate from one another.
Abstract
How do large language models (LLMs) organize moral knowledge? Models detect moral content broadly, but detection is a low bar. We ask whether they go further, distinguishing moral foundations from one another and organizing the relationships between them geometrically. We train six independent linear probes on open-weight language models, one per Moral Foundations Theory (MFT) category (care/harm, fair/cheat, lib/oppress, loy/betray, auth/subv, sanc/degrade), and examine how the resulting directions relate to each other in representation space. We find the directions neither collapse into a single moral detector nor isolate from one another. Rather, they span a near-maximal number of independent dimensions while sharing a positive common component. The shared component is the signature of integration, and it is moral-specific relative to a matched non-moral concept battery built identically (mean pairwise cosine 0.26 vs. 0.013). The geometry is consistent across architectures and scale and reaches its integration regime early in pre-training, well before probe accuracy saturates. The structure the model discovers shows no evidence of the individualizing/binding distinction predicted by Moral Foundations Theory (an underpowered test: only 20 candidate partitions exist) but rather reflects corpus statistics. Extending to moral dilemmas, each dilemma direction partially composes from its component foundations, at 2.7x a mismatched-pair baseline, while the majority of its variance encodes conflict-specific structure. The model represents moral tension itself, not a pre-resolved judgment.
This work presents the first systematic study of inverse relation directionality in LLMs, using a benchmark consisting of 5,457 instances spanning 27 distinct inverse relation labels and reveals systematic asymmetries in inverse relation classification across LLMs.
This work quantifies cross-lingual skill inconsistency orthogonally from knowledge and general benchmark performance via multilingual self-play, and shows that skill discrepancies are a measurable major roadblock in the development of truly multilingual models.
Bobby Cheng, Adam Gaber, Zhengzhe Liu et al.· 0 citations
As Large Language Models (LLMs) grow more capable across diverse tasks, their (in)ability to generalize remains difficult to quantify and poorly understood beyond limited domains. In particular, LLMs are known to struggle generalizing multilingually, to languages outside of English, and that are poorly attested in their training data. To understand why this may be, and what enables some models to perform better than others, we turn to a long history of work across the cognitive sciences, arguing that successful generalization derives from appropriate representations in similarity space. We look at how well LLMs'representations capture the hierarchical similarity structure between distinct languages. Strikingly, we show LLMs'latent representations largely recover the hierarchical structure of the Indo-European language family tree -- grouping languages that are members of the same subfamily closely together in representation space. Furthermore, we show that the degree to which models reflect the similarity structure of languages correlates with their performance on XNLI, a multilingual natural language inference benchmark. This extends classic work on similarity-driven generalization at scale, showing how models that represent similar languages similarly generalize better from one language to another.
Supantho Rakshit, Adele E. Goldberg, Henry Conklin· arXiv.org· 0 citations
Large language models (LLMs) have achieved remarkable gains in cognitive performance through attention mechanisms functionally inspired by human attention. This paper asks philosophically whether a comparable architectural insight could technically advance moral processing. We argue that current alignment techniques primarily shape outputs after representations have been formed. They therefore cannot realise, using Iris Murdoch’s loving attention approach, a just, reality-sensitive orientation toward others that operates at the level of representation. Drawing on Murdoch’s moral philosophy, we identify three substrate-neutral features of loving attention, namely locational, representational, and dispositional, that survive translation from human moral phenomenology to computational systems. On this basis, we propose that moral processing in LLMs should be studied not only as a problem of output control but also as a problem of representational architecture. We further formulate the loving attention geometry hypothesis, suggesting that if morally improved perception involves a systematic shift from ego distorted to more just representations, this transformation may leave detectable structure in LLM embedding and activation spaces. With this focus on Murdoch, the paper contributes a novel philosophical technical approach that links moral attention, representation learning, and AI alignment. It clarifies why architectural considerations matter for moral AI, distinguishes between the tractability and validation of moral geometry, and outlines design constraints for responsible research. We argue that exploring representational forms of moral attention is a promising and necessary direction for AI ethics research. If transformer attention has been able to scale intelligence, it is worth inquiring if representations and architectures can scale morality.
G. Bombaerts, Bram Delisse, U. Kaymak· AI and Ethics· 0 citations
Convergence is measured against a human reference nobody built for the purpose -- 2,523 reader mark sets across 120 web documents, produced by people highlighting for their own reasons on a platform where the overlay of others'marks is off by default.
B\"urger et al. (2024) demonstrated that truth representations in large language models are universal across statement polarity but reside within a multidimensional subspace. We extend this framework along three questions: how the dimensionality of the subspace depends on the model's knowledge, which architectural component builds the truth direction, and what the direction is a mixture of. In Part I, a training-free directional probe derived from the SVD of hidden-state minimal pairs shows that the dimensionality of truth is knowledge-dependent: the signal concentrates on a single axis for known facts and diffuses as knowledge decreases. In Part II, a relational law emerges across multiple model families: attention propagates truth frames, the feed-forward network opposes the current block's frame, and post-peak decay is causally attributed to the SwiGLU value stream. Furthermore, per-category truth axes form a semantically signed arrangement that converges across families. Stress tests expose a sign instability in this orientation, which we repair with a spectral consensus gauge to sharpen the convergence into a knowledge-gated law. Finally, a replication campaign on Gemma-2-2b, extending our decomposition tools to accommodate its sandwich normalization, confirms these laws and attributions. We quantify the knowledge gate as classical attenuation and isolate a stable, model-specific private geometry.
Francesco Vicidomini· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.