Skip to content

MAxBench: A Multinomial Concept Recovery Benchmark

Sep 2026 · 0 citations · 57 references
Computer Science

TL;DR

This work introduces MAxBench, a geometry-agnostic evaluation framework for multinomial concept representations based on sampling from the recovered concept representation, and finds that affine subspaces steer more reliably and have greater recall than rank-one or linear subspaces.

Abstract

Fine-grained control of language model behaviors (e.g., steering) is among the more actionable outcomes of interpretability research. For binary concepts such as refusal, a single direction in activation space often suffices for steering. However, many concepts are not binary: Animals and Countries contain many subcategories, each with multiple instances. For these concepts, the search space over possible representation geometries is far larger than for binary concepts; it is thus not clear what geometries are most appropriate, nor what methods are most effective at recovering them. In this work, we introduce MAxBench, a geometry-agnostic evaluation framework for multinomial concept representations based on sampling from the recovered concept representation. We use MAxBench to compare 10 localization methods (covering 5 geometry types) across 6 concepts and 4 models. Using this framework, we find that (i) affine subspaces steer more reliably and have greater recall than rank-one or linear subspaces; (ii) much of this advantage is due to better non-zero offsets rather than the choice of bases; (iii) manifold steering is competitive with the best methods when applicable; and (iv) no method consistently outperforms prompting, in alignment with prior findings on binary concepts. These findings underscore the importance of expanding the scope of interpretability research and meta-evaluation to concepts with more varied structure.

View source

Similar papers

#machine learning Preprint Sep 2026

PermuFormer: Multi-Task Pretraining for Permutation Representation in Algebraic Combinatorics

Diverse pretraining has been shown to be an effective method for learning reusable, domain-aware representations that provide a starting point for fine-tuning on downstream tasks. While much of the excitement in AI for math has been concentrated in the use of frontier reasoning models to solve well-specified problems t...

Henry Kvinge · 0 citations
Preprint Oct 2026

LLM Benchmarking via Representation Multi-task Learning

Quantifying and evaluating the capabilities of Large Language Models (LLMs) remains a fundamental challenge in modern data science and artificial intelligence. In this paper, we consider LLM evaluation based on their performance across items in multiple benchmark domains (e.g., mathematical reasoning and coding) within...

Yuqing Xie, Yuxuan Xu, Yang-Yang Feng et al. · 0 citations
#machine learning Preprint Aug 2026

Correlation-Aware Structured Pruning for Large Language Models

Structured pruning is a promising approach for reducing the substantial inference costs of Large Language Models (LLMs) while maintaining hardware efficiency. Many existing methods assess the importance of prunable units (e.g., channels or heads) in isolation, implicitly assuming that pruning errors are additive. This...

Si-Cheng Xu, Hao Shi, Wei Zhang et al. · 0 citations
#machine learning Preprint Sep 2026

DE-Venus: A Data-Efficient RLVR Framework for Large Language Models

Reinforcement learning with verifiable rewards (RLVR) improves large language model reasoning, but its practical scaling is constrained by expensive on-policy rollouts and the cost of obtaining reliable targets at scale. Existing methods address sample selection, incomplete supervision, or noisy labels separately, ofte...

Shen-Zhi Yang, Guang-Cheng Zhu, Kai Tang et al. · 0 citations
#artificial intelligence Preprint Aug 2026

A.X K2 Technical Report

To support long contexts efficiently, Sparse Gated Attention (SGA), which combines sparse attention with gated attention, and adopt Gated Norm (GN) to stabilize large-scale training is introduced, which keeps 4-bit NVFP4 serving within one point of FP8 accuracy.

Cheolseung Baek, Dhammiko Arya, Eunki Kim et al. · 0 citations
#artificial intelligence Preprint Oct 2026

When Are Concept Bottleneck Model Explanations Faithful and Compact?

Concept bottleneck models (CBMs) are neural classifiers that allow to explain their decisions via high-level concepts, potentially enabling understanding, steering and debugging. However, their explanations are often derived heuristically. Building on formal explainability, we argue they should also be faithful, i.e.,...

Stefano Teso, E. Marconato, Steve Azzolin et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 29, 2026

Who we become when we talk to machines

Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.

MIT News · Artificial Intelligence Aug 27, 2026

Looking beyond natural sequences

A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.