Skip to content

Controllability Analysis for Vision State Space Models: A Structural Interpretability Framework.

Aug 2026 · IEEE Transactions on Neural Networks and Learning Systems · Vol PP, pp. 1-15 · 2 citations
Medicine

Abstract

State space models (SSMs) have emerged as a compelling alternative to transformers for visual recognition, offering linear computational complexity while maintaining competitive accuracy. However, the lack of interpretability tools designed specifically for the recurrent dynamics of these models remains a significant gap: existing saliency methods originate from convolutional or attention-based architectures and do not account for the sequential state propagation that governs information flow in SSMs. We introduce controllability analysis, a framework grounded in control theory that quantifies the structural influence of each input position on the internal state dynamics of vision SSMs. We derive two complementary indices: a Jacobian-based measure that tracks sensitivity propagation through the full sequence via an efficient $O(LN)$ backward recursion, and a Gramian-based measure that admits a closed-form, fully parallelizable solution for diagonal state matrices. Unlike gradient-based attribution methods that produce class-specific explanations, controllability indices measure intrinsic model properties that are invariant to the target class. Experiments on seven image classification benchmarks spanning diverse visual domains (microscopy, optical coherence tomography (OCT), dermatoscopy, radiography, natural images, apparel, and remote sensing) demonstrate that controllability analysis produces perfectly class-agnostic explanations (cross-class correlation of 1.0 compared to -0.45 to 0.18 for Grad-CAM across datasets) and reveals layerwise information flow patterns that vary across imaging domains. On the dataset where the model is best trained (BloodMNIST, 98% test accuracy), the proposed Jacobian index also outperforms Grad-CAM on standard insertion/deletion faithfulness (Jacobian $0.63 {\,}\pm {\,}0.02$ versus Grad-CAM $0.47 {\,}\pm {\,}0.03$ , mean ± standard deviation across three training seeds, with nonoverlapping per-seed 95% bootstrap CIs); on the lower accuracy datasets Grad-CAM remains stronger on this class-specific metric, a gap we discuss as a diagnostic of incomplete alignment between the model's internal dynamics and the classification task rather than a failure of the structural framework. A complementary internal-attention IoU metric, measuring alignment between each method's saliency and the model's own L2-magnitude attention pattern, shows controllability outperforming Grad-CAM on six of the seven datasets. Directional decomposition of the controllability profiles further shows that influence progressively sharpens with depth, transitioning from diffuse responses in early layers to sparse, high-magnitude peaks in deep layers that align with semantically meaningful structures. The proposed framework provides a principled foundation for understanding, debugging, and improving vision SSMs across application domains.

View source

Similar papers

Jul 2026

On the Identifiability of Controlled World Models

A joint identifiability condition for controlled world models with Gaussian latent states with Gaussian latent states is presented, which consists of two coupled components: representation identifiability and transition identifiability, and it is proved that when this condition holds, minimizing the LeJEPA-style predictive objective can recover both latent states and controlled dynamics in the sense of orthogonal transformation.

Xiangteng Zhang, Yang Guan, Bo Zhang et al. · 0 citations
Jul 2026

How are linear representations learned? Exact solutions to the dynamics of abstraction

In artificial and biological neural networks, concepts are often encoded as consistent linear directions in representation space. In deep learning, this idea is known as the linear representation hypothesis and underpins many interpretability and control methods based on linear probes, from concept detection to activation steering. Yet while prior work has studied whether such directions should exist $\textit{after}$ training, the dynamics of how they emerge $\textit{during}$ training remain poorly understood. Here, we develop a framework to study the alignment of concept directions during training - a process we call"abstraction". In a minimal linear network setting, we obtain exact solutions for the full trajectory of abstraction. These solutions reveal key analytic principles governing abstraction: (i) data and target geometry jointly determine abstraction at the end-of-learning, (ii) abstraction improves with network depth, and (iii) initialization scale controls the maximum abstraction reached during training. Extending our theory to nonlinear networks, we analyze how the choice of nonlinearity affects abstraction dynamics: erf networks approximate the linear theory, while abstraction in ReLU networks depends less on target geometry and more on input geometry. Across both, we prove a striking attenuation law: both nonlinearities weaken abstraction in activations relative to preactivations. We find evidence for this law in open models (DINOv3, Gemma 4) and apply our theory to improve linear probe generalization in LLMs. Together, our results provide a dynamical theory of abstraction with implications for interpretability and control.

William Yang, Andrew M. Saxe, Peter E. Latham · 0 citations
Jul 2026

VisualPatchWorld: Code World Models as Latent Structured Representations for Planning

VisualPatchWorld is introduced, which represents world dynamics as code and first selects a qualitative dynamical form with short active probes, then fits that form's free parameters from recorded state-action traces by minimizing multi-step prediction error.

Jiaxin Bai, Jia–Jie Xiong · 0 citations
Jul 2026

The Geometry of Memorization: Finite-Time Spectral Sensitivity as a Diagnostic for Flow Matching Models

Continuous-time generative frameworks construct probability paths between base and target domains by optimizing time-dependent velocity fields. While theoretical targets favor straight trajectories, empirical networks develop complex path deformations. This paper presents the Finite-Time Spectral Sensitivity (FTSS) g(t), a gradient-free, forward-pass metric that exposes flow geometry by tracking the root-mean-square singular value of the state-transition matrix. Serving as a continuous proxy for stable rank, g(t) reveals a distinct geometric pathology under data scarcity: while generalizing models maintain stable effective dimensions, overfitting causes a spectral collapse. We leverage this structural phenomenon to develop an internal geometric audit based on g(t). Our framework detects generative memorization using purely internal trajectory dynamics, removing the need for external membership queries or baseline data comparison.

Shu-Chan Wang · 0 citations
Jul 2026

A modular state-space model of human perception, cognition, and decision dynamics

Human-centered adaptive systems require behavioral models that are both psychologically interpretable and mathematically analyzable. Many existing predictors either operate as black-box input-output mappings or provide limited access to latent internal dynamics. This paper addresses this gap by modeling behavior as a perception-cognition-decision pipeline. We propose a modular state-space model in which attentional selection, predictive inference, cognitive-state evolution, intention formation, and action selection are represented by coupled mathematical mappings. The model links sensory inputs to observable behavior through latent internal states while retaining interpretable connections to neuro-cognitive mechanisms. We establish sufficient conditions for boundedness, Lipschitz regularity, forward invariance, contraction of perceptual inference under constant input, and input-to-state stability of the cognitive state dynamics. Numerical sensitivity analyses show that the model yields interpretable changes in perceptual tracking, cognitive amplification, intention expression, and action decisiveness. We further demonstrate a closed-loop rehabilitation case study in which a receding-horizon controller uses the model to adapt movement difficulty from partial feedback. In this proof-of-concept setting, the model-based controller sustains simulated task participation and achieves lower realized cumulative cost than target-following and random baselines. Overall, the framework provides a white-box dynamical structure for estimation, validation, and model-based control in human-centered settings.

Sven Schoonebeek, C. Cenedese, A. Jamshidnejad · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.