Skip to content
Preprint

P3CA: Encoder-Agnostic Interpretation of Vision Foundation Model Embeddings via Spatial Probing

Aug 2026 · 0 citations · 22 references
Computer Science

TL;DR

Position-prompted PCA (P3CA), an encoder-agnostic method for local probing of channel-rich spatial tensors, is proposed and implemented in EmbedVision, an interactive 3D Slicer-based workflow, and evaluated across natural images, colorectal pathology foundation-model embeddings, and spatial transcriptomic tensors.

Abstract

Vision foundation models are increasingly used as reusable encoders in medical image computing, yet their high-dimensional spatial embeddings are difficult to inspect beyond downstream task performance or global dimensionality reduction. We propose position-prompted PCA (P3CA), an encoder-agnostic method for local probing of channel-rich spatial tensors. Given a user-selected spatial prompt, P3CA estimates the feature normalization and dominant covariance directions within that region, then applies the resulting projection to the full tensor to visualize where locally informative directions are expressed. This produces a region-conditioned representation lens without modifying the encoder, retraining, or requiring task-specific labels. We implement P3CA in EmbedVision, an interactive 3D Slicer-based workflow, and evaluate it across natural images, colorectal pathology foundation-model embeddings, and spatial transcriptomic tensors. Across these settings, prompted projections reveal local structure suppressed by global PCA, improve prompt-matched pathology discrimination from frozen three-dimensional projections, and support comparison between learned and measured spatial representations.

View source

Similar papers

Preprint Aug 2026

Representation-driven Endoscopic Visual Embedding Alignment for Latent Generation

Developing foundation generative models for endoscopy is limited by the gap between natural and clinical images and the computational cost of training large Diffusion Transformers. Although representation alignment has improved efficiency in general computer vision, its role within the highly specialized endoscopic ima...

Francisco Caetano, T. Jaspers, Haiko Middeljans et al. · 0 citations
#machine learning Preprint Aug 2026

EXPOSE: Explainable and Domain-Robust Embeddings from Pathology Vision Foundation Models using Sparse Autoencoders

This work proposes Explainable Probing of Cross-Domain Sparse Embeddings (EXPOSE), a framework that uses Sparse Autoencoders (SAEs) as an explainable bottleneck to identify and suppress domain-specific components in VFM embeddings.

Anja Witte, M. Lennartz, Jan Baumbach et al. · 0 citations
Preprint Sep 2026

Isotropic Embedding Perturbations for Robust Vision Language Encoders

Aether is introduced, a simple plug-in method that applies diffusion-style random perturbations in the embedding space via controlled alpha-mixing, specifically designed to provide isotropic regularization that remains semantically consistent.

Hyesong Choi, Daeun Kim, Song Park et al. · 0 citations
Open access Sep 2026

Mixture of Lie-Group Kernels for Equivariance-Inspired Feature Learning.

Deep visual models typically achieve robustness to geometric transformations through extensive data augmentation or increased model capacity, yet these empirical strategies do not guarantee explicitly equivariant or structurally constrained representations. While group-equivariant CNNs provide principled mechanisms for...

Yao-Xian Yang, Guipeng Lan, Shuai Xiao et al. · 0 citations
#artificial intelligence Preprint Sep 2026

RGSQ: Riemannian Geometry-Sensitive Quantization for Large Vision-Language Models

Riemannian Geometry-Sensitive Quantization (RGSQ), which formulates quantization as a reconstruction problem under a unified Fisher-Riemannian metric, enabling standard unimodal PTQ methods to evaluate multimodal quantization error under their original assumptions.

Zhi-Ping Wu, Dong-Dong Ren, Yang Zhou et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.