ABLE: Representing and Mapping LLMs via Attribution-Based Large-model Embedding
ABLE (Attribution-Based Large-model Embedding), a framework that leverages the interpretability space to construct model representations by aggregating gradient-based feature attributions via a tokenizer-agnostic word-level alignment, captures model-specific input-sensitivity patterns rather than only surface-level outputs.