Jul 2026
Sparse Inter-Layer Dependencies of Transformer FFN Neurons
A training-free attribution method that estimates the relative influence of upstream neurons and attention outputs on a target neuron's activation and identifies candidate sparse pathways with potential implications for efficient inference is introduced.
Johannes Knittel, Hanspeter Pfister
· arXiv.org · 0 citations