Training Transformer-based architectures with finite data augmentation has become an increasingly popular approach in geometric machine learning. Despite its empirical success, the interplay between the Transformer architecture, invariance to different symmetries, and augmentation budgets remains underexplored. In this...
Eduardo Santos-Escriche, Valerie Engelmayer, Ya-Wei Eileen Lin et al.· 0 citations
Training-free model merging aims to combine multiple fine-tuned models into a single model without further optimization on labeled data. Yet, in transformers, independently merging individual layers can affect a shared attention computation because the query-key and value-output operators depend on composed matrices, o...
Akansh Maurya, Y. Lin, Stefanie Jegelka et al.· 0 citations
This paper proposes Single-stage Sparse Retrieval (SSR), a paradigm shift that replaces expensive clustering with efficient sparse coding, and utilizes Sparse Autoencoder (SAE) to project token embeddings into a high-dimensional but highly sparse representation.
This paper presents the first polynomial-time algorithm for learning with exact group invariances that applies uniformly to finite and infinite groups, and shows that exact symmetries can be identified from data and exploited for learning in polynomial time.
Ashkan Soleymani, B. Tahmasebi, Patrick Jaillet et al.· Annual Conference Computatio...· 2 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.