Linear Vision Transformers (ViTs) are designed to replace the attention in Softmax ViTs with the linear-complexity attention operator for more efficient token routing, but they require from-scratch pre-training and typically underperform the original Softmax version. How to initialize linear ViTs both efficiently and e...
Huai-Yuan Qin, Mu-Li Yang, Gabriel James Goenawan et al.· 0 citations
This work proposes CADE (Contrastive Adaptive Debias Ensemble), a training-free, plug-and-play method that leverages modality-specific answer priors that yields significant gains on the proposed benchmark, which can foster the development of more fair and reliable AI systems for sustainable development.
Zihang Lin, Huaiyuan Qin, Mu Yang et al.· arXiv.org· 0 citations
DiD is introduced, a label-free conversion method that exclusively trains the linear-attention backbone by aligning detector-facing interface tensors with those of a frozen Softmax teacher, and substantially outperforms established baselines and matches supervised, fully trained linear models.
Huai-Yuan Qin, Gabriel James Goenawan, Zihang Lin et al.· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.