When Does Self-Supervised Pretraining Help Tabular Models? A Study of Label Scarcity and Missing Data
While SSL outperforms training from scratch on average and remains competitive with state-of-the-art tree ensembles, the SSL-vs-scratch gains exhibit high inter-task variance and lack significance, indicating the findings reflect general properties of tabular SSL rather than idiosyncrasies of one particular pretext task.