This work systematically pre-training and evaluating on many diverse datasets and analyzes what aspects of the data are most important for building a Tabular Foundation Model (TFM) generalizing across domains to show that the number and quality of tasks one can construct from a dataset is key to downstream performance.
Junwei Ma, Nour Shaheen, Alex Labach et al.· arXiv.org· 4 citations
This work adopts an alternate approach, sticking with row-based attention while incorporating long context pre-training to eliminate the need for retrieval in TabDPT-Turbo, a model that provides comparable default performance to TabDPT v1.1 on TabArena-Lite, CC18, and CTR23, at orders of magnitude faster.
Rasa Hosseinzadeh, Alex Labach, Zexin Xue et al.· 4 citations
A task-centric, retrieval-based perspective is offered for how TFMs generalize: it is believed that tabular in-context generalization is largely retrieval-based, and good models are those that learn to identify relevant examples in the provided context and aggregate them well.
Nour Shaheen, Junwei Ma, Alex Labach et al.· 3 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.