Computational prediction of crystal properties plays a pivotal role in materials science. With the accelerated progress in machine learning, crystal property prediction has seen remarkable advancements. Nevertheless, the utilization of machine learning in this context faces several challenges. First, existing methods that utilize the smallest repeatable unit cell of a crystal often have a limited receptive field. Second, as experiments measuring crystal properties are time-consuming, labeled data is often scarce. To address these challenges, we propose a
S
elect
I
ve mu
L
ti-
V
iew representation
A
ugmentation framework (SILVA) for crystal property prediction. To go beyond limited receptive fields, we introduce the notion of a crystal supercell, which enables more comprehensive explorations of crystal structure. To fully combine insights from multi-view structures, i.e., from unit cells and supercells, we propose a multi-view representation learning (MRL) module that features a representation space that enhances the learning of representative features specific to different views. To alleviate the limited availability of labeled data, we propose a selective representation augmentation (SRA) module. Given representations of labeled training data, we carefully select nearby representations in the representation space established by the MRL module so that labels can be reused. An experimental study offers evidence that SILVA is capable of outperforming state-of-the-art methods.
Haomin Yu, Jilin Hu, K. Tolborg et al.· ACM Transactions on AI for S...· 0 citations
We present a simple Gaussian approximation to the finite-sample distribution of the classical ridge regression estimator. Our approximation captures the fact that, in finite samples, the ridge regression estimator trades off bias and variance to reduce estimation and prediction error. Our approximation is based on nonstandard asymptotics where $i)$ we let the estimator's regularization parameter grow proportionally to the sample size; and $ii)$ we treat the population regression coefficients as \emph{local} to the reference vector that defines the estimator's direction of shrinkage. In contrast to other asymptotic approximations in the literature, we allow for general forms of heteroskedasticity and autocorrelation in the data generating process (at the cost of considering a low-dimensional model where the number of covariates is not allowed to grow with the sample size). We use our simple Gaussian approximation to propose two new strategies to select the regularization parameter for the ridge regression estimator. The suggested strategies select the regularization parameter to minimize either average or worst-case excess prediction risk, where risk is computed using our suggested Gaussian approximation.
J. M. Olea, Ryan Strong, Amilcar Velez et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.