This work introduces a new approach to steering neural networks based on partial dependence, such that their average response to certain features aligns with specific functional domain knowledge about the problem.
Abstract
Over the last few years, there has been an increased interest in making machine learning models more interpretable. Although a great deal of effort goes into developing techniques for interpreting the interactions learned by a given model, fewer studies focus on assessing the quality of such explanations. Even fewer focus on how to adjust the model to produce explanations faithful to prior knowledge, a process known as explanation-guided learning. Furthermore, most approaches in this area focus on classification problems and usually assume prior knowledge about which input features or regions are most important. In this work, we introduce a new approach to steering neural networks based on partial dependence, such that their average response to certain features aligns with specific functional domain knowledge about the problem. We empirically demonstrate on a range of regression problems, including dynamical systems forecasting, that models whose training has been controlled using our method perform better than unconstrained models and are more data-efficient. Moreover, we highlight that interpretations obtained from the former actually align with the user-provided knowledge, whereas those obtained from the latter do not.
The need to interpret the predictions obtained by machine learning models has become ever more important over the last two decades, mainly due to the omnipotent potential of such models, particularly deep learning models, to reach unprecedented levels of accuracy and high levels of performance. Several papers have been...
Tameem Adel· International Journal of Art...· 0 citations
A curious phenomenon called mode connectivity, the ability to connect neural networks in the loss surface, defies explanation entirely is elucidates, explains and exploits this special structure in the loss landscape.
This work introduces space explanations, a logic-based notion of explanation that represents sufficient conditions for a neural network to predict a given class over a (potentially large and geometrically complex) subset of the feature space and demonstrates that the interpolation-based explanations are more meaningful...
Faezeh Labbaf, Tomáš Kolárik, Martin Blicha et al.· CI-BD-SOQE@FLoC· 0 citations
This work proposes a method to sample from models'perceptual prior distributions directly, by steering a generative model to produce stimuli along controllable axes and running Gibbs sampling over that space with the model under study as the judge, and recovers both canonical biases and surprising novel priors invisibl...
Manuel Cherep, Pattie Maes, Nikhil Singh· 1 citation
As forecasts increasingly drive decisions in fields such as energy, transportation, and healthcare, understanding the historical data behind these predictions has become as crucial as the predictions themselves. Although existing interpretable-by-design forecasters reveal their internal structures, they offer no guaran...