Skip to content

Steering Neural Network Training through Interpretable Constraints Based on Partial Dependence

Jul 2026 · arXiv.org · Vol abs/2607.08641 · 0 citations · 40 references
Computer Science

TL;DR

This work introduces a new approach to steering neural networks based on partial dependence, such that their average response to certain features aligns with specific functional domain knowledge about the problem.

Abstract

Over the last few years, there has been an increased interest in making machine learning models more interpretable. Although a great deal of effort goes into developing techniques for interpreting the interactions learned by a given model, fewer studies focus on assessing the quality of such explanations. Even fewer focus on how to adjust the model to produce explanations faithful to prior knowledge, a process known as explanation-guided learning. Furthermore, most approaches in this area focus on classification problems and usually assume prior knowledge about which input features or regions are most important. In this work, we introduce a new approach to steering neural networks based on partial dependence, such that their average response to certain features aligns with specific functional domain knowledge about the problem. We empirically demonstrate on a range of regression problems, including dynamical systems forecasting, that models whose training has been controlled using our method perform better than unconstrained models and are more data-efficient. Moreover, we highlight that interpretations obtained from the former actually align with the user-provided knowledge, whereas those obtained from the latter do not.

View source

Similar papers

Jul 2026

EXPLAINING UNCERTAINTY ESTIMATES BASED ON GENERATIVE MODELS

The need to interpret the predictions obtained by machine learning models has become ever more important over the last two decades, mainly due to the omnipotent potential of such models, particularly deep learning models, to reach unprecedented levels of accuracy and high levels of performance. Several papers have been...

Tameem Adel · 0 citations
2026

Using Craig Interpolation for Explanation of Neural Networks (Abstract)

This work introduces space explanations, a logic-based notion of explanation that represents sufficient conditions for a neural network to predict a given class over a (potentially large and geometrically complex) subset of the feature space and demonstrates that the interpolation-based explanations are more meaningful...

Faezeh Labbaf, Tomáš Kolárik, Martin Blicha et al. · 0 citations
#artificial intelligence Preprint Aug 2026

Probing Perceptual Priors of MLLMs via Gibbs Sampling with Interpretable Generative Controls

This work proposes a method to sample from models'perceptual prior distributions directly, by steering a generative model to produce stimuli along controllable axes and running Gibbs sampling over that space with the model under study as the judge, and recovers both canonical biases and surprising novel priors invisibl...

Manuel Cherep, Pattie Maes, Nikhil Singh · 1 citation
Jul 2026

Information Bottleneck Learning for Faithful Time Series Forecasting Explanations

As forecasts increasingly drive decisions in fields such as energy, transportation, and healthcare, understanding the historical data behind these predictions has become as crucial as the predictions themselves. Although existing interpretable-by-design forecasters reveal their internal structures, they offer no guaran...

Xu Zheng, Wei Cheng, Zhuomin Chen et al. · 2 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.