Skip to content

Directional Influence Function: Estimating Training Data Influence in Constrained Learning

Jul 2026 · arXiv.org · Vol abs/2607.23388 · 0 citations · 19 references
Computer Science

TL;DR

DIF formulates the opti- mality conditions of constrained learning as a variational inequality (VI) and ana- lyzes how perturbing training data affects this VI, establishing DIF as an efficient and reliable tool for data attribution in constrained learning.

Abstract

As constrained learning becomes increasingly common, models are trained under explicit feasibility requirements to enforce fairness, safety, robustness, regulariza- tion, and physics or logic constraints. Understanding how training samples in- fluence the model solution (e.g., learned parameters) is crucial for interpretability and robustness. The classical influence function (IF) estimates sample contribu- tions via local sensitivity analysis, measuring how the solution changes when a specific training sample is perturbed or removed. However, IF becomes unreli- able in constrained settings: data perturbations can reshape both the objective and the feasible region, leading to estimates that violate feasibility. In response, we propose the Directional Influence Function (DIF), a novel estimator that explicitly incorporates these constraints into influence estimation. DIF formulates the opti- mality conditions of constrained learning as a variational inequality (VI) and ana- lyzes how perturbing training data affects this VI. We validate DIF on constrained linear regression and demonstrate that it recovers leave-one-out retraining results, whereas IF and penalty-based IF exhibit significant bias. We further apply DIF to fairness-constrained CNNs, where DIF accurately predicts test loss changes under data removal and aligns closely with actual retraining. Our results establish DIF as an efficient and reliable tool for data attribution in constrained learning.

View source

Similar papers

Preprint Aug 2026

Bellman Calibration for Marginalized Importance Weighting in Offline Reinforcement Learning

The introduction of isotonic Bellman calibration, a one-dimensional, model-agnostic post-processing method that reduces residual occupancy-balance violations while preserving the ranking information in any initial occupancy-ratio estimate, and establishes finite-sample calibration guarantees and a KL oracle inequality...

L. van der Laan, Nathan Kallus · 0 citations
#machine learning Preprint Sep 2026

Bilinear Optimization Divergence: Diagnosing Factor-Constrained LoRA Continual Learning

Orthogonality in a LoRA factor does not by itself specify what the composed update protects: the answer depends on the task-start state, the parameterization, and the realized optimizer displacement. We formalize this question through Bilinear Optimization Divergence (BOD), an anchor-relative diagnostic of effective-up...

Yong-Shun Wang, Jianlin Su, Yong-Yi Ma · 0 citations
Book Open access Aug 2026

Causal Structure-guided Distributionally Robust Optimization under Domain Shifts

Causal Structure-guided DRO (CS-DRO) is proposed, which estimates a directed acyclic graph (DAG) that encodes the predictive relationships between representations and labels, serving as a proxy for causal structure shared across source domains.

Seonggyeom Kim, Eunjung Choi, Dong-Kyu Chae · 0 citations
#software testing Preprint Aug 2026

Why and When Neural Networks Improve Local Approximation in Optimization

It is shown that the contradiction in experience with neural surrogates in derivative-free optimisation dissolves once three factors are stated, and that these, rather than the fit accuracy a training curve reports, are what delimit when a learned local model pays.

Cheng Bian, Peng-Cheng Xie · 0 citations
#artificial intelligence Preprint Sep 2026

Do Influence-Derived Data Perturbations Enable Machine Unlearning? A Controlled Study of Three Plausible Roles

We evaluate Deep Perturbation Learning (DPL), which perturbs training images and labels along influence-derived directions, in three roles in which prior work has positioned it for machine unlearning: a direct deletion signal (the strongest claim), a utility-preserving regularizer, and a warm start for adversarial unle...

Chen Wu, Chrispine Kambimbi, Qin-Yang Zeng et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.