DIF formulates the opti- mality conditions of constrained learning as a variational inequality (VI) and ana- lyzes how perturbing training data affects this VI, establishing DIF as an efficient and reliable tool for data attribution in constrained learning.
Abstract
As constrained learning becomes increasingly common, models are trained under explicit feasibility requirements to enforce fairness, safety, robustness, regulariza- tion, and physics or logic constraints. Understanding how training samples in- fluence the model solution (e.g., learned parameters) is crucial for interpretability and robustness. The classical influence function (IF) estimates sample contribu- tions via local sensitivity analysis, measuring how the solution changes when a specific training sample is perturbed or removed. However, IF becomes unreli- able in constrained settings: data perturbations can reshape both the objective and the feasible region, leading to estimates that violate feasibility. In response, we propose the Directional Influence Function (DIF), a novel estimator that explicitly incorporates these constraints into influence estimation. DIF formulates the opti- mality conditions of constrained learning as a variational inequality (VI) and ana- lyzes how perturbing training data affects this VI. We validate DIF on constrained linear regression and demonstrate that it recovers leave-one-out retraining results, whereas IF and penalty-based IF exhibit significant bias. We further apply DIF to fairness-constrained CNNs, where DIF accurately predicts test loss changes under data removal and aligns closely with actual retraining. Our results establish DIF as an efficient and reliable tool for data attribution in constrained learning.
The introduction of isotonic Bellman calibration, a one-dimensional, model-agnostic post-processing method that reduces residual occupancy-balance violations while preserving the ranking information in any initial occupancy-ratio estimate, and establishes finite-sample calibration guarantees and a KL oracle inequality...
Orthogonality in a LoRA factor does not by itself specify what the composed update protects: the answer depends on the task-start state, the parameterization, and the realized optimizer displacement. We formalize this question through Bilinear Optimization Divergence (BOD), an anchor-relative diagnostic of effective-up...
Yong-Shun Wang, Jianlin Su, Yong-Yi Ma· 0 citations
Causal Structure-guided DRO (CS-DRO) is proposed, which estimates a directed acyclic graph (DAG) that encodes the predictive relationships between representations and labels, serving as a proxy for causal structure shared across source domains.
Seonggyeom Kim, Eunjung Choi, Dong-Kyu Chae· Proceedings of the 32nd ACM...· 0 citations
It is shown that the contradiction in experience with neural surrogates in derivative-free optimisation dissolves once three factors are stated, and that these, rather than the fit accuracy a training curve reports, are what delimit when a learned local model pays.
We evaluate Deep Perturbation Learning (DPL), which perturbs training images and labels along influence-derived directions, in three roles in which prior work has positioned it for machine unlearning: a direct deletion signal (the strongest claim), a utility-preserving regularizer, and a warm start for adversarial unle...
Chen Wu, Chrispine Kambimbi, Qin-Yang Zeng et al.· 0 citations
DIEM is proposed, a principled and fully automated framework that makes data utilization adaptive throughout RFT and consistently outperforms strong static and dynamic baselines.
Haoru Tan, Sitong Wu, Yan-Feng Chen et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.