Skip to content

Leakage-Aware LLM Augmentation for Attrition Prediction: A DecisionCentric Evaluation

Sep 2026 · Artificial Intelligence and Internet Studies · 0 citations · 30 references
AI and HR Technologies

TL;DR

A leakage-aware, dual-network augmentation framework that integrates the semantic reasoning of Large Language Models with the distributional rigorousness of GAN discriminators is proposed, offering HR practitioners a robust, decision-centric tool to identify at-risk employees without sacrificing transparency.

Abstract

Employee attrition is a high-cost, asymmetric decision problem plagued by data imbalance. To address this, we propose a leakage-aware, dual-network augmentation framework that integrates the semantic reasoning of Large Language Models (LLMs) with the distributional rigorousness of GAN discriminators. Specifically, we employ an autoregressive transformer (GReaT) as a generative "actor" to synthesize diverse minority samples, coupled with a post-hoc "critic" derived from a conditional GAN to strictly filter low-fidelity outliers. This actor-critic inspired loop ensures that synthetic records broaden coverage without introducing noise. We evaluate this pipeline using a rigorous protocol where all generation and filtering occur strictly within training folds to prevent leakage. Experiments on HR datasets show that moderate, critic-guided oversampling yields significant recall gains (e.g., raising recall from 45% to 53%) and improved F₁ scores compared to baselines, while maintaining calibration and discrimination (stable ROC-AUC). Furthermore, the approach preserves interpretability (SHAP) and fairness (low TPR gaps), offering HR practitioners a robust, decision-centric tool to identify at-risk employees without sacrificing transparency.

Read PDF

Similar papers

#machine learning Preprint Sep 2026

Beyond Homoscedasticity: Decoupled Uncertainty Optimization for Deep Imbalanced Regression

Deep Imbalanced Regression (DIR) is pervasive in continuous prediction tasks across diverse modalities, such as age estimation, depth prediction, and protein mutation activity prediction, where label-scarce tail samples often carry higher practical value. However, most existing methods still learn deterministic point m...

Jun-Chen Zhou, Jia-Xi Lu, Wei-Jing Zeng et al. · 0 citations
Open access 2026

MSTabVAE: Multi-Step Latent Conditional Variational Autoencoder for Imbalanced Tabular Data Synthesis

MSTabVAE, a novel generative framework that extends TabNet, a deep learning architecture for tabular data, into a conditional variational autoencoder (CVAE) framework, and introduces a multi-step latent mapping strategy to capture complex feature relationships in heterogeneous tabular data.

Min-Ji Kang, Hyeryung Jang · 0 citations
Book Open access Aug 2026

NiWo: An Augmentation Framework to Enhance ML Performance and Interpretability for Tabular Data with Class Imbalance

NiWo optimizes the weights of influential neighborhood instances within an augmentation budget, thus preserving computational efficiency and offering interpretability, and outperforms other augmentation methods at enhancing ML performance, especially over datasets with class imbalance and scarce instances.

Asif Ahmed, Sakhawat Hossain Saimon, Jianhua Ruan et al. · 0 citations
#machine learning Preprint Sep 2026

SAGE: Subpopulation-Aware Generative Enhancement for Mitigating Spurious Correlations

Subpopulation-Aware Generative Enhancement (SAGE), a two-stage generative augmentation framework, is introduced, using cluster-derived sub-labels and class labels to fine-tune a conditional generative model and text encoder, generating targeted synthetic data to fill underrepresented regions in the training set and con...

Yi-Ming Luo, Rong-Qiang Zhao, Jie Liu · 0 citations

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.