Skip to content
Open access

MSTabVAE: Multi-Step Latent Conditional Variational Autoencoder for Imbalanced Tabular Data Synthesis

2026 · IEEE Access · Vol 14, pp. 126104-126116 · 0 citations · 36 references
Computer Science

TL;DR

MSTabVAE, a novel generative framework that extends TabNet, a deep learning architecture for tabular data, into a conditional variational autoencoder (CVAE) framework, and introduces a multi-step latent mapping strategy to capture complex feature relationships in heterogeneous tabular data.

Abstract

Recent advances in artificial intelligence have expanded its applications in the financial domain, particularly in fraud detection, a critical task for preventing losses for both customers and institutions. However, fraud detection is challenging due to severe class imbalance, which significantly degrades detection performance. Existing synthetic data generation methods for minority-class augmentation often fail to capture the heterogeneous structure and complex feature relationships of tabular data in the financial domain. To address these challenges, we propose MSTabVAE, a novel generative framework that extends TabNet, a deep learning architecture for tabular data, into a conditional variational autoencoder (CVAE) framework. MSTabVAE preserves TabNet’s step-wise feature-selection mechanism and introduces a multi-step latent mapping strategy to capture complex feature relationships in heterogeneous tabular data. Class-conditional information is further injected at every decision step to target minority-class synthesis. Experiments on four imbalanced financial datasets with two downstream classifiers show that augmenting training data with MSTabVAE-generated samples consistently improves classification performance over existing generative and traditional oversampling baselines, achieving up to a 45% relative improvement in Recall@1%FPR on the highly imbalanced BAF dataset. Additional experiments verify the fidelity of the generated data through inter-feature correlation and marginal distribution preservation, negligible privacy leakage as measured by the Distance to Closest Record metric, and up to $238\times $ speedup in generation time over diffusion-based baselines. Finally, MSTabVAE provides structural interpretability for the generation process by revealing the features selected at each decision step.

Read PDF

Similar papers

Aug 2026

SMOG: an adaptive hybrid oversampling framework using SMOTE and conditional GAN for difficulty-aware learning on imbalanced data

SMOG, an adaptive hybrid oversampling framework that integrates the Synthetic Minority Over-sampling Technique with a Conditional Generative Adversarial Network (GAN)-based difficulty-aware learning strategy, highlights the effectiveness of adaptive hybrid generative strategies for intelligent learning on imbalanced da...

Jatinder Kaur, Vimal Parmar, B. K. Rao et al. · 0 citations
Sep 2026

Information-Bottlenecked Variational Autoencoder for Top-N Recommendation via Gating Mechanism

Generative models for top- \(N\) recommendation have garnered significant attention, with Variational Autoencoder (VAE) emerging as a promising approach for modeling user preferences. Yet, traditional VAE-based models encounter two major challenges: simplistic priors may cause posterior collapse, resulting in ineffecti...

Xiao-Bo Guo, Shaoshuai Li, You-Ru Li et al. · 0 citations
#large language models Open access Sep 2026

Leakage-Aware LLM Augmentation for Attrition Prediction: A DecisionCentric Evaluation

A leakage-aware, dual-network augmentation framework that integrates the semantic reasoning of Large Language Models with the distributional rigorousness of GAN discriminators is proposed, offering HR practitioners a robust, decision-centric tool to identify at-risk employees without sacrificing transparency.

Wei-Quan Liao, Jia-You Xu, Ekaterina A. Panova · 0 citations
Open access 2026

Comparative Evaluation of Synthetic Data Generation Methods for Binary Classification Problems with Class Imbalance

Class imbalance is a critical challenge in the classification of tabular data, since it affects the diagnostic capacity of models in domains such as health and finance. This research compares four synthetic data generation paradigms: traditional interpolation (SMOTE-NC), deep generative models (CTGAN and TVAE), and a h...

Jhonatan Esquivel, Christian Humpiri, Jose M. Vega et al. · 0 citations
Conference 2026

ProReGen: Progressive Residual Generation under Attribute Correlations

ProReGen is presented, a progressive residual generation approach inspired by the classical Robinson’s transformation, to partial out from an image attribute x2 its component mx1 that is predictable by other image attributes x1, and the residual γ=x2-mx1 that is not.

Ruby Shrestha, Ajay Gopi, Casey Meisenzahl et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.