Skip to content
Review Open access

Automated Feature Engineering Techniques for Tabular Data

2023 · International Journal of Artificial Intelligence & Digital Transformation · Vol 6, pp. 01-15 · 0 citations

TL;DR

The effectiveness of automated feature engineering is proved by experimental results that show that the method can enhance the accuracy and robustness of models as well as improve generalization with respect to the multiple benchmark datasets.

Abstract

AFE is an automated feature engineering system that has become a key enabler to scalable and high-performing machine learning systems with scalable systems that run on tabular data. The conventional feature engineering makes excessive use of domain, trial and error and iterative optimization, which are computed consuming time, error prone and hard to repeat. As data-driven applications in finance, healthcare, manufacturing, and e-commerce have been exponentially increasing, there is an increasing need in automated, systematic, and reliable methods of features construction. The overall objective of automated Feature Engineering methods is to generate, transform, select, and optimize features adventurously, from raw tabular data, with minimal human intervention, and at a higher predictive efficiency. The paper contains a complete detailed analysis of automated feature engineering approaches to tabular data with references to their theoretical principles, algorithmic approaches, and real-world examples. The paper expounds on rule construction based feature construction, statistical construction, deep learning based representation, evolutionary learning, and reinforcing learning methods, and end to end AutoML. An intricate literature review shows major achievements, comparisons, and unresolved issues. The suggested methodology defines the three branches of feature generation, selection and evaluation as a single automated pipeline via mathematical representation and algorithmic processes. The effectiveness of automated feature engineering is proved by experimental results that show that the method can enhance the accuracy and robustness of models as well as improve generalization with respect to the multiple benchmark datasets. Lastly, issues of limitations, interpretability, computational trade-offs, and research directions are discussed in the paper. The given publication meets the IEEE publication standards and offers well-organized, high-quality information to a researcher or an organization practitioner dealing with tabular data analytics.

Read PDF

Similar papers

Open access Jul 2026

Bridging Scalability and Interpretability in AutoML Via Feature Engineering

Experimental evaluation on the Madelon dataset indicates that the automated and interpretable pipeline performs comparably to, and in some respects favourably against, baseline feature engineering approaches, demonstrating the practical effectiveness of combining scalable feature generation with interpretable AutoML.

CH. Vasavi, SK. Raqeeba · 0 citations
Review Open access 2026

A Survey on Feature Selection Techniques for Predictive Analytics

The feature selection is a crucial step in predictive analytics to determine which subset of features makes the most contribution to the high-dimensional data and remove irrelevant, redundant, or noisy features. The dimensionality of datasets keeps on growing, and, as contemporary data-driven applications produce large volumes of heterogeneous data, overfitting, computational complexity, worse model interpretability, and poorer generalization become issues as heterogeneous data increases. The feature selection methods are meant to address such challenges by improving predictive accuracy, minimizing training time and improving model robustness. This survey is a systematic and extensive overview of feature selection methods used in predictive analytics which are utilized in a variety of areas and fields, including healthcare, finance, bioinformatics, cybersecurity, and smart systems. In the paper, the features selection techniques have been classified as filter, wrapper, embedded, and hybrid techniques which give a comprehensive theoretical background of each of the techniques as well as a comparison of each of the techniques. Statistical, information-theoretic, similarity-based, and probabilistic filters are discussed in addition to the heuristic and metaheuristic wrapper methods, i.e. evolutionary, swarm-based etc. Also critically analyzed is embedded techniques that make use of regularization, decision trees, and ensemble learning. Moreover, this survey talks about the evaluation metrics, benchmark data, and design considerations of the experiment which are used in the evaluation of the effectiveness of the feature selection. Such practice issues as scalability, stability, data imbalance, and interpretability are mentioned, as well as new directions related to deep learning-based feature selection and multi-objective optimization and explainable artificial intelligence. This piece of work can be regarded as a useful source of information by the researcher and practitioners who want to develop effective, precise, and understandable predictive analytics systems.

Nesca Mthethwa, Thane Nkosi · 0 citations
#machine learning Preprint Aug 2026

SymboLLM-FE: LLM-Accelerated Symbolic Regression for Automated Feature Engineering on Tabular Data

This paper combines symbolic regression with LLMs for feature engineering (SymboLLM-FE) to solve the dual challenges of poor interpretability and numerous iterations by employing a statistical prior-grounded LLM refinement mechanism and single-digit LLM calls.

Zi-Jian Cheng, Zi-Yi Jia, Zhi Zhou et al. · 0 citations
Preprint Aug 2026

TS2TabPFN: Time Series Classification and Extrinsic Regression through Feature Extraction and a Tabular Foundation Model

Time series data are ubiquitous in practical applications, where classification (TSC) and extrinsic regression (TSER) have emerged as essential tasks for obtaining value from temporal sequences. While the literature has seen significant progress through feature-based and deep learning models, existing methods often focus either on the quality of feature extraction or on the intrinsic predictive power of complex architectures applied to raw data. This division creates a gap between the control offered by feature engineering and the automated performance of end-to-end models. This paper proposes TS2TabPFN, a framework that bridges this gap by integrating explicit feature extraction with TabPFN 2.5, a cutting-edge foundation model for tabular data, to leverage its predictive capabilities. Our extensive experimental evaluation demonstrates that TS2TabPFN significantly outperforms state-of-the-art models in TSER tasks with statistical significance, providing a robust and efficient alternative for TSC and surpassing most of the currently best-performing algorithms. These results suggest that combining foundation models with structured features overcomes single-paradigm limitations, establishing a new time series state-of-the-art.

G. Merlin, D. F. Silva · 0 citations
Open access 2020

Data-Centric Engineering: Integrating Simulation, Machine Learning, and Statistics

This article delves deep into the confluence of simulation, ML, and statistics, showcasing how they synergize to improve engineering workflows and emphasizes that DCE is not just a technological advancement but a foundational strategy for next-generation engineering solutions.

Benjamin Scott · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.