Human Factors in Data Preparation
Abstract
Data preparation is a cornerstone of the data science lifecycle, yet it remains one of the most time-consuming and human-dependent stages. Human factors are critical here, shaping how users interact with tools, perceive data quality, and make trade-offs under uncertainty. At the crossroads of Human–Computer Interaction (HCI), Visualization, and human-in/over-the-loop paradigms, recent research highlights new opportunities to make data preparation more usable, transparent, and trustworthy. Interactive interfaces and visual analytics systems allow users to iteratively refine transformations, explore provenance, and detect anomalies, while cognitive aspects such as workload, bias, and trust determine the reliability of human interventions. Automation increasingly complements these workflows: human-in-the-loop methods leverage user feedback to guide cleaning, labeling, or schema alignment, and over-the-loop frameworks provide oversight mechanisms to prevent over-reliance on automation. At the same time, modern generative AI introduces powerful new capabilities, such as automatically suggesting transformations, synthesizing metadata or documentation, and even generating realistic synthetic datasets. These advances promise efficiency and creativity but raise concerns about reliability, explainability, and provenance, reinforcing the need for careful human oversight. This chapter reviews these possibilities and the research challenges in developing a new class of hybrid, adaptive data-preparation solutions that blend human expertise with automated intelligence. Key challenges include balancing scalability with user effort, ensuring transparency in automated suggestions, preventing loss of situational awareness, and developing evaluation metrics that reflect both human experience and technical accuracy. Addressing these challenges will enable more robust, human-aware data preparation practices across domains.