Teaching Data Cleaning in Machine Learning: A Misconception-Driven Pedagogical Framework for Explainable and Ethical Data Preprocessing
Data cleaning — the process of detecting and remedying missing, erroneous, and inconsistent records in raw datasets — is widely recognized as the most time-consuming and consequential stage of the machine learning preprocessing pipeline. Yet despite its operational centrality, how data cleaning should be taught to unde...