Masking-Guided Structure and Texture Decoupling for Lightweight Blind Screen Content Image Quality Assessment
Abstract
Screen content images (SCIs) exhibit complex structural heterogeneity, rendering traditional statistics-based natural scene image quality assessment (NR-IQA) metrics ineffective. Although deep learning models achieve high prediction accuracy, their prohibitive computational demands preclude deployment in latency-sensitive industrial scenarios. While existing handcrafted lightweight SCI-IQA metrics reduce computational overhead, most rely on unsegmented global feature pooling or holistic edge statistics (e.g., edge histograms or Fisher vector coding), thereby diluting locally critical text-edge distortions in vast homogeneous backgrounds. To address this limitation, we propose an ultra-lightweight, deep-learning-free NR-IQA framework centered on human visual masking. Unlike existing lightweight methods, our approach explicitly employs dual-scale Canny edge operators to partition SCIs into edge-sensitive and flat background regions. Guided by this visual prior, structural degradations and micro-compression textures are extracted region-wise using Sobel gradients and uniform local binary patterns (LBPs) and aggregated with global Commission Internationale de I’Eclairage L*a*b*(CIELAB) color statistics into a compact 60-dimensional descriptor. A grid-search-optimized Support Vector Regression (SVR) maps these features to subjective quality scores. Extensive cross-validation on the SIQAD and SCID datasets demonstrates that our metric outperforms existing handcrafted lightweight SCI metrics and traditional NSS models, while achieving accuracy competitive with representative full-reference metrics. Consuming only 79.3 ms per image on a standard CPU, it offers a practical accuracy–efficiency trade-off for resource-constrained periodic quality monitoring.