Compact Occlusion-Robust Facial Expression Recognition via Clean-Anchored Hard Occlusion Fine-Tuning
Abstract
Facial occlusion removes expression-relevant evidence and remains a major source of error in camera-based affective sensing. Existing approaches to occlusion-robust facial expression recognition often rely on specialized attention, reconstruction, semantic, or geometric pipelines, whereas aggressive training using synthetic occlusions may impair discrimination on clean images or overfit to synthetic corruption patterns. We address this tension between cleanness and robustness through clean-anchored hard occlusion fine-tuning (CA-HOFT). HardMix samples structured and random occlusion modes according to a facial region-weighted distribution. An explicit classification branch for the non-HardMix source view preserves ground-truth supervision, while a fixed reference teacher initialized from the preceding mixed-occlusion stage supplies a stationary distribution for both paired student views. These signals jointly train a single classifier while retaining a single-backbone inference pathway. Across five independent training runs, the ResNet-18 student achieves 90.08% accuracy on the Real-World Affective Faces Database (RAF-DB) and 87.38% accuracy with 83.34% macro-F1 on Occlusion-RAF-DB, with corresponding sample standard deviations of 0.21, 0.12, and 0.32 percentage points. The retained model has 11.18 million parameters and requires 1.814 giga multiply–accumulate operations (GMACs). Controlled comparisons and ablations indicate that it has the most favorable observed cleanness–robustness trade-off among the tested epoch-matched alternatives; however, fixed-checkpoint comparisons on Occlusion-RAF-DB are not significant after Holm correction. AffectNet-8 and evaluations using natural occlusion provide supporting evidence, while broader cross-domain validation remains future work.