Synthetic-to-Authentic Data Mixing for Face Recognition: Impact of Backbone Depth and Loss Functions
Abstract
Data privacy regulations have restricted access to authentic face datasets, making synthetic alternatives increasingly necessary for face recognition training. Yet their joint effect with backbone depth and loss function on verification accuracy remains poorly understood. This paper presents a three-dimensional study using unbalanced subsets of CASIA-WebFace (authentic) and DCFace (synthetic), without demographic correction. The synthetic-to-authentic ratio is varied from 0 to 30 identities, across three backbone depths (ResNet34, ResNet50, ResNet100) and two loss functions (CosFace, Arc-Face), evaluated on six standard face verification benchmarks. Three results are reported: (i) synthetic and authentic unbalanced data produce equivalent verification accuracy at all backbone depths, while accuracy decreases with backbone depth under low-data conditions; (ii) the optimal synthetic proportion varies with backbone depth, shifting from 5 identities for ResNet34 and ResNet50 to 15 for ResNet100; (iii) CosFace outperforms ArcFace on synthetic-only data at all depths, while ArcFace outperforms CosFace on combined data for ResNet100, achieving the highest accuracy across all configurations. These results show that backbone depth, mixing ratio, and loss function interact and should not be optimized independently.