Skip to content
Preprint

Quantifying Training Membership Information in the Hyperspherical Embedding Geometry of Face Recognition Models

Jul 2026 · 0 citations · 37 references
Computer Science

TL;DR

The results indicate that the number of training identities has the largest effect on member/non-member separability, while backbone and loss head contribute far less, and that, on a same-domain held-out reference, the geometric membership signal decreases monotonically as more identities are added to training.

Abstract

Face recognition models represent each face as an embedding vector on the unit hypersphere by clustering embeddings of the same identity while pushing different identities apart through angular-margin losses. Because these losses act only on training identities, non-member identities may form clusters with different geometric properties. In this paper, we quantify the magnitude of this difference and what training-time factors control it. We compute four statistics based on cluster geometry across 180 face recognition models in a factorial design over IResNet backbone size, loss head, training duration, and the number of training identities, and evaluate each configuration on nine benchmarks. Our results indicate that the number of training identities has the largest effect on member/non-member separability, while backbone and loss head contribute far less, and that, on a same-domain held-out reference, the geometric membership signal decreases monotonically as more identities are added to training. We provide an analysis of cross-domain (pose, age, quality, ethnicity) non-member benchmarks and report that these inflate the apparent membership signal. Finally, we fuse all four statistics with a learned classifier to reveal additional membership information beyond the best individual statistic.

View source

Similar papers

Preprint Jul 2026

Private Face Recognition Training Dataset Publication via Identity-Decoupled and Geometry-Preserving Face Distillation

Private Face Distillation is proposed, an identity-decoupling and geometry-preserving framework that uses Orthogonal Geometry Preservation to construct decoupled proxy identities from private identity representations while maintaining hyperspherical geometry, and Relational Topology Alignment to preserve identity relations for recognition learning.

Shuhuan Chen, Xiangyu Zhu, Weisong Zhao et al. · 0 citations
2026

DiEL: Disentangled Evolutionary Learning for Identity-Preserving Face Enhancement and Recognition

The purpose of face enhancement tasks is to improve the recognition of faces, thus adapting to diverse visualization and recognition demands. However, the performance of the majority methods is drastically degraded under extreme conditions, including large pose variations, low resolution, blur, occlusion, and illumination changes, which can distort facial geometry and identity related details. In this work, we construct a simple and effective face robust enhancement method. In particular, in order to maintain the identity consistency of the reconstructed face, an evolutionary learning framework for face disentanglement representation is proposed, in which we disentangle the identity and pose information of the face and unite it with identity recognition as a multi-objective optimization problem, where reconstruction, adversarial, and identity-preserving objectives are adaptively balanced. Further, in order to maintain the pose consistency of reconstructed faces, we construct a unified face pose dictionary, which forms a robust and standard pose representation by statistics and induction of the geometric structure of a large number of face images. In the conditional generation architecture, the pose dictionary could accurately guide the model to realize face reconstruction with desired poses. Extensive benchmark experiments on MS1M, LFW, CPLFW, CFP-FF, CFP-FP, and AgeDB show that the proposed method not only significantly outperforms state-of-the-art methods, but also can further stimulate the discrimination potential of existing face recognition models. Specifically, DiEL achieves an average improvement of 4.66% over the SOTA methods across six benchmark datasets, with particularly significant gains on challenging cross-pose benchmarks such as CPLFW and CFP-FP.

Jingwei Xin, Tian Yang, Jun Hao et al. · 0 citations
Conference Jul 2026

Synthetic-to-Authentic Data Mixing for Face Recognition: Impact of Backbone Depth and Loss Functions

Data privacy regulations have restricted access to authentic face datasets, making synthetic alternatives increasingly necessary for face recognition training. Yet their joint effect with backbone depth and loss function on verification accuracy remains poorly understood. This paper presents a three-dimensional study using unbalanced subsets of CASIA-WebFace (authentic) and DCFace (synthetic), without demographic correction. The synthetic-to-authentic ratio is varied from 0 to 30 identities, across three backbone depths (ResNet34, ResNet50, ResNet100) and two loss functions (CosFace, Arc-Face), evaluated on six standard face verification benchmarks. Three results are reported: (i) synthetic and authentic unbalanced data produce equivalent verification accuracy at all backbone depths, while accuracy decreases with backbone depth under low-data conditions; (ii) the optimal synthetic proportion varies with backbone depth, shifting from 5 identities for ResNet34 and ResNet50 to 15 for ResNet100; (iii) CosFace outperforms ArcFace on synthetic-only data at all depths, while ArcFace outperforms CosFace on combined data for ResNet100, achieving the highest accuracy across all configurations. These results show that backbone depth, mixing ratio, and loss function interact and should not be optimized independently.

Mohamed Mehdi Ayeche, Marouane Ben Haj Ayech, Lotfi Chaouech et al. · 0 citations
#machine learning Preprint Aug 2026

Picture the Epsilon: Pursuing Identity-Level Privacy Guarantees for Images

A comparative study of four audits applicable to pre-trained, black-box face generators, which consistently reveal substantial identity distinguishability while reporting markedly different epsilon estimates that reflect each method's distinct assumptions and finite-sample treatment.

Arman Zareian Jahromi, Vishnu Bondalakunta, Mohammad Akbar Bin Shah et al. · 0 citations
Book Open access Aug 2026

Geometric Space, Architecture and Learning Objective for Large Pre-Trained Models

Large pretrained models have reshaped artificial intelligence, yet their Euclidean design assumptions often limit their ability to model hierarchy, curvature, symmetry, and heterogeneous relations in real-world data. The Geometric Space, Architecture and Learning Objective for Large Pre-Trained Models (GALOP) workshop is an accepted half-day KDD 2026 workshop that examines how geometric principles can make large pretrained models more expressive, robust, interpretable, and efficient. The workshop is organized around three complementary themes: (1) Geometric Space, which studies non-Euclidean representation spaces such as hyperbolic, spherical, and mixed-curvature manifolds; (2) Geometric Architecture, which designs model architectures that respect data symmetries, manifold structure, and relational geometry; and (3) Geometric Learning Objective, which develops objectives and optimization methods that preserve distances, angles, curvature, and topology during training. Through two invited talks and four contributed talks, the workshop brings together researchers from machine learning, data mining, and related fields to advance geometrically-informed foundation models for language, vision, graphs, knowledge discovery, and scientific discovery.

Menglin Yang, Jiahong Liu, Lucas Vinh Tran et al. · 0 citations
Jul 2026

Emergent Region-Level Facial Correspondence in Frozen Vision Foundation Models

These results establish frozen DINOv3 as a strong zero-shot representation for region-level facial correspondence and identify intermediate self-supervised features as the most useful layer for dense face analysis.

Izaldein Al-Zyoud, Abdulmotaleb El Saddik · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.