DenseFace: Bias Mitigation in Face Recognition via Density-Aware Probabilistic Matching
Authors & Institutions
Mansur Bultygov
VisionLabs
Vadim Seliutin
Amazon Web Services
VisionLabs (work performed there)
Dmitry Nekhaev
VisionLabs
Ivan Laptev
Mohamed bin Zayed University of Artificial Intelligence
What Problem It Solves
DenseFace asks whether racial false-match disparities can be reduced at comparison time, without changing or retraining the underlying face encoder. It also replaces subgroup accuracy standard deviation with a deployment-oriented protocol: set one threshold at a Caucasian false-positive rate of 0.001 and measure how far every other cohort moves from parity.
Key Result
Across CosFace and AdaFace models trained with Glint360K, MS1MV2, WebFace4M, WebFace12M, and BUPT-BalancedFace, the same global threshold becomes substantially more even. On the BUPT model, the African cohort's false-positive multiple falls from 6.15 to 1.43 and the Indian cohort from 4.01 to 1.03 relative to the Caucasian reference. Conventional verification accuracy is preserved: for a CosFace Glint360K model, average RFW accuracy rises from 98.61% to 98.68% while subgroup standard deviation falls from 0.53 to 0.44. The regressor adds about 0.2% memory and 1.75% end-to-end latency in the authors' GPU test; its optimized CPU score calculation is about 1.5 times raw cosine cost. Results still depend on a representative anchor or regressor training population.
Abstract
Despite steady progress in face recognition, current face recognition models still suffer from significant demographic biases. While approaches for bias mitigation have been proposed, existing methods often impose constraints on the training procedure and result in the degradation of recognition accuracy. To address this issue, we here introduce a method that reduces racial bias in pre-trained face recognition models without compromising their accuracy. To this end, we model face embeddings of each person by von Mises-Fisher (MF) distribution. We next observe the dependency between demographic attributes and the density of MF distributions, and propose DenseFace, a probabilistic face matching procedure that accounts for differences in MF distributions. Our extensive experiments demonstrate DenseFace to consistently reduce racial bias in strong face recognition models varying in network architectures, training datasets and loss functions. Notably, DenseFace preserves recognition accuracy and requires no retraining of the underlying face recognition model. Our work also investigates previously adopted bias measures and makes suggestions.
Research Starting Point
Face recognition systems often use one global decision threshold, yet impostor-score distributions differ across demographic groups. A model can therefore look accurate on RFW when each group receives its own cross-validated threshold while producing very different false-match rates in a real service. Most mitigation methods also require balanced retraining data, architecture changes, or fairness losses that may reduce recognition performance and are difficult to retrofit into an approved model.
Method
For each face embedding, the method finds nearest identities in a demographically balanced anchor set and estimates local inter-class density. It represents the embedding with a von Mises-Fisher distribution whose concentration reflects that density, then scores a pair by mutual likelihood rather than raw cosine similarity. A margin separates near-neighbor identities, and a lightweight regressor variant predicts the density directly so production matching does not need a large anchor lookup. The encoder, embedding dimension, and enrollment data remain unchanged.
Paper Summary
A comparison-layer fairness control can improve cohort parity without reopening model training, but it does not remove the need for representative calibration data and subgroup monitoring at the actual operating threshold.