← Back to Blog
Research RadarFace RecognitionarXivSeptember 2026

Monthly arXiv Radar

September 2026 Face Recognition Papers: Fairer Matching, Thermal Adaptation, and the Twin Test

September's face recognition research is less about another backbone race and more about difficult operating regimes. DenseFace changes the matching rule of an existing model to reduce racial false-match disparities. SynThermFace turns scarce visible-thermal pairs into a larger adaptation set without adding image translation at inference. The Celeb Twins Test Set then asks whether today's embeddings use the tiny local cues that humans rely on to distinguish monozygotic twins.

What This Month Signals

DenseFace shows that local embedding density can be used as a post-hoc probabilistic matching signal, bringing several non-Caucasian false-match rates much closer to a Caucasian reference while preserving verification accuracy. SynThermFace uses a diffusion model only during training, then deploys a single adapted EdgeFace forward pass; its synthetic-pair model reaches 0.99% EER on MCXFace but 14.70% on unseen Tufts data, so sensor transfer remains material. CTTS contributes 21,120 twin-focused pairs and finds that twelve modern matcher variants remain around 73-77% accuracy even though their usual benchmark average is about 97%. The shared lesson is to validate the matching rule, capture spectrum, and hard identity population separately rather than treating one aggregate score as production readiness.

Paper 012026-09-14cs.CV

DenseFace: Bias Mitigation in Face Recognition via Density-Aware Probabilistic Matching

Authors & Institutions

Mansur Bultygov

VisionLabs

Vadim Seliutin

Amazon Web Services

VisionLabs (work performed there)

Dmitry Nekhaev

VisionLabs

Ivan Laptev

Mohamed bin Zayed University of Artificial Intelligence

What Problem It Solves

DenseFace asks whether racial false-match disparities can be reduced at comparison time, without changing or retraining the underlying face encoder. It also replaces subgroup accuracy standard deviation with a deployment-oriented protocol: set one threshold at a Caucasian false-positive rate of 0.001 and measure how far every other cohort moves from parity.

Key Result

Across CosFace and AdaFace models trained with Glint360K, MS1MV2, WebFace4M, WebFace12M, and BUPT-BalancedFace, the same global threshold becomes substantially more even. On the BUPT model, the African cohort's false-positive multiple falls from 6.15 to 1.43 and the Indian cohort from 4.01 to 1.03 relative to the Caucasian reference. Conventional verification accuracy is preserved: for a CosFace Glint360K model, average RFW accuracy rises from 98.61% to 98.68% while subgroup standard deviation falls from 0.53 to 0.44. The regressor adds about 0.2% memory and 1.75% end-to-end latency in the authors' GPU test; its optimized CPU score calculation is about 1.5 times raw cosine cost. Results still depend on a representative anchor or regressor training population.

Abstract

Despite steady progress in face recognition, current face recognition models still suffer from significant demographic biases. While approaches for bias mitigation have been proposed, existing methods often impose constraints on the training procedure and result in the degradation of recognition accuracy. To address this issue, we here introduce a method that reduces racial bias in pre-trained face recognition models without compromising their accuracy. To this end, we model face embeddings of each person by von Mises-Fisher (MF) distribution. We next observe the dependency between demographic attributes and the density of MF distributions, and propose DenseFace, a probabilistic face matching procedure that accounts for differences in MF distributions. Our extensive experiments demonstrate DenseFace to consistently reduce racial bias in strong face recognition models varying in network architectures, training datasets and loss functions. Notably, DenseFace preserves recognition accuracy and requires no retraining of the underlying face recognition model. Our work also investigates previously adopted bias measures and makes suggestions.

Research Starting Point

Face recognition systems often use one global decision threshold, yet impostor-score distributions differ across demographic groups. A model can therefore look accurate on RFW when each group receives its own cross-validated threshold while producing very different false-match rates in a real service. Most mitigation methods also require balanced retraining data, architecture changes, or fairness losses that may reduce recognition performance and are difficult to retrofit into an approved model.

Method

For each face embedding, the method finds nearest identities in a demographically balanced anchor set and estimates local inter-class density. It represents the embedding with a von Mises-Fisher distribution whose concentration reflects that density, then scores a pair by mutual likelihood rather than raw cosine similarity. A margin separates near-neighbor identities, and a lightweight regressor variant predicts the density directly so production matching does not need a large anchor lookup. The encoder, embedding dimension, and enrollment data remain unchanged.

Paper Summary

A comparison-layer fairness control can improve cohort parity without reopening model training, but it does not remove the need for representative calibration data and subgroup monitoring at the actual operating threshold.

Paper 022026-09-09cs.CV

SynThermFace: Amplifying Limited Paired Data for Visible-Thermal Face Recognition via Synthetic Data Generation

Authors & Institutions

Anjith George

Idiap Research Institute, Switzerland

Adam Unal

Idiap Research Institute, Switzerland

Sebastien Marcel

Idiap Research Institute, Switzerland

University of Lausanne, Switzerland

What Problem It Solves

SynThermFace amplifies limited paired supervision for visible-to-thermal recognition and tests whether synthetic thermal pairs improve both in-database accuracy and transfer to a sensor/database that was never used for training. It separates the benefit of its adaptation objective from the benefit of scaling synthetic identities.

Key Result

On the disjoint MCXFace development split, real-pair PACT reaches 3.04% EER and 96.49% Rank-1, beating the compared real-pair adaptation baselines. With 1,000 CASIA-derived synthetic identities, EER falls to 0.99%, Rank-1 reaches 99.75%, and verification at 0.1% FAR reaches 93.48%; the unadapted EdgeFace baseline has 23.06% EER and 40.10% Rank-1. On unseen Tufts data, Digi2Real-derived training improves EdgeFace from 43.41% to 13.91% EER and from 12.03% to 56.37% Rank-1. The remaining MCXFace-to-Tufts gap is large, and the generator still depends on real MCXFace pairs, so the method reduces rather than eliminates real cross-spectral collection.

Abstract

Face recognition (FR) is a widely used modality for biometric authentication, but conventional models rely on visible-spectrum imagery and degrade when high-quality RGB images cannot be captured. Cross-spectral face recognition addresses this limitation by matching visible images with other modalities such as thermal imagery, enabling more reliable performance in low-light, nighttime, and unconstrained conditions. However, progress is limited by the scarcity of paired visible-thermal data, which is difficult and costly to collect at scale. We propose SynThermFace, a framework that amplifies limited real visible-thermal supervision into larger paired adaptation datasets for cross-spectral face recognition. A diffusion model is first adapted using a limited set of paired visible--thermal images and then used to generate large-scale paired visible--synthetic thermal data from existing real or synthetic visible face datasets. The generated pairs are used to adapt a pretrained visible-spectrum face recognition model into a CFR model. Unlike synthesis-based approaches that require image translation at test time, the proposed method shifts generation to the training stage and performs inference with a single forward pass through the adapted recognition model. Under the same MCXFace real-pair protocol, PACT improves over the evaluated CFR adaptation baselines, isolating the effect of the proposed adaptation objective. Training PACT on larger generated paired datasets provides additional improvements over both the unadapted model and the real-pair PACT configuration. Cross-database evaluation on the Tufts dataset provides evidence that the learned representation transfers to an unseen database. The source code and trained models will be made publicly available.

Research Starting Point

Thermal cameras can support authentication and surveillance when visible-light capture fails, but most identity galleries are RGB and paired visible-thermal faces are expensive to collect. Directly translating every thermal probe at runtime adds a generator to the critical path and can alter identity. The practical question is whether a small real paired set can supervise a much larger training set while leaving deployment as a normal embedding lookup.

Method

A diffusion generator is first adapted on the paired MCXFace training split, then converts either real CASIA-WebFace images or synthetic Digi2Real identities into matching thermal views. Preservation-Aware Cross-Spectral Tuning starts from a pretrained EdgeFace model and combines symmetric thermal-to-visible and visible-to-thermal contrastive learning with a loss that keeps visible embeddings near a frozen copy of the original model. Generation happens only while building the adaptation set; inference is one forward pass through the tuned recognizer.

Paper Summary

Offline synthetic thermal generation can turn a small paired collection into useful recognition supervision without runtime translation, but new sensors and populations still need their own transfer and fairness tests.

Paper 032026-09-01cs.CV

Revisiting Face Recognition for Monozygotic Twins: The Celeb Twins Test Set

Authors & Institutions

Michael Zang

Department of Computer Science and Engineering, University of Notre Dame

Haiyu Wu

Department of Computer Science and Engineering, University of Notre Dame

Mrinal Sharma

Department of Computer Science and Engineering, University of Notre Dame

Kevin W. Bowyer

Department of Computer Science and Engineering, University of Notre Dame

What Problem It Solves

The paper creates a verification-style benchmark that is large enough for twin-disjoint cross-validation and annotated for local distinguishing marks and possible mirror asymmetry. It then tests whether modern deep face matchers exploit those cues, and whether removing horizontal-flip assumptions or generating synthetic twins supplies a plausible path forward.

Key Result

The twelve matcher configurations achieve only 73.14-76.50% mean CTTS accuracy, versus roughly 97.00-97.69% averaged over five conventional verification sets. Erasing visible marks leaves embeddings and pair-distance distributions essentially unchanged, indicating that current models ignore cues humans can use. The asymmetry-aware retraining reaches 78.58% plus or minus 3.7 and is not statistically different from the original model; it helps some mirror-twin sets and hurts others. Generated twins do not preserve realistic within-person and between-twin distance structure, so they are not yet a validated training substitute.

Abstract

Past literature on face recognition for monozygotic (("identical") twins points to facial marks and mirror asymmetry as possible directions for improved accuracy of twins recognition. The Celeb Twins Test Set (CTTS) contains web-scraped image pairs for 80 sets of celebrity twins. It is the only twins test set with meta-data for twins with distinguishing skin marks and possible mirror asymmetry. CTTS is organized in the manner of face verification test sets such as LFW, CALFW, CPLFW, CFP-FP, and AgeDB-30. Current deep CNN matchers can achieve over 76% accuracy in classifying CTTS same-person / different-person image pairs. We show that current matchers do not make use of skin marks, or asymmetry, and discuss reasons for this. Finally, we discuss the feasibility of using generative AI tools such as Grok, ChatGPT and Gemini to create images of imagined monozygotic twins as a means to increase representation of twins in face recognition training sets.

Research Starting Point

Monozygotic twins expose an identity boundary that ordinary face benchmarks barely represent. NIST previously reported that, at a threshold chosen for a 0.0001 non-twin false-match rate, many algorithms accepted one twin as the other 98-99% of the time. Existing twin datasets are small, old, or lack metadata about freckles, moles, scars, and mirror-twin status, making it difficult to learn which cues modern embeddings actually use.

Method

CTTS-80 collects web images for 80 celebrity twin sets and constructs 21,120 balanced same-person and twin-impostor pairs across ten folds, keeping each twin set entirely within one fold. The authors evaluate ArcFace, AdaFace, UniFace, and MagFace ResNet-100 models trained on three common datasets. They erase visible skin marks in selected identities, compare original and edited embedding distributions, retrain ArcFace without horizontal-flip augmentation, and inspect synthetic twin sets requested from three generative systems.

Paper Summary

Twin verification remains a distinct risk class: standard accuracy says little about it, and today's embeddings appear not to use obvious local marks. High-assurance deployments should test twin impostors explicitly rather than infer performance from general benchmarks.